This merge introduces a memory file markdown preview feature (FN-3584) with corresponding documentation, refines the AgentDetailView and AgentLogViewer components in the dashboard, and adds defensive collision handling for worktree operations during manual task moves (FN-3583).
Fusion-Task-Id: FN-3584
Completes typing for the scheduled evaluator integration in the cron runner and project engine, with corresponding test updates in the evaluator test file.
Fusion-Task-Id: FN-3389
- Update enginePaused setting docs to specify stuck-task timers are suspended while pauses are active
- Document that paused wall-clock time does not count toward taskStuckTimeoutMs, including shared globalPause windows
- Clarify that unpausing restores scheduling and grants active sessions a fresh stuck-task grace window before detection resumes
Fusion-Task-Id: FN-3538
The merge introduces the auto-claim setting feature, enabling agents to automatically claim tasks on assignment, with corresponding UI in AgentDetailView and documentation. It also adds the agent heartbeat execution system to the engine, refinements to workflow results styling and design tokens, aut
Fusion-Task-Id: FN-3479
Added 95 lines of test coverage for merge scheduler recovery paths in the merge error recovery test suite, completing Step 4 of FN-3317.
Fusion-Task-Id: FN-3317
This merge lands FN-3231 across two steps: it preserves a merge-active fix when verification bounces occur (step 1) and ensures the fix is retained during board routing transitions (step 2). Changes span the dashboard Board routing logic and the engine executor, with corresponding test coverage adde
Fusion-Task-Id: FN-3231
- Add scheduler logic to create dependency-linked follow-up tasks when actionable PR feedback remains after a PR is merged or closed
- Update engine runtime/project wiring to support manual PR create flows and branch publish behavior for fusion/<task-id>
- Add dashboard route coverage for manual PR creation/linking behavior and corresponding engine/runtime tests
- Document manual PR branch conventions and follow-up behavior in task management and dashboard docs
Fusion-Task-Id: FN-3202
The merge restores the engine's unpause merge sweep logic in `project-engine.ts` and documents the soft-pause merge resume behavior across architecture and settings reference docs, with associated test coverage added.
Fusion-Task-Id: FN-3201
Demote the refresh message from console.error to debugMcp so it no
longer appears as an error in normal output.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Triage and the todo→in-progress scheduler already sorted by priority
(urgent→low, then createdAt ASC, then id ASC); the auto-merge queue
was strictly FIFO, so a backlogged low-priority task could merge
ahead of an urgent one. drainMergeQueue now picks the highest-
priority eligible task each iteration, and the four in-review sweeps
(startup, periodic, global unpause, engine unpause) sort by priority
before enqueueing so the single-item fast path also picks priority-
first. Picker is hardened against concurrent queue mutation by stop()
and pause-handler removal: it re-locates the chosen entry by id and
re-checks shuttingDown after awaiting getTask.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Merges external Tailscale funnel detection (FN-2976) — adds types, detection logic, status inclusion, kill flow, and a dedicated UI panel for funnel processes started outside Fusion — alongside custom AI providers API routes and a new settings UI section (FN-2965).
Fusion-Task-Id: FN-2976
* Replace placeholder /remote/qr SVG (URL drawn as text) with real QR
rendered via the qrcode package; add format=terminal returning ASCII
QR for the TUI.
* Resolve the public tailscale funnel URL from captured CLI output
instead of constructing http://<hostname>:<port> from a configured
hostname label — that label was never used by `tailscale funnel` and
produced a non-public URL in the auth/QR link.
* Drop hostname requirement from engine + UI; only target port matters.
* Tighten tailscale parseReadiness to require a URL on the matched line
so the tunnel manager doesn't lock in `running` before the URL line.
* TUI: poll remote status, show ● tunnel indicator + URL in MainHeader,
bind Ctrl+Q to a global QR overlay (terminal ASCII), and switch the
in-Settings K shortcut to render the same ASCII QR.
* Auto-poll remote status in the dashboard while in `starting`/`stopping`
so the UI flips to running without reopening the modal.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
FN-2910 surfaced concurrent reviewer + merger activity on the same task.
Root cause: asymmetric in-flight guards let an unpause-resume kick off a
fresh executor session while a recovery path was already running, and the
auto-merge handoff fired before the executor's finally block finished
cleanup. This sweeps the surrounding lifecycle paths for similar races and
tightens the reviewer pause gate against TOCTOU through runtime setup.
- Symmetric in-flight tracking across `executing`, `recoveringCompleted`,
and `resumingUnpaused`; `recoverCompletedTask` bails when any are set.
- Atomic claim of the recovery slot in the completed-task watchdog before
any awaited work.
- Workflow-rerun bounce returns "bounced" | "skipped-pending" so the
watchdog can no longer log a false-success retry when the original
bounce is still mid-flight.
- Self-healing's completed-task scan re-checks executing IDs inside the
loop instead of trusting a pre-await snapshot.
- 300ms grace period before auto-merge enqueue, giving the executor's
finally block (session disposal, child cleanup) time to drain and
eliminating the residual log-overlap symptom from FN-2910. Test uses
fake timers, no real sleep added.
- New AgentSemaphore.runNested for synchronously nested helper agents
(reviewers): bumps activeCount for honest observability while bypassing
the wait queue, preserving forward-progress fairness for the parent at
low maxConcurrent. Both createReviewStepTool and triage's
createReviewSpecTool now use it.
- New beforeSpawnSession hook on AgentRuntimeOptions/AgentOptions fired
inside createFnAgent immediately before createAgentSession, past every
awaited setup step. Reviewer wires a pause re-check that throws a
sentinel error converted to UNAVAILABLE, closing the TOCTOU window
where pause flipped during runtime resolution or resource loading.
All 2887 engine tests pass; engine + core + cli + dashboard + plugin-sdk
+ pi-claude-cli + desktop typecheck clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds a layered recovery cascade to the merger's pre-rebase stage so tasks
no longer get stuck in in-review when their declared dependency was
squash-merged to main and left orphan raw commits in the dependent's
history. Also prevents the orphan situation at the source for new tasks.
Why:
- 13 tasks were stuck in in-review for hours, all hitting the same
pre-merge rebase abort because they shared 6 raw commits inherited
from FN-2729's branch (declared baseBranch). FN-2729 was then
squash-merged to main, turning those raw commits into orphans whose
content is in main but in a different commit shape, conflicting with
later-merged tasks. The merger's `smart-prefer-main` strategy
correctly refused -X ours (which would silently re-introduce main's
deletions), but the only escape hatch was a 30-min cooldown loop
that retried the same impossible rebase forever.
Recovery cascade (merger.ts pre-rebase stage):
- Layer 1: surgical `git rebase --onto <main> <dep-tip> <branch>` when
task.baseBranch is set. Resolves the dep tip from the live branch ref
or recorded baseCommitSha; peels off the dep's inherited commits
cleanly. Captures the squash-merge-of-dep case end-to-end.
- Layer 2: generic patch-id duplicate-content stripping. Walks the last
500 main commits, computes patch-ids, then drops branch commits whose
patch-id matches and cherry-picks the remainder onto main. Captures
manual cherry-picks, double-merges, and any other duplicate-content
variant Layer 1 doesn't see. Restores the branch's pre-mutation SHA
on partial-failure so worst case leaves the worktree no worse than
before the recovery attempt.
- Layer 3: AI arbitration fall-through. If Layers 1+2 fail, log the
situation and proceed to the existing 3-attempt AI merge cascade
instead of throwing. The deterministic post-merge verification
(test + build) gates whatever the AI produces — that gate is what
enforces prefer-main's safety contract under fall-through (no silent
re-introduction of main's deletions).
- Critical: the unsafe `-X ours` Attempt 3 is suppressed under
fall-through. AI Attempts 1+2 are the only paths that can complete
the merge; if both fail and verification rejects them, the task
bounces back to in-progress via the existing engine path rather than
silently merging.
Prevention (executor.ts worktree creation):
- When a task declares a non-main `baseBranch`, branch the worktree
off main (origin/<defaultBranch> when worktreeRebaseBeforeMerge is
enabled and a remote is resolvable; otherwise local rootDir HEAD)
and `git merge --squash` the dep's content as a single import commit.
The dependent branch then carries main's history + 1 commit instead
of inheriting the dep's raw commits, so a future squash-merge of the
dep produces patch-id-matching content that rebases cleanly.
- Honors settings: respects `worktreeRebaseBeforeMerge`,
`worktreeRebaseRemote`, and falls back to local HEAD when no remote
is resolvable. Fully fail-soft: any squash-import error falls back to
the legacy fork-from-dep behavior so worktree creation still works
for setups where the squash flow can't run.
Engine-side last-retry fix (project-engine.ts):
- Changed conflict-retry condition from `currentRetries < MAX` to
`currentRetries + 1 < MAX` so the bounce-to-in-progress code fires
in the same engine tick as the failing attempt, rather than relying
on a setTimeout-scheduled Nth attempt that dies on engine restart.
Without this, a dev-time engine restart between the 3rd and 4th
retry left the task with mergeRetries=MAX and only the 30-min
cooldown sweep could try again.
Tests:
- New "Layer 1 recovery" test asserts the surgical --onto rebase fires
when baseBranch is set and primary rebase aborts, and that Layer 3
fall-through is NOT triggered when Layer 1 succeeds.
- Updated the "no silent fall-through to -X ours" test to cover the
new fall-through path: even after Layers 1+2 fail and the merge
cascade proceeds, -X ours must not run, and the task log must record
both the Layer 3 fall-through entry and the Attempt 3 suppression.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Tasks were getting stuck in `in-review` forever when auto-merge could not
resolve conflicts within MAX_AUTO_MERGE_RETRIES. The conflict-exhaustion
branch silently cleared `status` (no error, no log entry, no comment),
and the 30-min cooldown sweep would reset retries and re-attempt the
same impossible merge — looping silently with no user-facing surface.
Why:
- FN-2918 and FN-2903 both spent hours in this loop with no error/comment
visible on the task. The only log evidence was repeated
"Auto-merge retry cooldown elapsed (30m idle)" entries with no
follow-up outcome.
How to apply:
- Every merge failure now writes a `<Manual|Auto>-merge failed: <msg>`
entry to the task log so the dashboard surfaces the reason.
- Conflict-retry exhaustion now bounces the task back to `in-progress`
with a comment + log entry so the executor re-rebases against main
and retries — mirroring the verification-failure-bounce pattern.
- New `mergeConflictBounceCount` task field caps outer bounces
(`MAX_MERGE_CONFLICT_BOUNCES = 2`); past the cap, the task is parked
in `in-review` with `status="failed"` and a follow-up triage task is
created so a human can resolve the conflict manually.
- Non-conflict and non-direct-strategy errors now also set
`status="failed"` so the cooldown sweep can't re-pick them up.
- `canMergeTask` skips tasks with `status="failed"` so terminal
failures (verification cap, bounce cap, non-conflict error) are no
longer eligible for cooldown re-attempts.
Schema migration v52 adds the `mergeConflictBounceCount` column.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The prior commit re-enqueued tasks after stale-merge recovery but didn't
account for the engine's in-memory `mergeActive` set, which still held the
wedged task. `internalEnqueueMerge` silently no-ops when the entry is
present, so the re-enqueue had no effect.
The recovery callback now also aborts the active merge's signal and
disposes its session if the wedged attempt was the currently-active one.
This is what unsticks tasks where an AI provider call is hung mid-await
and the surrounding `try/finally` never gets a chance to run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Stale-merge recovery now calls back into ProjectEngine's auto-merge queue
directly instead of waiting on the 15s polling sweep — wired via a new
InProcessRuntime.setMergeEnqueuer hook so SelfHealingManager can re-enqueue
without leaking engine internals.
createWorktree mirrors the merge-time rebase: when worktreeRebaseBeforeMerge
is enabled, the new task branch is rebased onto <remote>/<defaultBranch>
right after creation, so executors start from origin's tip with local main
replayed on top. Best-effort — fetch/rebase failures abort cleanly and
leave the merge-time rebase as the backstop.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add notification service module with provider abstractions and ntfy provider implementation
- Refactor NtfyNotifier into a compatibility wrapper that delegates task-event delivery to NotificationService
- Initialize and stop NotificationService from ProjectEngine while preserving gridlock notifications via NtfyNotifier
- Export notification APIs from engine index and add focused unit coverage for provider, service, and project-engine wiring
- Improve automation startup diagnostics and route handling for manual execution steps
- Add support for full manual automation step execution in dashboard and engine flows
- Expand due-schedule coverage in automation store and dashboard route tests
- Add cron runner regression tests for edge cases and document the automation execution fix via changeset
Three fixes for the worktree-overflow / stuck-task incident:
1. Cap deterministic-verification-failure bounces (fix#2)
Auto-merge previously bounced an in-review task back to in-progress
on every verification failure with no upper bound. A single flaky test
could keep a task ping-ponging in-review→in-progress forever, holding
its worktree and consuming agent slots. Adds verificationFailureCount
on Task (DB migration v48), increments on each bounce, and after 3
failures marks the task failed and creates a follow-up triage task
so a fresh agent can investigate the underlying flake instead of
re-running the same fix loop.
2. Reap unregistered orphan worktree dirs even when recycle is on (fix#3)
cleanupOrphans previously bailed out entirely when recycleWorktrees
was true, leaving stale dirs (clear-hawk-broken, *-bak, leftover
crash debris) on disk forever. New reapUnregisteredOrphans pass
removes only directories that aren't registered git worktrees, so
the recycle pool keeps its warm worktrees but the trash gets cleared.
3. Idempotence guard on activity-log listener wiring (fix#6)
setupActivityLogListeners() was registering handlers on every call.
When init() ran twice, every task:created / task:moved event wrote
N rows to activityLog, producing the duplicate entries visible in
the DB. Added activityListenersWired flag so repeated calls no-op.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Hitting Stop (globalPause) disposed the AI merge agent session but left
the spawned `pnpm test` / `pnpm build` child processes running until
they finished naturally. With recurring flaky-test loops at Step 5,
that meant Stop had no visible effect — new test runs kept piling up
across multiple worktrees.
Two gaps:
- project-engine.ts onGlobalPause never called mergeAbortController.abort(),
so subsequent verification commands (gated by the signal) weren't cancelled.
- merger.ts execWithProcessGroup only listened to its own internal
timeout — passing an AbortSignal had no effect on the in-flight
child process group.
Fix: abort the controller on global pause, and have execWithProcessGroup
SIGTERM/SIGKILL the detached process group when its signal aborts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Remove the standalone remoteEnabled setting from CLI, core settings defaults/types, and dashboard settings APIs/UI
- Treat remote access as enabled when an active provider is selected and that provider is configured as enabled
- Update remote auth and engine lifecycle checks to gate on provider activation instead of a global flag
- Adjust tests and add a changeset documenting the remote access configuration simplification
- Add ProjectEngine restore lifecycle core to perform safe restarts and surface detailed restore state transitions
- Expose restore diagnostics through remote-access status types and settings/memory route context, including legacy API mapping updates
- Add comprehensive regression coverage for restore lifecycle behavior in engine and dashboard headless remote-access tests
- Document the restore lifecycle contract in architecture/settings docs and include a patch changeset for @runfusion/fusion
- Add remote-access contracts, provider adapters, and a tunnel process manager with lifecycle handling
- Wire tunnel manager into ProjectEngine startup/shutdown flow and export new remote-access modules
- Update settings modal UX for remote auth URLs, including wrapping and related UI test coverage
- Document tunnel manager behavior and remote settings sync details in architecture, CLI, and settings docs
- Add merger abort primitives and track active merge runs for coordinated cancellation
- Abort in-flight merges during engine shutdown and propagate AbortError through fallback catch paths
- Honor abort signals before commit, push, and dependency sync to prevent post-cancel side effects
- Expand merger and project-engine tests to cover abort propagation and merge-abort-on-stop behavior
- Preserve overdue nextRunAt when schedule updates only touch non-cadence fields
- Recompute nextRunAt only when cadence changes, schedules are re-enabled, or nextRunAt is missing
- Sync memory dreams automation during ProjectEngine startup before CronRunner begins ticking
- Add core/engine regression coverage and a patch changeset for @runfusion/fusion release notes
- Replace previously silent catch blocks in auto-merge and settings-listener flows with runtimeLog.warn messages
- Add warning logs for startup and periodic auto-merge sweeps plus poll interval fallback when settings reads fail
- Add regression tests that force each swallowed-error path and assert structured warnings are emitted
- Cover global/engine unpause resumeOrphaned failures and stuck-detector checkNow failures in project-engine tests
- Add Memory section controls for enabling auto-summarize with threshold and cron schedule inputs
- Wire ProjectEngine to sync auto-summarize automation on startup and when related settings change
- Reuse a single startup settings snapshot when syncing insight extraction and auto-summarize automations
- Add SettingsModal and ProjectEngine tests covering auto-summarize UI persistence and automation re-sync behavior
- Wire MessageStore into the in-process executor runtime and project engine
- Forward MessageStore events through dashboard SSE infrastructure for mailbox updates
- Close mailbox pipeline gaps across API routes, server wiring, and mailbox UI components
- Add regression coverage for messaging routes, SSE forwarding, and agent tool behavior
- Document the MessageStore SSE wiring pattern in .fusion/memory.md
- Expose getAgentStore from runtime and engine packages
- Add includeSystem filter to getOrgTree API endpoint
- Wire AgentStore into SSE endpoint for event forwarding
- Forward agent events through the SSE pipeline
- Exclude ephemeral agents from org tree API and UI
- Filter ephemeral agents in useAgents hook
- Update AgentsView test to expect includeSystem filter
- Document agent SSE event forwarding architecture
- Add scope field to automations and routines for granular control
- Update automation-store and routine-store with scope-aware query methods
- Extend database schema with scope column for automations and routines
- Update cron-runner and routine-scheduler to respect scope boundaries
- Add project context injection to in-process runtime for scoped execution
- Include changeset for minor version bump
- executor.test.ts: remove unused imports (Column, StuckTaskDetector),
replace Function type with EventListener, add MockTaskStore interface
- restart.integration.test.ts: replace require() with ESM import,
replace Function types with proper function signatures
- All tests pass
- Remove unused imports across 25 files in engine package
- Remove unused variable declarations in ipc-worker.ts, child-process-runtime.ts, and mission-autopilot.ts
- Clean up unnecessary imports in agent-instructions.ts, agent-tools.ts, cron-runner.ts, executor.ts, and other modules
- Minor cleanup in notifier.ts, peer-exchange-service.ts, pi.ts, plugin-runner.ts, and other files
- Improves code quality and reduces potential confusion from unused code
Multiple engine processes (dashboard + serve) share the same SQLite database
but each has its own in-memory merge queue. Without a cross-process check,
two processes can start merging different tasks simultaneously.
Added store.getActiveMergingTask() as a DB-level check before any merge
starts. The drainMergeQueue defers with pollIntervalMs delay, and both
aiMergeTask and processPullRequestMergeTask have safety-net checks.
Also moved stale merge status cleanup to run regardless of autoMerge setting.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause: if a merge crashed or the process restarted mid-merge, the
"merging" status was never cleared. On next startup the stale task kept
its "merging" status while the queue moved on to the next task, resulting
in two tasks appearing to merge at once.
Two fixes:
1. Add "merging"/"merging-pr" to BLOCKING_TASK_STATUSES so tasks with
active merge status are not re-enqueued by the retry sweep.
2. Clear stale "merging" statuses during startup merge sweep — no merge
is actually running at engine start, so any such status is a leftover.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The onMerge() path (dashboard "merge now" button) bypassed the
drainMergeQueue serialization, allowing two tasks to enter "merging"
status simultaneously within the same project. Route manual merges
through the same queue so only one merge runs at a time per project.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When a merge completes but auto-recovery moves the task back to
in-review, the retry gating (mergeRetries >= 3) blocked re-processing.
Now canMergeTask always accepts mergeConfirmed tasks and drainMergeQueue
fast-paths them directly to done without re-running the merge agent.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>