Addresses findings from a multi-agent review of the two prior fixes.
P0 (executor.ts): the stale-conflict recovery force-removed worktreePath with
no bounds check; that path can come from a git admin entry resolving outside
.worktrees/. Now refuses unless the path is inside the worktrees dir, not a
symlink (realpathSync), not a registered worktree, and not actively owned, and
re-verifies liveness in the catch instead of trusting the error string. Also
excludes spawn failures (spawn git ENOENT) from the stale-path classification.
worktree-pool.ts: resolveGitdirPointer -> dotGitPointerIsDangling. Reaps only
when a .git link's gitdir target is confirmed missing; a real .git dir,
unparseable pointer, or any read/stat failure is treated as NOT dangling
(conservative) so a transient read error on a live worktree can't trigger rm.
Drops the string|"directory"|null sentinel union.
core store.ts: bypass the reconcile recency window when the live task table is
empty (corruption/restore: surviving task.json keep old mtimes) and when
fusion.db was auto-recovered on startup, so .recover row loss isn't stranded.
Adds an ignoreRecencyWindow option.
Tests: executor recovery + out-of-bounds refusal, unparseable .git skip,
recency boundary, empty-DB/forced bypass. engine 135 + core 12 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Directories under .worktrees/ that survive with a dangling .git pointer
(present on disk, but their .git/worktrees/<name> admin entry is gone) are
invisible to `git worktree list`/`prune` yet collide with freshly generated
worktree names. The executor's conflict cleanup then fails with
"is not a working tree", failing the workflow graph at node 'execute' after
3 attempts.
- executor.ts: extend FN-4813 stale-conflict recovery to also treat
"is not a working tree" and ENOENT (not just "validation failed, cannot
remove working tree") as "no live worktree here" — prune the admin entry,
force-remove the leftover dir, and proceed with fresh creation.
- worktree-pool.ts: reapOrphanWorktrees skipped any dir on mere .git-file
presence, contradicting its own documented invariant. Resolve the .git
pointer and only skip when the gitdir target exists; reap dangling
pointers like any other orphan so they stop accumulating across runs.
- Tests for both the dangling (reaped) and valid (skipped) .git cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Agents stop heartbeating during long legitimate work (e.g. a verification
step blocked on a multi-minute test command). The 5-minute floor could
misread a busy agent as dead and reclaim its in-progress task mid-run.
Raise MIN_HEARTBEAT_STALENESS_MS to 10 minutes and strengthen the floor
test (7-minute-silent fast-interval agent stays healthy — would have read
stale under the old 5-minute floor). Engine typecheck clean; heartbeat
suite 148/148 pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The gate test asserted the OLD behavior — a paused graph exit in the `todo`
column parked `status:"failed"` with "operator action required". FN-6782
made the todo case benign (no failed park; benign log + cleared marker), so
split the parameterized test: `todo` now asserts the benign path (never
parked failed), `done` keeps the operator-action surfacing (log only, no
park). Full engine-core gate suite passes (644/644).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
coderabbit Major: reapLeakedConcurrencySlots captured executingIds once
before the loop, but each holder awaits getTask — a task could start
executing mid-sweep and have its worktree slot pulled. Refresh the
executing set immediately before clearPhantomExecutorBinding and skip if
the holder is now executing (same race the A1 recovery fix closed).
clearPhantomExecutorBinding's live-session refusal remains the last line
of defense; this avoids racing it. Added a mid-sweep race test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Substantive (A1 recoverPausedAbortFailures):
- Self-guard on globalPause/enginePaused at method entry (greptile P1) — the
public method must not requeue tasks an operator intentionally froze.
- Re-validate the FULL predicate with a FRESH executing set on the re-read
before the backward move (coderabbit Major + greptile): add fresh.userPaused
and column re-check so a task that became ineligible across awaits is skipped.
- Isolate audit emission in its own try/catch (coderabbit) so an audit throw
after a successful mutation can't log a false "recovery failed".
- Decouple the recovery predicate from the literal error text via shared
PAUSE_ABORT_PARK_ERROR_MARKER/OPERATOR_MARKER constants (greptile) — the
executor builds the parked message from the same constants.
- Use the wired clearPhantomExecutorBinding (live-session-guarded) instead of
the declared-but-never-wired releaseExecutorWorktreeOwnership, which no-op'd.
Nits:
- FNXC-prefix new comments in executor.ts, run-audit.ts, and the benign test
per repo comment policy.
- Fix a test-only type error on the clearPhantomExecutorBinding mock.
Added a test asserting the globalPause self-guard. Engine typecheck clean;
pause-abort/reaper/benign + regression suites pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
reapLeakedConcurrencySlots() reclaims in-memory worktree slots whose
holder is no longer in-progress (the FN-6756 "in todo yet still a
maxWorktrees holder" leak) without an engine restart — defense-in-depth
behind the source fix.
- executor: new listWorktreeHolders() read-only introspection over
activeWorktrees; wired through in-process-runtime to SelfHealingManager.
- reaper releases ONLY when every guard agrees: not executing, task
missing or in todo/triage, past a 60s grace, and clearPhantomExecutor
Binding itself refuses (returns false) if a live session surface is
registered — so it can never pull a worktree from a running agent.
- registered in maintenance batch 2 (respects globalPause/enginePaused
skip + FN-4962 ordering).
- widened the clearPhantomExecutorBinding option type to surface its
boolean refusal signal.
Engine typecheck clean; 19 tests pass (new reaper 7 cases + regression).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A global pause/resume cycle parked tasks that had re-queued to todo as
status:"failed" ("operator action required") and leaked their in-memory
worktree slot. The scheduler kept re-dispatching the todo task, the
genuine-pause-abort branch re-fired on the still-set pausedAborted marker,
and it re-parked instantly with no backoff — a retry storm (75x/hr) that
pinned maxWorktrees=3/3 and concurrency-starved the whole queue.
- R1+R2 (executor.ts handleGraphFailure): treat a pause-abort that left a
task in `todo` as benign (FN-6782) — don't park failed, clear the
pausedAborted marker so the next dispatch is clean, and release the
leaked activeWorktrees slot. Operator-action failure preserved for
genuinely stranded non-todo columns (FN-6478).
- A1 (self-healing.ts recoverPausedAbortFailures): new maintenance sweep
that auto-recovers any pause-abort park still on the board and requeues
it (status:null = schedulable) so the board self-heals.
- run-audit.ts: new mutation types for the recovery telemetry.
Corrected the spec's null-vs-queued assumption: the scheduler dispatch set
is column==="todo" && !paused (scheduler.ts:1288); status:"queued" is the
*blocked* marker, status:null is runnable — so recovered tasks are left null.
Deferred (documented): A2 leaked-slot reaper needs a new executor
listWorktreeHolders introspection API to reap in-memory worktree slots
safely; R1 closes the observed leak at its source.
Tests: self-healing-paused-abort-recovery.test.ts (3),
executor-paused-abort-todo-benign.test.ts (2). Engine typecheck clean;
106 existing pause/graph-failure/limbo tests still pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Recover wedged in-progress tasks by clearing stale executor bindings only after liveness proves the owner is gone.
- Add a guarded executor escape hatch that clears only stale in-memory task bookkeeping while refusing live session surfaces.
- Teach self-healing to identify phantom executor-active bindings using age, checkout, heartbeat, run-audit, and worktree liveness signals before requeueing preserved work.
- Record reclaim events in run audit and cover preserved-worktree recovery with reliability interaction tests.
- Document the recovery path and add a patch changeset for the published CLI package.
Files changed:
.changeset/fn-6736-phantom-executor-binding.md | 5 +
AGENTS.md | 1 +
docs/architecture.md | 1 +
.../reclaim-phantom-executor-binding.test.ts | 244 +++++++++++++++++++++
packages/engine/src/executor.ts | 35 +++
packages/engine/src/run-audit.ts | 2 +
packages/engine/src/runtimes/in-process-runtime.ts | 3 +-
packages/engine/src/self-healing.ts | 112 ++++++++++
8 files changed, 402 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-6736
Fusion-Task-Lineage: c76191ba-f4c3-4832-a790-67676e258ba2
Optimizes shared branch group lifecycle coverage by bypassing repeated mock merger sessions where full merge behavior is not under test.
- Add a deterministic fast integration helper that merges staged member branches into the shared group branch and records merge metadata.
- Keep routing and self-healing cases on the full aiMergeTask path while using the faster seam for promotion and gating assertions.
- Preserve FN-5820 shared branch completion and auto-merge-off expectations with less slow-test overhead.
Files changed:
.../shared-branch-group-lifecycle.slow.test.ts | 54 ++++++++++++++++++----
1 file changed, 46 insertions(+), 8 deletions(-)
Fusion-Task-Id: FN-6691
Fusion-Task-Lineage: 5f01318c-20b0-4f11-ad3c-c2e1b8cd0f3e
Route planning executor selection through a shared engine seam for model and CLI-agent sessions.
- Add a planning executor selection type with model and CLI-agent variants.
- Wrap CLI-agent planning as a one-shot interactive session that returns terminal complete, question, or error events.
- Export the resolver and wire the core interactive adapter through the model-backed default path.
- Cover resolver behavior for default model sessions, CLI-agent terminal events, and CLI-agent failures.
Files changed:
.../src/__tests__/interactive-ai-session.test.ts | 91 ++++++++++++++++++++
packages/engine/src/index.ts | 13 ++-
packages/engine/src/interactive-ai-session.ts | 97 ++++++++++++++++++++--
3 files changed, 190 insertions(+), 11 deletions(-)
Fusion-Task-Id: FN-6689
Fusion-Task-Lineage: b2fc0589-621e-44dd-a470-49107b0ad69f
The paused-after-completion graceful-exit path finalizes a fully completed task to in-review while leaving a non-user paused:true flag set (handoffToReview/applyInReviewEnterEffects clear status/blockedBy but not paused). handleGraphFailure's completion-finalized guards required paused!==true, so once the volatile completion markers were lost (execute() re-entry deletes completionFinalizedTaskIds; teardown overwrites provenance to hard-cancel) the trailing graph failure was misclassified as an operator-action pause abort and the completed task was parked status:failed (FN-6638 recurrence). Drop the paused!==true requirement from alreadyFinalizedToReview and suppressFinalizedCompletionAbort, and gate genuinePauseAbort's bare paused clause on the completion suppression. Genuine userPaused/global-pause/in-progress tasks are unaffected.
Fusion-Task-Id: FN-6648
Prevent completed no-commit executions that already advanced to review from being re-parked as pause-abort failures.
- Add completion-finalize pause-abort provenance and exclude it from genuine pause handling after review handoff.
- Mark paused-after-completion finalization paths with the new provenance before handing tasks to review.
- Cover the finalize-to-review abort recovery path with executor regression tests and document the lifecycle exception.
- Add a patch changeset for the published Fusion package.
Files changed:
.changeset/fn-6625-finalize-to-review-abort.md | 5 +
docs/architecture.md | 2 +-
.../engine/src/__tests__/executor-recovery.test.ts | 158 ++++++++++++++++++++-
packages/engine/src/executor.ts | 26 +++-
4 files changed, 185 insertions(+), 6 deletions(-)
Fusion-Task-Id: FN-6625
Fusion-Task-Lineage: 728f6fe5-4c27-4597-b17e-e16ff97b9277
Ensure task-detail comments are delivered to live executor threads and preserved for the next step prompt when no step session is active.
- Forward steering comments through legacy, step-session, and workflow-step executor targets with delivery status logging.
- Keep step-session task details updated and include pending steering comments in full and reduced step prompts.
- Track delivered steering comment IDs so comments are injected or queued exactly once across active and subsequent step sessions.
- Update step-session executor tests for live steering, queued prompt fallback, and reduced prompt behavior.
Files changed:
.../src/__tests__/executor-step-session.test.ts | 467 ++++++---------------
.../src/__tests__/step-session-executor.test.ts | 63 ++-
packages/engine/src/executor.ts | 34 +-
packages/engine/src/step-session-executor.ts | 69 ++-
4 files changed, 283 insertions(+), 350 deletions(-)
Fusion-Task-Id: FN-6590
Fusion-Task-Lineage: 18fffd41-7632-4f29-8721-daaf3c239a74
Resolves conflicts in the lazy-loaded heavy-views inventory. main independently
grew the curated list to 22 (adding AppModals lazy modals); this branch added the
Command Center view. Combined count is 23 — updated the AGENTS.md prose/inventory
and the lazy-loaded-views-docs test contract (count + length assertions) to 23,
keeping main's richer "App-level and AppModals" wording.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>