Added a verification guard for steps 1-4 in project-engine with companion regression tests covering both the merge-error recovery path and a post-finalize noop scenario using real git fixtures.
Fusion-Task-Id: FN-4944
Fusion-Task-Lineage: ca8fca22-c732-42e4-aa8d-a82981d2f36e
The in-line branch-cross-contamination auto-recovery in executor.ts had
two related bugs that produced transient "no-worktree-no-merge-confirmed"
stall signals in the dashboard while a live worktree was still mapped
on disk:
1. autoRecoverCrossContamination was called with repoDir=this.rootDir.
The recovery does: git checkout --detach <baseSha> → cherry-pick
wanted commits → git update-ref → git checkout <branch>. When the
branch is checked out in a worktree (the normal case), that final
recheckout in rootDir is blocked by git with "branch already used
by worktree at ...", so the in-line happy path silently failed for
every contaminated task that had a real worktree. Pass the task's
worktree as repoDir when available so the operations stay internal
to the worktree and the recheckout succeeds.
2. The successful-recovery branch called
moveTask(taskId, 'todo', { preserveResumeState: true })
without preserveWorktree:true. moveTask defaults to nulling
task.worktree on requeue. The worktree directory and git mapping
were still live — the dashboard's in-review-stall classifier
(no-worktree-no-merge-confirmed) and TaskChangesTab both keyed off
task.worktree being null and lied about the worktree being gone.
Sibling recovery paths (auto-recovery-handlers/contamination.ts,
tryBootstrapMisbindingRecovery, self-healing.ts:1639) all already
pass preserveWorktree:true; this site was inconsistent.
Updated the existing FN-4428 regression test and added a new
FN-4939 test asserting repoDir uses task.worktree (with fallback
to rootDir when the task has no worktree pointer).
Refs: packages/core/src/in-review-stall.ts:116
The [scope-leak] reviewLevel=N enforcement=warn warning was firing on
many in-progress tasks for off-scope .changeset/FN-XXXX-*.md files (the
production signature on FN-4789, FN-4801, FN-4818 \u2014 their branches all
contained .changeset/FN-4811-*.md files from the in-progress fix stack).
By convention every task may add its own changeset entry under
.changeset/ per AGENTS.md 'Finalizing Changes' section, so .changeset/
files are now treated as always-allowed by the scope-leak guard
regardless of the task's declared file scope.
Cross-task changeset leakage is still caught by stronger downstream
guards (file-scope invariant at squash, post-merge audit) at much
higher signal-to-noise. This change only suppresses the noisy
per-execution warning that was flooding logs without adding any
defensive value.
Adds a new exported helper isAlwaysAllowedScopeLeakPath() so the
allowlist surface is easy to extend. Test coverage in
scope-leak-changeset-allowlist.test.ts.
Also (incorporated from interrupted merge state): loosens the
executing-task-lock.test.ts assertion that one losing-instance store
sees zero work-log entries rather than the brittle exact-count of
mockedCreateFnAgent invocations (the no-fn_task_done retry path can
fire on the winning instance, so the count varies).
Fusion-Task-Id: FN-4811
Investigating FN-4814 + FN-4811 re-failures after commit 8bef30655 (which
added per-instance synchronous this.executing.add) revealed the per-instance
guard was insufficient. FN-4809 log at 02:48:17-18 UTC:
02:48:17 [-] Resuming execution after unpause
02:48:17 [-] Step 4 (Testing & Verification) -> pending
02:48:17 [-] Step 4 (Testing & Verification) -> pending
02:48:17 [6097725-y2nb] Executor detected stale merge state ...
02:48:18 [6097816-9gde] Executor detected stale merge state ...
Both runs y2nb and 9gde reached executor.ts:2661 (which is INSIDE execute(),
past the synchronous this.executing.add claim). The only viable explanation
is that there is more than one TaskExecutor instance in the same Node
process (engine restart race, multi-project hybrid runtime, or similar code
path). Each instance has its own executing Set, so the per-instance guard
doesn't help.
Fix: module-level singleton executingTaskLock in active-session-registry.ts,
shared across all TaskExecutor instances. execute() synchronously tryClaim()s
the lock; if false, bails. Every existing this.executing.delete() site also
calls executingTaskLock.release(). Per-instance this.executing kept because
many other call sites use it (this.executing.has at handler gates,
stuck-detector, resumeTaskForAgent, etc.).
Test setup (resetExecutorMocks in executor-test-helpers.ts) clears the lock
between tests so process-wide state doesn't leak (executor-pause and
executor-prompt tests would otherwise show 'expected 2 createFnAgent calls
but got 0' / 'expected not called but called 3 times' flakes).
Tests:
- executing-task-lock.test.ts: 2 cases. Key case creates TWO TaskExecutor
instances and races them on the same task ID, asserts only ONE actually
runs. Verified FAILS on prior code (8bef30655) and PASSES on fix.
Verification:
- Targeted suite (4 files, 170 tests): pass.
- pnpm --filter @fusion/engine build: clean.
- pnpm lint: clean.
Fusion-Task-Id: FN-4811
TaskExecutor.execute() had a classic JS async race window. Original:
async execute(task) {
if (this.executing.has(task.id)) return; // check
const assignedAgentId = task.assignedAgentId;
if (assignedAgentId && await this.shouldDeferForHeartbeat(...)) // AWAIT yields
return;
this.executing.add(task.id); // add (too late)
...
}
Two concurrent execute(task) calls (scheduler dispatch + task:moved event
handler + restart-recovery) both:
1. Pass the synchronous has() check (Set is empty).
2. Enter the awaited shouldDeferForHeartbeat call (yields the event loop).
3. Resume and both call this.executing.add(task.id).
4. Both proceed to create the same worktree path.
Production failure shape (FN-4814 + FN-4811, observed within minutes):
01:30:56 [runA-caoe] Worktree created at /...worktrees/bright-mesa
01:30:56 [runB-w23q] Worktree created at /...worktrees/bright-mesa
01:30:58 worktree liveness assertion failed: not_usable_task_worktree
01:31:48 [thirdRun] also fires liveness assertion fail
01:37:48 In-review stall surfaced [no-worktree-no-merge-confirmed]
This is the root cause of the entire FN-4781/FN-4804/FN-4814/FN-4811
cascade. Every other guard added today (FN-4811 active-session gate,
self-healing reclaim defer, validation-failed recovery, silent reclaim
recovery, integrity-warning dedup) was patching SYMPTOMS of the
duplicate-run race. With this fix, the symptoms stop appearing.
Fix: claim the slot synchronously immediately after the has() check,
release it on the heartbeat-defer early-return path. No await happens
between check and claim, so the race window is closed.
Test added under
packages/engine/src/__tests__/reliability-interactions/concurrent-execute-race.test.ts
verified to fail on the prior (a1b1f9aa0) executor.ts and pass on the
fixed version:
- Two concurrent execute() calls produce the SAME number of
createFnAgent invocations as one execute() call (no amplification).
- A second sequential execute() after the first completes IS allowed
(slot was released).
The task must have assignedAgentId set to exercise the race \u2014 without
it, the short-circuit `assignedAgentId && ...` evaluates the left side
to false synchronously, and no await happens.
Full engine suite: 5048+ tests pass. The 7 transient test-file failures
in the broad parallel run are pre-existing flaky real-git tests
(branch-conflicts-zero-unique, branch-conflicts-recovery,
merger-overlap-guard subprocess-guard contention) \u2014 all of them pass
when run alone or as a smaller group, none touch the executor.execute()
path.
Fusion-Task-Id: FN-4811
Two follow-ups stacked on the FN-4811 active-worktree liveness gate:
1. Stale conflict-path recovery (FN-4813 production failure)
When 'git worktree remove --force' fails with 'fatal: validation failed,
cannot remove working tree', the worktree directory is missing on disk
and the git admin entry is stale. Without this recovery, every retry of
tryCreateWorktree on a stale conflict path failed 3 times with
'automatic cleanup failed', leaving tasks unable to create worktrees.
cleanupConflictingWorktree now catches that specific error class, runs
'git worktree prune' to drop the stale admin entry, best-effort deletes
the branch, and returns success so the caller can proceed.
Implementation note: the original attempt used existsSync(worktreePath)
as a pre-check, but vitest's vi.clearAllMocks() can leave the existsSync
mock returning undefined, causing the new branch to fire inside tests
that didn't expect it and leading to worker OOM in
executor-worktree.test.ts. The error-class-based catch is robust against
mock state and matches the real production failure signal exactly.
2. Collapsed broken FN-4806 nested branches
The previous FN-4806 refactor (commit 087b1a766) accidentally nested the
genuine 'agent finished without calling fn_task_done after N retries'
failure path INSIDE the silent-recovery branch, meaning ordinary
failures were being silently requeued (no status=failed, no onError, no
retry-budget burn) instead of being surfaced.
Restored the clean two-branch structure:
} else if (retryAbortedDueToReclaim) {
// silent recovery (FN-4806)
} else {
// genuine no-fn_task_done exhaustion: mark failed, onError, burn budget
}
Also clears baseCommitSha on silent recovery (matches the parallel
session-start-failure path's metadata clearing).
Tests:
- Adds 'FN-4811 follow-up (FN-4813): recovers from validation failed'
case to active-worktree-removal-liveness.test.ts (12 total cases).
- executor-recovery.test.ts no-fn_task_done reclaim coverage now
asserts baseCommitSha is cleared.
- executor-recovery.test.ts 'does not mark task as failed when invalid
transition error occurs on completion' regression fixed by restoring
the failure-path branch.
- executor-core.test.ts 'still enforces fn_task_done requirement in
fast mode' restored.
Full engine suite: 307 files, 5037 tests pass, 1 skipped. Lint clean.
Fusion-Task-Id: FN-4811
The executor's conflict-recovery paths (cleanupConflictingWorktree,
handleBranchConflict, and tryCreateWorktree's live-foreign/stale-resolved
branches) could force-remove a worktree even when it was currently bound
to an active executor session. This caused the FN-4781/FN-4804 cascade:
- 'Execution blocked: assigned worktree path disappeared mid-task' as
git deleted the live agent's filesystem out from under it
- Two parallel runs for the same task alive simultaneously, with the
second run started in a fresh worktree while the first was still
holding the old session
- Cross-task log attribution (an FN-4804 runContext writing to FN-4781)
- Post-merge 'branch tip misbound but content found on main via trailer'
rescues firing on every successful merge as the bookkeeping was
corrupted mid-merge
Adds a hard liveness gate centralized in findActiveWorktreeOwner(), which
checks both the in-memory activeWorktrees map and the DB for non-done,
non-paused, in-progress tasks bound to the worktree. The gate fires at
two points:
1. cleanupConflictingWorktree returns false (refuses removal) when an
active owner is found, logging an FN-4811 refusal entry.
2. handleBranchConflict short-circuits to 'sticky' BEFORE invoking
inspectBranchConflict, because some inspection branches force-remove
unconditionally.
When cleanup is refused, the live-foreign and stale-resolved branches in
tryCreateWorktree now FALL THROUGH to the suffix-rename path (rather
than returning null) so the requesting task can still proceed without
disturbing the live owner.
Tests:
- New reliability-interactions backstop at
src/__tests__/reliability-interactions/active-worktree-removal-liveness.test.ts
covers findActiveWorktreeOwner (5 cases: in-memory match, requesting
task excluded, DB-level match, paused exclusion, terminal-column
exclusion, self-exclusion), cleanupConflictingWorktree gate (3 cases:
in-memory refuse, DB refuse, no-owner proceed), and handleBranchConflict
gate (2 cases: short-circuit + inspection-skipped, no-owner proceeds).
- Updates existing executor-worktree.test.ts assertion that was
documenting the bug behavior (force-removing active worktree) to match
the new contract (refuses + falls through to suffix-rename).
Full engine suite: 307 files, 5035 tests pass.
Fusion-Task-Id: FN-4811
The executor's no-fn_task_done retry loop has two reclaim signals
(retryAbortedDueToReclaim=true): (1) pre-retry liveness recheck where the
task DB shows worktree/branch was cleared, and (2) session-start failure
because the worktree path no longer exists. Both are engine self-heal
situations triggered by FN-4546 stale-active-branch reclaim, FN-4742
self-healing removals, or related housekeeping paths — the agent never
got a fair retry attempt.
Previously these surfaced as task status=failed with error
'Worktree/branch reclaimed during no-fn_task_done retry — requeueing',
fired onError, and burned the taskDoneRetryCount budget. Three legitimate
problems followed: tasks accumulated spurious failures in the UI, the
exhausted-budget branch escalated reclaimed tasks to in-review instead of
retrying, and the noise masked the underlying worktree-removal regression
(FN-4811).
Now the reclaim branch silently:
- clears stale worktree/branch metadata so the next pickup creates a fresh worktree
- requeues to todo with preserveProgress
- logs an informational 'engine self-heal, no failure' line
- does NOT set status=failed, does NOT bump taskDoneRetryCount, does NOT call onError
The genuine 'agent finished without calling fn_task_done after N retries'
exhaustion path (retryAbortedDueToReclaim=false) is unchanged.
Tests updated in
packages/engine/src/__tests__/reliability-interactions/executor-no-task-done-vs-worktree-reclaim.test.ts
to assert the new silent-recovery contract on all three reclaim paths.
Fusion-Task-Id: FN-4806