TaskExecutor.execute() had a classic JS async race window. Original:
async execute(task) {
if (this.executing.has(task.id)) return; // check
const assignedAgentId = task.assignedAgentId;
if (assignedAgentId && await this.shouldDeferForHeartbeat(...)) // AWAIT yields
return;
this.executing.add(task.id); // add (too late)
...
}
Two concurrent execute(task) calls (scheduler dispatch + task:moved event
handler + restart-recovery) both:
1. Pass the synchronous has() check (Set is empty).
2. Enter the awaited shouldDeferForHeartbeat call (yields the event loop).
3. Resume and both call this.executing.add(task.id).
4. Both proceed to create the same worktree path.
Production failure shape (FN-4814 + FN-4811, observed within minutes):
01:30:56 [runA-caoe] Worktree created at /...worktrees/bright-mesa
01:30:56 [runB-w23q] Worktree created at /...worktrees/bright-mesa
01:30:58 worktree liveness assertion failed: not_usable_task_worktree
01:31:48 [thirdRun] also fires liveness assertion fail
01:37:48 In-review stall surfaced [no-worktree-no-merge-confirmed]
This is the root cause of the entire FN-4781/FN-4804/FN-4814/FN-4811
cascade. Every other guard added today (FN-4811 active-session gate,
self-healing reclaim defer, validation-failed recovery, silent reclaim
recovery, integrity-warning dedup) was patching SYMPTOMS of the
duplicate-run race. With this fix, the symptoms stop appearing.
Fix: claim the slot synchronously immediately after the has() check,
release it on the heartbeat-defer early-return path. No await happens
between check and claim, so the race window is closed.
Test added under
packages/engine/src/__tests__/reliability-interactions/concurrent-execute-race.test.ts
verified to fail on the prior (a1b1f9aa0) executor.ts and pass on the
fixed version:
- Two concurrent execute() calls produce the SAME number of
createFnAgent invocations as one execute() call (no amplification).
- A second sequential execute() after the first completes IS allowed
(slot was released).
The task must have assignedAgentId set to exercise the race \u2014 without
it, the short-circuit `assignedAgentId && ...` evaluates the left side
to false synchronously, and no await happens.
Full engine suite: 5048+ tests pass. The 7 transient test-file failures
in the broad parallel run are pre-existing flaky real-git tests
(branch-conflicts-zero-unique, branch-conflicts-recovery,
merger-overlap-guard subprocess-guard contention) \u2014 all of them pass
when run alone or as a smaller group, none touch the executor.execute()
path.
Fusion-Task-Id: FN-4811
The reclaimSelfOwnedBranchConflicts sweep was force-pausing actively-running
tasks. Production failure shape on FN-4819:
1. Self-healing sweep runs every cycle and inspects branch conflicts.
2. For FN-4819, inspection classified the conflict as 'tip-already-merged'
(the task's branch tip was already on main).
3. Sweep called removeWorktree({ reason: SelfHealingBranchConflict }).
4. The FN-4811 active-session gate correctly refused: the worktree was
bound to FN-4819/executor (a live agent session was using it).
5. The thrown ActiveSessionWorktreeRemovalError was caught by the outer
reclaim catch block.
6. The catch escalated to AutoRecoveryDispatcher with class
'branch-conflict-unrecoverable'.
7. decision.action === 'pause' marked the task failed + paused +
pausedReason='branch-conflict-unrecoverable' + moved to in-review.
Net effect: the FN-4811 gate (which is correct \u2014 you can't yank a live
worktree) became a regression source because the self-healing sweep
interpreted the refusal as fatal. Tasks that were actively making progress
got paused with a misleading 'branch conflict unrecoverable' error.
Fix: at the top of the per-task reclaim loop in
reclaimSelfOwnedBranchConflicts, check
activeSessionRegistry.isPathActive(task.worktree) and continue for any
task whose worktree is currently bound to a live session. The reclaim
will retry on the next sweep (sweeps run every cycle) once the session
has finished using the worktree. No data is lost, no decision is forced.
Test added under
packages/engine/src/__tests__/reliability-interactions/reclaim-defers-on-active-session.test.ts
covering:
- The skip path: when activeSessionRegistry has a registration for
task.worktree, the sweep MUST NOT call inspectBranchConflict,
removeWorktree, or isUsableTaskWorktree. The task MUST stay in
in-progress, not be marked failed/paused, not be moved to in-review.
- Control: with no registration, the sweep DOES proceed and reaches
inspectBranchConflict (preserving existing behavior).
Full engine suite: 314 files, 5061 tests pass, 1 skipped. Lint clean.
Build clean.
Fusion-Task-Id: FN-4811
- Clear the active chat room when switching Quick Chat to a direct session
- Keep the hidden session dropdown value and initial-session state aligned with room selection changes
- Add dashboard regression coverage for switching from a room back to a direct chat and include a CLI patch changeset
Fusion-Task-Id: FN-4804