Correct scheduler dispatch diagnostics so capacity decisions use consistent non-negative slot counts.
- Clamp excess semaphore releases at zero and warn once when a slot is returned without an active holder.
- Recompute dispatch capacity at each queue decision, including tasks started earlier in the same scheduler tick.
- Update scheduler and semaphore tests for true binding gates, non-negative diagnostics, and workflow-step env stability.
- Add a patch changeset for the scheduler capacity fix.
Files changed:
.changeset/fn-6423-scheduler-capacity.md | 5 +
packages/engine/src/__tests__/concurrency.test.ts | 32 ++++++
.../src/__tests__/executor-step-session.test.ts | 10 +-
packages/engine/src/__tests__/scheduler.test.ts | 122 ++++++++++++++++++++-
packages/engine/src/concurrency.ts | 29 ++++-
packages/engine/src/scheduler.ts | 117 ++++++++++----------
6 files changed, 248 insertions(+), 67 deletions(-)
Fusion-Task-Id: FN-6423
Fusion-Task-Lineage: a6b2e668-a822-46e9-9cd8-ac267fbde804
- align triage prompt templates on impacted-first verification
- guard verification max lifetime when timeout is disabled
- share idle semaphore leak recovery across scheduler and triage
Fusion-Task-Id: FN-6043
FN-2910 surfaced concurrent reviewer + merger activity on the same task.
Root cause: asymmetric in-flight guards let an unpause-resume kick off a
fresh executor session while a recovery path was already running, and the
auto-merge handoff fired before the executor's finally block finished
cleanup. This sweeps the surrounding lifecycle paths for similar races and
tightens the reviewer pause gate against TOCTOU through runtime setup.
- Symmetric in-flight tracking across `executing`, `recoveringCompleted`,
and `resumingUnpaused`; `recoverCompletedTask` bails when any are set.
- Atomic claim of the recovery slot in the completed-task watchdog before
any awaited work.
- Workflow-rerun bounce returns "bounced" | "skipped-pending" so the
watchdog can no longer log a false-success retry when the original
bounce is still mid-flight.
- Self-healing's completed-task scan re-checks executing IDs inside the
loop instead of trusting a pre-await snapshot.
- 300ms grace period before auto-merge enqueue, giving the executor's
finally block (session disposal, child cleanup) time to drain and
eliminating the residual log-overlap symptom from FN-2910. Test uses
fake timers, no real sleep added.
- New AgentSemaphore.runNested for synchronously nested helper agents
(reviewers): bumps activeCount for honest observability while bypassing
the wait queue, preserving forward-progress fairness for the parent at
low maxConcurrent. Both createReviewStepTool and triage's
createReviewSpecTool now use it.
- New beforeSpawnSession hook on AgentRuntimeOptions/AgentOptions fired
inside createFnAgent immediately before createAgentSession, past every
awaited setup step. Reviewer wires a pause re-check that throws a
sentinel error converted to UNAVAILABLE, closing the TOCTOU window
where pause flipped during runtime resolution or resource loading.
All 2887 engine tests pass; engine + core + cli + dashboard + plugin-sdk
+ pi-claude-cli + desktop typecheck clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Add priority levels (PRIORITY_MERGE=2, PRIORITY_EXECUTE=1, PRIORITY_SPECIFY=0) to AgentSemaphore
- Update acquire() and run() to accept a numeric priority parameter with FIFO ordering within same level
- Wire priority constants into executor (PRIORITY_EXECUTE) and triage (PRIORITY_SPECIFY) callers
- Add comprehensive tests for priority ordering, FIFO within same priority, and dynamic limit interaction
- Include changeset for the priority-based agent scheduling feature