Two independent wedges kept cards silently stuck on the board.
1. Workflow principals were capped. `WorkflowAgentCapacity.acquire` enforced
`settings.maxConcurrent` as a project session budget plus an optional
per-agent `maxWorkflowSessions`, and `routeWorkflowPrincipal`'s availability
test applied the same per-agent ceiling. The workflow roles stand in for
STAGES, not workers, and there is typically one agent per role - so the cap
serialized the entire board behind a single Workflow Executor regardless of
maxConcurrent/maxWorktrees. Admission now always succeeds; the lease survives
as bookkeeping (it is what activeSessions counts and what the renewal timer
keeps warm). `maxProjectSessions` is removed from the input rather than
defaulted, so it cannot be reintroduced without deleting the contract, and
the agent-capacity re-route loops in triage and graph admission are deleted
with the refusal they existed to work around.
2. Continuations that stop in `running` or `held` were never re-polled. The
scheduler's due-poll takes only `runnable`/`retrying`; a row claimed through
a path that leaves `leaseExpiresAt` NULL keeps `state: "running"` forever
after its process dies, and `acquireWorkflowWorkItemLease` can only re-take a
`held` row whose blockedReason matches workflow-principal-%. Observed live:
seven cards `running` behind leases from a process that exited ~9h earlier,
two `held` with a NULL blockedReason for 46h, none emitting a single
run-audit row while stranded. A further 33 active-state rows belonged to
archived+soft-deleted tasks (the FK cascade only fires on hard delete).
New sweep `reconcileStrandedWorkflowContinuations` (startup + periodic)
re-queues both stranded shapes and retires dead tasks' rows, gated by the
canonical liveness triple, a 10-minute grace matching the capacity lease
duration, and a compare-and-set on the scanned state so a real claim wins.
The decision is the pure `evaluateStrandedContinuationReclaim`, shared with
its tests so coverage cannot drift from behavior - the drift that let the
FN-8923 sweep ship covering one ninth of this problem.
Verified: pnpm lint, engine typecheck, pnpm test:gate (606 tests), verify:fast,
and the new suite under mutation (removing either guard fails 3 cases). The two
pre-existing failures in self-healing-orphaned-pending-step-results.test.ts
reproduce identically at HEAD without these changes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>