The board stopped moving. Work items churned held -> running -> held at
~3.5/sec across every task, pinning a core and writing ~19k workflowWorkItem
audit rows/hour while nothing executed. Hold reason:
workflow-principal-role-pool-exhausted:executor.
provisionBuiltinWorkflowRoleAgents seeded the four permanent owners (triage,
executor, reviewer, merger) with runtimeConfig.enabled=false, while the
router's available() treats enabled===false as unavailable. The only permanent
principals for every built-in role were unroutable BY CONSTRUCTION — shipped
that way, so any instance without operator-created role agents deadlocks at its
first workflow node. Nothing self-recovers: a pool only changes by operator
action.
Routability of these four is an invariant, not a setting. Unlike an operator's
agent, disabling one does not opt an agent out — it removes the only thing that
can run that stage, and there is no fallback.
- seed built-ins enabled; converge existing rows on provisioning
- enforceBuiltinWorkflowRoleRoutability coerces enabled back at the durable
writeAgent seam, so no REST/UI/plugin/restore path can reintroduce the
deadlock. Other runtimeConfig keys are preserved; operator-owned agents keep
their off switch
- share the static routability predicate (isWorkflowPrincipalEligible) between
provisioning and the router so the two cannot drift apart again
Also fix the spin itself: a principal hold had no cooldown, so the scheduler
re-dispatched instantly and the run re-entered only to re-fence and re-park.
It now records a backoff ladder (15s -> 5m) checked before graph entry, and
logs once per distinct reason instead of every pass — the same self-recovering
shape as holdForSessionContention. The hold never increments `attempt`, so no
existing guard could ever fire.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>