provisionBuiltinWorkflowRoleAgents took a project-scoped pg_advisory_xact_lock
inside transactionImmediate, then ran provision() against the pool. The lock
holder therefore needed a SECOND pooled connection to finish while concurrent
callers occupied the remaining slots blocking on that same advisory lock. With
DEFAULT_POOL_MAX=3 this self-deadlocked: the holder could never complete, the
waiters could never take the lock, and every subsequent query -- that is, every
DB-backed API route -- queued forever behind an exhausted pool. Observed as a
dashboard that booted ("Ready in 6.5s") and then answered no /api request while
the event loop sat idle in kevent; pg_stat_activity showed one session idle in
transaction holding the lock and two active sessions waiting on it.
Thread an optional QueryHandle through listAgents, findAgentByName, createAgent,
and writeAgent so provisioning runs on tx and the lock and its work share one
connection.
Regression test bounds the pool to a single connection, which makes any second
checkout unsatisfiable and fails deterministically rather than racing. Verified
by reverting the one-line fix: the suite hangs past 300s instead of passing in
under 4s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>