Code review of ef8828f14 found the continuation handover it introduced was a
hand-rolled, non-atomic replacement for a primitive this repo already has, with
six P1 defects — two of which recreated the very deadlock it was written to fix.
The invariant: a task may hold ONE active kind="task" work item
(idx_workflow_work_items_one_active_task_continuation), and that partial unique
index is NOT what a plain upsert's ON CONFLICT targets. So a predecessor the run
has already left makes the write RAISE.
Every continuation write in the executor and triage now goes through
replaceActiveTaskWorkflowContinuation, which retires non-matching active rows
and installs the successor in ONE transaction under the task advisory lock:
- Sibling foreach instances share the template nodeId and differ only by runId,
so the old node-identity guard released nothing and instance #1 re-deadlocked.
- Reacting to a FAILED write could not tell an index conflict from a transient
database error, so it destroyed legitimate held continuations.
- Read-then-write across separate transactions let a concurrent engine lose a
live claim; the lock now serializes it.
- A failed retry left the task with zero active rows and no error, because the
hold then transitioned an already-terminal row and the throw was swallowed.
- The same unguarded write existed on the executor's hold path and at both of
triage's planning-continuation writes; a throw there degraded a recoverable
availability hold into a terminal graph failure.
Coverage moves from a fake store to the real index: the new PG suite proves the
bare upsert raises and that replace handles a different node, a sibling foreach
instance, a held predecessor, and re-entry, plus a drift guard tying the SQL
predicate to ACTIVE_WORKFLOW_WORK_ITEM_STATES. The hand-rolled handover is
tombstoned so it cannot return as a "conflict fix", and both new run-audit
events are documented in the AGENTS.md inventory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>