The merge node could not observe a graph abort. WorkflowPrimitiveContext
carried no signal, so requestMerge raced the merge only against its own
30-minute GRAPH_MERGE_TIMEOUT_MS using a controller it owned. A hard-cancel
(user cancel, engine restart, pause/resume) aborted the graph controller and
the walk kept sitting inside the merge node for the full timeout. When the
timeout finally fired it aborted the still-running AI merge -- surfacing as
"Manual-merge failed: Request was aborted" -- and the walk reported
value=merge-timeout for a cancellation it had missed half an hour earlier.
An abort landing between merger-ai's `worktree: null` write and
mergeConfirmed then stranded the card as no-worktree-no-merge-confirmed.
Thread the graph AbortSignal from WorkflowNodeExecutionContext (where it
already existed) through primitiveNodeContext/primitiveContextForNode into
the primitives, and honor it on both merge surfaces:
- requestMerge fails fast when the walk is already cancelled, before
ensureWorkflowMergeBoundaryTask mutates the row or the requester enqueues
a merge, and links the graph signal into its timeout controller via
AbortSignal.any -- raced separately so the walk returns on the abort
rather than waiting on a requester that may never settle.
- The legacy merge seam had the identical unguarded race and gets the same
treatment.
The timeout stays: it bounds a wedged merge queue, which is a different
failure from cancellation. Both signals must stay live -- dropping either
silently restores the stall with no type error.
Cancellation returns a distinct `merge-cancelled` rather than reusing
merge-timeout. Returning `data.status: "failed"` would let classifyMergeFailure
read the unknown reason as merge-failed and route the cancellation into
bounded auto-merge retry, re-requesting the merge the operator just cancelled.
Regression test covers both merge surfaces, both cancel timings (pre-flight
and mid-flight), the no-signal back-compat path, the signal plumbing itself,
and the classification boundary. Verified by removing the fix: 7 of 9 cases
fail, with the mid-flight cases hanging until timeout.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Removes the legacy workflow-step EXECUTION path now that the graph records results
(U2): delete runWorkflowSteps(), the workflow-step seam + runWorkflowStep primitive
(runtime-primitives, workflow-node-handlers, authoritative-driver), and the legacy
execute() step blocks. Keeps task.workflowStepResults + its store write path (the
graph's sink) and executeWorkflowStep/executeScriptWorkflowStep (reused by the graph).
- Watchdog recoverCompletedTask now re-enters via maybeExecuteWorkflowGraph so the
graph re-runs pending gates, records results, and owns the in-review/back-for-fix
transition (KTD-2).
- maybeExecuteWorkflowGraph fails CLOSED (parks) when a store lacks
getTaskWorkflowSelection AND the task has enabled pre-merge steps — closing the
FN-7039 silent-skip class without changing minimal-store implementation runs (KTD-5).
KNOWN GAP (follow-up): the FN-4343 per-step workflowStepScopeEnforcement leak check
lived only in runWorkflowSteps and is NOT yet replicated on the graph path. Merge-time
File Scope enforcement (FileScopeViolationError, squash overlap) is unaffected.
Plan U4.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>