Phase A (Foundation) of
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`.
Three units, one commit each. No operator-visible behavior change.
## U1 — Lifecycle-column resolution seam
`resolveLifecycleColumns(ir)` returns `{ intake, hold, wip, review,
complete, archived }` — the first column carrying each trait,
`undefined` for a role no column carries.
`resolveTaskLifecycleColumns(store, taskId, cache?)` is the store-aware
form; the cache is caller-owned so a sweep reads one IR per workflow
rather than one per card.
A v1/column-less IR resolves to `undefined` for the **whole struct**
rather than a struct of undefined roles. A caller must be able to
distinguish "this workflow declares no hold column" (a real shape to
honor) from "no column vocabulary at all" (skip and log) — only the
second licenses conservative fallback.
Nothing consumes the seam yet; Phases B–D convert the ~207 hardcoded
column literals onto it.
## U2 — Delete the pre-cutover parity machinery (delete-only)
**`workflow-columns-settings.ts`** — `isWorkflowColumnsEnabled` had the
body `return true`. Six live call sites branched on it, so every
flag-OFF arm was dead code that read as a supported configuration.
Deleted; surviving side inlined at self-healing's transitionPending
sweep, the scheduler's per-column capacity diagnostic, merge-trait's
policy resolver, the board-workflows payload, two task-workflow routes,
and the CLI TUI's column enrichment.
**`workflow-parity.ts`** — asserted the default workflow's adjacency
*equals* the legacy `VALID_TRANSITIONS`. U11 deliberately breaks that
equality by merging Todo into Planning, so this is not a stale assertion
to update; it is a contract against the target state. Its emitter
(`workflow-parity-observer.ts`) is already a tombstone, so
`getWorkflowParitySummary` and `computeWorkflowColumnsGraduationReport`
aggregated run-audit rows nothing writes and had no caller outside
`TaskStore`. Both store methods go with it.
`flagEnabled` stays on the board-workflows **wire** as a constant `true`
— shipped dashboard clients still branch on it, and changing the
response shape is not a deletion. U10 retires the field once no client
reads it.
The `legacy-tombstones` ratchet is extended to both files plus seven
symbols, each with the reason it is gone.
### ⚠️ Finding: the third listed deletion was NOT dead
The plan also lists "the flag-off inline move path" in
`task-store/moves.ts`. It is **not** deleted, per U2's execution note
("any behavior change found while removing a branch means the branch was
not dead").
That path is gated on `isWorkflowColumnsCompatibilityFlagEnabled`
(`store.ts:38`) — a **different** function from the always-true public
helper. It reads the raw `experimentalFeatures.workflowColumns` setting,
which nothing in production sets (`settings-schema.ts:396` — "no default
flags are emitted"; zero non-test writers; the operator's own
`~/.fusion/settings.json` has no such key). So `useWorkflow` is false
for effectively every real project: the flag-OFF inline side effects are
the **live** default move path and the flag-ON `default-workflow-hooks`
path is the dead one. The code says so itself at `moves.ts:638`.
Deleting that branch would swap every project onto an untravelled code
path — a behavior change, not a deletion.
**Carry this into Phases B and C, stated plainly so the plan's error is
not repeated:**
> **The inline move path in `moves.ts` is LIVE.
`default-workflow-hooks.ts` (the trait-hook path) is DEAD.** KTD-6
asserted the inverse. Until the convergence unit lands, **nothing may
assume trait hooks run** — a guard, sweep, or subscriber written against
`applyDefaultWorkflowMoveEffects` would never fire in production and
would still pass its tests.
Convergence is **not** attempted here. It is its own unit (Phase A2)
with a proper equivalence proof, per operator decision.
### U3's emit point is on the LIVE path — the seam is not born dead
Worth stating explicitly because it is the failure mode that would make
every later subscriber silently never fire: the `TaskTransitioned` emit
is **not** inside the `if (useWorkflow)` branch. That block closes at
`moves.ts:1212`; the emit sits at `:1214`, beside the existing
`store.emit("task:moved", …)`, on the unconditional post-commit path. It
therefore fires on **both** the live inline path and the dead hooks
path, and the convergence unit inherits the obligation to keep it firing
on whichever path survives — same events, same order, same payloads.
The graph-side emitters (`NodeEntered`, `RunSuspended`) carry the same
risk from a different direction: the bus refuses an invalid payload
*silently* by design, so an emitter regression would stop the event with
no test failure. They are asserted end-to-end through the real bus —
"did a subscriber actually receive it", not "was emit called" — because
a spy passes on a refused payload. The `moveTaskInternalImpl` emit does
**not** yet have that end-to-end assertion against a real store move;
that proof belongs to the convergence unit, which has to build the
both-paths fixture anyway.
## U3 — Post-commit event seam with a transactional outbox
**The bus is not a queue, not a transaction participant, and not a
delivery guarantee.** Durable follow-on work uses the transactional
outbox — a `workflow_work_items` row written *inside* the transition
transaction (the shape `createCompletionHandoffWorkflowWork` already
uses). "Emit after commit, let a subscriber enqueue the work" has a
crash window where a process dies between commit and subscriber, leaving
no event *and* no work-item row, so required work is skipped permanently
with nothing to recover from. Post-commit subscribers therefore carry
only losable reactions.
Emission is consequently lossy and isolated by design: a throwing or
rejecting subscriber is caught and logged, cannot roll back the
transition, and cannot stop the others. Deliveries append to one serial
chain, so two transitions on a task deliver in commit order.
The ids/outcomes-only rule is **mechanised, not documented** —
run-audit's equivalent lives only in prose and has been violated
repeatedly. A payload carrying an object body or a prose string is
refused at the emit boundary and never reaches a subscriber or log sink.
It degrades rather than throws: the emitter is post-commit, so a shape
bug must not become a lifecycle failure.
Emit points: `TaskTransitioned` from the single post-commit point in
`moveTaskInternalImpl`; `NodeEntered` and `RunSuspended` from the graph
column boundary, the latter *after* the durable continuation is
persisted so an observed suspension implies a resumable run.
`registerWorkflowEventSubscribers` (engine) is empty on purpose —
U7/U8/U10 move real reactions onto it, each with the characterization
test proving the reaction was non-authoritative first.
## Verification
- `pnpm test:gate` — green (2/10, 16/299, 1/71).
- `pnpm lint`, `pnpm build`, `tsc --noEmit` on core and engine — green.
- U1: 20 tests in `workflow-lifecycle-traits.test.ts`, including the
fully-renamed-workflow case (fails if the resolver falls back to a
literal) and a shared-cache read-count assertion.
- U2: `legacy-tombstones.test.ts` green with the extended ratchet;
`board-workflows`, `merge-trait`, `workflow-graph-executor-parity`, and
move-hook suites green with no expectation edits.
- U3: 20 bus-invariant unit tests (isolation, ordering, the allowed-key
and required-key halves of the ids-only rule, lossiness) plus 3
end-to-end emitter-delivery tests; 5 outbox tests against a **real
PostgreSQL** work-item table (crash survival, rollback, at-least-once
redelivery on lease expiry, idempotent handler → one effect,
dropped-subscriber vs. durable work). A hand-written fake of the lease
predicate would only prove the fake redelivers.
**Not verified:** the `moveTaskInternalImpl` emit is confirmed on the
unconditional post-commit path by structure and by the surrounding
tests, but is *not* yet asserted end-to-end against a real store move on
both flag settings — that is Phase A2's fixture. The engine subscriber
registry ships empty by design, so no production subscriber exercises
the bus end-to-end yet. `settings-defaults.test.ts` has one pre-existing
failure on `main` (a logger-prefix mismatch in the
`mergeIntegrationWorktree=cwd-main` warning) — confirmed present on a
clean tree, unrelated to this branch.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Workflow lifecycle columns are now derived from workflow definitions,
supporting renamed and custom workflows.
* Added post-commit lifecycle events for task transitions, node entry,
and run suspend/resume with validated payloads.
* Follow-on processing for lifecycle emissions is now more robust
(rollback-safe, at-least-once delivery, idempotent handling).
* **Bug Fixes**
* Workflow board responses, task enrichment, and promotion no longer
depend on workflow-columns feature-flag gating.
* Subscriber failures no longer impact committed workflow transitions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
440 lines
18 KiB
TypeScript
440 lines
18 KiB
TypeScript
/*
|
|
FNXC:WorkflowColumnBoundary 2026-07-18-20:35:
|
|
U1 — the graph is the sole lifecycle driver. As traversal crosses from a node in
|
|
column A into a node in column B, the card MOVES to column B through the store's
|
|
trait-hook `moveTask` path. This controller is the seam the graph executor calls
|
|
on every real-node entry; it owns the trait decision (move vs. park), the audit
|
|
emission (KTD-12), and the durable IR pin / drift guard (KTD-3).
|
|
|
|
Rules encoded here:
|
|
- KTD-1: `end` and failure-terminal arrivals never move — the executor never
|
|
calls `onNodeEntry` for a node without a column, and failure edges in the
|
|
benchmark terminate at `end` (columnless), so the card stays put and parks in
|
|
place. Only column-bearing node entries move.
|
|
- KTD-2: a hold→wip boundary is NEVER graph-moved. The graph parks the card at
|
|
the ready-for-release seam and the scheduler's capacity sweep (U4) is the sole
|
|
actor that performs the hold→wip move. This controller records the seam and
|
|
performs no move on that boundary.
|
|
- KTD-3: each node entry pins the resolved IR (version/content hash) durably on
|
|
the run state; `detectDrift` compares the stored pin against the current IR at
|
|
run start and parks with `task:reconcile-workflow-drift` on mismatch instead
|
|
of traversing a mutated graph.
|
|
- KTD-12: `task:column-transition` and `task:reconcile-workflow-drift` carry
|
|
ids/counts/outcomes-only metadata — no prose, no model ids.
|
|
|
|
Idempotency (U1 scenario 4 — kill/restart mid-transition settles exactly once):
|
|
the controller tracks the card's current column and only moves when the entered
|
|
node's column differs. A re-entered node (rework loop) or a re-run after restart
|
|
in the same column is a no-op, so exactly one move + one audit event is produced
|
|
per real boundary crossing. The crash-safe `transitionPending` marker written
|
|
in-txn by `moveTaskInternal` is the store-level half of that guarantee.
|
|
*/
|
|
|
|
import {
|
|
type WorkflowIr,
|
|
type WorkflowIrNode,
|
|
type WorkflowIrPin,
|
|
computeWorkflowIrPin,
|
|
detectWorkflowDrift,
|
|
findWorkflowColumn,
|
|
isHoldToWipBoundary,
|
|
resolveColumnFlags,
|
|
TransitionRejectionError,
|
|
emitWorkflowLifecycleEvent,
|
|
} from "@fusion/core";
|
|
|
|
/** Run-audit event emitted by the boundary controller (KTD-12, ids/counts only). */
|
|
export type WorkflowColumnBoundaryAuditEvent =
|
|
| {
|
|
type: "task:column-transition";
|
|
taskId: string;
|
|
workflowId: string;
|
|
fromColumn: string;
|
|
toColumn: string;
|
|
irHash: string;
|
|
nodeId: string;
|
|
}
|
|
| {
|
|
type: "task:reconcile-workflow-drift";
|
|
taskId: string;
|
|
workflowId: string;
|
|
pinnedNodeId: string;
|
|
reason: string;
|
|
};
|
|
|
|
/** The single store move seam. Wired to `store.moveTask(id, toColumn, {
|
|
* moveSource: "engine", workflowMoveSource: "workflow-graph" })`. */
|
|
export type WorkflowColumnMove = (
|
|
toColumn: string,
|
|
ctx: { fromColumn: string; nodeId: string },
|
|
) => Promise<void>;
|
|
|
|
export interface WorkflowColumnBoundaryDeps {
|
|
taskId: string;
|
|
workflowId: string;
|
|
ir: WorkflowIr;
|
|
/** The card's lifecycle column at run start (task.column). */
|
|
initialColumn: string;
|
|
/** The store move seam. Absent → the controller records intent but performs no
|
|
* move (byte-inert; used by unit tests that assert decisions only). */
|
|
moveTask?: WorkflowColumnMove;
|
|
/** Emit an ids/counts/outcomes-only run-audit event (KTD-12). */
|
|
emitAudit?: (event: WorkflowColumnBoundaryAuditEvent) => void | Promise<void>;
|
|
/** KTD-3: persist the per-node-entry IR pin (wired to the task row's
|
|
* `workflowIrPin*` fields via `createStoreIrPinPersistence` in production). */
|
|
pinNodeEntry?: (pin: WorkflowIrPin) => void | Promise<void>;
|
|
/*
|
|
FNXC:WorkflowIrPin 2026-07-19-21:10 (KTD-3 drift-park loop fix, PR #2342):
|
|
Clear the durable pin the moment drift is DETECTED. The pin that fired the drift
|
|
guard is by definition stale (its node/column/hash no longer exists in the current
|
|
IR); leaving it on the row made the park permanent — every requeue re-loaded the
|
|
same pin, re-fired detectDrift, and re-failed until an operator manually nulled
|
|
the fields. Clearing at drift-park time makes the park self-correcting: the next
|
|
ordinary requeue loads no prior pin, re-resolves the CURRENT IR fresh, and
|
|
proceeds — adopting the changed workflow is exactly the desired outcome.
|
|
Wired to `createStoreIrPinPersistence().clearPin` in production; absent → the
|
|
pre-fix posture (pin left in place) for minimal test harnesses.
|
|
*/
|
|
clearPin?: () => void | Promise<void>;
|
|
/** KTD-3: the pin recorded by a prior (possibly crashed) run, for drift check. */
|
|
priorPin?: WorkflowIrPin;
|
|
/** Optional diagnostics sink; never throws into the run. */
|
|
onWarn?: (message: string, detail: Record<string, unknown>) => void;
|
|
/** Persist a durable continuation before control returns to the scheduler. */
|
|
onSuspend?: (suspension: Extract<WorkflowColumnBoundaryEntryResult, { kind: "suspended" }>) => void | Promise<void>;
|
|
}
|
|
|
|
/** The seam the graph executor consumes. */
|
|
export interface WorkflowColumnBoundary {
|
|
/** The card's current lifecycle column (updated after each successful move). */
|
|
currentColumn(): string;
|
|
/** Cross into `node.column` when it differs from the current column. */
|
|
onNodeEntry(node: WorkflowIrNode): Promise<WorkflowColumnBoundaryEntryResult | void>;
|
|
/** KTD-3 drift guard — run once at graph start. Returns true when the pinned
|
|
* node/column is gone from the current IR (run must park, not traverse). */
|
|
detectDrift(): Promise<boolean>;
|
|
}
|
|
|
|
export type WorkflowColumnBoundaryEntryResult =
|
|
| { kind: "entered" }
|
|
| {
|
|
kind: "suspended";
|
|
reason: "capacity";
|
|
nodeId: string;
|
|
fromColumn: string;
|
|
toColumn: string;
|
|
irHash: string;
|
|
};
|
|
|
|
/*
|
|
FNXC:WorkflowIrPin 2026-07-19-18:30 (KTD-3 / U9b):
|
|
Store-backed KTD-3 pin persistence. The U9b schema landed the durable pin as three
|
|
task-row fields (`workflowIrPin` = content hash, `workflowIrPinNodeId`,
|
|
`workflowIrPinColumnId` — migration 0026), threaded through updateTask /
|
|
serialization / slim projections in core. This factory binds that row surface to
|
|
the boundary's `pinNodeEntry` / `loadPriorPin` seams:
|
|
- pinNodeEntry writes via `updateTask` ONLY when the pin actually changed
|
|
(same node re-entry / rework loops and same-hash chains are free — one row
|
|
write per real node entry, not per traversal step);
|
|
- loadPriorPin reads the fields back off the task row at run start so a
|
|
restart/re-entry compares the crashed run's pin against the CURRENT IR and
|
|
takes the drift-park path (`task:reconcile-workflow-drift`) on mismatch.
|
|
Degradation: a store lacking `updateTask`/`getTask`, or whose rows do not carry
|
|
the pin fields (in-memory fakes, pre-U9b DBs), yields no prior pin and no-op
|
|
writes — exactly the pre-wiring inert posture, so legacy harnesses are untouched.
|
|
*/
|
|
|
|
/** The minimal task-row surface the pin persistence needs. Structural on purpose
|
|
* so real stores and test fakes both fit without importing the full TaskStore. */
|
|
export interface WorkflowIrPinStoreSurface {
|
|
updateTask?: (
|
|
id: string,
|
|
updates: {
|
|
workflowIrPin?: string | null;
|
|
workflowIrPinNodeId?: string | null;
|
|
workflowIrPinColumnId?: string | null;
|
|
},
|
|
) => unknown;
|
|
getTask?: (id: string) => Promise<{
|
|
workflowIrPin?: string;
|
|
workflowIrPinNodeId?: string;
|
|
workflowIrPinColumnId?: string;
|
|
}>;
|
|
}
|
|
|
|
/** Build the store-backed `pinNodeEntry`/`loadPriorPin` pair for one task. */
|
|
export function createStoreIrPinPersistence(
|
|
store: WorkflowIrPinStoreSurface,
|
|
taskId: string,
|
|
): {
|
|
pinNodeEntry: (pin: WorkflowIrPin) => Promise<void>;
|
|
loadPriorPin: () => Promise<WorkflowIrPin | undefined>;
|
|
clearPin: () => Promise<void>;
|
|
} {
|
|
// Last pin known to be on the row (loaded or written) — the change-only gate.
|
|
let lastPin: WorkflowIrPin | undefined;
|
|
const samePin = (a: WorkflowIrPin, b: WorkflowIrPin): boolean =>
|
|
a.nodeId === b.nodeId && a.irHash === b.irHash && a.columnId === b.columnId;
|
|
|
|
return {
|
|
pinNodeEntry: async (pin) => {
|
|
if (typeof store.updateTask !== "function") return; // no-op degradation
|
|
if (lastPin && samePin(lastPin, pin)) return; // unchanged → no row write
|
|
await store.updateTask(taskId, {
|
|
workflowIrPin: pin.irHash,
|
|
workflowIrPinNodeId: pin.nodeId,
|
|
workflowIrPinColumnId: pin.columnId ?? null,
|
|
});
|
|
lastPin = pin;
|
|
},
|
|
loadPriorPin: async () => {
|
|
if (typeof store.getTask !== "function") return undefined;
|
|
try {
|
|
const row = await store.getTask(taskId);
|
|
if (!row?.workflowIrPin || !row.workflowIrPinNodeId) return undefined;
|
|
const pin: WorkflowIrPin = {
|
|
nodeId: row.workflowIrPinNodeId,
|
|
irHash: row.workflowIrPin,
|
|
columnId: row.workflowIrPinColumnId ?? undefined,
|
|
};
|
|
lastPin = pin;
|
|
return pin;
|
|
} catch {
|
|
// A store that cannot read the row degrades to "no prior pin" — the
|
|
// drift guard stays inert rather than failing the run on bookkeeping.
|
|
return undefined;
|
|
}
|
|
},
|
|
// FNXC:WorkflowIrPin 2026-07-19-21:10: null all three row fields (production
|
|
// task-update.ts treats null as clear-to-undefined) so a requeued run loads
|
|
// NO prior pin and re-resolves the current IR instead of re-firing drift.
|
|
clearPin: async () => {
|
|
if (typeof store.updateTask !== "function") return; // no-op degradation
|
|
await store.updateTask(taskId, {
|
|
workflowIrPin: null,
|
|
workflowIrPinNodeId: null,
|
|
workflowIrPinColumnId: null,
|
|
});
|
|
lastPin = undefined;
|
|
},
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Build the boundary controller for one graph run. Additive: when the executor
|
|
* has no `columnBoundary` dep wired it performs no lifecycle moves at all, so
|
|
* every legacy runner/executor test stays byte-identical.
|
|
*/
|
|
export function createWorkflowColumnBoundary(
|
|
deps: WorkflowColumnBoundaryDeps,
|
|
): WorkflowColumnBoundary {
|
|
let column = deps.initialColumn;
|
|
|
|
const warn = (message: string, detail: Record<string, unknown>): void => {
|
|
try {
|
|
deps.onWarn?.(message, detail);
|
|
} catch {
|
|
/* diagnostics must never affect the run */
|
|
}
|
|
};
|
|
|
|
const flagsFor = (columnId: string) => {
|
|
const col = findWorkflowColumn(deps.ir, columnId);
|
|
return col ? resolveColumnFlags(col) : {};
|
|
};
|
|
|
|
return {
|
|
currentColumn: () => column,
|
|
|
|
async detectDrift(): Promise<boolean> {
|
|
const pin = deps.priorPin;
|
|
if (!pin) return false;
|
|
const reason = detectWorkflowDrift(deps.ir, pin);
|
|
if (!reason) return false;
|
|
try {
|
|
await deps.emitAudit?.({
|
|
type: "task:reconcile-workflow-drift",
|
|
taskId: deps.taskId,
|
|
workflowId: deps.workflowId,
|
|
pinnedNodeId: pin.nodeId,
|
|
reason,
|
|
});
|
|
} catch (err) {
|
|
warn("drift audit emit failed", { error: err instanceof Error ? err.message : String(err) });
|
|
}
|
|
// FNXC:WorkflowIrPin 2026-07-19-21:10 (drift-park loop fix): the pin that
|
|
// fired the guard is stale — clear it NOW so the park self-corrects on the
|
|
// next requeue (fresh IR resolution) instead of looping forever. Fail-soft:
|
|
// a failed clear merely re-fires drift next run (the pre-fix behavior),
|
|
// never fails this run's bookkeeping.
|
|
try {
|
|
await deps.clearPin?.();
|
|
} catch (err) {
|
|
warn("stale drift pin clear failed", { error: err instanceof Error ? err.message : String(err) });
|
|
}
|
|
return true;
|
|
},
|
|
|
|
async onNodeEntry(node: WorkflowIrNode): Promise<WorkflowColumnBoundaryEntryResult> {
|
|
const toColumn = node.column;
|
|
|
|
/*
|
|
FNXC:WorkflowEvents 2026-07-27-15:20 (U3 / R5, PR #2467 review):
|
|
Announce the NODE ENTRY, not the column crossing — so this fires BEFORE the
|
|
columnless short-circuit and before the same-column no-op below. Traversal
|
|
genuinely entered the node in all three cases, and a subscriber tracking
|
|
graph progress must see rework loops and terminal `end` arrivals, which a
|
|
crossing-only signal hides. `column` is omitted for a columnless node,
|
|
which is exactly why `NodeEnteredEvent.column` is optional.
|
|
|
|
The paired `TaskTransitioned` comes from the store's own post-commit point,
|
|
so a real crossing produces both and every other entry produces only this
|
|
one. Neither is authoritative for any lifecycle decision.
|
|
*/
|
|
emitWorkflowLifecycleEvent({
|
|
type: "NodeEntered",
|
|
taskId: deps.taskId,
|
|
at: new Date().toISOString(),
|
|
workflowId: deps.workflowId,
|
|
nodeId: node.id,
|
|
...(toColumn ? { column: toColumn } : {}),
|
|
});
|
|
|
|
// KTD-1: a columnless node (e.g. `end`) never moves the card.
|
|
if (!toColumn) return { kind: "entered" };
|
|
|
|
// KTD-3: pin the resolved IR for this node-entry (durable seam).
|
|
try {
|
|
await deps.pinNodeEntry?.(computeWorkflowIrPin(deps.ir, node.id));
|
|
} catch (err) {
|
|
warn("ir pin write failed", { nodeId: node.id, error: err instanceof Error ? err.message : String(err) });
|
|
}
|
|
|
|
// Idempotent: a re-entered/rework node or a same-column node chain no-ops.
|
|
if (toColumn === column) return { kind: "entered" };
|
|
|
|
const fromColumn = column;
|
|
|
|
// KTD-2: never graph-move a hold→wip boundary — the scheduler is the sole
|
|
// mover there; the card parks at the ready-for-release seam (U4 performs the
|
|
// actual release). Record the seam but perform no move here.
|
|
if (isHoldToWipBoundary(flagsFor(fromColumn), flagsFor(toColumn))) {
|
|
warn("hold→wip boundary parked at ready-for-release seam (scheduler-owned)", {
|
|
fromColumn,
|
|
toColumn,
|
|
irHash: computeWorkflowIrPin(deps.ir, node.id).irHash,
|
|
nodeId: node.id,
|
|
});
|
|
const suspension = {
|
|
kind: "suspended",
|
|
reason: "capacity",
|
|
nodeId: node.id,
|
|
fromColumn,
|
|
toColumn,
|
|
irHash: computeWorkflowIrPin(deps.ir, node.id).irHash,
|
|
} as const;
|
|
await deps.onSuspend?.(suspension);
|
|
/*
|
|
FNXC:WorkflowEvents 2026-07-27-12:05 (U3 / R5):
|
|
Emitted AFTER `onSuspend` has persisted the durable continuation, so an
|
|
observed `RunSuspended` implies the continuation exists. The continuation
|
|
— not this event — is what the scheduler resumes from; dropping the event
|
|
costs a notification, never a stranded run.
|
|
*/
|
|
emitWorkflowLifecycleEvent({
|
|
type: "RunSuspended",
|
|
taskId: deps.taskId,
|
|
at: new Date().toISOString(),
|
|
workflowId: deps.workflowId,
|
|
nodeId: node.id,
|
|
reason: "capacity",
|
|
fromColumn,
|
|
toColumn,
|
|
});
|
|
return suspension;
|
|
}
|
|
|
|
// The single mover: the store's trait-hook moveTask path.
|
|
if (deps.moveTask) {
|
|
try {
|
|
await deps.moveTask(toColumn, { fromColumn, nodeId: node.id });
|
|
} catch (err) {
|
|
/*
|
|
FNXC:WorkflowReviewGates 2026-07-26-13:40:
|
|
A CAPACITY rejection on a real (non-hold→wip) boundary is transient, not a graph failure.
|
|
This became reachable once the pre-merge review gates moved into `in-review`: the paired
|
|
remediation node crosses in-review → in-progress, and that crossing re-enters a
|
|
capacity-bearing column. Capacity is enforced in-transaction and is never bypassable, so
|
|
if the pool filled while the gate ran, the move is rejected — and rethrowing here killed
|
|
the run at the remediation node, stranding the card in `in-review` behind a failed
|
|
pre-merge step with nothing scheduled to fix it.
|
|
Park it the same way the hold→wip seam parks instead: a `suspended` result unwinds to a
|
|
clean `outcome: "success"` with a suspension marker (no failure recorded, worktree/branch
|
|
and the durable failed gate result preserved), so the next graph re-run retries the
|
|
remediation move once a slot frees. Non-capacity rejections (invariant violations) are
|
|
real errors and still propagate.
|
|
*/
|
|
if (err instanceof TransitionRejectionError && err.rejection.code === "capacity-exhausted") {
|
|
warn("graph column move parked — column at capacity (will retry on next run)", {
|
|
fromColumn,
|
|
toColumn,
|
|
nodeId: node.id,
|
|
});
|
|
const suspension = {
|
|
kind: "suspended",
|
|
reason: "capacity",
|
|
nodeId: node.id,
|
|
fromColumn,
|
|
toColumn,
|
|
irHash: computeWorkflowIrPin(deps.ir, node.id).irHash,
|
|
} as const;
|
|
await deps.onSuspend?.(suspension);
|
|
// FNXC:WorkflowEvents 2026-07-27-12:06 (U3): see the hold→wip seam above.
|
|
emitWorkflowLifecycleEvent({
|
|
type: "RunSuspended",
|
|
taskId: deps.taskId,
|
|
at: new Date().toISOString(),
|
|
workflowId: deps.workflowId,
|
|
nodeId: node.id,
|
|
reason: "capacity",
|
|
fromColumn,
|
|
toColumn,
|
|
});
|
|
return suspension;
|
|
}
|
|
// A rejected move (invariant) leaves the card in its current column;
|
|
// routing/parking is U4/U5. Do not advance `column` and do not emit a
|
|
// transition audit for a move that did not happen.
|
|
warn("graph column move rejected", {
|
|
fromColumn,
|
|
toColumn,
|
|
nodeId: node.id,
|
|
error: err instanceof Error ? err.message : String(err),
|
|
});
|
|
throw err;
|
|
}
|
|
}
|
|
|
|
column = toColumn;
|
|
try {
|
|
await deps.emitAudit?.({
|
|
type: "task:column-transition",
|
|
taskId: deps.taskId,
|
|
workflowId: deps.workflowId,
|
|
fromColumn,
|
|
toColumn,
|
|
irHash: computeWorkflowIrPin(deps.ir, node.id).irHash,
|
|
nodeId: node.id,
|
|
});
|
|
} catch (err) {
|
|
warn("column-transition audit emit failed", {
|
|
fromColumn,
|
|
toColumn,
|
|
error: err instanceof Error ? err.message : String(err),
|
|
});
|
|
}
|
|
return { kind: "entered" };
|
|
},
|
|
};
|
|
}
|