fix: resume capacity-parked workflow continuations

The planning-continuation drain skipped every due `kind: "task"` row whose
`waitReason` was not "planning", on the premise that such rows "belong to a
different drain". No such drain exists: `listDueWorkflowWorkItems` has exactly
two callers, this pass and the self-healing reclaim sweep, and the sweep
deliberately leaves `runnable`/`retrying` rows alone as "the dispatcher's own
queue". A capacity-parked continuation was therefore owned by nobody — skipped
here every poll with no state change and no audit row, and passed over there by
design.

Observed on the Fusion board: eight cards sat runnable for up to 8h with the
engine unpaused, 0 tasks in progress, and 4 of 10 worktrees used. Three carried
`waitReason: "capacity"` from the capacity-suspend path; five carried NULL. The
09:04 reclaim sweep had just moved them held -> runnable, handing them to this
drain and simultaneously putting them out of its own reach, so the auto-resume
fix tightened the strand it repaired.

Dispatch stays admission-gated by `admitPlanningContinuation`, so a
capacity-parked card resumes only when a slot is genuinely free.

Also repairs two stale path allowlists in planning-claim-single-writer.ts: the
mission stores and replan-target.ts moved into subdirectories, leaving that
ratchet red on main and accusing the two modules it exists to exclude.

Verified: the patched classifier returns `actionable` for all 8 live stranded
rows; gate + lint green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
gsxdsm
2026-08-11 10:36:26 -07:00
parent eef68fe0b9
commit dfd88e7540
5 changed files with 107 additions and 18 deletions

View File

@@ -0,0 +1,7 @@
---
"@runfusion/fusion": patch
---
summary: Fix queued tasks never starting after the board fills up.
category: fix
dev: `resolvePlanningContinuationCandidate` no longer gates on `waitReason === "planning"`. That skip meant a continuation parked by the capacity-suspend path (`waitReason: "capacity"`) or carrying a NULL reason was owned by nobody — skipped by the only drain that dispatches `runnable` task continuations, and passed over by the self-healing reclaim sweep, which by design leaves `runnable` rows to that drain. Cards sat runnable indefinitely with no state change and no audit row while the engine was idle. Dispatch stays admission-gated by `admitPlanningContinuation`, so a capacity-parked card still resumes only when a slot is genuinely free. Also repairs two stale path allowlists in `planning-claim-single-writer.test.ts` (mission stores and `replan-target.ts` moved into subdirectories), which had left that ratchet red on main.

View File

@@ -550,13 +550,20 @@ describe("#4 an operator-parked item leaves the due window instead of starving t
expect(resolveParkedContinuationDeferral(resolved, NOW)).toBeNull(); expect(resolveParkedContinuationDeferral(resolved, NOW)).toBeNull();
}); });
it("never defers a non-planning item — that item belongs to a different drain", () => { /*
FNXC:WorkflowScheduling 2026-08-11-17:30:
This case previously asserted that a `capacity` item is SKIPPED as "not-planning" because it
"belongs to a different drain". No such drain exists, and that skip stranded eight cards for up to
8h on 2026-08-11. The deferral outcome is unchanged (still null) but for the opposite reason: the
item is actionable now, and deferring ready work would stall the lane.
*/
it("never defers a non-planning item — it is actionable, and deferring ready work stalls the lane", () => {
const resolved = resolvePlanningContinuationCandidate( const resolved = resolvePlanningContinuationCandidate(
dueItem({ waitReason: "capacity" }), dueItem({ waitReason: "capacity" }),
task(), task(),
); );
expect(resolved.kind === "skip" && resolved.reason).toBe("not-planning"); expect(resolved.kind).toBe("actionable");
expect(resolveParkedContinuationDeferral(resolved, NOW)).toBeNull(); expect(resolveParkedContinuationDeferral(resolved, NOW)).toBeNull();
}); });

View File

@@ -101,16 +101,27 @@ const PLANNING_CLAIM_WRITERS = ["packages/engine/src/triage.ts"];
* route around the write patterns, so a NEW entry deserves a look even when the * route around the write patterns, so a NEW entry deserves a look even when the
* module turns out, like this one, to be a reader. * module turns out, like this one, to be a reader.
*/ */
const PLANNING_CLAIM_BINDERS = ["packages/engine/src/replan-target.ts"]; /* FNXC:PlanningClaimSingleWriter 2026-08-11-17:30: relocated to `execution/` — same stale-path
drift as NON_TASK_STATUS_MODULES below, and red on main for the same reason. */
const PLANNING_CLAIM_BINDERS = ["packages/engine/src/execution/replan-target.ts"];
/** /**
* Mission planning is a different entity with its own `status` column. Excluded by * Mission planning is a different entity with its own `status` column. Excluded by
* path rather than by pattern, because a pattern loose enough to tell them apart is * path rather than by pattern, because a pattern loose enough to tell them apart is
* a pattern loose enough to miss a real task write. * a pattern loose enough to miss a real task write.
*/ */
/*
FNXC:PlanningClaimSingleWriter 2026-08-11-17:30:
PATHS, so a MOVE breaks them. Both mission stores were relocated into `missions/` and
`async-stores/` subdirectories, and this list kept naming the old flat paths — so the ratchet
had been failing on main, accusing two mission modules it was written to exclude. A guard that
is red for a reason nobody believes is a guard people stop reading; the drift is the cost of
excluding by path, accepted here because a pattern loose enough to tell mission status from
task status is loose enough to miss a real second writer (see the note above).
*/
const NON_TASK_STATUS_MODULES = [ const NON_TASK_STATUS_MODULES = [
"packages/core/src/mission-store.ts", "packages/core/src/missions/mission-store.ts",
"packages/core/src/async-mission-store.ts", "packages/core/src/async-stores/async-mission-store.ts",
]; ];
function sourceFiles(root: string, base: string = REPO_ROOT): string[] { function sourceFiles(root: string, base: string = REPO_ROOT): string[] {

View File

@@ -62,14 +62,43 @@ describe("resolvePlanningContinuationCandidate", () => {
).toEqual({ kind: "orphan", item, reason: "task-terminal" }); ).toEqual({ kind: "orphan", item, reason: "task-terminal" });
}); });
it("skips non-planning and paused planning items without cancelling", () => { /*
const capacity = workItem("cap", "capacity"); FNXC:WorkflowScheduling 2026-08-11-17:30:
expect(resolvePlanningContinuationCandidate(capacity, task("T-cap"))).toEqual({ The invariant, not the repro: EVERY waitReason a writer can persist must dispatch. The board strand
was found through `capacity` rows, but five of the eight stuck cards carried a NULL reason, so a
capacity-only assertion would have re-shipped the wedge for the majority case. Enumerated surfaces:
`"planning"` (`plan-review-continuation.ts`), `"capacity"` (`workflow-column-boundary-hooks.ts`), and
undefined (every upsert path that omits it). Skipping is now reserved for operator parks alone.
*/
it.each(["planning", "capacity", undefined] as const)(
"dispatches a due continuation whatever stopped it (waitReason=%s)",
(waitReason) => {
const item = workItem(`live-${waitReason ?? "none"}`, waitReason);
const live = task("T-live", { column: "todo" });
expect(resolvePlanningContinuationCandidate(item, live)).toEqual({
kind: "actionable",
item,
task: live,
});
},
);
/*
FNXC:WorkflowScheduling 2026-08-11-17:30:
An operator park still outranks the waitReason relaxation above — a capacity-parked card belonging
to a PAUSED task must stay skipped, or the relaxation would start dispatching work a human stopped.
*/
it("still skips an operator-parked task even on a non-planning continuation", () => {
const capacity = workItem("cap-paused", "capacity");
expect(resolvePlanningContinuationCandidate(capacity, task("T-cap", { paused: true }))).toEqual({
kind: "skip", kind: "skip",
item: capacity, item: capacity,
reason: "not-planning", reason: "paused",
});
}); });
it("skips paused planning items without cancelling", () => {
const paused = workItem("paused", "planning"); const paused = workItem("paused", "planning");
expect(resolvePlanningContinuationCandidate(paused, task("T-p", { paused: true }))).toEqual({ expect(resolvePlanningContinuationCandidate(paused, task("T-p", { paused: true }))).toEqual({
kind: "skip", kind: "skip",
@@ -90,12 +119,17 @@ describe("resolvePlanningContinuationCandidate", () => {
}); });
describe("selectActionablePlanningContinuations", () => { describe("selectActionablePlanningContinuations", () => {
it("retains only planning items whose tasks are present, unpaused, and non-terminal", () => { it("retains every continuation whose task is present, unpaused, and non-terminal", () => {
/* /*
FNXC:WorkflowScheduling 2026-07-21-22:31: FNXC:WorkflowScheduling 2026-07-21-22:31:
Regression for the FN-8470→FN-8471 starvation class: a deleted/archived Regression for the FN-8470→FN-8471 starvation class: a deleted/archived
earlier due row must not remain "actionable" and must not prevent a later earlier due row must not remain "actionable" and must not prevent a later
live planning continuation from being selected. live planning continuation from being selected.
FNXC:WorkflowScheduling 2026-08-11-17:30:
`capacity` and NULL-waitReason rows on live tasks now survive selection — they are this drain's
work too. Only the TASK's condition (missing, parked, terminal) removes a row; why the
continuation stopped never does.
*/ */
const selected = selectActionablePlanningContinuations([ const selected = selectActionablePlanningContinuations([
{ item: workItem("eligible", "planning"), task: task("T-1") }, { item: workItem("eligible", "planning"), task: task("T-1") },
@@ -113,6 +147,8 @@ describe("selectActionablePlanningContinuations", () => {
expect(selected.map(({ item, task: selectedTask }) => [item.id, selectedTask.id])).toEqual([ expect(selected.map(({ item, task: selectedTask }) => [item.id, selectedTask.id])).toEqual([
["eligible", "T-1"], ["eligible", "T-1"],
["capacity", "T-2"],
["no-wait-reason", "T-5"],
["later-live", "FN-8471"], ["later-live", "FN-8471"],
]); ]);
}); });

View File

@@ -169,7 +169,7 @@ const LEGACY_TERMINAL_PAIR: ReadonlySet<string> = new Set(["done", "archived"]);
/** Outcome of resolving one due work item for the planning-continuation drain. */ /** Outcome of resolving one due work item for the planning-continuation drain. */
export type PlanningContinuationResolution = export type PlanningContinuationResolution =
| { kind: "actionable"; item: WorkflowWorkItem; task: Task } | { kind: "actionable"; item: WorkflowWorkItem; task: Task }
| { kind: "skip"; item: WorkflowWorkItem; reason: "not-planning" | "paused" | "awaiting-approval" } | { kind: "skip"; item: WorkflowWorkItem; reason: "paused" | "awaiting-approval" }
| { | {
kind: "orphan"; kind: "orphan";
item: WorkflowWorkItem; item: WorkflowWorkItem;
@@ -180,7 +180,12 @@ export type PlanningContinuationResolution =
* FNXC:WorkflowScheduling 2026-07-21-22:31: * FNXC:WorkflowScheduling 2026-07-21-22:31:
* Classify a due work item after a per-item task load. Lookup failures and * Classify a due work item after a per-item task load. Lookup failures and
* terminal/missing tasks become orphans (cancel); paused planning items stay * terminal/missing tasks become orphans (cancel); paused planning items stay
* held without cancel; non-planning due rows are skipped by this drain. * held without cancel.
*
* FNXC:WorkflowScheduling 2026-08-11-17:30: `waitReason` no longer gates this —
* see the note at the `isTaskBlockedOnApproval` guard for the strand that gate
* caused. The remaining guards (terminal, approval, pause, dispatchable) apply
* to every continuation regardless of why it stopped.
*/ */
export function resolvePlanningContinuationCandidate( export function resolvePlanningContinuationCandidate(
item: WorkflowWorkItem, item: WorkflowWorkItem,
@@ -194,9 +199,30 @@ export function resolvePlanningContinuationCandidate(
if (task.deletedAt || terminal.has(task.column)) { if (task.deletedAt || terminal.has(task.column)) {
return { kind: "orphan", item, reason: "task-terminal" }; return { kind: "orphan", item, reason: "task-terminal" };
} }
if (item.waitReason !== "planning") { /*
return { kind: "skip", item, reason: "not-planning" }; FNXC:WorkflowScheduling 2026-08-11-17:30:
} EVERY due `kind: "task"` continuation is this drain's work, whatever its `waitReason`. This used to
skip anything but `waitReason: "planning"` as "belonging to a different drain" — but no such drain
exists. `listDueWorkflowWorkItems` has exactly two callers: this pass, and the self-healing reclaim
sweep, which deliberately refuses to touch `runnable`/`retrying` rows because they are "the
dispatcher's own queue" (`workflows/stranded-continuation-reclaim.ts`). So a runnable non-planning
row was owned by nobody: skipped here every ~2s poll with no state change and no audit row, and
passed over there by design. Silent, permanent, and invisible — the board simply looks idle.
Observed on the Fusion board 2026-08-11: eight cards (FN-8901/8902/8953/8955/8956/8958/8987/8988)
sat runnable for up to 8h while the engine was unpaused with 0 tasks in progress and 4 of 10
worktrees used. Three carried `waitReason: "capacity"` (written by the capacity-suspend path in
`workflow-column-boundary-hooks.ts` — the graph correctly parks a card when the board is full, and
nothing ever resumed it once capacity freed); five carried a NULL reason. The 09:04 reclaim sweep
had just moved them `held -> runnable`, which HANDED them to this drain and simultaneously put them
out of the sweep's own reach — so the auto-resume fix made the strand tighter than the wedge it
repaired.
Dispatch is node-agnostic (`executor.execute(task)` re-enters the durable graph at the card's own
node) and admission-gated by `admitPlanningContinuation`, which re-checks the real live-task cap.
A capacity-parked row therefore resumes only when a slot is genuinely free — accepting it here
cannot reintroduce the over-cap dispatch the suspend exists to prevent.
*/
/* /*
FNXC:PlanApprovalHold 2026-07-27-19:30 (U7 / R4): FNXC:PlanApprovalHold 2026-07-27-19:30 (U7 / R4):
Dispatching a planning continuation starts a Plan Review run, so a card blocked Dispatching a planning continuation starts a Plan Review run, so a card blocked
@@ -266,9 +292,11 @@ export const PARKED_CONTINUATION_DEFER_MS = 60_000;
* *
* Only the OPERATOR-PARK skips qualify (`awaiting-approval`, `paused`): those are * Only the OPERATOR-PARK skips qualify (`awaiting-approval`, `paused`): those are
* open-ended waits on a human, which is what makes them able to accumulate. * open-ended waits on a human, which is what makes them able to accumulate.
* `not-planning` is deliberately excluded — that item belongs to a different *
* drain, and deferring another owner's work would be this drain reaching outside * FNXC:WorkflowScheduling 2026-08-11-17:30: `not-planning` was the third skip
* its own lane. * reason here and is now gone — this drain owns every `kind: "task"` row, so a
* non-planning continuation is dispatched rather than skipped and has nothing
* left to defer.
* *
* Pure and separately exported so the deferral is testable without constructing a * Pure and separately exported so the deferral is testable without constructing a
* runtime, matching why `resolvePlanningContinuationCandidate` is exported. * runtime, matching why `resolvePlanningContinuationCandidate` is exported.