Commit Graph

3436 Commits

Author SHA1 Message Date
gsxdsm
89d6d76d60 Unowned: the R7 sweep guessed with another workflow's columns — its "do not guess" guard was unreachable dead code (#2600)
## Unowned: the R7 sweep's "do not guess a column" guard could not fire

Picked up from my own #2543 finding. Independent of my other PRs.

### The guard existed in comment form only

`reconcileUndeclaredTaskColumns` wraps IR resolution in a try/catch
whose comment reads:

> An unresolvable workflow is its own fault path; do not guess a column.

But `resolveWorkflowIrById` catches **every** failure and returns
`defaultCodingWorkflowIr()`, and `resolveWorkflowIrForTask` does the
same for a failed selection read. The resolver never rejects, so that
catch is **dead code**.

What actually happened to a card whose workflow could not be loaded: it
was judged against the **default** workflow, and if its column was not
one the default declares, the sweep re-homed it to the **default's**
rebound target. It guessed, using a workflow that is not the card's own
— the precise outcome the guard was written to prevent, in a **startup
recovery path that runs against every task**.

### How it was found, which is the part worth keeping

By being **unable to make a test of the guard fail**. Three separate
mutations all passed — deleting the `continue`, deleting the try/catch,
and simulating a whole-sweep abort at that very catch. I had written
that off once as "this case pins the outcome, not the mechanism". The
inability was the signal, not a limitation of the assertion: the branch
is unreachable.

This is the seventh instance of the program's core shape, and the first
I found in a guard I had just finished writing coverage for.

### The fix

The sweep now **proves the resolved IR belongs to the task** before
moving its card: it reads the task's workflow selection and confirms
that id resolves to a real definition (built-in or stored).

- A task with **no** selection legitimately resolves to the default
workflow — not treated as unresolvable.
- An unreadable selection **read** is itself grounds not to guess.

Placed at the **move site**, not at resolution, deliberately: it costs
one definition read only for a card already about to be moved — a
healthy board reaches that line for nobody — and it keeps the fix inside
the sweep instead of changing a resolver whose soft-failure many other
callers depend on. Changing `resolveWorkflowIrById` to reject would have
been the tidier-looking fix and a much wider blast radius.

### Revert-proof, both directions

- Remove the proof → the case fails `expected 2 to be 1`: the unloadable
card is re-homed on a guess.
- The same case asserts the neighbour **is** still repaired, so the fix
cannot be mistaken for letting one bad card disable the sweep for
everyone else. That is the per-task isolation property, and a
single-task fixture cannot distinguish it from a whole-sweep abort —
verified by injecting a throw at the loop head (`expected 0 to be 2`).

### Verification

`pnpm test:gate` (482 + 10 + 71), `pnpm lint`, engine typecheck green.
Sweep suite + `legacy-tombstones`: 13 passed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Prevented startup recovery from moving cards into incorrect columns
when their workflow cannot be loaded or resolved.
* Cards with unreadable workflow information now remain in place, while
other recoverable cards continue to be repaired correctly.
* Added safeguards to avoid guessing a fallback workflow during column
reconciliation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:21:04 -07:00
gsxdsm
534798dea0 test(U9): E2E evidence for the merge safeguards on a real PG store (completion-bar item 3) (#2615)
**U9 E2E evidence.** One new `.pg.test.ts`, 6 tests, green. No
production changes. `pnpm test:gate` green — `pgDescribe`-skipped
without PostgreSQL, so the gate is unaffected.

## What this closes

The U9 safeguard baseline verified all six merge safeguards by
**mutation at unit level**. The sibling `workflow-merge-family-live-e2e`
covers exactly **one** end-to-end. This drives
`finalizeProvenAutoMergeTask` — the last move a card makes — against a
real PostgreSQL `TaskStore`, asserting on the **persisted column** read
back after clearing the task cache. Never on "a function was called".

Only the merge **proof** is seeded (`mergeDetails.mergeConfirmed`),
which is what a real merger writes; there's no git and none is needed.
Column resolution, blocker evaluation, the move and its guards, and
persistence are all real. Includes the rename differential, where a
guard keyed on a literal goes silent.

## Three things I expected and measured wrong

Corrected in the file rather than worked around — each is a claim I
would otherwise have shipped:

**1. Dependency gating does not reach this seam.** My first draft
asserted a refusal. A proven-merged card with a live `blockedBy`
finalizes to the complete column anyway. That's coherent: dependency
gating lives in `getTaskCompletionBlocker` and gates whether work may be
*called* complete, while this seam runs after `mergeConfirmed` —
refusing would strand a merged card in review and misreport the
repository without un-merging anything. Now pinned as designed behavior
*with* that reasoning, not filed as a hole.

**2. The at-most-once outcome is `already-done`**, not the
`already-complete` I guessed.

**3. `expect(outcome).toBe("blocked")` cannot attribute a refusal.** The
finalizer has **three layered refusal gates**, and the two proof gates
emit the *same* reason (`missing-merge-confirmation`, also returned by
`validateWorkflowDoneMergeProof`). So removing either one left my
original assertion **green**:

| Mutation | Result |
|---|---|
| remove the durable-proof gate | 6 passed — invisible |
| remove the main-path proof gate | 6 passed — invisible |
| remove **both** | **2 failed** / 4 passed |

Fixed by pinning the **reason**, not just the refusal. The lesson
generalises: single-gate mutation cannot detect redundant
defense-in-depth from outside, so the unit-level attribution in the
baseline doc and this E2E are **complementary**, not duplicative. I
nearly labelled these tests as proving a specific gate they don't.

## Flagged, not changed — safeguard 1 at this seam

Written as open questions and answered by running them. **Both a
`paused` and a `userPaused` proven-merged card are moved to the complete
column.**

For `paused` that's documented design — `auto-merge-finalization.ts:243`
evaluates hard blockers with `paused: false` because the branch already
landed.

For `userPaused` it sits against the invariant re-ratified in #2486:
*never MUTATE lifecycle state of a user-paused card.* The mitigating
argument is the same one — the merge is durable, so the move is
bookkeeping that reflects reality, and refusing would leave an
operator's card permanently misfiled in review.

**Either reading may be right. What was not acceptable is that it was
untested.** Both are now explicit named assertions with the tension in
the comment, so tightening the pause contract becomes a decision rather
than a discovery. Resolution belongs to whoever owns the pause contract
— I'm not quietly changing merge behavior on a paused card.

## Safeguard coverage after this PR

| # | Safeguard | Unit (mutation) | E2E |
|---|---|---|---|
| 1 | user pause | ✅ | ✅ pinned as an exception at this seam — flagged
above |
| 2 | autoMerge:false | ✅ | ✗ gate lives upstream in `project-engine`,
not this seam |
| 3 | dependency gating | ✅ | ✅ pinned as *not* applying here, with
rationale |
| 4 | capacity single-flight | ✅ | ✗ in-memory pump, no store seam to
observe |
| 5 | merge-proof | ✅ | ✅ both vocabularies, reason-attributed |
| 6 | at-most-once | ✅ | ✅ second finalize classifies `already-done`, no
second move |

The two gaps are stated rather than implied: safeguard 2's gate is
`allowInReviewMergeProcessing` in `project-engine`, which needs an
engine harness rather than a store one, and safeguard 4 is an in-memory
single-flight latch with nothing persisted to assert on. Both are
covered by mutation at unit level and both are in the gate as of
#2526/#2569.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:20:17 -07:00
gsxdsm
592fd5c0c6 U11 [mission-feature-sync + spec-staleness]: convert the last two planner-lane guards (48 -> 46) (#2610)
**Taking: `engine/mission-feature-sync.ts`, `engine/spec-staleness.ts`**
— the last two planner-lane guards in my area.

## Census (comment-stripped, `=== "triage"` / `!== "triage"` in
`packages/*/src`, tests excluded)

| file | before | after |
|---|---:|---:|
| `packages/engine/src/mission-feature-sync.ts` | 1 | **0** |
| `packages/engine/src/spec-staleness.ts` | 1 | **0** |
| **repo total** | **48** | **46** |

## Both are real conversions, not seams

Each guard takes its vocabulary from the **caller**, which holds the
store — so unlike a defaulted parameter nothing passes, these can
actually be driven.

**`reconcileMissionFeatureState`** — a card back in a planner lane
returns the mission feature to `triaged`. Keyed on literals, a renamed
workflow left the feature reading `in-progress` forever: the roadmap
claims work is underway while the card waits to be re-planned. Nothing
errors; the rollup is just wrong. The vocabulary arrives via
`MissionFeatureSyncContext` rather than by widening this module's
deliberately narrowed `Pick<TaskStore, "getTask">`.

**`shouldSkipSpecStalenessForPreservedProgress`** — returning `false`
for a planner-lane card is what *keeps* staleness evaluation on. Miss
the lane and it falls through to the preserved-progress branch, so a
card with progress skips staleness and keeps a spec that should have
been re-validated.

## The two take different defaults — and I got it wrong first

I defaulted **both** to the `triage`/`todo` pair and broke the
pre-existing U11 proof in `spec-staleness.test.ts`, which states the
reason exactly:

> same column, different status, opposite correct answer

- **mission-feature-sync → the PAIR.** It asks "is this card waiting to
be planned?", true in either lane.
- **spec-staleness → the DEDICATED planner column only.** On a merged
lineage `todo` is *also* the hold lane, so the planner distinction there
is carried by **status** (`planning` / `needs-replan`), not by the
column. Treating the merged column as a planner lane stops a parked card
with preserved progress from skipping staleness. Its default is now the
single legacy id — byte-identical to the literal it replaced.

That asymmetry is now pinned by its own test rather than left for the
next reader to rediscover.

## Findings on the remaining census, from measuring it

Two of the 46 are **not lifecycle-column guards** and converting them
would be wrong:

- `tool-availability.ts:32` — `surface === "triage"` where `surface:
"triage" | "executor"` is an **agent lane**, not a column.
- `skill-resolver.ts:432` — `sessionPurpose === "triage"`, a **session
purpose**.

Also worth noting for the count: `replan-target.ts` reads as 2 in a raw
grep but is **0** — both hits are inside comments. `board-workflows.ts`
(2) and `archive-planning.ts` (1) are likewise comment-only. A raw grep
says 52; comment-stripped says 46.

## Not wired at the call sites yet

`scheduler.ts` / `mission-autopilot.ts` (mission sync) and `executor.ts`
/ `scheduler.ts` (staleness) still omit the new option, so behaviour is
byte-identical today. Deliberate: `executor.ts` belongs to u8's active
slice and I would rather not create a textual collision for a
pass-through. The seam is proven by tests and the count is real; wiring
is a follow-up.

## Verification

- **Mutation-verified:** restoring either literal fails a test
- 35 tests green across the three suites, merge gate green (482 + 132 +
10), tsc clean, lint clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-29 21:20:09 -07:00
gsxdsm
45e8b5f7ac U8: pin the completion-finalize ordering invariant before moving the last out-of-band exit (#2599)
Groundwork for moving `paused-after-completion`, the **last**
out-of-band exit. Stacked on #2590.

## What lands

1. **An indentation defect I introduced.** My bulk edit when the exit
vocabulary landed left the second `paused-after-completion` site
mis-indented inside a `finally` block. Cosmetic, but misleading
indentation in a `finally` is how a future reader misjudges scope.

2. **The adjacency ratchet now requires `markCompletionFinalized` before
the handoff, at every reporting site.** It previously checked only the
first occurrence, and only for the handoff itself.

That ordering is the invariant `handleGraphFailure` depends on and
**cannot check for itself**: `alreadyFinalizedToReview` /
`completionFinalized` exist to recognise this out-of-band move when a
later teardown re-marks the abort as `hard-cancel`. Without the durable
marker set first, a completed no-commit task is re-parked `failed` —
FN-6644/FN-6641.

It is asserted **structurally, and labelled as such in the test**. Both
call sites sit in pause and `finally` paths that cannot be driven
without mocking an entire agent session; presenting a source assertion
as behavioural coverage would repeat the overclaim I have been correctly
pulled up on twice in this unit.

Red-green: removing `markCompletionFinalized` from either site fails the
ratchet.

## Why the move itself is not in this PR

`paused-after-completion` is structurally harder than the pending-review
ending that #2590 moved, and the difference is worth recording before
someone assumes it is a copy-paste:

- it does **four** things, not one — `markCompletionFinalized`,
`handoffTaskToReview`,
`clearCompletedTaskWatchdog`/`signalTaskComplete`. Only the handoff is
lifecycle; the rest is substrate that must stay put.
- one of the two sites is inside a **`finally`**. Moving a transition
out of a `finally` is not the same operation as moving one out of a
branch: the graph may already be unwinding, so "report and let the graph
route" needs a defined answer for a run that is already ending.
- there is **no behavioural coverage of either site today** — the
closest tests only exercise the exit vocabulary. The pending-review move
succeeded on the fourth attempt precisely because FN-5436 existed to
catch each wrong version; this exit has no equivalent, so the move needs
that floor built first, and building it means real session mocking
rather than a shortcut.

## Verification

- exit-events + primitive-exit-events + step-session + ownership ledger
— green
- `pnpm lint` clean; `tsc --noEmit` clean
- No user-facing behaviour change, so no changeset

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
  - Improved handling of workflow steps that pause for review.
- Tasks now remain in review when a review request has no subsequent
decision.
  - Added clearer completion events for primitive prompt steps.
  - Preserved correct failure handling when later workflow steps fail.

- **Workflow Improvements**
- Built-in workflows now route pending reviews through a dedicated
review handoff.
- User-authored workflows retain compatible review parking behavior when
routing is unavailable.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:04:39 -07:00
gsxdsm
3f763cba87 U8: the graph owns the pending-review park — ownership ledger 28 → 27 (#2590)
The routing move this unit has been building toward, landing on the path
the engine actually runs. **Includes #2578's commit** (the live-path fix
it depends on) — merge that first, or this supersedes it.

## What changes

Three things together, because a half-routed move is a card that
silently does not advance:

1. The **live** implementation primitive (`runCodingSession`) returns
`{outcome: "failure", value: "review-pending"}` for that ending.
2. The primitive step handler stops flattening every ending to
`step-done`/`step-failed`, so the value survives the foreach —
`runForeach` propagates a failing instance's value as the node's own —
and reaches an edge.
3. The inline `handoffTaskToReview` in `runImplementation` is
**deleted**. The phase reports and stops, which is all an implementation
phase should do.

Built-in workflows route to the `review-pending-handoff` node added in
#2519/#2546, which performs the handoff and ends the run: the same two
effects in the same order, with the graph as the owner.

## Proof, end to end

FN-5436 — the test that blocked this move twice and was right both times
— now passes, with a **stronger** assertion than it had:

```ts
expect(store.moveTask).toHaveBeenCalledWith("FN-5436-B", "in-review",
  expect.objectContaining({
    workflowMoveSource: "workflow-graph",
    workflowMoveMetadata: expect.objectContaining({ nodeId: "review-pending-handoff" }),
  }));
```

The old two-argument `moveTask(id, "in-review")` could not distinguish a
graph-owned park from an out-of-band one — which is the entire
distinction this unit exists to make. The invariant (park in review,
never `failed`) is unchanged; the owner is now proven.

## Every ratchet fired, and each records a real change

| Ratchet | Before | After | Why |
|---|---|---|---|
| Ownership ledger — `runImplementation` review handoffs | 3 | **2** |
the handoff left the phase |
| Ownership ledger — `handleGraphFailure` | 0 | **1** | the named compat
classifier |
| Ledger headline — executor-owned dispositions | 28 | **27** | first
decrement of the unit |
| Out-of-band exit list | 2 | **1** | pending-review is graph-owned now
|
| Primitive routing pin | "must not reroute" | routes *only* the moved
ending | declared, not discovered |

None was relaxed. The `handleGraphFailure` 0 → 1 is the honest one: for
a user-authored graph without the edge this is a **relocation, not an
elimination** — the transition is still executor-performed, but from one
named classifier in the failure ladder rather than a call buried two
thousand lines into a session loop. The ledger says so rather than
letting the headline number imply more progress than there is.

## Why it took four attempts

Recorded because the reason is reusable: the value was being produced on
`createAuthoritativeWorkflowSeams`, a handler that never runs (#2578).
Every earlier attempt was correct code on a dead path, and the only
thing that showed it was instrumenting until a negative result was
proven observable rather than assumed.

## Verification

- step-session + exit-events + primitive-exit-events + ownership ledger
+ graph-requeue-gate + task-done-blocked — **83 tests green**
- `pnpm test:gate` green (10 / 482 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved handling of tasks awaiting review so they are correctly
routed to the review workflow.
* Tasks now remain in review instead of being marked as failed when no
follow-up review route is configured.
* Review handoffs now include workflow ownership and provenance details.
* Preserved standard failure handling for tasks that are not awaiting
review.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:59:06 -07:00
gsxdsm
d5f1ce7abd U11 [writes]: stop CREATING cards into a column the workflow no longer declares (9 -> 0, engine+cli) (#2603)
**Taking: `engine/triage.ts`, `engine/pr-comment-handler.ts`,
`engine/eval-followups.ts`, `cli/commands/task.ts`, `cli/extension.ts`**
(write class — no collision with the comparison backlog).

## A class the census does not count

The 48-guard work list tracks `=== "triage"` **comparisons**. These are
`column: "triage"` **writes** — and post-#2515 every one creates a card
directly into the state STALL 3 was about, except **manufactured
continuously** rather than left behind by the upgrade.

## Why they bite

`createTaskImpl` resolves the column as:

```ts
column: input.column || options?.resolvedEntryColumn || fallbackIntakeColumn || "triage"
```

`input.column` **wins**, so an explicit `column: "triage"` overrides the
workflow's resolved intake column entirely.
`store-create-intake-column.test.ts` already pins that a create with
**no** column lands in the default workflow's intake (now `todo`) —
these callers opted out of it.

The sharpest is `triage.ts`'s `fn_task_create` agent tool: it passed
`workflowId: params.workflow_id` **and** `column: "triage"` in the same
call. The caller chose a workflow and the column ignored it — a Coding
(Ideas) create landed in `triage` instead of `ideas`.

## Counts

**Comparison guards: unchanged by this PR.** This is the write class;
conflating the two would misreport convergence toward the zero bar.

| file | `column: "triage"` writes before | after |
|---|---:|---:|
| `packages/engine/src/triage.ts` | 1 | **0** |
| `packages/engine/src/pr-comment-handler.ts` | 1 | **0** |
| `packages/engine/src/eval-followups.ts` | 1 | **0** |
| `packages/cli/src/commands/task.ts` | 3 | **0** |
| `packages/cli/src/extension.ts` | 3 | **0** |
| **total** | **9** | **0** |

## A test that pinned the defect

`pr-comment-handler.test.ts` asserted `column: "triage"` in the
createTask call — so it would have **failed the fix and passed the
bug**. Rewritten to assert the invariant (the caller passes no column,
so the workflow's intake wins) plus an explicit `Object.hasOwn(arg,
"column") === false`, which is what actually catches a reintroduction.

## Interaction with #2591

My merged #2591 rescues these cards once created — they sit on a legacy
planner id their workflow doesn't declare and are still in planning
stage. So this isn't a *visible* stall today; the rescue absorbs it.
**That's the reason to fix it rather than leave it:** a self-healing
path silently absorbing a steady stream of malformed creates is exactly
how the underlying defect stays invisible.

## Deliberately not touched

- `{ id: "start", kind: "start", column: "triage" }` in the builtin
coding / PR / lead-generation IRs — workflow-internal **node
declarations** for workflows that still legitimately declare a `triage`
column, not lifecycle writes.
- Left for their owners: `core/task-store/project-store-ops.ts:210`,
`core/task-store/update-task-deps.ts:111` (main worker),
`dashboard/src/routes/register-gitlab.ts:108` (u12). Same defect, same
one-line shape.

## Verification

- 304 engine/CLI tests green across the affected suites
- merge gate green (482 + 132 + 10), engine + CLI tsc clean, lint clean

No changeset: `@fusion/engine` and `@fusion/core` are private; the CLI
change is a bug fix with no user-facing API change — happy to add one if
you'd rather it appear in release notes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-29 20:58:52 -07:00
gsxdsm
9c1c6f7479 docs(engine): close the unproven-sites ledger — one entry was wrong, the rest need two named lanes (#2544)
Comment-only change to the ledger. No test or production code moves.

## Why this is a PR and not a note

The ledger is the artifact that keeps *"the E2E covers the conversion"*
honest. It gets the same treatment as the code: claims verified by
mutation, not by reading.

## Correction: one entry was wrong

`core/task-store/reads.ts` was listed as **unproven**. It isn't. Core's
`store-stale-paused-renamed-hold.pg.test.ts` is a real-store test that
drives `listTasks` against a renamed hold column — and forcing the
hydration back to the `todo` literal **fails exactly that file's renamed
case**.

I had listed it as unproven because I assumed a separate E2E was needed.
Verified *before* removing it, since "already covered somewhere else" is
precisely the assumption that lets a gap hide.

## What remains, and why it is not another table row

**Lane 1 — real git.** `merger.ts`'s `resolveMergerLifecycleColumn` and
`executor.ts`'s `resolveReboundColumnFor` are module-private helpers
whose only callers sit inside merge/session machinery needing a real
worktree, branch and squash; `merger-ai.ts` is the same. Re-checked with
the lens that freed `auto-merge-finalization` and both self-healing
rebounds — **these genuinely need the lane.** The earlier over-broad
claim doesn't retroactively excuse them.

**Lane 2 — dashboard HTTP.** The four `register-task-workflow-routes`
sites sit behind `registerTaskWorkflowRoutes(ctx, deps)`, needing a full
`ApiRoutesContext` plus twelve injected deps. Standing that up is the
mock-the-world shell FN-5048 says not to add. The narrower alternative —
exporting the two private resolvers — yields **unit** evidence while
looking like E2E.

Deliberately not done rather than done badly and overclaimed.
`live-agent-count`'s `columnIsIntakeOrHold` is the same lane: only the
*waiting* predicate reads it, and its consumers are dashboard-side.

## Running total

**10 of 15 census sites proven end to end** across six suites; 5 remain,
each named with the lane it needs.

## Verification

- lifecycle suite 20/20; engine `tsc --noEmit` clean; `pnpm test:gate`
green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Updated end-to-end test coverage documentation to accurately reflect
verified workflow and task-store behavior.
* Clarified coverage gaps for live agent-count logic and dashboard
workflow routes.
  * Added a two-lane breakdown describing remaining coverage work.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:56:04 -07:00
gsxdsm
c9df4b9dee U11 migration proof: the path an operator actually hits, with all three caveats answered (#2597)
Tests only. Proves the upgrade path the existing E2E does not cover, and
answers the three caveats.

## Why the existing coverage was not enough

The existing cases strand a card in a synthetic
`a-column-no-workflow-declares` on a **fixture** vocabulary. The real
upgrade leaves cards in **`triage`**, on the **real `builtin:coding`**
workflow.

That difference is the whole point: `triage` is still a legal `ColumnId`
and is still declared by legacy-coding, Ideas and every linear built-in
(R11), so nothing rejects it and **nothing throws**. The card simply
sits in a column its *own* workflow no longer declares — where it
carries no trait flags and is invisible to every trait-driven sweep.

## What is proven, on a temp PostgreSQL project

- A card left in the deleted `triage` column on a default-workflow board
is re-homed to `todo`, the merged Planning column.
- **Revert check in-suite:** without the sweep running, the card stays
in `triage`. Without this, the case above could pass because some
*other* sweep or a store-open reconcile moved the card — and would keep
passing if the sweep were deleted outright.
- **Progress and the plan artifact survive.** `preserveProgress: true`
is asserted end-to-end rather than trusted from the option name.
- A `userPaused` card is skipped and stays in the deleted column.

**Mutation-verified:** stubbing `reconcileUndeclaredTaskColumns` to
`return 0` turns **5 of 11** tests red, including all three positive
migration cases. The sweep is demonstrably the mover.

## The three caveats — answered

**1. `userPaused` cards are skipped → caveat, not a stall.**
An operator park is authoritative and the sweep must not override it, so
the card does stay in a column its workflow no longer declares. But it
is reachable two ways: unpausing makes the next sweep re-home it, and
**U11's undeclared-source escape hatch in `resolveAllowedColumns`
(merged with #2515) lets an operator move it by hand meanwhile** — that
path returns the workflow's rebound target instead of `Valid targets:
none`. Recorded as a test so the behaviour is a decision rather than an
accident.

My recommendation: **leave it skipped.** Re-homing a paused card
silently moves work an operator deliberately froze, and the escape hatch
already gives them a way out. Overriding a park to fix a column is the
wrong trade.

**2. Sweep only runs when self-healing is enabled → caveat, not a stall,
for the same reason.**
The escape hatch lives in the **move-validation** path, not in
self-healing, so it works with self-healing off entirely. A card
stranded that way is draggable out of the deleted column by hand. Worth
knowing: before #2515's escape hatch this *would* have been a hard stall
— `resolveAllowedColumns` returned `[]` for an undeclared source, so the
card could not be moved anywhere at all, by anyone.

**3. Re-home targets the HOLD column → correct, and progress survives.**
Under U11 the hold column **is** the Planning column, so "everything
lands in Planning" is the intended destination rather than a compromise.
Asserted with real step progress on the row.

**None of the three is worse than a caveat.** The reason all three are
survivable is the same single mechanism — the undeclared-source escape
hatch — which is worth knowing because removing it would silently
promote all three to hard stalls.

## Incidental

Fixed two fixture-level PostgreSQL column-name errors found while
writing this: `currentStep` → `current_step`, and `user_paused` is an
**integer** flag rather than a boolean. Both would have made a future
test here fail confusingly.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:31:15 -07:00
gsxdsm
131feb243c U8: the exit announcement was on a dead code path — move it to the handler the engine actually runs (#2578)
A merged behavior of mine has never executed. This fixes it and adds the
ratchet that would have caught it.

## The finding

`createDefaultNodeHandlers` chooses the prompt-node handler like this:

```ts
const promptLike = deps?.primitives
  ? createPrimitivePromptLikeHandler(deps.primitives, runCustomNode)
  : createPromptLikeHandler(seams, runCustomNode);
```

`executeWorkflowGraph` always passes `primitives:
this.createAuthoritativeWorkflowPrimitives(settings)`
(`executor.ts:6051`). **So `createPromptLikeHandler` — and with it every
`execute` / `step-execute` function in
`createAuthoritativeWorkflowSeams` — is unreachable for prompt nodes.**
Both objects are passed to the graph executor and only one is consulted.

The `NodeCompleted.exit` announcement added in #2507 was wired into that
seam. It type-checks, its tests pass (they call the seam object
directly), and it has never run in production. `runCodingSession` in the
primitives is the live twin, and that is where it emits now.

## How it was found — and why the negative is trustworthy

Instrumenting `createAuthoritativeWorkflowSeams.stepExecute` produced no
output for a run that demonstrably visits `steps#0:step-execute`. So did
instrumenting `createPromptLikeHandler`'s dispatch. A negative result
from instrumentation is worthless until the instrumentation is shown to
be observable, so: a `process.stderr.write` at module load of the same
file **did** appear, exactly once, in the same run. The two negatives
were real, not swallowed output.

This is also the answer to the open question I left in #2546 — the
pending-review routing move kept failing because the seam value it
depends on is never produced. **That move is still not landed here.**
This commit only relocates the announcement, so it stays small and
separately revertable; the routing move follows once its value
originates on the live path.

## The ratchet

A source assertion pins the dispatch rule: `deps?.primitives ?
createPrimitivePromptLikeHandler` and the executor's wiring of
`primitives`. Inverting or conditionalising that preference would
silently disable every behavior attached to the primitives path — the
same failure in the other direction — and **a seam-level unit test
cannot tell the two apart**, which is precisely how this survived review
twice.

## Red-green

Removing the emit fails 2 of the 4 new tests (`Tests 2 failed | 2 passed
(4)`). The other two are the regression floor: an ordinary completion
emits `success` with no `exit`, and the returned routing outcome is
unchanged — announcing must not reroute.

## Scope note

I did **not** delete the now-known-dead seam wiring in this PR.
`createAuthoritativeWorkflowSeams` is still passed to the graph executor
and its non-prompt entries (`stepReview`, `merge`) are reached through
other handlers, so deciding what is genuinely dead there is a deletion
audit of its own — and this program's rule is that deletions never ride
along with behavior changes. Filed as the next slice.

## Verification

- 4 new tests + exit-events + step-session + triage audit + ownership
ledger — **54 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:54 -07:00
gsxdsm
2aa68867e5 U11 follow-up: usage-limit parking silently stopped covering the planning lane (#2567)
Second of the 39 audited `triage` sites from #2515's safety audit.
Unlike the first, this one is **real breakage**, not a proof of safety.

## The defect

The usage-limit pauser decides which tasks are on a rate-limited
provider by asking, per lane, whether the card sits in that lane's
column. The **planning** lane asked for the literal `triage`.

Now that Todo is merged into Planning, a card being planned on the
default workflow rests in `todo`. The branch resolves to an empty
provider list, so the card is not recognised as using the planning
provider — and is **neither parked when that provider hits its limit nor
resumed when it recovers**. It runs into the limit and fails.

Silent by construction: the detector reports nothing, it simply matches
no tasks.

## The test caught my own first attempt at testing it

The initial version asserted on the task that **triggered** the
usage-limit hit — and **passed against unfixed code**, because the
trigger is always parked directly without consulting `taskUsesProvider`.
Only a **bystander** card reaches the lane/column branch.

All three assertions now use a separate trigger, and the comment says
why, because the obvious test shape is the one that proves nothing.

## Why a paired literal rather than trait resolution

`taskUsesProvider` is a synchronous predicate over a task and settings,
with no IR in scope and no call site that could supply one without a
signature change reaching several callers.

Both ids name a pre-implementation column in every built-in — `triage`
for the split shape, `todo` for the merged one and for Coding (Ideas) —
so the pair covers the planning lane in all of them. Flagged for U12's
ratchet allowlist with that reason attached.

**Over-inclusion is the safe direction and is deliberate.** On a split
workflow a `todo` card is capacity-parked rather than actively planning,
so it may now be parked during an outage it was not using. Parking one
extra idle card is recoverable; failing to park a card whose provider is
rate-limited is not.

Regression direction asserted: widening the **column** match must not
widen the **provider** match — a planning card on a different provider
is still not parked.

## Audit progress

39 exclusive `triage` sites (from
`docs/solutions/architecture-patterns/u11-triage-literal-safety-audit.md`,
merged in #2515):

| status | sites |
|---|---|
| proven safe as-is | `spec-staleness.ts` — the guard is carried by
**status**, not column; the mechanical conversion was tried and is
*wrong* |
| confirmed safe by inspection | `mission-feature-sync.ts` (already
OR-pairs), `TaskContextMenu.tsx` (already trait-paired) |
| **fixed here** | `usage-limit-detector.ts` |
| remaining | 35, with the owners named in the audit |

Gate 309/309, lint clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:47 -07:00
gsxdsm
b0b9614fd5 U12 part 10: pin the R7 undeclared-column sweep — the repair three earlier PRs cited had no test of its own (#2543)
## U12 part 10 — the R7 sweep everything else leans on was itself
unpinned

`reconcileUndeclaredTaskColumns` re-homes a card resting in a column its
workflow no longer declares. It is the shipped answer to **R7**, and it
is the reason several earlier U12 deletions were safe — I cited it when
deleting the superseded `runWorkflowColumnsIntegrityPass` (#2500), and
again when arguing that a torn workflow switch leaves *recoverable*
state (#2512).

Its only coverage was **incidental**: two live PostgreSQL e2e suites
that exercise it in passing. A repair the rest of the unit leans on had
no test of its own — a guarantee everyone cites and nobody checks, which
is the exact shape this unit keeps finding.

### Six cases

The plan names three scenarios for U12; those are the three ways this
sweep can be wrong, plus I added the over-fire direction:

- repairs the stranded card to its workflow's **own** rebound target
(not a hardcoded legacy id)
- leaves a **user-paused** card alone
- leaves an **unresolvable-workflow** card alone
- is **idempotent** — a second run does not move the card again
- ignores a card already resting in a declared column
- repairs one stranded card **without disturbing** healthy or paused
neighbours

The leave-alone cases matter more than the repair. A sweep that
over-fires rewrites an operator's board, and this one runs at startup
against every task.

It also asserts `recoveryRehome: true` explicitly, because that flag is
load-bearing rather than incidental: the stranded card's *source* column
is undeclared too, so adjacency resolves to `[]` and every target is
rejected without it. Its absence once made this sweep a repair that
never repaired anything (#2462).

### Mechanism coverage — measured, and one case that isn't

Verified by mutation rather than asserted:

| mutation | result |
|---|---|
| delete the user-pause guard | **2 cases fail** |
| delete the already-declared short-circuit | **2 cases fail** |
| delete the unresolvable-workflow `continue` | still green |

That last row is stated at the assertion rather than hidden. The
unresolvable-workflow case pins the **outcome**, not the mechanism:
every mutation I could construct — dropping the `continue`, dropping the
try/catch so the throw reaches the outer handler — also ends in "no
move". So it is a regression guard on observable behaviour, not proof
the specific guard is reached, and I am not claiming otherwise.

### A decision I made

Store double rather than PostgreSQL. The sweep's decisions are pure
functions of the task list and the resolved IR, and a double makes the
"did **not** move" assertions exact rather than inferred from an absence
of change. It also keeps the suite off the slow lane, per the standing
rule against adding slow tests.

### Verification

`pnpm test:gate` (414 + 10 + 71), `pnpm lint`, engine typecheck green.
New suite: 6 passed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added coverage for automatically restoring tasks stranded in
undeclared workflow columns.
* Verified paused tasks, unresolved workflows, and tasks already in
valid columns remain unchanged.
  * Confirmed repairs are idempotent and affect only the intended task.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:27 -07:00
gsxdsm
88c7502eae test(engine): prove the recovered-lease rebound AND its audit on a renamed board (#2539)
Test-only. Sixth E2E family. Closes the `mesh-lease-manager` ledger
entry.

## Two things to prove, and only one is where the card lands

The conversion note records the defect precisely:

> They were previously two independent `=== "todo"` comparisons that
could disagree, which is how **the audit came to claim a card landed in
`todo` when the workflow has no such column**.

1. the card rebounds to the renamed workflow's own rebound column
2. the unreachable-owner **audit** reports the column the card actually
reached

**(2) is the half that rotted silently, and it is the worse one.** The
audit is what an operator reads to find out where a recovered card went.
Confidently wrong is worse than absent — and on a renamed board it named
a column the workflow does not even declare.

## One thing deliberately NOT renamed

`decisionPath` keeps its legacy `lease-recovered-to-todo` wording. The
code explains why: it is a stable discriminator that existing queries
and dashboards match on, and renaming it would break them in order to
describe the same decision. The column actually used travels in
`newColumn`.

I've pinned that split with an explicit assertion so a future vocabulary
"cleanup" cannot quietly rename a field that **is not a column at all**.
Stating it here so it reads as a decision rather than an oversight.

## Mutation-verified

Forcing the legacy literal fails **exactly the three renamed cases**,
leaving the default-vocabulary floor and the fresh-lease negative green.

## Negative half

A lease renewed just now is not recoverable — "rebound anything with a
checkout" would tear live work off its owner.

## Fixture guard

The stale-lease seed writes lease bookkeeping through the admin client
and then **asserts the seed took effect**. A silently-dropped write
would make the recovery look correctly declined — the same trap that
produced a vacuous paused-park test earlier in this program, so it is
now guarded by default.

## Verification

- six live-E2E suites green together: **58/58**
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:19 -07:00
gsxdsm
a68785a41d P0: two silent triage guards in the executor's ownership — one strands a card with nothing to rescue it (#2572)
P0 audit of the executor's assigned `triage` sites after the
Planning-column merge. **One of them can strand a card**, so leading
with that.

## The stall — `handleDepAbortCleanup`

`executor.ts` moved a dependency-aborted task to the **literal**
`triage`. The default coding lineage no longer declares that column.

A card that gains a dependency mid-execution has its work discarded and
is then parked in a column its own workflow does not define. Nothing in
the graph routes a card out of an undeclared column. The only rescue is
`reconcileUndeclaredTaskColumns`, which runs on the **next engine
start** — so between the abort and a restart the card is stalled with no
automatic recovery. It does not throw, so it would have surfaced as a
user report, not a red test.

Fixed to `resolveReboundColumnFor`, the helper the other ~16 executor
rebounds already use.

## The silent skip — `UsageLimitPauser.taskUsesProvider`

The planning lane was identified by the same literal. For a default card
the lane resolved to **no providers**, so when a provider hit a usage
limit during a *planning* session, the fan-out that pauses peers on that
provider skipped every default-workflow card and they kept hammering the
rate-limited provider.

Not a stall: the triggering task is still paused by the explicit
fallback below the filter. What was lost is blast-radius containment. A
planning session runs while the card is pre-implementation, and the
caller has already excluded `done`/`archived`, so that is exactly "not
the implementation column and not the review column" — which matches
`todo`, `triage`, `ideas`, and a renamed planner alike.

## Full audit table for my assigned sites

| Site | (a) Still fires for a default card? | (b) What silently stops |
(c) Action |
|---|---|---|---|
| `executor.ts:16395` `moveTask(id, "triage")` | **No** — writes an
undeclared column | Card parked where nothing routes it; rescue only at
next engine start | **Fixed** — `resolveReboundColumnFor` |
| `usage-limit-detector.ts:126` `column === "triage"` | **No** |
Usage-limit fan-out skips every default card; peers keep hitting the
limited provider | **Fixed** — pre-implementation predicate |
| `executor.ts:3409` `from === "todo" \|\| from === "triage"` | **Yes**,
via the `todo` arm | — | Unchanged; `triage` arm still live for
legacy-coding |
| `executor.ts:4951` `originColumn === "todo" \|\| === "triage"` |
**Yes**, via the `todo` arm | — | Unchanged |
| `executor.ts:4963` `originColumn === "triage"` double-hop | No, and
correctly so | Nothing — the extra hop exists only for shapes that
declare `triage` | Unchanged; still required by legacy-coding |
| `executor.ts:1110` `Type.Literal("triage")` | n/a | — | **Not a
column** — an agent ROLE in `spawnAgentParams` |

Counts for my ownership: **6 sites audited, 2 defects, 2 fixed, 3
correct as-is, 1 false positive.**

## Red-green

Reverting each fix fails its own test:

```
Tests  2 failed | 2 passed (4)
  × dependency-abort cleanup requeues to a DECLARED column
  × usage-limit fan-out … pauses a peer card sitting in the merged Planning column (id `todo`)
```

The other two are the regression floor and pass both ways by design: a
legacy workflow that **does** declare `triage` still fans out, and an
in-progress card is still **not** swept into the planning lane (the
guard must stay narrow — "any non-wip column" would have been the easy
wrong fix).

## Verification

- New audit suite + graph-boundary + step-session + ownership ledger —
**45 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 19:03:10 -07:00
gsxdsm
beb33b5dd1 P0 STALL 3: rescue cards stranded in a column their workflow no longer declares (fixes 8 red tests on main) (#2591)
Based on `main`. **Fixes STALL 3 — and it needs no data migration.**

## The stall

#2515 removed `triage` from the default lineage while leaving the id
legal for stored rows, and shipped **no migration**. Planning discovery
resolves a card's lanes from its own workflow, and for a default card
`intake` and `hold` **both** resolve to `todo` — so a card *sitting* in
`triage` matched neither branch and was admitted by nothing.

`triage` was the default intake column before #2515, so **every existing
project has cards there.**

Nothing else rescued them. #2515's escape hatch makes an undeclared
source column resolve to the workflow's rebound target, but every path
that *uses* it (executor, agent-heartbeat, merger) is triggered by
**active work**, and a parked card has none. The card sat until an
operator dragged it by hand.

## Proof this is a real regression, not a stale test

**8 tests in `triage.test.ts` were RED on clean `origin/main`** —
verified by swapping main's `triage.ts` into this tree and re-running.
**All 8 pass with this change.** The sharpest:

```
expected "specifyTask" to be called 4 times, but got 0 times
```

Discovery was admitting zero triage cards.

## The fix

A card resting on a legacy pre-implementation id that its own workflow
no longer declares is **unowned by construction** — no lane's rules
apply to it. Admitting it to **planning** heals it through the normal
path: it gets planned, and finalize releases it to the workflow's hold
column, **re-homing the row as a side effect of ordinary work**. No
migration, no backfill, no operator action.

## The narrowing is the load-bearing part

My first version rescued **any** undeclared column, and it was wrong. A
card can also sit in a column its workflow genuinely owns while the
**selection** fails to resolve — the resolved default IR then doesn't
declare that column either. That version re-specified a parked Coding
(Ideas) `ideas` card, breaking **FN-7596's manual-intake rule** (an
ideas card is promoted by an *operator*, never auto-planned).

`triage.test.ts` caught it. The rescue is now scoped to the legacy
planner ids, so a workflow-specific column name is never second-guessed.
That distinction — healing #2515's orphans vs. overruling a workflow
about its own board — is the whole design.

## A user-pause hole this would have opened

`couldBeCandidate` screens `paused` but not `userPaused`, so a row
carrying `userPaused` alone slipped through. Harmless before (an
undeclared-column card was admitted by nothing) and **reachable the
moment admission widens**. Planning a card mutates its lifecycle state,
which the ratified safeguard forbids for a user-paused card — so the
guard is now explicit rather than inherited. Covered by a test and
mutation-verified.

## Cost

Resolution now derives roles **and** declared column ids from one
`resolveWorkflowIrForTask` call, replacing
`resolveTaskLifecycleColumns`. Same call, same `irCache`, same bounded
concurrency window — **cost unchanged**, no added read.

## Verification

- **Mutation-verified three ways**, each failing a different test:
remove the rescue; widen it back to any undeclared column; drop the
user-pause guard
- 8 previously-red-on-main tests now green
- 263 triage/scheduler tests green, merge gate green (482 + 10 + 71),
tsc clean, lint clean

## What this does NOT do

It does not re-home rows that are past the planning stage. Admission
still requires `isTaskStillInPlanningStage`, so a card that advanced
past planning in an undeclared column stays with self-healing's
advanced-recovery sweep rather than being re-specified here. If such
rows exist and are also stranded, that is a separate sweep and a
separate PR.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-29 19:02:55 -07:00
gsxdsm
cf7b1a3d46 Drift review (unowned): gridlock detection + autopilot retries resolve the hold column — main 103→101 (#2561)
> **Based on `main`, not on my U7 stack** — merges in any order, no
dependency on #2517.

My assigned files (`triage.ts`, `replan-target.ts`) are at zero, so this
picks up two lifecycle-column literals **no unit's file list claims**.
Both ask *"is this card in the hold column?"* by the id `todo`, and both
are broken **today** for any workflow that renamed it.

## gridlock-detector — the worse of the two

`column !== "todo"` decides which cards count as **schedulable**, and an
empty schedulable set is an **early return**. On a renamed board the
detector concluded *"no gridlock"* at exactly the moment a real one
would be visible.

> A detector that goes quiet on the boards it cannot parse is worse than
one that is absent, because its silence reads as health.

**Converting only the `todo` half would have shipped a still-broken
detector**, and the test caught it. The `active` filter is equally
literal (`in-progress` / `in-review`) — and an empty active set is
*also* an early return. Two literals, one silence.

The `in-progress` half sits **outside the drift review's `todo|triage`
pattern**, which is precisely why a count-driven sweep would have left
it behind and declared the file done. Converted here rather than
deferred as out of scope. Worth flagging to the other workers: the
convergence metric is a good *tracker* but a bad *definition of done* —
an adjacent literal in the same predicate can preserve the whole bug at
a lower score.

## mission-autopilot

The retry compared against `todo` **and moved to the literal `todo`** —
so on a renamed workflow it relocated the card into a column the
workflow may not declare (R7) on **every retry**. Now resolves the hold
role; when the workflow declares none it leaves the card in place and
says so, because the error/status clear still runs, so the retry is not
lost — the card just stays in its own lane.

## Two fixture defects of my own, both caught by the tests failing
wrongly

**My first autopilot tests re-implemented the decision** and asserted on
the copy — proving only that the copy works. That is the anti-pattern
named in
`docs/solutions/store-fake-defects-that-masquerade-as-production-bugs.md`
(#2534) and in the #2527 ratchet review, and I had no excuse: the
constructor takes two stores and `handleTaskFailure` is public.
Rewritten to drive the real method.

**My first gridlock fixture failed on both vocabularies** — the detector
needs three preconditions and I supplied one. A test that fails on its
*no-regression* half is a broken fixture, not a discovered bug. The
"both halves failed" heuristic from that same doc is what flagged it.

That is eight fixture defects across this unit, every one caught by
reading *why* a test failed rather than making it pass.

## Revert proofs, each isolated to one literal

| Restored | Result |
|---|---|
| gridlock hold filter | **1 of 5 fails** (renamed case) |
| autopilot move target | **1 of 5 fails** (renamed case) |

Default-vocabulary halves pass either way — the correct signature for
conversions that change no existing behavior.

## Convergence

Measured against `origin/main` with a comment-stripped scan of `column
=== / !== "todo" | "triage"` in `packages/*/src`, excluding tests:

**103 → 101.**

(The gridlock `active` filter is a third site fixed here that this
pattern does not count.)

## Verification

| Check | Result |
|---|---|
| new suite | 5/5 |
| pre-existing gridlock + autopilot suites | 85/85, **no expectation
edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (414 + 10 + 71) |
| `pnpm check:changesets` | clean |

## Still unowned after this

`mission-feature-sync.ts` (1: a planning-lane check) and
`auto-claim-snapshot.ts` (1: `isRunnableAutoClaimCandidate`, a **pure
sync** predicate that needs the injected-lane pattern from #2551, not a
resolve). `notification-service.ts` has one more with a different
semantic — *"has progressed past"* — which needs its own thinking rather
than a mechanical swap. I will take these next unless someone claims
them.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:54:35 -07:00
gsxdsm
bbaa254dc3 test: add the missing debug to 27 logger mocks (206 → 4 failures) (#2573)
**Test-infrastructure fix.** 29 test files. No production code, no
altered assertions, no widened timeouts.

Now **3 commits** (#2584 merged into this branch): the logger-mock
sweep, a cron-runner follow-up from review, and the 4 residual failures
the sweep deliberately deferred.

**Whole branch: 764 tests, 0 failures** across the touched set.

---

## Commit 1 — the missing `debug` on 27 logger mocks

`createLogger`'s real shape is `{ log, debug, warn, error }`. 27 engine
test files mock `../logger.js` with logger-shaped literals that **omit
`debug`**, so any production path reaching `log.debug` threw:

```
TypeError: schedulerLog.debug is not a function
TypeError: runtimeLog.debug is not a function
TypeError: log.debug is not a function      (SelfHealingManager.start)
```

Measured, same commit, same 27 files:

| | Failed | Passed |
|---|---|---|
| before | **206** | 558 |
| after | **4** | 760 |

**202 failures fixed by one missing mock export.** Per-file: `notifier`
36→0, `plugin-runner` 56→0, `grok-runtime-routing` 14→0,
`self-healing-completion-fanout` 1→0. That last one also leaked an
unhandled rejection out of `startMaintenance`, which vitest warns "might
cause false positive tests" elsewhere in the file.

*A note on the number:* a full `engine-default` run went 283 → 106
across my two sessions, but `main` moved in between (U11 landed), so
that spread is **not** attributable here. 206 → 4 is the honest figure:
same commit, same file set, only this diff varying.

## Commit 2 — cron-runner's factory (greptile P1)

My regex required `log: vi.fn()`; `cron-runner.test.ts` uses `log:
cronLoggerSpies.log`, so the `createLogger` factory's returned literal
never matched and the logger production received still lacked `debug`.

**Measured before claiming a live fix, and the numbers don't support
that part:** `cronLoggerSpies.debug.mock.calls.length` is **0** across
all 155 tests, and the suite is 155 passed both before and after. The
described failure mode — `tick()` hitting `log.debug`, throwing, and
being swallowed by its own error handler — is **not reachable today**,
because no test exercises those three branches (`cron-runner.ts:377`,
`:385`, `:410`). The fix is defensive, not curative. The real gap it
surfaced is **missing coverage** for schedule dedupe / scope mismatch /
lost atomic claim, which I did not write blind to close a thread.

## Commit 3 — the 4 residuals

**`notification-service` (3):** messages moved to DEBUG in production
(`:580`, `:846`) while tests asserted `schedulerLog.log`.

The token case needed more than a relocation. It asserted
`expect(schedulerLog.log).not.toHaveBeenCalledWith(containing("new-token"))`.
Moving only the *positive* assertion to `debug` would leave the secrecy
check watching a channel the message no longer uses — a token could leak
through `debug` and the test would still pass. The negative now runs
across all four channels. **Verified it bites:** interpolating the token
into the debug line fails the test.

**`openclaw-runtime-integration` (1):** `../pi.js` mock missing
`wrapToolsWithOutputBudget` (same class as #2547); this suite exercises
a non-pi runtime, exactly where that wrapper applies.

**Not swept repo-wide, and the measurement is why.** 37 `pi.js` mocks
omit that export. Patching 30 moved the set from **11 failed to 10** —
thirty files of churn for one test. Reverted. Commit 1 earned its
27-file diff with 202 fixes; this one earned nothing, and a no-op sweep
is just future merge conflicts for other workers on this program.

---

## Why none of this is appeasement

AGENTS.md forbids making a red test pass by loosening it. This does the
opposite: the mocks were **wrong** — they claimed to stand in for
`createLogger` while missing part of its interface. Nothing was relaxed;
stubs were completed, and the one assertion I did move got **stronger**
(four channels instead of one).

## Also deliberately not done

Extending `scripts/check-mock-completeness.mjs` to catch this class.
Measured first: a naive rule over relative intra-package mocks flags
**147** factories of which **146 are green** — almost pure false
positives. The barrel heuristic works because `cliSrc` gives a tight
import surface; that doesn't transfer. A gate that noisy gets ignored,
which is worse than no gate.

## How this was found

While characterizing U9's review lane. These files were pre-existing
baseline noise under mutation runs — and that noise is exactly what made
my own safeguard baseline (#2511, corrected in #2520) report two false
verdicts. **A red suite does not merely lack coverage; it makes every
nearby measurement untrustworthy.**

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:50:46 -07:00
gsxdsm
d2ce1ba8b5 U11: resolve the scheduler's event-handler columns by trait (10 live sites, sync resolution) (#2518)
Based on `main`. Ten live `"todo"` sites in `scheduler.ts` now resolve
the column by trait.

## Four groups, converted together

They fail **independently**, and a half-conversion is indistinguishable
from a working system:

| group | sites | failure mode |
|---|---:|---|
| **Wake triggers** | 4 | **Latency** — snapshot invalidation,
mission-failure tracking, engine requeue tracking, move-to-backlog wake.
The wake doesn't fire and the card waits up to a poll interval. Exactly
why it would go unnoticed indefinitely. |
| **Parked wakes** | 2 | Latency — unpause and planning-finished, keyed
on hold OR intake. |
| **Dependency** | 3 | **Not latency.** After a blocker completes or is
soft-deleted, the query returns nothing, so the dependent is *never*
unblocked and waits on a blocker that already finished. |
| **Agent link** | 1 | `rollbackRunningAgentsForQueuedTodoTask` passes a
synthetic `{ column: "todo" }`. Wrong here **drops a running agent's
task link** — the worse direction of that safeguard. Resolved
`parkedColumns` is now passed through too, rather than letting the
helper fall back to its legacy default. |

## Resolution is synchronous, deliberately — the part worth reading

My first cut used the async resolver and made the `task:updated`
listener `async` to suit it. **That broke 5 pre-existing tests, and the
tests were right:** introducing a new `await` *before* a listener's
existing synchronous work defers everything after it to a microtask and
reorders handlers relative to a synchronous emitter.

A conversion must not change event ordering. It now uses the store's
sync IR path (`resolveTaskWorkflowIrSync`), so **no new suspension point
is introduced anywhere**.

That's the fifth time in this program a change that looked like a move
quietly altered behavior — and the first time the existing suite caught
it before review.

## Verification

- **Mutation-verified:** forcing the resolver back to the literals fails
**4 of the 6** new tests
- 110 tests green across all 8 scheduler suites (6 new)
- Fail-soft to the legacy pair: an unresolvable workflow behaves exactly
as before rather than losing the wake
- merge gate green (309 + 10 + 71), tsc clean, lint clean

## Measured

10 of my unit's 68 remaining code sites converted.

`scheduler.ts` now has **one** `"todo"` literal left in live code:
`isRunnableQueuedOverlapCandidate`, which is **exported but has no
production caller** — its only consumer was the legacy dispatcher
deleted in #2505. That's a **deletion, not a conversion**, so it is
deliberately not in this PR.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Scheduling now correctly recognizes workflow-specific hold and intake
columns, including renamed columns.
* Tasks entering a hold column reliably trigger scheduling and wake-up
behavior.
  * Dependency recovery now finds blocked tasks in renamed hold columns.
* Planning, unpausing, task completion, deletion, and requeue flows now
respect each workflow’s configured parked columns.
* Prevented unnecessary scheduling for moves between unrelated workflow
columns.

* **Tests**
* Added coverage for renamed hold-column scheduling, wake-up, and
dependency-unblocking scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 10:44:12 -07:00
gsxdsm
fb7ab6df26 test: re-green self-healing, worktree-pool and DB-corruption assertions (#2592)
**Test-only.** Three files, two commits. No production changes.

| File | Before | After |
|---|---|---|
| `self-healing-db-corruption` | 5 failed / 1 passed | **6 passed** |
| `self-healing` | 1 failed / 411 passed | **412 passed** |
| `worktree-pool` | 2 failed / 57 passed | **59 passed** |

All three are the same underlying story in different costumes: **the
assertion is watching a channel production stopped using**, or a step
that aborts before it can log at all.

## Commit 1 — the fake store was missing the health refreshers

`surfaceDbCorruption` *refreshes* health before reading the snapshot
(`FNXC:IncompletePgPorts 2026-07-26-20:45`, so PG connectivity is
re-checked instead of trusting an always-healthy sentinel). The fake
carried **neither** refresher, so the async branch fell through to
`this.store.refreshDatabaseHealth()` — undefined — and the step threw
before reaching dispatch. **Every assertion in the file was measuring
zero calls against a step that had already aborted.**

Both stubs are **no-ops on purpose.** Production ignores the refresh
return and reads `getDatabaseHealth()` immediately after, so the
snapshot mock stays the single source of truth. My first attempt
delegated them to `getDatabaseHealth`, which consumed a *second* value
per pass from the test that queues three `mockReturnValueOnce` snapshots
(one per `runMaintenance`) and broke its corruption → clear → corruption
ordering. Faithful beats convenient.

## Commit 2 — two more debug-level assertions

- **`self-healing`**: `"auto-archive: archived …"` is emitted at DEBUG
(`self-healing.ts:2747`); the test asserted `.log`. The mock already had
`debug` (from #2573), so only the target was stale.
- **`worktree-pool`**: both checkout-failure cases assert on
`console.error`, which is *correct* — `createLogger`'s `debug` writes
there. But debug is **gated on `FUSION_DEBUG`** (`logger.ts:43`), unset
under vitest, so the line was never emitted. One test is literally named
*"logs checkout -- failure at debug level"* while asserting a channel
debug could not reach.

Fixed by enabling `FUSION_DEBUG="worktree-pool"` for the suite and
deleting it in `afterEach` so the flag can't leak into sibling files.
**Deliberately not** fixed by re-pointing the assertions at another
channel — that describes whatever the code happens to do rather than the
behavior the test names.

## Verified each actually guards

A test that merely stops failing can still assert nothing, so every fix
was mutation-checked:

| Mutation | NEW failures |
|---|---|
| `surfaceDbCorruption` returns early | **5** |
| remove the auto-archive debug line | **1** — that test, only it |
| remove the checkout-failure debug line | **2** — both cases, only them
|

## Known residual, stated rather than hidden

`self-healing-db-corruption` **still exits non-zero** with 9 unhandled
`this.store.listTasks is not a function` rejections from
`openSurfacingCycle` (`self-healing.ts:7737`). These **predate this
change** — identical count before and after. The maintenance pass opens
one shared surfacing cycle up front, independent of which steps
`stubMaintenance` stubs.

I tried to clear them and backed it out, twice:
- adding `listTasks: async () => []` lets the cycle open, but then
*other* unstubbed sweeps run for real — an orphaned-planning-segment
audit fires and breaks 3 assertions expecting `recordRunAuditEvent`
never to be called;
- stubbing the four `surface-*` siblings didn't help either, because the
cycle is opened by the **pass**, not by the steps.

Making that file honestly green needs a fake complete enough for the
whole maintenance registry — a bigger change than the bug in front of
me, and one that would bury the fix above. Flagging it rather than
shipping a half-sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:43:14 -07:00
gsxdsm
21497b23db P0: approved plans never released after #2515 + triage.ts 11 -> 0 (#2549)
Rebased onto post-#2515 `main`. **This PR is the fix for a P0 stall**,
not just a conversion.

## Stall 1 — approved plans were never released

`recoverApprovedTask` opened with a bare `task.column !== "triage"`.
#2515 merged Todo into Planning on the default lineage, so every default
card now sits in `todo` and **this guard rejected all of them**. An
approved plan whose finalize was interrupted was never released, and
nothing else owns that card. Callers: `triage.ts:1296` (stuck-kill
recovery) and `in-process-runtime.ts:1460`.

`triage` stayed a legal id, so nothing threw — the guard just stopped
matching.

I verified the fix **mechanism** rather than assuming it.
`resolveLifecycleColumns` on the IR #2515 actually shipped returns:

```
{ intake: "todo", hold: "todo", wip: "in-progress", review: "in-review", complete: "done", archived: "archived" }
```

so the converted guard admits default cards. The new regression test
asserts the **return value**, because on a merged lineage the card is
already where the release would send it — "no move issued" is what
*both* the broken and the fixed code do, so only the outcome
discriminates.

**Mutation-verified:** restoring the literal `!== "triage"` fails 2 of 5
tests.

## A defect of my own, found while auditing — same shape as the P0

`clearStaleSpecifyingStatuses` is a board-wide startup sweep with no
single task to resolve lanes against, and I had resolved **both** its
queries from the default workflow. Post-#2515 that workflow's `intake`
and `hold` are the **same** column, so both queries collapsed onto
`todo` and **nothing ever swept `triage`**. A legacy or Coding (Ideas)
card holding a stale `planning` status would then occupy a planning
admission slot permanently — exactly the failure the 2026-07-04 note
above that function warns about.

Now queries the **union** of the legacy planner ids and the resolved
lanes, deduped by task id. Querying extra columns is free here: the
sweep only reads, and every row is filtered on `status === "planning"`
before anything is written.

Caught by `triage.test.ts`, **not by my own tests** — worth recording,
since it is the same collapse the P0 is about.

## Rebase note

The discovery conflict was resolved **in favour of `main`**. Main's
version is strictly better than mine: it resolves lanes with the
**async** `resolveTaskLifecycleColumns` (so it is not subject to the
sync-resolver limitation below), keeps the two admission branches
disjoint for a merged column, and bounds concurrency. My sync version
was dropped.

## Measured

| file | comparisons before | after |
|---|---:|---:|
| `packages/engine/src/triage.ts` | **11** | **0** |

## Known red, NOT from this PR

8 tests in `triage.test.ts` fail on **clean `origin/main`** — confirmed
by swapping main's `triage.ts` into this tree and re-running (same 8).
They are reporting the upgrade stall, not stale expectations: a card
*sitting* in `triage` is admitted by nothing after #2515 (`expected
"specifyTask" to be called 4 times, but got 0 times`), and #2515 shipped
no data migration re-homing those rows. Left untouched here — the fix is
a data migration, not a conversion. Reported to the coordinator
separately.

## Verification

- merge gate green (414 + 10 + 71), tsc clean, lint clean
- mutation-verified as above

## Separate finding — affects every worker

`resolveTaskWorkflowIrSync` **cannot resolve a task's selection in
production.** `getTaskWorkflowSelectionImpl` is `return undefined`
unconditionally and `getTaskWorkflowSelectionAsyncImpl` is *"always
PostgreSQL path"*, so the sync resolver **always** returns the DEFAULT
workflow IR. `moves.ts` already hit this and fixed it by going async.
Consequence for `resolvePlannerLanes` here: correct for default-lineage
cards (the default IR is exactly what comes back — which is why Stall 1
is genuinely fixed) and **inert for custom workflows**. Not papered
over; the async path is main's discovery code, and converting the
remaining event-listener sites needs the handler-reordering problem
solved first.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## P0 audit table — every `triage` site in my assigned files

(a) does it still fire for a default-workflow card after #2515? (b) if
not, what silently stops happening? (c) fix.

| site | (a) still fires? | (b) what silently stops | (c) disposition |
|---|---|---|---|
| `triage.ts:613` wake handler | **yes** | — OR-shaped (`todo \|\|
triage`), still matches | converted anyway |
| `triage.ts:651` evacuation guard | **yes** | — OR-shaped, still
matches | converted anyway |
| `triage.ts:741` stale-planning sweep | **yes** | — OR-shaped, still
matches | converted anyway |
| `triage.ts:1088` `recoverApprovedTask` | **NO** | **STALL 1** —
approved plan never released; nothing else owns the card | **fixed +
regression test + mutation-verified** |
| `triage.ts:1396` advanced-recovery discovery | **NO** | that recovery
never matches a default card | fixed by the same conversion |
| `clearStaleSpecifyingStatuses` (mine) | **NO** | **my own defect** —
both queries collapsed onto `todo`, `triage` never swept; stale
`planning` holds an admission slot forever | **fixed** (union of legacy
+ resolved lanes) |
| `replan-target.ts:177` / `:185` | **NO** | **STALL 2** — see #2552 |
fixed in #2552 |
| `spec-staleness.ts:95` | **NO** | narrow: a Planning card with null
status and `currentStep > 0` now skips staleness where it previously did
not | **recorded, not fixed** — see below |
| discovery (`isAtIntakeColumn`) | **NO** | **STALL 3** — a card
*sitting* in `triage` is admitted by nothing | **reported, not fixed** —
needs a data migration |

**Why `spec-staleness.ts:95` is not fixed here.** The guard already
returns `false` for `status === "planning"` and `needs-replan`, so an
*actively* planning card is still covered by status. The `column ===
"triage"` arm only added coverage for a planner-lane card with **no**
status — and post-merge that case is genuinely ambiguous, because `todo`
is now both the planning lane and the hold lane, so a card with progress
there may legitimately be a released card that *should* skip. Guessing
either way is a behaviour change without evidence, so I recorded it
rather than picking one.
2026-07-29 10:30:18 -07:00
gsxdsm
c92bce2f8c test: delete 2 project-engine-manager tests for the deleted cross-project cap (#2575)
**Test-only.** One file, 2 obsolete tests + 1 dead import removed. No
production change.

`project-engine-manager.test.ts` has been **red on main: 2 failed / 44
passed** → now **44 passed**.

## The failures

Both threw `TypeError: Cannot read properties of undefined (reading
'acquire')`, because both reach `(manager as any).globalSemaphore` — a
private field that no longer exists.

`project-engine-manager.ts:88` records why (`FNXC:CapacityModel
2026-07-28-20:10`, *"drop the cross-project cap"*):

> The shared cross-project semaphore, its mutable limit and the
`concurrency:changed` subscription are **DELETED**. Capacity is two
numbers per project; a machine-wide cap was a third limiter with its own
separate authority (a central-DB singleton row), and reconciling it
against the per-project gates is exactly the multi-limiter arbitration
this simplification removes.

So both tests assert residual-slot accounting on a shared pool that was
**deliberately** removed — not a regression.

## Why deleted rather than repaired

There is no shared semaphore left for them to describe. Reconstructing
one inside the test would assert a capacity model the engine no longer
has — a test that passes while describing fiction, which is worse than
the red it replaces.

Also drops the now-dead `ScopedAgentSemaphore` import (these were its
only uses). Lint does not flag unused imports here, so it would
otherwise have sat as quiet dead code.

## What I did NOT take, and why

`workflow-graph-optional-step-fix.test.ts` — the other red file adjacent
to this lane, 5 failures. Its failures are **U11 column-vocabulary
drift**: the replan rebound now resolves to `todo` where the test
expects `triage`, and one case gets a hard-cancel pause-abort log
instead of the Plan Review replan message.

That is the U11/U12 owner's semantics to settle. Picking whichever
column makes the assertion pass could silently encode the wrong
lifecycle target — and per the graph-entry contract doc, a rebound
landing in a column the workflow does not declare is precisely the
failure mode that "does not fail a test; it disables a recovery path in
production." Flagging it rather than guessing.

## Running tally of this cleanup thread

| File | Before | After |
|---|---|---|
| 27 logger mocks (#2573) | 206 failed | 4 failed |
| `merge-error-recovery` (#2559) | 10 failed | 0 |
| `reviewer` (#2547) | 2 failed | 0 |
| `project-engine-manager` (this) | 2 failed | 0 |

Every one was a test describing behavior that had moved or been deleted,
or a mock that had drifted from its real shape — none was a product
defect. That pattern is worth naming: on this repo a red non-blocking
suite has mostly meant *stale tests*, which is exactly what makes it
easy to ignore, and exactly why it silently corrupted my own safeguard
measurements in #2511.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:22:46 -07:00
gsxdsm
1c9f6c546b P0: default Planning cards read as ADVANCED after #2515 (FN-8596 stranding re-opened) + 6 red tests repaired (#2552)
Rebased onto `main` (post-#2515) and **upgraded from a conversion to a
P0 stall fix**.

## What changed since review

Greptile's P1 on this PR said the `plannerColumn` seam was unused —
*"every current production caller omits `plannerColumn`, so this default
still compares against `triage`."* That was correct, and **#2515 turned
it from an unused seam into a live stall.**

## The stall

#2515 merged Todo into Planning on the default lineage (one
pre-implementation column, id `todo`, display "Planning"). `triage`
stayed a legal id, so nothing throws — the bare `column === "triage"`
guards in `hasAdvancedPastPlanning` just **stopped matching for
default-workflow cards**.

A default card in `todo`, status cleared to null by the stale-status
sweep, carrying execution stamps from a previous pass, now returns
**ADVANCED**. So `isTaskStillInPlanningStage` is false and nine guarded
call sites refuse planning updates, finalize, delete and handoff:

`triage.ts` 3108 / 3117 / 3160 / 3438 / 3937 — `self-healing.ts` 12126 /
12448 / 12454

The file's own FNXC note at `:150` already records what that costs:

> Nobody owned the card and it sat indefinitely.

This is that same FN-8596 stranding, re-opened by the column merge.
Confirmed empirically — the rescue test fails on pre-fix code.

## The fix is an asymmetry, and that's the point

The two guards are **not the same rule**:

1. the FN-8596 **arrival-order rescue** — a stamp predating arrival in
the planner lane means replanning, not advancement
2. **"the planner column itself is never advanced"**

Rule 1 must recognise the merged Planning column. **Rule 2 must not** —
on the merged lineage `todo` is *also* the released/hold lane, so making
it blanket "not advanced" would strand the release path instead: a
released card with steps would read as still-planning and
`hasAdvancedPastPlanning(t) || releasedToTodo` would stop distinguishing
anything.

Rule 1 is already gated on the stamp predating arrival, so a released
card later claimed by execution keeps its newer stamp and still reads as
advanced.

**Closed via the default** (`mergedPlanningColumn = "todo"`) rather than
by wiring call sites — the stall closes everywhere at once, with no
call-site change and nothing to collide with another worker's slice.
Dedicated-planner workflows (Coding (Ideas), and every workflow still
declaring `triage`) are byte-identical.

## Second commit: 6 tests left RED on main by #2515

Verified pre-existing by stashing every local change and re-running —
same 6 failures on a clean branch. `resolveReplanTargetColumn` reads the
IR rather than a literal, so it **self-healed** to the correct
post-merge answer (`todo`); the expectations were the stale half.
Updated to the post-merge truth, not loosened — each still pins one
exact column.

## Audit table for this file (P0 sweep)

| site | still fires for a default card? | what silently stopped |
disposition |
|---|---|---|---|
| `replan-target.ts:177` `inPlannerLane` | **NO** | FN-8596 rescue —
planning writes no-op, card strands | **fixed** (rule 1) |
| `replan-target.ts:185` never-advanced | NO | nothing — must stay
dedicated-planner-only | **deliberately unchanged** (rule 2) |
| `resolveReplanTargetColumn` | yes (IR-driven) | — self-healed to
`todo` | tests repaired |
| its two `return "triage"` fallbacks | n/a | reachable only for
workflows declaring neither column | **recorded, not fixed** —
column-policy decision, has its own covering test |

## Verification

- **Mutation-verified both directions:** dropping the merged lane from
rule 1 fails **2** tests; wrongly extending rule 2 to the merged lane
fails **1**
- 50 replan-target tests green (7 new)
- merge gate green (414 + 10 + 71), tsc clean, lint clean

## Separate finding — affects every worker

`resolveTaskWorkflowIrSync` **cannot resolve a task's selection in
production.** `getTaskWorkflowSelectionImpl` is `return undefined`
unconditionally and `getTaskWorkflowSelectionAsyncImpl` is *"always
PostgreSQL path"*, so the sync resolver **always** returns the DEFAULT
workflow IR. `moves.ts` already hit this and fixed it by going async.
Any conversion built on the sync resolver is inert for **custom**
workflows — harmless for default cards, since the default IR is exactly
what comes back. Reported to the coordinator for the other workers; not
actionable in this PR, which uses no sync resolution.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-29 10:15:54 -07:00
gsxdsm
41031dbe2c Drift review (unowned): auto-claim candidacy resolves hold + completion roles — three literals, two opposite failures (#2565)
> **Based on `main`** — independent of my U7 stack and of #2561; merges
in any order.

Third unowned drift-review site. `isRunnableAutoClaimCandidate` is the
single source of truth for *"may an agent claim this task?"* (FN-6873),
and it carried **three** lifecycle literals that fail in **opposite
directions**.

## The two failures

**`column === "todo"` gated candidacy** on the hold role. Keyed on the
literal, a renamed workflow's candidate set was **permanently empty** —
agents were never offered its work, and nothing anywhere reported it.
Silence, not an error.

**`dependency?.column === "done" || "archived"` gated dependency
satisfaction**, and this is the more dangerous half: a dependency that
finished in a renamed **complete** column was never recognised as done,
so the dependent stayed **blocked forever**.

One makes work invisible; the other makes it permanently ineligible.
Both are silent.

## Roles resolve per task, not per pass

The non-obvious part: **a dependency may sit on a different workflow
from the claimant.** A single per-pass answer is wrong for one of them
on any mixed board — so the map is keyed by task id, and the dependency
check reads the *dependency's* roles, not the claimant's.

Asserted directly: a dependency completed in `done` (default vocabulary)
satisfying a claimant waiting in `drafting` (renamed).

## Shape

Both callers already have the store and are async, so they resolve for
real rather than taking the injected-lane fallback the *synchronous*
predicates needed (#2551). The predicate itself stays synchronous — a
resolved-roles map is passed in — because it runs inside two
`filter`/`flatMap` bodies.

Tasks absent from the map keep the legacy ids, so a partially-resolvable
board degrades to today's behavior instead of silently emptying the
candidate set.

**Type narrowing preserved.** The two callers take `Pick<TaskStore,
"listTasks">`, which is what makes them testable without a real store.
Rather than widening to the whole `TaskStore`, they now take
`Pick<TaskStore, "listTasks"> & WorkflowIrResolverStore` — the minimal
additional shape resolution needs.

## Revert proofs, isolated per literal

| Restored | Result |
|---|---|
| hold literal only | **3 of 6 fail** |
| dependency-completion literals only | **1 of 6 fails** |

The three default-vocabulary cases pass under both. Splitting the proof
matters here: it confirms the two halves are **independently**
load-bearing rather than one masking the other — a single combined
revert would have shown 3 failures and told me nothing about the
dependency half.

## Convergence

Measured on `main`, comment-stripped scan of `column === / !== "todo" |
"triage"` in `packages/*/src` excluding tests:

- this file alone: **103 → 102**
- with #2561: **103 → 100**

The `done` / `archived` literals fixed here sit outside that pattern and
are not counted — same caveat as #2561's gridlock `active` filter. Two
PRs now where the real fix is larger than the metric shows.

## Verification

| Check | Result |
|---|---|
| new suite | 6/6 |
| pre-existing auto-claim suite | 17/17, **no expectation edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (414 + 10 + 71) |
| `pnpm check:changesets` | clean |

## Remaining unowned in my area

`mission-feature-sync.ts` (1, a planning-lane check) and
`notification-service.ts` (1, *"has progressed past"* — a different
semantic needing its own thinking, not a mechanical swap). Taking those
next unless claimed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:01:14 -07:00
gsxdsm
8578a1d27d U8 PR5: thread the implementation exit to the step seam, and declare the stepwise pending-review park (inert) (#2546)
Follows **#2519** (U8 PR4). Both halves are inert — **no behavior
change** — and this removes the blocker PR4 documented.

## What was blocking

PR4 could only land its IR half because the pending-review ending could
not reach a graph edge on the **default** workflow. Three links in the
chain:

| Link | Problem |
|---|---|
| `runGraphTaskStep` | awaited the memoized implementation pass and
**discarded** its result |
| `RunTaskStepResult` / `RunSingleStep` | had nowhere to carry an exit |
| `stepExecute` seam | flattened every ending to `step-done` /
`step-failed` |

All three are fixed. The outcome stays `failure` (the step genuinely did
not complete) while the **value** now names the ending — which is what
`runForeach` propagates upward, since it returns a failing instance's
value as the foreach node's own. Every other ending keeps `step-failed`
byte-identically.

One design note: the exit is a property of the **pass**, not of a step.
A single memoized pass serves every foreach instance, so all instances
report the same ending — correct, because the ending is what stopped the
whole session.

With the value surviving, the stepwise IR declares the same
`review-handoff` park node and `steps --outcome:review-pending-->
review-pending-handoff --success--> end` edge the plain-`execute` shape
got in PR4, inherited by the final-review and Ideas variants that clone
it.

## A bug my own threading introduced, and what caught it

The first threading commit covered **one of the two** paths out of
`runProjectedGraphTaskStep`. The early-return branch carried the exit;
the main path goes through `runTaskStep` in `step-runner.ts`, which
builds its own result and dropped it — i.e. it worked on the path I
happened to read, and not on the path the default workflow actually
takes.

**FN-5436's regression test caught it, not code review.** That is the
second time this test has stood between this unit and a silent
regression, which is worth recording somewhere durable:
`executor-step-session.test.ts > FN-5436: pending-review skip on
no-fn_task_done exit` is the load-bearing test for this area.

## Why the seam flip is still not here

With the threading complete I applied the behavior half again — flip the
execute seam to return `review-pending`, delete the inline
`handoffTaskToReview`, add a named compat classifier for user-authored
graphs. **FN-5436 still failed**: the card did not reach `in-review`, so
something between the seam value and the park node is not routing under
that harness. I have not isolated whether that is the mock store's IR
resolution (it exposes no `getWorkflowDefinition`, so the run resolves
the built-in through a different path), a foreach aggregation detail, or
the park node's own seam.

I stopped rather than keep guessing, and reverted the behavior edits so
this lands green and inert. Shipping a half-routed move is exactly the
failure this unit exists to remove — a lifecycle transition that
silently does not happen. The alternative on offer was to relax
FN-5436's assertion, which would have been appeasing a test that is
telling the truth.

### What the instrumentation showed (done after opening this PR)

I ran the bounded next step rather than leaving it as a note. Two facts,
both measured:

1. **The IR is correct.** Resolving
`BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR` at runtime shows the
node and the edge survive the final-review variant's edge rewiring:

```
EDGES [{"from":"steps","to":"browser-verification","condition":"success"},
       {"from":"steps","to":"review-pending-handoff","condition":"outcome:review-pending"},
       {"from":"steps","to":"end","condition":"failure"}]
HAS NODE true
```

That matters because the variant does `template.edges = [ ... ]` (a
wholesale replacement) and filters outer edges touching `review` —
`review-pending-handoff` is not `review`, so it survives. Worth knowing
before anyone adds another node near it.

2. **The `stepExecute` seam is never invoked in that harness**, even
though the run terminates at `steps#0:step-execute` and the
implementation session demonstrably runs (`"Agent finished without
calling fn_task_done but Step 0 is blocked on pending review"` is in the
task log). A `console.log` at the seam's value computation produced no
output. So the exit is threaded correctly and the IR can route it, but
under this harness the value never originates.

3. **Nor is `createPromptLikeHandler`'s returned handler.**
Instrumenting its dispatch (`node.id` + resolved seam) produced nothing
either — so the node is not reaching the prompt-like path at all.

**Control experiment, because a negative result from instrumentation is
worthless until you prove the instrumentation is observable.** A
`process.stderr.write` at module load of the same file appears exactly
once in the same run, so writes from that module *are* captured under
this harness and the two negatives above are real, not artifacts of
swallowed output.

That narrows the remaining work to one question — what actually drives
`steps#0:step-execute` in this run, if neither the prompt-like handler
nor the `stepExecute` seam does — and rules out the IR, the foreach
propagation, the threading, and the instrumentation as suspects.

**Next step, now much narrower:** find the handler registration this run
resolves for a foreach instance node (the graph executor's handler map,
not the seam table), then flip the seam, delete the inline handoff, and
update the three ratchets that will correctly fire — PR3's routing pin,
the out-of-band adjacency check, and PR1's ownership ledger
(`runImplementation` 3 → 2; `handleGraphFailure` 0 → 1 for custom graphs
only).

## Verification

- `executor-step-session` + exit-events + ownership ledger +
graph-boundary — **56 tests green**
- `builtin-workflows` + `builtin-coding-workflow-ir` — green. The
layout-completeness contract required a layout entry for the new node in
all four stepwise-derived workflows; placed off the main line, because a
park is an exit and not a stage.
- `pnpm test:gate` green (10 / 309 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:54:39 -07:00
gsxdsm
3681a9f9a5 test(U9): re-green merge-error-recovery.test.ts (10 stale tests deleted, replacement contract covered) (#2559)
**U9, PR8.** Test-only, two commits (deletion and new coverage
deliberately separate).

`merge-error-recovery.test.ts` has been **red on main: 10 failed / 23
passed**. Now **24 passed**.

## Commit 1 — the 10 failures test a feature that no longer exists

All 10 assert that `ProjectEngine` creates recovery follow-up **tasks**
and dedupes them by parent/branch. Evidence this was deliberate, not a
regression:

- `project-engine.ts` contains **zero** `createTask` calls.
- The string the dedupe tests assert on — `"follow-up already exists"` —
exists **only in the test file**; no production code emits it.
- `project-engine.ts:4801` documents it outright
(`FNXC:AutostashRecovery 2026-07-26`): *"This used to file an automated
recovery follow-up card via the shared follow-up engine; that engine was
deleted ... So the card is replaced by a durable log entry AND an
operator comment on the parent."*

Deleted rather than repaired — there is nothing left for them to assert.

## Commit 2 — cover the contract that replaced them

The production comment is explicit that `record.label` *"must never be
dropped from the message or truncated"* — it is the handle `git stash`
recovery needs, and the parent may already be `done`, so the notice is
the only trace of real uncommitted work.

**That invariant had no working assertion.** The file was red, so every
claim it made was inert.

The new test asserts one log entry + one comment for a `live` orphan (a
`subsumed` record stays silent), and that label, short sha, detecting
task and source phase all survive into the comment, with the label in
both the log message and its detail field.

| Mutation | Result |
|---|---|
| replace the label with `(omitted)` | 1 failed / 23 passed — this test
|
| notify on non-live orphans too | `NEW-failures=1` — this test |

## A tooling bug this uncovered, which matters beyond this PR

The new test originally reported **zero** new failures under mutation
while passing normally — i.e. it looked vacuous.

It was not. A thrown assertion left the engine running, which **crashed
the vitest worker**, and a crashed run emits no parseable `FAIL` lines —
so my mutation harness parsed zero failures and printed **NOT COVERED
for a guard that had just correctly failed**.

Two fixes:
- The test stops the engine in a `finally`, so a failure reports as an
assertion instead of killing the worker.
- The harness now treats *non-zero exit with zero parsed failures* as
**INCONCLUSIVE**, never as a coverage verdict, and prints the crash
signature.

This is the **second** time a blind spot in my own tooling manufactured
a false "uncovered" result — after the `|project|` regex that matched
nothing for `@fusion/core`. Both had the same shape: the measuring
instrument reported success without checking anything, which is
precisely the defect class this program is chasing. Worth stating
plainly rather than quietly fixing.

## Why this file matters to U9

Its 10 pre-existing failures are what corrupted my own safeguard
measurements in #2511 — an absolute-count mutation run credited them to
the mutation. **A red file in the merge lane does not merely lack
coverage; it poisons the measurement of everything near it.**

## Wider context, measured

`engine-default` on clean `main` is **283 failed / 9062 passed across 28
files**. This PR clears one of those files. I did not attempt the rest:
most are outside the merge/review lane and plausibly owned by other
workers on this program. Also measured and abandoned: extending
`check:mock-completeness` to relative intra-package mocks — the naive
rule flags **147** factories of which **146 are green**, so it would be
almost pure false positives; the barrel heuristic does not transfer.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:54:04 -07:00
gsxdsm
67904f8a2c U11: merge Todo into Planning on the default lineage (+ the migration mechanism, and a measured safety audit that cuts the work list 32%) (#2515)
**Merges Todo into Planning on the operator's real default workflow.**
Held from merge pending the `triage` literal audit below — see *Gating*.

## The board change

`builtin:coding` → `BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR` →
clones `BUILTIN_STEPWISE_CODING_WORKFLOW_IR`. That IR now declares
**five** columns, and `plan`, `plan-review`, `plan-replan` and `start`
all live in the merged Planning column:

```
columns: todo="Planning", in-progress, in-review, done, archived
  start -> todo      plan       -> todo
  plan-review -> todo  plan-replan -> todo
  parse -> in-progress            (first implementation node)
```

The id stays `todo`, the display name becomes "Planning". That is the
cheaper half: `todo` was already the hold column, so every trait lookup,
task row, stored selection and the 121 `column === "todo"` guards keep
their meaning, and **no stored row needs re-homing**. Promoting `triage`
instead would have produced the same board while making those guards
workflow-*dependent* — live for Coding (Ideas), silently dead for
Coding.

`builtin:legacy-coding` keeps its six-column shape, per the operator's
decision. It exists to be the old thing.

## Entry contract, before and after each IR edit

| | result |
|---|---|
| before the default-lineage edit | **15 passed** |
| after the edit | **13 passed, 2 failed** |
| after reading both | **15 passed** |

Neither failure was routed around. One was a genuine expectation change
(two planning entry points became one); the other was my own
`mergeTodoIntoPlanning` helper throwing *"source IR is not the
split-column shape this merge transforms"* — because production **is**
the merged shape now. I **deleted** the helper rather than making it
tolerant: a transform that has silently become a no-op asserts nothing.

## The safety argument, proven not asserted

Entering at `start` is exactly what dragged cards backward in the three
earlier reverted attempts. `merged-planning-start-node-no-move.test.ts`
proves against the **real** boundary controller and **real** default IR
that entering `start` performs no move (`moveTask` is never *called*),
reaches no hold→wip capacity seam, and **still moves on a genuine
crossing** so the no-op is same-column rather than a disabled boundary.
Removing the controller's same-column short-circuit turns exactly the
two no-move tests red.

## The migration mechanism

A card can outlive its column. `resolveAllowedColumns` derives targets
from graph adjacency, and an undeclared source has none — so it returned
`[]` and **every** move was rejected with "Valid targets: none",
including the one that would rescue the card. An undeclared source now
resolves to the workflow's rebound target. Escape hatch, not relaxation:
declared columns are untouched, and it offers the rebound target *only*,
so a stranded card gets back **into** the lifecycle rather than a free
jump past review.

## A real regression this surfaced

`isDefaultWorkflowColumns` matched the legacy **six** ids as a set. The
merged default declares five, so the match stopped firing and the
default board fell through to neighbor-only adjacency, which **drops
legal moves and invents an illegal one**:

| edge | effect |
|---|---|
| `in-progress → done` | **dropped** — the mission-validation cross edge
|
| `in-review → todo` | **dropped** — review work back to planning |
| `todo/done → archived` | **dropped** — the FN-4892 direct-archival
edges |
| `done → in-review` | **invented** — a backward edge no rule allows |

Adjacency now derives from lifecycle **roles**. The load-bearing
assertion: the legacy six still reproduce `VALID_TRANSITIONS`
**verbatim**. Applied only when a workflow declares the full role set,
so custom boards keep neighbor adjacency.

## Failure accounting (core package, vs a 49-failure baseline)

| stage | failed | new |
|---|---:|---:|
| after the merge | 65 | 18 |
| after the escape hatch | 52 | 5 |
| after role-derived adjacency | 53 | 4 |

The 4 remaining are 3 `builtin-workflows` expectations encoding the
pre-merge shape and 1 create-intake expectation naming `triage` on
`builtin:coding`.

Two `schema-applier` and two `workflow-reconciliation-production-shape`
failures appeared in intermediate runs and are **not mine** — both files
pass in isolation (75/75 and 7/7). I re-ran each before attributing
them, which is why the earlier "priority" flag on the reconciliation
pair was withdrawn.

Gate: **309/309**. Lint clean.

## Gating: the `triage` audit
(`docs/solutions/architecture-patterns/u11-triage-literal-safety-audit.md`)

Program tracking cited **58** `triage` comparisons. Measured with the
same pattern:

| | count |
|---|---:|
| raw comparisons | 87 |
| inside comments | 1 |
| **not a lifecycle column at all** | **15** |
| column comparisons | 71 |
| OR-paired with `"todo"` in the same expression | 32 |
| **exclusive `triage` — the real work list** | **39** |

**15 do not compare a column.** `role === "triage"`, `surface ===
"triage"`, `sessionPurpose === "triage"`, `entry.agent === "triage"`
name the planning **agent**. Converting them would be actively wrong,
and the failure — a planning agent that can't resolve its prompt
template — would look nothing like a column bug.

**One site changes an operator-visible affordance**, which is why
per-site review beat a sweep:

`TaskCard.tsx:1927` — `taskColumnFlags?.intake === true && task.column
!== "triage"`. The literal is a **narrowing**, not a match. After the
merge a Planning card has `intake === true` and `column === "todo"`, so
the narrowing stops applying and **Start begins rendering on default
Planning cards where it previously did not.** A sweep would have
"converted" the literal and shipped the new affordance silently.

These guards do not go **dead**, they go **workflow-dependent** —
`triage` stays live for legacy-coding, Ideas, every linear built-in and
any user workflow (R11) — which is harder to detect than dead.

Work list and ownership are in the audit doc.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:39:20 -07:00
gsxdsm
69790dc3e7 test(U9): revive two permanently-red testMode guards in reviewer.test.ts (#2547)
**U9, PR7.** One test file, +15 lines, no production change.

## Two safety tests that could never pass

`reviewer.test.ts`'s `vi.mock("../pi.js")` is missing
`wrapToolsWithOutputBudget`, which `wrapCustomToolsForPluginRuntime`
(`agent-session-helpers.ts:104`) calls as the outermost tool wrapper.
Both test-mode-forcing cases therefore threw:

```
No "wrapToolsWithOutputBudget" export is defined on the "../pi.js" mock
```

They have been **permanently red on main** — dead enforcement on the
invariant that **testMode never issues real AI calls**.
`reviewer.test.ts`: 83 passed | 2 failed → **85 passed**.

Found while characterizing the reviewer lane for U9: they surfaced as
pre-existing baseline failures under an unrelated mutation run. This is
exactly why the delta harness records a baseline — under the old
absolute-count method these two would have been silently credited to
whatever mutation was running.

**Not a product bug.** testMode forcing works correctly; its guard did
not.

## Verified the revived tests actually guard something

A dead test can also be a vacuous one, so passing again is not
sufficient evidence. Mutating `isTestModeActive` in
`model-resolution.ts` to ignore `settings.testMode` fails **exactly
these two** (`NEW-failures=2`). Both assert
`expect(mockedCreateFnAgent).not.toHaveBeenCalled()` — no live agent
spawn.

## Why the existing gate didn't catch it

`pnpm check:mock-completeness` runs in the merge gate and passes. It
inspects only the `@fusion/engine` and `@fusion/dashboard` **barrels**,
under `cli/` and `dashboard/` test dirs — never a relative intra-package
mock like `"../pi.js"`. So the whole class of engine-internal mock drift
is outside it.

**Deliberately not fixed here.** Extending the checker is its own change
and I want the violation count measured before proposing it, rather than
opening a PR that turns out to touch dozens of files. That's the next
PR.

## Also observed, stated rather than buried

Mutating `useMockRuntime` in `agent-session-helpers.ts` produces **no**
failure in this file — the reviewer path routes through model resolution
instead. That downstream seam has its own coverage question which I have
not answered; flagging it rather than implying this PR closes it.

## Scope note

This was initially committed onto #2541's branch. I split it onto its
own branch so each PR stays independently revertable — #2541 is now one
commit (the FN-7720 verdict assertion) and this is one commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:32:34 -07:00
gsxdsm
c2705f292f U11: delete the dead isRunnableQueuedOverlapCandidate export (scheduler.ts now has zero live todo literals) (#2542)
Based on `main`. **Pure deletion — zero production callers.**

`isRunnableQueuedOverlapCandidate`'s only consumer was the legacy
pull-from-todo dispatcher deleted in #2505. The three remaining
references were all in tests.

This was `scheduler.ts`'s **last `"todo"` literal in live code**, so
removing it rather than converting it is what actually finishes the file
— converting a dead predicate would have added a trait lookup nothing
calls, and reported U11 progress for a site that cannot execute.

## Why deleting its tests does not lose coverage

Worth checking, because the function carried a real invariant — *"a busy
merge lane must not block unrelated dispatch"* — and its own doc comment
claims it's a shared contract with self-healing and repair paths.

Two facts settle it:

1. **The overlap logic is still live**, implemented inline inside
`runHoldReleaseSweepPass` (`activeScopes`, `overlapIgnorePaths`,
`getFilteredFileScope`). The behavior didn't die with the predicate;
only this copy of it did.
2. **`scheduler-overlap-starvation.test.ts` exercises that live path**
through `scheduler.schedule()`, including *"does not defer ready work
behind queued overlap blocked by an active lease"* — the same invariant
the deleted test asserted, against code that actually runs.

The doc comment's claim that self-healing *"must use this same
predicate"* is **stale**: no self-healing path imports it. That claim
outlived the coupling it described.

## The three test references were not equal

Treating them identically would have been wrong:

- **Two were incidental trailing assertions** in tests about other
subjects (stuck-loop exhaustion parking; transient merge-error
classification). Only the assertion line is removed — each test keeps
its real subject.
- **One test's entire subject was this function** (*"does not block
unrelated executor dispatch when merge lane is busy"*), so it goes with
it; its invariant is covered on the live path per (2).

## Measured

`scheduler.ts` **2,840 → 2,820 = −20**, and it now holds **zero `"todo"`
literals in live code**.

Combined with #2505's −929, `scheduler.ts` is down **949 lines** across
this unit — all genuine removal, not relocation.

## Verification

582 tests green across the reliability-interactions suite and four
scheduler suites; merge gate green (414 + 10 + 71); tsc clean; lint
clean.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Simplified internal task scheduling logic by removing obsolete overlap
coordination checks.
* Preserved existing task progress, parking behavior, logging, and
review handling.

* **Tests**
* Updated reliability checks to align with the streamlined scheduler
behavior.
* Continued validating transient errors, non-progress handling, and
correct task dispatch without changing the end-user experience.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 09:25:16 -07:00
gsxdsm
7fd1c7f124 P0 fix: stop reaping worktrees out from under live planners (FN-6756) (#2531)
User-reported: worktrees deleted while a planning agent was still
working in them. Small, isolated, ahead of all remaining capacity work.

## Mechanism

`clearPhantomExecutorBinding` is documented as *"the last line of
defense against pulling a worktree out from under a running agent"*. It
computed liveness from four sets — `activeSessions`,
`activeStepExecutors`, `activeWorkflowStepSessions`,
`activeCliTaskSessions` — **all TaskExecutor-owned**. A triage PLANNING
session is owned by `TriageProcessor`, lives in *its own*
`activeSessions` map, and registers in the module-level
`activeSessionRegistry`. It matched none of the four.

Worse: the method **writes** to that registry (unregistering the task’s
paths) but never **read** it as a liveness signal. It destroyed the very
evidence that proved the planner alive.

Under plan-in-place a card is specified while it sits in
`todo`/`triage`, and `reapLeakedConcurrencySlots` treats both as
reapable on a rationale written *before* planning moved there (“a task
waiting to run must not pin a worktree”). Every gate ahead of the last
one passes for a planner:

| Gate | Saves a planner? |
|---|---|
| in `listWorktreeHolders()`? | **No** — `ensureTaskWorktreeForPlanning`
→ `ensureGraphCustomNodeWorktree` → `addActiveWorktree`
(`executor.ts:8581`) |
| reapable column? | **No** — plan-in-place keeps the card in
`todo`/`triage` |
| in the executor’s `executing` set? | **No** — a planner is
triage-owned |
| 60 s `LEAKED_WORKTREE_SLOT_GRACE_MS` | **No** — keyed on
`columnMovedAt`, and planning routinely runs for minutes |

So the broken guard decided alone.

## This is FN-8600 recurring through a second sweep

That fix registered planning paths in the registry and taught the
**self-owned-branch reclaim** sweep to consult `isPathActive`. The
leaked-slot reaper never got the same signal — fixed at one surface, not
enumerated across all. Exactly what the AGENTS.md Surface Enumeration
rule exists to prevent.

## Fix

The refusal now also fires when
`activeSessionRegistry.pathsForTask(taskId)` is non-empty. Keyed on
**any** registered path rather than on kind: the point is that a
registered surface of any kind means someone is working in that
worktree.

## Enumeration — the part that stops a third recurrence

The guard is a **chokepoint**, so this covers every caller rather than
just the reported one:

- `reapLeakedConcurrencySlots` — the reported path
- `recoverPausedAbortFailures` — **had the identical executor-only
pre-gate**
- the `preserveWorktrees: true` reclaim

Audited the rest of self-healing’s liveness gates: the self-owned-branch
reclaim, worktree-metadata reconcile and PR-branch sweeps already
consult `isPathActive`/`lookupByPath`. The three that read only
`getExecutingTaskIds` — `checkStuckBudget`, `recoverCompletedTasks`,
`recoverStrandedCompletedTodoTasks` — move columns and never destroy a
worktree, so they are noted rather than changed.

## Trade-off, stated plainly

A leaked registry entry now blocks this sweep instead of a live planner
losing its worktree. That is the strictly safer failure and the one the
“last line of defense” wording already promises. The registry is
process-local and in-memory, so a leak cannot outlive the process, and
stale entries have their own reconciler. **A test pins that a genuine
phantom — no executor surface AND no registration — still clears**, so
this is not a blanket refusal that would trade this bug for a wedged
queue.

**The 60 s grace is deliberately unchanged.** Raising it would only make
the bug rarer and harder to reproduce; the liveness gate was the defect.

## Verification

Revert-proof, measured: removing the registry term turns **3 of the 4**
new tests red, including the end-to-end sweep case (card in `triage`,
past the grace, executor sets empty → asserts the slot is not reaped and
the worktree survives). The 4th stays green both ways *by design* — it
is the anti-overcorrection guard.

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (309 +
10 + 71) · new suite 4/4. The 2 failures in `self-healing.test.ts` /
`-completion-fanout.test.ts` are **pre-existing** — identical with this
change stashed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Prevented active planning worktrees from being mistakenly deleted or
reclaimed while related planning sessions are still active.
* Enhanced session liveness checks so phantom executor bindings are not
cleared when a live session is registered.
* Updated paused abort recovery to defer or abort safely when a live
planning session is detected, avoiding unintended task/worktree
mutations.
* **Tests**
* Added regression coverage for leaked-slot reaping, paused abort
recovery behavior, phantom binding refusal, and end-to-end sweep
outcomes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:04:51 -07:00
Phil Larson
72391c90b2 fix(engine): route workflow reviews through validator models (#2533)
## Summary

- classify review-type workflow steps with the existing review-step
classifier
- resolve their primary, fallback, and thinking-level settings from the
validator model lane
- retain per-step model overrides and executor-purpose workflow-step
tooling
- keep ordinary workflow steps on the execution lane
- make missing-fallback diagnostics identify the correct lane

## Why

Code Review, Plan Review, verification, and inline-review gates were
executed through the implementation model lane merely because they run
inside `executeWorkflowStep()`. That defeats configured reviewer-model
separation and can make the same model implement and validate its own
work.

This changes model selection—not the workflow-step session/tooling
contract—so review steps remain executor-purpose sessions while using
validator lane models.

## Verification

- `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/engine
exec vitest run src/__tests__/executor-workflow-step-model.test.ts` — 14
passed
- `corepack pnpm@10.33.0 --filter @fusion/engine typecheck`
- `corepack pnpm@10.33.0 changeset status --since=origin/main`
- `git diff --check origin/main...HEAD`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Review-type workflow steps now route through the configured validator
model lane (instead of the execution lane).
* Validator primary/fallback and thinking-level settings are applied
correctly for review steps.
  * Step/task overrides still take priority over lane-based resolution.
* Fallback retry sessions now use the appropriate validator/executor
configuration, with lane-specific fallback guidance when fallback
settings are missing.
* **Tests**
* Expanded executor workflow-step model resolution and routing/fallback
precedence assertions for validator-lane behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:05:04 -07:00
gsxdsm
18d654a5ff capacity, part 3: delete the globalMaxConcurrent setting, API and UI (#2529)
Part 3 of the capacity simplification, and the half that removes the
**knob**. Enforcement (shared semaphore, runtime wiring) went in #2509;
this removes everything an operator or API client can still see, so
nothing is left readable-but-ignored.

## Deleted

Settings key + schema default · CentralCore’s
`getGlobalConcurrencyState` / `updateGlobalConcurrency` /
`acquireGlobalSlot` / `releaseGlobalSlot` and the `concurrency:changed`
event · the whole Global Concurrency block in `async-central-core` ·
`PUT /api/global-concurrency` · the Scheduling · Global settings section
· the footer and Command Center global sliders · the dead
`getGlobalConcurrencyLimit` reader whose only caller went in #2509.

## Kept, deliberately

**`GET /api/global-concurrency` survives as telemetry only** — live
`currentlyActive` / `projectsActive` from CentralCore’s side-effect-safe
source. “How busy is this machine?” is still a real question once the
cap that used to answer it is gone. It no longer reports
`globalMaxConcurrent`/`queuedCount`: those came from the deleted cap and
from slot bookkeeping production code never incremented, so publishing
them was publishing zeros dressed as state.

**`useGlobalConcurrency` becomes read-only.** Everything that existed to
*persist* went with the cap — the 500 ms debounce, the save-state
machine, the commit-on-close/unmount flush, the slider clamp, the
`interactive` gate. The module-level shared store is **kept**: its
original justification (two mounted consumers drift apart with private
copies) holds for a polled read exactly as it did for a cap, and one
fetch now serves both.

The live “N running (all projects)” readout survives in both surfaces,
moved onto the per-project row.

## Two sections become one

Scheduling · Global existed to host exactly one control. With it deleted
the section renders an empty pane, so the Global/Project pair merges
back into **“Scheduling”**. An empty nav entry is a promise of settings
that are not there.

## One real fix found on the way

`SchedulingSection`’s `concurrencyLoading` gated the **project**
concurrency inputs on the **global**-concurrency fetch — never the right
source, since `maxConcurrent` and `maxWorktrees` come from the settings
form. It is repointed at the form’s own load, preserving the invariant
it existed for: a concurrency input stays disabled until its live value
arrives, so an operator cannot overwrite a resolved limit with a blank
fallback.

## Migration

A stored `globalMaxConcurrent` is **ignored** — it is a project-blob key
nothing reads, so dropping it needs no schema change. The
`central.global_concurrency` **table** is dropped in a follow-up; this
slice stops seeding and reading it first, so that drop has no live
writer to race.

## Verification, and how the wider suite was controlled

`pnpm lint` clean · core/engine/dashboard `tsc` clean · `pnpm test:gate`
green (309 + 10 + 71) · dashboard settings/footer/command-center/hooks
**2237/2237** · core `central-core-backend` 9/9.

The broader dashboard suite shows failures, and I checked rather than
assumed: running the suspect files on **clean main** reproduces
`api-git` (49), `TaskDetailModal.rendering` (28) and `settings-mobile`
(17) identically. Two were genuinely mine —
`SettingsModal.scheduling-merge` (0 on main, 17 on this branch: my nav
rename) and one `settings-mobile` picker case asserting `scheduling` is
a scoped pair — and both are fixed.

Tests for deleted behaviour are removed with it (footer
confirm/cancel/flush/dedupe, global marker geometry, the hook’s PUT
case, the CentralCore slot cases), each carrying a note on what it
guarded and where the surviving **project-side** equivalent lives.
Fixture-only references were updated, not deleted.

Nothing booted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:13 -07:00
gsxdsm
919f68f9bc test(U9): cover the two unguarded merge safeguards and admit them to the gate (#2526)
**U9, PR5.** Closes the gap #2520 measured. Tests + gate config only; no
production behavior change.

## The gap

#2520 found that safeguards **1 (user pause)** and **4 (capacity
single-flight)** had **zero test coverage**. Deleting either guard
produced no new failure anywhere in the merge, project-engine,
self-healing, or concurrency suites. Both guards work correctly today —
nothing would have noticed if they stopped. U9 moves merge behind graph
nodes, so this is exactly the state not to convert on top of.

## Two tests

- **`merge admission excludes a user-paused card`** — safeguard 1, the
pause invariant re-ratified in #2486. Without the `paused || userPaused`
filter, the admission provider offers a user-paused card to the merge
pump.
- **`drainMergeQueue is single-flight`** — safeguard 4. Asserted via
`reconcileStaleMergeActive`, the first statement *inside* the guard, so
the probe isolates the guard rather than dispatching a real merge.
(Driving a real drain crashed the vitest worker; probing the guard
directly is both safer and more precise.)

**Both are two-sided** — they assert the guard blocks *and* permits. A
one-sided test would still pass against a guard that rejects everything,
which is a real failure mode for a filter.

## Proven by mutation delta

Baseline fail-set vs mutated fail-set on the identical selection, NEW
failures only:

| Mutation | NEW failures |
|---|---|
| remove the pause filter | **1** — the pause test, and only it |
| remove the single-flight guard | **1** — the single-flight test, and
only it |
| filter rejects *everything* | **1** — proves not one-sided |
| drain *always* refuses | **1** — proves not one-sided |

## Gate admission

`project-engine.test.ts` joins the `engine-core` allow-list. **One file
proves five safeguards** — user pause, `autoMerge:false`, capacity
single-flight, the pre-enqueue merge-proof consult, and at-most-once
enqueue.

Before this, **none of the six safeguards was defended by blocking CI**.
A regression surfaced only in non-blocking full-suite, after the merge.

Measured, not assumed:

| | Files | Tests | Wall (3 runs) |
|---|---|---|---|
| before | 17 | 309 | 5.19 / 5.51 / 5.19s |
| after | 18 | 412 | 6.19 / 6.24 / 6.21s |

**+~1.0s against a ~60s ceiling.**

**Verified the gate fires**, rather than assuming the allow-list edit
took — the failure mode greptile caught in #2494:

- remove safeguard 1 → `pnpm test:gate` **exits 1** (1 failed / 411
passed)
- remove safeguard 4 → **exits 1** likewise
- restored → **exits 0**

Deterministic: store, runtime, merger and notifier all mocked; no real
git, no network, no real timers in these two cases.

## Reversible calls I made rather than asking

- **Added to `project-engine.test.ts` rather than a new file.** A
dedicated file would need ~200 lines of duplicated `vi.mock`
scaffolding; reusing the existing harness also means one gate admission
covers five safeguards instead of two.
- **Did not wait for U8.** These guard code that exists today and the
conversion needs them in place first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:20:04 -07:00
gsxdsm
e9bfd0d313 test(engine): prove the admission-control agent count on a renamed board — and two of my own cases were vacuous (#2516)
Test-only. Fifth E2E family. Closes three of the four `live-agent-count`
classifications.

## Why this one is a scheduler bug, not a display bug

`live-agent-count.ts` classifies a card's column, and
`persistedTopLevelAgentSlotsFromStore` turns that into **the number
admission control compares against the cap**. So a mis-classified column
fails in whichever direction hurts:

| mis-classification | consequence |
|---|---|
| wip column not recognised | under-count → **over-admits past the
operator's cap** |
| complete column not recognised | a finished card counts forever →
**board silently stalls** |

Both are silent, and both land only on a renamed board.

Everything in the path is real: PostgreSQL store, real persisted
workflows, cards walked through the real transition policy, and the real
counting function resolving each card's own IR. Nothing about counting
is reimplemented here.

## Two of my own cases were vacuous — mutation-testing caught it

This is the more useful half of the PR.

**1. "does not count a card in the COMPLETE column" passed with the
terminal classification hardcoded to `done`.** `isRunningAgentTask`
rejects that card at the *wip* check anyway, so the test was really
asserting "shipped isn't a wip column". `terminalKind` short-circuits
**first**, so it only changes the answer for a card whose status would
otherwise make it count. Now covered by a complete card carrying a
live-looking `planning` status — the state a crashed run leaves behind,
which on a renamed board consumes a slot forever.

**2. The review/merge lane had no case at all.** A review status is
deliberately *not* globally live (a stale `fixing` in wip must not
consume capacity), so it is gated on `columnIsReviewOrMerge`. If the
renamed review lane isn't recognised, a genuinely-active reviewer stops
counting and admission control lets another agent in over the cap. Now
covered by a review card with an active merge-pipeline status.

Both new cases assert their fixture took effect first, so they can't
degrade back into the weaker version silently.

## Mutation-verified independently

| classification | mutation | result |
|---|---|---|
| `countsTowardWip` | → `"in-progress"` | 3 renamed cases fail |
| `complete` | → `id === "done"` | exactly the new terminal case fails |
| `mergeBlocker` | → `id === "in-review"` | exactly the new review case
fails |

## What this does NOT cover, stated plainly

The fourth classification, `columnIsIntakeOrHold`, is read only by the
**waiting** predicate, which the admission count never calls. It stays
in the ledger as unproven rather than being claimed by proximity — the
mistake I made last slice with `resolveMergeOrchestrationColumn`.

## Verification

- five live-E2E suites green together: **52/52**
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (309 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added end-to-end coverage for live agent-count admission across
workflow lanes and lifecycle states.
* Verified slot handling for active, completed, held, and mid-review
tasks, including stale statuses.
* Confirmed mixed-lane counts and renamed board vocabularies produce
consistent results.
* **Documentation**
* Expanded coverage notes for live agent-count classifications and
waiting-state behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:19:22 -07:00
gsxdsm
2e39763930 test(engine): prove agent-link hygiene on a renamed board — a leaked agent slot, not a stale link (#2514)
Stacked on #2510. Test-only. Closes the `task-agent-sync` ledger entry.

## The defect this reproduces, in the code's own words

`task-agent-sync.ts`'s conversion note:

> a move into a renamed terminal column matched nothing and this handler
returned early — so the agent kept a `taskId` pointing at a finished
card and stayed `running`, **with no error and no failing test**.

"No error and no failing test" is the whole problem — and the cost is
not a stale link. **The scheduler counts `running` agents against its
cap**, so on a renamed board every completed task permanently consumes
an agent slot until a human notices. A board would just get slower and
slower.

## Everything in the path is real

Real PostgreSQL `TaskStore`, real `AgentStore`, the real
`attachAgentLinkSync` subscribed to the store's real `task:moved` event
(the same call `in-process-runtime` makes), and a real `moveTask` to
trigger it. Assertions read the **agent row** back out of the store —
never "the handler was called".

## Mutation-verified

Forcing the legacy literal sets (the pre-conversion behavior) fails
**exactly the two renamed cases**, leaving the default-vocabulary floor
and both negatives green. So this reproduces the original defect rather
than merely covering the file.

## Negative half

An ordinary mid-lifecycle move (`wip → review`) must **not** release the
agent — given the same time to run as the positive case. "Clear the link
whenever the card moves" would drop the binding the moment work started,
a louder failure than the leak it fixes.

## Two anti-flake, anti-vacuity details

- **Async delivery.** `task:moved` is a plain EventEmitter and the
handler is async, so the assertions **poll the persisted row** to a
bounded deadline and fail with the row's actual contents. A fixed sleep
would flake in both directions.
- **The fixture asserts itself.** The link and `running` state are
verified *before* the move, so an agent that was never linked cannot
make this pass for the wrong reason — the failure mode I hit twice
already in this program.

## Verification

- four live-E2E suites green together: **39/39**
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Improved agent-link cleanup when tasks reach workflow completion.
* Ensured completed task links are released even when workflow columns
have been renamed.
* Preserved active agent links when tasks move through non-terminal
workflow stages.
* Improved reporting of link cleanup outcomes and handling of
synchronization errors.

* **Tests**
* Added live PostgreSQL end-to-end coverage for completion, in-progress
moves, and renamed-column scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:59 -07:00
gsxdsm
3badc244a7 U12 part 2: bind the three U5 reconciliation guards — USER-VISIBLE (and one path that couldn't run under PostgreSQL at all) (#2512)
## U12 part 2 — the three U5 reconciliation guards now actually fire

USER-VISIBLE. Taken on standing authority; here is exactly what changed
for operators.

All three read the RAW `experimentalFeatures.workflowColumns` key via
`store.workflowColumnsFlagOn()`. Nothing in production writes it, so all
three have been inert since the workflow-columns cutover.

| Guard | Before (every real project) | After |
|---|---|---|
| Workflow edit removing an **occupied** column | Save succeeded; cards
left in a column the workflow no longer declares | Save fails with
`OccupiedColumnsError` unless `rehomeTo` is supplied |
| Workflow **delete** | Occupant capture returned `[]`; cards sat in the
deleted workflow's columns until the next engine start | Cards move to
the default workflow's entry column as part of the delete |
| Workflow **switch** | Never reconciled; the `reconciliation` field in
the declared return type was never populated | Card in an undeclared
column moves to the resolved target; a declared column is preserved |

Both consumers already handle the new outcomes and needed no change:
`register-workflow-routes.ts` maps `OccupiedColumnsError` to a
structured 409 carrying per-column occupant counts, and
`fn_workflow_update` returns a retryable structured result. The
dashboard editor's `rehomeTo` retry flow becomes reachable for the first
time. I only updated two stale "flag-ON" comments there — that code was
correct all along and simply never fired.

### What an operator actually sees (USER-VISIBLE — read this bit)

Four changes to what the board and the API do. Nothing here is silent.

1. **Editing a workflow to remove a column that has cards in it now
FAILS.** Previously the save succeeded and the cards were left in a
column their workflow no longer declared. The dashboard shows the
existing 409 with per-column occupant counts and prompts for a re-home
target; retrying with `rehomeTo` moves the cards and saves. Removing an
EMPTY column is unaffected.
2. **Deleting a workflow moves its cards immediately** to the default
workflow's entry column, instead of leaving them until the next engine
start.
3. **Switching a task's workflow moves the card** when the new workflow
does not declare its current column. A card whose column IS declared
stays exactly where it is. The API response now carries the
`reconciliation` summary it always promised.
4. **A switch whose re-home would be REJECTED is now refused before
anything is written.** If the destination column is at its WIP limit,
the switch fails with a structured 409 (`workflow-switch-rehome-failed`)
naming the task, both columns and the reason — and **nothing changes**:
the task keeps its current workflow AND its current column. Retry after
making room. Previously this combination committed the selection and
then silently reported a move that never happened, leaving selection and
column disagreeing.

**Can a torn card still happen? Yes, in one narrow case, and here is how
you recover.** If the destination fills in the window between the
pre-flight and the move, the selection is already committed and the card
ends up in a column its new workflow does not declare. That case is not
silent: it writes a `task:workflow-switch-torn` run-audit row, and the
error carries `selectionCommitted: true` with both columns. Recovery:
make room in the destination and move the card there, or switch the task
back — and if neither happens, the R7 startup sweep
`reconcileUndeclaredTaskColumns` re-homes it on the next engine start.
The card is never lost; it is visible in a lane the board may not draw
until one of those runs.

The one thing to watch after merge: (1) converts a previously-silent
success into a visible failure, so an operator mid-edit on a busy
workflow will start seeing a 409 they never saw before. That is the
point — the alternative was stranding their cards — but it is the change
most likely to generate a "this used to work" report.

### The thing that made this more than a gate removal

Un-gating the switch guard surfaced that
`selectTaskWorkflowAndReconcileImpl` read the task through
`store.readTaskFromDb` — the **synchronous SQLite** reader, which throws
under PostgreSQL:

```
TaskStore.db: SQLite Database is not available in backend mode
```

The flag returned before that line, so the gate was hiding a path that
**could not execute at all in the production backend**, not merely a
disabled feature. Ported to the async `readTaskRow`. Found by the new
tests, not by reading the code.

### Review round 2 (both findings real, both fixed)

**Torn write with no alarm — fixed by ORDERING, not by a louder
message.** My first attempt only made the error loud, which left the
torn state intact. The real fix is that the deterministic rejection
cause (destination at its WIP limit) is now checked BEFORE
`selectTaskWorkflow` commits, by resolving the target IR straight from
`workflowId` instead of through the task's selection. Nothing commits on
that path.

For the residual race the failure is loud AND recorded: `rehomeOccupant`
now returns `{ moved, error? }` (additive; sweep callers ignore it), the
switch writes a `task:workflow-switch-torn` run-audit row, and throws
`WorkflowSwitchRehomeFailedError` with `committed: true`. Consumers
translate it: the dashboard route returns a structured 409 with
`selectionCommitted`, and `fn_task_set_workflow` returns the same fields
— no more generic "something went wrong".

**Fabricated column for a deleted task.** My first fix fell back to
`fromColumn` when the final read found no row, so a task soft-deleted
mid-switch was reported as having its old column *preserved*. Absent now
reads as absent (the optional `reconciliation` is omitted). Extracted as
the pure `buildSwitchReconciliation` seam because the window is not
reachable through the public call — `selectTaskWorkflow` rejects an
already-deleted task up front — so it is a genuine race, and I test the
decision directly rather than asserting it from reading the code.

### Revert-proof, measured

New `workflow-reconciliation-production-shape.pg.test.ts` — 6 cases,
with the flag **never written**, which is the configuration every real
project has. Each flip reverted individually:

- re-gate the edit guard → **2 failures** (OccupiedColumnsError case;
rehomeTo re-home case)
- re-gate the delete capture → **1 failure** (card stays in
`custom-hold`)
- restore the switch early return → **2 failures** (`reconciliation`
undefined; card does not move)
- all three in place → **6/6 green**

Round-2 fixes, also measured:
- restore the `fromColumn` fallback → the "row is gone" case fails
(reports `preserved: true` for a deleted task)
- drop the `!outcome.moved` throw → the capacity-blocked case fails
(resolves instead of raising)
- **move the capacity pre-flight back AFTER the commit → the case fails
on the SELECTION assertion** (expected `WF-002`, received `WF-001`),
i.e. it proves the ordering, not the wording

The pre-existing coverage in `workflow-authoritative-reads.pg.test.ts`
reached the occupied-column guard by **writing the flag ON itself** —
same pattern as the ListView/Board suites in part 1. Its flag write is
removed; it now runs in the production shape.

### Where I nearly got this wrong

My first revert harness was buggy and I briefly concluded the delete
re-home was **redundant** — I had probed the stored column and seen
`triage` with what I thought was the flip reverted. It wasn't.
`workflow-ops.ts` contains two identical `const occupantTaskIds = await
store.listWorkflowOccupantTaskIds(id, false)` lines (field-reconcile
block, delete path), so my first-match edit reverted the wrong one.
Re-run anchored on surrounding context, the delete case fails as
predicted. Recorded in the test header as a caution. I also chased and
**refuted** a scarier hypothesis along the way — that an unrelated
`updateTask` coerces a custom column back to `triage`. It does not; the
column survives.

### Deliberately NOT in this PR

The v1-IR rollback-compat persistence (`downgradeIrToV1IfPure`) on the
workflow UPDATE path. It shared the same `flagOn` variable, which is how
it surfaced: **one flag read was feeding two unrelated decisions, so the
flag has more decision sites than call sites** — my earlier 9-site
inventory undercounted. It chooses the stored *shape* of the graph
rather than gating a guard, so it is a persistence-format change with a
different blast radius. It now reads the flag explicitly, behaviour
unchanged, for a follow-up.

The `moves.ts` group remains U2b's.

### Verification

`pnpm test:gate` (307 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (17
steps), both typechecks green. Full `packages/core` PostgreSQL suite:
**1042 passed, 3 failed** — `central-archive-secrets.test.ts`
(log-prefix assertion) and
`workflow-settings-project-identity.pg.test.ts` (×2, project-id
resolution). I confirmed the identical 3 failures on a stashed clean
tree: pre-existing, unrelated. No Fusion instance booted.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Workflow edits now prevent removal of occupied columns unless cards
are moved to a specified destination.
* Cards are automatically re-homed when workflows are deleted or
switched.
* Workflow switches now check destination capacity before committing and
provide clear conflict details when re-homing fails.
* Reconciliation results now indicate whether cards were moved or
preserved.


<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:46 -07:00
gsxdsm
3bbb6ffc6b capacity, part 2: delete the cross-project concurrency cap (enforcement half) (#2509)
Stacked on #2502 — review that first; this branch contains its three
commits.

Operator: two capacities **per project**. `globalMaxConcurrent` is a
machine-wide *third* limiter kept in a separate authority (a central-DB
singleton row) that every runtime had to subscribe to and periodically
re-reconcile. It goes.

This slice removes **enforcement and wiring only**. The setting key,
central DB state, API route and Settings UI come out in part 3, so each
half lands green and independently revertable.

**Deleted:** the shared `AgentSemaphore` instance in `ProjectManager`
and `ProjectEngineManager`; the per-project `ScopedAgentSemaphore` in
`InProcessRuntime`; the `globalSemaphore` runtime-config field; both
`concurrency:changed` subscriptions; ProjectManager’s 30s limit-refresh
poll; the residual-slot return on project stop.

The scheduler/triage semaphore gate is now simply **absent** — same
shape as the worktrees-off gate in #2502. `semaphoreGate?` was already
optional, so no gate object is constructed rather than one holding an
infinite limit. Absence cannot start binding again by accident.

---

## Two findings that changed the shape of this slice

**1. `AgentSemaphore` the class stays — my earlier estimate was wrong
and I withdraw it.**

I previously told the coordinator that ~75% of `concurrency.ts` (≈662 of
886 lines) was semaphore machinery that could go with this cap. That was
line-range arithmetic, and it was wrong. `AgentSemaphore` is a general
primitive with four consumers unrelated to the global cap:

| Consumer | Governs |
|---|---|
| `verification-concurrency.ts` | `maxConcurrentVerifications` |
| `research-orchestrator.ts` | research `maxConcurrentRuns` |
| `experiment-executor.ts` | `maxConcurrentExperiments` — **a knob
absent from my original inventory** |
| `step-session-executor.ts` | parallel workflow steps |

What goes is the global **instance** and its wiring, not the class. I
will report the measured `concurrency.ts` delta after part 3 rather than
repeat an estimate.

**2. `acquireGlobalSlot` / `releaseGlobalSlot` had no production callers
— only tests.**

So the cross-project cap had *two* mechanisms: the in-memory semaphore
(live) and a durable central-DB `currentlyActive` counter (dead — never
incremented by real work). Both deleted, along with the tests that
pinned the dead passthrough.

## The regression this almost introduced

`runWithMergeAdmission` in `project-engine.ts` opened with:

```ts
if (!semaphore) return await start();
```

Unreachable while a global semaphore always existed. With the semaphore
gone it would have fired on **every** merge and skipped
`projectAdmissionCoordinator.admitOldest` entirely — silently stopping
merges from counting against the **per-project** agent count.

That is the opposite of the intent: a merge *is* an agent and still
consumes one of the project’s slots; it just no longer consumes a
machine-wide one. So the early return is **deleted rather than left to
fire**. `admitOldest` already declares `semaphore` as optional and
enforces `maxConcurrent` independently of it (`claimed() + reservations
>= maxConcurrent`), so dropping the argument preserves per-project
admission and oldest-first fairness exactly.

Worth flagging as a pattern: this is the third time in this unit that a
branch which was *unreachable* became *always-taken* once a limiter was
removed. The type system caught the worktree one; this one was only
visible by reading the branch, because the semaphore was reached through
an `any` cast.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (309 +
10 + 71) · project-manager + hybrid-executor + merge-single-flight +
scheduler 93/93.

Nothing booted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:39 -07:00
gsxdsm
743df98aa4 capacity, part 1: merge pinned at 1, worktrees-off mode, and one dead knob deleted (#2502)
First slice of the capacity simplification. Operator: *"just have two
capacity — overall per project agent count and max worktrees. Remove all
other capacities and counts."* Plus two later additions: **merge is
always 1, fixed**, and **worktrees off ⇒ limit by total agents only**.

Three independently revertable commits. No limiter is added anywhere;
one is deleted, one is made structurally absent, and one is pinned.

---

## 1. Merge concurrency ratcheted at 1 (test-only)

I was asked to add a limiter if merge concurrency could be raised. **It
cannot** — there is no setting, workflow property, pool or trait config
anywhere that raises it, so this adds no code and pins what already
holds.

Serialization lives in the **pump**: `drainMergeQueue`’s `mergeRunning`
re-entrancy latch, `activeMergeTaskId` as a single-slot identity, the
`mergeBodyInFlight` next-generation latch, and one `ProjectEngine` per
projectId.

**Not** in the merge-queue lease, which is a per-task ROW (`primaryKey
[projectId, taskId]`) — two tasks can hold leases simultaneously by
construction, and it has exactly one caller (the worktree-reuse
handoff). Ordinary merges never take it. A lease-level test would have
been describing an invariant that layer has never held.

The second half guards the other direction: a merge-concurrency
*setting* would not fail the pump ratchet — it would sit unread until
someone wired it up.

**Revert-proof:** deleting the latch → `expected 1 times, but got 2
times`; deleting the `finally` → latch-stuck; injecting
`maxConcurrentMerges: 2` → fails naming the key; injecting a
`maxParallelLanes` merge-trait field → fails naming the field. Sources
restored byte-identical after each injection.

## 2. `worktreesEnabled` — off means the worktree limit cannot bind

No worktrees-off mode existed (no
`worktreesEnabled`/`useWorktrees`/`worktreeMode` anywhere — only
worktree *configuration*).

**Why not `maxWorktrees: 0`, which needs no new key:** it deadlocks. `??
4` keeps `0` (not nullish), the gate is `used >= limit`, so `0 >= 0`
holds **on an empty board** and nothing ever dispatches — while the
operator-visible reason reads `gate=maxWorktrees; used=0/0`, a limiter
that looks like it is working while the board is dead. It also needs the
Command Center `{min:1}` clamp relaxed. So `0` costs the gate rewrite
*and* the clamp change *and* encodes a mode as a magic value.

**Off is absence, not a big number.** `resolveWorktreeCapacityLimit`
returns `number | null`; `ConcurrencyGateDiagnostic.maxWorktreesGate` is
now optional, so consulting a worktree limit in OFF mode does not
type-check. A gate holding `Infinity` can start binding again the moment
someone "fixes" a comparison; an absent gate cannot.

That paid for itself immediately: making it nullable surfaced a
**second, independent** worktree gate (`activeWorktrees >= maxWorktrees`
early-return) that a skip-by-convention approach would have missed
silently.

**Scope, deliberately:** this is a statement about *counting*, not
isolation. It does not make concurrent agents safe to share one checkout
and builds nothing toward that — the non-worktree paths that exist today
are fallbacks to the operator’s own tree, one of which caused FN-8600.

**Revert-proof:** a resolver ignoring the flag turns both OFF scheduler
tests red while every ON test stays green — they reuse the *same*
fixture (5 in-progress, limit 4) that pre-existing tests prove blocks,
so the pair moves in opposite directions. Removing `disabled:` reddens
the UI test.

## 3. `maxTriageConcurrent` deleted — it controlled nothing

**Measured: zero enforcement reads.** The only `.maxTriageConcurrent`
reference in the repo was a route echoing it back in `/config`. FN-8453
removed the pool it gated and left the knob shipping in
`DEFAULT_SETTINGS`, the settings type, the section registry, the API
response and six i18n catalogs, doing nothing, for releases.

Historical FNXC comments are **updated, not deleted** — they explain a
real past incident; they now say "planning admission slot" so they stop
implying a live setting. Tombstoned so it cannot return.

`/config` loses a field; safe in-repo since `fetchConfig`’s own return
type never declared it.

---

## Two corrections worth recording

- I earlier reported `maxWorktrees` had **no** Settings UI. Wrong —
`WorktreesSection.tsx:47`; my grep was truncated by `head`. It changed
the placement (toggle beside it, rather than a duplicate key in
Scheduling).
- I planned to assert the queued-reason string is rewritten in OFF mode.
Measured that it is **unreachable**: when `maxConcurrent` binds, the
sweep bails before the per-task reason and logs nothing. The test
asserts absence instead.

Two near-misses caught before commit: a pre-existing FN-7505 guard
caught my *new* key missing a description mapping; and editing i18n via
`json.load/dump` silently dropped unrelated duplicate keys
(`autoUpdateAndRestart` in `fr`) — Python keeps only the last of a
duplicated key. Redone textually, every catalog re-validated.

## Verification

`pnpm lint` clean · core/engine/dashboard/i18n typecheck clean · `pnpm
test:gate` green (309 + 10 + 71) · capacity/worktree suites 11/11 ·
engine merge-invariant + scheduler 45/45 · dashboard settings 114/114.
Rebased onto current main and re-verified.

Nothing was booted at any point.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a project setting to enable or disable running tasks in
worktrees.
* Disabling worktrees removes worktree capacity limits from task
scheduling.
* The “Max Worktrees” setting is disabled when worktree execution is
turned off.

* **Changes**
* Removed the unused triage concurrency setting from configuration and
dashboard responses.
* Updated scheduling diagnostics and queue messages to reflect disabled
worktree capacity limits.
  * Added localized labels and help text for the new setting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:31 -07:00
gsxdsm
35b0df1838 U11 PR2: entry contract under the merged column + a real intake-column bug the audit surfaced (#2503)
Second small PR for **U11**. Two commits: a tests-only entry-contract
pin, then a **real present-day bug fix** the audit surfaced.

## The audit you asked for, finished — no design fork

You named four surfaces as the remaining risk. All four can take a
combined `intake` + `hold` column. One needed a code change; here it is.

| Surface | Verdict | Evidence |
|---|---|---|
| `isUnplannedForExecution` | Safe | PR1 (#2495) — passed unmodified; a
mutation now fails exactly the merged-column test |
| Capacity hold / release | Safe | PR1 — `hold-release.ts:260` already
accepts intake **or** hold |
| `start`'s column / entry contract | Safe | commit 1 — all 6 assertions
passed unmodified |
| `createTask` intake wiring | **Broken today** | commit 2 — fixed,
revert-proven |
| *(also found)* triage auto-discovery | Needs conversion |
`triage.ts:1382` — deferred to PR3, see below |

## Commit 1 — entry contract under the merged column (tests only)

All 6 new assertions passed on the first run. **Regression floor, not
evidence of a fix** — I could not make them fail and am not claiming
otherwise.

They pin one real behavioral **difference** rather than asserting
sameness everywhere: the merged shape answers `start` where the split
shape answers `plan`, because `start` becomes the first node in that
column once the columns collapse. That is equivalent *only* because
`start` reaches the specification node by a single unconditional success
edge — asserted, so if a node is ever inserted between them this fails
instead of silently admitting an unspecified card into implementation.

Also pinned: past planning both shapes agree exactly; a card past the
merged column still never resumes at a planning node (the backward drag
that fires `abort-on-exit`); and a row persisted in the **deleted**
`triage` column resolves to `undefined`, safe only while the executor's
start-node fallback exists.

## Commit 2 — a real bug, found by the audit

The intake column was resolved **only** as a by-product of materializing
workflow steps. A create supplying `enabledWorkflowSteps` without an
explicit `workflowId` takes **neither** materialization branch, so
`resolvedEntryColumn` stays `undefined` and `column:` falls through to
the hard-coded `|| "triage"`.

Today, on Coding (Ideas), that lands the card in `triage` — **a column
that workflow does not declare.** Created straight into a phantom lane.
Measured: the new test fails `expected 'triage' to be 'ideas'` against
unmodified sources.

**Why it blocks U11.** Once `triage` leaves the coding IRs this stops
being an Ideas edge case and becomes the default workflow's behavior for
every create down this path: the card lands in an undeclared column
**and** — because `isIntakeColumn` keys on the same `"triage"` literal —
gets `generateSpecifiedPrompt` instead of the bootstrap seed. Triage
admits a card for planning only when its `PROMPT.md` reads as a seed, so
a placeholder spec is classified "already planned" and never planned.
The card sits in Planning forever with no log line in any lane —
**FN-8587's exact failure mode, promoted from one edge case to every new
card.**

The fix resolves the intake column **side-effect-free** (read the IR,
ask which column carries `intake`). It deliberately does *not* call
`materializeDefaultWorkflowSteps`, which would persist step rows the
caller explicitly opted out of by supplying its own toggles.
Unresolvable workflow returns `undefined` and each call site keeps its
legacy fallback, so no path loses behavior when the IR cannot be read.

Applied to both create paths. Branch ordering preserved in both — the
explicit empty-toggle case (`length === 0` hydrating back as `[]`) still
runs, now nested rather than sequential.

**Revert check:** with `task-creation.ts` reverted, *"lands a Coding
(Ideas) task in ideas even when enabledWorkflowSteps is supplied"* fails
`expected 'triage' to be 'ideas'`. The companion bootstrap-`PROMPT.md`
assertion passes either way today — it is correct **by accident of the
`"triage"` literal** — and is kept precisely because that accident
disappears with U11.

## Verification

37 tests green across the three intake/create suites; 119 across the
entry-contract, merged-column and lifecycle suites; `pnpm test:gate`
green (307 + 10 + 71); lint and core typecheck clean. Changeset added.

## Deferred to PR3, with the line numbers

`discoverReadyPlanningTasks` has two hardcoded branches:

```ts
(t) => t.column === "triage" && isTaskStillInPlanningStage(t)   // triage.ts:1382
(t) => t.column === "todo"   && !this.processing.has(t.id) …    // triage.ts:1389
```

Delete `triage` and branch 1 matches nothing for coding cards; branch 2
then does all the work and is **narrower** (it admits only
`needs-replan` or bootstrap-stub cards). Commit 2 is what makes branch 2
sufficient — every new card now gets a real bootstrap seed. They cannot
double-fire: a card is in `todo` xor `triage`.

Two adjacent sites are already merged-shape-ready: `triage.ts:3899`
skips the redundant same-column move for a plan-in-place card, and
`triage.ts:753`'s stale-status sweep already scans both columns.

Then the ~10-line IR change, then the migration proof.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:14 -07:00
gsxdsm
9d3e53d0c5 U8 PR3: the implementation phase announces HOW it ended — including when the executor moved the card itself (#2507)
Third PR of **U8 — the graph owns execution**. Independent of everything
merged so far; small, green, revertable on its own.

## The problem this makes visible

`result.taskDone` is the entire language the execute seam has for
talking to the graph:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
return { outcome: "failure", value: paused ? "implementation-paused" : "implementation-incomplete" };
```

The endings that one bit cannot express are exactly the ones the
implementation phase **transitions itself**:

- a session that paused *after* the work was already complete →
finalizes to review inline;
- a session that stopped because a step is blocked on a pending review →
hands off to review inline (a pending-review block is a wait, not a
failure; marking it failed deadlocks a row that is both `in-review` and
`failed`).

The graph then sees `taskDone === false`, reports
`implementation-incomplete`, and `handleGraphFailure` compensates with
`alreadyFinalizedToReview` / `completionFinalized` — classifiers whose
entire job is recognising a move the graph did not make.

**That was invisible.** An out-of-band transition and a genuine
implementation failure were indistinguishable in logs, in events, and in
tests. You cannot remove a transition you cannot see, and you cannot
prove you removed it either.

## What lands

A closed `ImplementationExit` enum
(`engine/executor/implementation-exit.ts`) reported from six
completion-adjacent exits in `runImplementation`, announced by the
execute seam as `NodeCompleted.exit` on the U3 lifecycle bus. Two ids
are flagged as out-of-band — the ones where the executor, not the graph,
performs the transition.

**Routing is unchanged, and that is the point.** The seam returns
byte-identically what it returned before for every exit, so this PR
cannot move a card. The routing move needs new IR edges and lands
separately; splitting them is what keeps both independently revertable.
Per R5 an exit id is a **reaction** — nothing branches on one, and
dropping every subscriber must change no outcome (a named U8 test
scenario, asserted here).

`NodeCompleted.exit` is added to the event key allow-list deliberately —
which is exactly what that allow-list is for — and carries closed enum
ids only, never prose.

## Revert-proofs, each observed failing

| Injected change | Result |
|---|---|
| Remove the emit entirely | **6 failures** |
| Let an exit change the returned outcome | **2 failures** (the
routing-unchanged pins) |
| Delete one `reportImplementationExit(...)` call site | **1 failure**
(the wiring ratchet) |

**The third proof exists because of a hole I found in my own tests.**
These tests stub `runImplementationPhase` — the only way to reach all
six exits deterministically — which means deleting a real call site left
the entire file **green**. A stubbed seam can only prove the seam. I'd
also written "every exit is reported — the signal is real, not a
placeholder" in the header, which the tests did not support. Both are
fixed: there is now a ratchet asserting every enum id is wired at a real
call site and that each out-of-band id sits adjacent to the handoff it
describes, and the header says what the tests actually prove.

## Scope

**6 of `runImplementation`'s ~28 dispositions** (per the ownership
ledger merged in #2490), chosen as the ones the routing move needs. The
remaining ~22 report nothing yet — the ledger, not this enum, stays the
record of that gap, and the module says so.

## Verification

- 15 new tests + ledger + graph-boundary + task-done-blocked +
graph-requeue-gate + step-session + review-verdicts + tool-failure-retry
— **9 files, 115 tests green**
- `@fusion/core` `workflow-events` — 20 tests green (allow-list change
covered)
- `pnpm test:gate` green (17/307, 2/10, 1/71); `pnpm lint` clean; `tsc
--noEmit` clean on both packages
- Changeset included (`patch`, `internal`), passes `check:changesets`

## Next

PR4 is the routing move itself: `review-handoff-pending-review` becomes
a graph outcome with its own IR edge, and `alreadyFinalizedToReview`
becomes provably unreachable for that path. The IR edge change will be
its own commit, separate from the seam change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:07 -07:00
gsxdsm
d5030c55ea test(engine): prove BOTH rebound paths on a renamed board — 2 more ledger entries closed (#2510)
Stacked on #2508. Test-only.

## Why this site matters more than most

`resolveReboundTarget` answers one question: **where does a recovered
card go back to?**

Keyed on the literal `todo`, a recovered card on a renamed board is
requeued to a column that board **does not declare**. That is not
cosmetic — an undeclared column carries no trait flags, so `findColumn`
returns undefined and the card becomes invisible to every trait-driven
sweep: nothing schedules it, nothing releases it, the board does not
draw the column. **The "recovery" strands the card harder than the
failure it was recovering from.**

One of the two covered paths, `reconcileUndeclaredTaskColumns`, exists
*specifically* to repair that state — which makes it the worst possible
place for this bug to live.

## Covered, each mutation-verified independently

| site | mutation | result |
|---|---|---|
| `reconcileUndeclaredTaskColumns` | target → `"todo"` | exactly the 2
renamed cases fail |
| `autoRecoverWorktreeSessionStartFailure` | rebound → `"todo"` |
exactly the renamed requeue fails |

Neither needs git — the corrected ledger lens from #2508 (*what the
function touches*, not *what family it sits in*) made that obvious
rather than assumed.

## Both negatives included

"Re-home anything whose column looks wrong" would be a louder failure
than the strand it repairs, so: a card whose column **is** declared is
left alone, and an operator `userPaused` park is never undone.

## Fixture finding, kept in-file

`updateTask({ userPaused: true })` leaves the field `undefined` on both
`getTask` and `listTasks({slim:true})`. Seeding it that way produced a
card the sweep **correctly** saw as unpaused — a broken fixture that
would have read as a broken guard, and would have looked like a real
safety hole in the paused-park protection.

Found by probing the persisted row rather than trusting the write. Now
seeded through the integer column directly, and the test asserts the
seed took effect *before* exercising the sweep, so this cannot silently
regress into a vacuous pass.

## Verification

- three live-E2E suites green together (lifecycle, merge-family,
rebound-family)
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:01 -07:00
gsxdsm
4eaa509024 test(engine): prove merge finalization on a renamed board — 3 ledger entries closed, and the ledger itself corrected (#2508)
Test-only. Closes three `auto-merge-finalization` entries from the
unproven-sites ledger.

## Why this one first

It is the **last move a card makes**. Keyed on the literal `done`, a
renamed board's proven-merged card is moved to a column its own workflow
does not declare — or refused and left stranded in review **with the
work already landed**. That is the most expensive failure shape in the
lifecycle, and nothing had run it against a renamed workflow.

## The ledger was wrong, and is corrected in this PR

My own ledger said this family *"needs a REAL git worktree, branch, and
squash … an engine-slow real-git lane, not another table row"*.

`finalizeProvenAutoMergeTask` **needs no git at all** — the merge proof
is a field on the row. It was reachable the whole time. The inference
came from *the family the code sits in* rather than from what the
function actually touches, and it parked reachable coverage for a slice.
The correction is written into the ledger so the remaining entries get
re-checked the same way rather than inheriting the assumption.

## A second correction, from mutation-testing rather than reading

I first claimed `resolveMergeOrchestrationColumn` as covered because it
sits in the same resolver as the other two. **All cases passed with it
hardcoded.** It changes only whether finalization records a
column-mismatch *repair* — never where the card lands, which is why the
other cases are blind to it.

It got its own case. Keyed on `in-review`, a renamed board's card
resting in `checking` compares unequal, so **every ordinary finalization
would be audited as repairing a mismatch that never existed** — a
healthy board reads as one constantly self-healing, and the audit trail
operators use to spot real strandings fills with false positives.

Sitting next to covered code is not coverage.

## Mutation-verified independently

| mutation | result |
|---|---|
| `completeColumn` → `"done"` | 3 fail — both renamed cases + the
differential |
| `isCompleteColumn` → `id === "done"` | exactly the already-done case
fails |
| `mergeColumn` → `"in-review"` | exactly the new audit case fails |

## Shared fixture extracted (pure move)

The vocabulary + IR builder moved to `_workflow-vocabulary-fixture.ts`
so the two suites cannot drift into testing different workflows — two
copies of a differential fixture is precisely how a renamed-workflow
test starts passing for reasons unrelated to the code under test. The
lifecycle suite is unchanged: **20/20 before and after**. The
`mergeOrchestration` trait is an opt-in option so the existing suite's
IR stays byte-identical.

## Fixture note worth keeping

Seeding needed **completed steps**: task creation parses three pending
steps out of the bootstrap PROMPT even with `applyDefaultWorkflowSteps:
false`, and `getTaskHardMergeBlocker` refuses on them (`"task has
incomplete steps"`). Found by the suite blocking on **both**
vocabularies — the signature of a broken fixture rather than a broken
guard.

## Verification

- 51/51 across the lifecycle, merge-family, ratchet and hold-release
suites
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:15:35 -07:00
gsxdsm
8288e4a8ab U7 PR3: the specification reaction acts on what finalize DID, not on the fact that planning stopped (#2506)
Completes the pair started in #2498. That landed the outcome; this makes
the engine's reaction consume it.

## The bug

`onSpecifyComplete` fired on **every** finished specification, because
the seam announcing it fired unconditionally. So a card parked at the
manual plan-approval gate — finalize writes `status:
"awaiting-approval"` and **returns early**, before the release move —
was logged as `Specified X → todo` and had a Plan Review run armed for a
plan the operator had not approved.

#2491 stopped the **seeder** from acting on that, defensively, at the
seeder. This removes the reason it was ever asked. Both layers are
deliberate and neither is redundant:

- the seeder guard covers **every caller**, including self-healing's
re-seed;
- this one stops the engine doing work nobody asked for, and stops it
telling the operator something false about their own board.

`released` is the only outcome that licenses arming a run — the only one
meaning the card crossed into the hold column (or was already resting
there, plan-in-place) and is the graph's now. `parked` belongs to a
human; `withheld` belongs to the caller's retry budget.

## The event still fires on every outcome

Deliberately. Dropping the reaction for a non-release would also drop
the runtime's `recordActivity()` idle signal, and a reaction that
silently does not happen is harder to reason about than one that happens
with an accurate payload. R5's division of labour: **the seam announces,
the subscriber decides what a given outcome licenses.**

## Why there is a new extracted function

`reactToSpecificationComplete` is pulled out of the inline
`InProcessRuntime` callback for the same reason the continuation drain
was in #2491: the callback is built inside a class whose construction
attaches to the real central project registry, so no test could
distinguish *"the reaction respects the outcome"* from *"the reaction
ignores it"*.

**Revert proof:** with the outcome gate removed from the reaction, **5
of 8 fail**.

## Two call-site decisions worth naming

**`tryFinalizeExplicitDuplicateMarker` reports through a mutable ref,
not a widened return type.** Its boolean answers a *different* question
— "was this a duplicate marker at all?" — and 16 existing tests assert
it directly. I tried the widened return first and it turned all 16 red.
Expectation edits are exactly how a behavior change travels disguised as
churn, so I backed it out. **This diff touches zero existing test
expectations.**

**A duplicate-marker redirect reports `parked`**, which is accurate: it
deletes, flags, or clears the marker; it never releases the card into
the hold column.

## A fixture note — third of this shape on the program

My "task vanished between release and reaction" case passed `undefined`,
which triggered the harness **default parameter** and silently handed
the reaction a live task — making it a duplicate of the control rather
than the case it claimed to be. It now passes `null`, with a comment
saying why.

Running tally of near-false-greens on this unit, all the same family: a
fake that ignores its predicate (#2491), a stub that ignores its
callback (#2498), a default parameter that swallows the interesting
input (here). Each was caught by the test failing for the *wrong reason*
and being read rather than fixed.

## Verification

| Check | Result |
|---|---|
| new suite | 8/8 |
| 15 triage / planning / continuation suites | 361/361, **no expectation
edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (307 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:15:29 -07:00
gsxdsm
a2b4ca76ac U11: delete the unreachable legacy todo dispatcher from scheduler.schedule() (-929 lines, pure deletion) (#2505)
Based on `main`. **Pure deletion — no behavior change**, because the
deleted code cannot execute.

## Found while trying to convert it

This started as a U11 slice to make the scheduler's dispatch path
resolve its column by trait. Per the lesson from the dependency-blocked
feature I checked reachability *before* converting:

```ts
function shouldRunWorkflowColumnScheduler(_settings: Settings): boolean {
  return true;                       // parameter UNUSED, body a literal
}
...
if (shouldRunWorkflowColumnScheduler(settings)) {
  await this.runHoldReleaseSweepPass(tasks, settings);
  ...
  return;                            // UNCONDITIONAL, at the block's own depth
}
<929 lines of legacy pull-from-todo dispatcher>   // unreachable
```

The guard takes an **unused** parameter and returns a **literal**, so
the branch is statically always taken, and it ends in an **unconditional
`return`**. Everything after it in `schedule()` is unreachable.

`tsc` doesn't flag it because the condition is a function call rather
than a literal — which is exactly why 929 lines survived the U6 cutover.
The replacement was added *in front of* the old dispatcher rather than
*instead of* it, and the in-file comment says so outright:

> the hold/release sweep owns todo→in-progress pickup, so do not fall
through into the legacy pull-from-todo dispatcher after the sweep runs

## Why this matters beyond line count

**4 of the 15 `"todo"` literals in `scheduler.ts` live in this dead
region.** Converting them would have been pure waste — and worse, it
would have reported progress against the U11 critical path while
changing nothing. 11 live sites remain and are the real work.

## Corroborating evidence

Six imports became unused and are removed with it:
`resolveDependencyOrder`, `sortTasksByPriorityFanoutThenAgeAndId`,
`buildUnblockWeightMap`, `TransitionRejectionError`,
`isUnplannedSeedPrompt`, `DEFAULT_WORKFLOW_POOL_ID`.

That the dead region was their **only** consumer in this file is itself
evidence: a live dispatcher would still need dependency ordering and
priority sorting.

## Why no new test

The proof here is **static, not behavioral** — an unconditional `return`
before the code. A test cannot demonstrate absence of execution more
strongly than the control flow already does, and one that passed both
before and after would be theatre.

The evidence that nothing depended on it: **all 100 scheduler tests and
the full merge gate pass unchanged.**

## Measured

`scheduler.ts` **3,726 → 2,797 = −929 lines.**

Unlike every consolidation in this program, this is a **genuine net
reduction** — nothing was moved elsewhere.

## Verification

100 scheduler tests green across all 7 scheduler suites; merge gate
green (307 + 10 + 71); tsc clean; lint clean.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-28 17:15:23 -07:00
gsxdsm
e4004c8694 U9 baseline: pin merge-region IR config as a dead policy authority (test-only) (#2494)
**U9, PR1 of several.** Test-only, no production code touched. This is
the characterization baseline the plan's Execution note asks for before
the merge lane converts.

## The finding

`builtin-coding-workflow-ir.ts` declares merge-region policy that **no
engine code reads**:

| IR declaration | Consumed by |
|---|---|
| `merge-retry` → `{ policy: "merge", maxAttempts: 3 }` | nothing —
`retry-backoff` handler is `async () => ({ outcome: "success" })`
(`workflow-node-handlers.ts:728`) |
| `merge-manual-hold` → `{ release: "manual" }` | nothing — returns a
constant `manual-required` |
| `branch-group-*` → `{ maxReworkCycles: 3 }` | nothing — returns a
constant `success` |

Live merge policy authority is elsewhere, on two separate axes:
- **conflict** retries — `settings.maxAutoMergeRetries` (default 3),
already covered by `auto-merge-retry-cap-settings.test.ts`
- **transient** retries —
`ProjectEngine.MAX_AUTO_MERGE_TRANSIENT_RETRIES = 5`
(`project-engine.ts:545`)

So the IR is a **third, dead authority**. These are different axes, not
a same-axis contradiction — but a reader looking at the IR would
reasonably take the declared numbers as live, and nothing currently says
otherwise. U9's acceptance criterion is "merge policy changes via IR
config alone, with no code change"; that fails today and this pins why.

## Why characterization rather than a fix

Making these handlers config-driven is a **merge behavior change**, and
the `requestMerge` primitive it routes through lives at
`executor.ts:7383` — inside U8's blast radius. U9 is sequenced behind U8
precisely so the merge lane converts onto an executor that is already
substrate. Landing the behavior change now would change merge semantics
on an executor about to be reshaped. It lands inside U9 proper.

When U9 wires a node kind onto its IR config, the matching case here
goes **red** and the U9 commit must move that kind out of
`CONFIG_BLIND_MERGE_REGION_KINDS`. That is the ratchet working.

## Proof it fails when reverted

A test that passes with the change reverted is not a test. The "change"
here is the test itself, so the honest analogue is mutating the
characterized production behavior. Three independent mutations, each
reverted after measuring:

| Mutation | Result |
|---|---|
| `retry-backoff` honours `config.maxAttempts` (what U9 will do) | **2
failed** / 6 passed |
| `manual-merge-hold` honours `config.release === "external-event"` |
**2 failed** / 6 passed |
| IR declaration drift: `maxAttempts: 3` → `7` | **1 failed** / 7 passed
|

Measured: 8 tests, 4.16s. `pnpm lint` clean. Tree restored to clean
after each mutation.

The assertions are behavioral, not string matches: each handler is
invoked with two contradictory configs (opposite budgets, opposite
release modes, disjoint surfaces) and asserted to return deep-equal
results.

## Six safeguards

This PR changes no production behavior, so no safeguard is altered by
it. The full six-row table with test attribution is the required
artifact for the **conversion** PR, not this one. Baseline located so
far, to be completed and verified by mutation before any conversion
lands:

| # | Safeguard | Consulted at (today) | Test attribution |
|---|---|---|---|
| 1 | user pause | `project-engine.ts:645` (`task.paused \|\|
task.userPaused`) | not yet verified |
| 2 | `autoMerge:false` | `allowsAutoMergeProcessing` —
`project-engine.ts:2797`, `merger.ts:7178` | not yet verified |
| 3 | dependency gating | not yet located | not yet verified |
| 4 | capacity | not yet located |
`workflow-column-boundary-capacity.test.ts` (unverified) |
| 5 | merge-proof | `getTaskMergeBlocker` — `project-engine.ts:2609` |
`merger-file-scope-invariant.test.ts`,
`merger-diff-volume-gate.slow.test.ts` (unverified) |
| 6 | at-most-once merge | `activeMergeTaskId` single-flight —
`project-engine.ts:693`/`:2729` | not yet verified |

Rows 3, 4 and all attributions are honestly incomplete rather than
asserted — I will not present a table I have not earned.

## Also found, for the coordinator

- **Slice statuses are stale.** S02/S03/S04 in
`docs/plans/workflow-owned-merge-stack/` are all marked
`draft-stack-handoff` but S04 has **landed** (the merge-region IR nodes
above), S03's `claimDueWorkflowWorkItem` is implemented and wired via
`workflow-work-processor.ts`, and S02's
`projectMergeRequestToWorkflowWorkItem` is implemented with **zero
production callers**. S06/S07/S08 are genuinely not started. Doc
correction coming as its own small PR.
- **S1 prerequisite verified present, not assumed** — all four store
methods live in `store.ts`, migration `0031` in tree. No S1-completion
gap.
- **Second control plane into the merge lane:** `self-healing.ts:3198`
and `:7200` call `enqueueMerge` directly, bypassing the graph. That
needs to become a recovery-fact/wake (the stack's R6) during U9.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added a new test suite to cover U9 merge-region behavior across
supported workflow node types.
* Verified merge-region results are consistent across built-in,
contradictory, and missing configuration inputs.
* Documented current behavior for retry backoff (always succeeds) and
manual merge hold (fails as manual-required).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:55 -07:00
gsxdsm
eaea082259 U8 PR2: the execution-policy ladder resolves its own workflow's columns (the wip literal made retry, escalation and loop protection unreachable) (#2497)
Second PR of **U8 — the graph owns execution**, independent of
[#2490](https://github.com/Runfusion/Fusion/pull/2490) and of every
other unit. Small, green, independently revertable.

## The defect

`handleGraphFailure`'s execution-policy ladder — FN-7863/FN-7926
dispatch-loop terminalization, FN-7996 tool-failure retry, FN-7998
escalation — decided a task's own lifecycle by naming `"todo"` and
`"in-progress"` **literally, at 9 sites**. U5b converted the executor's
*rebounds* to `resolveReboundColumnFor`; these were left behind, each
sitting somewhere an awaited resolver could not reach: inside
synchronous `updateTaskAtomic` mutators, inside fire-and-forget resume
closures, and in conditions evaluated before any resolution happened.

**The severe one is the wip gate, and it fails silently in the worst
direction:**

```ts
if (live.column !== "in-progress") {
  // "Workflow graph run ended after task already advanced — no further action needed"
  return;
}
```

Under a workflow that renames the implementation column, that is true of
a card sitting in **its own wip column**. So the graph failure was
swallowed whole — no terminal park, no status, no error, nothing on the
board — and the scheduler re-dispatched the same doomed run. Every later
branch sits behind that gate, which is why the retry budgets, the
escalation, and the bounded terminalization were **unreachable rather
than mistargeted**.

This is precisely the failure the program's problem frame predicts: *a
guard that stops matching disables a recovery path invisibly and the
suite stays green.* I found it because my first renamed-column test for
the escalation site could not reach the escalation code at all.

Two further sites misbehave once the gate is passable:

- **FN-7998 node escalation** wrote `column: "todo"` inside the atomic
claim — parking the card where no workflow declares it, which is on the
plan's **"Stop implementation if"** list and what R7 exists to clean up
after. The scheduler's effective-node resolution, the entire point of a
node escalation, never runs.
- **FN-7863/FN-7926's `live.column === "todo"` arm** is the classic
guard that stops matching. In-process the `executeNodeSelfRequeued`
marker covers the same case, so this degrades only on the **durable**
arm — after a restart, or for a second `TaskExecutor` instance in the
process, where the column read is the only evidence the inner executor
requeued. A progressing card then falls through to the terminal sink and
is parked `failed`.

## The fix

Resolve hold and wip **once per graph failure** through U1's
`resolveTaskLifecycleColumns` and thread the pair through the ladder.
Both fall back to the legacy literal when the workflow cannot be
resolved, so an unresolvable workflow keeps exactly its pre-conversion
behavior rather than guessing. One IR read on a terminal recovery path —
not an enumeration loop.

## Red-green, measured

**3 of the 8 new tests fail with this commit's executor change
reverted:**

```
FAIL  FN-7998 … > requeues a node escalation to the RENAMED hold column, not the literal todo
FAIL  FN-7998 … > still does not move the card for a MODEL-target escalation
FAIL  FN-7863/FN-7926 … > recognises an inner-executor requeue that landed in the RENAMED hold column
      Tests  3 failed | 5 passed (8)     ← reverted
      Tests  8 passed (8)                ← with the fix
```

The other **5 pass both ways by design**, and I am not claiming them as
red-green — they are the regression floor:

- default coding workflow still resolves hold → `todo`, wip →
`in-progress` (byte-identical);
- an unresolvable workflow still uses the legacy literals;
- the in-process self-requeue marker still works when no workflow
resolves;
- and a **negative case** proving the dispatch-loop gate stays narrow —
a card still in its wip column with no marker is a genuine execute
failure and must NOT be swallowed as a benign recovery. Widening that
gate to "any column" would have been the easy wrong fix.

## Scope

Deliberately the execution-policy ladder only. **20 further column
literals remain in the same method's pause-abort, merge, and in-review
regions** — they belong to U5's executor slice (B4, not started) and
U9's merge lane, and are untouched here. Flagging the overlap: this PR
edits `executor.ts`, so whoever takes U5-B4 should rebase onto it rather
than converting these 9 sites again.

## Verification

- 8 new tests + the preserved-behavior suites
(`executor-tool-failure-retry`, `executor-graph-requeue-gate`,
`executor-task-done-blocked`, `executor-graph-boundary`,
`executor-stuck-requeue-preserve-progress`,
`executor-paused-abort-todo-benign`, `executor-abort-provenance`) — **9
files, 112 tests, green**
- `pnpm test:gate` — green (2/10, 16/299, 1/71); `pnpm lint` clean; `tsc
--noEmit` on `@fusion/engine` clean
- Changeset included (`patch`, category `fix`), passes `pnpm
check:changesets`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed execution recovery for workflows with renamed lifecycle columns
so retry, escalation, and loop-protection behaviors correctly follow the
workflow’s declared hold/WIP columns.
* Preserved legacy behavior for default workflows and continued safe
handling when lifecycle columns can’t be resolved.
* **Tests**
* Added a Vitest suite validating execution-policy “ladder” behavior for
renamed columns, including node escalation, dispatch-loop gating, and
fail-closed scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:47 -07:00
gsxdsm
2934cccad8 U7 PR2: finalize reports what it did with the card — a refused planning handoff is retried, not counted as recovered (#2498)
## The bug

`finalizeApprovedTask` has ~25 exit points and returned `void`, so no
caller could tell *"the card was handed off"* from *"finalize gave up"*.
Both callers assumed success.

`recoverApprovedTask` returned `true` **unconditionally** after
finalize, and `handleStuckAbortRequeue` treats `true` as "recovery done,
stop here". So when the release move was **refused by the planning-stage
guard** (FN-8361), or the store could not perform the move at all,
recovery reported success and the card's stuck-retry budget was skipped
— nothing re-planned it, nothing escalated it, and it sat in the planner
column holding a finished spec.

The refusal was already logged loudly by FN-8596's visibility work. The
return value was the part still lying.

## Three states, not a boolean

This is the load-bearing decision in the PR:

| Outcome | Meaning | Retry? |
|---|---|---|
| `released` | crossed into the hold column, or already resting there
(plan-in-place) | n/a — handed off |
| `parked` | deliberate, terminal-for-now: awaiting manual plan
approval, duplicate decision, operator pause, deleted duplicate | **no**
— a human owns it |
| `withheld` | finalize could not complete the handoff, nothing waiting
on a human | **yes** — caller's budget owns it |

`recoverApprovedTask` returns `outcome !== "withheld"`, so **`parked`
still returns `true`**. Narrowing to `=== "released"` is the tempting
simplification and it is wrong: it would send the stuck handler down its
draft path and stamp `needs-replan` over a plan a human is mid-review on
— a worse bug than the one being fixed. That is asserted, and the
assertion fails under exactly that narrowing.

## Why a mutable report, not a return at each exit

Threading a return through 25 exits is 25 chances to mis-classify a
branch, and mis-classifying turns a truthfulness fix into a lifecycle
bug. The report defaults to `parked`, which is equivalent to today's
observable behavior at every exit — so the plumbing is **inert
everywhere except the three sites explicitly classified**. Adding a
state to an exit is then a deliberate, reviewable act rather than a
diff-wide judgement call.

Only **two** exits are marked `withheld`, both in the release block,
both already warning loudly. Deliberately *not* marked:

- the `updatePlanningStateIfStillCurrent` guard — FN-8024 says a normal
scheduler advance legitimately lands there; the card has moved on, so a
retry would be wrong.
- `recoverMissingPromptBeforeRelease` — it owns its own recovery budget;
retrying would double up.

## Revert proofs (measured)

| Reverted | Result |
|---|---|
| `recoverApprovedTask` back to unconditional `true` | `Tests 2 failed
\| 3 passed (5)` |
| narrowed to `outcome === "released"` | `Tests 1 failed \| 4 passed
(5)` — the approval-park control |

The second row is the point: the park case is load-bearing, not
decoration.

## A fixture note that nearly produced a false green

A `vi.fn()` stub for `updateTaskAtomic` that ignores its callback makes
**every** finalize report "no longer in the planning stage" and return
before the release — silently collapsing every case into the same
uninteresting early exit. My first run was 3 failures for that reason,
not the reason I expected. The fake now applies the patch, and the
control asserts `moveTaskIf` was actually reached. Same class as the
`moveTaskIf` fake caught on #2491; recording it so the next person
recognises the shape.

## Scope

The other caller — `specifyTask`'s unconditional `onSpecifyComplete` —
is **not** gated here. Reaching it needs a live planning session, so
gating it without first extracting the reaction would be a change I
cannot prove, which is exactly the finding review caught on #2491's
deferral. That lands next, on this plumbing.

## Verification

| Check | Result |
|---|---|
| new suite | 5/5 |
| 11 triage/planning suites (triage, finalize-duplicate-lineage,
stuck-requeue-preserve-draft, explicit-duplicate-marker, preflight,
plan-artifact-writeback, refinement-routing, planning-wake,
planning-evacuation, …) | 326/326 |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (299 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:36:20 -07:00
gsxdsm
8aba310d78 U11: resolve the worktree-acquisition requeue column by trait (2 sites, both branches) (#2496)
Based on `main`. First of my U11 conversion PRs — small, green,
independently revertable.

Both heartbeat worktree-acquisition requeue sites hardcoded `"todo"`.

## Why this is critical path, not a renamed-workflow nicety

**U11 deletes the `todo` column from the builtin workflows.** After
that, these two sites would requeue every acquisition-failed card into a
column that no longer exists.

## Both sites converted together

They are different branches of the same failure:
- the **bounded-retry** requeue, and
- the **retry-cap-exhausted** terminal park.

Converting one and not the other would leave the rarer path — which
fires only after three consecutive failures, so it's the one least
likely to be noticed — still writing the literal.

Target is the KTD-10 ordering via `resolveReboundTarget` (hold → intake
→ first column): the same helper `self-healing` and `mesh-lease-manager`
already use for "requeue a recovered card", so the recovery paths cannot
drift apart.

## What is deliberately untouched

`preserveStatus: true` on the exhausted path. It exists because
reopen-to-todo semantics would otherwise wipe the `status: "failed"`
written immediately before (FN-7721) — changing the column must not
disturb that flag. A test asserts the full options object, not just the
column.

Fail-soft to the legacy id: a requeue must not be abandoned because a
workflow lookup failed, or the card is left holding a worktree it could
not acquire. Covered by a regression-floor test.

## Verification

- **Mutation-verified:** restoring the literal fails 2 of the 3 new
tests
- 7 tests green (3 new + the 4 pre-existing worktree tests, unchanged)
- tsc clean, lint clean, merge gate green (299 + 10 + 71)

## Measured progress

**2 of the 74** code-level `"todo"` sites in my unit (engine
recovery/scheduling core) are now trait-resolved.

Remaining in-unit: `self-healing` 48, `scheduler` 15, `triage` 8,
`replan-target` 1.
`stuck-task-detector` needs **no work** — all 4 of its occurrences are
comments, not code.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Tasks now return to the workflow’s configured hold column when
heartbeat worktree acquisition fails, including workflows that use a
renamed hold column.
* Retry and retry-limit handling now preserves task progress and, when
applicable, status.
* Added a safe fallback to the default “todo” column when workflow
details cannot be resolved.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-28 11:36:13 -07:00
gsxdsm
8492278fdd U11 PR1: pin the merged intake+hold column contract before the IR moves (a mutation proved the first 7 tests insufficient) (#2495)
First of several small PRs for **U11** (merge Todo into Planning).
**Tests only — no production change.** It lands the precondition so the
IR edit arrives on proven substrate instead of an assumption.

## Decision taken (reversible, proceeding on it)

**The surviving Planning column keeps the id `todo`; `triage` is
deleted.** Same board the operator asked for — one column labelled
"Planning", no "Todo" — via the cheaper and safer half.

Measured, comments excluded, non-test, `packages/*/src` +
`dashboard/app`:

| | guards | writes | fallbacks | total |
|---|---:|---:|---:|---:|
| `"todo"` | 121 | 68 | 9 | 323 |
| `"triage"` | 90 | 30 | 21 | 304 |

Deleting `triage` instead of `todo` also means **no data migration**
(every live card in `todo` is already in the surviving column) and **no
guard changes meaning** (`column === "todo"` still denotes the hold
column). Under the plan's letter the opposite is true, and worse than
"dead": because Coding (Ideas) keeps `todo` per R10/R11, a surviving
`column === "todo"` guard would stay live for Ideas cards while silently
never matching for Coding cards — workflow-dependent, not dead.

This is also a proven in-tree pattern rather than a new idea:
**`builtin:coding-ideas` already ships this exact merge** — id `todo`,
display name "Planning", `hold(capacity)` + `reset-on-entry`,
plan-in-place.

Consequence worth flagging: **U11 no longer waits on Phase B.** The 121
`todo` guards keep their meaning, so converting them becomes U12 cleanup
rather than a U11 blocker.

## What this PR pins

Nothing in tree has ever carried `intake` and `hold` on one column.
Every built-in splits them. KTD-1 asserts the merged shape works; that
assertion was untested.

## The result, reported as found

**All 11 assertions passed on the first run against unmodified
sources.** The merged column is already supported by trait resolution,
the capacity sweep, and the release gate. **I could not make the first
seven fail**, so they are a regression floor — not evidence of a fix,
and I am not claiming them as one.

What makes them worth keeping is that they are *differential*: the same
scenario runs against the split-role vocabulary and the merged one and
asserts the role-level outcomes are **equal**, so a literal creeping
into any path fails the merged half while the split half stays green.

## The finding

**The first seven tests were not enough, and proving that is the point
of this PR.**

A mutation encoding the plausible-but-wrong belief *"an intake column
has no releaser"*:

```diff
- if (currentFlags.intake !== true && currentFlags.hold !== true) return false;
+ if (currentFlags.intake === true) return false;
+ if (currentFlags.hold !== true) return false;
```

left **all seven green**.

That belief is not hypothetical — it is stated verbatim in
`builtin-plan-review-group.ts`'s own FNXC comment as the reason Plan
Review lives in `todo` rather than `triage` today. Under U11 the
planning column **is** an intake column, so any code encoding it
silently stops holding unplanned cards and they release into
implementation with a bootstrap stub for a spec.

The gap: nothing reached `isUnplannedForExecution`. The mock store had
no `getTasksDir`, so both halves of the gate returned early — the tests
were exercising less than they appeared to. The fourth block drives it
with a real temp dir and a real bootstrap `PROMPT.md`.

**Re-running the same mutation now fails exactly one test — the
merged-column one — while its split-shape twin stays green.** That
discrimination is what the suite is for.

## Verification

26 tests green across this file plus `hold-release-renamed-columns`,
`hold-release-instrumentation`, and `pre-release-plan-review`. Lint
clean. No production file touched, so there is nothing to regress.

## Next PRs in this unit

1. Entry-contract test for `start` in a hold-carrying column (the
specific interaction the earlier, reverted attempt got wrong).
2. The ~10-line IR change itself — deliberately last, per KTD-7.
3. The intake-lane `triage` conversion: the 21 `?? "triage"` creation
defaults are the dangerous ones, since they would silently create cards
into a column that no longer exists.

`self-healing.ts` (11 triage guards + 4 writes) is the main worker's
file — I am not touching it and will hand over the line list rather than
race them.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added coverage for merged Planning column behavior across split and
merged workflow configurations.
* Verified intake and hold resolution, rebound targeting, capacity
hold/release outcomes, and execution gating.
* Confirmed planned cards are released appropriately while cards already
in progress are not unnecessarily held.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:36:05 -07:00
gsxdsm
fbe7eb5c5a U7 PR1: the manual plan-approval gate was bypassable (3 planning-lane surfaces, 8/13 revert-proof) (#2491)
## What this is

The first slice of **U7 — the graph owns planning**. Characterizing the
planning lane's dual ownership turned up a live defect in the exact seam
the unit exists to remove, so this PR fixes that first and reports the
measured map of what U7 still has to move.

## The defect

The manual plan-approval gate parks a card by writing `status:
"awaiting-approval"` and **returning early** from `finalizeApprovedTask`
— before the release move. `specifyTask` then calls `onSpecifyComplete`
**unconditionally** afterwards. Three automated surfaces went on to
advance the parked card, each having re-derived its own weaker "may I
advance this?" check from `paused`/`userPaused` alone.

`isTaskBlockedOnApproval` (`packages/core/src/task-merge.ts`) already
declares itself *"the single shared predicate core and engine code must
consult before rebounding, requeuing, resuming, re-planning, or
otherwise advancing a task"*. **Measured: it had exactly one production
consumer** (`overseer-human-control-policy.ts`). Now four.

Reachable end to end for a **plan-in-place** card — one whose column
already equals the plan-review node's column (Coding (Ideas), or any
`needs-replan` revision resting in the default workflow's `todo`):

```
park at awaiting-approval
  → onSpecifyComplete fires anyway
  → a runnable plan-review continuation is seeded
  → the drain dispatches it
  → Plan Review runs on a plan the operator never approved
  → its evidence satisfies isUnplannedForExecution
  → the capacity sweep releases the card into In progress
```

Blast radius: projects that have manual plan approval switched on.
`planApprovalMode` defaults to auto-approve (FN-7557), so unset projects
have no gate to skip — but the operator who turns it on is precisely the
one who cares.

## Surface enumeration

Per AGENTS.md — fix the invariant, not the repro.

| # | Surface | Fix |
|---|---|---|
| 1 | `issueRelease` — the choke point for the sweep, `promoteHeldTask`,
`releaseHeldTaskByEvent`, and the scheduler's `reserveSlot` guard |
Guarded there rather than inside `isUnplannedForExecution`, because an
approval-held card is not "unplanned". Guarded **again** inside the
`moveTaskIf` predicate so a park landing mid-sweep cannot lose the race
(R6 — only the in-txn check is authoritative). Operator force-promote
(`allowUnplanned`) still waives it: that *is* a human decision about
this card. |
| 2 | **Both** continuation seeders —
`seedPreReleasePlanReviewContinuation` (normal completion) and
`evaluateStrandedHoldContinuation` (FN-8592 self-healing re-seed) |
Guard at the seam, not in the callers: the seeder itself checked
nothing, and its two callers each pre-checked a different subset. |
| 3 | `resolvePlanningContinuationCandidate` (drain classifier) |
**Skip, never orphan.** Cancelling terminalizes the item, so an approval
landing a minute later would have nothing left to resume and would need
a second repair to come back. |

## Measured, not assumed

The two hold shapes `isTaskBlockedOnApproval` accepts were **not equally
broken**. The `paused` + `pausedReason` shape was already refused by the
sweep and the drain — they happen to test `paused` — so it was refused
*for the wrong stated reason*, not advanced. Every genuine advance gap
is on the **status-only** shape, which is exactly what the gate writes.
Both are covered anyway, plus an `ORDINARY_PAUSE` counter-case so the
new check cannot quietly become a catch-all for every operator park.

## Revert proof

With the three production files reverted: **8 of 13 tests fail.** The 5
that still pass are the 3 controls and the 2 pause-shape rows the
pre-existing `paused` checks already covered.

```
·x··xxxxx·xx·      → Tests 8 failed | 5 passed (13)
```

## Verification

| Check | Result |
|---|---|
| new suite | 13/13 |
| hold-release (×2) + plan-review (×3) + pre-release-plan-review +
promote-force-unplanned | 43/43 |
| stranded-hold-continuation (×2) + continuation-selection +
planning-finished-wake + planning-service | 27/27 |
| scheduler-trait-dispatch | 9/9 |
| `pnpm --filter @fusion/engine exec tsc --noEmit` | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green |
| `pnpm check:changesets` | clean |

## Two findings for the coordinator

**1. `triage.ts` is absent from the Phase B census.** The plan's
per-file table (535 sites) covers `self-healing.ts` (U4), the
executor/scheduler cluster (U5), and the core policy modules (U6).
`triage.ts` appears in none of them, so its lifecycle-column literals
are unowned scope — U7 absorbs them.

Measured with the plan's own methodology (block and line comments
stripped, code lines only): a naive quoted-literal grep of `triage.ts`
reports **50** sites, but **35 of those are the agent *role* string
`"triage"`**, not the column. The genuine lifecycle-column surface is
**15 sites**, of which 12 are planning-lane and 3 are `column !==
"done"` in duplicate search. The 50 figure would over-count by 3.3×.

**2. The graph's planning seam is a rubber stamp, in triplicate.**
`createAuthoritativeWorkflowSeams().planning` returns `{ outcome:
"success", value: "pre-specified" }`;
`WorkflowPlanningService.runPlanningSession` returns the same;
`createNoopLegacySeams().planning` is a bare success. The real
specification is ~1,000 lines of `triage.specifyTask`, entirely outside
the graph. That is the flip U7's remaining slices have to make, and it
is the reason the planning lane has two owners at all.

## Deliberately not in this PR

Triage's unconditional `onSpecifyComplete` call. That is the
**ownership** half — `finalizeApprovedTask` must report whether it
released, and the reaction must key on that outcome — and it belongs
with the seam flip, where finalize's outcome becomes the graph's edge
condition anyway, rather than as a half-measure now. With the three
guards above in place, the downstream damage is already contained; what
remains is a reaction firing for a non-event and an operator-visible log
line (`Specified X → todo`) that is untrue for a parked card.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks awaiting manual plan approval are no longer automatically
planned, reviewed, started, or released into active work.
* Approval-held items are consistently skipped across planning
continuations and related workflows.
* Approval-held due work is deferred to prevent starvation while
waiting, and operator force-promotion still bypasses the gate.
* **Tests**
* Added regression coverage to ensure the manual approval hold behavior
remains invariant across multiple continuation scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:39:21 -07:00
gsxdsm
319e051c65 U8 PR1: pin the execution-lifecycle ownership ledger (measured: 28 executor-owned dispositions vs 3 graph handbacks) (#2490)
First PR of **U8 — the graph owns execution** (plan
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`,
line ~436). The plan states this unit "is expected to land as several
commits; it must not be attempted as one sweep", and
Execution-note-first: **characterization before ownership moves**. This
is that floor. **No behavior change.**

## Why a ledger and not a refactor

U8's goal is "the executor stops deciding *what happens next*" — and
that had no measurable form.

- **Executor line count does not measure it.** A 3,178-line
`runImplementation` can shrink substantially with every lifecycle
decision still exactly where it was.
- **A green suite measures it least of all.** Every disposition counted
below already has passing tests, because each one was *correct behavior*
when it was written. What is wrong is the **owner**, not the behavior.

So the unit needs a number, and the number has to exist *before* the
migration — a ratchet written afterwards cannot prove the migration
happened.

## The measured baseline

Counted from source, comments stripped, method bodies extracted by brace
matching:

| Method | `store.moveTask` | `handoffTaskToReview` | terminal
`status:"failed"` | `graphCompletion` handbacks |
|---|---:|---:|---:|---:|
| `runImplementation` (3,178 lines) | 16 | 3 | 9 | **3** |
| `handleGraphFailure` (~930 lines) | 0 | 0 | 7 | — |

**The implementation phase decides its own lifecycle 28 times and asks
the graph 3 times.**

These are measured, not estimated. My first `handleGraphFailure`
estimate was **wrong** (2 moves / 4 parks); the extractor corrected it
to 0 / 7 — the `moveTask` calls that read as belonging to that method
sit past its closing brace, in the recovery helpers below it. The
correction is in the ledger comment so the next reader does not repeat
the misread.

## The finding this makes concrete

`createAuthoritativeWorkflowSeams.execute` collapses that entire
implementation phase to one boolean:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
```

The graph has no vocabulary for *"the agent stopped because a step is
blocked on a pending review"* or *"the session paused after the work was
already complete"*. So the implementation phase performs those
transitions itself (`executor-exit-while-review-pending`,
`paused-after-completion`) and the graph finds out afterwards.

That is why `handleGraphFailure` carries `alreadyFinalizedToReview` /
`completionFinalized` — **classifiers whose entire job is to recognise a
move the graph did not make.** They are compensation for dual ownership,
and they are U8's acceptance test: they become unreachable, and then
deletable, exactly when the last out-of-band transition is gone. This PR
records that contract in source at the seam (FNXC comment), which is
where the next PR starts.

## Proof the guard fails on the defect

A ratchet that reports success without checking anything is worse than
no ratchet. Both failure modes were injected and observed:

1. **The defect it exists to catch** — injected one `await
this.store.moveTask(task.id, "in-review", {})` into
`runImplementation`'s completion path → ledger fails, `16 -> 17`.
2. **A broken guard** — injected a string literal containing `}` so
naive brace matching ends the body early → the size self-check fails at
**13 lines**, instead of silently reporting a comfortable zero for every
count.

Both injections were reverted; `git diff` against the pre-injection copy
is empty.

## Direction of travel

Executor-owned counts may only go **down**, and a decrement must land
with the disposition visible as a **graph outcome** — not merely
deleted. An increment is a new out-of-graph lifecycle decision and needs
a stated justification in its PR, not a quiet edit to the constant.

This is the precursor to U12's planned
`no-out-of-graph-lifecycle-writes.test.ts`; when the counts reach their
floor the assertion becomes "zero, outside the allowlist", and this file
is where that allowlist grows up.

## Preserved behaviors

Untouched, and re-run green as the regression floor for everything that
follows: FN-8141 honest-blocked exit
(`executor-task-done-blocked.test.ts`), FN-7996/FN-7998 tool-failure
retry + escalation (`executor-tool-failure-retry.test.ts`), FN-7863
dispatch-loop terminalization and FN-7926 completed-blocked parking
(`executor-graph-requeue-gate.test.ts`).

## Verification

- `pnpm --filter @fusion/engine exec vitest run` on the ledger + the
four preserved-behavior suites + `legacy-tombstones` — **6 files, 49
tests, green**
- `pnpm test:gate` — **green** (2/10, 16/299, 1/71)
- `pnpm lint` — clean; `tsc --noEmit` on `@fusion/engine` — clean

No changeset: test-only plus a source comment, no `@runfusion/fusion`
behavior change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added a lifecycle-ownership “source-scanning” test that analyzes the
executor’s task disposition patterns to ensure counts remain consistent
across execution and graph-failure flows.
  * Added safeguards to catch unintended changes to lifecycle handling.

* **Documentation**
* Documented the lifecycle-ownership boundary for task disposition
handling, including how completion and failure transitions are
consolidated and how related failure classifiers are affected.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:59 -07:00
gsxdsm
7871b28766 fix(core): bind the in-transaction capacity gate — one shared pool-id convention (NOT user-visible yet — see R2) (#2488)
## The bug

`moves.ts` asked `countActiveInCapacitySlotAsync` for occupants of pool
`"builtin:coding"`, while the counter buckets selection-less rows under
`DEFAULT_WORKFLOW_POOL_ID` (`"__default-workflow__"`). Nothing ever
landed in the pool being asked about, so the count came back **0** and a
finite limit could never bind.

## Root fix, not a literal swap

A shared *constant* would not have prevented this:
**`DEFAULT_WORKFLOW_ID` was already imported in `moves.ts` and the code
still wrote a literal.** So both sides now call a shared **function**,
`resolveCapacityPoolId` — "which pool does a selection-less task belong
to" has exactly one answer and no call site is in a position to disagree
with it.

The one variable serving two masters is split: a capacity **pool key**
(a bucketing sentinel that must not collide with a workflow id) and a
**workflow id** (telemetry, must stay a real id). The emitted
`TaskTransitioned` payload is byte-identical.

## Checked, not assumed: no second copy

`scheduler.ts:2514` and `:2536` do carry `?? "builtin:coding"` — but as
an **IR resolution key** (`resolveWorkflowIrById`), where a real
workflow id is required and the pool sentinel would not resolve at all.
Same literal, different concept, correctly used. A blanket replace would
have broken it.

## Something did depend on the gate being dead — exactly one thing

`move-path-equivalence.pg.test.ts` → *"UNPROVEN: in-transaction column
capacity did NOT reject on EITHER path in this fixture"*. It left the
cause open —

> something further in (`resolveColumnCapacity`'s limit resolution, or
what `countActiveInCapacitySlotAsync` counts as an occupant — a task
with no session/agent may not count) keeps the check from firing … This
suite does not establish which.

— and predicted its own obsolescence (*"if a future change makes this
reject, that is the capacity gate coming alive"*). **Neither guess was
right; it was the pool id.** Updated to assert the divergence with the
answer recorded — **not weakened**. Its fixture also had to start each
phase from an empty wip column: once the gate binds, the inline phase's
leftovers trip the cap on the *holder* move before the contended move
under test runs.

`schema-applier.test.ts` failed only in the full-suite run and passes in
isolation both with and without the fix — cross-file contamination, not
mine.

## Before / after — measured, both directions

`maxConcurrent: 1`, real PG store, real `moveTask`:

| | flagOFF / no selection | flagOFF / selection | flagON / no selection
| flagON / selection |
|---|---|---|---|---|
| **before** | ADMITTED | ADMITTED | **ADMITTED** ← the bug | REJECTED |
| **after** | ADMITTED | ADMITTED | **REJECTED** | REJECTED |

The E2E acceptance row asserts **held at cap 1 and admitted at cap 2 on
the same fixture**, so it cannot pass by simply never admitting
anything. **With the fix reverted that row fails**; the `admitted` case
still passes, as it should. The Phase A3 ratchet's two flipped
assertions also fail with the fix reverted.

Ratchet flipped exactly as its author specified: `DEFECT (R1)` becomes a
rejection, and `it.fails` on the invariant becomes a plain `it`.

## ⚠️ This is NOT user-visible yet — please read before merging

The premise this was approved on ("once it binds, cards that currently
slip through will start being held") **does not hold for this change
alone.** The whole capacity block sits inside `if (useWorkflow &&
workflowIr && fromColumn !== toColumn)`, and `useWorkflow` is
`experimentalFeatures.workflowColumns === true` — absent from
`DEFAULT_GLOBAL_SETTINGS`, with **no writer anywhere outside tests**.
That is Phase A3's R2, still live and now retitled `DEFECT (R2, STILL
LIVE)` with the measured matrix recorded in it.

So on merge: nothing changes for any real project. Making it actually
bind means **also** removing the `useWorkflow` condition — a materially
larger, genuinely user-visible change that I have not made unilaterally.
Escalated for a decision; if that lands, the changeset here should be
re-categorised.


## Review follow-up (48e79ffd9): the convention was still duplicated —
swept and ratcheted

The first pass added the resolver and routed the transactional gate +
counters, but **hold-release still derived the pool independently**.
Swept the repo: six sites name the sentinel, **five derive the
convention** and now call `resolveCapacityPoolId`
(`hold-release.ts:116/118/442/576`, `task-store-helpers.ts:290`). The
sixth, `scheduler.ts:1558`, names the default pool as a literal in a
capacity *diagnostic* — no selection input, nothing to disagree with —
so it keeps the constant.

**Does this change hold-release behavior? No, and it was never releasing
against the wrong pool.** hold-release computed `x ??
DEFAULT_WORKFLOW_POOL_ID`, which is exactly what the counter buckets
under; `moves.ts` (`?? "builtin:coding"`) was the sole disagreeing site,
and the first commit moved *it* into agreement with hold-release, not
the reverse. `resolveCapacityPoolId(x)` **is** `x ??
DEFAULT_WORKFLOW_POOL_ID`, so every routed site computes an identical
value for every input. **No second user-visible change rides along with
this PR** — the only behavior delta remains the gate binding on the
flag-ON path, which per R2 is still not the path production takes.
Evidence: hold-release + capacity suites **43/43 identical before and
after**.

**The resolver is now the only way to compute a pool id, not merely the
newest way.** `scripts/check-capacity-pool-id.mjs` fails on any inline
`?? DEFAULT_WORKFLOW_POOL_ID` outside `workflow-capacity.ts`, wired into
**both `pretest` and the blocking `test:gate`**. A review note would not
have sufficed: the original defect landed in a file that *already
imported* the canonical constant. Verified both ways — clean run scans
1124 files and passes; reintroducing the old hold-release expression
exits 1 and names the line.


## Review follow-up (a5b675503): the ratchet was rebuilt because it
would not have caught the bug

The first ratchet matched one spelling (`?? DEFAULT_WORKFLOW_POOL_ID`)
and the real defect used another (`?? "builtin:coding"`). **Verified:
reintroducing the original defect and running the old checker exits 0.**
A guard that reports success without checking is worse than no guard —
it stops anyone looking.

Rebuilt on the TypeScript AST with two rules. **Rule 1 (sink):** a value
reaching a capacity counter's `workflowId` must come from
`resolveCapacityPoolId`, or a local initialized from it — so it fires on
the original defect regardless of which literal was used, on one line or
twenty. **Rule 2 (sentinel):** no `??` onto the sentinel at any
qualification depth or as its raw value; multiline is one AST node and
caught by construction. `?? "builtin:coding"` is deliberately *not*
banned outright — it is the legitimate default for a *workflow* id in ~8
places, and is only a bug when it reaches a capacity pool.

**Fails closed three ways** that previously reported success without
inspecting: unreadable file, unparseable file, and an empty file listing
(the old script would have printed a green tick off a broken glob).

**Acceptance was not "passes on main".** Each form was reintroduced into
the real source and confirmed to fail: the original defect in
`moves.ts`, a multiline fallback, and a deeply qualified sentinel. All
are pinned in `capacity-pool-id-check.test.ts` (12 cases: 7 must-catch
starting with the reduced actual pre-fix `moves.ts`, 4 must-not-flag, 1
fail-closed) so the guard cannot silently narrow again.

Also added to `pretest:full`, which had omitted it.


### Follow-up (0be8df6ea): a dead rule found by fixing a test title

Splitting the mislabelled fail-closed test surfaced more than a
mislabel: **`ts.createSourceFile` is error-tolerant and does not throw
on malformed syntax**, so the `try/catch` behind the `unparseable` rule
was unreachable and that rule could never fire. The earlier "fails
closed three ways" claim was overstated — the guard advertised a
capability it did not have. Detection now reads `sf.parseDiagnostics`; a
partial AST can silently lack the `??` nodes and sink calls the rules
look for, so "did not parse" must not read as "inspected and clean".
Mutation-verified: reverting the detection fails that case and only that
case.

Test-file exclusion also moved to the repo's `{test,spec}.{ts,tsx}`
guideline shape — a `.spec.ts` under `packages/<pkg>/src/` was being
scanned as production source. Verified both ways: the `.spec.ts` is
skipped, and the identical content in a non-test file is still caught,
so the exclusion is scoped rather than a hole.

## Verification

- engine + core `tsc --noEmit` clean
- `pnpm test:gate` green (299 + 10 + 71)
- E2E 20/20; capacity + move-path suites 14/14
- full core PG: **1037 passed / 3 failed** — all three reproduce with
the fix stashed (pre-existing)
- engine-default: **279 failed** vs **280 at baseline** with the fix
stashed — pre-existing red lane, no regression
- hold-release + capacity suites: **43/43 identical before and after**
the resolver routing
- `check-capacity-pool-id` ratchet: 14/14 regression cases; clean over
1124 files; exits 1 on the original defect, a multiline fallback, and a
deeply qualified sentinel reintroduced into real source

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed capacity-limit accounting when workflow selection is missing by
consistently deriving the correct capacity pool id.
* Made capacity enforcement align across move and hold/release paths,
rejecting over-limit moves with `capacity-exhausted`.
* **Tests**
* Updated PostgreSQL and added an E2E scenario to verify the corrected
in-transaction gating behavior at `maxConcurrent` limits of 1 and 2.
* **Chores**
* Added an automated guard to detect inconsistent capacity pool id
fallback patterns in code.
* **Public API**
* Exposed `resolveCapacityPoolId` for consistent capacity pool id
derivation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:51 -07:00