Commit Graph

3691 Commits

Author SHA1 Message Date
gsxdsm
5795d70b27 fix(engine): assignment load must be resolved per task — #2787 P1 follow-up (#2796)
Fix-forward for the P1 that arrived on **#2787 after it merged** — so it
lands as its own PR rather than a thread reply on merged code.

## The finding

`selectPermanentAgentForTask`'s `activeColumns` was resolved from the
**candidate** task's workflow and then applied to every row `listTasks`
returned. On a project running several workflows — the normal case —
assignments living in another workflow's load-bearing lanes vanished
from the tally, and the already-loaded-agent-wins bug returned through a
different door.

**A column id means something only relative to its OWN workflow.**
`blocker-fanout.ts` documents exactly this and offers a per-task
`classify`; the option is now that same shape rather than a third
invention:

```ts
countsAsAssignmentLoad?: (task: Task) => boolean
```

The scheduler resolves each assigned row against its own IR, sharing one
cache for the selection, so a board spanning three workflows reads three
IRs — not one per assigned card.

## Why this is the third round on the same parameter, stated plainly

1. I added the parameter and **never wired the caller** — inert in
production.
2. I wired it as a **union of wip+review**, which dropped hold/intake
and made it a *regression* for backlog work.
3. I resolved it from **one workflow** and applied it to all — this fix.

Each round was a smaller version of the same error: treating a lane
answer as global when it is per-task, and per-role when it is
per-membership. Worth recording because the first two rounds both looked
correct and both passed their tests — the tests asserted the renamed
case I was thinking about, not the shape of the data.

## Verification

- new cross-workflow case; reverting the predicate to a single
workflow's lanes **fails it**
- `agent-assignment` suite **14 passed**
- `pnpm test:gate` — **161 / 13 / 487 / 71** · lint clean · census
`--strict` exits 0

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 11:06:20 -07:00
gsxdsm
6bb5e4f787 test(engine): measure the optional-role-parameter conversion class (#2798)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Follows #2795, which
found the first instance of this pattern.


`packages/engine/src/__tests__/workflow-optional-role-param-caller-audit-live-e2e.pg.test.ts`

## The finding

#2795 showed a conversion pattern the lifecycle-column census cannot
see: a role question migrated into an **optional parameter whose default
is the legacy literal**, converted at some call sites and not others.
This shows it is not a one-off, and measures it.

| seam | call sites passing the resolved answer |
|---|---|
| `shouldHoldActiveFileScopeLease` | **2 of 4** (both `scheduler.ts`;
neither `self-healing.ts`) — #2795 |
| `evaluateParkedAgentTaskLink` | **2 of 6** (`scheduler.ts`,
`task-agent-sync.ts`; neither `agent-heartbeat.ts` ×2 nor
`self-healing.ts` ×2) — this PR |

The second is the more damaging, and the callee's own FNXC note already
names the outcome: without the resolved columns "the card would be
treated as unparked and its live agent link cleared" — **a stale-link
bug turned into a dropped-link bug**. Driven here: a card parked in a
renamed board's hold column, with live execution proof, has its agent
link dropped.

### Why the census is blind to it

The callee is converted and its default is correctly marked
`DELIBERATE-LITERAL` — for an unconverted caller that default genuinely
*is* the intended behaviour. **The unconverted call sites contain no
column literal at all**; it lives one function away. So the census
counts the callee's annotated literals and sees nothing at the call
sites, and the conversion reads as complete from every angle except
running it.

This is a *class*, not two bugs. The same shape exists at roughly twenty
seams (`revertableColumns`, `plannerColumns`, `roleColumn`,
`terminalColumns`, `activeColumns`, …). Two are now measured. I checked
two others I flagged as unknown in #2795 —
`restart-recovery-coordinator.ts`'s `isReviewColumn?` and the
`isRecoverableMissingWorktreeReviewFailure` family — and **their callers
are fully converted** (`extension.ts:1924`, `task.ts:1390`,
`self-healing.ts:12087`), though the doc comment claiming `extension.ts`
"still asks with the literal" is now stale. The rest are unaudited; the
audit case is written so adding a seam is a small edit.

## Scope, stated honestly

Three cases are driven end to end: real persisted rows from a live
store, the real exported predicate, both call shapes. The **call-site
split is asserted against source text** — reaching all six sites needs
the heartbeat and self-healing harnesses, which I did not build, and the
audit case says so in its own comment rather than dressing it up.

It is an alarm in **both** directions: a new unconverted caller pushes
the count up and fails; converting an existing one pushes it down and
also fails. The second is deliberate — that is the moment someone should
read the three behavioural cases and update the number on purpose.

## Mutation-verified

Flipping the callee's default from the legacy parked pair to
`["backlog"]`:

| case | result |
|---|---|
| CONTROL (default board, no options) | **fails** |
| CHARACTERIZATION (renamed board, no options) | **fails** |
| BOUND (renamed board, options passed) | passes — correct, the argument
overrides the default |
| AUDIT | passes — correct, it is a source assertion |

## Not done, and why

**No fix.** Passing the resolved columns at the four unconverted sites
means resolving each linked task's traits inside the heartbeat and
self-healing paths — async work in loops that already hold locks — and
both files belong to other workers. The differential says exactly what
the fix should make true.

## Verification

- new suite — **4/4 passed**, mutation matrix above
- full live-PG E2E surface — **137/137 passed** (133 on main + 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 11:03:09 -07:00
gsxdsm
60054aab0a test(engine): live-PG evidence of an inert conversion at the CALL SITE (#2795)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit.


`packages/engine/src/__tests__/workflow-file-scope-lease-caller-gap-live-e2e.pg.test.ts`

## Why this is a different finding, not a sixth of the same one

#2789/#2791/#2792/#2793/#2794 all concern **one** mechanism: a site
resolves the workflow synchronously and silently gets the default board.
This is a **second** mechanism, and neither the lifecycle-column census
nor the sync-resolver allow-list can see it.

`shouldHoldActiveFileScopeLease` was converted by turning its two role
questions into optional parameters with literal defaults:

```ts
const isWipColumn    = options?.isWipColumn    ?? task.column === "in-progress";
const isReviewColumn = options?.isReviewColumn ?? task.column === "in-review";
```

A caller that resolved the traits passes the answer; a caller that has
not gets exactly the pre-conversion behaviour. That is a deliberate
migration device and the source says so — correctly marked
`DELIBERATE-LITERAL`.

**But the migration was only half made:**

| call site | passes the resolved answer? |
|---|---|
| `scheduler.ts:1986` | ✅ `{ isWipColumn: true }` |
| `scheduler.ts:2006` | ✅ `{ isReviewColumn: true }` |
| `self-healing.ts:4525` | ❌ neither |
| `self-healing.ts:5443` | ❌ neither |

So the same predicate is right on the scheduler's path and wrong on
self-healing's. The harm is the one the function's own FNXC note
describes: on a renamed board both branches fall through, the predicate
returns false for every card, `activeScopes` stays empty, and the
dispatch path sees no overlap — *two agents editing the same files*,
which is what the overlap machinery exists to prevent. At the
self-healing sites the consequence is narrower but identical in shape: a
stale-lease reconciler concludes a live blocker holds no lease and
proceeds to clear state the scheduler would have honoured.

### Why the existing instruments are blind to it

**There is no column literal at the self-healing call sites.** The
literal lives inside the callee's default, one function away — and there
it is correct, because for an unconverted caller it *is* the intended
behaviour. A census counting `=== "in-progress"` occurrences sees the
callee's two (properly marked) and nothing at all at the call sites. The
conversion reads as complete from every angle except running it.

This generalizes: **any conversion that migrates behaviour behind an
optional parameter leaves a residue the census scores as done.** Worth a
sweep for the same shape elsewhere — `agent-assignment.ts`'s
`activeColumns?` and `restart-recovery-coordinator.ts`'s
`isReviewColumn?` are the same pattern; I have not checked whether their
callers supply them.

## Scope, stated honestly

Three cases are driven end to end: real persisted rows from a live
store, the real exported predicate, both call shapes. The **call-site
fact is asserted against source text, not driven** — reaching those
sites needs the full dependency-lease reconcile harness, which I did not
build. The last case reads the file and says so in its own comment
rather than dressing it up as an end-to-end result. It doubles as an
alarm: when those call sites are converted it fails and points at the
three cases above, which describe exactly what changes.

## Mutation-verified

Flipping the callee's default from `"in-progress"` to `"building"`:

| case | result |
|---|---|
| CONTROL (default board, no options) | **fails** |
| CHARACTERIZATION (renamed board, no options) | **fails** |
| BOUND (renamed board, option passed) | passes — correct, the option
overrides the default |
| SOURCE-LEVEL | passes — correct, it is a source assertion |

The two default-dependent cases bind to the default; the bound case
proves the override; nothing passes for the wrong reason.

## Not done, and why

**No fix.** Passing the resolved answers at the two self-healing sites
requires resolving each blocker's column traits there — an async
resolution inside a reconcile path that already holds locks, and
`self-healing.ts` is another worker's file. Flagging with a differential
that says exactly what the fix should make true.

## Verification

- new suite — **4/4 passed**, mutation matrix above
- full live-PG E2E surface — **137/137 passed** (133 on main + 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:59:57 -07:00
gsxdsm
1496ba9658 test(engine): bound the inert-sync-resolution class on a live store (#2794)
## What

One new live-PostgreSQL E2E suite, 3 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Closes the series:
#2789 (scheduler), #2791 (planner lanes), #2792 (custom fields), #2793
(terminal node).


`packages/engine/src/__tests__/workflow-sync-selection-blast-radius-live-e2e.pg.test.ts`

## Why this one is different

The four PRs above each proved a site broken because it resolved a
task's workflow synchronously. Read together they invite a conclusion
that is **false and would be expensive**: that every synchronous
consumer of the workflow selection is inert.

Most are not. The difference is one line of shape:

```ts
// GUARDED (correct)
store.getTaskWorkflowSelectionAsync
  ? await store.getTaskWorkflowSelectionAsync(id)
  : store.getTaskWorkflowSelection(id)

// UNGUARDED (inert)
store.resolveTaskWorkflowIrSync(id)
```

The real PostgreSQL store **does** implement the async reader, so every
guarded site takes the async arm and resolves the card's own workflow.
Only the sync IR helper — which has no async arm to fall to — is stuck
with the default.

Observed on one live store, one persisted workflow, one task:

```
hasAsyncReader   = function
SYNC  selection  = undefined
ASYNC selection  = { workflowId: "WF-001", stepIds: [] }
EFFECTIVE planReviewMaxRevisions = 9   <- the custom workflow's declared default
```

## The point

"The ternary saves them" is an inference from reading, and the whole
premise of this program is that reading is what let the class survive in
the first place. The guarded sites are exactly the ones a fleet worker
would otherwise "fix": converting a correct site costs review time,
risks behaviour, and produces a diff that looks like progress. This
makes the bound checkable in the same lane as the defects.

Guarded call sites (correct today): `workflow-settings-resolver.ts`,
`workflow-ir-resolver.ts`, `executor.ts`,
`workflow-graph-task-runner.ts`, `workflow-task-runtime.ts`, and
`board-workflows.ts` in the dashboard.

## The allow-listed family is now closed

| site | status |
|---|---|
| `scheduler.ts` | proven broken — #2789 |
| `replan-target.ts` | proven broken — #2791 |
| `task-store-helpers.ts` | proven broken — #2792 |
| `branch-and-pr-entities.ts` | proven broken — #2793 |
| `workflow-task-create-ops.ts` | **legitimately correct** — creation
runs before any selection exists, so the default IR is the right answer
|
| `lifecycle-ops.ts` | **NOT proven, stated as such** |

`lifecycle-ops.ts`'s stale-transition-pending recovery re-runs plugin
column-transition hooks against the sync IR. Driving it needs a
registered plugin hook plus a crash-simulated marker; I did not build
that harness and I am not substituting a unit test for it. Named in the
file so it is not mistaken for covered.

## Evidence discipline

- **Observed state.** Both readers called on one live store against one
persisted workflow, plus a real resolved settings value — not a spy on
which arm ran.
- **The settings default is `9`**, deliberately not the builtin's, so
the value can only have come from this workflow.
- **Mutation-verified.** Rewriting the guarded consumer to call the sync
reader directly fails **exactly** the bound arm; the two structural arms
are correctly unaffected, which is what a bound should do.

## Verification

- new suite — **3/3 passed**, mutation-verified
- full live-PG E2E surface — **136/136 passed** (133 on main + 3; #2791
landed while this branch was in flight)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:56:48 -07:00
gsxdsm
90f6319b79 batch-engine tail: re-land the ASYNC half; the sync-resolved half was inert (engine −15) (#2785)
Tail of `batch-engine` (#2773). That PR merged as a squash while later
engine work was still in flight, so `self-healing.ts`, `executor.ts` and
`worktree-pool.ts` landed at their pre-conversion counts. This re-lands
**only the half that is real**, and the reason the other half is not
here is the substance of this PR.

## Census, per file (measured, `--strict` verified)

| file | main | here |
| --- | ---: | ---: |
| `engine/src/self-healing.ts` | 107 | 97 |
| `engine/src/executor.ts` | 15 | 12 |
| `engine/src/worktree-pool.ts` | 3 | 2 |
| `engine/src/ephemeral-worker-manager.ts` | 1 | 0 |
| `engine/src/agent-tools.ts` | 5 | **0** |
| `engine/src/gridlock-detector.ts` | 3 | **0** |
| `engine/src/triage.ts` | 4 | 1 |
| `engine/src/mission-execution-loop.ts` | 2 | **0** |
| **net** | | **−28** |

Baseline re-recorded; `--strict` tightened exactly these 4 entries and
no others.

## Finding: a whole class of conversions in this program is INERT, and
the census scores it as progress

`resolveTaskWorkflowIrSync` returns the **default** workflow IR for
every task in production. The sync selection reader behind it is a
PostgreSQL-cutover stub:

```ts
// packages/core/src/task-store/workflow-definitions.ts:505
export function getTaskWorkflowSelectionImpl(_store, _taskId) {
  return undefined;   // "Backend mode cannot synchronously read PostgreSQL"
}
```

So a guard written as
`resolveLifecycleColumns(store.resolveTaskWorkflowIrSync(id))?.hold`
resolves an IR, asks for a trait, and answers **from the default
workflow for every custom board** — silently. It reads as converted and
the census counts it as converted. `main` gained
`sync-workflow-ir-callsite-allowlist.test.ts` for exactly this after my
branch point; it is what caught me.

I had built three sync resolvers on that reader — `resolveMoveLanesSync`
(self-healing, executor) and a widened `resolveTaskParkedColumnsSync`
(scheduler) — reasoning that a *synchronous* `task:moved` listener needs
a *synchronous* reader. That reasoning was sound about the shape and
never checked whether the reader reads anything.

**Dropped from this PR, deliberately, and NOT re-landed anywhere:**

- `scheduler.ts` 12 → 1 (the widening; the pre-existing narrow helper on
main is untouched)
- the executor `task:moved` handler, incl. the Move-Task hard-cancel
lane comparison
- self-healing's `task:moved` fan-out,
`classifyPausedAbortWorkflowRecovery`, `reconcileInReviewBranchRebind`,
`recoverWedgedActiveMerge`, `recoverPausedAbortFailures`, and 12
single-row lane conversions

Those sites are back to their literals. The allow-list's own guidance is
the standard I applied:

> An unconverted `=== "todo"` is strictly better, because it is at least
honest about being a literal.

I did not add my call sites to the allow-list. Six entries would have
turned the gate green in two minutes and buried the defect; the list's
contract requires proving the async resolver is genuinely unreachable,
and for a fire-and-forget listener it is not — the listener can `void`
an async lane resolution the same way `NotificationService` already
does. That is the correct fix and it is a behaviour-shaped change, so it
is out of scope here.

**Fleet-wide consequence:** any conversion routed through
`resolveTaskWorkflowIrSync` is fake progress, and the census cannot see
the difference. `pnpm test:gate` can: the allow-list test is the
detector. Its passing here (161/161) is this PR's evidence that nothing
inert survived the split.

## What IS in this PR — all async-resolved

1. **`self-healing.clearStaleBlockedBy`** — lanes resolved per
**REFERENCED** task, not per iterated task. A blocker's own workflow
decides whether it is still blocking.
2. **`executor` dependency satisfaction** — resolved per **DEPENDENCY**
via `columnsWithFlag`. Preserves the load-bearing asymmetry that a
dependency in *review* already satisfies a dependent; a bulk sweep
flattens that to complete-only and deadlocks the board.
3. **`agent-tools` — the agent task tools listed FINISHED cards as
active.** `fn_task_list` says it lists "tasks that aren't done or
archived"; `fn_task_search` offers `includeDone: false`. Both filtered
on `task.column !== "done"`, so a renamed complete lane returned
finished cards as outstanding work **to an agent**, which then reasons
and acts on them. `includeArchived` was always enforced by the QUERY and
survived a rename; `"done"` was only ever a TS predicate, which is why
exactly that half broke.

Plus the two **dedup** guards in the same file. The cross-parent
diagnostic filter kept a *shipped* card as a candidate on a renamed
board, so the guard adopted it as canonical and returned `wasDuplicate:
true` — absorbing new diagnostic work into a task nobody is working on
(the eval-followup defect shape again). The defined-feature bootstrap
preflight is **not** the query-filter class: its query passes
`includeArchived: true`, so the TS predicate is the *only* archived
guard there; on a renamed archive lane the archived sibling became the
bootstrap canonical and `claimDefinedFeatureTask` then rejects the
non-live row, so a valid first task fails to be created at all.

Both dedup invariants **already had tests** — asserted against the
legacy ids only, so both passed for the very comparison being replaced.
Extended in place into vocabulary differentials rather than added as
parallel files. Two helpers rather than one parameterised one: "is this
finished?" and "is this archived?" are different questions, and merging
them would make the archived-only guard also reject completed rows.

The list/search half re-landed **with the test it originally shipped
without.** No suite exercised either tool, so the original commit's
"304/304 green" said nothing about the change — the optional-flags
failure mode exactly. Both call sites are covered; converting two copies
and testing one is the Surface Enumeration failure this program has
already hit twice.

4. **`gridlock-detector` — FALSE dependency alarms.** The gate compared
each blocker against `done`/`in-review`/`archived`; on a renamed board
all three are true for a *finished* blocker, so no dependency ever
counted as met and the detector reported dependency gridlock for tasks
that are not blocked — `notifyGridlock` then pages the operator.
Resolved per dependency using the **same five flags** as the executor's
gate (`complete`, `archived`, `mergeOrchestration`, `mergeBlocker`,
`humanReview`) — `review` is not a trait, and two gates answering "is
this dependency satisfied?" differently is a split brain. Every
pre-existing case in that file omits a workflow, so none could detect
the change; added the renamed case plus a non-vacuous companion.

5. **`triage` — its OWN copies of the same two tools.**
`createTriageTools` carries a `fn_task_list` and `fn_task_search`
byte-identical in intent to the agent-tools pair, plus a third site
filtering duplicate candidates. Same defect on all three. Reused the
(now exported) agent-tools helper rather than adding a third copy —
deliberately stronger than the two-parallel-tests reading of Surface
Enumeration, since the copies now share one implementation and cannot
drift. **Not claiming call-site coverage:** `createTriageTools` is
private and not drivable without standing up a TriageAgent; the helper
is revert-proofed, those two call sites are covered only through it.

6. **`mission-execution-loop` — a finished fix task read as LIVE,
stalling remediation.** The comment above that line states the rule it
implements: *only an open task makes duplicate triage safe to suppress.*
On a renamed board the rule inverts — a finished fix task is not
`done`/`archived`, so it reads as live, remediation for a fresh
validation failure is suppressed indefinitely, and the mission stalls
with no error surfaced.

**Not revert-proven, and I am not claiming it is.** No test reaches the
`hasLiveFixTask` branch, and the only case that mints a fix feature is
git-gated and heavyweight; building that fixture is larger than the
conversion. The change strictly *widens* the finished set (resolved
roles ∪ the two legacy ids), so default boards are byte-identical — that
is the argument for shipping it unproven, not a substitute for coverage.

7. **Four census-invisible membership guards**, each inverted on a
renamed board — `worktree-pool` (merger-managed branch reclaim could
delete a branch out from under an in-flight merge), `agent-assignment`
(assignment load counted nothing), `ephemeral-worker-manager`
(`isAgentIdle` inverted on both sides), and the dead constants their
conversion orphaned. These are `SET.has(task.column)` shapes the census
does not count, so the −15 understates them.

## Revert results (measured, each run)

| conversion | reverted → |
| --- | --- |
| `clearStaleBlockedBy` per-referenced lanes | renamed-vocabulary case
fails; stale `blockedBy` never clears |
| executor dependency satisfaction | dependent never unblocks on a
renamed review lane |
| `worktree-pool` merger-managed set | reclaim proceeds against an
in-flight merge |
| `ephemeral-worker-manager.isAgentIdle` | idle agent reads busy on a
renamed board |
| `fn_task_list` terminal filter | RENAMED case fails — shipped card
listed as active |
| `fn_task_search` terminal filter | RENAMED case fails — same,
independently |
| cross-parent diagnostic dedup | RENAMED case fails — `wasDuplicate:
true`, new work absorbed |
| bootstrap preflight archived guard | RENAMED case fails — `validate`
called with the archived sibling |
| gridlock dependency gate | RENAMED case fails — false gridlock raised
for an unblocked task |

`agent-assignment`'s widened `taskStore` type is compile-time; its
revert is a tsc failure, not a test failure — stated rather than claimed
as coverage.

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71, all green (161 includes
`sync-workflow-ir-callsite-allowlist`)
- `npx tsc -p packages/engine/tsconfig.json --noEmit` — clean
- `pnpm lint` — clean

One commit is a pure import restore: `columnsWithFlag` arrived in a
sibling commit that built on the inert resolver and was left behind. The
engine tsconfig excludes `src/__tests__/**`, so the gate was green while
tsc was not — worth knowing that on this package a green gate is not a
green build.


## Verified NOT a gap — measured, so the next worker does not re-open
them

- **`restart-recovery-coordinator` (5 counted).** Four already take an
optional `reviewColumns` set and the counted literals are the documented
**fallback** arm, which must stay for the same reason `columnRoles.ts`
keeps its id fallback. The sole production caller
(`self-healing.ts:12151-12154`) already passes the resolved set. The
fifth is documented at the site as a re-assertion behind a `listTasks({
column: "in-progress" })` query filter. Nothing to convert.
- **`notification/notification-service` (5 counted).** Already
documented in-file as deliberately counted with no exemption marker: the
wedge-episode site needs per-task serialisation of wedge handling (a
delivery-semantics change to operator notifications), and
`isManualMergeHold` needs a pre-resolved `LifecycleColumns` threaded
through `handleTaskUpdated`, which would pay resolution on every task
update. Both are behaviour/placement judgements, not conversions.
- **`planner-overseer` (3 counted).** `resolveWatchedStage`'s two
literals are fed by `pollPlannerOverseer`, which calls `listTasks({
column: "in-progress" })` and `{ column: "in-review" }` — hardcoded
**query** filters. On a renamed board those queries return no rows, so
the predicate never sees a renamed column. Converting it alone would
drop 3 from the census and change nothing an operator can observe. The
real fix is at the query layer; that is the tracked query-filter-bounded
class, not this PR.
- **`triage:695`** reads `resolvePlannerLanes` → the allow-listed sync
IR reader. Left as an honest literal per the rule above.

**Still open in `packages/engine`, deliberately not in this PR:**
`self-healing.ts` (97, of which ~31 are the query-filter-bounded class
and the rest need per-site classification in a 13k-line file),
`scheduler.ts` (12, blocked on the sync reader above), `executor.ts`
(12), and a tail of ~13 more copies of the "is this task finished?"
question across eight small files (`agent-reflection`,
`auto-merge-finalization`, `merger-scope-auto-widen`,
`backlog-pressure-reporter`, `merger-orphan-rehome`,
`merger-integration-worktree`, `plugin-runner`, `cli-agent/*`). That
tail is a clean follow-up: one question, eight call sites, and the
exported `resolveTerminalColumnsForTasks` helper already exists for it.

That is the same discipline as the sync-resolver finding: a census
number that drops without a behaviour change is not progress, and four
of these files would have handed over exactly that.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:47:44 -07:00
gsxdsm
6bdde6f246 fix: five lifecycle gates the census cannot see — incl. live ephemeral workers reaped and duplicate follow-up cards (#2787)
Five lifecycle-column fixes the census **structurally cannot see**. Each
gate is a `Set` or array literal — a *definition*, not a comparison — so
no backlog entry ever pointed at any of these files. Found by grepping
for lane-shaped list literals after the same shape surfaced in
`duplicate-intake` and `blocker-fanout` (both merged via #2780), then
confirmed by reading each USE site.

**On opening this:** I offered twice to fold these into a PR and kept
them on handoff refs to respect one-open-PR-per-worker. They have now
sat unadopted across several cycles while `main` moved, and two of them
destroy or duplicate work. Opening is the reversible call — **close it
if it breaks queue policy** and I will keep them on the branch.

## What is in it

| commit | defect on a renamed board | severity |
|---|---|---|
| `beb107a7bc` | assignment load-balancing **defeated** —
`assignmentLoad` stays empty, every candidate reads as load 0, the sort
falls through to its stable `createdAt` tiebreak, so **one agent wins
every assignment** while the rest idle | distribution |
| `cf4b59e1cb` | the zombie sweep **deletes LIVE ephemeral workers** |
**destroys work** |
| `5fe004ae64` | eval follow-up dedup sees **zero open tasks**, so every
run re-files follow-ups it already filed | **duplicate cards** |
| `a1021de8b2` | agents keep a **"working on" indicator for finished
cards** | stale UI |
| `86680d1220` | the **Files tab never loads** — the fetch never fires |
silent empty |

### The one that destroys work

`shouldDeleteOnSweep` tested a hard-coded terminal `Set`, then fell
through to `return task.column !== "in-progress"`. On a renamed board
**both halves miss, and they compound in the worst order**: the terminal
test fails, control reaches the fallthrough, and `"building" !==
"in-progress"` is `true`. An ephemeral worker **actively executing a
task** is classified as a zombie and deleted. Nothing logs.

Its fallback is **deliberately asymmetric**, and the comment says why:
an unresolvable workflow keeps the legacy literals rather than guessing.
Failing to reap a dead worker costs a slot; reaping a live one destroys
work in flight. Those are not symmetric, so uncertainty fails toward
keeping the worker.

## Verification

Verified **as a set**, not only per-branch:

- `pnpm test:gate` — **161 / 13 / 487 / 71**
- engine suites (assignment, ephemeral, eval-followups) — **44 passed**
- dashboard suites (agent-task-link, useSessionFiles) — **16 passed**
- `tsc` engine + dashboard server + dashboard app — clean
- `pnpm lint` clean · census `--strict` exits 0

**Revert-proven individually.** Restoring each literal fails its own
case: the renamed-wip zombie case, the renamed-wip assignment case, the
renamed-lane dedup case, the sanitizer ratchet, and both
`useSessionFiles` role cases.

## Two honesty notes, flagged rather than buried

**`a1021de8b2`'s guard is STRUCTURAL, not behavioural.**
`sanitizeAgentTaskLinks` is a closure inside `createApiRoutes`,
reachable only by standing up the full express app. The ratchet asserts
the source — resolver threaded per task, bare literal call gone, cache
shared, fallback retained — and **fails on revert**, verified. It is not
a substitute for a behavioural test; whoever owns the dashboard server
should add one if that seam grows.

**`useSessionFiles`'s negative case passed in isolation and failed in
the suite.** Hooks are not unmounted between cases there, so a prior
case's in-flight fetch landed inside it. That is the classic shape of a
test that gets "fixed" by reordering; it now asserts a **delta** against
the pre-render call count, which is independent of what leaks in.

## Deliberately NOT included

`worktree-pool.ts:1205` — the sixth site from the same sweep. It **fails
safe**: a missed match means the skip does not fire, so the branch is
added to `activeBranches` and *protected* from cleanup. The cost is
stale branches accumulating, not deletion. It also sits in the merger's
branch-reaping path, where the opposite error destroys work, so it
deserves its owner's judgement rather than a drive-by conversion.
Flagged, not guessed.

Also still open and unclaimed: roughly 69 untriaged literal-list sites
across engine/dashboard/cli. The grep is one line and the file list is
on #2775 — with the measured caveat that about half are false positives
on shape alone (`LEGACY_*` names, seeds unioned with resolved values,
and `roles: ["triage"]`, which is an `AgentCapability`, not the deleted
column). Only the use site settles it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:35:31 -07:00
gsxdsm
dcf9900d61 test(engine): live-PG differential between resolvePlannerLanes and its async twin (#2791)
## What

One new live-PostgreSQL E2E suite, 5 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Follows #2789 (same
defect class, different site).


`packages/engine/src/__tests__/workflow-planner-lanes-sync-vs-async-live-e2e.pg.test.ts`

## The finding

`replan-target.ts` exports two functions with identical logic and
identical fallbacks, differing only in how they obtain the task's IR:

```
resolvePlannerLanes(store, taskId)              -> store.resolveTaskWorkflowIrSync(taskId)
resolvePlannerLanesForTaskAsync(store, taskId)  -> await resolveWorkflowIrForTask(store, taskId)
```

Under PostgreSQL the sync selection reader answers `undefined` for every
task, so the sync twin resolves the **default** workflow for every card
regardless of the board it is on. **Nine production call sites use it**
(2 × `executor.ts`, 7 × `triage.ts`); one uses the async twin.

The module's own doc comment argues this, and the store-level fact is
proven in `sync-workflow-ir-is-always-default.pg.test.ts`. What had no
executable evidence is the consequence **at this seam** against a real
store with a real persisted workflow. That is this file — a pure
differential: both twins, same store, same task, same call.

### Two harms, different severity

1. **Wrong lanes, labelled authoritative.** `resolvedFromWorkflow`
exists to tell a caller "these came from the workflow, not the
fallback". The sync twin sets it `true` — an IR did come back — while
handing over the default board's ids. A caller that correctly checks the
flag before trusting the lanes is misled *precisely by checking it*,
which is strictly worse than the honest `false` an unresolvable store
would give.
2. **Invented forward lanes.** `wip`/`review`/`complete` are optional so
a caller *refuses* rather than moving a card into a column the board
does not declare (PR #2628's review). The sync twin defeats that
contract without touching it: never having seen the real board, it
reports the default board's forward lanes as present. The optionality is
intact in the type and unreachable in practice. The sharpest arm: a
board declaring **no** review lane gets `undefined` from the async twin
and `"in-review"` from the sync twin.

### A correction worth carrying forward

"It falls back to the legacy lanes" is the wrong mental model **twice
over**. `LEGACY_PLANNER_LANES` (`hold: "todo", intake: "triage"`) is
reached only when no IR resolves at all — under PostgreSQL, never. What
a caller actually receives is the **post-U11 merged default**, whose
intake and hold are one `todo` lane. So the sync twin does not return
`triage` for intake; it returns `todo`, and a caller reading `intake`
gets not merely a wrong id but a lane that is not a dedicated intake at
all.

This is also why the control arm uses `MERGED_VOCAB`: the shape the
twins agree on is the merged one. `DEFAULT_VOCAB`, which splits intake
out as `triage`, already separates them.

## Evidence discipline

- **Observed state.** These are exported pure functions over a live
store; the observation is their return value against a persisted
workflow definition. No spies, no mock IR anywhere in the file. Contrast
the unit coverage in `planner-lanes-async-resolution.test.ts`, which
must supply a mock `resolveTaskWorkflowIrSync` and therefore cannot see
this divergence at all.
- **Control arm.** On the post-U11 default shape the twins agree exactly
— which is why this survived: every default-board test passes and only a
renamed board separates them.
- **Characterization, not endorsement.** The four renamed arms assert
the wrong-but-current values deliberately; they flip when the call sites
move to the async twin, and that flip is the point.

### Mutation-verified, including a round that found weak arms

Replacing the sync twin's whole body with `return LEGACY_PLANNER_LANES`:

| | arms failing |
|---|---|
| first draft | **3 of 5** |
| after strengthening | **5 of 5** |

Two arms originally asserted only `wip`/`review`/`complete`, which are
identical in the merged default IR and in `LEGACY_PLANNER_LANES` — so
they proved the lanes were wrong without proving *why*, and survived the
mutation. Each now also pins `intake` (`todo` merged vs `triage`
legacy), the single field that separates "resolved the wrong board" from
"took the fallback". Recorded in the file next to the assertions.

## Not done, and why

**No fix.** The async twin already exists and is documented as a drop-in
("identical logic and identical fallbacks — the ONLY difference is
awaiting the authoritative resolver"), so the migration is mechanical
*where the caller is already async*. It is not universally so: several
`triage.ts` sites are inside synchronous paths, and `triage.ts:831`
calls `resolvePlannerLanes(this.store, "")` with an empty task id — a
sweep-wide lane read that has no single task to resolve against and
needs a decision, not a mechanical swap. Both are behaviour calls in
files another worker owns; flagging, not smuggling.

## Verification

- new suite — **5/5 passed**, mutation-verified 5/5
- full live-PG E2E surface, 20 suites — **126/126 passed** (was 121/121)
- `pnpm lint` — clean
- `pnpm check:lifecycle-columns` — exit 0

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added end-to-end coverage comparing synchronous and asynchronous
workflow lane resolution.
* Validated lane consistency for renamed, non-default, and custom boards
using persisted workflow data.
* Added checks for incorrect fallback lanes, workflow resolution
indicators, and absent review lanes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:32:24 -07:00
gsxdsm
87442b9664 test(engine): live-PG evidence that declared custom fields cannot be written (#2792)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Third in the series
after #2789 and #2791; same root cause, materially worse consequence.


`packages/engine/src/__tests__/workflow-custom-fields-sync-resolution-live-e2e.pg.test.ts`

## The finding

**A workflow that declares custom fields cannot have any of them
written.**

`TaskStore.resolveTaskCustomFieldDefsSync` reads a task's field
definitions through `store.resolveTaskWorkflowIrSync`, which under
PostgreSQL answers `undefined` for every task and therefore resolves the
**default** workflow IR. The default declares no `fields`, so the
function returns `[]` for every task on every board. `task-update.ts`
validates every write against that empty list:

```ts
const defs = store.resolveTaskCustomFieldDefsSync(id);
const result = validateCustomFieldPatch(defs, updates.customFields);
if (!result.ok) throw new CustomFieldRejectionError(result.rejection);
```

Observed against a real store with a real persisted workflow declaring
one `text` field:

```
STORED fields = [{"id":"risk","name":"Risk","type":"text"}]
SYNC   defs   = []
WRITE  threw  = CustomFieldRejectionError
                custom field 'risk' rejected (no-fields-defined):
                the resolved workflow declares no custom fields; no values may be written
```

The rejection message is a true statement about the workflow that got
resolved and a false one about the workflow the card is on.

### The two halves of the feature disagree in production

The executor resolves the same definitions through the **async**
resolver (`executor.ts` → `resolveTaskCustomFieldDefs` →
`resolveWorkflowIrForTask`) and sees the real field. So an agent can be
prompted to supply a value that the store will then refuse to store. The
last case asserts both answers against **one store, one task, one
workflow** — which is why this cannot be dismissed as a fixture
artefact.

This is a different severity from the previous two PRs in the series.
#2789 and #2791 are wrong-lane defects, mostly latency, one of them
unbounded. This one is a declared feature that does not function off the
default board.

## Scope on record

Three write paths share the sync resolver: `task-update.ts` (driven
here), `workflow-task-create-ops.ts:394`, and `workflow-ops.ts:488`.
Only the first is exercised; the other two are named in the file so the
surface is recorded rather than implied.

Also worth stating plainly: because the empty list *is* the default IR's
`fields`, the same rejection is what a default-board card gets too. The
feature is not merely renamed-board-broken.

## Evidence discipline

- **Fixture integrity first.** The opening case asserts the stored
workflow really does declare the field via the async resolver. Every
other assertion is about a *missing* definition and would pass just as
well against a workflow that never declared one — that case is what
makes the rest mean something.
- **Observed state.** The thrown typed rejection plus the **absence** of
a persisted value on a re-read row. No spy on the validator.
- **Mutation-verified.** Replacing the sync resolver's body with a
hardcoded `[{id:"risk",…}]` fails **3 of 4** arms. The fourth is the
fixture-integrity case, which exercises the async path by design and
correctly survives.

## Not done, and why

**No fix.** The async resolver already exists and is already used by the
executor for the same data, so the shape of the fix is clear — but
`task-update.ts`'s validation runs inside a synchronous update path, and
making it async is a behaviour decision in `@fusion/core` that belongs
to that file's owner, not to a smuggled edit in an evidence PR. The
call-site allow-list entry for `task-store-helpers.ts` ("Synchronous
helper shared by txn-hot paths") should cite this suite either way: the
entry is accurate about the constraint and silent about the cost.

## Verification

- new suite — **4/4 passed**, mutation-verified 3/4 (fourth by design)
- full live-PG E2E surface, 20 suites — **125/125 passed** (121 on main
+ 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:32:12 -07:00
gsxdsm
755ada91ac test(engine): live-PG evidence that the terminal-node guard fires on the wrong node (#2793)
## What

One new live-PostgreSQL E2E suite, 3 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Fourth in the series
after #2789 (scheduler), #2791 (planner lanes), #2792 (custom fields).


`packages/engine/src/__tests__/workflow-terminal-node-sync-resolution-live-e2e.pg.test.ts`

## The finding

FN-7641 Signature 2 exists because setting `nodeId` to the terminal node
used to be written verbatim and silently do nothing — the card sat in
review with every step done, unadvanced and unexplained. The contract: a
terminal override **with** durable merge proof finalizes the card;
**without** proof it is rejected with an actionable error; non-terminal
overrides are untouched.

On a board whose terminal node is not called `end`, **both halves
invert**:

| write | contract says | actually observed |
|---|---|---|
| `nodeId: "end"` — an ordinary planning node here | written, untouched
| **rejected** with a merge-proof error about finalizing a card the
operator was not finalizing |
| `nodeId: "finish"` — this board's real `end`-kind node | finalize, or
reject | **written verbatim**, no error, card left in review |

The second row is the original FN-7641 bug, restored on every custom
board.

## The correction the mutation runs forced

My first draft blamed `isTaskTerminalNodeIdImpl`'s sync IR resolution
alone. Mutating it changed only one of the two cases, which is how I
found there are **two** guards:

```
branch-and-pr-entities.ts:568   validateNodeOverrideChange(task, nodeId, { isTerminalNodeId })
                                -> sync IR resolution (the default board, under PostgreSQL)
task-update.ts:53               validateNodeOverrideChange(task, nodeId)
                                -> NO options, so `defaultIsTerminalNodeId` — the bare literal
                                   `nodeId === "end"`
```

The inner one is an unconverted literal sitting behind a converted call
site, and it silently overrides it. **Converting the outer guard alone
changes nothing an operator can see.** A column census cannot find the
inner one either — `end` is a node id, not a column. This is the "a
guard survives in a branch of the same function" shape, one function
apart.

### Mutation matrix

| corrected | `end` rejected | `finish` silent |
|---|---|---|
| *(nothing — main)* | pass | pass |
| outer sync-IR guard only | pass | **fail** |
| inner `defaultIsTerminalNodeId` only | pass | **fail** |
| **both** | **fail** | **fail** |

Two different failure structures, which is why the cases are kept apart:

- **`end` rejected is over-determined** — both guards independently call
it terminal, so it survives a mutation of either one. Not a weak
assertion: a faithful record of a defect with two independent causes,
and the reason a partial fix here is invisible.
- **`finish` silent is under-determined** — both guards must miss the
id, so correcting either flips it. This is the arm that notices a
partial fix.

The fixture-integrity case exercises the async resolver by design and
correctly survives every mutation.

## Fixture

The shared builder's terminal node is `end`, so it cannot express this
shape. This file derives from it: one `lifecycleIr`, node ids shifted so
the `end`-kind node is `finish` and the non-terminal planning node takes
the name `end`. Columns, traits, edges and structure are otherwise the
builder's, so the only variable is which node ids carry which kind. The
first case asserts that shift really happened — both characterizations
are claims about which node is terminal and would read as defects if the
fixture had quietly kept the builder's ids.

## Evidence discipline

- **Observed state.** Whether `updateTask` throws, and what the re-read
row's `nodeId` and `column` actually are. No spies.
- **Characterization, not endorsement.** Both cases assert the
wrong-but-current behaviour deliberately, and the matrix above says
exactly which fix flips which.

## Not done, and why

**No fix.** It needs two coordinated edits in `@fusion/core` — threading
the resolved terminal check into `task-update.ts:53`, and making the
outer resolution async — and the second is the same synchronous-path
constraint as #2792. Both are behaviour decisions in another worker's
files. Worth flagging that fixing only the allow-listed sync site would
look like progress and deliver none, which the matrix above makes
checkable.

## Verification

- new suite — **3/3 passed**, mutation matrix above
- full live-PG E2E surface — **124/124 passed** (121 on main + 3)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:29:04 -07:00
gsxdsm
e467d939a5 fix(tests): the last 3 engine reds — a pause guard asserted at the wrong layer (#2779)
## What was red

The final 3 failures in `executor-prompt.test.ts` ("global pause
behavior"), all the same assertion:

```ts
expect(mockedCreateFnAgent).not.toHaveBeenCalled();
```

made after calling `executor.execute(task)` **directly** on a paused
todo row.

## It is not a regression, and not a live safety hole

I initially flagged this in #2778 as a possible live hole — "a
user-paused todo task **now** reaches `createFnAgent`". **That framing
was wrong**, and the bisect is what corrected it:

| commit | result |
|---|---|
| `origin/main` (HEAD) | fail |
| `main~40` | fail |
| `main~80` | fail |
| `main~150` | fail |
| `main~250` | fail |

Red 250 commits back. It never described shipped behaviour, so nothing
regressed.

**`execute()` holds no pause gate.** Neither `executeCore` nor the
workflow-graph executor consults `paused`/`userPaused` before starting a
session — I checked both. Refusing to dispatch a parked row is the
**scheduler's** invariant, enforced twice:

1. Candidacy is keyed on both flags (`scheduler.ts:138`) — `userPaused`
is a durable operator stop even when legacy `paused` is false.
2. The row is **re-read immediately before dispatch** and refused if it
comes back parked (`scheduler.ts:2086`) — this closes the race the first
check cannot.

The test called `execute()` directly, stepping around the component that
owns the guarantee, then asserted the bypassed layer enforced it. A true
statement about the system was being made to look false.

Every protective outcome #2371 documented **does** hold and stays
asserted: `fn_task_done` never completes the card, it is never handed to
`in-review`, no completion watchdog is armed, the pause is never
cleared, and the run narrates the benign paused park. Only *"no session
was created"* was false. The 3 sibling assertions in `resumeOrphaned`
are untouched — that path genuinely does refuse.

## The invariant moves to the layer that owns it

Rather than delete an assertion and lose the coverage,
`scheduler-paused-dispatch-refusal.test.ts` pins it through real
`schedule()` passes:

- **control** — an unparked ready card IS dispatched
- refuses a row parked with legacy `paused`
- refuses a row parked with `userPaused` alone
- refuses when the operator pauses **after selection, before dispatch**

Driven through `schedule()` rather than by calling the predicate
directly: a test that calls the guard cannot tell whether the dispatch
path still *consults* it — which is precisely how the executor-prompt
version came to assert a layer that had stopped being asked.

## The control earned its place on the first run

It failed immediately, and twice over: the hold-release gate refuses a
card still carrying a bootstrap seed (fixed with the shared
`seedPlannedSpec`), and a `moveTaskIf` stub returning `moved: false`
makes the release unobservable. Without the control, all three refusals
would have passed **vacuously** — a scheduler that dispatches nothing
refuses everything.

## Evidence

- **106/106** across both files.
- **Mutation:** removing the pre-dispatch pause re-read → the race case
fails. The other two are caught earlier by candidacy (defence in depth);
the passing control makes their refusal attributable to the flag alone,
since the same store dispatches without it.
- Gate **732 green** · lint clean · engine `tsc --noEmit` **0 errors**.
Test-only (the scheduler mutation was reverted; `git diff` clean).

## Engine suite status

Measured baseline on `origin/main`: **39 failures / 10835 passed**. With
#2776 (32, notifier harness) and #2778 (4), this last set takes the
engine suite to **0 failures**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:04:52 -07:00
gsxdsm
145c022af4 test(engine): live-PG evidence that the scheduler's sync parked-column read is inert (#2789)
## What

One new live-PostgreSQL E2E suite. **No production file is touched** —
this is evidence, per the E2E worker's remit.


`packages/engine/src/__tests__/workflow-scheduler-parked-columns-live-e2e.pg.test.ts`
(2 tests)

## The finding

`scheduler.ts`'s `resolveTaskParkedColumnsSync` resolves a task's
hold/intake columns through `store.resolveTaskWorkflowIrSync`. Under
PostgreSQL that reader answers `undefined` for **every** task, so the
resolver returns the **default** IR and the function yields `{ hold:
"todo", intake: "triage" }` on every board — byte-identical to the
literals it was converted away from. It is an **inert conversion**, and
it is currently allow-listed
(`sync-workflow-ir-callsite-allowlist.test.ts`) on the grounds that
these handlers are synchronous.

Five call sites read it. Four groups of handler fail as **latency** — a
wake that does not fire costs up to one poll interval, which is why the
class hid. The `task:deleted` dependency reconciliation is different: it
queries `listTasks({ column: hold })` **and** re-checks
`dependent.column === hold` before clearing `blockedBy`. On a renamed
board both tests are against `"todo"`, a column that board does not
contain, so **a dependent parked in the renamed hold column is never
unblocked and waits forever on a blocker that is already gone.**
Persisted, operator-visible, unbounded.

### It contradicts a passing unit test

`scheduler-renamed-hold-events.test.ts` asserts the opposite and passes,
because its mock supplies `resolveTaskWorkflowIrSync: vi.fn(() =>
renamedIr())` — an answer the real store provably never gives. That test
is not wrong about the *scheduler* (given a working resolver the
handlers do resolve the renamed lane); it is wrong about the *resolver*.
Flagging rather than editing it: it is still the right unit test for its
own subject, and it is not my file.

### The mechanism is not the one the code reads like

The obvious reading blames the fail-soft `?? "todo"`. It is **not** that
— `lifecycle` is never nullish, a real default IR comes back and real
traits resolve off it, so both `??` arms are dead in production.
Established by mutation, not by reading:

| mutation to `resolveTaskParkedColumnsSync` | control arm | renamed arm
|
|---|---|---|
| *(none — main)* | unblocks ✅ | never unblocks ✅ |
| both `??` fallbacks → renamed vocabulary | unblocks (unchanged) |
never unblocks (unchanged) → **dead branch** |
| returned object → renamed pair | fails | fails → **both arms decided
here** |

This is the sharpest form of the defect class: the site resolves an IR
and reads a trait off it, so it looks converted at every level except
the one that decides the answer.

## Evidence discipline

- **Observed state, not spies.** Each arm asserts the dependent's
persisted `blockedBy` after a real soft-delete on a real store with real
stored workflow definitions.
- **The negative is self-validating.** The handler's work is
fire-and-forget, so observing it needs a bounded wait — and a bounded
wait proving a negative is normally worthless. The default-vocabulary
arm is the control: same store, same window, and it *does* unblock (297
ms against a 2 000 ms window). If this ever flakes the control fails
first; the fix is the quarantine ledger, never a larger number.
- **Differential.** Both boards come from the one shared vocabulary
builder and differ only in their column ids.
- The renamed arm is a **characterization** test — it asserts the
wrong-but-current behaviour deliberately, and is expected to flip when
the read is fixed.

## Two fixture traps found on the way (both would have made this
vacuous)

1. `updateTask({ column })` is not a column move — `column` is not in
the update payload, so it typechecks as an unknown key and leaves the
card where it was. Use `moveTask`.
2. **Order matters.** Writing the blocked state *after* placing the card
re-homes it to the intake column (observed `todo -> triage`), and the
reconciliation re-checks `dependent.column === hold` — so the control
fails for a fixture reason that looks exactly like the defect. Block
first, then place. Both are written down in the file.

## Not done, and why

**No fix.** The honest fix is to make the read async, and that is not
free: these run inside synchronous `task:moved` / `task:updated`
listeners, where an added `await` defers the rest of the handler to a
microtask and reorders handlers against a synchronous emitter. That is a
behaviour decision in a file another worker owns, so it belongs to
whoever owns `scheduler.ts` — not to a smuggled edit in an evidence PR.
The allow-list entry should cite this suite either way.

## Verification

- `workflow-scheduler-parked-columns-live-e2e.pg.test.ts` — **2/2
passed**, mutation-verified in both directions (table above)
- full live-PG E2E surface, 19 suites — **121/121 passed** (was 119/119)
- `pnpm lint` — clean
- `pnpm check:lifecycle-columns` — exit 0

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:04:34 -07:00
gsxdsm
698bded476 fix(tests): 4 engine reds on main — each was green for a reason the fleet removed (#2778)
## Context

A full `@fusion/engine` run on `origin/main` (`9b61d795c9`) reports **39
failures / 10835 passed**. 32 are the notifier harness, fixed in #2776.
This PR takes 4 of the remaining 7.

All three files share one shape: **each case was passing off something
the lifecycle conversions have since correctly taken away.** In every
one, the product is fine and a good change landed as a red test.

---

### 1. `executor-graph-failure-lanes-resolved.ts` — an equality that
fails on its own fix

The guard forbids resolving a lifecycle *guard* through the synchronous
`resolvePlannerLanes` (a no-op under the shipped PostgreSQL backend, so
the census counts the site as converted while it behaves like the
literal). It asserted `expect(callSites).toBe(3)`.

#2764 converted the promotion-path site to
`resolvePlannerLanesForTaskAsync` — exactly the direction this guard
wants. Count went **3 → 2** and the assertion failed.

The guard's own comment states the invariant as *"Any FOURTH is a new
sync resolution"* — one-directional. Coded as equality, it fails on
removal, which is the change it exists to encourage. Now
`toBeLessThanOrEqual(2)`.

**Mutation:** adding a third sync call site → `expected 3 to be less
than or equal to 2`. Still load-bearing.

### 2. `restart.integration.test.ts` — a fixture matching a fallback
constant

`recoverCompletedTask` re-homes intake → hold → wip only when the origin
is the board's **intake** lane; otherwise it hands straight to review.
The failure showed the 1st move as `in-review` with no re-home.

Nothing regressed. The fixture put the card in `triage` and resolved
lanes through the sync resolver, so it fell through to
`LEGACY_PLANNER_LANES` — where `intake` is literally `"triage"`. **It
was matching a hardcoded fallback, not a declared lane.** #2764 made the
site await the real resolver; the mock selects `builtin:coding`, and
**U11 merged intake and hold onto one Planning column (`todo`)**, so
`triage` is not a lane on that board and the two-hop correctly
collapses.

The invariant the test is named for — completed work in a distinct
intake lane is re-homed along a legal path, not moved intake → review,
which role adjacency rejects — is still real. So the fixture now
**declares** a board with intake separate from hold, the only shape
where the two-hop is reachable.

**Mutation:** removing the re-home hop from the product → fails with the
expected `todo` first-move. Load-bearing.

### 3. `executor-abort-provenance.test.ts` — a call one argument short

Both provenance cases returned `false` for a clean completed in-review
row. This reads as an FN-6796 regression stranding rows that are already
handed off for review.

It is not. #2703 added a 7th `reviewLane` parameter so the lane is
resolved by the caller. **The call goes through `as any`, so the missing
argument was not a type error** — it arrived `undefined`, `live.column
!== reviewLane` held for every row, and the classifier answered false
for everything.

Passed explicitly rather than defaulted inside the classifier: a default
would restore the literal the parameter exists to remove. Added a
**differential** — a card resting in a *renamed* review lane classifies
the same, a mismatched one does not — so the parameter cannot be
re-literalized while still looking converted.

**Mutation:** `live.column !== "in-review"` → the differential fails.
The other cases pass, which is precisely why it was worth adding.

---

## Evidence

| file | result |
|---|---|
| `executor-graph-failure-lanes-resolved` | **24 passed** |
| `restart.integration` | **48 passed** |
| `executor-abort-provenance` | **16 passed** |

Gate **732 green** · `pnpm lint` clean · engine `tsc --noEmit` **0
errors**. Test-only — no product file is touched by this PR (the
mutations above were run and reverted; `git diff` confirms clean).

## Deliberately NOT fixed here

3 cases in `executor-prompt.test.ts` ("global pause behavior") remain
red on main: **a user-paused todo task now reaches `createFnAgent`**.
That is a safety invariant rather than a stale fixture, and neither
`executeCore` nor the graph executor holds a pause gate — the refusal
#2371 documented is not where its note implies. It gets its own change;
editing the fixture to match current behaviour would hide it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:40:49 -07:00
gsxdsm
c643d62e85 fix(executor): wipDeclared must ask ALL six lifecycle roles, not two (#2777)
## What this fixes

`resolveResumeLanes` returns `wipDeclared`, which gates whether
`routeGraphFailureToExecutionResume` may route a graph failure back into
execution resume. Getting it wrong terminalizes tasks on boards that
should resume.

Two prior versions were wrong, both caught in review rather than by me
at write time:

**1. Two-state (`lifecycle?.wip !== undefined`)** — greptile P1 on
#2760. A v1-upgraded board terminalizes: `synthesizeDefaultColumns`
emits `{ id, name: id, traits: [] }`, so *every* role resolves
`undefined` even though those columns literally are the legacy lanes.
Verified by parsing a real v1 IR.

**2. Proxying "synthesized" as "hold and review are both undefined"** —
my own fix for (1), and also wrong. I caught this against #2765 rather
than shipping it. A **v2** board that declares only `intake` +
`complete` has hold and review undefined too, so it would be misread as
synthesized and treated as declaring wip when it deliberately does not.

The failure mode both versions share: reading a *sample* of the roles
and treating the answer as a verdict about the whole IR. #2765 says it
directly — an empty result has two meanings, and you cannot tell them
apart from a subset.

## The rule

```ts
wipDeclared: lifecycle?.wip !== undefined || !declaresAnyLifecycleRole(lifecycle),
```

Three states, asking all six roles:

- **wip declared** → true, the board says so.
- **some role declared but not wip** → false. A v2 board that omits wip
means it; do not resume into a lane it did not define.
- **no role declared at all** → true. That is the
synthesized/v1-upgraded shape, whose columns *are* the legacy lanes; the
pre-existing behaviour is correct there and must not regress.

`declaresAnyLifecycleRole` iterates `Object.values(lifecycle)` rather
than naming roles, so a seventh role added later is included
automatically instead of silently falling into the wrong branch.

## Evidence

- `executor-resume-lanes-resolved.test.ts`: **7 passed**, +23 lines
covering the v1-synthesized board and the declares-some-but-not-wip
board.
- **Mutation:** restoring the naive two-state rule → **1 failed / 6
passed**. The added coverage is load-bearing and pins exactly the
regression greptile caught.
- Gate **732 green** · `pnpm lint` clean · engine `tsc --noEmit` **0
errors**.
- Rebased on current main.

## Scope

`executor.ts` (+22) and its test (+23). One predicate; no other behavior
touched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 09:28:41 -07:00
gsxdsm
8c9b84ae38 batch-core: packages/core + dashboard/src lifecycle conversion (129 → 92) (#2780)
## batch-core — `packages/core` + `packages/dashboard/src`

Shared branch: two workers are converting into it. Opening the PR
because the branch was green with none, and a branch without a PR merges
nothing.

### Census

Measured with `node scripts/lifecycle-column-census.mjs --json`.

| | guards |
|---|---|
| batch-core scope at branch point | 129 |
| batch-core scope now | **92** (51 files) |
| repo total now | 358 |

Files closed so far: `store.ts` 11→0, `task-merge.ts` 6→0,
`live-agent-count.ts` 6→0 (marked, not converted — see #2762),
`task-update.ts` 3→0, display-ordering + Wake Delta ranking 5→0,
`register-git-github.ts` 4→0.

### The `register-git-github.ts` slice

Three PR routes — `pr/create`, `pr/push-branch`, `pr/resolve-conflicts`
— plus the `CHANGES_REQUESTED` handler each compared `task.column !==
"in-review"`. On a renamed board **none** of them matched, so every PR
affordance the dashboard offers was refused for a card sitting in the
lane that board calls review, and the refusal named a column that does
not exist there.

All four now share one helper, `reviewColumnsForTask`, which gets two
things right that this program has repeatedly gotten wrong:

- **Membership, not a single id.** It takes the broad review set
(`mergeOrchestration ∪ mergeBlocker ∪ humanReview`).
`resolveLifecycleColumns` returns the *first* column per trait, so a
single-id answer silently ignores a board that declares a merge lane
**and** a separate human sign-off lane. These guards only refuse or
permit — they never move the card — so over-admitting costs nothing
while under-admitting refuses a request that should have worked.
- **An empty resolved set means UNEXPRESSED, not absent.**
`synthesizeDefaultColumns` upgrades a v1 graph by emitting every default
column with `traits: []`, so a v1-upgraded workflow resolves to an empty
review set while its `in-review` column plainly exists and holds the
card. Reading empty as "this board has no review lane" would refuse
these routes on **every pre-v2 project** — a worse regression than the
one being fixed, and invisible to any v2 test.

This is the dashboard twin of the `fn pr create` guard in
`packages/cli/src/commands/pr.ts` (#2775). The two surfaces answer the
same question and now agree — FN-5893 surface enumeration.

### Testing note: why the seam and not the routes

I wrote route-level HTTP tests first and **deleted them**. An express
fixture over `registerGitGitHubRoutes` hangs — every case, including the
pure refusals, times out at 4s, because registering the router starts
background work the fixture never satisfies. Making it run would mean
mocking git, the GitHub client, and the pollers: a mock-the-world shell,
which is what the project's do-not-add-slow-tests rule (FN-5048) says to
avoid in favour of a narrow seam.

`reviewColumnsForTask` *is* the narrow seam — it holds the entire
decision, and the four call sites now do nothing but ask it and render
its answer. Six cases pin it: the renamed lane is returned and
`in-review` is not, a two-lane board returns both, a v1-upgraded board
falls back, an unresolvable workflow falls back, and the refusal renders
lanes an operator can act on.

**Mutation-verified, both directions:** reverting the helper to the
legacy literal fails 2 of 6; treating an empty set as an answer fails 1
of 6.

One fixture bug worth recording, since it would have made the two-lane
case vacuous: the trait id is kebab-case `human-review`, not
`humanReview`, and the built-in traits must be registered via `import
"@fusion/core"` before flags resolve.

### Verification

- `pnpm --filter @fusion/dashboard exec tsc --noEmit -p tsconfig.json` →
0 errors
- `pnpm lint` → 0 errors
- `register-git-github.review-lanes.test.ts` → 6 passed

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:28:27 -07:00
gsxdsm
b42b40aa48 fix: make flushAsyncWork actually drain — clears 32 notifier failures on main (#2776)
## Red on main, fix-forward

batch-engine (#2773) left 32 failing cases on main:
`notifier.runtime.test.ts` (19) and `notifier.test.ts` (13), all
`expected "vi.fn()" to be called 1 times, but got 0 times`.

## The failure is in the harness, not the notifier

`notifier.ts` is untouched by #2773.
`notification/notification-service.ts` was converted, and
`handleTaskMoved` is fire-and-forget (`void
this.handleTaskMovedAsync(data)`). The async path now awaits
`resolveLifecycleColumnsForTask` and then `resolveReviewColumnsForTask`
(which awaits `resolveWorkflowIrForTask`) — several more await hops
after `store.emit(...)` returns.

The harness had no slack to absorb them:

```ts
export async function flushAsyncWork(): Promise<void> {
  await vi.waitFor(() => { expect(true).toBe(true); });
}
```

The condition is true on the first tick, so `waitFor` resolves
immediately. **It never waited for anything.** It worked only while the
handler completed within a single turn — and it reported the resulting
breakage as a notifier defect rather than as its own.

Another entry in the recurring pattern this program keeps hitting: a
cheap check that reads as authoritative. A `waitFor` looks like
synchronization at the call site; this one was a no-op.

## Evidence

| run | result |
|---|---|
| before | 32 failed / 68 passed (100) |
| after | **100 passed (100)**, 3.25s |
| after, mutated back to a single `await Promise.resolve()` | 32 failed
/ 68 passed — the same 32 |

The mutation run is the point: the fix is load-bearing, not a
coincidence of timing.

Gate 732 green · `pnpm lint` clean · engine `tsc --noEmit` clean.

## Reversible decision, noted: microtasks only

A `setTimeout(0)` drain also turns all 100 green, and it was my first
version. Rejected on measurement:

- it costs real wall-clock at every call site — the two files went **~2s
→ over 2 minutes**;
- it **stalls under the fake timers** `notifier.test.ts` installs (lines
355/381/589), where a pending `setTimeout` never fires — 4 cases hung.

The awaits being drained are promise-based (workflow-IR resolution), so
microtask turns are the right currency, they work identically under real
and fake timers, and they cost nothing. Per AGENTS.md *"Do Not Add Slow
Tests"* — prefer fake timers over real time waits.

16 turns is slack, not a tuned number; the chain is ~4 deep today.

## Scope

One test-harness file. No product code, no behavior change. Tests
asserting a specific outcome should still prefer `vi.waitFor` on *that
outcome* — this helper covers the "let the fire-and-forget handler
finish" case, and now actually does it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:25:21 -07:00
gsxdsm
1fb53f9924 docs(scheduler): the task:moved arms are blocked by a synchronous prologue — measured, and deliberately left counted (#2771)
## No behaviour change, and deliberately **no markers**

The ten `from`/`to` comparisons in `scheduler.ts`'s `task:moved` handler
stay **counted** in the census. They are genuinely wrong on a renamed
board — real backlog. Marking them `DELIBERATE-LITERAL` would claim
*"reviewed, correct"* when the truth is *"reviewed, still broken,
blocked on an ordering question"*, and that is the opposite of what
#2767's markers were for.

This records the blockage instead. I have deferred these twice citing
risk; this is the analysis that deferral was standing in for.

## The measured constraint

**The handler is `async`, but its prologue is not.** There is no `await`
anywhere between the handler's first line and the terminal-blocker
branch ~55 lines down. The snapshot invalidation, the PR-monitor
start/stop pair, the mission hand-off and the failed-task tracking all
run in the **same tick as the emitter**.

So hoisting a resolution to convert those arms does not cost "one await"
— it converts the **whole prologue into a microtask**, reordering this
listener against every other synchronous `task:moved` subscriber and
against the emitter's own continuation.

That makes `resolveTaskParkedColumnsSync`'s *"SYNCHRONOUS on purpose"*
note **load-bearing rather than stale** — verified by measurement, not
assumed. I had been treating it as possibly-stale boilerplate.

## Why the two obvious workarounds don't apply

- **Resolve lazily inside the branch.** Doesn't help: the *condition* is
what needs the lanes, and it is evaluated in the prologue.
- **A cheap sync superset prefilter** — the shape that worked in
`usage-limit-detector` — needs a literal predicate that cannot wrongly
*exclude* on an unknown vocabulary. For *"is `to` the terminal lane?"*
no such predicate exists: a renamed board's terminal id is unknown by
construction. (That is precisely why the prefilter *was* safe there —
literals can only fail to exclude, never over-exclude.)

## What would actually unblock it

1. **Audit the ordering**, then hoist one await and convert all ten
together. That is an audit across every `task:moved` emitter and
subscriber — not a scheduler-local change, and not something to do
speculatively.
2. **Carry the resolved lanes on the event payload**, so no listener
resolves at all. This is the only option that scales to the *other*
synchronous listeners with the same problem, and it removes the class
rather than one instance.

I'd recommend (2) if this is worth funding — it is the same shape as the
fix that removed the sync-resolution class in #2759, one layer up.

## Verification

11 scheduler suites — **130 passed** · `pnpm test:gate` **158 / 487 / 10
/ 71** · `pnpm lint` clean · engine `tsc --noEmit` **0 errors** ·
`--strict` exits 0, census unchanged at 12 for this file (which is the
point).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added internal documentation explaining synchronous event-ordering
requirements when resolving task lanes.
* Clarified why asynchronous resolution must not be introduced in this
scheduling flow.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:10:02 -07:00
gsxdsm
9b61d795c9 fix(engine): heartbeat asked 'is this task finished?' with legacy ids — and one of the two sites writes status:failed onto completed work (#2769)
## Two heartbeat sites asked "is this task finished?" with the legacy
ids

`agent-heartbeat.ts` **4 → 0**. Neither site is cosmetic.

**Linked-task clear.** The heartbeat clears an agent's assignment once
its card is finished. Keyed on the literals, an agent on a renamed board
stayed bound to a **completed** card indefinitely — every later
heartbeat ran with stale task context instead of picking up new work,
and nothing else clears it.

**Worktree-acquisition gate.** Its failure bookkeeping runs only for a
**non-terminal** task. A card in a renamed complete lane read as
non-terminal, so an acquisition failure could stamp `status: "failed"`
and an error message onto work that was **already done**.

That second site *writes*, which drives the fallback direction: an
unresolvable workflow degrades toward "terminal", because treating a
finished card as unfinished is the expensive mistake here.

Both sit in async paths — the first has `await taskStore.getTask(...)`
three lines above — so this is an `await`, not a restructure. Extracted
to one predicate rather than converted twice: they are the same
question, and the two must not drift when one of them acts
destructively.

## How this was found, and the part worth recording

Generalising #2767. That PR marked a documented false positive the
census kept advertising, so I swept for **other** files whose lifecycle
literals were reasoned about in prose but still counted — to find out
whether the trap was systemic.

**It is not.** Of twelve candidate files, only this one carried real
unconverted guards, and its "false positive" mentions turned out to be
unrelated (detection heuristics, not column literals). The sweep mostly
came back **negative**, and that is worth saying so nobody repeats it
expecting a haul.

## Revert proof

| reverted | result |
|---|---|
| neuter the resolution (predicate → literals) | **3 failed** / 2 passed
|
| shipped | **5 passed** |

The two that survive the revert are the degraded-mode pair, which is
correct — they assert the *legacy* answer, so they must pass either way.
One case also pins that a legacy `done` id is **not** terminal on a
board that does not declare it, which is what a board-wide union would
get wrong.

## Verification

- 13 heartbeat suites — **501 passed** (496 before, +5 new)
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
engine `tsc --noEmit` **0 errors** · `--strict` exits 0

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:42:00 -07:00
gsxdsm
7119432c79 fix(engine): the stranded-completed recovery never resolved on a renamed board — and its suite fed the broken reader the right answer (#2764)
## The "recovery of last resort" never resolved anything

`recoverCompletedTask` carries this note in its own source:

> *This is the recovery of last resort — a literal here means the last
resort does not exist off the default lineage.*

It was resolving through `resolvePlannerLanes`, which reads
`resolveTaskWorkflowIrSync` — whose selection reader returns `undefined`
**unconditionally** in PostgreSQL mode, the shipped backend. So it
resolved the **default** workflow for every card,
`promotedFromPlannerColumn` was `false` on every renamed board, and the
recovery never fired.

That is precisely the stranding it exists to fix — completed work
sitting in a planning lane with nothing left to rescue it — **with the
conversion in place and the census counting it as done.**

The call site is inside an async method that has already awaited store
reads, so the fix is an `await`, not a restructure.
`resolvePlannerLanesForTaskAsync` is the async twin: identical logic,
identical fallbacks, one `await`. Answers are unchanged on the default
lineage and correct everywhere else.

## The existing suite could not see any of it — the more important half

`executor-planner-lanes-resolved.test.ts` injected **only**
`resolveTaskWorkflowIrSync`:

```ts
(store as { resolveTaskWorkflowIrSync: ... }).resolveTaskWorkflowIrSync = () => ir;
```

It fed the broken reader **the right answer**. Every case proved the
promotion *logic* while being structurally blind to whether production
resolves at all — and it was green the entire time. A suite that cannot
fail for the reason the code is broken is the same defect as the code,
one level up.

The harness now feeds the sync reader the **default lineage** (what it
actually returns) and the async readers the task's real workflow.

| | reverting the call site to sync |
|---|---|
| before this PR | **0 failed** — suite blind |
| after | **5 failed** / 13 passed |

Two cases opt back in via `syncResolvesIr`, and only those two: they
cover `isPlannerColumnFor` and `isBackwardMoveOutOfPlanning`, which are
still synchronous, so there the sync reader genuinely *is* the input
path and feeding it the IR tests the classifier rather than the reader.

## Not converted, deliberately

**Those two classifiers.** They sit in an else-if chain whose next arm
is `from === "in-progress"`, so deferring the decision into an async
body changes which arm runs. That branch's own comment records a
previous half-conversion there:

> *a half-conversion turned a missed rescue into active damage. Third
time this program has produced that shape — gates converted,
destinations left literal.*

That needs the chain enumerated first, not a fast restructure at the end
of a sweep. Their production inertness is held by the
`resolveTaskWorkflowIrSync` call-site allow-list in #2759, so they
cannot be forgotten.

## Verification

- new suite **4 passed** · strengthened suite **14 passed** (18
together)
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
engine `tsc --noEmit` **0 errors** · `--strict` exits 0

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:38:43 -07:00
gsxdsm
7f3eee9db7 docs(engine): mark the usage-limit terminal filters DELIBERATE-LITERAL — a documented false positive that has now baited two workers (#2767)
## No behaviour change. This is the work order retracting a documented
false positive.

I went to convert the `done`/`archived` filters in
`usage-limit-detector.ts`, reasoning that on a renamed board a provider
rate limit would pause already-finished work. I wrote the conversion —
and only then read the note a previous worker had left directly above
it:

> *"The FIRST thing I suspected there — the `done`/`archived` terminal
filter — turned out to be a **FALSE POSITIVE**: its revert stayed green,
because the lane check already excludes finished cards."*

**They are right and I was wrong.** A terminal card is already excluded
downstream: `taskUsesProvider` resolves the task's active lane, a
finished card matches no active lane, so it resolves no providers and
cannot be affected. The suite pins exactly this — `pauses a PEER
executing in the renamed WIP column` asserts `FN-SHIPPED` is not paused.

My conversion is reverted. It changed nothing at runtime and would have
lowered the census count while behaviour stayed identical — the precise
shape this program keeps warning about, produced by me this time.

## Why a marker and not just the existing prose

The note was already there and I walked into it anyway, because **the
census kept listing this file as 4 unconverted guards**. The work order
advertised the work; the reasoning against it lived in a comment you
only reach after you have started. Prose informs a reader who is already
looking; a marker informs the *instrument*, so the file drops out of the
work order.

Two distinct reasons are recorded rather than one blanket marker,
because they are not the same argument:

- **the prefilter** is a deliberate cheap **superset** (#2672 review).
Converting it reintroduces the whole-board resolution that review
removed. Literals are safe here in the direction that matters — a
renamed board declares no `done`/`archived` id, so nothing is wrongly
*excluded*.
- **the final filter** is redundant with the lane check, and that
redundancy is already proven by an existing test.

## Census

| | before | after |
|---|---|---|
| `usage-limit-detector.ts` | 4 | **0** |
| repo backlog | 437 | **433** |
| DELIBERATE-LITERAL (reviewed) | 38 | **42** |

Every one of the 4 is a marker, not a conversion. The backlog moved
because reviewed literals left it honestly, not because behaviour
changed.

## Verification

`usage-limit-detector.test.ts` **58 passed**, unchanged before and after
· `pnpm test:gate` **10 / 158 / 487 / 71** · `pnpm lint` clean · engine
`tsc --noEmit` **0 errors** · `--strict` exits 0.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:31:47 -07:00
gsxdsm
d86c1f9d29 batch-engine: packages/engine lifecycle-column conversions (capacity worker's mega-batch) (#2773)
The engine mega-batch. Folds my four engine PRs and will absorb the
remaining `packages/engine` guards as commits on this branch.

**Superseded and closed:** #2722, #2741, #2766, #2770.

## Census — files converted so far

| file | before | after |
|---|---:|---:|
| `notification/notification-service.ts` | 9 | **5** |
| `runtimes/in-process-runtime.ts` | 6 | **1** |
| `eval-followups.ts` | 2 | **0** |
| `pr-comment-handler.ts` | 1 | **0** |
| `task-revert.ts` | 2 | **0** |

The last two are **census-invisible** (`Set.has(task.column)`
membership) — the class measured in #2763, which a comparison-based scan
cannot count. So the backlog number moves less than the work does,
deliberately.

## What each one actually fixed — all silent, none cosmetic

- **Notifications stopped entirely.** `handleTaskMovedAsync` compared
`data.to` to `in-review`/`done`, so on a renamed board the two
notifications operators rely on most were never sent.
- **A finished card's plan review could re-enter.** The continuation
drain's terminal test matched nothing, so a completed card's planning
continuation was handed to the executor.
- **The revert route admitted and the service refused.** The route
resolved terminal lanes; the service compared to a hardcoded pair. The
operator got a dead end from an affordance the UI and route both
offered.
- **Follow-up dedup blocked new cards forever.** A finished follow-up in
a renamed complete lane read as *open*, so the dedup matched it
permanently — defeating the intent the code documents in the line above
it.
- **The mission requeue wrote a column that may not exist**, and its
guard never matched.

## Flagged, not fixed — deliberately

- **`concurrency.ts` idle semaphore leak recovery** — the last live
caller of the running-agent predicate that does not enrich. On a renamed
board it under-counts and can reclaim a legitimately-held slot. The
enriching variant is async and this is a synchronous repair path whose
failure mode is reclaiming live work.
- **The archival `task:moved` listener** — runs on every move with no
cheap gate ahead of it; converting costs an IR resolution per move to
decide most moves are not archival.

## Notes carried from the folded PRs

Two conflicts resolved in main's favour because **main's version was
better**: `in-process-runtime`'s seam uses `terminalColumns:
ReadonlySet` (membership) where mine used `LifecycleColumns`
(first-per-role), and the test is rewritten against main's API. That
arity trap has now caught me four times, so membership is the default
shape in everything new here.

Review fixes from the folded PRs are included: the notifier's review
set, the second human-review site, the second dedup copy, the workspace
revert surface, and the file-content assertions.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **224 passed** across
the touched engine suites · engine and dashboard `tsc` clean · `pnpm
lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:24:50 -07:00
gsxdsm
86639f2ce4 fleet: planning drain + archive writers 12 → 4 — one stale row starves planning, and a finaliser that wrote an undeclared column (#2742)
**Claimed on #2733 before starting.** `in-process-runtime.ts` +
`task-artifacts-ops.ts` — **12 → 4**.

## 1. The planning drain: one stale row stops planning for the whole
project

FN-8470's own note on this code says it: **one orphan earlier in
created_at FIFO prevented every later planning continuation from
dispatching.** So on a renamed board the literal terminal pair did not
mis-handle one card — an archived or completed card's stale work item
read as live, stayed in the due set, and **starved the drain behind
it**.

The two classifiers take an **optional** terminal set, which is this
file's own injection idiom (the specification-complete reaction already
takes a `resolveIr` dependency so the pure passes are testable without
constructing a runtime that would attach to the real project registry).

**Optional is load-bearing:** a *required* parameter would have compiled
at every existing caller and then answered "not terminal" for
everything. That is the silent direction, and both halves are asserted
in the test.

## 2. `moveToDoneImpl` writes `task.column` directly

This is the store's own finaliser, not a `moveTask` caller — so its
literal is **not** caught by `moveTask`'s unknown-column validation the
way every converted call site in this program is. It silently persisted
`done` on a board that does not declare it, and then emitted `to:
"done"` to every listener.

**This is one of the few sites where a literal writes bad state rather
than merely failing to act.** A workflow declaring no complete lane now
throws instead of inventing one — #2733's rule: a missing field on a
resolved struct *is* an answer, and `?? legacy` discards it.

## 3. The unarchive destination — three decisions in four lines, all
literal

| pre-archive column | lands in |
|---|---|
| unusable / archived | the **complete** lane |
| the **wip** or **review** lane | the **hold** lane (its worktree and
session are long gone) |
| anything else | back where it was |

The second is the expensive one: a card archived *from* the wip lane was
restored straight back *into* it **with no worktree**, and the scheduler
then counts it as a live holder **occupying a slot**. Made async — its
one production caller already is, and the sync alternative is the
PostgreSQL no-op documented in #2703.

## Also

- **The mission-error requeue** (guard *and* destination in one change):
an errored mission task stayed in the wip lane holding a slot, because
the guard never matched.
- **The planner-chat retention cutoff on archive** — the quiet direction
of this defect class: nothing breaks, data that should be deleted simply
accumulates, and the only symptom is storage growth nobody attributes to
a column name.

## The live defect is not where the census points

`reliability-metrics.ts`'s 6 guards are **pure historical readers** over
activity-log entries, and **the dashboard does not call them**. The live
path is `server.ts`'s `getTaskMovedCountsByDay({ toColumn: "in-review"
})` — a **SQL query filter**, the class the census counts separately.

So the operator's reliability panel reads zero on a renamed board
because of a *query* literal, and converting the six guards the census
reports **would change nothing an operator sees**. Converting historical
readers also risks reinterpreting past events under today's traits,
which is a different decision from converting a live guard — I am not
making it inside a vocabulary sweep.

Worth generalising for the fleet: **a file's census count and its live
exposure are different numbers.** This is the second file where the
reported guards are the inert copy and the real one is a query
(`executor.ts:5805` was the first).

## Verification

`pnpm test:gate` **10 / 158 / 487 / 71** · 31/31 continuation suites ·
8/8 archive PG suites · in-process-runtime PG suite green · 5 new cases,
**2 red on revert** · `tsc` clean in core and engine · `pnpm lint`
clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:53:28 -07:00
gsxdsm
c53d3aec38 fix(executor): re-land the no-wip-lane fix — #2757 merged a snapshot that predated it (#2760)
## Why this exists

#2757 merged as `9a2a033b9e`, but **the third of its three fixes is not
on main**:

```
$ git show origin/main:packages/engine/src/executor.ts | grep -c wipDeclared
0
```

The merge captured my branch *before* commit `68381db72f`, so the
no-wip-lane fix was dropped while the other two landed.
`executor-execution-policy-renamed-columns` → *"a workflow with no wip
column terminalizes visibly instead of claiming the card advanced"* is
still red on main, still returning `status: null, error: null`.

This is that commit, cherry-picked cleanly onto current main. No new
work — the review discussion is in #2757.

## What it fixes (recap)

`resolveResumeLanes` defaulted `wip: lifecycle?.wip ?? "in-progress"`,
collapsing two different states:

1. **the IR failed to resolve** — defaulting is right; the `catch` arm
wants exactly this
2. **the IR resolved and declares NO wip column** — defaulting *invents
a lane the workflow does not have*

`routeGraphFailureToExecutionResume` then admitted a card resting in
that workflow's **hold** lane with incomplete steps, rehomed it,
returned `true` — and the terminalize branch never ran. Resuming into a
workflow with no implementation lane *is* "claiming the card advanced"
when nothing did.

The fix adds `wipDeclared` (declared, as opposed to defaulted) and
declines the resume when it is false — the same fail-closed rule the
sibling branch already applies with `wipColumn !== undefined` before
calling a card "already advanced". That path failed closed; this one
failed open.

IR-unavailable deliberately keeps today's behaviour (`catch` reports
`wipDeclared: true`), so an infrastructure error does not start refusing
legitimate resumes.

## Verification on current main

| check | result |
|---|---|
| the three affected files | **23 passed** |
| `pnpm test:gate` | **726** |
| engine `tsc --noEmit`, `pnpm lint` | clean |
| remove the guard | 1 failed — the fail-closed case goes red again |
| always decline | 1 failed — a legitimate resume breaks |

Full-suite blast radius was measured on #2757 before it merged: 834
files / 10,845 tests, 5 failures, both files pre-existing
(`executor-prompt`'s pause-guard 3 and `executor-abort-provenance`'s 2,
byte-identical to baseline). Zero new failures.

## Note

Worth checking whether other PRs merged in that window lost their final
commits the same way — I only noticed because I re-verified main after
the merge rather than assuming a merged PR contains what the branch
held.
2026-07-30 07:44:35 -07:00
gsxdsm
7c408ef650 fleet: merge path 10 → 2 — a merged PR never advanced its task on a renamed board (#2733)
**Claimed on #2728 before starting.** The merge path:
`merge-queue-ops-2.ts` + `merger.ts` — **10 → 2**, both survivors
flagged with reasons.

## A merged PR never advanced its task on a renamed board

`applyPrMergedTransition` is what moves a card when GitHub reports a PR
merged. Every guard in it was a default-lineage literal, and they all
failed **in the same direction**:

| guard | renamed board |
|---|---|
| `column === "done"` → skip as already-done | never matched, so a
complete card was re-processed |
| `column !== "in-review"` → bail `wrong-column` | always matched, so a
card **sitting in review** bailed |

Net effect: **a PR merged on GitHub never advances its Fusion task.**
The operator sees a merged PR whose card sits in review forever — which
reads as a broken webhook, so it gets debugged in the wrong place
entirely. That is the most expensive property of this defect class: it
does not just fail, it misdirects.

One snapshot now covers the pre-check, the deliberate **re-read** (a
merge can land between checks), and the **move target**. The target is
asserted in the test alongside the guards, because converting guards
alone would admit the card and then move it to a column the board does
not declare.

## merger.ts

- **The orphan-stash liveness guard** classified every finished task as
unfinished on a renamed board, so orphaned stashes were never cleaned
up. Unioned with the legacy ids: too strict here leaves clutter, too
loose **discards a stash whose task is still running**, so
over-inclusion is the safe direction.
- **The worktree-conflict scan** filters by worktree *path* before
resolving lanes. The naive order — resolve, then filter — is exactly
what made the github-tracking reconciler scan proportional to task
history (#2714 review). Lesson transferred rather than re-learned.
- The deprecated `aiMergeTask` already-finalized guard.

## Two flagged, not converted

**`merge-queue-ops-2`'s sync enqueue guard** runs inside
`store.db.transactionImmediate`. A synchronous lane resolution reads
`getTaskWorkflowSelectionImpl`, which returns `undefined`
**unconditionally in PostgreSQL mode** — so a "conversion" there would
drop the census by one and behave exactly as the literal (the finding
from #2703). Converting it properly means making the path async or
pushing the trait read into SQL: store architecture, not a call site.
Left literal **with that note**, so the next worker does not turn it
into a false green.

`merger.ts`'s last comparison is the same class.

## Pre-existing red, reported not folded

**22 failures in
`packages/dashboard/src/__tests__/routes-github.test.ts`** — spec
revise/rebuild and approve/reject-plan, all asserting moves to
**`triage`, the column U11 deleted**. Verified by reverting my diff and
re-running: identical 22. Same stale-literal-in-a-test class as the two
assertions #2720 fixed, and it is 22 tests pinning a column that does
not exist — worth someone owning deliberately rather than as a rider
here.

## Verification

census **10 → 2** · `pnpm test:gate` **487 / 71** · 23/23 across three
merger suites · 4 new cases, **2 red on revert** · `tsc` clean in core
and engine · `pnpm lint` clean.

## Also examined and deliberately left alone

- **`live-agent-count.ts` (6 guards)** — every literal there is the
*documented degradation path* for a task shape that was not enriched,
and both production callers already enrich (`useExecutorStats`, `fn
project`). Converting them converts nothing; deleting them removes the
fallback that fixtures rely on. The invariant that matters is **caller
enrichment**, which is not a literal at all.
- **`task-merge.ts` (6 guards)** — `getTaskMergeBlocker` is a **pure**
function with no store; its callers inject `resolveTask`. Resolving
lanes needs a matching injected resolver, which is an interface change
across every caller. Also worth a decision first: its dependency check
accepts `in-review` as satisfied while the store's `blockedBy`
computation (#2720) does not — **two definitions of "dependency
satisfied" in one codebase**, and I am not settling that one silently
inside a vocabulary sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:08:16 -07:00
gsxdsm
94d88f1d6f fix(census): the work order was sending fleet workers at non-columns (722 -> 714) (#2692)
Found while claiming `TaskDetailModal.tsx` — its census entry included
`session.agentState === "done"`, an **agent state, not a lane**.
Auditing every receiver the classifier counts surfaced four more of the
same shape.

## The misclassified receivers

| site | receiver | what it actually is |
|---|---|---|
| `register-chat-routes.ts` | `event.type === "done"` | an SSE event
type |
| `useTaskDiffStats.ts` | `mode === "done"` | a cache-key mode |
| `async-mission-store.ts` | `evidence.kind === "done"` | an evidence
kind |
| `telemetry-hub.ts` | `event.kind === "done"` | a telemetry event kind
|
| `TaskDetailModal.tsx` | `session.agentState === "done"` | an agent
state |

Each shares a **word** with a column id and nothing else. Converting one
asks the trait registry what lane an SSE event is in, which has no
answer — the same failure class as converting `role === "triage"`, which
this list already exists to prevent.

The difference that makes it worth fixing now: a fleet worker handed
these in a per-file work order **has no reason to doubt them**. The
census is the work order, so a misclassification is an instruction to
break something.

## What I did not exclude

`state` is deliberately kept. `state === "archived"` in `audit-ops.ts` /
`comments-ops.ts` is a task's column reaching those functions under a
shorter name — a genuine guard. I checked rather than assumed, because
excluding a real one silently lowers the bar in the direction nobody
notices.

## Census effect

```
column  722 -> 714
role      5 -> 14
```

Those 8 are **reclassified, not converted** — this PR changes no
production code. The baseline is re-recorded so `--strict` agrees.

## Verification

`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0. `pnpm lint` clean.

No changeset: instrument accuracy, no user-facing change.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:01:52 -07:00
gsxdsm
9a2a033b9e fix(test): two engine reds from fleet churn — and stop the shellout guard breaking on line drift (#2757)
Both failures are fleet-churn fallout on `main`, not defects in the
conversions. Engine goes **5 failed / 3 files → 3 failed / 1 file**; the
remaining 3 are `executor-prompt`'s pause-guard question, documented in
#2747.

## 1. The shellout guard was coupled to line numbers — third time

Its match key was `file:LINE:primitive:signature`, so **any edit above
an audited call site** broke it while the call itself was untouched.
Recent fleet conversions shifted `executor.ts` and `self-healing.ts`,
and 5 sites drifted at once.

**This is the third time it has gone red this way, and the third hand
re-pin.** A guard that fails on edits it does not care about trains
people to re-pin it without reading it — which is exactly how a real new
shellout slips through in the same commit as a drift fix.

So this removes the coupling rather than updating the numbers again.
Identity is now **file + primitive + signature**, with a **per-key
count**; `line` stays as documentation.

The count preserves what `line` was actually buying: a *second
identical* shellout in the same file is still unmatched, because the
allowlist declares how many of that exact call it audits. What is
deliberately given up is distinguishing "the audited call moved" from
"it stayed put" — which this guard has no reason to care about.

**Measured, all three directions:**

| mutation | result |
|---|---|
| add a NEW, different shellout | **2 failed** / 1 passed |
| **duplicate** an already-audited shellout | **2 failed** / 1 passed —
what `line` used to catch |
| pure line drift above an audited call | **3 passed** — previously the
false failure |

## 2. A test that predicted its own flip

`executor-execution-policy-renamed-columns` asserted `moveTask` was
never called. It now rehomes the card to `inbox`.

That is not a surprise — **the test's own comment called it**:

> *"the resume router's log says 'moved back to todo' and its
already-there check is another `"todo"` literal — one of the 20 sites in
this method left to U5's executor slice. It is why the card stays put
here rather than being rehomed to `inbox`."*

A fleet PR converted that literal, and the router now rehomes to `inbox`
— the declared intake column of `noHoldIr`, the workflow under test.
**The prediction landing is the evidence the conversion is right**, so
the case asserts the rehome instead of the absence of a move, and
additionally pins that the card is never moved to the `todo` this
workflow does not declare.

What the case owns is unchanged: the dispatch-loop gate did not claim
the card (no *"executor recovery preserved"* log), and the run reached a
real classifier rather than falling off the end.

## Verification

`pnpm test:gate` **726**, `pnpm lint` clean, engine `tsc --noEmit`
clean.

## Note on PR count

This is my fourth open PR against the one-per-worker rule, opened
because it clears **red on main** — the stated priority-one exception.
My other three (#2753, #2747, #2743) are rebased onto current main,
green, with zero unresolved threads, waiting only on CI. I am opening
nothing further until they land.
2026-07-30 06:54:44 -07:00
gsxdsm
e18a6cf00c fleet: executor.ts 57 → 15 on top of #2689 — the review/wip lanes, 4 half-conversions, 8-of-19 revert proof (#2703)
**Supersedes #2691, which I am closing.** #2689 landed the terminal-pair
batch on `executor.ts` while my PR was open on the same file — we
collided, that PR won the race, and 30 of my 70 conversions are now
identical to its work. Rather than resolve 30 conflict hunks in a
20k-line lifecycle file (unreviewable, and the wrong artifact to hand
you), I rebuilt from `origin/main`.

**`executor.ts` 57 → 15.** Repo backlog 679 → **650**.

## The four that are defects, not vocabulary

**1. `isReentrantPausedAbortedInFlightNode` resolved lanes at the END,
for its return value, while its four `in-review` eligibility gates were
literals.** On a renamed board those gates all read false — so a review
card skipped the global-pause recheck, the `autoMerge === false`
refusal, the shared-branch-member arbitration **and** the
merge-confirmed refusal — and then the lane-resolved final line answered
*"re-entrant"*. FN-7214's own comment says an auto-merge-off review row
must stay terminal.

**2. The REVERSE half-conversion.**
`routeGraphFailureToExecutionResume`'s destination was already resolved
(U7's `resolveReboundColumnFor`) behind a gate that was still three
literals — so the router refused before reaching its own working move.

| direction | what happens | visible? |
|---|---|---|
| resolved gate → literal destination | card admitted, move rejected by
a board with no such column | **yes** — the move errors |
| literal gate → resolved destination | card refused; the working
recovery never runs | **no** |

Only the second is silent, which is exactly why it survived U7's own
conversion of that destination. **When you convert a destination, check
the gate in front of it in the same commit.**

**3. `routeUnusableWorktreeGraphFailureToRecovery` skipped FN-5147's
auto-merge-off gate** on a renamed board — an automatic recovery moving
a human-review-terminal card backward. #2689 converted the terminal
guard at the top of that method; this is the other half of the same
decision, which is the general risk when two people split one file.

**4. `handleGraphFailure`'s `alreadyFinalizedToReview` /
`suppressFinalizedCompletionAbort`** read `column !== "in-progress"`, so
a completed, already-finalized row looked still-in-wip: FN-6644 /
FN-6647's suppression never fired and the row was re-parked as an
operator-action pause abort — the durability gap those tickets closed.

## Two patterns worth carrying to other files

**An inert guard rarely reports "renamed board" — it reports something
that sounds like a different problem.** `finalizeAlreadyReviewedTask`
returned `"missing"` for a card sitting in review. The completion
handoff logged *"no longer active"* for a card that was actively
executing. The stuck-requeue cleanup logged *"recovered concurrently"*
about a recovery that had not happened. Three different false
explanations, one cause.

**Directions differ inside one family, so convert per method, not per
pattern.** Most wip guards read `!== "in-progress"` and REFUSE on
no-match (renamed board → silently disabled). The rerun watchdog reads
`=== "in-progress"` and SKIPS on match — there the literal never
matched, so a rerun could fire on a card **mid-execution**. A mechanical
sweep of `!== "in-progress"` fixes the refusals and leaves that
admission in place.

Also: the resolver choice inverts within a few lines. *"Is this card in
the ONE column finalize targets?"* needs the **complete** column — the
terminal union carries the legacy ids, so a card in a column merely
*named* `done` reads as already finalized and the finalize is
**skipped**. *"Is this card already finished, so do not move it?"* needs
the **union** — over-inclusion only skips a move, under-inclusion moves
a finished card out of its terminal column. Both are recorded at their
sites.

## Revert proof

19 cases in `executor-graph-failure-lanes-resolved.test.ts`, on a board
sharing **no** column id with the default lineage (on the default board
these guards are correct by coincidence — the literals *are* the board).
**8 fail on revert.** The rest are labelled **in the file** as paired
positives, default-board no-change cases, or — in one instance — a guard
that is genuinely redundant with a later lane check. I would rather
label a case as non-evidence than count it.

Two fixture corrections are recorded at their sites, both my own
assertion failing to touch the behaviour it named: asserting a router's
return value (which was already false for an unrelated reason — fixed by
spying on the recovery call), and `allowsAutoMergeProcessing` keying on
the **global** setting rather than `task.autoMerge` (fixed the fixture,
not the assertion).

## The 15 that remain, each with a reason

- **7 `to`/`from` move-effect parameters** — a move's endpoints, not a
card's resting column. Trait-hook territory.
- **2 enumeration scans** — one is a `listTasks({ column: "in-progress"
})` query whose filter cannot be converted without the query (converting
the filter alone reads as done and changes nothing); the other loops
every task, so per-task resolution is a real cost wanting a shared memo.
- **`12325`, the dependency guard** — *"is this dependency satisfied?"*
is not any single lane role. The same question exists at
`register-task-workflow-routes.ts:3995`; both should be decided once,
together.
- **`14484`** (`fromColumn === "in-review" && toColumn === "in-review"`)
— a same-lane move check that belongs with the move-effect group above.

## Verification

`pnpm test:gate` **158 / 10 / 487 / 71** · **135/135** across the 17
suites covering these paths · `tsc -p packages/engine` clean · `pnpm
lint` clean · census `--strict` exit 0, baseline re-recorded.

No changeset: `@fusion/engine` is private and the behaviour change is
confined to renamed boards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved workflow execution across boards with renamed lifecycle lanes
by resolving lane targets per board instead of using fixed column names.
* Fixed review, WIP, completion, and failure-recovery behaviors to
respect the correct board snapshot (including auto-merge and terminal
work states).
* Improved artifact-recovery protection timing and tightened
execution-resume gating for failure scenarios.
* **Tests**
* Added a new lifecycle invariant test suite covering renamed-lane
recovery, resume, pause/abort, and router-gating behavior.
  * Updated lifecycle column census baseline data.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:48:25 -07:00
gsxdsm
61b82a2737 fleet: pure lifecycle predicates 17 → 5 — a monitoring signal that went quiet, and a blocker that waited forever (#2745)
**Claimed on #2742 before starting.** Four pure modules — **17 → 5**,
every survivor flagged with a reason.

All four are **pure functions with no store**, so the fix shape is the
injected-set contract established in #2728, not an in-function resolve.

## Three failures that never error

| predicate | what a renamed board got |
|---|---|
| `getTaskAgeStalenessSignal` | `undefined` for **every** card —
age-staleness reported nothing |
| `isStaleBlockedByBlocker` | "not stale" for a blocker that was
finished, paused in review, or retry-exhausted |
| `areAllDependenciesDone` | "not satisfied" for a dependency that had
landed |

The first is the one to sit with: **a monitoring signal that goes quiet
is indistinguishable from health.** The board looks fine while cards sit
for days, and nobody investigates a metric that isn't alarming. The
signal also chose its *threshold pair* by wip-vs-review, so both halves
were literal.

The second means the blocked card **waited forever**, silently — "not
stale" is the answer that produces no event.

The third is the **third place** "satisfied" is asked. It now gives the
same answer as the store's `blockedBy` computation (#2720) and the merge
blocker: complete or archived, unioned with the legacy ids. Three
surfaces, one rule — which is exactly why I refused to settle it inside
a vocabulary sweep the first two times it came up.

## Optional is load-bearing

Both halves are asserted for every predicate: supplying lanes makes a
renamed board work, **omitting them preserves every existing caller**.

A *required* parameter would have compiled at every call site and then
answered "not active" / "not stale" / "not satisfied" for everything.
That is the silent direction, and **no type checker catches it** — which
is the argument for optional-plus-legacy-default over a clean signature.

The restart-recovery classifiers (with-progress / no-progress /
merge-active) take the same set, and **the combiner threads it to all
three**, so a caller cannot convert the outer question and leave an
inner one literal. `isInReviewMissingWorktreeSessionStartFailure` is
deliberately untouched — #2728 converts it and duplicating that would
conflict.

## The five that remain

- **3 are the ternary trait-fallback branches** (`lanes ? … : legacy`) —
the documented degradation path the census counts by design, not
unconverted guards. I am not marking them `DELIBERATE-LITERAL` to move
the number; that marker means "a lifecycle literal reviewed and kept",
and mislabelling to flatter a count is how the instrument stops meaning
anything.
- **`recoverInterruptedRuns`' filter sits behind a `listTasks({ column:
"in-progress" })` query.** The query is the live filter, so converting
the redundant predicate moves the census and changes nothing an operator
sees. **Third file** where the reported guard is the inert copy and the
real one is a query.
- **`resolveWorkflowBypassGuards` is sync and receives only column
strings** — no task, no store. Converting it means adding lanes to
`MoveTaskOptions` and threading them from the moves path, which another
worker owns. Marked `DELIBERATE-LITERAL` as an explicit hand-off, with
the consequence named: on a renamed board the operator's drag out of the
wip lane was rejected by the transition validator, so **a card could not
be cancelled from the board at all** (AGENTS.md's Move-Task hard-cancel
contract).

## Verification

`pnpm test:gate` **10 / 158 / 487 / 71** · 9 new cases, **5 red on
revert** · 13/13 with the archive PG suite · `tsc` clean in core and
engine · `pnpm lint` clean · census **17 → 5**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:48:00 -07:00
gsxdsm
3da8b90ed9 fleet: scheduler.ts 26 → 12 — five quiet wrong answers on a renamed board (finished deps blocked forever, PRs unwatched, missions stalled) (#2729)
## Census

| | before | after |
|---|---|---|
| `packages/engine/src/scheduler.ts` | 26 | **22** |
| repo backlog | 657 | **653** |

Baseline re-recorded in this PR; `--strict` exits 0. Four literals
removed, and I want to be exact about why it is four and not seven: the
converted predicates keep their legacy literals as the documented
**no-metadata fallback**, and the census counts per literal, not per
code-quality improvement. Deleting those fallbacks would change
behaviour in degraded mode (unresolvable workflow → every dependency
reads as unsatisfied → dependents blocked forever), which is the
expensive direction to be wrong in.

## What was actually broken

**1. Dependency satisfaction was keyed on three column ids.**

```ts
return !!dep && (dep.column === "done" || dep.column === "in-review" || dep.column === "archived");
```

On a board whose complete column is `shipped`, a **finished** dependency
matched none of the three. `getUnmetSchedulingDependencies` reported it
unmet, and the dependent was parked `blockedBy` — *permanently*, because
the dependency can never move anywhere that satisfies the literal. Work
stops and nothing rescues it.

Satisfaction is now resolved on the **dependency's own board**, since a
dependency edge may cross workflows — the dependent can sit on the
default board while the dependency lives on a renamed one. Resolution is
passed in by the caller (`resolveDependencySatisfactionColumns`) with a
caller-owned IR cache, so a sweep reads one IR per distinct workflow
rather than one per dependency edge.

**2. The review half of the file-scope lease was never converted.**

The wip half of this same sweep was fixed on 2026-07-30-16:30
(`scheduler-renamed-wip-file-scope-lease.test.ts`). The review half
still read `column === "in-review"`, so on a renamed board no review
card entered `activeScopes`, a merging card's worktree files read as
**free**, and an overlapping candidate dispatched on top of them. One
registry, two halves, disagreeing. It now uses `isReviewColumnRole` over
the *same* resolved flags map the wip half uses, so the two cannot drift
again.

## Finding I am reporting rather than fixing

**The two satisfaction rules in this one function genuinely disagree,
and the live one is the broader.**

| rule | satisfied when |
|---|---|
| legacy (**live**) | complete ∪ archived ∪ **review lane** |
| marker (shadow) | complete ∪ archived |

#2720 settled "satisfied = complete or archived" for
`update-task-deps.ts` — which matches the **marker** rule, not the live
one. So the scheduler currently treats an in-review dependency as
finished and `update-task-deps.ts` does not.

Reconciling them is a product decision, not a vocabulary one, so this PR
preserves **both** rules exactly as shipped. Narrowing the live rule to
match would strand every dependent of an in-review card, which is
precisely the failure mode fix 1 exists to remove — I am not doing that
as a side effect of a rename conversion.

## Revert proof

Each fix reverted **alone**, tests re-run:

| reverted | result |
|---|---|
| dependency satisfaction → literals | **2 failed** / 5 passed |
| review lane → `column === "in-review"` | **2 failed** / 5 passed |
| neither (shipped) | **7 passed** |

Each revert fails exactly its renamed case *and* its "both vocabularies
reach the same outcome" invariant, while the default-vocabulary controls
stay green — so the failures are attributable to a surviving column-id
literal and not to a generally broken path. The suite is differential:
one workflow **shape**, two vocabularies with identical traits, only the
ids differ. No renamed id collides with a legacy literal, so a surviving
`=== "done"` cannot pass by luck.

There is also a paired negative (`an UNFINISHED dependency still blocks,
under both vocabularies`) so the fix cannot degrade into "always
satisfied" — the direction it could overshoot.

## Verification

- `pnpm --filter @fusion/engine exec vitest run <19
scheduler/dependency/hold-release/overlap suites>` — **174 passed**, no
regressions
- new suite — **7 passed**
- `pnpm test:gate` — **158 / 487 / 10 / 71 passed**
- `pnpm lint` clean · `npx tsc -p packages/engine/tsconfig.json
--noEmit` clean · `check:lifecycle-columns --strict` exits 0

## The remaining 22, triaged

Not guessed at — grouped by what they actually ask:

- **6 fallback branches already converted** (L258, L265, L439, L440,
L1712 and the review twin): trait-first with the literal as documented
degraded-mode fallback. These reach 0 by *marking*, not converting.
- **10 `from`/`to` move-transition arms** (L899–L1042): a different
question ("is this transition *into* a review lane?"), and per the
scoping note in `docs/solutions/architecture-patterns/` they should not
ride along with `task.column` conversions.
- **6 `task.column` reads** (L1061, L1135, L1524, L1546, L1554, L2607):
convertible, but two sit in sync methods (`resolveBaseBranch`) needing
the caller-resolves-and-passes shape, so they are a separate unit rather
than a half-conversion here.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:15:05 -07:00
gsxdsm
72d42652e5 fleet: CLI surface 16 → 0 — 'active=0' on a busy board, and a retry gate that disagreed with the dashboard (#2728)
**Claimed on #2714 before starting.**
`packages/cli/src/commands/task.ts` (8) + `dashboard.ts` (8) — **16 →
0**.

## The finding that matters: `active=0` on a busy board

The same four-line aggregation appears **four times** in `dashboard.ts`
— the TUI stats refresh, the serve summary, the status line, the
agent-stats pass. Each compared the default lineage's two ids, so on a
renamed board every one reported `active=0` while the board was plainly
busy.

**This is worse than an inert internal guard.** A recovery path that
silently stops firing is invisible until something breaks. A stats line
that says zero is **read, believed, and acted on** — *"nothing is
running, so I can restart the engine."*

The four copies are now one helper, and that is the other half of the
fix: four independent copies of a lifecycle decision is how they drift,
and these were identical **by accident, not by construction**. One IR
read per *workflow*, asserted by call count — because the returned
number is identical either way, so only counting the work can see it.

## The retry gate exists twice, and #2713 converted one of them

After #2713, `POST /tasks/:id/retry` accepted a renamed board's stalled
review card while `fn task retry` refused it with *"not in a retryable
state"* — **one operator action answering differently depending on the
surface**.

The rule, stated at the site: **converting one copy of a duplicated gate
creates a disagreement that is harder to diagnose than the original
inert guard.** Grep the classifier by name before calling a lane
converted.

## The rest

- **`fn task set-node` / `clear-node`** rewrote the node override of an
*actively executing* card, because the "is in progress" check never
matched. That guard exists because the rewrite races the run.
- **The duplicate-guard candidate filter** kept completed cards in the
comparison set on a renamed board, so a new task was reported as a
duplicate of work that had already landed — the opposite of useful.
- **The duplicate-lineage `(archived)` marker** never printed, so the
operator could not tell a live duplicate from a filed one.

## Two DELIBERATE-LITERALs, with reasons

The board-render glyph compares `col` taken from the legacy `COLUMNS`
enum **that loop iterates** — the literal matches its own receiver by
construction. The real defect is already named in the code above it: a
card in a renamed column **is not rendered at all**, which is the R8/U10
surface change, not this glyph. Converting it would hide that behind a
trait lookup while the loop still cannot see the card.

## Pre-existing, not mine

5 failures in `commands/__tests__/task.test.ts` (GitHub import) **fail
on `origin/main`** — verified by stashing this change and re-running.
Someone owns that; it should not ride in here.

## Verification

census **16 → 0** · `pnpm test:gate` **10 / 71** · `pnpm smoke:boot`
**PASS** · `tsc -p packages/cli` clean · `pnpm lint` clean ·
`task-retry` 3/3 · 4 new cases with **2 red on revert**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 06:04:23 -07:00
gsxdsm
ddba730a59 fix(core): 8 reds across 4 files — incl. a real FN-8603 contract violation and a ratchet row pinning deleted code (#2725)
## Measured

Full `@fusion/core` suite: **8 failed / 6 files → 1 failed / 1 file**
(4611 passed). `pnpm typecheck` exit 0 across every package, `pnpm lint`
clean, gate **726**.

The one remaining failure is **not mine to fix** — see the last section.

## Four causes; two are product-side, not test drift

**1. A real FN-8603 contract violation.** `tool-output-budget.ts:116`
had a bare `console.warn`, breaking the rule that production diagnostics
route through the shared logger so severity markers and `FUSION_DEBUG`
gating survive. `log-severity-spam-contract` caught it exactly as
designed. Now `createLogger("tool-output-budget")`, kept at `warn` — an
invalid operator-supplied budget is a real misconfiguration, not routine
chatter.

**2. A ratchet row pinning deleted code.** The manifest pinned a `local
reattached project ${project.id}` demotion in `central-core.ts` whose
call site was deleted by `5ae6332563` ("collapse dead SQLite dual-path
code"). Verified absent from **all** of `packages/core/src`, not merely
moved. A manifest row for deleted code can only ever fail — it ratchets
nothing — so it is removed with that provenance recorded in place.

**3. An intentional settings overlap.** `agentToolOutputMaxChars` now
appears in both scopes. Admitted to the parity list because
`settings-schema.ts:462` states the intent outright: *"Project settings
participate in the existing effective-settings merge, allowing a
project-specific tool-output cap … to override global policy."* Placed
in `GLOBAL_SETTINGS_KEYS` order, as that test requires.

**4. `maxPostReviewFixes` 3 → 10 — the third file pinning the stale 3.**
Driven off the exported `DEFAULT_MAX_POST_REVIEW_FIXES` rather than a
fourth literal copy. That constant exists *because* the declaration
default and two inline `3`s had already drifted apart once; adding
another copy would guarantee a fourth drift.

## duplicate-guard: a narrow seam instead of a rebuilt mock

Its 3 failures were `Cannot read properties of undefined (reading
'projectId')` — the fake modelled the **deleted SQLite path**
(`db.prepare().all()`) and recovered the window by parsing a captured
cutoff string. It broke when the query moved to `asyncLayer` + Drizzle.

Rebuilding a Drizzle chain to recover a number the policy already
returns would be mock-the-world for no gain, so the window policy is now
one exported pure function — `resolveFingerprintWindowMs`, the
**byte-identical** expression — that both the store query and the tests
call. Two side benefits: the ±5s timing tolerance is gone (exact
assertions), and the `Math.max(1, …)` floor now has coverage the old
cutoff-parsing shape could not see.

**Load-bearing, verified by mutation:** restoring the old 5-minute
ceiling fails 3 of them; deleting the floor fails the new case.

## The remaining failure is a deliberately-deferred product decision

`agent-logs-and-monitor.pg.test.ts > aggregateActivityAnalytics …`
expects funnel stage `todo` count 2 and gets 0. This is **already
diagnosed and deferred by another worker**, in
`activity-analytics.ts:604`:

> *"The merged column landing in `triage` while the `todo` stage stays
empty is a SEPARATE and larger question — it makes the funnel show a
phantom 100% drop between Triage and Todo on every default board since
U11 — and it is deliberately not settled here. Changing which stage the
Planning column reports would retroactively alter how historical
analytics read… Flagged for a product decision on PR #2669."*

The merged Planning column carries `["intake","hold","reset-on-entry"]`
and `stageForTraits` prefers the earliest stage, so `intake` wins.
Either fix — remapping the stage, or changing the expectation — silently
settles how historical analytics read. I left it alone rather than pick
a side inside a test-repair PR.

## Two "flaky" files that are NOT flaky — and I nearly mislabelled them

`pg-test-harness-template-concurrency.pg.test.ts` and
`moves-intake-only-hard-cancel.pg.test.ts` each failed in one full-suite
run and not another, which reads as flake and would have earned a
quarantine entry plus a 14-day deletion clock under the standing rule.

Measured in isolation instead:

| File | alongside other PG suites | alone |
|---|---|---|
| `pg-test-harness-template-concurrency` | fails intermittently | **4
passed, 3/3 runs** |
| `moves-intake-only-hard-cancel` | failed once | **2 passed** |

So this is **shared-PG-template contention between concurrently running
suites**, not an inherent flake in either test. Quarantining them would
have started a deletion clock on healthy coverage and hidden a real
harness-parallelism interaction. Flagged for whoever owns the PG
harness; no quarantine entry added.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved duplicate-detection window handling with consistent defaults,
limits, and minimum values.
* Invalid tool output limits now produce standardized warning messages
while preserving fallback behavior.
* Updated settings and workflow validation to accurately reflect
supported configuration defaults and scopes.

* **Tests**
* Strengthened coverage for duplicate-detection windows and
configuration parity.
* Removed an outdated logging severity expectation tied to a
no-longer-applicable message.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 05:33:27 -07:00
gsxdsm
277a034e4b test(engine): two files missed the logger-mock debug sweep (38 red → 0) + a deleted API still pinned (#2716)
## What was red

| File | Failures | Error |
|---|---:|---|
| `notification/__tests__/notification-service.test.ts` | 26 |
`schedulerLog.debug is not a function` |
| `runtimes/__tests__/child-process-worker.test.ts` | 12 |
`runtimeLog.debug is not a function` |

`debug` is part of the logger surface (`logger.ts:25`) — the channel the
noisy-line demotion moved subsystem chatter onto, gated on
`FUSION_DEBUG`. A mock that omits it throws on the **first** demoted
call, failing every case in the file for a reason unrelated to what any
of them assert. These two were missed by the earlier sweep across 27
engine files.

## Measured

| Check | Result |
|---|---|
| the two files | 38 failed → **38 passed** |
| `debug` removed from the mocks again | **38 failed** — the entries are
load-bearing |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint` | clean |

Census unchanged — test files only.

## Two real drifts under the mock gap

Both re-pointed at what the product actually does, not relaxed:

**1. Suppression lines are `debug`, not `log`.**
`notification-service.ts` routes all five of its `"suppressed ..."`
messages through `schedulerLog.debug`. Two assertions looked on `.log` —
where the product no longer writes. Verified by grepping the product for
the message before editing the test, rather than assuming the mock was
the whole story.

**2. `centralCore.getGlobalConcurrencyState` is DELETED, not missing.**
The cross-project cap was dropped deliberately (`central-core.ts:2124`,
`FNXC:CapacityModel 2026-07-28-23:30`) together with
`updateGlobalConcurrency`, `acquireGlobalSlot`, `releaseGlobalSlot` and
the `concurrency:changed` event — capacity is two numbers **per
project** now. That comment also records the slot pair was already dead:
no production caller ever invoked it, so `currentlyActive` was never
incremented by real work.

The worker's stub provides `getLiveRunningAgentCounts` and
`recordTaskCompletion`, so the case is re-pinned to the former. It had
been pinning an API the product removed on purpose — the assertion would
have kept "passing" a shape that no longer exists if the stub had
happened to retain a same-named field.

## Flagged, deliberately not changed

`ipc/__tests__/ipc-host.test.ts` and `ipc/__tests__/ipc-worker.test.ts`
carry the **same incomplete logger mock** but are currently green (56
passed) because no demoted line is reached on their paths.

Adding `debug` there cannot be shown to fail today, so it is recorded
here rather than slipped in as an unfalsifiable edit. They go red the
moment any code they exercise demotes a line — which is how the 29 files
before them broke. Found by scanning every engine logger mock for the
pattern, not by guessing.
2026-07-30 05:18:19 -07:00
gsxdsm
3b618f2530 fleet: mission-execution-loop.ts 10 → 2 (one rule, five copies; and why these converted where store.ts's look-alikes could not) (#2711)
Claiming **`packages/engine/src/mission-execution-loop.ts`** (10). Every
one of its 10 census sites is the **same rule written five times**:

```ts
linkedTask.column === "done" || linkedTask.column === "archived"
```

## Census before/after

| | before | after |
|---|---:|---:|
| `mission-execution-loop.ts` | **10** | **2** |

Converted 4 of the 5 copies (8 of 10 sites) to the complete/archived
roles via core's `resolveTaskLifecycleColumns`. Each site already had
`this.taskStore` in scope inside an async method **and already had the
linked task fetched**, so the resolution rides along with a read that
was happening anyway.

## Why these converted where `store.ts`'s look-alikes could not (#2709)

Both read **another task's** column. The difference is not whose column
it is:

- **Here** — async methods, store at hand, one task per call. A
resolution is already affordable.
- **`store.ts`** — synchronous `filter` callbacks over a prefetched
`taskById` map, where per-dep resolution means N awaits inside a sync
predicate on a path that prefetches precisely to avoid per-item I/O.

Same-looking code, opposite verdicts. The distinguishing question is
**"is a resolution already affordable here"**, not "whose column is it"
— worth stating because a fleet worker pattern-matching on the receiver
alone would get both wrong.

## Not deduped, deliberately

The right shape is one predicate used five times rather than five inline
copies — and the FNXC comments above each copy show their intent has
already drifted apart. But introducing that predicate is a **new
abstraction**, which the fleet rules exclude, and it would fold five
reviewable substitutions into one design change. Flagged as the obvious
follow-up instead of smuggled in.

## Remaining: 2

The fifth copy, at the `hasLiveFixTask` site, where the terminal check
is one clause of a longer `Boolean(...)` expression whose other clauses
I would have had to reflow. Reviewability, not difficulty.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · the four mission
suites **150/150** · `pnpm lint` clean · engine `tsc` clean · `--strict`
exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 05:08:10 -07:00
gsxdsm
39523403c7 test(engine): isolate the fast-mode red — v1 seam→column normalization skips start (1 → 0) + an unanswered pause-guard question (#2719)
## The fast-mode failure, finally isolated

```
expected [ 'review' ] to deeply equal [ 'start', 'review' ]
```

This resisted **three earlier diagnosis attempts** because nothing about
the assertion, the fast-mode flag, or the node kind points at the real
mechanism: **column normalization of a v1 IR.**

**Isolated by mutation, not by reading.** Swapping the node's `config: {
seam: "review" }` for the sibling test's `config: { executor: "skill",
... }` makes `start` reappear in `visitedNodeIds`. So `config.seam` is
the trigger — which is not a thing the failure text suggests looking at.

The chain:

1. v1 normalization places nodes into synthesized default columns **by
seam** (`workflow-ir.ts:150` — `review` → `in-review`, seam-less →
`todo`).
2. The card rests in `in-progress`, which this three-node graph has **no
node for**.
3. `resolveColumnResumeNode` (`workflow-graph-executor.ts:473`)
therefore resumes at the next node **forward** — the review node —
instead of re-entering at `start`.

That resolver's `>=` comparison is commented for exactly this case: *"a
card can rest in a column the pipeline has no node for … and must then
resume at the next node forward."*

**So the product is right and the expectation was stale.** Visited is
`["review"]`. `pnpm test:gate` **726 passed**, lint clean, census
unchanged.

I also recorded in-file *why the sibling skill-executor case
legitimately still expects `start`*: its seam-less node normalizes into
`todo`, which is **behind** `in-progress`, so there is no forward match
and entry falls back to `start`. That contrast is an accident of config
rather than a deliberate difference — so the two expectations must
**not** be "aligned", which is the obvious-looking wrong move for the
next person here.

## Flagged, NOT fixed: `executor-prompt`'s 3 failures are a real design
question

These assert that `execute()` refuses to dispatch a user-paused row
during a global pause:

```
expect(mockedCreateFnAgent).not.toHaveBeenCalled();   // actually called 1 time
```

The test's own comment (from #2371) states the behaviour as "a paused
todo task is no longer dispatched at all". **The guard lives in the
scheduler, not in `execute()`** — `scheduler.ts:1515` is the
`globalPause` gate, and it is a hard stop that never reaches
`.execute(`. These three tests call `executor.execute(task)`
**directly**, so no guard fires.

That is not merely a test-layer mismatch, which is why I am not "fixing"
the assertions. `execute()` has **five call sites outside the
scheduler**:

- `executor.ts:3333`, `:3495`, `:5702`, `:5858` — internal re-dispatch
paths
- `runtimes/in-process-runtime.ts:2255` — `void
this.executor.execute(task)`

Whether each of those is separately gated determines whether work can be
dispatched on a user-paused row, or during a global pause, by a path
that never consults the scheduler. Two legitimate resolutions exist and
they differ in behaviour:

1. give `execute()` its own pause guard (defence-in-depth — but must not
break legitimate internal re-dispatch), or
2. retire the direct-`execute()` assertions and cover the invariant at
the scheduler layer.

Picking either silently inside a test-repair PR would either change
lifecycle behaviour around **user pause** — a safeguard this program has
re-ratified and told me not to narrow — or delete coverage of it. So it
stays flagged with the call-site evidence for whoever owns the pause
contract.

This is the same file and question I flagged much earlier in the
program; it is now backed by the specific line numbers rather than a
suspicion.
2026-07-30 04:41:06 -07:00
gsxdsm
a59576607a test(engine): 11 reds from two intended product changes the tests still pinned (#2717)
## The 11 failures, two causes

Four files, all confirmed red on clean `origin/main`, all unowned.
Neither cause is a product defect.

**A. U11 merged the two pre-implementation columns.** `builtin:coding`,
`builtin:stepwise-coding` and `builtin:brainstorming` now declare
`todo,in-progress,in-review,done,archived` — **no `triage`**. So the
entry column is `todo`, the former `triage → todo` graph hop no longer
exists (there is no boundary to cross), and the replan rebound
(`executor.ts:4392`, via `resolveReboundColumnFor`) targets `todo`.

Verified by resolving each built-in IR and printing its column ids, not
inferred from the failure text.

**B. `maxPostReviewFixes` was raised 3 → 10** via
`DEFAULT_MAX_POST_REVIEW_FIXES` (`builtin-workflow-settings.ts:555`).
Two files still pinned 3.

| File | Before | After |
|---|---:|---:|
| `workflow-graph-optional-step-fix` | 5 failed | **39 passed** |
| `builtin-workflows-lifecycle` | 3 failed | **94 passed** |
| `agent-tools-intake-column` | 2 failed | **4 passed** |
| `workflow-settings-fallback-alignment` | 1 failed | **3 passed** |

`pnpm test:gate` **726 passed**; lint clean; census unchanged (test
files only).

## Three choices so these don't re-break on the next rename

Swapping `"triage"` for `"todo"` everywhere would have worked and been
wrong in three places:

1. **The replan log assertion no longer embeds a column id.**
`executor.ts:5366` *interpolates* the resolved column into the message,
so a literal there pins a column name inside prose — guaranteed to break
again. It now matches the message shape plus the attempt/budget counter,
while the destination column stays pinned by the `moveTask` assertion in
the same test.
2. **The budget case drives off the imported
`DEFAULT_MAX_POST_REVIEW_FIXES`.** That constant exists *because* the
declaration default and two inline literal `3`s had already drifted
apart once — its own comment says raising the declaration alone "would
have left every unset-settings path on the old value". A third copy in a
test would repeat exactly that mistake.
3. **`agent-tools-intake-column`'s two guards are RE-PINNED, not
deleted.** They hold the default workflow's landing column stable and
fired on an intended change, so they still have a job. Deleting them
removes the only check; leaving them on `triage` pinned a column that no
longer exists.

## What I deliberately did not touch

`builtin-workflows-lifecycle` has **18** expectations and only **3**
were failing. I changed only those three: the other 15 pass unchanged,
which proves those workflows genuinely still declare `triage`. A blanket
rename across the file would have broken them — the same trap that bit
me earlier in `starved-refinement`, where renaming a shared fixture
default broke a test that had been passing.

Similarly in `workflow-settings-fallback-alignment`: case (b), which
scans engine source for literal `?? <n>` fallbacks, **already passed** —
there are no literal fallbacks left for that key because the read sites
import the constant. Only the human-readable audit table had lagged.
That table is the drift detector, so updating it keeps it honest rather
than pinning a value the product abandoned.

## Still red in this area, not in this PR

`executor-fast-mode-workflows` (1) and `executor-prompt` (3) remain.
They are not column- or budget-drift; `executor-prompt`'s three call
`execute()` directly and raise a real design question (whether
`execute()` should carry its own pause guard, or whether those
direct-call assertions should retire) that I am not answering silently
in a test-repair PR.
2026-07-30 04:28:50 -07:00
gsxdsm
3577cb6adf fleet: project-engine.ts 12 → 5 — auto-merge silently declined every card on a renamed board (#2706)
**Claim announced before the work** (on #2689, alongside
`register-task-workflow-routes.ts`):
`packages/engine/src/project-engine.ts`.

**12 → 5.** Repo backlog → **685**.

## The failure mode here has no error signature

Every merge guard in this file spelled the lane `in-review`. On a
renamed board nothing throws, nothing logs a warning — **auto-merge
simply declines every card**:

| guard | what a renamed board gets |
|---|---|
| `requestInterpreterMerge` | returns `noOp: true` — *"parked cleanly in
review, awaiting human merge"* — for a card that was in review and fully
eligible |
| the merge-queue snapshot | returns an **empty list** for a queue full
of review cards, so the coordinator sees nothing to admit |
| the `taskMoved` auto-merge handoff | never fires, so nothing reaches
auto-merge in the first place |
| the pause-interruption tracker | drops every card from its
paused-review set on the next update, so a merge paused mid-flight is
never interrupted |

The operator sees cards resting in review with auto-merge **on**, and
every log line says the system did the right thing. There is no string
to search for — which is the argument for the census being a parse
rather than a grep over error messages.

## Implementation notes

- **Core's `resolveTaskLifecycleColumns` directly** — the canonical
helper, so no new abstraction and no fourth local resolver in a file
that had none.
- **The merge-queue snapshot resolves per task through a shared
`irCache`**, because a merge queue can hold cards from *different*
workflows. Per-workflow, not per-card: one IR read each.
- **The handoff and its post-grace recheck share one snapshot.** They
are halves of one decision — "did this card just enter the merge lane,
and is it still there?" — and that is exactly the split that produced
the defects in `executor.ts`.

## Revert proof

**1 of 3 cases reddens** with the literal restored.

The suite invokes the real `requestInterpreterMerge` via `.call()` on a
minimal `this` (`runtime.getTaskStore`, `allowInReviewMergeProcessing`,
`onMerge`) instead of standing up a whole `ProjectEngine` runtime. The
body under test is the shipped one, and *reaching* `onMerge` is the
assertion. I would rather explain that seam than either skip the proof
or spend the test budget booting a runtime. The default-board case is
labelled in the file as no-change evidence, not counted as coverage.

## The remaining 5, flagged not guessed

All five are `column === "done"` **merge-confirmation reads** — "did the
merge land?". That is a different question from any lane role, and it
shares its answer with the dependency guards I flagged in `executor.ts`
(`12325`) and `register-task-workflow-routes.ts` (`3995`). Three files,
one open question: **what does "landed / satisfied" mean on a board
whose terminal column is not named `done`, and is it the complete column
or the terminal union?** Deciding it once and applying it to all three
is right; swapping it three times independently is how the resolver
choice ends up inconsistent — which already happened once inside
`executor.ts`, where the correct resolver inverts between two guards a
few lines apart.

## Verification

`pnpm test:gate` **158 / 10 / 487 / 71** · **56/56** across the six
auto-merge / project-engine suites · 3/3 in the new suite · `tsc -p
packages/engine` clean · `pnpm lint` clean · census `--strict` exit 0.

No changeset: `@fusion/engine` is private, and the behaviour change is
confined to renamed boards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 04:25:19 -07:00
gsxdsm
0c7dc8c8ae feat(census): report MIXED-VOCABULARY files — the shape behind four half-conversion findings in one day (#2704)
## The pattern

Four review findings dispatched to me in a single day were the **same
defect**: a guard converted to role resolution while the function it
*feeds* still filters on the literal. The resolved guard admits a custom
column, the literal collaborator rejects it, and **nothing errors** —
the endpoint returns `repaired: 0` and reads as converted.

| PR | resolved side | literal collaborator |
|---|---|---|
| #2700 | review guard | `reconcileInReviewBranchRebind` filters `===
"in-review"` |
| #2700 | retry guard | `isInReviewMissingWorktreeSessionStartFailure`
likewise |
| #2698 | role-aware tabs | reconciliation effects still compare
`"done"` / `"in-review"` |
| #2688 | role-derived flags | memos and a `useState` capture keyed on
the stale value |

Since opening this I have been handed **two more** of the identical
shape (#2701, #2702). It is not a coincidence; it is what a conversion
phase produces by default.

## What this adds

A file where **both vocabularies are live** is where that can happen, so
the census now names those files.

**Measured: 23 of 134 guard-bearing files, holding 311 of 686 guards** —
and the top of the list is exactly where the findings landed:

```
MIXED-VOCABULARY files (a role resolver AND legacy literals): 23, holding 311 guards
   110  packages/engine/src/self-healing.ts
    57  packages/engine/src/executor.ts
    26  packages/engine/src/scheduler.ts
    20  packages/dashboard/src/routes/register-task-workflow-routes.ts
```

## Report-only, deliberately

A partially converted file is the **expected** state during a conversion
phase. Gating this would punish correct in-progress work and would be
routed around within a day. What it buys is that a reviewer of a listed
file knows to check the collaborators of anything converted — which is
what this repo's **Surface Enumeration** rule already requires, and what
each of those PRs missed.

The rule exists. The fleet work order does not mention it, so reviewers
are catching these one site at a time.

## Verification

Five tests, both directions: flags a mixed file; does **not** flag a
fully literal one (or the entire backlog lights up and the signal
carries no information); does **not** flag a fully converted one; does
not match a resolver name inside a longer identifier (the
`hold`-inside-`threshold` trap from #2677); survives an unreadable file.
**Mutation: dropping the resolver condition fails 2 of 38.**

The helper lives in the **lib**, not the CLI — importing the CLI
executes it and calls `process.exit`, so nothing defined there is
reachable from a test. I found that by trying.

38 census tests green · `--strict` and `--compare` exit 0 · lint clean ·
gate green (487 + 158 + 10 + 71). **No census numbers change.**
2026-07-30 04:02:19 -07:00
gsxdsm
a037ca93c7 test(engine): un-red the reliability-interactions tier (5 → 0) + a guard that could not fire (#2707)
## How this was found

Full-suite **shard 3/4** reports no `Tests N failed` summary at all —
its log ends mid-`@fusion/engine [1/2]` on a watchdog heartbeat, so the
red reads as infrastructure noise. It is not.

Facts that ruled out the infrastructure explanations, before touching
any test:

- watchdog budget is **1500s**; the engine slice died after **~6.5–10.5
min** (varies run to run) — not a timeout, and not a fixed one
- job `timeout-minutes: 60`, ran **9.7 min** — not the job timeout
- `concurrency.cancel-in-progress: false` — not cancellation
- annotation says **exit code 1**, not 137 — not an OOM kill
- across **four consecutive runs** the last file named is always
`reliability-interactions/explicit-duplicate-marker-sweep.test.ts`

Running that file locally reproduces real failures. The summary line is
simply missing from the CI log's final chunk (the last ~5s of output
never appears), which is what disguised a normal test failure as a
crash.

## Four causes, none a product defect

| File | Failures | Cause |
|---|---:|---|
| `explicit-duplicate-marker-sweep` | 2 | fixture seeds `column:
"triage"` |
| `starved-refinement-x-approval-gate` | 1 | same |
| `starved-refinement-x-triage-poll` | 1 | same |
| `executor-pending-review-skip-retry` | 1 | review handoff now passes
move **options** |

`triage` is no longer declared on any workflow post-U11, and these
sweeps filter by **role** — so cards seeded there carried no intake role
and the sweeps reported 0.

## Measured

| Check | Result |
|---|---|
| `reliability-interactions` | 5 failed → **0** (103 files, **530
passed**) |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint`, engine `tsc --noEmit` | clean |

Census unchanged — test files only.

## The real find: "honors the disable flag" could not fail

Forcing `enabled = true` in `resolveExplicitDuplicateMarkerTasks` — i.e.
making the sweep **ignore the disable flag entirely** — left all 16
cases **green**.

The fixture never set `triageDuplicateResolution`, so the sweep had no
resolution action to take and the duplicate survived whether the flag
was honoured or ignored. The assertion held for a reason unrelated to
the test's name. It had been red only because of the column literal,
which would have made "fix the literal, go green" a repair that left a
guard guarding nothing.

Setting the resolution mode makes that same mutation delete the
duplicate and the case fail. Verified both directions:

| | mutation applied |
|---|---|
| before | 16 passed — **guard cannot fire** |
| after | **1 failed** / 15 passed |

Recorded in-file with the measurement, so nobody strips the setting back
out as redundant.

## Two assertions strengthened rather than relaxed

- The disable-flag and failed-delete cases now assert the column is
**unchanged from a value read before the sweep**, rather than equal to a
literal. They cannot pass because a seed happened to land where the
assertion looked, and they survive the next column rename. The
invariants those cases own are "the flag stops the sweep" and "the task
whose delete threw survives" — the column id was always incidental.
- The handoff move asserts its **provenance** (`nodeId:
"review-pending-handoff"`, `preserveProgress: true`) instead of the
`expect.anything()` its siblings in that file use, so a move to the same
column by another path cannot satisfy it.

## Still open for whoever owns CI

Shard 3/4's log loses its final chunk, which is why a plain test failure
presented as a crash and stayed unexplained across at least four runs.
Anyone triaging full-suite from shard conclusions alone will keep
mis-reading this one; the failure has to be reproduced locally to be
visible.
2026-07-30 03:59:13 -07:00
gsxdsm
339f6e7830 fix(census): stop the baseline serialising the fleet — every fleet PR conflicted with every other one (#2699)
## The problem

Every fleet PR conflicts with every other fleet PR in
`lifecycle-column-census-baseline.json` — **even when they convert
entirely different files**. I have rebased **six** of my own branches
for nothing but this file, and the resolution was *always* "take main's,
re-run `--update-baseline`". Never once a real merge.

That makes a generated artifact the serialisation point for the whole
fleet phase.

## The cause

`totals`, `byColumnId`, `properties` and `queryByColumnId` are
**derived** — recomputable from the per-file maps — and **`--strict`
never reads any of them**. It compares `byFile`, `deliberateByFile` and
`queryByFile`, and nothing else.

But every conversion changes at least one aggregate line. So those lines
were a **shared write on a file whose real content is per-file and
disjoint**. Removing them, two PRs converting different files touch no
common lines.

## Trade-off, stated because it undoes a deliberate choice

An earlier note kept the totals in the pin *"so the new number lands in
the diff where a reviewer sees it"*. That was a good reason. The signal
survives elsewhere:

- the CLI prints the totals on every run;
- `--update-baseline` prints each tightened entry by name;
- the fleet rules already require a census before/after **in the PR
body**.

Reversible if the diff-visible number proves to matter more than the
conflicts.

## Cost, stated too

Merging this makes every in-flight fleet PR re-record once. That is one
more instance of an operation they are already performing on every
rebase — a one-time cost against a recurring one.

## Verification

The end-to-end test that asserted the write via `totals.column` now
asserts the same claim via the per-file entry: the stale pin says 1, the
rewritten pin must carry the tree's real higher count for that file.
**Mutation: suppressing the `--update-baseline` write still fails it**,
so the assertion did not weaken.

71 census tests green · `--strict` and `--strict --exact` exit 0 · lint
clean · gate green (487 + 158 + 10 + 71).

## Not done

A merge driver. `.gitattributes` can name one, but registering it needs
`git config` per clone and this repo has no `postinstall`/`prepare` hook
to do that — so it would silently not apply for most people. Removing
the shared lines fixes the conflicts without needing any local setup.
2026-07-30 03:56:12 -07:00
gsxdsm
78b6b5ba37 fleet: packages/engine/src/executor.ts 85 → 57 (in progress; 4 batches, plus the structural measurement this cluster needs) (#2689)
**Claiming `packages/engine/src/executor.ts`** — the largest unclaimed
cluster (self-healing.ts and scheduler.ts are taken).

## Census

| | before | after |
|---|---:|---:|
| `executor.ts` | 85 | **75** |
| repo total | 722 | **712** |
| `done` | 195 | 190 |
| `archived` | 147 | 142 |

Baseline re-recorded in the same commit; it shrinks by exactly the
converted count (10 literals across 5 sites).

## Batch 1 — terminal-lane guards

Five identical *"this card is already finished, refuse"* guards, all the
literal pair `live.column === "done" || live.column === "archived"`. On
a renamed board neither matches, so the refusal falls through — the same
inert-guard shape as #2670.

Converted to `resolveTerminalColumnsFor`, **the helper this file already
established** at line 4509 — no new abstraction. It unions the resolved
terminal columns with the legacy pair, so each converted guard is a
strict **superset** of the literal: it can refuse in more cases, never
fewer. That is what makes this batch safe without per-site behavior
review.

## The structural measurement this cluster needs

#2683 found self-healing.ts unsafe to batch because of **sync** workflow
reads — a converted guard there would resolve through a sync path that
cannot resolve a selection in production, silently falling back to
defaults. I measured whether executor.ts has the same problem, per guard
(not per line):

| context | guards |
|---|---:|
| **async** — safe, can `await resolveWorkflowIrForTask` | **71** |
| **sync** — needs threading or is not convertible in place | **14** |
| module scope | 0 |

The 14 sync-context guards are at lines 3455, 3479, 3530, 3540, 4611,
5501, 5502, 5504, 5777, 10213, 12306 (×3), 15782 — `in-progress` 4,
`in-review` 4, `archived` 3, `done` 2, `todo` 1. **I am not converting
those in place**, and I will flag rather than guess if threading
resolved data changes behavior.

So: unlike self-healing, this cluster is **83% safely convertible**,
which is why it is worth working as a batch.

## Note on #2685

Engine code converts through core's resolvers
(`resolveLifecycleColumns`, `resolveTerminalColumns`), not the dashboard
`columnRoles` helpers. So the 680-guard helper gap #2685 fixes is
**dashboard-side** — this cluster is not blocked on it.

## Verification

engine `tsc` clean · lint clean · gate green (487 + 158 + 10 + 71).

**Pre-existing failures, not caused by this change:** five tests in
`src/__tests__/reliability-interactions` fail, all in
`SelfHealingManager.recoverStarvedRefinementTriageTasks`. I confirmed by
stashing this change and re-running on a clean tree — they fail there
too. This change touches only `executor.ts` and does not go near that
path. Flagging rather than fixing: it is someone's cluster and not mine
to alter mid-flight.

## Not done

Batches 2+ (the remaining 75). I will keep working this file in this PR
with small commits, per the fleet rules.
2026-07-30 03:25:21 -07:00
gsxdsm
3d28e264c1 test(self-healing): three suites seeded a column id the product stopped declaring (7 red → 0) (#2695)
## What was red

`self-healing-advanced-triage`, `self-healing-agent-link-drift`, and
`self-healing-starved-refinement` — **7 failed / 19 passed** on clean
`origin/main`. All three fail the same way: the sweep returns `0`
recoveries where the test expects `1`.

## Cause

Those three sweeps were converted from `listTasks({ column: "triage" })`
to **role** filters (`filterByPreWipRole(..., ["intake"])`). Each store
fake has no workflow-selection readers, so the sweep resolves the
**default IR** — in which `triage` is not a declared column at all. The
seeded cards carried no intake role, the filters returned no candidates,
and the sweeps did nothing.

The product change was correct; the fixtures were asserting against a
column id the product had stopped declaring. `self-healing.ts` even
warns about this exact shape in its own comment: *"Converting a
predicate while leaving its source query on a literal produces a sweep
that LOOKS converted and does nothing."* The fixtures are the mirror
image — a seed on the old literal makes a correctly-converted sweep look
broken.

**Verified, not assumed.** I resolved the default IR and printed the
flags rather than reasoning from column names:

```
todo        {"intake":true,"hold":true,"resetOnEntry":true}
in-progress {"countsTowardWip":true,"abortOnExit":true,"timing":true}
in-review   {"mergeBlocker":true,"humanReview":true,...}
done        {"complete":true}
archived    {"archived":true,"hiddenFromBoard":true}
```

`todo` is the intake column post-U11 (the merged Planning column).
`triage` does not appear.

## Measured

| Check | Before | After |
|---|---|---|
| the three files | 7 failed / 19 passed | **26 passed** |
| all 40 `self-healing-*` files | 3 failed files / 7 failed tests | **40
passed / 695 tests** |
| `pnpm test:gate` | — | **726 passed** |
| `pnpm lint` | — | clean |

Restoring the deleted column id reintroduces the failures, so the seeds
are load-bearing rather than cosmetic.

**Pre-existing, untouched:** the same 9 vitest *"unhandled errors"* and
the identical non-zero exit appear on clean main. Verified by reverting
only these three test files and re-running — unrelated to the fixtures,
so flagged rather than folded in.

## One deliberately surgical spot

In `starved-refinement` only the starved refinement's own seed moved. My
first pass renamed the file's default seed and **broke a test that was
passing** — the auto-approve-all case calls `recoverApprovedTask`
directly (no role filter) and asserts a move **into** `todo`, so seeding
it in `todo` makes that move degenerate. I reverted and moved one seed
instead.

That leaves a real question I did **not** answer: post-U11 intake and
hold are one column, so a `triage → todo` move may no longer encode
anything. Deciding that means changing what the test is *about*, which
is a behaviour judgement on someone else's assertion. **Flagged in-file,
not guessed.**

Also worth noting for whoever converts the remaining backlog: line 43 of
that file pairs a `triage` seed against an explicit `todo` seed as a
contrast. U11 collapsed those two columns, so the contrast the fixture
was drawing no longer exists in the product — not a defect, but any
fixture built on intake-vs-hold being distinct is now suspect.

Census unchanged (722) — test files are not scanned by the census.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Updated self-healing workflow test scenarios to reflect the current
intake column.
* Improved test coverage for triage, agent-link drift, and starved
refinement handling.
* Added documentation clarifying workflow column and role-based
filtering assumptions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 03:25:06 -07:00
gsxdsm
dfbab18fd5 fix(scheduler): a renamed wip column held NO file-scope lease — two agents could edit the same files (#2693)
## The defect

`activeScopes` is the file-scope lease registry the dispatch path reads
(`scheduler.ts:2167`) to decide whether a candidate overlaps work
already in flight. Two column-id literals kept it empty on any board
whose columns are renamed:

1. the lease loop gated on `task.column !== "in-progress"`;
2. `shouldHoldActiveFileScopeLease` keyed **both** its branches on
`in-progress` / `in-review`, so it returned `false` for *every* card on
a renamed board.

Forty lines above that loop, the same sweep resolves `countsTowardWip`
from the workflow IR for capacity arithmetic. **The scheduler was
simultaneously right about capacity and wrong about leases.**

Consequence: a second task sharing a file scope **dispatched instead of
queueing** — two agents editing the same files, which is precisely what
`groupOverlappingFiles` exists to prevent.

## Measured, differential

Same workflow *shape* under two vocabularies with identical traits; only
the column ids differ, so any difference is attributable to a surviving
literal. No renamed id collides with a legacy one, so a surviving `===
"in-progress"` cannot pass by luck.

| | default vocabulary (control) | renamed vocabulary |
|---|---|---|
| fix reverted | queued on lease ✓ | **dispatched into the wip column**
✗ |
| fix applied | queued on lease ✓ | queued on lease ✓ |

`2 of 3 fail` reverted → `3 of 3 pass` applied. The control passes on
**both** sides, so a change that breaks overlap protection generally
cannot hide behind this test.

I checked the test wasn't vacuous before trusting it: instrumented the
run to print the actual `moveTask` calls, and confirmed the renamed case
really produced `[["FN-CAND","building"]]` — a genuine dispatch — rather
than the candidate simply never being considered. Both failure modes
look identical in the assertion.

## Why optional booleans, not a flags object

`shouldHoldActiveFileScopeLease` is **exported** and shared with the
self-healing / repair paths (`self-healing.ts:4488`, `:5406`) — its own
comment says those "must use this same predicate so stale
`overlapBlockedBy` cleanup does not preserve blockers the scheduler
would ignore". So the role questions became optional parameters that
**default to today's literals**: a caller that resolved the traits
passes the answer, a caller that has not gets exactly current behaviour.
No existing call site changes meaning, and no dependency on #2690.

## Verification

| Check | Result |
|---|---|
| scheduler / capacity / hold-release / overlap / self-healing | **59
test files green** |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint`, engine `tsc --noEmit` | clean |

`self-healing-advanced-triage`, `-agent-link-drift`,
`-starved-refinement` are **7 failed / 19 passed both before and after**
— verified pre-existing on clean `origin/main` by reverting only
`scheduler.ts` and re-running. Flagged, not fixed, and not in scope
here.

## Census

**722 → 721**, `scheduler.ts` 28 → 27. Baseline re-recorded in the same
commit.

To be precise about what that −1 is: the *loop* literal is gone, while
the two literals **inside** the predicate remain by design as the
documented defaults. So this is not "scheduler is now trait-aware" — it
is one site, plus the seam that lets callers be.

## Merge-order note

**#2690 also records `scheduler.ts` 28 → 27**, converting a *different*
site (`isWipColumnTask`'s hand-rolled flags-first copy, `:1690`). The
two are independent and do not double-count: if both land,
`scheduler.ts` is **26**, and whichever merges second will conflict on
`scripts/lib/lifecycle-column-census-baseline.json` and must re-record
to 26 rather than resolve to 27. Flagging so the merger does not take
one side blindly.

## Still broken, flagged for an owner

The **in-review** half. `activeScopes` is also populated for review-lane
cards via `t.column === "in-review"` (`scheduler.ts:1751`, `:1757`), and
this PR leaves those literals in place: the sweep's flags map holds only
`countsTowardWip`, so no review-role answer is available to pass in.
Fixing it needs the flags-object change in #2690, after which the same
optional parameter added here carries it. Until then a renamed review
column still holds no lease.
2026-07-30 03:19:02 -07:00
gsxdsm
e9e63d8e0f consolidate/capacity: --strict was red on main (my #2621), 14 stale baselines, routines seeding a deleted column, worktrees-off audit (#2652)
Capacity unit consolidation. Three coherent themes, small commits
inside.

## Census before/after (`node scripts/lifecycle-column-census.mjs`)

| | before | after |
|---|---:|---:|
| triage column guards (the bar) | 10 | **10** |
| `--strict` on main | ❌ **RED** | ✅ green |
| baseline staleness | 14 files stale | **0** |

This branch does **not** move the triage bar — its remaining 10 are
moves.ts (dies with the flag), the dashboard cluster, and one deliberate
site. It fixes the instrument that measures the bar, plus a live defect
the comparison count cannot see.

---

## 1. `--strict` was RED on clean `origin/main`, and it was my fault

```
packages/dashboard/src/routes/register-task-workflow-routes.ts: 22 -> 23
```

My merged #2621 added a v1-IR pre-WIP fallback answering a greptile P1
and shipped no marker or baseline update, so the program's measuring
instrument has been failing on main since it landed.

Fixed **at the site** with a `DELIBERATE-LITERAL` marker, not by bumping
the baseline. That branch runs only when the IR declares no columns and
no nodes, so there is no role to resolve — `resolveLifecycleColumns`
returns nothing and the legacy pre-implementation ids are the only
pre-WIP signal that exists there. It is *unconvertible*, not unfinished;
the sibling `else` two lines down is the trait path for every IR that
can answer. A rise that is genuinely correct belongs where a reader will
see it.

## 2. The baseline was stale for 14 files — a hole, not cosmetics

A stale allowance lets converted guards return while the check stays
green. Measured gaps:

```
self-healing.ts          allows 126, tree has 111
executor.ts              allows 112, tree has 104
moves.ts                 allows  44, tree has  39
default-workflow-hooks   allows  25, tree has   7
mission-feature-sync     allows   5, tree has   0
MissionControlPanel      allows   4, tree has   0        (+8 more)
```

**Only two of the fourteen are mine.** The other twelve are
already-merged conversions by other workers where nobody re-recorded.
Re-recorded all fourteen here rather than waiting for twelve PRs,
because until it happens the ratchet is not holding the 779 it exists to
hold. Flagging it plainly: those drops are other people's work being
locked in, not mine being claimed.

## 3. Routines created tasks into the column U11 deleted

The routine editor's "Target Column" defaulted to `triage`. That value
is submitted as the create step's `taskColumn`, and an **explicit**
column bypasses the workflow entry-column resolution added for
column-less creates (#2589) — so every routine saved with the untouched
default seeded its tasks into a column the board does not declare.

Defaulting to `todo` would be the same mistake one column over: a custom
workflow declaring no `todo` is seeded into an undeclared column just as
surely, because an explicit column overrides entry resolution whatever
its value. So the default sends **nothing** and each workflow's own
intake resolution decides.

The `triage` **option** is removed too, not merely un-defaulted — fixing
the initializer alone left the operator able to pick the deleted column
one click away, and it was the option labelled "Planning", the name the
merged `todo` column now displays. Removing it retires that label
inversion as well.

Found by scanning **membership** forms rather than comparisons: the
comparison census cannot see a `?? "triage"` default, so no count showed
this and nobody was looking. Revert-proof — restoring the default fails
with *"the default must not name a column at all"*.

## 4. "Worktrees off is INERT" had one unaudited reader

The constraint was that `maxWorktrees` become genuinely inert, "not set
very high and not skipped by convention". `resolveWorktreeCapacityLimit`
returns `null` for that, and its unit tests can only prove the
**resolver** is right — they cannot see a second reader, which is the
only way the constraint breaks.

Audited every `maxWorktrees` read that bounds anything. **Exactly two:**
`scheduler.ts` (the admission gate, via the resolver, single call site,
optional gate snapshot) and `self-healing.ts`'s `enforceWorktreeCap` —
`(settings.maxWorktrees ?? 4) * 2`, a **raw** read.

The second is **not a bug** and is left alone: it bounds worktree
*directories on disk* and only removes *idle* ones. Worktrees still
exist in OFF mode, so that bound must keep applying or idle directories
accumulate unbounded. Recorded consequence: in OFF mode the number still
governs disk retention while gating no admission — an edge you scoped
out. The note says explicitly **not** to unify the two readers: routing
hygiene through the resolver returns `null` in OFF mode and silently
removes the disk bound, which is a leak dressed as a simplification.

New ratchet requires every file bounding on `maxWorktrees` to be named
with a reason, and rejects a **stale** allowlist entry. Proven by
injecting `active >= (settings.maxWorktrees ?? 4)` into
`hybrid-executor.ts`.

---

## Deliberately NOT included

- **My own census script.** #2633 landed the canonical one, and it is
better than mine — an AST classifier *plus* an independent text
classifier with `--compare`, and a baseline that fails on unrecorded
**drops** as well as rises. Mine only caught rises. I deleted mine
rather than ship a second measuring instrument; three copies of "strip
comments" is the drift shape this program keeps paying for, so the
worktree ratchet now imports #2633's `stripComments`.
- **My TaskContextMenu fix.** Superseded, and by a better answer: main's
`isPureIntakeColumn` (intake *without* hold) keeps the merged Planning
column shown and suppresses only a bare Ideas capture, which resolves
the exact hold-lane objection coderabbit raised against my version. I
briefly clobbered that merged work by checking my old file out
wholesale, caught it in the diff, and reverted.

## Verification

`pnpm lint` clean · core + dashboard `tsc` clean · census suite 23/23 ·
worktree ratchet 8/8 · RoutineEditor 49/49 ·
`routes-task-retry-planning-column` 16/16 · `lifecycle-column-census
--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Added after review (all four greptile threads were real, and two of
them mattered)

**The routine fix was half a fix.** `routine-runner.ts:515` *and*
`cron-runner.ts:982` both did `column: (step.taskColumn as Column) ||
"triage"` **after** the step is read, so every routine — including ones
saved through the fixed editor — still created tasks into the deleted
column. Both now omit it.

**The advanced steps editor MANUFACTURED the defect.**
`ScheduleStepsEditor.tsx` had three `triage` defaults: the new-step
template (`:64`), the per-step initializer (`:95`), and the select still
offering it (`:344`). So the path I had *not* fixed produced the bug by
default, on fresh data. Template names no column; initializer coerces a
persisted `triage`; `triage` removed from the options; empty submits
`undefined`.

**Four pre-existing tests pinned the defect** and are rewritten to the
corrected invariant rather than appeased:

| test | asserted |
|---|---|
| `cron-runner`: "defaults column to triage when taskColumn is not set"
| `column: "triage"` |
| `ScheduleStepsEditor`: "adds a create-task step..." | `taskColumn`
toBe `"triage"` |
| `ScheduleStepsEditor`: "allows saving create-task step..." | the
legacy column is **resubmitted** |
| plus the explicit-column case added beside each, so the fix cannot
swallow a deliberate choice |

**The allowlist hole was the worst finding.** `AUDITED_BOUNDS` was keyed
by FILE, so every bounding expression in an allowlisted file was exempt
— a second raw bound in `scheduler.ts` stayed green, the one case that
ratchet exists for. Per-expression now, and making it so **immediately
surfaced a real second bound the file-level version was hiding**
(`maxWorktreesGate.used >= maxWorktreesGate.limit`, safe by construction
since the snapshot is `undefined` in OFF mode). Proven by injection.

## Found while re-reading my own deletion, not reported

A **rendered tooltip** still named a deleted cap. The "Queued to plan"
badge read *"planning starts when a concurrency slot frees up
(maxConcurrent / globalMaxConcurrent)"*. The cross-project cap is gone —
capacity is two numbers per project — so it told operators their
planning waited on a limiter they can no longer find a setting for.
Names the surviving dimension only now.

## Coding (Ideas): enforcing #2651 rather than repeating it

I took the unowned coding-ideas IR merge, concluded it must not be done,
then found **#2651 had already implemented, reverted and documented
exactly that** — with better grounding than my own argument. It added no
test, so nothing stops the next person reaching the same dead end.

So this ships their reasoning as a ratchet, not a second opinion: triage
discovery keys on the column's `autoTriage`, so a merged column is
either never scanned (cards sit on a bootstrap stub until the **capacity
hold** releases them, sending **unplanned** work into in-progress —
worse than stalling) or scanning wins and the manual gate is gone. Their
scope caveat is kept: `autoTriage` is a general trait field, so only
*this preset's* collapse is dead, not manual intake as a concept. The
registry does not reject the merged shape, which is why prose was not
enough.

## Verification (re-run)

`pnpm lint` clean · core + engine + dashboard-app `tsc` clean ·
`lifecycle-column-census --strict` exits 0 ("every file matches its
baseline exactly") · routine-runner 24/24 · cron-runner 156/156 ·
ScheduleStepsEditor 41/41 · RoutineEditor 49/49 · worktree +
coding-ideas 12/12. TaskCard has 2 failures **pre-existing on main** —
confirmed identical with my changes stashed.

---

## Bears directly on the closing bar: this PR already removes the
67-guard ratchet slack

Measured on current `origin/main` with the census itself:

```
tree total: 787   baseline total: 854   SLACK: 67

FILES ABOVE BASELINE (1):
   +1  packages/dashboard/src/routes/register-task-workflow-routes.ts  (22 -> 23)

FILES BELOW BASELINE: 13, totalling 68 unrecorded conversions
   -18  core/default-workflow-hooks.ts (25->7)   -15  engine/self-healing.ts (126->111)
    -8  engine/executor.ts (112->104)             -5  core/task-store/moves.ts (44->39)
    -5  engine/mission-feature-sync.ts (5->0)     -4  core/live-agent-count.ts (10->6)
```

**The slack is not regression — it is 13 files of merged conversions
nobody re-recorded**, against exactly **one** rise. This PR re-records
the baseline **854 → 782 across 140 files**, which closes it.

**And the "+3 that slipped in" is +1, and it is mine.**
`register-task-workflow-routes.ts 22 → 23` is the v1-IR pre-WIP fallback
my #2621 added; it is justified (that branch runs only when the IR
declares no columns or nodes, so there is no role to resolve) but it
shipped with no marker and no baseline update — which is why `--strict`
has been **red on main since it merged**. Fixed here at the site with a
`DELIBERATE-LITERAL` marker rather than by bumping the baseline, because
a rise that is genuinely correct belongs where a reader will see it.

Sequencing note for the auto-lowering change: if this lands first, that
work is purely the mechanism (auto-lower, or fail with tighten
instructions) rather than a cleanup, and the two re-records will not
collide in the same file.

Also worth carrying into that mechanism, from building the same guard
here: **`--update` must refuse to RAISE.** An earlier version of mine
wrote current counts verbatim, so a developer who added a literal and
ran the documented update command locked the regression in as the new
ceiling — the mirror of the high-water problem. Lowering can be
unattended; raising should be a hand edit with the reason recorded.

## Third piece of residue from my own deletion

`updateGlobalConcurrency` in the dashboard API client PUT to
`/api/global-concurrency`, a route removed when the machine-wide cap
went. Zero callers; the only reference was the `legacy.ts` barrel
re-export. Deleted both. `fetchGlobalConcurrency` **survives on
purpose** — the GET route remains and serves live utilization telemetry
to the footer and Command Center; nothing gates on it.

That is the third: after the second raw `maxWorktrees` reader and the
"Queued to plan" tooltip. A deletion is not finished when the
enforcement goes — the client, the label and the tooltip outlive it.

---

## Re-greened the dashboard API tests: 117 failures on main, ONE root
cause

These would have polluted the closing verification pass, and nobody
owned them.

`api()` builds headers via `new Headers(...)` and returns
`Object.fromEntries(headers.entries())` — and `Headers.entries()`
**lowercases every key**, so the object reaching `fetch` is
`content-type`, not `Content-Type`. `ab87d0d80` then added
`x-fusion-client: dashboard-ui` for run-audit attribution. Both changes
are correct; neither is visible at a call site, so **114 assertions
across 7 files** kept asserting the old shape and went red together.

Fixed by naming the shape **once** in `app/test/apiRequestHeaders.ts`
rather than patching 114 literals — restating a shared fact 114 times is
what made a two-line client change look like 117 failures. Deliberately
not a loose `objectContaining`: these tests are the only thing pinning
that the attribution header is sent *at all*.

**117 → 4.** The remaining 4 are unrelated pre-existing CSS failures
(`task-detail-modal-tablet-width` ×3, `space-token-defined` ×1) —
confirmed identical on clean main with my changes stashed.

### A gap this surfaced, recorded not papered over

Three routes failed in the *opposite* direction — they send the old
shape because they call `fetch()` **directly**, bypassing `api()`, so
they never get the attribution header. `client.ts` claims the opposite:

> "Applied once here rather than per-call so no future mutation route
has to remember it."

That does not hold for a route that bypasses the helper it is applied
in. **Measured in `app/api/`: 8 files make direct `fetch()` calls and 7
include mutations (POST/DELETE)** — among them `ai-sessions.ts`'s
DELETE, which is the same class as the four-delete incident the header
was added for. So the attribution fix has a hole in exactly its
motivating case.

Not fixed here: routing those onto `api()` is a behaviour change across
the API layer and belongs to its owner, not to a test re-green. Those
assertions use a separate `API_JSON_HEADERS_NO_ATTRIBUTION` constant so
the gap stays **visible** — if a route is later moved onto `api()`, its
test fails and points at the note explaining why.

---

## This branch takes the triage bar 10 → 5, and makes `--strict` green

`node scripts/lifecycle-column-census.mjs` on this branch reports
**triage 5**, against **10** on `origin/main`. The five removed are the
ScheduleStepsEditor template/initializer/option and the RoutineEditor
default/option — the automation paths that were creating tasks into the
deleted column.

**`--strict` was also RED on clean main, twice over, and both causes
were the same mistake:** a thorough written rationale the tool cannot
read, because the marker was not where the census looks. The census
reads a comparison node's **leading comments**; a `DELIBERATE-LITERAL`
in the JSDoc above the enclosing function or declaration does not reach
the comparison inside it.

| site | why it is legitimate | why the tool could not see it |
|---|---|---|
| `columnRoles.ts:80` `isHoldColumnRole` | degrades to `columnId ===
"todo"` only when a column has **no resolved traits** — identical in
kind to `LEGACY_PRE_IMPLEMENTATION_COLUMN_IDS` directly above, which
escapes counting only because a Set is a membership form | rationale
written, **no marker token** |
| `MissionControlPanel.tsx` ×3 | the SDLC funnel **alias table** — maps
`to-do`/`ready`/`review`/`shipped` onto one display stage with an
explicit `other` bucket, and nothing branches on it | marker in the
JSDoc; the comparisons are arrow bodies **inside the array literal**,
which it does not reach |

The second only surfaced because converting the `triage` stage to a Set
removed its count and exposed the siblings — red gate, justification
sitting three lines above, unreachable.

Both are markers, no behaviour change. Neither is a conversion
candidate: resolving the funnel table to traits would **drop the
non-column aliases it exists to accept**.

**For the auto-lowering work:** the marker-placement rule is now the
recurring trap — three instances, three different authors, including me.
A marker that does not register is indistinguishable from no marker, and
the failure mode is a red gate with a written explanation nobody can act
on. If the census accepted a marker anywhere in the enclosing
declaration's comments, none of the three would have happened.

Baseline re-recorded per the tool's own instruction ("Re-record the
baseline in the SAME PR that lowered the count").


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Task “Actions” menus no longer appear on bare cards in the Planning
column.
- Routines, scheduled tasks, and create-task steps now respect each
board’s configured workflow intake column instead of using a retired
default.
- Legacy tasks saved with the retired intake column are migrated to
automatic workflow resolution.
- Target-column selection now offers only “Automatic (workflow intake)”
and “Planning,” removing the obsolete option.
- Capacity/planning messaging and related UI tooltip text were
clarified; concurrency cap updates are managed per project.

- **Tests**
- Added/updated coverage for workflow intake resolution, create-task
target column behavior (including legacy coercion), capacity safeguards,
and API request consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 03:10:35 -07:00
gsxdsm
aa02db5782 fleet: scheduler.ts 28 → 27 + make the column-role predicates reachable from the other 80% of the backlog (#2690)
## Census

**722 → 721**; `packages/engine/src/scheduler.ts` **28 → 27**. Exactly
the one site converted. Baseline re-recorded in the same commit —
`--strict` flagged the stale allowance itself, and `in-progress` went
138 → 137.

## The unblocker (commit 1)

The role helpers live in `packages/dashboard/app/utils/columnRoles.ts`,
a dashboard-**app** module. Measured against the census:

| Location | Guards | Share | Helpers importable? |
|---|---:|---:|---|
| `packages/engine/**` | 316 | 43% | no |
| `packages/dashboard/app/**` | 150 | 20% | **yes** |
| `packages/core/**` | 148 | 20% | no |
| `packages/dashboard/src/**` | 78 | 10% | no |
| `packages/cli/**` | 24 | 3% | no |

**Only 20% of the backlog can call them at all.** #2685 widens the
helper *set* correctly; that is coverage, not location.
`packages/core/src/column-roles.ts` is the same flags-first /
legacy-id-fallback predicate placed where the other 80% can reach it —
core already exports `resolveColumnFlags`, so no new resolution
machinery comes with it.

**Semantics are mirrored from #2685, not invented**, so the two sets
cannot answer the same question differently: `complete` EXCLUDES
`archived`; `wip` keys on `countsTowardWip` (the same flag capacity
arithmetic uses); `review` accepts `mergeBlocker` OR `humanReview`. One
addition — `isTerminalColumnRole` for the `!== "done" && !== "archived"`
union, the most repeated shape in the backlog.

10 tests cover both modes of all 8 predicates, including the **degraded
no-flags fallback** — the half with no coverage when these lived only in
the dashboard app — plus the two cases that prove the predicate does
something rather than nothing: a renamed column carrying the right trait
answers yes, and a legacy id carrying the WRONG trait answers no.

## A trap every fleet worker converting engine code will hit

A new core export must be added to **both** `index.ts` and
`index.gate.ts`.

The `engine-core` gate project resolves `@fusion/core` to a bundle built
from `index.gate.ts` (`scripts/build-engine-core-gate-bundle.mjs`). An
export present only in `index.ts` is `undefined` at runtime under the
gate: 13 `scheduler-workflow-cutover` tests failed with `isWipColumnRole
is not a function`, in a file that does not mock `@fusion/core` at all.
The symptom points at the consumer, the cause is the barrel.

I nearly mis-attributed this. Baseline first:
`scheduler-workflow-cutover` is **42 passed on clean main**, so the 13
were mine — not pre-existing. That measurement is the only reason I
looked at the barrel instead of "fixing" the tests.

## The conversion (commit 2)

`scheduler.ts:1690`'s `isWipColumnTask` was a hand-rolled copy of
`isWipColumnRole` — it stored only `countsTowardWip` as a boolean and
re-implemented flags-first-then-legacy-id inline. It now stores the
resolved flags object and lets the shared predicate decide.

Behaviour is identical in all four states: column present with the flag
true or false (flags win), column absent from a resolved IR, and IR
resolution failed (both defer to the legacy id).

| Check | Result |
|---|---|
| `scheduler-workflow-cutover` | **42 passed** before and after |
| 21 scheduler/capacity/hold-release files | **372 passed** |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint` | clean |

## Flagged and skipped, not guessed

**`scheduler.ts:1736` — a latent legacy-vocabulary defect, not a
conversion.** `if (task.column !== "in-progress") continue;` gates the
file-scope-lease loop on the literal, ~40 lines below capacity
arithmetic that is trait-aware. On a renamed WIP column the loop
silently does nothing while capacity counts the same cards correctly.
Converting it *changes behaviour* on renamed boards (from wrong to
right), which the fleet rules put out of scope — so it is flagged here
for whoever owns that fix. It is the same class as U10's six
legacy-vocabulary defects.

**`hold-release.ts:343`** — already marked `DELIBERATE-LITERAL`. It is
the legacy half of FN-5719's dual-accept pair; converting it would make
both halves compute the same answer, deleting the compatibility signal
*and* its divergence detector while looking like a cleanup. Untouched.

**`task-merge.ts:254`** — the documented fallback for callers that have
not proven lane identity; trait-aware callers pass
`skipColumnIdentityCheck`. Untouched.

**The other 14 `scheduler.ts` sites** have no flags in scope (e.g.
`isLegacyDependencySatisfied(dep: Task | undefined)`,
`shouldHoldActiveFileScopeLease(...)` — task-only pure functions).
Threading an IR in changes signatures and call graphs: behaviour change,
out of scope. This is why the cluster is 28 → 27 and not 28 → 0, and the
reachability measurement behind it is #2687.

No changeset: `@fusion/*` are private and no `@runfusion/fusion`
behaviour changes.
2026-07-30 03:01:36 -07:00
gsxdsm
bb3bdab999 The ratchet follows the count down — a drop tightens instead of reddening the gate (coordinator item 2) (#2679)
Taken after asking twice for reassignment with no reply, and after the
same failure bit a **third** time. No open PR touches the census CLI, so
this is unowned in practice — **U12, say so if you have started and I
will close this in favour of yours.**

## What changed

A **drop** now tightens the baseline instead of failing. Failing hard
was defensible in isolation — a stale allowance is a hole, since those
guards can return up to the old count while the check stays green. What
it missed:

**The drop is almost never the failing author's to fix.** Eleven files
dropped during one merge wave, none of those PRs re-recorded, and none
of their authors did anything wrong. Measured three times since CI began
gating this: `columnRoles.ts` 0 → 1, then `executor.ts` twice.

A permanently-red gate is a bigger hole than a stale allowance, because
it gets ignored and then nothing is guarded at all. **The rise check —
the ratchet's actual purpose — is untouched and still fails hard.**

## The residual, named rather than glossed

In CI the write is discarded with the runner, so the committed baseline
stays stale until someone commits a tightened one. The exposure is
bounded (regrowth only up to the old count), printed on every run, and
strictly smaller than the exposure from a check people route around.
`--strict --exact` restores hard failure for the pinned end state.

**One writer:** the write is now a named `writeBaseline()` shared by the
tighten path and `--update-baseline`, rather than a second
`writeFileSync`. Two writers for one artifact is how they drift — a
lesson this file already learned once.

## Exercised end to end

| scenario | result |
|---|---|
| drop, `--strict` | exit **0**, `TIGHTENED`, allowance rewritten 9 → 6
|
| drop, `--strict --exact` | exit **1**, baseline untouched |
| rise, `--strict` | exit **1** |
| clean | exit **0** |

Pinned through the real CLI with an isolated baseline. Revert proof:
restoring the hard failure fails **1 of 32**.

## Two of my own mistakes, recorded

**A vacuous assertion, in the case that guards against vacuity.** I
first wrote `expect(allowedAfter).toBeLessThan(4 + allowedAfter)` — true
for every number. Replaced with a comparison against the inflated value
the fixture started from. This file documents that trap repeatedly and I
still walked into it, which is the argument for the mechanical revert
check over careful reading.

**The env override is `FUSION_CENSUS_BASELINE_PATH`**, not the
`FUSION_CENSUS_BASELINE` I used in the first draft — so the first
version of these cases silently ran against the **real** baseline and
passed for the wrong reason. A test whose fixture never took effect is
the same failure as a test whose fixture can't fail.

## Verification

32/32 census suites, `pnpm test:gate` **71/71**, `--strict` exits 0,
`pnpm lint` clean, `docs/testing.md` updated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Update — the base-ref ratchet (review round 2, commit `4895845579`)

The first version of this PR shipped a **named residual**: the
tightening write dies with the CI runner, so the committed allowance
stays high and a later PR can regrow guards up to it while `--strict`
prints green. I called the exposure bounded and moved on. Greptile
flagged it P1 and was right — naming a hole is not closing one.

`--strict` now stops trusting the committed number for files the branch
touched. It measures each **changed** file at the base commit
(`FUSION_CENSUS_BASE_REF`, else the PR base branch, else `origin/main`)
and fails if the file carries more guards than the base ref has. **The
enforced ceiling is what main has today**, so a stale, missing, or
long-unrecorded baseline no longer opens a window.

| decision | why |
|---|---|
| changed files only, `<ref>...HEAD` | untouched files have main's
counts by construction; censusing all ~400 at the base ref is ~400 `git
show` calls to re-derive numbers that cannot have moved. Three-dot also
stops charging this branch for guards that landed on main after the
fork. |
| a new file's base allowance is **0** | "absent at the base ref" as
unbounded would make a new file the cheapest place to hide a fresh guard
|
| fails **open** on an unresolvable ref, printing `SKIPPED` | a shallow
clone cannot produce an honest comparison; a degraded run must not read
as a clean one. The baseline comparison still applies. |
| merged into the existing `regressions` list | one failure per file,
and `--update-baseline` keeps working as the deliberate escape hatch. No
new exit path. |

**Revert proof, measured both ways.** With the base-ref block removed,
the regrowth fixture — base commit 2 guards, HEAD 5, baseline allowing 9
— exits **0** with `TIGHTENED`, which is precisely the reported
scenario. With it: exit **1**, `column-guard count ROSE`, `above its
count on the base ref`, baseline left at 9. **3 of the 4** end-to-end
cases go red on revert. The fourth passes without the fix by design — it
is the genuine-conversion case the auto-tighten exists to keep green,
and a case that reddens either way proves nothing.

The end-to-end suite builds a throwaway two-commit `git init` repo under
the temp dir, because this exploit is a property of the **plumbing**,
not of the comparison: resolving a ref, working out the changed set,
reading base source through `git show`. The comparator itself is pure
with the reader injected (`findRegrowthAgainstBase`), with its own cases
in `lifecycle-column-census-ast.test.ts` — including the one that would
silently pass everything, looking up the wrong key in
`summarize().byFile`.

**Rebased onto `origin/main` @ bc782d8d92** (the branch was forked
before the recent merge wave; its baseline read 746 against a tree of
722).

Verification on the rebased branch: census **722** / `--strict` exit 0 ·
**70/70** across both census suites · `pnpm test:gate` **71/71** · `pnpm
lint` clean.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:56:08 -07:00
gsxdsm
da77e61118 A rate-limited provider kept getting hammered — the executor and merger lane checks were still literals (#2672)
`taskUsesProvider` resolves a task's **active lane** to decide which
providers it is running on. The **planner** half was converted to traits
— its own note in that function describes this exact failure and says it
was fixed — and the **executor** and **merger** halves were left as
`task.column === "in-progress"` / `=== "in-review"`.

So on a renamed board an actively-executing card resolved **no
providers**, and a provider rate limit never paused it: the engine kept
sending work to the limited provider with that card.

## Measured

Renamed board (`building` = wip, `checking` = review), limit triggered
by a peer:

| lane | before | after |
|---|---|---|
| executor | `["FN-TRIGGER"]` | `["FN-TRIGGER","FN-PEER"]` |
| merger | `["FN-TRIGGER"]` | `["FN-TRIGGER","FN-PEER"]` |

The trigger was paused either way — but only through the
*always-include-the-trigger* fallback, and **that is what made the hole
quiet**: one task always gets paused, so the behaviour looks like it
works.

Resolved from the **same per-workflow IR cache** the planner lane
already uses, so this adds no reads. Both halves fail soft to the legacy
literal, so an unresolvable workflow behaves exactly as before.

## How it was found — the part worth keeping

A scan for *"legacy literal within a few lines of a role-resolved call"*
flagged this file. That is the same heuristic that produced #2670, and
it is now 2 for 3.

The first thing I suspected here was the `done`/`archived` terminal
filter. I wrote that fix, and **its revert stayed green.** That is not a
dead end — chasing *why* it would not go red showed the lane check
already excludes finished cards, so the terminal literal there is
genuinely redundant, and the real defect was one line over in the lane
check itself. **A revert that stays green is information: either the
guard is vacuous, or you are looking at the wrong line.**

I dropped the unprovable change and kept the provable one. The
`done`/`archived` filter is deliberately unchanged, with that reason
recorded in the test.

## Revert proofs, isolated

- wip lane back to the literal → **1 of 55 fails** (the executing peer)
- review lane back to the literal → **1 of 55 fails** (the merger peer)

Three paired cases pass under both, so "pause everything with that
provider" cannot pass for "resolve the lane": a card parked in the
renamed planning lane is **not** paused on an executor limit, a wip card
is **not** paused on a merger limit, and the legacy vocabulary still
pauses peers when no workflow resolves.

## Verification

- 55/55 `usage-limit-detector`; `pnpm test:gate` **71/71**; engine
typecheck clean; `pnpm lint` clean
- census unchanged at 748 — the fail-soft literals remain by design,
which is why the count is not the measure of this fix

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:50:30 -07:00
gsxdsm
3e80dcb8ef fix(test): a Vite prefix-match alias silently unresolved a core subpath (greens full-suite shard 1) (#2686)
## What

`full-suite.yml` shard 1 on main fails with **zero test failures** — it
dies on a resolution error:

```
Failed to resolve import "@fusion/core/task-delete-attribution" from "packages/dashboard/app/api/client.ts"
```

**Root cause.** Vite string aliases match by **PREFIX**. So `find:
"@fusion/core"` → `core/src/index.ts` rewrites
`@fusion/core/task-delete-attribution` into
`core/src/index.ts/task-delete-attribution`, which cannot resolve. The
narrower subpath alias has to come *first*.

The module exists and *is* correctly declared in
`packages/core/package.json` exports — this is purely a test-config
trap, and `packages/dashboard/vitest.config.ts` already documents it in
a comment. Six configs alias `@fusion/dashboard` (whose
`app/api/client.ts` imports that browser-safe leaf) while lacking the
narrower alias, so they inherited the trap. This carries the same
one-line pattern to all six.

## Measured

`dependency-graph` — the project actually red on main:

| | Test files | Tests collected |
|---|---|---|
| before | 3 failed \| 17 passed | 147 |
| after | **20 passed** | **180** |

**33 tests were never collected** — neither passing nor reported as
failing. That is the part worth flagging: an unresolved import removes
tests from the run silently, and the shard's own summary printed no
`Tests N failed` line at all, which is why this red looked like
infrastructure noise rather than a real defect.

No regressions: `reports` 110, `cli-printing-press` 41,
`compound-engineering` 317, **gate 726** — all green. `pnpm lint` clean.

`@fusion/desktop` is `1 failed | 264 passed` **both before and after**;
verified pre-existing on clean `origin/main` by reverting just that one
config and re-running. Cause is `@fusion-plugin-examples/roadmap` entry
resolution, unrelated — **flagged, not fixed.**

## Deliberately not changed

Engine's *second* `@fusion/core` alias (the `.gate-bundle/core.mjs`
entry) is untouched: that lane bundles core on purpose, and pointing it
at source would defeat the isolation the gate bundle exists to provide.

## Full-suite triage this came out of (for whoever owns the rest)

Reading the four red shards of the last completed run on main
(`30523568756`):

| Shard | Real cause | Owner |
|---|---|---|
| 1/4 | **this PR** — resolution error, 0 test failures | — |
| 2/4 | 23 failed: `store-wedge-resolution.pg`,
`central-archive-secrets`, `task-delete-caller-attribution`,
`task-delete-nonblocking-cleanup` | #2669 / #2675 cover the first two |
| 3/4 | **watchdog SIGKILL** mid-`@fusion/engine [1/2]` — no test
failures, no summary | unowned |
| 4/4 | 17 failed, all in `@runfusion/fusion` CLI (`project.test.ts` 8,
`task.test.ts` 5, `extension.test.ts` 2, +2) | unowned |

Two of the four shard reds contain **no failing test at all**, so
"main's full-suite failure count" cannot be read off the shard
conclusions — it has to be read off `Tests N failed` summary lines, and
shards 1 and 3 emit none.
2026-07-30 02:47:14 -07:00
gsxdsm
dc50425e98 docs: correct 104 future-dated FNXC timestamps across 61 files (#2680)
## What

The FNXC convention exists so a reader can place a note against the
change that motivated it. A stamp dated *after* the edit landed defeats
exactly that.

This is program-wide drift, not one author's slip — I contributed to it
in my own commits this week, which is how I noticed it.

## Measured, on this tree

**104 stamps across 61 files** dated later than the day they were
written, from one day ahead to **2026-10-19 (81 days)**:

| count | date | count | date | count | date |
|---|---|---|---|---|---|
| 50 | 2026-07-31 | 6 | 2026-08-05 | 3 | 2026-08-13 |
| 17 | 2026-08-01 | 1 | 2026-08-07 | 1 | 2026-08-19 |
| 7 | 2026-08-02 | 1 | 2026-08-12 | 2 | 2026-08-26 |
| 11 | 2026-08-03 | | | 3 | 2026-10-19 |

An earlier number I circulated was ~70. That came from a narrower
pathspec and was wrong; **104** is the measurement.

## How

Each stamp is rewritten to the date of the commit that introduced **that
line**, via per-line `git blame` — deliberately *not* stamped uniformly
with today's date. A uniform stamp swaps a wrong date for a different
wrong date and flattens the ordering that makes these comments
navigable; blame preserves it. Times of day are untouched, and a blame
date in the future is clamped rather than trusted.

## Why the verification is listed

A docs sweep across 61 files is precisely where a stray edit hides, so
the safety claims are mechanical rather than asserted:

- every changed line begins with a comment marker — **no code touched**;
- **no test asserts an FNXC date later than today**, so no `toContain`
assertion on embedded source text can be silently invalidated (several
such assertions do exist);
- CSS files, which carry several of those assertions, are outside the
pathspec.

## Verified

lint clean · merge gate green (487 + 158 + 10 + 71) · `census --strict`
exit 0 · tsc clean for core, engine, and dashboard
(`tsconfig.app.json`).

**No behavior change.** Comment text only.

## Not done here

A guard preventing recurrence. A check that rejects an FNXC stamp dated
after the commit would stop this returning, but it needs a decision
about where it runs (lint rule vs. gate) and it is a behavior change to
CI — it does not belong riding inside the sweep it would police.
2026-07-30 02:38:58 -07:00
gsxdsm
543f4a556c Tell an already-converted fallback literal from an unconverted guard — 19 of 19 dashboard scan hits were the former (#2677)
The batch phase is about to hand per-file guard lists to cheap workers,
and the census currently cannot distinguish **"not yet converted"** from
**"converted, with a documented degradation."**

## The measurement that makes this a class, not a preference

A proximity scan for *"legacy literal near a role-resolved call"* — the
heuristic that produced #2670 and #2672 from the engine — returned **19
hits across the dashboard and zero defects.** Every one was:

```ts
if (flags) return flags.hold === true || flags.countsTowardWip === true;
return column === "todo" || column === "in-progress";      // reachable only without traits
```

That literal is **correct**: it answers for callers with no resolved
column metadata, which is the case `resolveLifecycleColumns` returns
`undefined`-for-the-whole-struct to preserve. A worker told to "convert"
it would delete the only answer available when traits are absent.

## And the difference is structural, so the parser can see it

In **both** engine defects the literal sat in a **separate statement
beside resolved data**, not in a fallback branch. Proximity cannot tell
those apart; an AST can.

`traitFallback` flags the ternary form and the **early-return** form
(which is how most are actually written), and deliberately does **not**
flag a fallback whose test is itself a column-*name* check — otherwise
any `if/else` over column names would launder itself.

## Reported beside the backlog, not subtracted from it

```
COLUMN guards (the backlog):   746
  of the column guards, 9 are trait-fallback branches (already converted)
```

A fallback literal is still a literal and should go when the trait path
becomes unconditional. This only says which **kind** of work it is.

**Advisory, and structurally so:** `traitFallback` never changes `kind`,
and the count lives *outside* `totals`. My first attempt put it in
`totals` and broke two existing suites that correctly deep-equal that
shape — an advisory number does not belong in the structure that defines
the bar.

## Revert proof

Forcing `traitFallback: false` fails **3 of 33** (both fallback forms,
plus the kind-unchanged case). The two *negative* cases pass under the
revert — which is the point: they assert what must **not** be flagged,
and a classifier that flags nothing satisfies them trivially. Worth
stating, because a revert proof that only counts failures would look
stronger than it is.

## Baseline

Re-recorded: `executor.ts` 87 → 85 was **main's own drift** from #2568
landing, so `--strict` was red on main again. #2668 made the re-record
possible; the auto-tighten (coordinator item 2) is still open, and this
is the third time in this program that a legitimate merge has left the
gate red for everyone else.

## Verification

62/62 across both census suites, `pnpm test:gate` **71/71**, `--strict`
exits 0, `pnpm lint` clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added reporting for legacy column comparisons found in trait-fallback
branches.
* Census results now include a separate count for these fallback-related
column guards.
* Human-readable reports display the new metric alongside the existing
backlog totals.

* **Tests**
* Added coverage for fallback detection across ternary, early-return,
and conditional patterns.
* Added safeguards to prevent false positives in resolved-data and
column-name checks.

* **Maintenance**
  * Updated baseline census metrics to reflect revised classifications.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:33:19 -07:00
gsxdsm
87a4dbdc60 test(u9): a release-leg E2E fixture that diagnoses itself (same defect found 3x independently) (#2678)
## What

The planned-spec release-leg fixture defect has now been diagnosed
**three times independently** — #2634 (`workflow-lifecycle`), #2643
(`workflow-merged-board`), and again in `workflow-planning-lane`. Each
time it cost real time, because it presents as a *scheduler* bug rather
than a fixture bug.

The mechanism: `createTaskWithReservedId` leaves a bootstrap seed (`#
<id>\n\n<description>`), `isUnplannedForExecution` (`hold-release.ts`)
refuses to move an unplanned card out of any intake- or hold-trait
column, and the sweep reports `held: [{ reason:
"move-rejected-or-no-slot" }]` while releasing nothing. That is **the
gate working.**

This extracts the write into `seedPlannedSpec`
(`_planned-spec-fixture.ts`) which **self-checks against the real
predicate the gate uses** (`isUnplannedSeedPrompt`, imported — not
restated) and throws naming the fixture as the cause. Three call sites
converted; their ~12-line comments collapse to a pointer, so the
diagnosis and both dead-end hypotheses live once, at the seam that
causes them.

## Measured (real PostgreSQL, not estimated)

| Check | Result |
|---|---|
| 3 converted families + new ratchet | **41 pass / 41** |
| Guard neutered (`if (false)`) | **4 of 5** ratchet tests fail |
| Original defect reproduced faithfully | **3 of 7** planning-lane tests
fail |
| `pnpm lint`, engine `tsc --noEmit` | clean |

**The ratchet fails on the original defect.** It drives the two real
seed shapes through the production builders (`buildBootstrapPrompt`,
`buildRefinementSeedPrompt`) rather than local imitations, so it cannot
keep passing if the seed shape drifts — which is exactly the drift the
fixture absorbs. The one test that survives the neutered guard is the
happy path, which should not move.

**The fixture is load-bearing, not decorative.** Proven by reproducing
the defect end-to-end: production `buildBootstrapPrompt` on the created
row with the guard bypassed → 3 planning-lane tests fail, including its
own control case.

A false mutation is worth recording, because it nearly produced a wrong
"not load-bearing" verdict: a *hand-written* seed passed 7/7.
`isUnplannedSeedPrompt` is **byte-equality** against a prompt built from
the task's own title/description, so only a byte-exact seed reproduces
it. A *missing* prompt does not either — `isUnplannedForExecution`
catches the read error and returns `false`.

## Deliberately not converted

`workflow-rebound-family`'s `PROMPT.md` write. It is a
content-preservation artifact asserted byte-identical across a re-home,
not a release-gate fixture; the helper would overwrite the very bytes
under assertion.

## Reversible decisions taken (per standing authority)

- **`opts.content` seam.** Exists so the ratchet can drive a known seed
and prove the throw fires. Without it the guard could only ever be
observed passing — the "guard that reports success without checking
anything" failure mode. Documented as test-only.
- **`title`/`description` optional.** The check is shape-based: both
recognised seed forms are `<heading>\n\n<description>` with no section
headings, so the written spec cannot match either for *any* description.
Omitting them cannot mask a positive, and a test asserts that directly.
- **`merged-board`'s spec text lengthened** to match the other two (both
were already non-seed, so behaviour is unchanged; verified by the
41-pass run).

## Method correction worth propagating

My collision scan was wrong and I nearly acted on it. `git diff
origin/main origin/<branch> -- <file>` reports a difference when a
branch is merely **stale** (the file did not exist at its base), so it
flagged dashboard and CLI PRs as touching engine E2E files. Diffing
against each branch's **merge base** is correct. Re-run under the fixed
method: all five files here are uncontested, and
`feature/code-organization-wave17` genuinely does touch
`packages/engine/src/triage.ts` (so that one stays hands-off).

## Lane

`.pg.test.ts` under engine-default, `pgDescribe`-skipped without
PostgreSQL — the merge gate is unaffected. Throwaway per-file database,
never port 4040, no temp-root walk.
2026-07-30 02:19:17 -07:00
gsxdsm
07c29757a3 The archived half of the terminal pair, end to end (#2568's fix landed first and is better) (#2670)
**Live on `main`.** `parkCompletedBlockedTask` opens with *"is this card
already finished?"* and answered it with:

```ts
if (task.column === "done" || task.column === "archived") return false;
```

On a renamed board neither matches, so the guard was **inert** — and
inert here is not a missed rescue, it is active damage: the very next
block rebounds the card to its planning lane. **A completed card sitting
in a renamed complete or archived column was moved backwards out of
it.**

## Why this survived two reviews

#2644 (mine) converted the **rebound** half to resolve its target by
role. This **terminal** half stayed a literal. A role-resolved rebound
behind a name-matched guard means the renamed board takes the rebound
and never the guard — the same half-conversion shape as the evacuation
branch, except the two halves were owned by different changes, so
neither review saw both.

Worth keeping as a review heuristic for the remaining conversions:
**when a converted site sits next to an unconverted one, the conversion
can make the neighbour worse — and a diff showing only one of them looks
complete.**

## Relationship to #2568

The fix exists there, stranded four deep in a stack whose bottom (#2544)
has not merged, so nothing in that chain has reached `main`. This
re-lands **only** the guard, directly against `main`. I have noted it on
#2568 so its author can drop that hunk rather than resolve it twice; I
took no other part of that PR (its extraction and the terminal-pair
ownership change are still theirs).

## Deliberate choices

- **Fail-soft to `["done", "archived"]`** — an unresolvable workflow
behaves exactly as the literal pair did.
- **An unclassifiable column is NOT terminal.** Being unable to prove a
card is finished must not be the same as proving it is; that is what
keeps a stranded card moving.

## Revert proof

Restoring the literal pair fails **2 of 5** — the renamed complete and
archived cases. Three paired cases pass under the revert, so neither
"never park" nor "always park" can pass for resolving the lanes: a
genuinely mid-pipeline card is still parked, the legacy pair still
answers when no workflow resolves, and an unclassified column still gets
the park.

## Verification

- 5/5 new; 22/22 with `executor-rebound-already-there` and
`executor-task-done-blocked`
- `pnpm test:gate` **71/71**; engine typecheck clean; `pnpm lint` clean
- census: `executor.ts` **87 → 85** column guards; baseline re-recorded
in the same commit

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Reliability**
* Improved handling of completed/blocked tasks when workflow column
naming changes, ensuring tasks aren’t parked or moved incorrectly in
archived or non-terminal lanes.
* **Tests**
* Added Vitest coverage to verify terminal-lane decisions, including
cases where workflow resolution occurs mid-operation and lane placement
changes during the wait.
* **Maintenance**
* Updated lifecycle column census baseline numbers to match current
workflow usage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:16:30 -07:00