Commit Graph

12610 Commits

Author SHA1 Message Date
gsxdsm
5795d70b27 fix(engine): assignment load must be resolved per task — #2787 P1 follow-up (#2796)
Fix-forward for the P1 that arrived on **#2787 after it merged** — so it
lands as its own PR rather than a thread reply on merged code.

## The finding

`selectPermanentAgentForTask`'s `activeColumns` was resolved from the
**candidate** task's workflow and then applied to every row `listTasks`
returned. On a project running several workflows — the normal case —
assignments living in another workflow's load-bearing lanes vanished
from the tally, and the already-loaded-agent-wins bug returned through a
different door.

**A column id means something only relative to its OWN workflow.**
`blocker-fanout.ts` documents exactly this and offers a per-task
`classify`; the option is now that same shape rather than a third
invention:

```ts
countsAsAssignmentLoad?: (task: Task) => boolean
```

The scheduler resolves each assigned row against its own IR, sharing one
cache for the selection, so a board spanning three workflows reads three
IRs — not one per assigned card.

## Why this is the third round on the same parameter, stated plainly

1. I added the parameter and **never wired the caller** — inert in
production.
2. I wired it as a **union of wip+review**, which dropped hold/intake
and made it a *regression* for backlog work.
3. I resolved it from **one workflow** and applied it to all — this fix.

Each round was a smaller version of the same error: treating a lane
answer as global when it is per-task, and per-role when it is
per-membership. Worth recording because the first two rounds both looked
correct and both passed their tests — the tests asserted the renamed
case I was thinking about, not the shape of the data.

## Verification

- new cross-workflow case; reverting the predicate to a single
workflow's lanes **fails it**
- `agent-assignment` suite **14 passed**
- `pnpm test:gate` — **161 / 13 / 487 / 71** · lint clean · census
`--strict` exits 0

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 11:06:20 -07:00
gsxdsm
6bb5e4f787 test(engine): measure the optional-role-parameter conversion class (#2798)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Follows #2795, which
found the first instance of this pattern.


`packages/engine/src/__tests__/workflow-optional-role-param-caller-audit-live-e2e.pg.test.ts`

## The finding

#2795 showed a conversion pattern the lifecycle-column census cannot
see: a role question migrated into an **optional parameter whose default
is the legacy literal**, converted at some call sites and not others.
This shows it is not a one-off, and measures it.

| seam | call sites passing the resolved answer |
|---|---|
| `shouldHoldActiveFileScopeLease` | **2 of 4** (both `scheduler.ts`;
neither `self-healing.ts`) — #2795 |
| `evaluateParkedAgentTaskLink` | **2 of 6** (`scheduler.ts`,
`task-agent-sync.ts`; neither `agent-heartbeat.ts` ×2 nor
`self-healing.ts` ×2) — this PR |

The second is the more damaging, and the callee's own FNXC note already
names the outcome: without the resolved columns "the card would be
treated as unparked and its live agent link cleared" — **a stale-link
bug turned into a dropped-link bug**. Driven here: a card parked in a
renamed board's hold column, with live execution proof, has its agent
link dropped.

### Why the census is blind to it

The callee is converted and its default is correctly marked
`DELIBERATE-LITERAL` — for an unconverted caller that default genuinely
*is* the intended behaviour. **The unconverted call sites contain no
column literal at all**; it lives one function away. So the census
counts the callee's annotated literals and sees nothing at the call
sites, and the conversion reads as complete from every angle except
running it.

This is a *class*, not two bugs. The same shape exists at roughly twenty
seams (`revertableColumns`, `plannerColumns`, `roleColumn`,
`terminalColumns`, `activeColumns`, …). Two are now measured. I checked
two others I flagged as unknown in #2795 —
`restart-recovery-coordinator.ts`'s `isReviewColumn?` and the
`isRecoverableMissingWorktreeReviewFailure` family — and **their callers
are fully converted** (`extension.ts:1924`, `task.ts:1390`,
`self-healing.ts:12087`), though the doc comment claiming `extension.ts`
"still asks with the literal" is now stale. The rest are unaudited; the
audit case is written so adding a seam is a small edit.

## Scope, stated honestly

Three cases are driven end to end: real persisted rows from a live
store, the real exported predicate, both call shapes. The **call-site
split is asserted against source text** — reaching all six sites needs
the heartbeat and self-healing harnesses, which I did not build, and the
audit case says so in its own comment rather than dressing it up.

It is an alarm in **both** directions: a new unconverted caller pushes
the count up and fails; converting an existing one pushes it down and
also fails. The second is deliberate — that is the moment someone should
read the three behavioural cases and update the number on purpose.

## Mutation-verified

Flipping the callee's default from the legacy parked pair to
`["backlog"]`:

| case | result |
|---|---|
| CONTROL (default board, no options) | **fails** |
| CHARACTERIZATION (renamed board, no options) | **fails** |
| BOUND (renamed board, options passed) | passes — correct, the argument
overrides the default |
| AUDIT | passes — correct, it is a source assertion |

## Not done, and why

**No fix.** Passing the resolved columns at the four unconverted sites
means resolving each linked task's traits inside the heartbeat and
self-healing paths — async work in loops that already hold locks — and
both files belong to other workers. The differential says exactly what
the fix should make true.

## Verification

- new suite — **4/4 passed**, mutation matrix above
- full live-PG E2E surface — **137/137 passed** (133 on main + 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 11:03:09 -07:00
gsxdsm
60054aab0a test(engine): live-PG evidence of an inert conversion at the CALL SITE (#2795)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit.


`packages/engine/src/__tests__/workflow-file-scope-lease-caller-gap-live-e2e.pg.test.ts`

## Why this is a different finding, not a sixth of the same one

#2789/#2791/#2792/#2793/#2794 all concern **one** mechanism: a site
resolves the workflow synchronously and silently gets the default board.
This is a **second** mechanism, and neither the lifecycle-column census
nor the sync-resolver allow-list can see it.

`shouldHoldActiveFileScopeLease` was converted by turning its two role
questions into optional parameters with literal defaults:

```ts
const isWipColumn    = options?.isWipColumn    ?? task.column === "in-progress";
const isReviewColumn = options?.isReviewColumn ?? task.column === "in-review";
```

A caller that resolved the traits passes the answer; a caller that has
not gets exactly the pre-conversion behaviour. That is a deliberate
migration device and the source says so — correctly marked
`DELIBERATE-LITERAL`.

**But the migration was only half made:**

| call site | passes the resolved answer? |
|---|---|
| `scheduler.ts:1986` | ✅ `{ isWipColumn: true }` |
| `scheduler.ts:2006` | ✅ `{ isReviewColumn: true }` |
| `self-healing.ts:4525` | ❌ neither |
| `self-healing.ts:5443` | ❌ neither |

So the same predicate is right on the scheduler's path and wrong on
self-healing's. The harm is the one the function's own FNXC note
describes: on a renamed board both branches fall through, the predicate
returns false for every card, `activeScopes` stays empty, and the
dispatch path sees no overlap — *two agents editing the same files*,
which is what the overlap machinery exists to prevent. At the
self-healing sites the consequence is narrower but identical in shape: a
stale-lease reconciler concludes a live blocker holds no lease and
proceeds to clear state the scheduler would have honoured.

### Why the existing instruments are blind to it

**There is no column literal at the self-healing call sites.** The
literal lives inside the callee's default, one function away — and there
it is correct, because for an unconverted caller it *is* the intended
behaviour. A census counting `=== "in-progress"` occurrences sees the
callee's two (properly marked) and nothing at all at the call sites. The
conversion reads as complete from every angle except running it.

This generalizes: **any conversion that migrates behaviour behind an
optional parameter leaves a residue the census scores as done.** Worth a
sweep for the same shape elsewhere — `agent-assignment.ts`'s
`activeColumns?` and `restart-recovery-coordinator.ts`'s
`isReviewColumn?` are the same pattern; I have not checked whether their
callers supply them.

## Scope, stated honestly

Three cases are driven end to end: real persisted rows from a live
store, the real exported predicate, both call shapes. The **call-site
fact is asserted against source text, not driven** — reaching those
sites needs the full dependency-lease reconcile harness, which I did not
build. The last case reads the file and says so in its own comment
rather than dressing it up as an end-to-end result. It doubles as an
alarm: when those call sites are converted it fails and points at the
three cases above, which describe exactly what changes.

## Mutation-verified

Flipping the callee's default from `"in-progress"` to `"building"`:

| case | result |
|---|---|
| CONTROL (default board, no options) | **fails** |
| CHARACTERIZATION (renamed board, no options) | **fails** |
| BOUND (renamed board, option passed) | passes — correct, the option
overrides the default |
| SOURCE-LEVEL | passes — correct, it is a source assertion |

The two default-dependent cases bind to the default; the bound case
proves the override; nothing passes for the wrong reason.

## Not done, and why

**No fix.** Passing the resolved answers at the two self-healing sites
requires resolving each blocker's column traits there — an async
resolution inside a reconcile path that already holds locks, and
`self-healing.ts` is another worker's file. Flagging with a differential
that says exactly what the fix should make true.

## Verification

- new suite — **4/4 passed**, mutation matrix above
- full live-PG E2E surface — **137/137 passed** (133 on main + 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:59:57 -07:00
gsxdsm
1496ba9658 test(engine): bound the inert-sync-resolution class on a live store (#2794)
## What

One new live-PostgreSQL E2E suite, 3 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Closes the series:
#2789 (scheduler), #2791 (planner lanes), #2792 (custom fields), #2793
(terminal node).


`packages/engine/src/__tests__/workflow-sync-selection-blast-radius-live-e2e.pg.test.ts`

## Why this one is different

The four PRs above each proved a site broken because it resolved a
task's workflow synchronously. Read together they invite a conclusion
that is **false and would be expensive**: that every synchronous
consumer of the workflow selection is inert.

Most are not. The difference is one line of shape:

```ts
// GUARDED (correct)
store.getTaskWorkflowSelectionAsync
  ? await store.getTaskWorkflowSelectionAsync(id)
  : store.getTaskWorkflowSelection(id)

// UNGUARDED (inert)
store.resolveTaskWorkflowIrSync(id)
```

The real PostgreSQL store **does** implement the async reader, so every
guarded site takes the async arm and resolves the card's own workflow.
Only the sync IR helper — which has no async arm to fall to — is stuck
with the default.

Observed on one live store, one persisted workflow, one task:

```
hasAsyncReader   = function
SYNC  selection  = undefined
ASYNC selection  = { workflowId: "WF-001", stepIds: [] }
EFFECTIVE planReviewMaxRevisions = 9   <- the custom workflow's declared default
```

## The point

"The ternary saves them" is an inference from reading, and the whole
premise of this program is that reading is what let the class survive in
the first place. The guarded sites are exactly the ones a fleet worker
would otherwise "fix": converting a correct site costs review time,
risks behaviour, and produces a diff that looks like progress. This
makes the bound checkable in the same lane as the defects.

Guarded call sites (correct today): `workflow-settings-resolver.ts`,
`workflow-ir-resolver.ts`, `executor.ts`,
`workflow-graph-task-runner.ts`, `workflow-task-runtime.ts`, and
`board-workflows.ts` in the dashboard.

## The allow-listed family is now closed

| site | status |
|---|---|
| `scheduler.ts` | proven broken — #2789 |
| `replan-target.ts` | proven broken — #2791 |
| `task-store-helpers.ts` | proven broken — #2792 |
| `branch-and-pr-entities.ts` | proven broken — #2793 |
| `workflow-task-create-ops.ts` | **legitimately correct** — creation
runs before any selection exists, so the default IR is the right answer
|
| `lifecycle-ops.ts` | **NOT proven, stated as such** |

`lifecycle-ops.ts`'s stale-transition-pending recovery re-runs plugin
column-transition hooks against the sync IR. Driving it needs a
registered plugin hook plus a crash-simulated marker; I did not build
that harness and I am not substituting a unit test for it. Named in the
file so it is not mistaken for covered.

## Evidence discipline

- **Observed state.** Both readers called on one live store against one
persisted workflow, plus a real resolved settings value — not a spy on
which arm ran.
- **The settings default is `9`**, deliberately not the builtin's, so
the value can only have come from this workflow.
- **Mutation-verified.** Rewriting the guarded consumer to call the sync
reader directly fails **exactly** the bound arm; the two structural arms
are correctly unaffected, which is what a bound should do.

## Verification

- new suite — **3/3 passed**, mutation-verified
- full live-PG E2E surface — **136/136 passed** (133 on main + 3; #2791
landed while this branch was in flight)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:56:48 -07:00
gsxdsm
90f6319b79 batch-engine tail: re-land the ASYNC half; the sync-resolved half was inert (engine −15) (#2785)
Tail of `batch-engine` (#2773). That PR merged as a squash while later
engine work was still in flight, so `self-healing.ts`, `executor.ts` and
`worktree-pool.ts` landed at their pre-conversion counts. This re-lands
**only the half that is real**, and the reason the other half is not
here is the substance of this PR.

## Census, per file (measured, `--strict` verified)

| file | main | here |
| --- | ---: | ---: |
| `engine/src/self-healing.ts` | 107 | 97 |
| `engine/src/executor.ts` | 15 | 12 |
| `engine/src/worktree-pool.ts` | 3 | 2 |
| `engine/src/ephemeral-worker-manager.ts` | 1 | 0 |
| `engine/src/agent-tools.ts` | 5 | **0** |
| `engine/src/gridlock-detector.ts` | 3 | **0** |
| `engine/src/triage.ts` | 4 | 1 |
| `engine/src/mission-execution-loop.ts` | 2 | **0** |
| **net** | | **−28** |

Baseline re-recorded; `--strict` tightened exactly these 4 entries and
no others.

## Finding: a whole class of conversions in this program is INERT, and
the census scores it as progress

`resolveTaskWorkflowIrSync` returns the **default** workflow IR for
every task in production. The sync selection reader behind it is a
PostgreSQL-cutover stub:

```ts
// packages/core/src/task-store/workflow-definitions.ts:505
export function getTaskWorkflowSelectionImpl(_store, _taskId) {
  return undefined;   // "Backend mode cannot synchronously read PostgreSQL"
}
```

So a guard written as
`resolveLifecycleColumns(store.resolveTaskWorkflowIrSync(id))?.hold`
resolves an IR, asks for a trait, and answers **from the default
workflow for every custom board** — silently. It reads as converted and
the census counts it as converted. `main` gained
`sync-workflow-ir-callsite-allowlist.test.ts` for exactly this after my
branch point; it is what caught me.

I had built three sync resolvers on that reader — `resolveMoveLanesSync`
(self-healing, executor) and a widened `resolveTaskParkedColumnsSync`
(scheduler) — reasoning that a *synchronous* `task:moved` listener needs
a *synchronous* reader. That reasoning was sound about the shape and
never checked whether the reader reads anything.

**Dropped from this PR, deliberately, and NOT re-landed anywhere:**

- `scheduler.ts` 12 → 1 (the widening; the pre-existing narrow helper on
main is untouched)
- the executor `task:moved` handler, incl. the Move-Task hard-cancel
lane comparison
- self-healing's `task:moved` fan-out,
`classifyPausedAbortWorkflowRecovery`, `reconcileInReviewBranchRebind`,
`recoverWedgedActiveMerge`, `recoverPausedAbortFailures`, and 12
single-row lane conversions

Those sites are back to their literals. The allow-list's own guidance is
the standard I applied:

> An unconverted `=== "todo"` is strictly better, because it is at least
honest about being a literal.

I did not add my call sites to the allow-list. Six entries would have
turned the gate green in two minutes and buried the defect; the list's
contract requires proving the async resolver is genuinely unreachable,
and for a fire-and-forget listener it is not — the listener can `void`
an async lane resolution the same way `NotificationService` already
does. That is the correct fix and it is a behaviour-shaped change, so it
is out of scope here.

**Fleet-wide consequence:** any conversion routed through
`resolveTaskWorkflowIrSync` is fake progress, and the census cannot see
the difference. `pnpm test:gate` can: the allow-list test is the
detector. Its passing here (161/161) is this PR's evidence that nothing
inert survived the split.

## What IS in this PR — all async-resolved

1. **`self-healing.clearStaleBlockedBy`** — lanes resolved per
**REFERENCED** task, not per iterated task. A blocker's own workflow
decides whether it is still blocking.
2. **`executor` dependency satisfaction** — resolved per **DEPENDENCY**
via `columnsWithFlag`. Preserves the load-bearing asymmetry that a
dependency in *review* already satisfies a dependent; a bulk sweep
flattens that to complete-only and deadlocks the board.
3. **`agent-tools` — the agent task tools listed FINISHED cards as
active.** `fn_task_list` says it lists "tasks that aren't done or
archived"; `fn_task_search` offers `includeDone: false`. Both filtered
on `task.column !== "done"`, so a renamed complete lane returned
finished cards as outstanding work **to an agent**, which then reasons
and acts on them. `includeArchived` was always enforced by the QUERY and
survived a rename; `"done"` was only ever a TS predicate, which is why
exactly that half broke.

Plus the two **dedup** guards in the same file. The cross-parent
diagnostic filter kept a *shipped* card as a candidate on a renamed
board, so the guard adopted it as canonical and returned `wasDuplicate:
true` — absorbing new diagnostic work into a task nobody is working on
(the eval-followup defect shape again). The defined-feature bootstrap
preflight is **not** the query-filter class: its query passes
`includeArchived: true`, so the TS predicate is the *only* archived
guard there; on a renamed archive lane the archived sibling became the
bootstrap canonical and `claimDefinedFeatureTask` then rejects the
non-live row, so a valid first task fails to be created at all.

Both dedup invariants **already had tests** — asserted against the
legacy ids only, so both passed for the very comparison being replaced.
Extended in place into vocabulary differentials rather than added as
parallel files. Two helpers rather than one parameterised one: "is this
finished?" and "is this archived?" are different questions, and merging
them would make the archived-only guard also reject completed rows.

The list/search half re-landed **with the test it originally shipped
without.** No suite exercised either tool, so the original commit's
"304/304 green" said nothing about the change — the optional-flags
failure mode exactly. Both call sites are covered; converting two copies
and testing one is the Surface Enumeration failure this program has
already hit twice.

4. **`gridlock-detector` — FALSE dependency alarms.** The gate compared
each blocker against `done`/`in-review`/`archived`; on a renamed board
all three are true for a *finished* blocker, so no dependency ever
counted as met and the detector reported dependency gridlock for tasks
that are not blocked — `notifyGridlock` then pages the operator.
Resolved per dependency using the **same five flags** as the executor's
gate (`complete`, `archived`, `mergeOrchestration`, `mergeBlocker`,
`humanReview`) — `review` is not a trait, and two gates answering "is
this dependency satisfied?" differently is a split brain. Every
pre-existing case in that file omits a workflow, so none could detect
the change; added the renamed case plus a non-vacuous companion.

5. **`triage` — its OWN copies of the same two tools.**
`createTriageTools` carries a `fn_task_list` and `fn_task_search`
byte-identical in intent to the agent-tools pair, plus a third site
filtering duplicate candidates. Same defect on all three. Reused the
(now exported) agent-tools helper rather than adding a third copy —
deliberately stronger than the two-parallel-tests reading of Surface
Enumeration, since the copies now share one implementation and cannot
drift. **Not claiming call-site coverage:** `createTriageTools` is
private and not drivable without standing up a TriageAgent; the helper
is revert-proofed, those two call sites are covered only through it.

6. **`mission-execution-loop` — a finished fix task read as LIVE,
stalling remediation.** The comment above that line states the rule it
implements: *only an open task makes duplicate triage safe to suppress.*
On a renamed board the rule inverts — a finished fix task is not
`done`/`archived`, so it reads as live, remediation for a fresh
validation failure is suppressed indefinitely, and the mission stalls
with no error surfaced.

**Not revert-proven, and I am not claiming it is.** No test reaches the
`hasLiveFixTask` branch, and the only case that mints a fix feature is
git-gated and heavyweight; building that fixture is larger than the
conversion. The change strictly *widens* the finished set (resolved
roles ∪ the two legacy ids), so default boards are byte-identical — that
is the argument for shipping it unproven, not a substitute for coverage.

7. **Four census-invisible membership guards**, each inverted on a
renamed board — `worktree-pool` (merger-managed branch reclaim could
delete a branch out from under an in-flight merge), `agent-assignment`
(assignment load counted nothing), `ephemeral-worker-manager`
(`isAgentIdle` inverted on both sides), and the dead constants their
conversion orphaned. These are `SET.has(task.column)` shapes the census
does not count, so the −15 understates them.

## Revert results (measured, each run)

| conversion | reverted → |
| --- | --- |
| `clearStaleBlockedBy` per-referenced lanes | renamed-vocabulary case
fails; stale `blockedBy` never clears |
| executor dependency satisfaction | dependent never unblocks on a
renamed review lane |
| `worktree-pool` merger-managed set | reclaim proceeds against an
in-flight merge |
| `ephemeral-worker-manager.isAgentIdle` | idle agent reads busy on a
renamed board |
| `fn_task_list` terminal filter | RENAMED case fails — shipped card
listed as active |
| `fn_task_search` terminal filter | RENAMED case fails — same,
independently |
| cross-parent diagnostic dedup | RENAMED case fails — `wasDuplicate:
true`, new work absorbed |
| bootstrap preflight archived guard | RENAMED case fails — `validate`
called with the archived sibling |
| gridlock dependency gate | RENAMED case fails — false gridlock raised
for an unblocked task |

`agent-assignment`'s widened `taskStore` type is compile-time; its
revert is a tsc failure, not a test failure — stated rather than claimed
as coverage.

## Verification

- `pnpm test:gate` — 161 + 487 + 13 + 71, all green (161 includes
`sync-workflow-ir-callsite-allowlist`)
- `npx tsc -p packages/engine/tsconfig.json --noEmit` — clean
- `pnpm lint` — clean

One commit is a pure import restore: `columnsWithFlag` arrived in a
sibling commit that built on the inert resolver and was left behind. The
engine tsconfig excludes `src/__tests__/**`, so the gate was green while
tsc was not — worth knowing that on this package a green gate is not a
green build.


## Verified NOT a gap — measured, so the next worker does not re-open
them

- **`restart-recovery-coordinator` (5 counted).** Four already take an
optional `reviewColumns` set and the counted literals are the documented
**fallback** arm, which must stay for the same reason `columnRoles.ts`
keeps its id fallback. The sole production caller
(`self-healing.ts:12151-12154`) already passes the resolved set. The
fifth is documented at the site as a re-assertion behind a `listTasks({
column: "in-progress" })` query filter. Nothing to convert.
- **`notification/notification-service` (5 counted).** Already
documented in-file as deliberately counted with no exemption marker: the
wedge-episode site needs per-task serialisation of wedge handling (a
delivery-semantics change to operator notifications), and
`isManualMergeHold` needs a pre-resolved `LifecycleColumns` threaded
through `handleTaskUpdated`, which would pay resolution on every task
update. Both are behaviour/placement judgements, not conversions.
- **`planner-overseer` (3 counted).** `resolveWatchedStage`'s two
literals are fed by `pollPlannerOverseer`, which calls `listTasks({
column: "in-progress" })` and `{ column: "in-review" }` — hardcoded
**query** filters. On a renamed board those queries return no rows, so
the predicate never sees a renamed column. Converting it alone would
drop 3 from the census and change nothing an operator can observe. The
real fix is at the query layer; that is the tracked query-filter-bounded
class, not this PR.
- **`triage:695`** reads `resolvePlannerLanes` → the allow-listed sync
IR reader. Left as an honest literal per the rule above.

**Still open in `packages/engine`, deliberately not in this PR:**
`self-healing.ts` (97, of which ~31 are the query-filter-bounded class
and the rest need per-site classification in a 13k-line file),
`scheduler.ts` (12, blocked on the sync reader above), `executor.ts`
(12), and a tail of ~13 more copies of the "is this task finished?"
question across eight small files (`agent-reflection`,
`auto-merge-finalization`, `merger-scope-auto-widen`,
`backlog-pressure-reporter`, `merger-orphan-rehome`,
`merger-integration-worktree`, `plugin-runner`, `cli-agent/*`). That
tail is a clean follow-up: one question, eight call sites, and the
exported `resolveTerminalColumnsForTasks` helper already exists for it.

That is the same discipline as the sync-resolver finding: a census
number that drops without a behaviour change is not progress, and four
of these files would have handed over exactly that.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:47:44 -07:00
gsxdsm
4184fde08d batch-cli-plugins: 7 guards — 3 were a foreign enum, and fn pr create refused every card on a renamed board (#2775)
`batch-cli-plugins` — the u7 worker's mega-batch: `packages/cli` +
`plugins` + anything left.

## The batch is 7 guards, and 3 of them are not guards at all

The census's per-file list gives this batch seven sites. Reading them,
**three are a foreign vocabulary the census matches on the string
alone**:

| file | site | verdict |
|---|---|---|
| `plugins/fusion-plugin-reports/store/report-store.ts` | `next ===
"archived"` ×2 | **not a column** — `next` is a `ReportStatus` |
| `plugins/fusion-plugin-reports/store/report-types.ts` | `to ===
"failed" \|\| to === "archived"` | **not a column** — same enum, its own
terminal states |

The reports plugin has its own status lineage (`draft → generating →
review_* → approved → published`, plus `failed`/`archived`) that shares
two spellings with the lifecycle vocabulary. A report is not on a board
and has no workflow, so resolving an IR there would answer a question
nobody asked. All three are marked `DELIBERATE-LITERAL` with the reason
at the site.

**This cuts the other way from #2763.** That PR establishes the census
total as a *floor* (25 membership predicates it structurally cannot
see). This is the opposite error in the same number: a foreign enum
inflating it. The total is neither a ceiling nor a floor — it is an
estimate with error in both directions, and the per-file list is worth
reading before trusting a file's count.

## Converted (census before → after, per file)

| file | before | after |
|---|---|---|
| `packages/cli/src/commands/pr.ts` | 1 | **0** |
| `plugins/…/even-realities-glasses/notifications/diff.ts` | 1 | **0** |
| `plugins/…/reports/store/report-store.ts` | 2 | **0** (deliberate) |
| `plugins/…/reports/store/report-types.ts` | 1 | **0** (deliberate) |

### `fn pr create` refused every card on a renamed board

The live defect in this batch. The gate was `task.column !==
"in-review"`, and its error told the operator to move the task to a
column their board does not have:

```
Error: Task must be in 'in-review' column to create a PR (current: signoff)
```

There is no way to satisfy that short of renaming the workflow back. Now
resolved through core's `resolveReviewColumns`, and the message names
the lanes that actually exist.

**The SET, not `lifecycle.review`.** A board may declare more than one
review lane, and a card parked in a `humanReview`-only lane is still a
card you can open a PR from. A single-id answer keeps refusing those —
the same narrowing #2728's review caught in the CLI retry gate, which is
why the test pins both lanes.

## Skipped, with the reason

**`plugins/fusion-plugin-even-cards` (2 guards) — blocked on packaging,
not on analysis.** The defect is real: `boardToDeck` filters with
`column !== "archived" && column !== "done"`, so on a renamed board
every finished card stays in the deck, fills `maxCards`, and pushes the
active cards off the display. The wearer sees a board that never
finishes anything.

I implemented the fix and **reverted it**: this plugin is not in
`pnpm-workspace.yaml` and depends only on `@fusion/plugin-sdk` — it has
no `@fusion/core` dependency, so the route cannot reach
`resolveTaskLifecycleColumns`. Adding one is a packaging change, which
this program's rules put out of scope. Shipping only the injected
parameter without a caller was the alternative, and that is precisely
the decorative conversion #2759 documents: the census would drop by 2
and the deck would keep the bug.

Flagged for whoever owns the plugin's dependency surface. The glasses
plugin next door *does* depend on `@fusion/core`, so this is a
one-plugin problem, not a plugin-wide one.

## Honest note on the glasses conversion

`diff.ts`'s completion branch is **currently unreachable** — the only
production caller (`notifier.ts`) passes `alsoNotifyOnDone: false`. So
that conversion changes nothing at runtime today. It is converted rather
than marked deliberate because the literal is not deliberate: it is
wrong, and would ship the bug the day someone turns the flag on. Stated
here rather than left for a reviewer to discover.

## Verification

- new CLI suite **4 passed**; `pr-command` + `pr-automerge-cleanup` +
`bin-pr-router` **35 passed**
- glasses plugin **181 passed (19 files)** · reports plugin **110 passed
(23 files)**
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
`--strict` exits 0

**Revert proof, measured.** Restoring `if (task.column !== "in-review")`
fails 3 of the 4 new cases (`process.exit:1` on both renamed lanes, and
the refusal message reverts to naming `in-review`). The
unresolvable-workflow case keeps passing — it is the legacy path — so
the negative cases alone do not pin the fix and all four are required.

## Handoff to `batch-engine`

`packages/engine/src/project-engine.ts` **5 → 0** is finished, green,
and pushed as `handoff/project-engine-lanes-for-batch-engine`
(`34dbb35209`) for the capacity worker to cherry-pick — it is
engine-owned, not mine to land.

It fixes two live defects: a card that **had merged** reported as a
failed merge to `fn task merge` and the dashboard button (`merged:
finalTask?.column === "done"`), and the three post-finalize `column ===
"done" && mergeConfirmed` fast-path checks, which on a renamed board
sent an already-landed card down the bounce path — re-queued,
retry-counted, and in the capped branch parked `failed` with its merge
sitting on main. Plus `hasAutoHealableVerificationBufferFailure`, which
returned false for every card on a renamed board, so a buffer-overflow
verification failure was never auto-healed.

8 new tests, revert-proven (restoring the literal fails 4 of 8), gate
green.

---

## Completion pass (u7) — the batch is now closed

Two workers converged on this branch. I rebased onto the first-landed
commit rather than force-pushing over it, took its wording wherever the
conclusion was identical, and added what was missing.

### What this pass added

1. **`even-cards` (2 sites)** — the only in-scope file the first pass
left open. Marked DELIBERATE-LITERAL: the package depends on
`@fusion/plugin-sdk` only, and the SDK does not re-export the lifecycle
role helpers, so there is no IR, no store, and no trait flags to resolve
*from*. Fixing it properly means the SDK exposing role flags on the task
shape it hands plugins — a structural change, out of scope, and recorded
at the site as the correct home. Live consequence is cosmetic: a
finished card on a renamed board shows as active in the glasses deck.

2. **A red test in the `fn pr create` conversion.** The incoming version
rendered `Task must be in 'in-review' to create a PR`, dropping the word
`column`. `task.test.ts:3422` pins `must be in 'in-review' column`, so
that hunk failed `runTaskPrCreate > exits with error when task not in
in-review column`. Restoring the word makes the single-lane message
**byte-identical** to the pre-conversion one, which is what a vocabulary
conversion should be — the guard's own test now passes unmodified.
Marked at the site so it is not "simplified" back.

3. **Duplicate imports** — the two independent conversions each added
`resolveWorkflowIrForTask`/`resolveReviewColumns`, which does not
compile. Deduped in its own commit.

### Census

Measured with `--json` on `origin/main` and on this branch.

| file | before | after | action |
|---|---|---|---|
| `packages/cli/src/commands/pr.ts` | 1 | 0 | converted |
| `plugins/fusion-plugin-reports/src/store/report-types.ts` | 1 | 0 |
marked |
| `plugins/fusion-plugin-reports/src/store/report-store.ts` | 2 | 0 |
marked |
| `plugins/fusion-plugin-even-cards/src/cards/board-cards.ts` | 2 | 0 |
marked |
| `plugins/fusion-plugin-even-realities-glasses/.../diff.ts` | 1 | 0 |
marked |

Backlog **415 → 408** (−7, exactly the in-scope count). Deliberate **40
→ 46** (+6 marked); 6 + 1 converted = 7. `--strict` exits 0. **Nothing
remains in `cli` + `plugins` + everything-else — there is no follow-up
batch behind this one.**

### One note on the `even-realities-glasses` site

Worth recording beyond "cannot resolve": its only production caller
(`notifier.ts:80`) passes `alsoNotifyOnDone: false`, so that arm is
**unreachable today**. Converting it could not have changed observed
behaviour either way.

### Verification (measured, on the merged branch)

- `pnpm --filter @runfusion/fusion exec tsc --noEmit` → exit 0
- `pnpm lint` → 0 errors
- CLI `task.test.ts` → 144 passed, including the `runTaskPrCreate` guard
test
- `@fusion-plugin-examples/reports` → 110 passed;
`even-realities-glasses` → 181 passed

**Pre-existing failures, not from this change:** the 5
`runTaskImportFromGitHub` / `runTaskImportGitHubInteractive` tests fail
identically on `origin/main` — verified by stashing this diff and
re-running (5 failed / 144 passed both ways).

---

## Census audit (unowned follow-on)

After closing the batch scope I audited whether the **392**
column-backlog number is inflated by foreign vocabularies — the class
this batch found in the reports plugin, where `"archived"` is a
`ReportStatus` rather than a board lane. If that class were widespread,
every remaining batch would be chasing sites that must not be converted.

**It is not. The number is real.** A receiver-level pass over all 392
column-category sites found exactly **3** false positives, all in
`plugins/fusion-plugin-reports` (`next`, a `ReportStatus`), all now
marked in this PR.

What was checked and cleared:

- **Property-reached foreign enums** (`step.status`, `feature.status`,
`mission.status`) — already correctly bucketed into the separate
`status` category (185), not the column backlog. Verified against
`merge-queue-ops.ts`: 11 lifecycle-spelled literals in the file, census
counts **1**, and that 1 is the genuine `.column` guard.
- **Bare step-status variables** (`status`, `currentStatus`,
`liveStatus` compared to `"done"`/`"skipped"`) — likewise excluded.
- **Every other receiver in the backlog** — `to`, `from`, `column`,
`fromColumn`, `toColumn`, `latestColumn`, `state`, `preArchiveColumn`.
All resolve to genuine task columns. `executor.ts`'s 15 sites were
spot-checked line by line: all 15 are real.

The gap the classifier genuinely cannot close is a foreign enum held in
a **bare variable** — the receiver name carries no type information, so
`next === "archived"` is indistinguishable from a lifecycle guard by AST
alone. That is why the reports sites need a marker rather than a
classifier fix, and it is now documented in
`lifecycle-column-census-ast.mjs`'s header alongside the measured scope,
so the remaining batches do not re-run this hunt.

Census tests: **43 passed**. The change is comment-only.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:38:39 -07:00
gsxdsm
6bdde6f246 fix: five lifecycle gates the census cannot see — incl. live ephemeral workers reaped and duplicate follow-up cards (#2787)
Five lifecycle-column fixes the census **structurally cannot see**. Each
gate is a `Set` or array literal — a *definition*, not a comparison — so
no backlog entry ever pointed at any of these files. Found by grepping
for lane-shaped list literals after the same shape surfaced in
`duplicate-intake` and `blocker-fanout` (both merged via #2780), then
confirmed by reading each USE site.

**On opening this:** I offered twice to fold these into a PR and kept
them on handoff refs to respect one-open-PR-per-worker. They have now
sat unadopted across several cycles while `main` moved, and two of them
destroy or duplicate work. Opening is the reversible call — **close it
if it breaks queue policy** and I will keep them on the branch.

## What is in it

| commit | defect on a renamed board | severity |
|---|---|---|
| `beb107a7bc` | assignment load-balancing **defeated** —
`assignmentLoad` stays empty, every candidate reads as load 0, the sort
falls through to its stable `createdAt` tiebreak, so **one agent wins
every assignment** while the rest idle | distribution |
| `cf4b59e1cb` | the zombie sweep **deletes LIVE ephemeral workers** |
**destroys work** |
| `5fe004ae64` | eval follow-up dedup sees **zero open tasks**, so every
run re-files follow-ups it already filed | **duplicate cards** |
| `a1021de8b2` | agents keep a **"working on" indicator for finished
cards** | stale UI |
| `86680d1220` | the **Files tab never loads** — the fetch never fires |
silent empty |

### The one that destroys work

`shouldDeleteOnSweep` tested a hard-coded terminal `Set`, then fell
through to `return task.column !== "in-progress"`. On a renamed board
**both halves miss, and they compound in the worst order**: the terminal
test fails, control reaches the fallthrough, and `"building" !==
"in-progress"` is `true`. An ephemeral worker **actively executing a
task** is classified as a zombie and deleted. Nothing logs.

Its fallback is **deliberately asymmetric**, and the comment says why:
an unresolvable workflow keeps the legacy literals rather than guessing.
Failing to reap a dead worker costs a slot; reaping a live one destroys
work in flight. Those are not symmetric, so uncertainty fails toward
keeping the worker.

## Verification

Verified **as a set**, not only per-branch:

- `pnpm test:gate` — **161 / 13 / 487 / 71**
- engine suites (assignment, ephemeral, eval-followups) — **44 passed**
- dashboard suites (agent-task-link, useSessionFiles) — **16 passed**
- `tsc` engine + dashboard server + dashboard app — clean
- `pnpm lint` clean · census `--strict` exits 0

**Revert-proven individually.** Restoring each literal fails its own
case: the renamed-wip zombie case, the renamed-wip assignment case, the
renamed-lane dedup case, the sanitizer ratchet, and both
`useSessionFiles` role cases.

## Two honesty notes, flagged rather than buried

**`a1021de8b2`'s guard is STRUCTURAL, not behavioural.**
`sanitizeAgentTaskLinks` is a closure inside `createApiRoutes`,
reachable only by standing up the full express app. The ratchet asserts
the source — resolver threaded per task, bare literal call gone, cache
shared, fallback retained — and **fails on revert**, verified. It is not
a substitute for a behavioural test; whoever owns the dashboard server
should add one if that seam grows.

**`useSessionFiles`'s negative case passed in isolation and failed in
the suite.** Hooks are not unmounted between cases there, so a prior
case's in-flight fetch landed inside it. That is the classic shape of a
test that gets "fixed" by reordering; it now asserts a **delta** against
the pre-render call count, which is independent of what leaks in.

## Deliberately NOT included

`worktree-pool.ts:1205` — the sixth site from the same sweep. It **fails
safe**: a missed match means the skip does not fire, so the branch is
added to `activeBranches` and *protected* from cleanup. The cost is
stale branches accumulating, not deletion. It also sits in the merger's
branch-reaping path, where the opposite error destroys work, so it
deserves its owner's judgement rather than a drive-by conversion.
Flagged, not guessed.

Also still open and unclaimed: roughly 69 untriaged literal-list sites
across engine/dashboard/cli. The grep is one line and the file list is
on #2775 — with the measured caveat that about half are false positives
on shape alone (`LEGACY_*` names, seeds unioned with resolved values,
and `roles: ["triage"]`, which is an `AgentCapability`, not the deleted
column). Only the use site settles it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:35:31 -07:00
gsxdsm
dcf9900d61 test(engine): live-PG differential between resolvePlannerLanes and its async twin (#2791)
## What

One new live-PostgreSQL E2E suite, 5 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Follows #2789 (same
defect class, different site).


`packages/engine/src/__tests__/workflow-planner-lanes-sync-vs-async-live-e2e.pg.test.ts`

## The finding

`replan-target.ts` exports two functions with identical logic and
identical fallbacks, differing only in how they obtain the task's IR:

```
resolvePlannerLanes(store, taskId)              -> store.resolveTaskWorkflowIrSync(taskId)
resolvePlannerLanesForTaskAsync(store, taskId)  -> await resolveWorkflowIrForTask(store, taskId)
```

Under PostgreSQL the sync selection reader answers `undefined` for every
task, so the sync twin resolves the **default** workflow for every card
regardless of the board it is on. **Nine production call sites use it**
(2 × `executor.ts`, 7 × `triage.ts`); one uses the async twin.

The module's own doc comment argues this, and the store-level fact is
proven in `sync-workflow-ir-is-always-default.pg.test.ts`. What had no
executable evidence is the consequence **at this seam** against a real
store with a real persisted workflow. That is this file — a pure
differential: both twins, same store, same task, same call.

### Two harms, different severity

1. **Wrong lanes, labelled authoritative.** `resolvedFromWorkflow`
exists to tell a caller "these came from the workflow, not the
fallback". The sync twin sets it `true` — an IR did come back — while
handing over the default board's ids. A caller that correctly checks the
flag before trusting the lanes is misled *precisely by checking it*,
which is strictly worse than the honest `false` an unresolvable store
would give.
2. **Invented forward lanes.** `wip`/`review`/`complete` are optional so
a caller *refuses* rather than moving a card into a column the board
does not declare (PR #2628's review). The sync twin defeats that
contract without touching it: never having seen the real board, it
reports the default board's forward lanes as present. The optionality is
intact in the type and unreachable in practice. The sharpest arm: a
board declaring **no** review lane gets `undefined` from the async twin
and `"in-review"` from the sync twin.

### A correction worth carrying forward

"It falls back to the legacy lanes" is the wrong mental model **twice
over**. `LEGACY_PLANNER_LANES` (`hold: "todo", intake: "triage"`) is
reached only when no IR resolves at all — under PostgreSQL, never. What
a caller actually receives is the **post-U11 merged default**, whose
intake and hold are one `todo` lane. So the sync twin does not return
`triage` for intake; it returns `todo`, and a caller reading `intake`
gets not merely a wrong id but a lane that is not a dedicated intake at
all.

This is also why the control arm uses `MERGED_VOCAB`: the shape the
twins agree on is the merged one. `DEFAULT_VOCAB`, which splits intake
out as `triage`, already separates them.

## Evidence discipline

- **Observed state.** These are exported pure functions over a live
store; the observation is their return value against a persisted
workflow definition. No spies, no mock IR anywhere in the file. Contrast
the unit coverage in `planner-lanes-async-resolution.test.ts`, which
must supply a mock `resolveTaskWorkflowIrSync` and therefore cannot see
this divergence at all.
- **Control arm.** On the post-U11 default shape the twins agree exactly
— which is why this survived: every default-board test passes and only a
renamed board separates them.
- **Characterization, not endorsement.** The four renamed arms assert
the wrong-but-current values deliberately; they flip when the call sites
move to the async twin, and that flip is the point.

### Mutation-verified, including a round that found weak arms

Replacing the sync twin's whole body with `return LEGACY_PLANNER_LANES`:

| | arms failing |
|---|---|
| first draft | **3 of 5** |
| after strengthening | **5 of 5** |

Two arms originally asserted only `wip`/`review`/`complete`, which are
identical in the merged default IR and in `LEGACY_PLANNER_LANES` — so
they proved the lanes were wrong without proving *why*, and survived the
mutation. Each now also pins `intake` (`todo` merged vs `triage`
legacy), the single field that separates "resolved the wrong board" from
"took the fallback". Recorded in the file next to the assertions.

## Not done, and why

**No fix.** The async twin already exists and is documented as a drop-in
("identical logic and identical fallbacks — the ONLY difference is
awaiting the authoritative resolver"), so the migration is mechanical
*where the caller is already async*. It is not universally so: several
`triage.ts` sites are inside synchronous paths, and `triage.ts:831`
calls `resolvePlannerLanes(this.store, "")` with an empty task id — a
sweep-wide lane read that has no single task to resolve against and
needs a decision, not a mechanical swap. Both are behaviour calls in
files another worker owns; flagging, not smuggling.

## Verification

- new suite — **5/5 passed**, mutation-verified 5/5
- full live-PG E2E surface, 20 suites — **126/126 passed** (was 121/121)
- `pnpm lint` — clean
- `pnpm check:lifecycle-columns` — exit 0

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added end-to-end coverage comparing synchronous and asynchronous
workflow lane resolution.
* Validated lane consistency for renamed, non-default, and custom boards
using persisted workflow data.
* Added checks for incorrect fallback lanes, workflow resolution
indicators, and absent review lanes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:32:24 -07:00
gsxdsm
87442b9664 test(engine): live-PG evidence that declared custom fields cannot be written (#2792)
## What

One new live-PostgreSQL E2E suite, 4 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Third in the series
after #2789 and #2791; same root cause, materially worse consequence.


`packages/engine/src/__tests__/workflow-custom-fields-sync-resolution-live-e2e.pg.test.ts`

## The finding

**A workflow that declares custom fields cannot have any of them
written.**

`TaskStore.resolveTaskCustomFieldDefsSync` reads a task's field
definitions through `store.resolveTaskWorkflowIrSync`, which under
PostgreSQL answers `undefined` for every task and therefore resolves the
**default** workflow IR. The default declares no `fields`, so the
function returns `[]` for every task on every board. `task-update.ts`
validates every write against that empty list:

```ts
const defs = store.resolveTaskCustomFieldDefsSync(id);
const result = validateCustomFieldPatch(defs, updates.customFields);
if (!result.ok) throw new CustomFieldRejectionError(result.rejection);
```

Observed against a real store with a real persisted workflow declaring
one `text` field:

```
STORED fields = [{"id":"risk","name":"Risk","type":"text"}]
SYNC   defs   = []
WRITE  threw  = CustomFieldRejectionError
                custom field 'risk' rejected (no-fields-defined):
                the resolved workflow declares no custom fields; no values may be written
```

The rejection message is a true statement about the workflow that got
resolved and a false one about the workflow the card is on.

### The two halves of the feature disagree in production

The executor resolves the same definitions through the **async**
resolver (`executor.ts` → `resolveTaskCustomFieldDefs` →
`resolveWorkflowIrForTask`) and sees the real field. So an agent can be
prompted to supply a value that the store will then refuse to store. The
last case asserts both answers against **one store, one task, one
workflow** — which is why this cannot be dismissed as a fixture
artefact.

This is a different severity from the previous two PRs in the series.
#2789 and #2791 are wrong-lane defects, mostly latency, one of them
unbounded. This one is a declared feature that does not function off the
default board.

## Scope on record

Three write paths share the sync resolver: `task-update.ts` (driven
here), `workflow-task-create-ops.ts:394`, and `workflow-ops.ts:488`.
Only the first is exercised; the other two are named in the file so the
surface is recorded rather than implied.

Also worth stating plainly: because the empty list *is* the default IR's
`fields`, the same rejection is what a default-board card gets too. The
feature is not merely renamed-board-broken.

## Evidence discipline

- **Fixture integrity first.** The opening case asserts the stored
workflow really does declare the field via the async resolver. Every
other assertion is about a *missing* definition and would pass just as
well against a workflow that never declared one — that case is what
makes the rest mean something.
- **Observed state.** The thrown typed rejection plus the **absence** of
a persisted value on a re-read row. No spy on the validator.
- **Mutation-verified.** Replacing the sync resolver's body with a
hardcoded `[{id:"risk",…}]` fails **3 of 4** arms. The fourth is the
fixture-integrity case, which exercises the async path by design and
correctly survives.

## Not done, and why

**No fix.** The async resolver already exists and is already used by the
executor for the same data, so the shape of the fix is clear — but
`task-update.ts`'s validation runs inside a synchronous update path, and
making it async is a behaviour decision in `@fusion/core` that belongs
to that file's owner, not to a smuggled edit in an evidence PR. The
call-site allow-list entry for `task-store-helpers.ts` ("Synchronous
helper shared by txn-hot paths") should cite this suite either way: the
entry is accurate about the constraint and silent about the cost.

## Verification

- new suite — **4/4 passed**, mutation-verified 3/4 (fourth by design)
- full live-PG E2E surface, 20 suites — **125/125 passed** (121 on main
+ 4)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:32:12 -07:00
gsxdsm
755ada91ac test(engine): live-PG evidence that the terminal-node guard fires on the wrong node (#2793)
## What

One new live-PostgreSQL E2E suite, 3 tests. **No production file is
touched** — evidence, per the E2E worker's remit. Fourth in the series
after #2789 (scheduler), #2791 (planner lanes), #2792 (custom fields).


`packages/engine/src/__tests__/workflow-terminal-node-sync-resolution-live-e2e.pg.test.ts`

## The finding

FN-7641 Signature 2 exists because setting `nodeId` to the terminal node
used to be written verbatim and silently do nothing — the card sat in
review with every step done, unadvanced and unexplained. The contract: a
terminal override **with** durable merge proof finalizes the card;
**without** proof it is rejected with an actionable error; non-terminal
overrides are untouched.

On a board whose terminal node is not called `end`, **both halves
invert**:

| write | contract says | actually observed |
|---|---|---|
| `nodeId: "end"` — an ordinary planning node here | written, untouched
| **rejected** with a merge-proof error about finalizing a card the
operator was not finalizing |
| `nodeId: "finish"` — this board's real `end`-kind node | finalize, or
reject | **written verbatim**, no error, card left in review |

The second row is the original FN-7641 bug, restored on every custom
board.

## The correction the mutation runs forced

My first draft blamed `isTaskTerminalNodeIdImpl`'s sync IR resolution
alone. Mutating it changed only one of the two cases, which is how I
found there are **two** guards:

```
branch-and-pr-entities.ts:568   validateNodeOverrideChange(task, nodeId, { isTerminalNodeId })
                                -> sync IR resolution (the default board, under PostgreSQL)
task-update.ts:53               validateNodeOverrideChange(task, nodeId)
                                -> NO options, so `defaultIsTerminalNodeId` — the bare literal
                                   `nodeId === "end"`
```

The inner one is an unconverted literal sitting behind a converted call
site, and it silently overrides it. **Converting the outer guard alone
changes nothing an operator can see.** A column census cannot find the
inner one either — `end` is a node id, not a column. This is the "a
guard survives in a branch of the same function" shape, one function
apart.

### Mutation matrix

| corrected | `end` rejected | `finish` silent |
|---|---|---|
| *(nothing — main)* | pass | pass |
| outer sync-IR guard only | pass | **fail** |
| inner `defaultIsTerminalNodeId` only | pass | **fail** |
| **both** | **fail** | **fail** |

Two different failure structures, which is why the cases are kept apart:

- **`end` rejected is over-determined** — both guards independently call
it terminal, so it survives a mutation of either one. Not a weak
assertion: a faithful record of a defect with two independent causes,
and the reason a partial fix here is invisible.
- **`finish` silent is under-determined** — both guards must miss the
id, so correcting either flips it. This is the arm that notices a
partial fix.

The fixture-integrity case exercises the async resolver by design and
correctly survives every mutation.

## Fixture

The shared builder's terminal node is `end`, so it cannot express this
shape. This file derives from it: one `lifecycleIr`, node ids shifted so
the `end`-kind node is `finish` and the non-terminal planning node takes
the name `end`. Columns, traits, edges and structure are otherwise the
builder's, so the only variable is which node ids carry which kind. The
first case asserts that shift really happened — both characterizations
are claims about which node is terminal and would read as defects if the
fixture had quietly kept the builder's ids.

## Evidence discipline

- **Observed state.** Whether `updateTask` throws, and what the re-read
row's `nodeId` and `column` actually are. No spies.
- **Characterization, not endorsement.** Both cases assert the
wrong-but-current behaviour deliberately, and the matrix above says
exactly which fix flips which.

## Not done, and why

**No fix.** It needs two coordinated edits in `@fusion/core` — threading
the resolved terminal check into `task-update.ts:53`, and making the
outer resolution async — and the second is the same synchronous-path
constraint as #2792. Both are behaviour decisions in another worker's
files. Worth flagging that fixing only the allow-listed sync site would
look like progress and deliver none, which the matrix above makes
checkable.

## Verification

- new suite — **3/3 passed**, mutation matrix above
- full live-PG E2E surface — **124/124 passed** (121 on main + 3)
- `pnpm lint` — clean

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:29:04 -07:00
gsxdsm
59b5e61fa2 fix(tests): the last CLI reds — import assertions still required the column U11 removed (#2788)
## What was red

All 5 failures in a full `@runfusion/fusion` run on `origin/main` (`5
failed / 1673 passed`), in `src/commands/__tests__/task.test.ts`:

```
- "column": "triage",
```

## The product change is intentional and documented

**#2603 (U11)** removed the hardcoded `column: "triage"` from the
GitHub/GitLab import writes so `createTaskImpl` resolves the
**workflow's** intake column instead. Passing `column` would override
that resolution and, post-U11, name a lane the default workflow no
longer declares.

`task.ts` still carries the note at three sites:

> `createTaskImpl` resolves the WORKFLOW'S intake column, and
`input.column` would override it. Hard-coding `"triage"` created the
card in a column the default [workflow does not declare].

Six `toHaveBeenCalledWith` assertions still required the removed
literal, so a correct product change surfaced as five CLI failures.

## Scoped deliberately

Only the **six assertion-side** occurrences are removed. The other
**16** `column: "triage"` literals in this file are mock *return* values
and `makeTask` fixtures, and they stay — what a created task comes
*back* as is a different question from what the import *asks for*, and
blanking them would weaken unrelated cases.

## Evidence

- Full CLI package: **1678 passed / 106 skipped, 125 files green** (was
5 failed).
- **Mutation:** reintroduce `column: "triage"` into the import write →
**2 failed**. The assertions still pin the invariant rather than having
been loosened into always-true — the thing worth checking when a fix is
"delete an expectation".
- Gate **732 green** · `pnpm lint` clean. Test-only (mutation reverted;
`git diff` clean).

## Ownership

`packages/cli` belongs to the **batch-cli-plugins** owner (u7) under the
mega-batch split. This is fix-forward on a red rather than a conversion,
confined to one test file, and touches no production code.

With #2779 and #2786 this leaves engine, core and CLI at **0 failures**
on main. The remaining known reds are the 123 dashboard failures
documented in #2784.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:20:01 -07:00
gsxdsm
02d0f80068 fix(tests): update the archived-gate inventory after #2780's conversions (#2786)
## New red on main

`archived-column-gate-parity.test.ts` fails with **"TypeScript encoding
changed"** after batch-core (#2780). It is the only failure in a full
`@fusion/core` run (`1 failed / 4748 passed`).

## The conversions are right — this is their missing half

#2780 moved four files off raw `column === "archived"` comparisons onto
`isTerminalColumnRole`, exactly the intended role-based pattern:

| file | before → after |
|---|---|
| `assigned-task-ranking.ts` | 1 → 0 |
| `duplicate-intake.ts` | 1 → 0 |
| `near-duplicate-canonical.ts` | 1 → 0 |
| `store.ts` | 2 → 1 |

(The literal still in `duplicate-intake.ts` is a `moveTask`
**destination**, not a gate comparison — a different question, like the
planner-lane move targets.)

The guard's own failure text asks for the inventory to be updated **in
the same commit** as a conversion. That did not happen, so the ratchet
went red on main. This PR is only that bookkeeping.

## Verified NOT a split-brain

That is the thing this file exists to catch — TypeScript moving to the
resolved role while the SQL sides keep comparing the raw string, so on a
renamed board one says a task is archived and the others return it as
live.

The Drizzle and raw-sql inventories are **unchanged and both pass**.
Worth stating explicitly because those two assertions run *after* the
TypeScript one: a plain red tells you nothing about them, they had to be
re-run green to know.

## The guard still bites

Appending a real `task.column === "archived"` to an audited file fails
it immediately.

**Recorded because it nearly fooled me:** my first two mutation attempts
*passed*, which looked like a guard blind to new comparisons — a much
worse finding than a stale inventory. Both had been inserted at line 2
of a file whose line 1 opens a JSDoc block, so my "code" was comment
text and was never compiled. **A mutation that does not compile is not
evidence of anything.** Appended at end of file instead, the guard fails
on the first run.

Gate **732 green** · lint clean. Test-only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:13:55 -07:00
gsxdsm
e467d939a5 fix(tests): the last 3 engine reds — a pause guard asserted at the wrong layer (#2779)
## What was red

The final 3 failures in `executor-prompt.test.ts` ("global pause
behavior"), all the same assertion:

```ts
expect(mockedCreateFnAgent).not.toHaveBeenCalled();
```

made after calling `executor.execute(task)` **directly** on a paused
todo row.

## It is not a regression, and not a live safety hole

I initially flagged this in #2778 as a possible live hole — "a
user-paused todo task **now** reaches `createFnAgent`". **That framing
was wrong**, and the bisect is what corrected it:

| commit | result |
|---|---|
| `origin/main` (HEAD) | fail |
| `main~40` | fail |
| `main~80` | fail |
| `main~150` | fail |
| `main~250` | fail |

Red 250 commits back. It never described shipped behaviour, so nothing
regressed.

**`execute()` holds no pause gate.** Neither `executeCore` nor the
workflow-graph executor consults `paused`/`userPaused` before starting a
session — I checked both. Refusing to dispatch a parked row is the
**scheduler's** invariant, enforced twice:

1. Candidacy is keyed on both flags (`scheduler.ts:138`) — `userPaused`
is a durable operator stop even when legacy `paused` is false.
2. The row is **re-read immediately before dispatch** and refused if it
comes back parked (`scheduler.ts:2086`) — this closes the race the first
check cannot.

The test called `execute()` directly, stepping around the component that
owns the guarantee, then asserted the bypassed layer enforced it. A true
statement about the system was being made to look false.

Every protective outcome #2371 documented **does** hold and stays
asserted: `fn_task_done` never completes the card, it is never handed to
`in-review`, no completion watchdog is armed, the pause is never
cleared, and the run narrates the benign paused park. Only *"no session
was created"* was false. The 3 sibling assertions in `resumeOrphaned`
are untouched — that path genuinely does refuse.

## The invariant moves to the layer that owns it

Rather than delete an assertion and lose the coverage,
`scheduler-paused-dispatch-refusal.test.ts` pins it through real
`schedule()` passes:

- **control** — an unparked ready card IS dispatched
- refuses a row parked with legacy `paused`
- refuses a row parked with `userPaused` alone
- refuses when the operator pauses **after selection, before dispatch**

Driven through `schedule()` rather than by calling the predicate
directly: a test that calls the guard cannot tell whether the dispatch
path still *consults* it — which is precisely how the executor-prompt
version came to assert a layer that had stopped being asked.

## The control earned its place on the first run

It failed immediately, and twice over: the hold-release gate refuses a
card still carrying a bootstrap seed (fixed with the shared
`seedPlannedSpec`), and a `moveTaskIf` stub returning `moved: false`
makes the release unobservable. Without the control, all three refusals
would have passed **vacuously** — a scheduler that dispatches nothing
refuses everything.

## Evidence

- **106/106** across both files.
- **Mutation:** removing the pre-dispatch pause re-read → the race case
fails. The other two are caught earlier by candidacy (defence in depth);
the passing control makes their refusal attributable to the flag alone,
since the same store dispatches without it.
- Gate **732 green** · lint clean · engine `tsc --noEmit` **0 errors**.
Test-only (the scheduler mutation was reverted; `git diff` clean).

## Engine suite status

Measured baseline on `origin/main`: **39 failures / 10835 passed**. With
#2776 (32, notifier harness) and #2778 (4), this last set takes the
engine suite to **0 failures**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:04:52 -07:00
gsxdsm
145c022af4 test(engine): live-PG evidence that the scheduler's sync parked-column read is inert (#2789)
## What

One new live-PostgreSQL E2E suite. **No production file is touched** —
this is evidence, per the E2E worker's remit.


`packages/engine/src/__tests__/workflow-scheduler-parked-columns-live-e2e.pg.test.ts`
(2 tests)

## The finding

`scheduler.ts`'s `resolveTaskParkedColumnsSync` resolves a task's
hold/intake columns through `store.resolveTaskWorkflowIrSync`. Under
PostgreSQL that reader answers `undefined` for **every** task, so the
resolver returns the **default** IR and the function yields `{ hold:
"todo", intake: "triage" }` on every board — byte-identical to the
literals it was converted away from. It is an **inert conversion**, and
it is currently allow-listed
(`sync-workflow-ir-callsite-allowlist.test.ts`) on the grounds that
these handlers are synchronous.

Five call sites read it. Four groups of handler fail as **latency** — a
wake that does not fire costs up to one poll interval, which is why the
class hid. The `task:deleted` dependency reconciliation is different: it
queries `listTasks({ column: hold })` **and** re-checks
`dependent.column === hold` before clearing `blockedBy`. On a renamed
board both tests are against `"todo"`, a column that board does not
contain, so **a dependent parked in the renamed hold column is never
unblocked and waits forever on a blocker that is already gone.**
Persisted, operator-visible, unbounded.

### It contradicts a passing unit test

`scheduler-renamed-hold-events.test.ts` asserts the opposite and passes,
because its mock supplies `resolveTaskWorkflowIrSync: vi.fn(() =>
renamedIr())` — an answer the real store provably never gives. That test
is not wrong about the *scheduler* (given a working resolver the
handlers do resolve the renamed lane); it is wrong about the *resolver*.
Flagging rather than editing it: it is still the right unit test for its
own subject, and it is not my file.

### The mechanism is not the one the code reads like

The obvious reading blames the fail-soft `?? "todo"`. It is **not** that
— `lifecycle` is never nullish, a real default IR comes back and real
traits resolve off it, so both `??` arms are dead in production.
Established by mutation, not by reading:

| mutation to `resolveTaskParkedColumnsSync` | control arm | renamed arm
|
|---|---|---|
| *(none — main)* | unblocks ✅ | never unblocks ✅ |
| both `??` fallbacks → renamed vocabulary | unblocks (unchanged) |
never unblocks (unchanged) → **dead branch** |
| returned object → renamed pair | fails | fails → **both arms decided
here** |

This is the sharpest form of the defect class: the site resolves an IR
and reads a trait off it, so it looks converted at every level except
the one that decides the answer.

## Evidence discipline

- **Observed state, not spies.** Each arm asserts the dependent's
persisted `blockedBy` after a real soft-delete on a real store with real
stored workflow definitions.
- **The negative is self-validating.** The handler's work is
fire-and-forget, so observing it needs a bounded wait — and a bounded
wait proving a negative is normally worthless. The default-vocabulary
arm is the control: same store, same window, and it *does* unblock (297
ms against a 2 000 ms window). If this ever flakes the control fails
first; the fix is the quarantine ledger, never a larger number.
- **Differential.** Both boards come from the one shared vocabulary
builder and differ only in their column ids.
- The renamed arm is a **characterization** test — it asserts the
wrong-but-current behaviour deliberately, and is expected to flip when
the read is fixed.

## Two fixture traps found on the way (both would have made this
vacuous)

1. `updateTask({ column })` is not a column move — `column` is not in
the update payload, so it typechecks as an unknown key and leaves the
card where it was. Use `moveTask`.
2. **Order matters.** Writing the blocked state *after* placing the card
re-homes it to the intake column (observed `todo -> triage`), and the
reconciliation re-checks `dependent.column === hold` — so the control
fails for a fixture reason that looks exactly like the defect. Block
first, then place. Both are written down in the file.

## Not done, and why

**No fix.** The honest fix is to make the read async, and that is not
free: these run inside synchronous `task:moved` / `task:updated`
listeners, where an added `await` defers the rest of the handler to a
microtask and reorders handlers against a synchronous emitter. That is a
behaviour decision in a file another worker owns, so it belongs to
whoever owns `scheduler.ts` — not to a smuggled edit in an evidence PR.
The allow-list entry should cite this suite either way.

## Verification

- `workflow-scheduler-parked-columns-live-e2e.pg.test.ts` — **2/2
passed**, mutation-verified in both directions (table above)
- full live-PG E2E surface, 19 suites — **121/121 passed** (was 119/119)
- `pnpm lint` — clean
- `pnpm check:lifecycle-columns` — exit 0

Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is
reachable, so the merge gate is unaffected. Throwaway per-file database;
never port 4040.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 10:04:34 -07:00
gsxdsm
31b9fafe11 fix(tests): the last core red — assert the funnel COUNTS the move, not which stage owns it (#2781)
## What was red

A full `@fusion/core` run on `origin/main` reports **1 failed / 4700
passed**. This is that one: the SDLC funnel case in
`agent-logs-and-monitor.pg.test.ts`.

```ts
expect(result.funnel.stages.find(({ stage }) => stage === "todo")?.entered).toBe(2);
// expected 2, received 0
```

## The move is not lost

Post-U11 the default Planning column is **one** column carrying
`["intake","hold","reset-on-entry"]`. `stageForTraits` prefers the
earliest stage in flow order, so `intake` wins and a move to `todo` is
attributed to the **`triage`** stage. Analytics is working correctly;
the column vocabulary underneath it merged.

## Why not just re-point the assertion

Which stage the merged Planning column *should* report is an open
product question — I flagged it on #2669 while adding the `hold`
mapping, and it is visible to users: **the funnel shows a phantom 100%
drop between Triage and Todo on every default board since U11.**

- Re-pointing the **test** at `"triage"` quietly blesses the phantom
drop.
- Re-pointing the **mapping** retroactively changes how historical
analytics read — not a reversible call.

Neither belongs in a change whose job is clearing a red.

So the assertion now pins what is true under **either** resolution: both
moves are counted exactly once, in the single pre-implementation stage
the Planning column resolves to. When #2669 is decided, this test does
not need rewriting.

## Evidence

**6/6 passed.** Mutation-proved for the failure that actually matters:

| mutation | result |
|---|---|
| unmap every pre-implementation trait (move falls to `OTHER`) |
**fails** — `expected +0 to be 2` |
| unmap `intake` alone (attribution shifts triage → todo) | **passes, by
design** — that is the open question, not a defect |

The second row is the point of the rewrite: the test is indifferent to
the unsettled question and strict about the invariant. It still catches
a dropped, double-counted, or split move.

Gate **732 green** · lint clean. Test-only — no production file touched
(the mutations above were reverted; `git diff` clean).

## Ownership

`packages/core` belongs to the **batch-core** owner under the mega-batch
split. This is fix-forward on a red rather than a conversion, so it is
deliberately confined to one assertion in one test file and touches no
production code — it should not conflict with the batch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:46:55 -07:00
gsxdsm
698bded476 fix(tests): 4 engine reds on main — each was green for a reason the fleet removed (#2778)
## Context

A full `@fusion/engine` run on `origin/main` (`9b61d795c9`) reports **39
failures / 10835 passed**. 32 are the notifier harness, fixed in #2776.
This PR takes 4 of the remaining 7.

All three files share one shape: **each case was passing off something
the lifecycle conversions have since correctly taken away.** In every
one, the product is fine and a good change landed as a red test.

---

### 1. `executor-graph-failure-lanes-resolved.ts` — an equality that
fails on its own fix

The guard forbids resolving a lifecycle *guard* through the synchronous
`resolvePlannerLanes` (a no-op under the shipped PostgreSQL backend, so
the census counts the site as converted while it behaves like the
literal). It asserted `expect(callSites).toBe(3)`.

#2764 converted the promotion-path site to
`resolvePlannerLanesForTaskAsync` — exactly the direction this guard
wants. Count went **3 → 2** and the assertion failed.

The guard's own comment states the invariant as *"Any FOURTH is a new
sync resolution"* — one-directional. Coded as equality, it fails on
removal, which is the change it exists to encourage. Now
`toBeLessThanOrEqual(2)`.

**Mutation:** adding a third sync call site → `expected 3 to be less
than or equal to 2`. Still load-bearing.

### 2. `restart.integration.test.ts` — a fixture matching a fallback
constant

`recoverCompletedTask` re-homes intake → hold → wip only when the origin
is the board's **intake** lane; otherwise it hands straight to review.
The failure showed the 1st move as `in-review` with no re-home.

Nothing regressed. The fixture put the card in `triage` and resolved
lanes through the sync resolver, so it fell through to
`LEGACY_PLANNER_LANES` — where `intake` is literally `"triage"`. **It
was matching a hardcoded fallback, not a declared lane.** #2764 made the
site await the real resolver; the mock selects `builtin:coding`, and
**U11 merged intake and hold onto one Planning column (`todo`)**, so
`triage` is not a lane on that board and the two-hop correctly
collapses.

The invariant the test is named for — completed work in a distinct
intake lane is re-homed along a legal path, not moved intake → review,
which role adjacency rejects — is still real. So the fixture now
**declares** a board with intake separate from hold, the only shape
where the two-hop is reachable.

**Mutation:** removing the re-home hop from the product → fails with the
expected `todo` first-move. Load-bearing.

### 3. `executor-abort-provenance.test.ts` — a call one argument short

Both provenance cases returned `false` for a clean completed in-review
row. This reads as an FN-6796 regression stranding rows that are already
handed off for review.

It is not. #2703 added a 7th `reviewLane` parameter so the lane is
resolved by the caller. **The call goes through `as any`, so the missing
argument was not a type error** — it arrived `undefined`, `live.column
!== reviewLane` held for every row, and the classifier answered false
for everything.

Passed explicitly rather than defaulted inside the classifier: a default
would restore the literal the parameter exists to remove. Added a
**differential** — a card resting in a *renamed* review lane classifies
the same, a mismatched one does not — so the parameter cannot be
re-literalized while still looking converted.

**Mutation:** `live.column !== "in-review"` → the differential fails.
The other cases pass, which is precisely why it was worth adding.

---

## Evidence

| file | result |
|---|---|
| `executor-graph-failure-lanes-resolved` | **24 passed** |
| `restart.integration` | **48 passed** |
| `executor-abort-provenance` | **16 passed** |

Gate **732 green** · `pnpm lint` clean · engine `tsc --noEmit` **0
errors**. Test-only — no product file is touched by this PR (the
mutations above were run and reverted; `git diff` confirms clean).

## Deliberately NOT fixed here

3 cases in `executor-prompt.test.ts` ("global pause behavior") remain
red on main: **a user-paused todo task now reaches `createFnAgent`**.
That is a safety invariant rather than a stale fixture, and neither
`executeCore` nor the graph executor holds a pause gate — the refusal
#2371 documented is not where its note implies. It gets its own change;
editing the fixture to match current behaviour would hide it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:40:49 -07:00
gsxdsm
9da2674fa0 fix(tests): 2 TaskCard reds — the assertions pinned a jsdom detail, not the CSS (#2782)
## What was red

Both failures in the dashboard `app:components-b` lane. **Neither is a
style regression** — the CSS is byte-unchanged and correct in both
cases.

## Root cause: a jsdom upgrade, not a CSS change

jsdom does not substitute `var()`. What it does *instead* changed under
us in **4819c2634 (jsdom 27.4.0 → 29.1.1)**:

| | jsdom 27 | jsdom 29 |
|---|---|---|
| unresolvable shorthand (`padding`) | echoes raw text `var(--space-xs)
var(--space-sm)` | computes to `"0"` |
| single-value longhand (`gap`) | echoes | still echoes |

Tests asserting the **echoed string** were pinning a jsdom
implementation detail. The bump turned them red with nothing changed in
the product.

### 1. `renders a promote action when onPromote is provided`

`expected 'var(--space-xs) var(--space-sm)', received '0'`.
`.card-promote-action` still declares exactly that padding
(`TaskCard.css:1129`).

### 2. `FN-4511 keeps GitHub badge and timer chip geometry in parity`

`expected '1px' to be 'medium'`. The chips **are** in parity:

- badge: `border: 1px solid transparent`
- timer chip: `border: var(--btn-border-width) solid transparent`
- `--btn-border-width: 1px` (`styles.css:183`)

Here jsdom **discards** the unresolvable width rather than echoing it,
so `borderTopWidth` falls back to the initial value `medium`. The
existing `|| "1px"` fallbacks could not save it — `medium` is a
non-empty string, so it was the fallback that never ran, not the value
that was missing.

## The fix

Both assertions now read the **declared** value from the mounted
stylesheet's CSSOM and resolve a single `var()` against `:root`. That is
stable across jsdom versions and is what the assertions always meant.

Via the CSSOM rather than a regex over the CSS text **on purpose**: a
hand-rolled matcher over grouped selectors silently matches the wrong
rule and still reports success. Everything jsdom *can* resolve
(font-size, line-height, gap, padding parity) stays asserted against
computed style.

## Evidence

`components-b`: **1688/1688** (was 2 failed). Mutation-proved — both
fail as they should:

| mutation | result |
|---|---|
| timer chip border `1px → 2px` | `expected '2px' to be '1px'` |
| promote padding tokens changed | `expected 'var(--space-sm)
var(--space-lg)' to be 'var(--space-xs) var(--space-sm)'` |

**My first border mutation passed**, which would have read as a vacuous
guard. It had patched the wrong one of five identical `border:
var(--btn-border-width)` lines in the file. Re-run against
`TaskCard.css:1094` it fails correctly. Recorded because the mutation,
not the guard, was the thing that was wrong — a passing mutation is a
claim that needs checking too.

`pnpm lint` clean. Test-only — no production file or CSS touched
(mutations reverted; `git diff` clean).

## Ownership

`packages/dashboard/app` belongs to the **batch-dashboard-app** owner
(u12) under the mega-batch split. This is fix-forward on a red rather
than a conversion, confined to one test file, and touches no production
code — it should not conflict with the batch.

## Not fixed here

The `app:app` lane has **10 pre-existing failures** in `App.test.tsx`
(deep-link handling, board branch filters, FN-5817 mobile shell). They
were masked by this lane failing first — the runner skips remaining
lanes after the first failure, so they only became visible once
components-b went green. Separate change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:34:42 -07:00
gsxdsm
c643d62e85 fix(executor): wipDeclared must ask ALL six lifecycle roles, not two (#2777)
## What this fixes

`resolveResumeLanes` returns `wipDeclared`, which gates whether
`routeGraphFailureToExecutionResume` may route a graph failure back into
execution resume. Getting it wrong terminalizes tasks on boards that
should resume.

Two prior versions were wrong, both caught in review rather than by me
at write time:

**1. Two-state (`lifecycle?.wip !== undefined`)** — greptile P1 on
#2760. A v1-upgraded board terminalizes: `synthesizeDefaultColumns`
emits `{ id, name: id, traits: [] }`, so *every* role resolves
`undefined` even though those columns literally are the legacy lanes.
Verified by parsing a real v1 IR.

**2. Proxying "synthesized" as "hold and review are both undefined"** —
my own fix for (1), and also wrong. I caught this against #2765 rather
than shipping it. A **v2** board that declares only `intake` +
`complete` has hold and review undefined too, so it would be misread as
synthesized and treated as declaring wip when it deliberately does not.

The failure mode both versions share: reading a *sample* of the roles
and treating the answer as a verdict about the whole IR. #2765 says it
directly — an empty result has two meanings, and you cannot tell them
apart from a subset.

## The rule

```ts
wipDeclared: lifecycle?.wip !== undefined || !declaresAnyLifecycleRole(lifecycle),
```

Three states, asking all six roles:

- **wip declared** → true, the board says so.
- **some role declared but not wip** → false. A v2 board that omits wip
means it; do not resume into a lane it did not define.
- **no role declared at all** → true. That is the
synthesized/v1-upgraded shape, whose columns *are* the legacy lanes; the
pre-existing behaviour is correct there and must not regress.

`declaresAnyLifecycleRole` iterates `Object.values(lifecycle)` rather
than naming roles, so a seventh role added later is included
automatically instead of silently falling into the wrong branch.

## Evidence

- `executor-resume-lanes-resolved.test.ts`: **7 passed**, +23 lines
covering the v1-synthesized board and the declares-some-but-not-wip
board.
- **Mutation:** restoring the naive two-state rule → **1 failed / 6
passed**. The added coverage is load-bearing and pins exactly the
regression greptile caught.
- Gate **732 green** · `pnpm lint` clean · engine `tsc --noEmit` **0
errors**.
- Rebased on current main.

## Scope

`executor.ts` (+22) and its test (+23). One predicate; no other behavior
touched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 09:28:41 -07:00
gsxdsm
8c9b84ae38 batch-core: packages/core + dashboard/src lifecycle conversion (129 → 92) (#2780)
## batch-core — `packages/core` + `packages/dashboard/src`

Shared branch: two workers are converting into it. Opening the PR
because the branch was green with none, and a branch without a PR merges
nothing.

### Census

Measured with `node scripts/lifecycle-column-census.mjs --json`.

| | guards |
|---|---|
| batch-core scope at branch point | 129 |
| batch-core scope now | **92** (51 files) |
| repo total now | 358 |

Files closed so far: `store.ts` 11→0, `task-merge.ts` 6→0,
`live-agent-count.ts` 6→0 (marked, not converted — see #2762),
`task-update.ts` 3→0, display-ordering + Wake Delta ranking 5→0,
`register-git-github.ts` 4→0.

### The `register-git-github.ts` slice

Three PR routes — `pr/create`, `pr/push-branch`, `pr/resolve-conflicts`
— plus the `CHANGES_REQUESTED` handler each compared `task.column !==
"in-review"`. On a renamed board **none** of them matched, so every PR
affordance the dashboard offers was refused for a card sitting in the
lane that board calls review, and the refusal named a column that does
not exist there.

All four now share one helper, `reviewColumnsForTask`, which gets two
things right that this program has repeatedly gotten wrong:

- **Membership, not a single id.** It takes the broad review set
(`mergeOrchestration ∪ mergeBlocker ∪ humanReview`).
`resolveLifecycleColumns` returns the *first* column per trait, so a
single-id answer silently ignores a board that declares a merge lane
**and** a separate human sign-off lane. These guards only refuse or
permit — they never move the card — so over-admitting costs nothing
while under-admitting refuses a request that should have worked.
- **An empty resolved set means UNEXPRESSED, not absent.**
`synthesizeDefaultColumns` upgrades a v1 graph by emitting every default
column with `traits: []`, so a v1-upgraded workflow resolves to an empty
review set while its `in-review` column plainly exists and holds the
card. Reading empty as "this board has no review lane" would refuse
these routes on **every pre-v2 project** — a worse regression than the
one being fixed, and invisible to any v2 test.

This is the dashboard twin of the `fn pr create` guard in
`packages/cli/src/commands/pr.ts` (#2775). The two surfaces answer the
same question and now agree — FN-5893 surface enumeration.

### Testing note: why the seam and not the routes

I wrote route-level HTTP tests first and **deleted them**. An express
fixture over `registerGitGitHubRoutes` hangs — every case, including the
pure refusals, times out at 4s, because registering the router starts
background work the fixture never satisfies. Making it run would mean
mocking git, the GitHub client, and the pollers: a mock-the-world shell,
which is what the project's do-not-add-slow-tests rule (FN-5048) says to
avoid in favour of a narrow seam.

`reviewColumnsForTask` *is* the narrow seam — it holds the entire
decision, and the four call sites now do nothing but ask it and render
its answer. Six cases pin it: the renamed lane is returned and
`in-review` is not, a two-lane board returns both, a v1-upgraded board
falls back, an unresolvable workflow falls back, and the refusal renders
lanes an operator can act on.

**Mutation-verified, both directions:** reverting the helper to the
legacy literal fails 2 of 6; treating an empty set as an answer fails 1
of 6.

One fixture bug worth recording, since it would have made the two-lane
case vacuous: the trait id is kebab-case `human-review`, not
`humanReview`, and the built-in traits must be registered via `import
"@fusion/core"` before flags resolve.

### Verification

- `pnpm --filter @fusion/dashboard exec tsc --noEmit -p tsconfig.json` →
0 errors
- `pnpm lint` → 0 errors
- `register-git-github.review-lanes.test.ts` → 6 passed

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:28:27 -07:00
gsxdsm
b42b40aa48 fix: make flushAsyncWork actually drain — clears 32 notifier failures on main (#2776)
## Red on main, fix-forward

batch-engine (#2773) left 32 failing cases on main:
`notifier.runtime.test.ts` (19) and `notifier.test.ts` (13), all
`expected "vi.fn()" to be called 1 times, but got 0 times`.

## The failure is in the harness, not the notifier

`notifier.ts` is untouched by #2773.
`notification/notification-service.ts` was converted, and
`handleTaskMoved` is fire-and-forget (`void
this.handleTaskMovedAsync(data)`). The async path now awaits
`resolveLifecycleColumnsForTask` and then `resolveReviewColumnsForTask`
(which awaits `resolveWorkflowIrForTask`) — several more await hops
after `store.emit(...)` returns.

The harness had no slack to absorb them:

```ts
export async function flushAsyncWork(): Promise<void> {
  await vi.waitFor(() => { expect(true).toBe(true); });
}
```

The condition is true on the first tick, so `waitFor` resolves
immediately. **It never waited for anything.** It worked only while the
handler completed within a single turn — and it reported the resulting
breakage as a notifier defect rather than as its own.

Another entry in the recurring pattern this program keeps hitting: a
cheap check that reads as authoritative. A `waitFor` looks like
synchronization at the call site; this one was a no-op.

## Evidence

| run | result |
|---|---|
| before | 32 failed / 68 passed (100) |
| after | **100 passed (100)**, 3.25s |
| after, mutated back to a single `await Promise.resolve()` | 32 failed
/ 68 passed — the same 32 |

The mutation run is the point: the fix is load-bearing, not a
coincidence of timing.

Gate 732 green · `pnpm lint` clean · engine `tsc --noEmit` clean.

## Reversible decision, noted: microtasks only

A `setTimeout(0)` drain also turns all 100 green, and it was my first
version. Rejected on measurement:

- it costs real wall-clock at every call site — the two files went **~2s
→ over 2 minutes**;
- it **stalls under the fake timers** `notifier.test.ts` installs (lines
355/381/589), where a pending `setTimeout` never fires — 4 cases hung.

The awaits being drained are promise-based (workflow-IR resolution), so
microtask turns are the right currency, they work identically under real
and fake timers, and they cost nothing. Per AGENTS.md *"Do Not Add Slow
Tests"* — prefer fake timers over real time waits.

16 turns is slack, not a tuned number; the chain is ~4 deep today.

## Scope

One test-harness file. No product code, no behavior change. Tests
asserting a specific outcome should still prefer `vi.waitFor` on *that
outcome* — this helper covers the "let the fire-and-forget handler
finish" case, and now actually does it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:25:21 -07:00
gsxdsm
c428eed9d7 fix(test): two dashboard-api reds — conversions changed the log CHANNEL and the error MESSAGE (#2774)
Both red on main. `api:curated` goes **2 failed → 34 files / 1599
passed**. Neither is a product defect — both are conversions the tests
had not followed.

## 1. The log channel moved

`sse.test.ts` spied `console.log`. `sseDebug` routes through
`createLogger("sse").debug` (`sse.ts:50-53`), and the shared logger
writes debug lines to **`console.error`** carrying a `\0fnlvl=info\0`
severity marker — that is the point of FN-8603's adapter.

So the spy saw nothing, and the failure read `expected false to be
true`, naming neither the channel nor the logger. The stderr in the run
output showed the lines being emitted the whole time:

```
fnlvl=info [sse] [sse] + connection (active=1, hwm=2)
fnlvl=info [sse] [sse] - connection (active=0)
```

## 2. The error message is now built from resolved lanes

`routes-tasks` asserted the substring `"in-review or in-progress"`. The
message is now:

```ts
const allowed = [...prFeedbackReviewColumns, prFeedbackWipColumn]
  .map((column) => `'${column}'`).join(" or ");
throw badRequest(`PR feedback can only be addressed for tasks in ${allowed}`);
```

so it reads `'in-review' or 'in-progress'` — quoted, and derived from
the resolved columns.

**Asserted each lane separately rather than re-pinning the joined
string.** The join order and separator are presentation; the lanes being
the resolved review + wip columns is the fact this case owns. Re-pinning
the punctuation would break again on the next formatting change *and*
would not have caught a wrong lane — which is the failure this test
exists to catch on a renamed board.

## Verification

| check | result |
|---|---|
| `test:quality:api:curated` | 2 failed → **34 files / 1599 passed** |
| `sse.test.ts` | **24 passed** |
| `routes-tasks.test.ts` | **99 passed** |
| `pnpm lint`, dashboard `tsc` | clean |

## Scope

Fix-forward only, per the u9 lane. Found by re-scanning the packages
after #2739 / #2744 / #2754 merged, rather than by waiting for a report.

For the record on the other groups at the same commit: `components-a`
**1195 passed**, core is **2 failed** — both already accounted for
(`archived-column-gate-parity` is #2768's target,
`agent-logs-and-monitor.pg` is the deferred funnel/analytics decision on
#2669).
2026-07-30 09:22:05 -07:00
gsxdsm
1fb53f9924 docs(scheduler): the task:moved arms are blocked by a synchronous prologue — measured, and deliberately left counted (#2771)
## No behaviour change, and deliberately **no markers**

The ten `from`/`to` comparisons in `scheduler.ts`'s `task:moved` handler
stay **counted** in the census. They are genuinely wrong on a renamed
board — real backlog. Marking them `DELIBERATE-LITERAL` would claim
*"reviewed, correct"* when the truth is *"reviewed, still broken,
blocked on an ordering question"*, and that is the opposite of what
#2767's markers were for.

This records the blockage instead. I have deferred these twice citing
risk; this is the analysis that deferral was standing in for.

## The measured constraint

**The handler is `async`, but its prologue is not.** There is no `await`
anywhere between the handler's first line and the terminal-blocker
branch ~55 lines down. The snapshot invalidation, the PR-monitor
start/stop pair, the mission hand-off and the failed-task tracking all
run in the **same tick as the emitter**.

So hoisting a resolution to convert those arms does not cost "one await"
— it converts the **whole prologue into a microtask**, reordering this
listener against every other synchronous `task:moved` subscriber and
against the emitter's own continuation.

That makes `resolveTaskParkedColumnsSync`'s *"SYNCHRONOUS on purpose"*
note **load-bearing rather than stale** — verified by measurement, not
assumed. I had been treating it as possibly-stale boilerplate.

## Why the two obvious workarounds don't apply

- **Resolve lazily inside the branch.** Doesn't help: the *condition* is
what needs the lanes, and it is evaluated in the prologue.
- **A cheap sync superset prefilter** — the shape that worked in
`usage-limit-detector` — needs a literal predicate that cannot wrongly
*exclude* on an unknown vocabulary. For *"is `to` the terminal lane?"*
no such predicate exists: a renamed board's terminal id is unknown by
construction. (That is precisely why the prefilter *was* safe there —
literals can only fail to exclude, never over-exclude.)

## What would actually unblock it

1. **Audit the ordering**, then hoist one await and convert all ten
together. That is an audit across every `task:moved` emitter and
subscriber — not a scheduler-local change, and not something to do
speculatively.
2. **Carry the resolved lanes on the event payload**, so no listener
resolves at all. This is the only option that scales to the *other*
synchronous listeners with the same problem, and it removes the class
rather than one instance.

I'd recommend (2) if this is worth funding — it is the same shape as the
fix that removed the sync-resolution class in #2759, one layer up.

## Verification

11 scheduler suites — **130 passed** · `pnpm test:gate` **158 / 487 / 10
/ 71** · `pnpm lint` clean · engine `tsc --noEmit` **0 errors** ·
`--strict` exits 0, census unchanged at 12 for this file (which is the
point).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added internal documentation explaining synchronous event-ordering
requirements when resolving task lanes.
* Clarified why asynchronous resolution must not be introduced in this
scheduling flow.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 09:10:02 -07:00
gsxdsm
3092c9c2bb fleet: live-agent-count.ts 6 → 0 — the no-enrichment fallback, marked not converted (#2762)
Unclaimed file, no overlap with any open fleet PR. Previous PR (#2756)
is merged, so this is my one open PR.

## Census

| | before | after |
|---|---|---|
| backlog | 447 | **441** |
| reviewed | 38 | 44 |
| this file | 6 | **0** |

`--strict` exit 0, baseline re-recorded in the same commit.

## Why marked, not converted

All six literals sit after a `??` or a `flags ? … :`. Each is reached
**only** when the caller supplied no trait flags and no enriched shape —
precisely the case `enrichRunningAgentTaskShape` (takes the IR) and
`enrichRunningAgentTaskShapeFromFlags` (takes board flags) exist to
remove. There is nothing to resolve from there, so the choice is not
convert-vs-literal; it is **known legacy answer vs a different guess.**

**And the guess is not neutral.** Running and Waiting are *complements*
over the same rows:

```ts
isWaitingAgentTask = !running && (columnIsIntakeOrHold ?? isLegacyPreImplementationColumn(column))
```

A card matching neither arm is reported as **neither running nor
waiting**, so the footer's queued total silently under-reports it.
Guessing "not WIP" or "not review" loses cards from the count; the
legacy id at least matches every pre-rename board. That is why this file
already carries a `DELIBERATE-LITERAL` marker above
`isLegacyPreImplementationColumn` with the same argument — this PR
extends it to the three functions holding the remaining fallbacks
(`enrichRunningAgentTaskShapeFromFlags`, `terminalKind`,
`isRunningAgentTask`).

**The fix for a renamed board is at the CALLER** — pass flags, or use
the IR-taking enricher. Noted at the site.

## Pattern note for the fleet

This is the third file I have taken where `N → 0` is reached by marking
rather than converting, and they share a shape worth naming: **a literal
after `??` or in the `else` of a `flags ?` ternary is a degraded-mode
answer, not an unconverted guard.** The trait path is already there and
already correct; the literal is what runs when the trait path has no
input. Deleting it does not remove a decision — it substitutes a
different one, silently, in exactly the states where nobody is looking
(first paint, un-enriched callers, pre-rename data).

## Verification

Core typecheck clean · `live-agent-count.test.ts` 11/11 · `--strict`
exit 0 · comments only, no behavior change.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:57:52 -07:00
gsxdsm
6386be6626 test(dashboard): four more causes in components-b (14 → 0) — incl. a parity guard jsdom 27 made vacuous (#2743)
## Four distinct causes

| # | Cause | Files | Fixed |
|---|---|---|---:|
| 1 | **Portal** — `container` is empty; modal renders via
`createPortal` | `TaskDetailModal` | 6 |
| 2 | **Renamed testid** — `wf-add-step-modal` no longer exists |
`WorkflowNodeEditor` | 3 |
| 3 | **Unresolvable CSS** — jsdom can't resolve `currentColor` |
`SecretsView` | 4 |
| 4 | **jsdom 29 initial values** — `auto` vs `""` |
`TaskCard.badge-wrap` | 1 |

**1. Portal (6).** Same cause and fix as #2735 — 13 queries moved to
`document`. The symptom pointed away from it: assertions failed with
*"received value must be an HTMLElement / Received has value: null"* on
the **element**, while `screen.getByTestId` in the same test kept
working, because `screen` queries `document`.

**2. Renamed testid (3).** `wf-add-step-modal` exists nowhere in app
source — verified by grep, not inferred. The add-step dialog is a
`FloatingWindow` now (`WorkflowAddStepModal.tsx:145`,
`windowKey="workflow-add-step"`), so the stable id is
`floating-window-workflow-add-step`. It still scopes the
`within(dialog)` queries, so those keep their precision.

**3. Unresolvable CSS (4).** `SecretsView` compared
`getComputedStyle(svg).stroke` against the button's background.
`SecretsView.css:227` sets `stroke: currentColor`, which jsdom does not
resolve — every icon returned `rgba(0, 0, 0, 0)`, **equal to** the
transparent button background. The comparison was two unresolved values
matching each other, not a visibility check.

`currentColor` *is* the element's `color`, which jsdom does compute, so
it now asserts the same invariant through a property that resolves.
**Load-bearing, measured:** adding `color: rgba(0,0,0,0)` to the icon
rule fails exactly those 4.

**4. jsdom 29 initial values (1).** `.card-menu-btn` declares no
`min-height`, and `auto` is the CSS **initial** value — jsdom 29 reports
it where 27 returned `""`. The intent ("nothing constrains the button's
height") is what `auto` states; `""` was pinning a jsdom-27 quirk.

| Check | Result |
|---|---|
| `TaskDetailModal` / `WorkflowNodeEditor` / `SecretsView` /
`badge-wrap` | **51 / 179 / 14 / 20 passed** |
| `pnpm lint`, dashboard app `tsc` | clean |

## Flagged, not forced — and it's the interesting one

`TaskCard.test.tsx`'s 2 remaining failures. *"FN-4511 keeps GitHub badge
and timer chip geometry in parity"* reads border widths through
`githubStyles.borderTopWidth || "1px"`.

**Under jsdom 27 both sides returned `""` and both defaulted to `"1px"`
— so the parity assertion passed while comparing nothing.** jsdom 29
resolves them and they differ:

- the chip's `border: var(--btn-border-width) solid transparent`
(`TaskCard.css:1082`) reports `medium`, because jsdom cannot resolve
`var()` inside a shorthand;
- `.card-github-badge` — which has **no rule in TaskCard.css**, only the
class in `TaskCard.tsx` — reports `1px` from elsewhere.

So jsdom cannot adjudicate this parity at all, and whether the two
genuinely differ *visually* is a question for the e2e screenshot suite.
Restoring a `|| fallback` would rebuild a vacuous guard; changing the
CSS to satisfy a test limitation would alter the product to fit its
harness. Left for someone who can answer it in a real browser.

Worth noting the general shape: the jsdom 27 → 29 bump did not "break"
these tests so much as **stop hiding** what two of them were failing to
check.

## Branch arithmetic

This branch is off `main`, so `components-b` still shows 52 failures
here: **50 are `inline-editing`, fixed by #2735** on its own branch,
plus the 2 flagged above. Once both land, the group is at 2.

Across #2735, #2740 and this PR, dashboard goes from **88 failures / 9
files** to **3** — the 2 above plus the GitHub-tracking affordance
question flagged in #2735. Per #2732 these lanes are still never
executed in CI, since the shard aborts on the first failing package.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Updated dashboard component tests to account for portal-based modal
rendering, ensuring queries target the global document.
* Refined jsdom assertions for icon visibility/styling and card header
control sizing to match real intended behavior.
* Adjusted workflow editor tests for the new add-step dialog, using
updated stable identifiers and updated interaction/close checks.
* Added clarifying comments to document jsdom-specific limitations and
expected outcomes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 08:51:44 -07:00
gsxdsm
9b573268a4 docs(u9): audit every non-scheduler execute() call site — exactly ONE ignores pause (#2747)
## What this closes

`executor-prompt.test.ts` has 3 failures asserting that `execute()`
refuses to dispatch a user-paused card during a global pause. The guard
lives in the **scheduler** (`scheduler.ts:1515`, a hard stop that never
reaches `.execute(`), so those tests call a layer that has never
enforced it.

I flagged this in #2719 without being able to say how large the real gap
was. This answers that — and the answer is much narrower than "execute()
is unguarded".

## Every non-scheduler call site

Enclosing method and guard status resolved **programmatically**, not by
reading nearby lines:

| Site | Enclosing method | Guarded? | Reading |
|---|---|---|---|
| `executor.ts:3333` | `dispatchUnpauseResume()` | no | **Correct
as-is** — it *is* the unpause path; a guard here is self-contradictory |
| `executor.ts:3495` | `constructor()` — `task:moved` sub, `to ===
"in-progress"` | **no** | **The one real gap** |
| `executor.ts:5702` | `resumeTaskForAgent()` | yes | `globalPause \|\|
enginePaused` + `!task.paused` |
| `executor.ts:5858` | `resumeOrphaned()` | yes | same guard earlier in
the method |
| `in-process-runtime.ts:2255` | `drainWorkflowContinuations()` |
indirect | gated by `status !== "active"`; engine pause is expected to
leave the runtime non-active |

**Exactly one path can reach `execute()` without consulting pause
state.** So the open question is not *"does `execute()` need a guard"*
but *"can a `task:moved` → `in-progress` event fire while paused"* —
narrow, and answerable by whoever owns the pause contract.

## A measurement correction worth recording

My first pass used a 40-line window above each call and **mis-attributed
two sites**: `:5702` and `:5858` looked unguarded because their guard
sits earlier in the same method, above the window. Resolving the
enclosing method properly flipped both to guarded and cut the apparent
gap from three sites to one.

That is the difference between reporting "3 of 5 paths ignore pause" and
the truth. A proximity heuristic is not an enclosing-scope analysis.

## Still not decided, deliberately

Three options, and they are not equivalent:

1. **Guard the `task:moved` subscription** — narrowest; keeps the
scheduler as the single pause authority. Does *not* make the three tests
pass, since they call `execute()` directly.
2. **Give `execute()` its own guard** — makes the tests pass, but must
not refuse the legitimate internal re-dispatch paths.
`dispatchUnpauseResume()` would break outright: it exists to resume a
card the operator just unpaused.
3. **Retire the three direct-`execute()` assertions**, covering the
invariant at the scheduler layer where it is enforced.

(2) and (3) both touch coverage of **user pause** — a safeguard this
program re-ratified and told workers not to narrow. Choosing either
silently inside a test-repair PR is how a safeguard gets weakened by
accident.

## What is NOT verified

Whether that subscription is actually **reachable** while paused.
Proving it needs a trace of who emits `task:moved` with `to ===
"in-progress"` under a global pause; if every emitter is itself gated,
the gap is theoretical. That trace is the remaining work before
preferring option 1 over the status quo.

Docs-only — no source, no tests. Lint clean.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added an audit documenting workflow pause behavior and identifying an
event-driven execution path that can create sessions during a global
pause.
* Recorded findings from existing test failures and reviewed available
enforcement points for pause handling.
* Documented trade-offs and the remaining decision on where pause guards
should be applied.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 08:51:32 -07:00
gsxdsm
fdd958efc9 fix(test): the archived-gate parity guard could not see an aliased table (red on main) (#2768)
`archived-column-gate-parity` is red on main after #2745 converted three
TypeScript sites. Fixing the stale inventory is the small half. **The
guard had a hole, and it is the interesting part.**

## The hole

The SQL scan required the predicate's receiver to be literally
`<x>.tasks`:

```ts
ne(schema.project.tasks.column, "archived")   // seen
const table = schema.project.tasks;
ne(table.column, "archived")                  // INVISIBLE
```

So `branch-group-ops.ts:58` was **never audited**. The inventory claimed
six files; seven exist.

A parity guard that cannot see one of the encodings reports agreement it
never checked — the exact failure mode this file was written to prevent,
occurring inside the file itself.

## And that site is not hypothetical

`#2745` converted `branch-group-ops`'s **TypeScript** half to the
resolved role. Its **SQL** half still compares the raw string. On a
board whose archived lane is renamed, the two disagree — one says a task
is archived, the other returns it as live. That is the split-brain
described in the guard's own failure message, and the guard could not
see it.

## The fix

Alias bindings (`const <id> = <...>.tasks`) are collected per file in a
first pass — first pass because the binding can appear *after* its uses
inside nested closures — and accepted as the receiver.
`branch-group-ops.ts` joins `AUDITED_SQL_SITES` as **newly visible, not
newly written**.

Also drops the three TypeScript entries #2745 converted
(`blocker-fanout`, `branch-group-ops`, `task-store-helpers`) — that is
the red itself.

## Measured, both directions

| mutation | result |
|---|---|
| alias set emptied (the old, alias-blind scan) | **1 failed** —
"Drizzle encoding changed" |
| product SQL half converted to a resolved lane | **1 failed** —
"Drizzle encoding changed" |
| as shipped | **2 passed** |

The first proves the scanner fix is load-bearing. The second proves the
newly-audited site is genuinely *counted*, not merely listed in an
inventory.

Two earlier mutation attempts produced no output and I discarded them
rather than reading them as passes — they had broken the file's syntax,
so nothing ran. A mutation that fails to compile proves nothing, and
looks identical to a clean run when output is filtered.

Gate **726**, core `tsc` clean, lint clean.

## Left for the owner of #2745

Whether `branch-group-ops.ts:58`'s SQL half should now be converted too.
The guard's own header explains why the SQL halves cannot simply be
converted (`ne(tasks.column, ...)` needs the resolved id as a value,
which the call sites do not all have), so this is a real design question
rather than a mechanical follow-up — and it is now *visible* and
*audited* instead of silently absent.
2026-07-30 08:48:21 -07:00
gsxdsm
9b61d795c9 fix(engine): heartbeat asked 'is this task finished?' with legacy ids — and one of the two sites writes status:failed onto completed work (#2769)
## Two heartbeat sites asked "is this task finished?" with the legacy
ids

`agent-heartbeat.ts` **4 → 0**. Neither site is cosmetic.

**Linked-task clear.** The heartbeat clears an agent's assignment once
its card is finished. Keyed on the literals, an agent on a renamed board
stayed bound to a **completed** card indefinitely — every later
heartbeat ran with stale task context instead of picking up new work,
and nothing else clears it.

**Worktree-acquisition gate.** Its failure bookkeeping runs only for a
**non-terminal** task. A card in a renamed complete lane read as
non-terminal, so an acquisition failure could stamp `status: "failed"`
and an error message onto work that was **already done**.

That second site *writes*, which drives the fallback direction: an
unresolvable workflow degrades toward "terminal", because treating a
finished card as unfinished is the expensive mistake here.

Both sit in async paths — the first has `await taskStore.getTask(...)`
three lines above — so this is an `await`, not a restructure. Extracted
to one predicate rather than converted twice: they are the same
question, and the two must not drift when one of them acts
destructively.

## How this was found, and the part worth recording

Generalising #2767. That PR marked a documented false positive the
census kept advertising, so I swept for **other** files whose lifecycle
literals were reasoned about in prose but still counted — to find out
whether the trap was systemic.

**It is not.** Of twelve candidate files, only this one carried real
unconverted guards, and its "false positive" mentions turned out to be
unrelated (detection heuristics, not column literals). The sweep mostly
came back **negative**, and that is worth saying so nobody repeats it
expecting a haul.

## Revert proof

| reverted | result |
|---|---|
| neuter the resolution (predicate → literals) | **3 failed** / 2 passed
|
| shipped | **5 passed** |

The two that survive the revert are the degraded-mode pair, which is
correct — they assert the *legacy* answer, so they must pass either way.
One case also pins that a legacy `done` id is **not** terminal on a
board that does not declare it, which is what a board-wide union would
get wrong.

## Verification

- 13 heartbeat suites — **501 passed** (496 before, +5 new)
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
engine `tsc --noEmit` **0 errors** · `--strict` exits 0

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:42:00 -07:00
gsxdsm
7119432c79 fix(engine): the stranded-completed recovery never resolved on a renamed board — and its suite fed the broken reader the right answer (#2764)
## The "recovery of last resort" never resolved anything

`recoverCompletedTask` carries this note in its own source:

> *This is the recovery of last resort — a literal here means the last
resort does not exist off the default lineage.*

It was resolving through `resolvePlannerLanes`, which reads
`resolveTaskWorkflowIrSync` — whose selection reader returns `undefined`
**unconditionally** in PostgreSQL mode, the shipped backend. So it
resolved the **default** workflow for every card,
`promotedFromPlannerColumn` was `false` on every renamed board, and the
recovery never fired.

That is precisely the stranding it exists to fix — completed work
sitting in a planning lane with nothing left to rescue it — **with the
conversion in place and the census counting it as done.**

The call site is inside an async method that has already awaited store
reads, so the fix is an `await`, not a restructure.
`resolvePlannerLanesForTaskAsync` is the async twin: identical logic,
identical fallbacks, one `await`. Answers are unchanged on the default
lineage and correct everywhere else.

## The existing suite could not see any of it — the more important half

`executor-planner-lanes-resolved.test.ts` injected **only**
`resolveTaskWorkflowIrSync`:

```ts
(store as { resolveTaskWorkflowIrSync: ... }).resolveTaskWorkflowIrSync = () => ir;
```

It fed the broken reader **the right answer**. Every case proved the
promotion *logic* while being structurally blind to whether production
resolves at all — and it was green the entire time. A suite that cannot
fail for the reason the code is broken is the same defect as the code,
one level up.

The harness now feeds the sync reader the **default lineage** (what it
actually returns) and the async readers the task's real workflow.

| | reverting the call site to sync |
|---|---|
| before this PR | **0 failed** — suite blind |
| after | **5 failed** / 13 passed |

Two cases opt back in via `syncResolvesIr`, and only those two: they
cover `isPlannerColumnFor` and `isBackwardMoveOutOfPlanning`, which are
still synchronous, so there the sync reader genuinely *is* the input
path and feeding it the IR tests the classifier rather than the reader.

## Not converted, deliberately

**Those two classifiers.** They sit in an else-if chain whose next arm
is `from === "in-progress"`, so deferring the decision into an async
body changes which arm runs. That branch's own comment records a
previous half-conversion there:

> *a half-conversion turned a missed rescue into active damage. Third
time this program has produced that shape — gates converted,
destinations left literal.*

That needs the chain enumerated first, not a fast restructure at the end
of a sweep. Their production inertness is held by the
`resolveTaskWorkflowIrSync` call-site allow-list in #2759, so they
cannot be forgotten.

## Verification

- new suite **4 passed** · strengthened suite **14 passed** (18
together)
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean ·
engine `tsc --noEmit` **0 errors** · `--strict` exits 0

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:38:43 -07:00
gsxdsm
18641ba5d2 fleet: the age-staleness hydration site #2746 missed (3rd time in this file), and a blocker that blocked forever (#2749)
## Census

| | before | after |
|---|---|---|
| `packages/core/src/task-age-staleness.ts` | 4 | **0** |
| `packages/core/src/blocker-fanout.ts` | 4 | **1** (the marked
no-metadata fallback) |
| repo backlog | 581 | **573** |

Baseline re-recorded; `--strict` exits 0.

## Both were the unconverted sibling in an already-converted file

That is the shape this program keeps re-finding, and both files here
even carry notes about *previous* P1s on the same question.

### 1. Age staleness never fired

`getTaskAgeStalenessSignal` returns `undefined` unless the card is in
wip **or** review, then picks its warning/critical thresholds by which
of the two it is. Keyed on the literals, a renamed board produced **no
age-staleness badge at all**.

That is the worst shape a monitoring failure can take: **a missing
warning is indistinguishable from health**. Nothing looks broken — "this
card has been sitting in progress for a day" simply stopped being said.

`reads.ts` already resolves `holdColumn` and `reviewColumn` per row for
the sibling signals. Its own comments record a P1 where exactly this
role was threaded into a helper but **omitted at both hydration sites**
— "same defect, same file, one role over". So this adds the third
resolver (`resolveWipColumnForTask`, mirroring the review twin) and
threads **both** lanes at **both** sites, off the same per-pass IR
cache.

The signal's reported `column` deliberately stays on the legacy id: that
field is its public shape, which consumers switch on, so renaming it is
a separate breaking change rather than part of resolving a guard.

### 2. A blocker that blocked forever

`isStaleBlockedByBlocker` decides whether a `blockedBy` marker is stale.
Keyed on the literals, a **finished** blocker on a renamed board never
read as stale — so the dependent kept its marker permanently and its
"waiting on" badge pointed at work that shipped days ago. Every path
that clears a stale marker consults this predicate first, so nothing
else rescues it.

`computeBlockerFanoutMap` — the **only** production caller — already
takes `terminalColumns`/`holdColumn`/`classify`, and the file documents
two separate P1s about getting this right. The predicate sat on the
literals and the call passed nothing.

**Both are fixed, and that matters more than it sounds:** converting the
predicate alone would have changed *nothing at runtime* while the census
scored it as a 4-site win. That is the half-conversion trap, and it is
why the wiring gets its own revert proof below.

## Revert proof — each reverted alone

| reverted | result |
|---|---|
| age-staleness lanes → literals | **3 failed** / 8 passed |
| blocker predicate → literals | **2 failed** / 9 passed |
| fanout **wiring** (`classify` not consulted) | **1 failed** / 10
passed |
| none (shipped) | **11 passed** |

Each group also carries a paired negative — a non-active lane still
raises no staleness signal, and a live blocker is still not stale — so
neither fix can degrade into "always fires".

## Verification

- 4 core suites (blocker/staleness/age/reads) — **33 passed**, no
regressions
- new suite — **11 passed**
- `pnpm test:gate` — **10 / 158 / 487 / 71** · `pnpm lint` clean · core
`tsc --noEmit` **0 errors**

## The 1 remaining

`blocker-fanout.ts` keeps one `DELIBERATE-LITERAL`: the no-metadata
fallback for an unconverted caller. Deleting it makes an unresolved
caller read every blocker as non-terminal, so stale markers would never
clear **at all** — strictly worse than the legacy behaviour it would
replace.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:35:24 -07:00
gsxdsm
ae4ff9c111 test(core): sync workflow resolution always returns the default IR — the proof and a call-site ratchet (#2759)
## A whole class of "conversions" in this program is inert, and the
census scores them as done

`resolveTaskWorkflowIrSync` returns the **default** workflow IR for
**every** task in production. `getTaskWorkflowSelectionImpl` is a
PostgreSQL-cutover stub that returns `undefined` unconditionally
(`workflow-definitions.ts:505` — *"Backend mode cannot synchronously
read PostgreSQL, so return undefined and let the sync reader fall back
to its default"*), so the sync resolver always takes its `!workflowId`
branch. Its return type is **non-optional**, so no caller can detect the
substitution.

**Proven, not argued.** The PG suite binds a task to a workflow whose
lanes are `drafting`/`building`/`checking`/`shipped`, asserts the
**async** reader sees that binding — or the next assertion would be
vacuous — and then shows the sync resolver answering `hold: "todo"` for
the same task. A third case pins the cause directly: the sync selection
reader returns `undefined` while the async one returns the workflow id.

The async resolver on the same task, in the same test, returns the real
lanes. So the remedy for any affected site is always *reach the async
resolver*, never *resolve synchronously and hope*.

### Why this is worse than an unconverted literal

```ts
resolveLifecycleColumns(store.resolveTaskWorkflowIrSync(id))?.hold
```

That **reads** as converted. It resolves an IR, asks for a trait, and
the lifecycle-column census counts it as **progress** — while being
wrong for every custom workflow, silently. A plain `=== "todo"` is
strictly better, because it is at least honest about being a literal.

That is why this needs a guard rather than a comment: the failure mode
is code that *looks right in review*.

### The ratchet earned its place immediately

It allow-lists the six call sites, in the shape this repo already uses
for `engine-no-blocking-shellout` and `check-no-nohup`. On its first run
it corrected the list I had seeded by grep:

- **Found `replan-target.ts`**, which my grep missed — it calls through
an optional-property cast, so no textual search for
`store.resolveTaskWorkflowIrSync` matches it. **Its hazard is the
sharpest of the six:** `resolvePlannerLanes` returns
`resolvedFromWorkflow: true` whenever an IR came back, so on a renamed
board a caller branching on that flag is told the lanes are
workflow-resolved while being handed the **default** ones.
- **Rejected `executor.ts`**, which my grep had matched on a *comment*
with no real call site.

It also carries a completeness case (fails if the scan finds nothing,
rather than passing vacuously) and a staleness case (an entry whose file
stops using the primitive must be removed in the same change, so the
list can't rot into unreviewed permission).

### Verified to fire

Adding a new unlisted call site fails with the offending file named:

```
+   "packages/engine/src/gridlock-detector.ts"
      Tests  1 failed | 2 passed (3)
```

My first attempt at that proof landed the probe **inside a block
comment** and the guard correctly reported zero — the methodology was
wrong, not the guard. Worth stating, since a guard I couldn't make fail
is exactly the thing this PR is about.

### Scope

**No source changes.** This pins a fact and guards a primitive. The six
existing sites are deliberately left alone: each needs its own
async-reachability analysis, which is per-site work with real behaviour
risk, not a sweep. `scheduler.ts`'s entry is the honest case — it is
called from synchronous `task:moved` listeners where adding an `await`
would reorder handlers against a synchronous emitter.

### Verification

- new PG suite **3 passed** · new ratchet **3 passed**
- `pnpm test:gate` — **158 / 10 / 487 / 71** · `pnpm lint` clean · core
`tsc --noEmit` **0 errors** · `--strict` exits 0

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:35:12 -07:00
gsxdsm
7f3eee9db7 docs(engine): mark the usage-limit terminal filters DELIBERATE-LITERAL — a documented false positive that has now baited two workers (#2767)
## No behaviour change. This is the work order retracting a documented
false positive.

I went to convert the `done`/`archived` filters in
`usage-limit-detector.ts`, reasoning that on a renamed board a provider
rate limit would pause already-finished work. I wrote the conversion —
and only then read the note a previous worker had left directly above
it:

> *"The FIRST thing I suspected there — the `done`/`archived` terminal
filter — turned out to be a **FALSE POSITIVE**: its revert stayed green,
because the lane check already excludes finished cards."*

**They are right and I was wrong.** A terminal card is already excluded
downstream: `taskUsesProvider` resolves the task's active lane, a
finished card matches no active lane, so it resolves no providers and
cannot be affected. The suite pins exactly this — `pauses a PEER
executing in the renamed WIP column` asserts `FN-SHIPPED` is not paused.

My conversion is reverted. It changed nothing at runtime and would have
lowered the census count while behaviour stayed identical — the precise
shape this program keeps warning about, produced by me this time.

## Why a marker and not just the existing prose

The note was already there and I walked into it anyway, because **the
census kept listing this file as 4 unconverted guards**. The work order
advertised the work; the reasoning against it lived in a comment you
only reach after you have started. Prose informs a reader who is already
looking; a marker informs the *instrument*, so the file drops out of the
work order.

Two distinct reasons are recorded rather than one blanket marker,
because they are not the same argument:

- **the prefilter** is a deliberate cheap **superset** (#2672 review).
Converting it reintroduces the whole-board resolution that review
removed. Literals are safe here in the direction that matters — a
renamed board declares no `done`/`archived` id, so nothing is wrongly
*excluded*.
- **the final filter** is redundant with the lane check, and that
redundancy is already proven by an existing test.

## Census

| | before | after |
|---|---|---|
| `usage-limit-detector.ts` | 4 | **0** |
| repo backlog | 437 | **433** |
| DELIBERATE-LITERAL (reviewed) | 38 | **42** |

Every one of the 4 is a marker, not a conversion. The backlog moved
because reviewed literals left it honestly, not because behaviour
changed.

## Verification

`usage-limit-detector.test.ts` **58 passed**, unchanged before and after
· `pnpm test:gate` **10 / 158 / 487 / 71** · `pnpm lint` clean · engine
`tsc --noEmit` **0 errors** · `--strict` exits 0.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:31:47 -07:00
gsxdsm
2e4905fa0e refactor: one definition of "which columns are review" — three copies deleted onto core's resolver (#2751)
**#2730 added `resolveReviewColumns` to core. This deletes the three
copies that predated it.**

Measured on `origin/main` before this change — three in-tree
definitions, **none of which agreed**:

| site | definition |
|---|---|
| `core/workflow-lifecycle-traits.ts` (#2730, authoritative) |
mergeOrchestration ∪ mergeBlocker ∪ humanReview — **all** columns |
| `dashboard/routes/register-task-workflow-routes.ts` | mergeBlocker ∪
humanReview ∪ **first** mergeOrchestration |
| `cli/src/extension.ts` | mergeBlocker ∪ humanReview ∪ **first**
mergeOrchestration |
| `cli/src/commands/task.ts` | all three, full union |

**Both `.slice(0, 1)` variants are mine**, from #2723's review round: I
narrowed to core's then-single `.review` because the reviewer was right
that a superset let the dashboard act on a lane the engine did not own.
#2730 answered that question authoritatively in the other direction, so
the narrowing is obsolete.

Worse, and the part that makes this urgent rather than tidy: **the two
CLI copies had already drifted apart inside #2728.** `fn_task_retry`
refused a card in a second merge lane that `fn task retry` accepted —
two surfaces, one operator action, two answers, from two copies of one
definition written days apart by me.

All three now call core. The dashboard keeps its thin store→IR wrapper
(its callers hold a store and a task id, not an IR) but the **body** is
core's.

## One assertion inverted, deliberately

My #2723 case asserted that a **second** `mergeOrchestration` column is
**refused**. Core says every merge lane is review, so the behaviour
legitimately changed and the assertion flips with it.

**Kept rather than deleted**, because the invariant under test — *the
routes agree with core* — is unchanged. Deleting the case would have
hidden that its answer moved; inverting it records which decision moved
and why. A test whose expectation quietly disappears is
indistinguishable from a test that was wrong.

## A footgun found while rebasing

The shipped signature is
`isInReviewMissingWorktreeSessionStartFailure(task, isReviewColumn?:
boolean)` — the merged version takes the **answer**, not the lanes. My
branch had passed a `ReadonlySet`, and because the parameter is `boolean
| undefined` with a `??` default, **a truthy object makes it answer
`true` for every column**.

TypeScript stops typed callers; my test only reached it through an `as
never` cast, which is how I found it. All three production call sites
correctly pass `retryReviewColumns.has(task.column)` — now asserted
structurally so a fourth surface cannot omit it.

The boolean is arguably the better shape, and I'd keep it: there is
nothing left for the callee to re-derive, so it cannot disagree with the
caller's own membership test.

## The ratchet

No surface may reintroduce a local review union (`columnsWithFlag(…,
"mergeBlocker" | "humanReview")`). Those three copies appeared because
each was added **in good faith, in a different review round, by someone
reading only their own call site** — which no amount of care prevents
and a ratchet does.

## Verification

census **553** · `pnpm test:gate` **487 / 10 / 71** · `tsc` clean in cli
and dashboard · `pnpm lint` clean · 10/10 in each touched suite.

**Pre-existing, not mine:**
`register-task-workflow-routes.move-bypassguards.test.ts` fails on
`origin/main` (400 vs 200) — already reported on #2723.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:28:32 -07:00
gsxdsm
ea008b4064 docs(core): an empty role result has TWO meanings — correcting a rule I applied across three PRs (#2765)
## A correction to a rule I introduced, prompted by #2760 and then
measured

`synthesizeDefaultColumns` (`workflow-ir.ts`) upgrades a v1 graph by
emitting `{ id, name: id, traits: [] }` for the five default ids —
**placement only, by design**, with the real trait set living in
`BUILTIN_CODING_WORKFLOW_IR`.

So a v1-upgraded board arrives at every role resolver looking exactly
like a v2 board that declares nothing. Measured on such an IR:

| resolver | result |
|---|---|
| `resolveLifecycleColumns` | `{}` — every role undefined |
| `resolveReviewColumns` | `[]` |
| `columnsWithFlag(…, "countsTowardWip")` | `[]` |
| `resolveTerminalColumns` | `["done","archived"]` — its own legacy
fallback saves it |

## The correction I owe

I introduced **"resolved and empty means this board declares no such
lane"** deliberately across #2731, #2733 and #2734, to fix the
*opposite* bug — a legacy fallback masking a genuinely absent lane. I
argued for it repeatedly and applied it in several files.

It is right for hand-written v2. It is **wrong for the upgrade path**,
where it withdraws every role at once. And the two cases are
indistinguishable at the call site: both arrive as an empty array.

The contrast the test pins is the sharpest evidence —
`resolveTerminalColumns` survives the identical IR purely because it
never adopted that reading. **The data is the same; the reading is what
differs.**

## Affected, named rather than left to be rediscovered

- **on `main` today**: `default-workflow-hooks.ts` → `inRole` (from
#2734)
- **in my open PRs**: the store bypass guard (#2709) and the
tracking-state terminal classifier (#2754)

Consumers that kept a `length > 0 ? resolved : legacy` guard are
unaffected — including the notifier (#2722), which I checked.

## Not fixed here

The root repair is for the upgrade to carry real traits instead of
placeholders. That changes behaviour for **every persisted v1
workflow**, so it wants its own change with its own measurement — not
something slipped into a doc commit. Flagged at the source so the next
consumer chooses knowingly.

## Verification

34 trait tests green · core `tsc` clean · lint clean (0 errors) · gate
green (487 + 158 + 10 + 71). No production behaviour change; no census
movement.
2026-07-30 08:25:08 -07:00
gsxdsm
d86c1f9d29 batch-engine: packages/engine lifecycle-column conversions (capacity worker's mega-batch) (#2773)
The engine mega-batch. Folds my four engine PRs and will absorb the
remaining `packages/engine` guards as commits on this branch.

**Superseded and closed:** #2722, #2741, #2766, #2770.

## Census — files converted so far

| file | before | after |
|---|---:|---:|
| `notification/notification-service.ts` | 9 | **5** |
| `runtimes/in-process-runtime.ts` | 6 | **1** |
| `eval-followups.ts` | 2 | **0** |
| `pr-comment-handler.ts` | 1 | **0** |
| `task-revert.ts` | 2 | **0** |

The last two are **census-invisible** (`Set.has(task.column)`
membership) — the class measured in #2763, which a comparison-based scan
cannot count. So the backlog number moves less than the work does,
deliberately.

## What each one actually fixed — all silent, none cosmetic

- **Notifications stopped entirely.** `handleTaskMovedAsync` compared
`data.to` to `in-review`/`done`, so on a renamed board the two
notifications operators rely on most were never sent.
- **A finished card's plan review could re-enter.** The continuation
drain's terminal test matched nothing, so a completed card's planning
continuation was handed to the executor.
- **The revert route admitted and the service refused.** The route
resolved terminal lanes; the service compared to a hardcoded pair. The
operator got a dead end from an affordance the UI and route both
offered.
- **Follow-up dedup blocked new cards forever.** A finished follow-up in
a renamed complete lane read as *open*, so the dedup matched it
permanently — defeating the intent the code documents in the line above
it.
- **The mission requeue wrote a column that may not exist**, and its
guard never matched.

## Flagged, not fixed — deliberately

- **`concurrency.ts` idle semaphore leak recovery** — the last live
caller of the running-agent predicate that does not enrich. On a renamed
board it under-counts and can reclaim a legitimately-held slot. The
enriching variant is async and this is a synchronous repair path whose
failure mode is reclaiming live work.
- **The archival `task:moved` listener** — runs on every move with no
cheap gate ahead of it; converting costs an IR resolution per move to
decide most moves are not archival.

## Notes carried from the folded PRs

Two conflicts resolved in main's favour because **main's version was
better**: `in-process-runtime`'s seam uses `terminalColumns:
ReadonlySet` (membership) where mine used `LifecycleColumns`
(first-per-role), and the test is rewritten against main's API. That
arity trap has now caught me four times, so membership is the default
shape in everything new here.

Review fixes from the folded PRs are included: the notifier's review
set, the second human-review site, the second dedup copy, the workspace
revert surface, and the file-content assertions.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **224 passed** across
the touched engine suites · engine and dashboard `tsc` clean · `pnpm
lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:24:50 -07:00
gsxdsm
245086dad6 docs: the census total is a floor — 25 membership predicates it structurally cannot see, one a live defect (#2763)
Docs only, extending the entry #2748 landed. Opening it because the
fleet reads the census total as its completion bar, and that total
excludes a whole predicate class — a measurement that should not live in
a chat reply.

## Measured on `origin/main`

- **47** array/Set literals of two or more lifecycle ids, in 35 files.
- **25 are membership predicates against a task's column** —
`SET.has(task.column)` / `ARRAY.includes(task.column)` — in 19 files.
Two are documented fallbacks behind a resolved primary, so **~23 are
unconverted guards**.
- The census scans `===` / `!==` against a column. **None of these is a
comparison, so none is counted.**

| file | constant |
| --- | --- |
| `cli/src/commands/task.ts` (3) | `retryReviewColumns` |
| `dashboard/app/components/TaskCard.tsx` (2) | `TIME_INDICATOR_COLUMNS`
|
| `engine/src/eval-followups.ts` (2) | `OPEN_COLUMNS` |
| `engine/src/merger.ts` (2) | `sourceTerminal` |
| `engine/src/task-revert.ts` (2) | `REVERTABLE_COLUMNS` |
| `core/src/agent-role-policy.ts` (1) | `IMPLEMENTATION_TASK_COLUMNS` |

## One is a proven live defect

`isImplementationTask` is
`IMPLEMENTATION_TASK_COLUMNS.has(task.column)`, and
`evaluateImplementationTaskBind` short-circuits to `allowed: true` when
it returns false. **On a renamed board every agent is bind-compatible
with every task** — the role check that stops a liaison being handed
implementation work (the NEXT-871 loop FN-7851 fixed) does not apply.

It surfaced only because a reviewer questioned a coverage claim in one
of my dispatch tests (#2739). Passing an agent wasn't proof the
evaluator ran, so I asserted a `custom`-role agent must be *refused* —
and that test failed against production. Flagged at the site in #2739,
not fixed: `isImplementationTask` is a sync pure predicate with no
store, and making the routing policy async is a behaviour change to
agent admission.

## What this does and does not argue

The census is the right instrument — AST-based, honest about what it
measures, and it has caught real drift in both directions (it failed on
me in #2724 when merged conversions moved an inventory *down*). This is
not an argument against it.

It is an argument against reading **"backlog: N" as "N guards remain"**.
The same shape already appeared in the archived gate (#2724), where the
rule is additionally encoded in Drizzle predicates and raw `sql`
templates that no comparison scan can see. Two independent classes now,
found the same way — by looking at what the instrument's definition
excludes.

**Extending the census to count membership predicates is deliberately
left to you, not done here.** It would move every worker's number
mid-fleet, and deciding which sets are lifecycle guards versus
board-config definitions or type unions is exactly the judgement
`DELIBERATE-LITERAL` exists for — 47 collections would each need that
call.

## Verification

`pnpm lint` clean · census `--strict` exits 0 · no code changes.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:12:51 -07:00
gsxdsm
86a2b48968 fleet: branch-group-ops 6 → 0 — an agent asking for its next task was told there was none (#2739)
Second application of the sync-filter pattern decided in #2737.
`branch-group-ops.ts` 6 → **0**.

## The failure

`selectNextTaskForAgentImpl` picks an agent's next task by filtering the
board for its WIP lane, then its hold lane — both `task.column ===
"<literal>"`.

On a renamed board **both filters match nothing**, so an agent asking
for work is told there is none, with its own assigned tasks sitting in
the list it just fetched. No error, no log line. The agent idles.

`pauseTaskImpl` had the same shape: pausing a running card on a renamed
board left its `status` untouched, so the UI kept showing it as working.

## Consumer, not a gate — checked rather than assumed

Applying the #2724 test to this file, since it sits closer to the
persistence layer than the reconciler did: its **only** SQL predicate is
`eq(table.projectId, ...)`. Nothing here compares a column to a literal
in SQL, so there is no second encoding of these questions to diverge
from. The list arrives from `store.listTasks` and the filters select
among rows already in hand.

Async predicates were the alternative and would have turned these filter
chains into sequential awaits inside the dispatch path. One prefetch,
one IR read per distinct workflow, filters stay synchronous.

`pauseTaskImpl` resolves for the single task it holds rather than
joining a map — different entry point, one id in scope, and a map would
have exactly one entry.

## Revert proof

| reverted | result |
|---|---|
| the wip literal | renamed WIP case fails: `expected null to be truthy`
|
| the hold literal | renamed hold case fails identically |

The new test calls the impl **directly** with a store fake resolving a
renamed IR. The existing `selectNextTaskForAgent` coverage drives a real
store harness, so exercising a renamed vocabulary there means
registering a real custom workflow and moving cards through it — heavier
than the question, which is only which lane the filters name. The bind
evaluator runs for real; only the store is faked.

A third case pins that the hold filter keeps its `userPaused` exclusion,
so a filter matching every column would not satisfy the other two.

**Related coverage checked before writing a new file:**
`agent-heartbeat-worktree-renamed-hold.test.ts` covers the requeue
**target** on a renamed board, not the dispatcher's **selection**
filters — different branch of the same subsystem, so a case added there
would have read as duplicate coverage of the wrong thing.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **20 passed** across
the routing-policy and new dispatch suites · core `tsc` clean · `pnpm
lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Improved agent task selection when workflow lane names or IDs have
been renamed.
- Agents now correctly resume assigned in-progress or queued tasks
across customized workflows.
- Prevented agents from selecting tasks paused by users, including on
boards with renamed lanes.
- Updated task pausing behavior to correctly reflect lifecycle stages
beyond default lane names.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:09:36 -07:00
gsxdsm
76d77da4d1 fleet: taskSorting + TaskReviewTab 8 → 0 — and Board was faking a column id to force done-sorting (#2744)
Two app-side clusters, 8 → **0**, plus a caller-side hack retired.

## What was broken

**`TaskReviewTab.tsx`** — three of its four questions were `task.column
=== "in-review"`, driving the **Create-PR button**, the **"frozen on
entry to review"** auto-merge hint, and **PR-feedback addressing**. On a
renamed review lane all three took their non-review branch: the button
was absent, and the hint claimed the effective auto-merge value was
*not* frozen when it was.

**`taskSorting.ts`** — `isReviewColumn` decides whether merging cards
float to the top of a lane. Keyed on the id it silently stopped doing
that on any renamed review lane, so the operator loses the "what is
merging right now" ordering with nothing failing.

Both follow the shape this code already established: **caller supplies
the trait, default to the legacy id**. `columnFlags` on the review tab
is optional and wired from `TaskDetailModal`, which already resolved it
for `canEdit` and the actions menu.

## A synthetic column id, retired

`Board.tsx` forced done-sorting by passing the **literal `"done"`** as
the column argument for any complete-flagged lane:

```ts
grouped[column.id] = isWorkflowDoneLikeColumn
  ? sortTasksForDisplayColumn(grouped[column.id] ?? [], "done", doneSortMode)
  : sortTasksForDisplayColumn(grouped[column.id] ?? [], column.id as ColumnType, ...);
```

A synthetic id standing in for a trait — so a custom complete lane
sorted correctly only because its caller **lied about its name**. Both
call sites now pass the real column id and state the trait. (Board's own
census count stays at 2: those two literals were the synthetic ids and
are gone; the 2 remaining are different sites.)

## Revert proof

| reverted | failure |
|---|---|
| `task.column === "in-review"` on the Create-PR guard | `Unable to find
an element by: [data-testid="task-review-create-pr"]` |
| same, on the auto-merge hint | `expected 'Effective: Auto-merge off'
to contain 'frozen on entry to review'` |

A third case pins that the widened test does not treat *every* column as
review.

**None of the 45 existing `TaskReviewTab` cases could have caught this**
— `columnFlags` is optional and they all omit it, so they assert the
legacy fallback. That is the same blind spot as the reconciler's 33 in
#2737, and it keeps recurring: an optional-flags seam means the existing
suite stays green through the conversion *and* through a broken one.

## A process failure worth recording

**I lost this conversion once and had to redo it.** I overwrote four
files with their `origin/main` versions to check whether a failing test
was pre-existing, then "restored" with `git checkout HEAD -- <dir>`.
HEAD was still `origin/main` because I had not committed, so that
**discarded the work**.

Same class as the shared-stash incident two PRs back: an implicit or
positional restore reference. The fix is ordering, not care — **commit
before any baseline comparison**, so `git checkout HEAD -- <file>`
restores my work rather than main's. This PR's commit was created before
the comparison for exactly that reason, and the note is in the commit
message so the next person hits it there too.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **232 passed** across
TaskReviewTab / taskSorting / Board suites · dashboard `tsc -p
tsconfig.app.json` clean · `pnpm lint` clean · census `--strict` exits
0.

The 1 `board-mobile` failure is **pre-existing** — verified by swapping
in clean `origin/main` copies of all four files and reproducing it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:06:04 -07:00
gsxdsm
54d1a29621 fix(dashboard): the github-tracking-state classifier seam was never wired — a renamed terminal column still never closed its issue (#2754)
Not a conversion. A **conversion that was never connected to
production**, found by a parity sweep rather than by the census.

## How it surfaced

An AST pass over all **43** github/gitlab-named files, paired by name
with counts compared:

```
github-tracking-comments.ts=9  vs  gitlab-tracking-comments.ts=4   ASYMMETRIC
github-tracking-state.ts=2     vs  gitlab-tracking-state.ts=0      ASYMMETRIC
github-issue-comment.ts=1      vs  gitlab-issue-comment.ts=1
...8 more pairs, all symmetric at 0
```

The first is the pair #2715 fixes. The second pointed here — and the
asymmetry turned out not to be the interesting part.

## The defect

`decideIssueAction` has accepted an injectable `classify` since U12/R2,
and that file's own header states the bug the seam fixed:

> "A user-authored workflow whose terminal column is called something
else never closed its linked GitHub issue, and a custom archive column
never mapped to `not_planned`."

Its **only production caller** passed no classifier:

```ts
const decision = decideIssueAction(event.from, event.to);
```

So every real move fell through to `legacyColumnLifecycleClass`, and
**the documented bug was still live**. The seam was reachable from unit
tests only — which is why all 68 cases in that file were green while the
behaviour they document did not work.

Same shape as this branch's earlier finding on the tracking-comment
guard, where the guard returned *before* resolving. **Adding a seam and
wiring it are two changes; only the second one fixes anything.** Worth
watching for elsewhere in this program: a file can read as fully
converted, pass its suite, and still take the legacy path on every call.

## Ordering, inverted on purpose

`decideIssueAction` ran first, before the tracking-enabled check,
because comparing two strings is free. Resolving a workflow is not — so
the cheap property read now short-circuits and only tracked tasks
resolve, the ordering `github-tracking-comments.ts` and its GitLab twin
already settled on. Untracked tasks returned without acting before and
still do.

The two remaining literals **are** `legacyColumnLifecycleClass`, that
seam's named default, now marked `DELIBERATE-LITERAL` — and marked only
in the same commit as the wiring. While the default was the live path on
every move, exempting it would have hidden the real defect behind a
marker.

## Revert proof — it detects an *unwired* seam, not a missing one

Dropping the resolved classifier while **leaving the seam intact** fails
both new cases with 0 `setIssueState` calls. That is the whole point:
the new cases drive the **service**, not the pure decision function, so
they fail for exactly the reason the 68 existing cases could not. Those
pass either way.

## Two self-inflicted errors, recorded because both are recurrences

- **I hand-edited the baseline with python and wrote a raw NUL byte into
the JSON**, breaking the census parse. The `deliberateByFile` keys use a
real `\0` separator and must be written through `JSON.stringify`, never
string interpolation.
- **I then restored the baseline from a newer `origin/main` than my
branch point**, which made `--strict` report a `scheduler.ts: 12 → 26`
rise that was pure version mixing. Rebase first, then edit. (The real
`scheduler.ts` baseline staleness is already owned by #2712 — I checked
before assuming it was mine.)

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **71 passed** in the
tracking-state suite · dashboard `tsc` clean · `pnpm lint` clean ·
census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 08:05:50 -07:00
gsxdsm
28c069ac0c fix(dashboard): TaskDetailModal — one root cause that produced six separate review findings (#2761)
## The pattern

`workflowMoveMetadata` outlives a task switch, so while the modal is
open its flags describe the **previous** task for a render. This
component gates editability, the execution-mode replan decision, the
intake affordance, the actions menu and the review tab on them.

**The guard already existed** at line 961:

```ts
const detailFlagsAreForThisTask = workflowMoveMetadata?.taskId === task.id;
const detailColumnFlags = detailFlagsAreForThisTask ? workflowMoveMetadata?.currentColumnFlags : undefined;
```

Five call sites read **around** it. Each was found separately — #2744
(review tab), #2696 (`handleDelete` deps), and the four converted here:

| line | consumer |
|---|---|
| 1827 | `canEdit` |
| 2147 | execution-mode replan on save |
| 2285 | execution-mode replan on mode change |
| 3156 | `isIntakeColumn` |
| 3782 | actions-menu model |

**Six review rounds for one root cause.** Converting them together
retires the class instead of paying a round per site — the same
arithmetic as the review-lane family, which took #2730 and #2750 to end
rather than eight per-file patches.

Only the definition line still reads the raw value, which is the point
of it.

## Not covered by a new test — stated rather than papered over

is internal state populated by a fetch, not a prop, so the stale-flags
scenario needs the workflow-metadata request mocked **plus** a task
switch mid-render. That is a real test worth writing. It is not a line I
can add honestly in passing, and I have shipped four tests today that
passed with the bug fully in place — I would rather flag the gap than
repeat that.

This file's suite also carries **6 pre-existing failures**, confirmed
identical on clean by stashing and re-running, which would muddy the
signal from a new case.

## What is verified

- the narrowing is mechanical and **total** — one remaining raw read,
the definition;
- pre-existing failure count **unchanged at 6/45** with and without this
change;
- 106 passed + 5 skipped across both `TaskDetailModal` suites;
- dashboard `tsc` clean (`tsconfig.app.json`), lint clean (0 errors),
gate green (487 + 158 + 10 + 71).

No census movement — this changes which value is read, not whether a
column id is compared.
2026-07-30 07:59:25 -07:00
gsxdsm
58f1e41aa0 test(dashboard): keep the renamed-board tracking suite — its source changes landed as #2715 and #2737 (#2714)
**Claim announced on #2706 before starting.**
`github-tracking-comments.ts` + `github-tracking-reconciler.ts` — one
coherent subsystem, 9 guards each.

**18 → 7 by census, 18 → 0 behaviourally.** The gap is explained at the
bottom and it is not hand-waving.

## Both halves failed quietly, in the way that suppresses its own
evidence

| surface | what a renamed board got |
|---|---|
| the comment poster | returned early for **every** move — the tracked
issue silently stopped receiving both its "in progress" and its "done"
comment. The operator sees a linked GitHub issue that never updates. |
| the reconciler (3 scan passes) | matched **zero** tasks, so completed
work's issues were never closed — and the pass reported a clean
`scanned: 0`. |

That second one is the shape worth internalising: **the number that
would have revealed the problem is the number the bug suppresses.** No
error, no warning, a green sweep.

## The comment poster needed a derivation, not a swap

`event.to` was **both** compared against the two literals **and** passed
into `formatTrackingComment` as its `transition` argument (typed
`"in-progress" | "done"`). One value carrying two meanings: a lane id
and a comment kind.

Eight independent swaps would have had to keep agreeing with each other
forever — and a ninth site (the template's own `transition === "done"`)
is *not* a column at all, so a mechanical sweep would have converted it
wrongly. Resolving the lanes once and deriving the kind separates the
two meanings permanently. Log details still print the real column, so
the operator reads their own board's name.

## The reconciler

Per task through **one shared IR cache per scan**, resolved into a `Set`
of terminal ids rather than an async predicate inside `.filter(...)` —
`Array.filter` ignores promises, so an async predicate there silently
keeps **every** row. That is a trap worth naming for other fleet workers
converting list filters.

The archived-vs-complete distinction keeps its own resolver rather than
reusing the terminal pair: it decides GitHub's `state_reason`, and
closing a finished issue as `not_planned` is operator-visible and wrong
— as is the reverse.

## Revert proof

**5 of 8 new cases redden.** 201/201 across the nine `github-tracking`
suites (193 were already there and still pass).

## Why the census says 7 and not 0

The reconciler goes **9 → 0**. The comment poster still reports **7**,
and every one of those is `transition === "in-progress" | "done"` — the
derived comment **kind**, not a column. There is no lane comparison left
in the file.

That is exactly the vocabulary-collision class **#2692** is fixing (it
already lists five misclassified receivers: an SSE event type, a
cache-key mode, an evidence kind, a telemetry event kind, an agent
state). **`transition` is a sixth and I have reported it there.** Until
that lands the census counts them, so I am reporting both numbers rather
than the flattering one.

I deliberately did **not** mark them `DELIBERATE-LITERAL` to move the
count: that marker means "a lifecycle literal reviewed and kept", and
these are not lifecycle literals at all. Using it as a census-silencer
would put a wrong reason in the code to make a number look better.

## Verification

`pnpm test:gate` **487 / 71** · **201/201** github-tracking suites ·
`tsc -p packages/dashboard` clean · `pnpm lint` clean · census
`--strict` exit 0 (it also tightened three entries other workers' merges
left stale — the #2679 auto-tighten working).

No changeset: `@fusion/dashboard` is private and this is internal
behaviour on renamed boards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:53:45 -07:00
gsxdsm
86639f2ce4 fleet: planning drain + archive writers 12 → 4 — one stale row starves planning, and a finaliser that wrote an undeclared column (#2742)
**Claimed on #2733 before starting.** `in-process-runtime.ts` +
`task-artifacts-ops.ts` — **12 → 4**.

## 1. The planning drain: one stale row stops planning for the whole
project

FN-8470's own note on this code says it: **one orphan earlier in
created_at FIFO prevented every later planning continuation from
dispatching.** So on a renamed board the literal terminal pair did not
mis-handle one card — an archived or completed card's stale work item
read as live, stayed in the due set, and **starved the drain behind
it**.

The two classifiers take an **optional** terminal set, which is this
file's own injection idiom (the specification-complete reaction already
takes a `resolveIr` dependency so the pure passes are testable without
constructing a runtime that would attach to the real project registry).

**Optional is load-bearing:** a *required* parameter would have compiled
at every existing caller and then answered "not terminal" for
everything. That is the silent direction, and both halves are asserted
in the test.

## 2. `moveToDoneImpl` writes `task.column` directly

This is the store's own finaliser, not a `moveTask` caller — so its
literal is **not** caught by `moveTask`'s unknown-column validation the
way every converted call site in this program is. It silently persisted
`done` on a board that does not declare it, and then emitted `to:
"done"` to every listener.

**This is one of the few sites where a literal writes bad state rather
than merely failing to act.** A workflow declaring no complete lane now
throws instead of inventing one — #2733's rule: a missing field on a
resolved struct *is* an answer, and `?? legacy` discards it.

## 3. The unarchive destination — three decisions in four lines, all
literal

| pre-archive column | lands in |
|---|---|
| unusable / archived | the **complete** lane |
| the **wip** or **review** lane | the **hold** lane (its worktree and
session are long gone) |
| anything else | back where it was |

The second is the expensive one: a card archived *from* the wip lane was
restored straight back *into* it **with no worktree**, and the scheduler
then counts it as a live holder **occupying a slot**. Made async — its
one production caller already is, and the sync alternative is the
PostgreSQL no-op documented in #2703.

## Also

- **The mission-error requeue** (guard *and* destination in one change):
an errored mission task stayed in the wip lane holding a slot, because
the guard never matched.
- **The planner-chat retention cutoff on archive** — the quiet direction
of this defect class: nothing breaks, data that should be deleted simply
accumulates, and the only symptom is storage growth nobody attributes to
a column name.

## The live defect is not where the census points

`reliability-metrics.ts`'s 6 guards are **pure historical readers** over
activity-log entries, and **the dashboard does not call them**. The live
path is `server.ts`'s `getTaskMovedCountsByDay({ toColumn: "in-review"
})` — a **SQL query filter**, the class the census counts separately.

So the operator's reliability panel reads zero on a renamed board
because of a *query* literal, and converting the six guards the census
reports **would change nothing an operator sees**. Converting historical
readers also risks reinterpreting past events under today's traits,
which is a different decision from converting a live guard — I am not
making it inside a vocabulary sweep.

Worth generalising for the fleet: **a file's census count and its live
exposure are different numbers.** This is the second file where the
reported guards are the inert copy and the real one is a query
(`executor.ts:5805` was the first).

## Verification

`pnpm test:gate` **10 / 158 / 487 / 71** · 31/31 continuation suites ·
8/8 archive PG suites · in-process-runtime PG suite green · 5 new cases,
**2 red on revert** · `tsc` clean in core and engine · `pnpm lint`
clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:53:28 -07:00
gsxdsm
c53d3aec38 fix(executor): re-land the no-wip-lane fix — #2757 merged a snapshot that predated it (#2760)
## Why this exists

#2757 merged as `9a2a033b9e`, but **the third of its three fixes is not
on main**:

```
$ git show origin/main:packages/engine/src/executor.ts | grep -c wipDeclared
0
```

The merge captured my branch *before* commit `68381db72f`, so the
no-wip-lane fix was dropped while the other two landed.
`executor-execution-policy-renamed-columns` → *"a workflow with no wip
column terminalizes visibly instead of claiming the card advanced"* is
still red on main, still returning `status: null, error: null`.

This is that commit, cherry-picked cleanly onto current main. No new
work — the review discussion is in #2757.

## What it fixes (recap)

`resolveResumeLanes` defaulted `wip: lifecycle?.wip ?? "in-progress"`,
collapsing two different states:

1. **the IR failed to resolve** — defaulting is right; the `catch` arm
wants exactly this
2. **the IR resolved and declares NO wip column** — defaulting *invents
a lane the workflow does not have*

`routeGraphFailureToExecutionResume` then admitted a card resting in
that workflow's **hold** lane with incomplete steps, rehomed it,
returned `true` — and the terminalize branch never ran. Resuming into a
workflow with no implementation lane *is* "claiming the card advanced"
when nothing did.

The fix adds `wipDeclared` (declared, as opposed to defaulted) and
declines the resume when it is false — the same fail-closed rule the
sibling branch already applies with `wipColumn !== undefined` before
calling a card "already advanced". That path failed closed; this one
failed open.

IR-unavailable deliberately keeps today's behaviour (`catch` reports
`wipDeclared: true`), so an infrastructure error does not start refusing
legitimate resumes.

## Verification on current main

| check | result |
|---|---|
| the three affected files | **23 passed** |
| `pnpm test:gate` | **726** |
| engine `tsc --noEmit`, `pnpm lint` | clean |
| remove the guard | 1 failed — the fail-closed case goes red again |
| always decline | 1 failed — a legitimate resume breaks |

Full-suite blast radius was measured on #2757 before it merged: 834
files / 10,845 tests, 5 failures, both files pre-existing
(`executor-prompt`'s pause-guard 3 and `executor-abort-provenance`'s 2,
byte-identical to baseline). Zero new failures.

## Note

Worth checking whether other PRs merged in that window lost their final
commits the same way — I only noticed because I re-verified main after
the merge rather than assuming a merged PR contains what the branch
held.
2026-07-30 07:44:35 -07:00
gsxdsm
b0b9d1b373 fleet: store.ts 12 → 11 + names the sync-dependency-loop class blocking ~10 sites across 3 clusters (#2709)
Claiming **`packages/core/src/store.ts`** (12). One conversion and a
triage — because **10 of the 12 share a single blocking shape** that is
worth naming once rather than rediscovering per file.

## Census before/after

| | before | after |
|---|---:|---:|
| `store.ts` column guards | **12** | **11** |

Baseline re-recorded; `--strict` exits 0.

## Converted: 1

**1386** — the in-review guard inside `withTaskLock(id, async () => …)`.
Already async, and `this` **is** the store, so
`resolveTaskLifecycleColumns(this, task.id)` resolves the review role
with `in-review` as the fallback. Import added; nothing else in the
method changes.

## The blocking class — 6 sites, and it is not specific to this file

**1772, 1791 ×2, 1874 ×2, 1916, 1917, 1933** all read **another task's**
column — a dependency's, a blocker's, an overlap candidate's — inside
**synchronous callbacks over a prefetched `taskById` map**:

```ts
const unresolvedDeps = (task.dependencies ?? []).filter((depId) => {
  const dep = taskById.get(depId);
  return dep && !dep.deletedAt && dep.column !== "done" && dep.column !== "archived";
});
```

This is not a substitution. Each dependency may belong to a **different
workflow**, so the role must be resolved *per dep* — N async resolutions
inside a sync `filter`, on a path that deliberately prefetches into a
map precisely to avoid per-item I/O.

Two honest options:

1. **Prefetch lifecycle columns alongside `taskById`** and pass a
resolved map into these predicates. Keeps them synchronous, one
resolution per distinct workflow rather than per dep. This is the one
I'd argue for.
2. Accept per-dep resolution and make the callbacks async — changes the
shape of dependency evaluation.

Both are design changes with real cost, so this is flagged rather than
guessed.

**The same shape appears in at least two other clusters I've worked**:
`TaskDetailModal`'s `overlapBlockerTask.column` (#2696) and
`register-task-workflow-routes`' dependency-summary pair (#2700), both
flagged for this exact reason. **Worth one decision covering all three**
rather than three separate judgement calls by three workers.

## Also flagged: 3

**1610** and **1739** — enclosing-scope async-ness and store access not
established at those points, so not guessed. **1933** belongs to the
sync-filter family above.

## Verification

`pnpm test:gate` **GREEN** (158 + 487 + 10 + 71) · `pnpm lint` clean ·
core `tsc` clean · `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Updated failed pre-merge review bypass validation to support custom
workflow boards.
* Tasks can now bypass the step when placed in the board’s configured
review lane.
* Improved error messages to identify the correct review column when
bypassing is not allowed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:36:00 -07:00
gsxdsm
ceca08b1c3 fleet: github-tracking-reconciler 9 → 0 — deciding the sync-filter class (prefetch a resolved map), and the reconciler closed NO issues on a renamed board (#2737)
`github-tracking-reconciler.ts` 9 → **0**, and the reference
implementation for the `.filter((task) => task.column === "<id>")` shape
I have been flagging across four files.

## I stopped waiting and decided it

I flagged this class in #2709, #2696, #2700 and #2715 as "needs one
decision" and left ~25 sites unconverted. That decision was mine to make
and I should have made it three PRs ago.

**Prefetch a resolved map, then filter synchronously.** The alternative
— async predicates — forces every caller into `for await` and turns a
list comprehension into a sequential walk. Prefetching keeps the filters
synchronous, puts the awaits in one bounded place, and lets the IR cache
do the job it was explicitly built for:

> "A self-healing pass over 400 cards spanning three workflows must read
three IRs, not 400."

The cache is **instance-scoped and shared across all four passes**, so
each distinct workflow's IR is read once for the whole run rather than
once per pass. `resolveLifecycleColumns` is pure and *not* memoized by
that cache, so this still costs one cheap struct build per task — fine
in a background reconcile, and stated rather than hidden.

No new abstraction: `resolveTaskLifecycleColumns` already takes a
caller-owned cache. The only new code is a local map builder and two
named predicates.

## What it cost before

On a board with renamed terminal lanes, **every filter here matched
nothing**. The reconciler closed **no** GitHub issues and reported
`scanned: 0` — a clean-looking pass that did nothing.

## Why this is not the split brain #2724 documents — checked, not
assumed

#2724 proves the archived gate in `packages/core` is enforced in three
encodings, so converting one alone diverges them. I checked whether that
applies here before converting:

- This file contains **zero SQL** — measured: no drizzle, no `sql`
template, no `eq`/`ne`.
- It calls `listTasks({ includeArchived: true })`, so the SQL half has
already been told to include archived rows. The filter **selects among
rows it was handed** rather than deciding liveness a second time.

**Gate versus consumer** is the distinction, and a consumer can be
converted alone.

The fourth pass needed its own check because its list comes from
`listTasksForGithubTrackingReconcile`, which *is* SQL — but that impl
filters on `deletedAt IS NOT NULL` and `githubTracking IS NOT NULL`,
**never on the column**, so there is no SQL-side encoding of this
question to diverge from.

## Why the 33 existing tests stayed green through the conversion

Their fake store has **no workflow reader**, so
`resolveTaskLifecycleColumns` catches and returns `undefined` and every
case asserts the legacy fallback — exactly what it always asserted.
**None of them could have caught this being wrong.** `workflowIr` is now
an opt-in on that fake, which is what makes the new cases real tests
rather than restatements.

| reverted | result |
|---|---|
| terminal filter back to the ids | "closes issues on a RENAMED complete
lane" fails, no `setIssueState` |
| same | renamed archived-heuristic case fails, no `setIssueState` |

## A reachability finding, recorded not acted on

In backend mode `reconcileDeletedAndArchived` returns only
**soft-deleted** rows — its own comment says the archived-tasks fallback
is a separate `AsyncArchiveLineage` subsystem, skipped there — and
`task.deletedAt` is tested *first* in the `stateReason` chain. So its
archived arm is **effectively unreachable today**. I converted it rather
than deleting it: it is the documented FN-5577 done-heuristic, and
whether that fallback should be wired here is a separate question from
what vocabulary it speaks.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **35 passed** across
the three reconciler suites · dashboard `tsc` clean · `pnpm lint` clean
· census `--strict` exits 0.

Remaining files in this class (`branch-group-ops.ts`, `store.ts`, and
the dependency pairs) can now follow this pattern instead of waiting —
with the gate-versus-consumer check applied to each, since
`branch-group-ops.ts` sits closer to the persistence layer than this one
does.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:14:54 -07:00
gsxdsm
1322a1bb11 docs(solutions): the optional-flags seam kept four green suites blind to their own conversion — and why I did not ship a ratchet for it (#2748)
Docs only. No code, no census movement.

## The finding, measured

Four consecutive files in this program had **fully green suites at
conversion time** that could not have detected the conversion — correct
or broken:

| file | pre-existing cases blind to the change |
| --- | --- |
| `github-tracking-reconciler.ts` | **33** (fake store had no workflow
reader) |
| `TaskReviewTab.tsx` | **45** (`columnFlags` omitted everywhere) |
| `plan-approval-hold-invariant` drain | **25** (`opts.lifecycle`
omitted everywhere) |
| `task-age-staleness.ts` | **12** (`context.lifecycle` omitted
everywhere) |

The cause is structural. Every conversion here uses the same seam — the
caller passes resolved flags, the helper falls back to the legacy id
when they are absent — and every pre-existing test omits the flags. So
the suite passes **before** the conversion, **after a correct one**, and
**after a wrong one**, as long as the fallback is intact. "The suite is
green" carries no information about the change.

I reported this observation four times in PR bodies. Restating it a
fifth time is worth less than writing it where the next worker will
actually find it.

## It also corrects the obvious test

The natural property is "hold the traits fixed, change the id, behaviour
is identical". That is only half the invariant. It does not catch:

```ts
// Not a fallback — an OVERRIDE. The id wins even when traits disagree.
return column === "in-review" || flags?.mergeBlocker === true;
```

Renaming `in-review` → `checking` leaves that correct, because the trait
arm answers. The defect appears in the **converse** direction — a column
that still *carries* a lifecycle name while its traits say otherwise,
which is what you get by repurposing a default column rather than
renaming one. That is the direction that found a live **"Merge & Close"
offered on a mid-implementation card** in #2718.

## Why this is not a ratchet — a negative result, recorded

I tried to automate it, and I am shipping the reason it failed rather
than a guard I do not trust.

The **consumer scan is sound**: AST-based, 31 files, 66 role-helper call
sites. The **coverage half is not**. The renamed ids this program uses —
`building`, `checking`, `converted`, `published`, `backlog` — are
ordinary English words that appear in unrelated test prose, and a test
merely *importing* the module under test does not prove it exercises the
role path. My scan reported `TaskCard.tsx` as covered by
`Column.test.tsx` on a **filename coincidence**.

A guard built on that reports coverage that does not exist, which is
worse than no guard, so it is not shipped.

A sound alternative — pin the consumer set and make each new file
declare its status — was also rejected: a 31-entry status inventory
would conflict with every concurrent fleet PR that adds coverage. That
is the same churn already removed from the census baseline by dropping
its derived aggregates.

The attempt is written down so the next person does not repeat it from
scratch, and the requirement lives as a review criterion until someone
finds a sound signal.

## What it asks for

1. **A flags-supplying case** — if every case omits the new parameter,
the conversion is untested in both directions.
2. **Both directions where both are reachable** — renamed lane, and
repurposed column.
3. **A non-vacuous companion** — assert what the widened predicate must
still *exclude*, or a predicate matching every column satisfies your new
cases. (Both `TaskReviewTab` and the dispatch filters needed this.)
4. **Run the revert and record the failure text.** Twice in this program
a new case passed with the change reverted: once because the branch was
gated behind an unwired handler (`refine` needs `onOpenRefine`), once
because the hook was dispatched by trait and the test IR did not declare
that trait, so it never ran at all.

Cross-linked both ways with the adjacent `store-fake-defects` entry,
with the distinction stated so the two are not confused: **there** a
fake is missing a method so a branch never runs and production looks
wrong; **here** the fake is complete and the test is correct, but a
parameter is absent so production takes its documented fallback.

## Verification

`pnpm lint` clean · census `--strict` exits 0 (unmoved — this PR changes
no code).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:11:22 -07:00
gsxdsm
3066948110 fleet: async-comments-attachments.ts NOT converted — the archived gate lives in three encodings (52 sites), pinned by a guard that fails on each (#2724)
Claimed the **async-comments-attachments** cluster (9) and did not
convert it. This PR is the evidence for why, as a guard rather than a
note — **no production change, census unmoved.**

## What the cluster actually is

Every other file in the backlog converts on its own: resolve the task's
lifecycle columns, compare against the role. `archived` doesn't, and not
by a matter of degree.

**Measured in `packages/core` — one rule, three encodings, 52 sites:**

| encoding | sites | files | what it decides |
|---|---:|---:|---|
| TypeScript `=== "archived"` | **37** | 23 | what code does with a row
it already has |
| Drizzle `ne(tasks.column, 'archived')` | **7** | 6 | which rows a
query returns |
| raw `` sql`…column != 'archived'` `` | **8** | 5 | same — and
invisible to both scans above |

Convert only the TypeScript half and a board whose archived lane is
renamed **splits**: `getLiveTaskColumn` correctly reports the task
archived (it resolved the role) while `readLiveTaskRows` still hands it
back as live. A document write is rejected by its gate while its parent
is listed as live — a state neither gate alone can produce today.

And nothing would catch it. **Every builtin workflow spells that column
`archived`**, so all three encodings agree by accident on every board we
ship.

## Why the SQL sides aren't just converted too

`ne(tasks.column, ...)` needs the resolved id as a **query-build
value**, so the IR must be resolved *before* the query — including
inside the `for update` document/artifact transactions, which today
receive a `db`/`tx` handle and **no store and no workflow reader**. One
raw site is a hand-written `SELECT` string, so its comparison isn't even
a Drizzle expression that could take a bound value without rewriting the
query.

The two real options are: thread a resolver into the persistence layer,
or **declare `archived` a non-renameable system column** and mark all 52
sites deliberate. Both are decisions with blast radius. Neither is a
fleet conversion, so I didn't pick one.

## I was wrong about the shape, and that's the strongest part

I wrote the third case as an assertion that **no** raw `sql` template
compares a column to `'archived'` — I assumed two encodings. **It failed
on the first run with five files.**

Nothing in the repo was counting them: they aren't comparisons
(invisible to the column census) and aren't `eq`/`ne` calls (invisible
to the Drizzle scan). A partial conversion doesn't have to miss one
encoding — it can miss two. That case is now an inventory, with a note
on why asserting absence was the wrong invariant: an absence assertion
has to be deleted by whoever adds the next raw template, and deleting a
red guard is how a class of sites stops being tracked.

## Injection proof — all three run

| injected | result |
|---|---|
| convert one TS comparison in `async-comments-attachments.ts` | fails:
**"TypeScript encoding changed"** |
| remove one Drizzle predicate in `async-lifecycle.ts` | fails:
**"Drizzle encoding changed"** |
| remove one raw template in `reads.ts` | fails: **"Raw-sql encoding
changed"** |

Each failure carries the split-brain explanation and the two real
options, so the next worker hits the reason rather than a bare count
mismatch.

## Two scan bugs I fixed in my own guard

- **It audited documentation.** The raw-template scan reported three
sites in `async-archive-lineage.ts` that were prose in a JSDoc block.
Now comment-stripped through the census's shared `stripComments` rather
than a second implementation.
- **It missed half the cluster.** Requiring a receiver (`x.column ===
"archived"`) misses the list paths that hold the value in a bare local
(`if (column === null || column === "archived") return []`) — **four of
the eight sites in the target file**, measured. The matcher now accepts
a bare identifier named `column`.

Counts here are per file, not per line: line numbers churn on unrelated
edits, and a ratchet that cries wolf gets deleted.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · new guard **2/2** ·
`pnpm lint` clean · core `tsc` clean · census `--strict` exits 0 with
the **baseline untouched** — this PR converts nothing and claims
nothing.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:08:28 -07:00
gsxdsm
7c408ef650 fleet: merge path 10 → 2 — a merged PR never advanced its task on a renamed board (#2733)
**Claimed on #2728 before starting.** The merge path:
`merge-queue-ops-2.ts` + `merger.ts` — **10 → 2**, both survivors
flagged with reasons.

## A merged PR never advanced its task on a renamed board

`applyPrMergedTransition` is what moves a card when GitHub reports a PR
merged. Every guard in it was a default-lineage literal, and they all
failed **in the same direction**:

| guard | renamed board |
|---|---|
| `column === "done"` → skip as already-done | never matched, so a
complete card was re-processed |
| `column !== "in-review"` → bail `wrong-column` | always matched, so a
card **sitting in review** bailed |

Net effect: **a PR merged on GitHub never advances its Fusion task.**
The operator sees a merged PR whose card sits in review forever — which
reads as a broken webhook, so it gets debugged in the wrong place
entirely. That is the most expensive property of this defect class: it
does not just fail, it misdirects.

One snapshot now covers the pre-check, the deliberate **re-read** (a
merge can land between checks), and the **move target**. The target is
asserted in the test alongside the guards, because converting guards
alone would admit the card and then move it to a column the board does
not declare.

## merger.ts

- **The orphan-stash liveness guard** classified every finished task as
unfinished on a renamed board, so orphaned stashes were never cleaned
up. Unioned with the legacy ids: too strict here leaves clutter, too
loose **discards a stash whose task is still running**, so
over-inclusion is the safe direction.
- **The worktree-conflict scan** filters by worktree *path* before
resolving lanes. The naive order — resolve, then filter — is exactly
what made the github-tracking reconciler scan proportional to task
history (#2714 review). Lesson transferred rather than re-learned.
- The deprecated `aiMergeTask` already-finalized guard.

## Two flagged, not converted

**`merge-queue-ops-2`'s sync enqueue guard** runs inside
`store.db.transactionImmediate`. A synchronous lane resolution reads
`getTaskWorkflowSelectionImpl`, which returns `undefined`
**unconditionally in PostgreSQL mode** — so a "conversion" there would
drop the census by one and behave exactly as the literal (the finding
from #2703). Converting it properly means making the path async or
pushing the trait read into SQL: store architecture, not a call site.
Left literal **with that note**, so the next worker does not turn it
into a false green.

`merger.ts`'s last comparison is the same class.

## Pre-existing red, reported not folded

**22 failures in
`packages/dashboard/src/__tests__/routes-github.test.ts`** — spec
revise/rebuild and approve/reject-plan, all asserting moves to
**`triage`, the column U11 deleted**. Verified by reverting my diff and
re-running: identical 22. Same stale-literal-in-a-test class as the two
assertions #2720 fixed, and it is 22 tests pinning a column that does
not exist — worth someone owning deliberately rather than as a rider
here.

## Verification

census **10 → 2** · `pnpm test:gate` **487 / 71** · 23/23 across three
merger suites · 4 new cases, **2 red on revert** · `tsc` clean in core
and engine · `pnpm lint` clean.

## Also examined and deliberately left alone

- **`live-agent-count.ts` (6 guards)** — every literal there is the
*documented degradation path* for a task shape that was not enriched,
and both production callers already enrich (`useExecutorStats`, `fn
project`). Converting them converts nothing; deleting them removes the
fallback that fixtures rely on. The invariant that matters is **caller
enrichment**, which is not a literal at all.
- **`task-merge.ts` (6 guards)** — `getTaskMergeBlocker` is a **pure**
function with no store; its callers inject `resolveTask`. Resolving
lanes needs a matching injected resolver, which is an interface change
across every caller. Also worth a decision first: its dependency check
accepts `in-review` as satisfied while the store's `blockedBy`
computation (#2720) does not — **two definitions of "dependency
satisfied" in one codebase**, and I am not settling that one silently
inside a vocabulary sweep.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:08:16 -07:00
gsxdsm
ab715cbd39 fleet: default-workflow-hooks.ts 7 → 0 — every duration display read ZERO on a renamed board (#2734)
Claiming `default-workflow-hooks.ts` (7 → **0**), verified free against
every open PR's diff first.

## Not a vocabulary tidy — three silent zeroes

This file's header names it for the default workflow, but the store runs
it on the flag-ON path for **every** workflow: the trait registry
resolves each hook by **trait id**, not by workflow.
`reopen-semantics-by-role.test.ts` already documents that exact hazard
for the reopen predicates. The **timing, completion and in-review hooks
had the same defect** and were not part of that conversion.

On a renamed board, with nothing thrown and nothing logged:

- **`applyTimingEffects`** accrues `cumulativeActiveMs` while a card
sits in the WIP lane. With the lane named, the exit test never fires —
so **no active time is ever accrued**, and `productivity-analytics.ts`,
`task-timing.ts` and every duration display read **zero**.
- **`applyCompletionTimingEffects`** never stamps
`executionCompletedAt`, so a finished card looks unfinished to anything
reading that field.
- **`applyInReviewEnterEffects`** returns early, leaving the recovery
counters set.

The file already had the idiom — `ctx.lifecycleColumns`,
`planningColumnsOf`, `liveWorkColumnsOf` with `LEGACY_` fallbacks — so
this adds no abstraction.

One deliberate detail: `applyTimingEffects` resolves the WIP lane **once
into a local** rather than reading it twice. The exit test and the
re-entry test have to agree about which column is WIP, or a rename makes
the accounting count an interval twice, or not at all.

## A test that would have lied to me

I wrote the new cases through `applyDefaultWorkflowMoveEffects` first,
and **all three failed on the DEFAULT lineage too**. The dispatcher
resolves hooks by trait, and neither test IR declares the `timing`
trait, so those hooks never ran at all.

That failure looks exactly like a conversion bug. Going through the
dispatcher would have been testing the trait registry's wiring rather
than this change — so the cases call the converted functions directly,
and the reason is recorded in the test.

## Revert proof — all three, each naming the renamed lineage

| reverted | failure |
|---|---|
| the `in-progress` literals | `renamed lineage accrued no active time:
expected undefined to be 300000` |
| the `done` literal | `renamed lineage did not stamp completion:
expected undefined to be '2026-07-30T00:00:00.000Z'` |
| the `in-review` literal | `renamed lineage kept its recovery counter:
expected 3 to be undefined` |

Every case runs on **both** lineages and the default one passes either
way — which is the point of running it.

## A finding I did not act on

**`evaluateMergeBlockerGuard` appears exactly once in the repo — its own
definition.** And the file header says it is "implemented as the
`evaluateDefaultWorkflowGuards` reader", which does not exist either.
The merge-blocker guard hook is **defined and never consulted**.

I converted it (trailing optional lifecycle param, matching
`DefaultWorkflowMoveContext`) but did not delete it: the header states
this file is a deliberate parallel of `store.ts`'s flag-off path so the
two can be parity-checked, which makes removing it a scope call for
whoever owns that convergence — not something to decide inside a
conversion.

## Verification

`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **22 passed** across
`default-workflow-hooks` + `reopen-semantics-by-role` · core `tsc` clean
· `pnpm lint` clean · census `--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:05:05 -07:00
gsxdsm
94d88f1d6f fix(census): the work order was sending fleet workers at non-columns (722 -> 714) (#2692)
Found while claiming `TaskDetailModal.tsx` — its census entry included
`session.agentState === "done"`, an **agent state, not a lane**.
Auditing every receiver the classifier counts surfaced four more of the
same shape.

## The misclassified receivers

| site | receiver | what it actually is |
|---|---|---|
| `register-chat-routes.ts` | `event.type === "done"` | an SSE event
type |
| `useTaskDiffStats.ts` | `mode === "done"` | a cache-key mode |
| `async-mission-store.ts` | `evidence.kind === "done"` | an evidence
kind |
| `telemetry-hub.ts` | `event.kind === "done"` | a telemetry event kind
|
| `TaskDetailModal.tsx` | `session.agentState === "done"` | an agent
state |

Each shares a **word** with a column id and nothing else. Converting one
asks the trait registry what lane an SSE event is in, which has no
answer — the same failure class as converting `role === "triage"`, which
this list already exists to prevent.

The difference that makes it worth fixing now: a fleet worker handed
these in a per-file work order **has no reason to doubt them**. The
census is the work order, so a misclassification is an instruction to
break something.

## What I did not exclude

`state` is deliberately kept. `state === "archived"` in `audit-ops.ts` /
`comments-ops.ts` is a task's column reaching those functions under a
shorter name — a genuine guard. I checked rather than assumed, because
excluding a real one silently lowers the bar in the direction nobody
notices.

## Census effect

```
column  722 -> 714
role      5 -> 14
```

Those 8 are **reclassified, not converted** — this PR changes no
production code. The baseline is re-recorded so `--strict` agrees.

## Verification

`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0. `pnpm lint` clean.

No changeset: instrument accuracy, no user-facing change.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:01:52 -07:00
gsxdsm
15f90706e6 fleet: reliability-metrics.ts 6 → 0 — historical log values, marked not converted (#2756)
Unclaimed file, no overlap with any open fleet PR — deliberately picked
to avoid adding conflicts to the queue.

## Census

| | before | after |
|---|---|---|
| backlog | 539 | **533** |
| reviewed (DELIBERATE-LITERAL) | 31 | 36 |
| this file | 6 | **0** |

`--strict` exit 0, baseline re-recorded in the same commit.

## Why these are marked, not converted

All six ids come from `metadataColumn(entry, "from"|"to")` — the columns
**recorded on a past move event** in the activity log, not a task's
current column.

There is no workflow to resolve them against. The event was written
under whatever the board looked like at the time, and **a column renamed
since leaves every older entry carrying the old id forever.** Converting
them to a trait read would ask *"what role does the column named X play
today?"* about a record written months ago, possibly under a different
workflow — a different question with a different answer.

The failure mode matters: a trait-converted reader on a renamed board
would **zero the series** rather than fix it, silently dropping history
out of `tasksEnteredInReviewPerDay`, `tasksBouncedToInProgressPerDay`,
and `inReviewDurationMetrics`. That is worse than the literal, which at
least keeps matching the data that exists.

**The real fix for renamed boards is at the WRITER** — emit a role
alongside the id when the move event is recorded — not at this reader.
Noted at the site so whoever does that work finds it.

## A rule this generalises to

**Any reader of activity-log or run-audit metadata is a mark, not a
convert.** The census cannot distinguish `task.column === "in-review"`
(a live question, convert it) from `metadataColumn(entry, "to") ===
"in-review"` (a historical record, match it as recorded) — both are just
literals to the AST. Other fleet workers hitting log/audit readers
should expect the same call.

## Placement trap, third occurrence

My first pass marked the `const from`/`const to` declarations and moved
the count by **1 of 6** — the census excuses the construct a marker is
attached to, and the guards live in **sibling `if` statements**. Moved
the markers to the enclosing functions.

This has now caught #2645's author, me on `TaskContextMenu`, and me
again here. **Verify a marker by the count moving, not by the comment
existing** — and until every worker does, a batch reporting "N → 0" can
be off by most of N.

## Verification

Dashboard typecheck clean · reliability suites green (11 passed) ·
`--strict` exit 0 · no behavior change (comments only).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 07:01:02 -07:00