Commit Graph

11336 Commits

Author SHA1 Message Date
gsxdsm
f8fb9b1473 test(engine): pin the unmet-dependency rebound on a renamed board (11th resolver from the coverage map) (#3176)
Eleventh resolver from the verified coverage map on #3115.

`unmetDepReviewColumns` was uncovered: the existing FN-6778/FN-6779 case
uses `in-review`, where the literal is correct, so blinding the resolver
left the file green.

## What the literal costs

The sweep selects **no card**. A review card whose dependency is still
unmet is never rebounded — it sits in review, **eligible for merge,
ahead of the work it depends on**. That is precisely the ordering
violation this sweep exists to prevent, and it fails silently: no error,
no audit event, nothing to notice.

## Measured

3 pass; blinding `unmetDepReviewColumns` fails exactly the new case.

## Map status

**11 of 26 resolvers pinned** across 10 merged PRs. The remaining 15
need real harness work — I threw away two probes earlier today that
passed while proving nothing (`reconcileInReviewBranchRebind` never
entered its loop; `recoverAgentsRunningOnInactiveTasks` stayed green
under both blindings), and recorded them on #3164 rather than shipping
green decoration.

## Verification

`in-review-unmet-dependency-reconcile` **3 passed** · `pnpm test:gate`
full pass · lint — green.
2026-07-31 07:59:02 -07:00
gsxdsm
2a57820dd2 chore(gate): normalize the last future-dated stamp in task-update.ts (tightens the allowance 1 → 0) (#3168)
**Main is red on the FNXC gate again** — third occurrence of this class
today, third different file.

```
packages/core/src/task-store/task-update.ts: 2 future-dated FNXC stamp(s), baseline allows 1
```

Stamps dated **2026-08-01** while UTC is **2026-07-31-14:24**. Every
open PR inherits the failure; #3164 merged carrying it.

## The fix

Date only, to today. Clock times preserved exactly — they were real
times on the wrong day — and no comment text touched, so the record
reads identically, just in order:

```
-FNXC:StateMachine 2026-08-01-10:20 (PR #2793's finding — the INNER half, merged with #2821):
+FNXC:StateMachine 2026-07-31-10:20 (PR #2793's finding — the INNER half, merged with #2821):
```

Baseline **tightened** as a side effect (`1 → 0`): one future stamp was
grandfathered, normalizing the file cleared it too, and the gate refuses
a stale allowance on the way down. Re-recorded in the same commit.

## The recurrence is the point, not this fix

Three separate files have tripped this in one day — `scheduler.ts`, the
scheduler PG test, and now `task-update.ts` — plus the midnight-rollover
variant this morning that reddened everyone's baseline.

**Stamps are written from a local clock and validated against UTC.** A
worker behind UTC writes what is genuinely "today" for them and produces
a future stamp the moment UTC has already rolled. Nothing in the local
loop catches it: `pnpm lint` passes locally because the local date
agrees.

The durable fix is to generate the stamp from `date -u` rather than a
wall clock — one line in whatever produces these, and the class
disappears. I have patched the symptom three times today; someone should
take the cause. I have not done it myself because the stamps are
authored by hand across every worker's flow, so the change belongs
wherever that convention is documented, not in a file I happen to be
touching.

## Verification

`check-fnxc-future-dates` green (TZ=UTC CI=true) · `pnpm test:gate` 13 +
161 + 487 + 71 · lint · core typecheck clean · diff is date
substitutions only.
2026-07-31 07:56:09 -07:00
gsxdsm
10a0c5848f fix(executor): planner-evacuation lanes come from the emitter — executor leaves the inert list (16 → 12) (#3137)
`executor.ts` was the last file besides `triage.ts` and `scheduler.ts`
on `check-inert-sync-lanes`, holding **4 guards that read as converted
and behave as literals**. Neither cause turned out to be "needs an async
resolver".

## 1. Two of the four were in code with no caller

`isPlannerColumnFor` is a **private method with zero production
callers**. `tsc` reports it unused; the only things reaching it were two
tests casting through `executor as unknown as { … }`, which is exactly
what let it look alive. Its doc comment described the
planning-evacuation branch — but that branch calls
`isBackwardMoveOutOfPlanning` and never called this.

Deleted, along with the two tests whose subject it was. Converting
guards in unreachable code would have "fixed" behaviour that cannot run
and left two more sites to maintain; a test whose subject has no caller
pins nothing.

## 2. The other two no longer need to resolve anything

`isBackwardMoveOutOfPlanning` resolved its own lanes via
`resolvePlannerLanes`, whose selection reader returns `undefined`
unconditionally under PostgreSQL — so it answered with the **default
board for every task**, and both its guards were inert.

Its comment justified the sync resolver by the synchronous `task:moved`
emitter. That was true and **is no longer binding**: the emitter now
resolves lanes once, asynchronously (`moves.ts` →
`resolveWorkflowIrForTask`), and hands them on the payload — which #3112
already reads in this same listener. Reading a parameter is as
synchronous as reading `from`, so nothing reorders and no listener
resolves.

`lanes` is **required, not optional**. An optional parameter that the
one production caller happens to pass is the seam-with-no-supplier shape
this program keeps finding; required means a future caller fails
typecheck instead of silently getting a default board. When the emitter
itself could not resolve, the legacy ids answer — exactly what
`resolvePlannerLanes` degraded to anyway.

## Measured

| | before | after |
|---|---|---|
| `check-inert-sync-lanes` | **16** guards, 3 files | **12** guards, 2
files |
| `executor.ts` on that list | 4 | **0 — off the list** |
| census | 18 | 18 (`--strict`: every file matches baseline exactly) |

**The census is deliberately unchanged.** This targets the inert
population, which the census cannot see by construction: those guards
already read as converted. That gap is the argument in #3082 — 12 guards
still behave as literals while the census shows them as done.

## The producer half, which I nearly shipped without

The predicate's own suite covers it thoroughly — and every case calls it
**directly**. Mutation testing exposed that this proves nothing about
the listener: replacing the listener's `lanes` argument with `undefined`
left `planning-evacuation` at **20/20 green**. That is the fifth failure
shape in this program's learnings verbatim — a converted consumer with
an unconverted producer passing every instrument.

So there is now a case driving the **real listener** on a board whose
planner lanes share no id with the legacy pair (`queued` holds,
`drafting` intakes), withdrawing a card to a non-lifecycle column — the
reported symptom (`todo -> Ideas`) in that board's vocabulary.

## Verification

- engine `tsc` — **0 errors**
- `executor-planner-lanes-resolved` — **12 passed**
- `executor-archive-releases-active-session` — **14 passed**; listener
passing `undefined` → **1 failed | 13 passed**
- `planning-evacuation` + `triage-planning-wake` + archive suite — **47
passed**
- `check-inert-flag-seams`, `check-fnxc-future-dates`, census `--strict`
— exit 0
- `eslint` on changed files — 0 errors

The predicate tests are also **stronger than before**, not merely
adapted: they now build lanes with `toTaskMoveLanes`, the same function
`moves.ts` uses for the payload. Previously they reached the predicate
through the store-backed sync reader, so renamed-lane assertions passed
in the harness while the real path could never see a renamed lane.

## Not done here

The inert baseline still reads 29 against a tree of 12 and the gate
advises re-recording. I left it: a stale allowance is a real hazard, but
re-recording is a one-line change that conflicts with every lane, and it
should land once rather than in each of our branches.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:30:43 -07:00
gsxdsm
b3d009edde fix: main is red on check-fnxc-future-dates — one stamp dated tomorrow (blocks every open PR) (#3166)
`check-fnxc-future-dates` runs in `pr-checks.yml`, so while `main` is
red **every open PR fails this check** regardless of what it touches.
Measured on a clean detached `origin/main`:

```
[check-fnxc-future-dates] FNXC stamp population changed:
  packages/core/src/task-store/task-update.ts: 2 future-dated FNXC stamp(s), baseline allows 1
    FNXC:StateMachine   2026-08-01  (dated after today)
    FNXC:WorkflowEvents 2026-08-01  (dated after today)
```

## One line, scoped by blame

Two stamps in the file are future-dated; only one is **new**:

| line | stamp | commit | action |
|---|---|---|---|
| 86 | `FNXC:StateMachine 2026-08-01-10:20` | `e5c9ea38709` (07-30) |
**baselined — left alone** |
| 964 | `FNXC:WorkflowEvents 2026-08-01-05:10` | `71f459c2a5d` (07-31) |
corrected → `2026-07-31-23:10` |

The baselined one is not what turned main red, and rewriting it would
register as a **drop** — which is exactly how I did collateral damage in
the #3124 cycle by rewriting two stamps I had not authored. Exact-match
replacement on the distinct new string; the older stamp is verified
still present afterwards.

## What is not changed

**The baseline file is untouched.** The fix is the stamp, not the
allowance — re-recording would clear the red while leaving tomorrow's
date in the tree, which is the false green this gate exists to prevent.

## Process note

I checked for an existing fix PR **before** writing this one. My #3143
was a duplicate of #3139 because I skipped that step on the last red,
and a red `main` is the single most likely thing for two lanes to notice
simultaneously.

## Verification

- `check-fnxc-future-dates` — **exit 1 on `origin/main`, exit 0 here**
- diff is one line; baseline file confirmed unmodified

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:30:21 -07:00
gsxdsm
27741d0e2f fix(core): an archived child kept blocking its parent's delete on a renamed board (#3162)
Second **LANE** site from the archived triage (#3154), same additive
shape as #3160.

## The bug

`liveLineageChildFilter` is the lineage-integrity gate (VAL-DATA-010)
behind `deleteTask` and `archiveTask`: a parent with **live** children
is refused with `TaskHasLineageChildrenError`.

It excluded children in the `archived` column **by id**. So on a board
that renames that lane, an archived child still counted as live and the
parent could not be deleted — with an error naming a child the operator
had **already filed away**.

## The fix is permissive, and that is the correct direction

The gate exists to protect **live** children; an archived child is not
one. Resolving makes fewer rows block, which is what the gate always
meant.

I am flagging this explicitly because *"a conversion makes a delete gate
stop firing"* deserves a second look. The second look is that it was
firing on rows it was never meant to protect.

## LANE, not STATE

The two STATE sites in this inventory are marked at their own sites
(#3157) and must never be resolved — one of them deletes directories.
This one asks about the board.

## Wired at every caller, not left as an optional seam

`findLiveLineageChildrenImpl` (has `store`) and both
`archive-lifecycle-2.ts` gates resolve and pass it.
`hasLiveLineageChildren` takes the same parameter so the two readers
**cannot disagree** about which children are live — the half-conversion
shape this program keeps finding, and the reason #3129's earlier attempt
was reverted for leaving a seam unsupplied.

## Parity gate satisfied

Same reason as #3160: the conversion is **additive** — it keeps the
literal as the fallback, so no encoding's literal count moves and a
caller supplying no set gets byte-identical SQL.

## Measured

- 4 new cases; parity test still **2/2**, inventories unmoved.
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control**, the **fail-soft** case, and the
parent/project-scope **negative** green.
- lineage / archive / soft-delete / archived suites — **6 files / 19
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — the literal remains the fallback arm, by design.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:18:54 -07:00
gsxdsm
4365a3b10b test(engine): pin the paused-scope-decay lane filter (plus two probes I threw away) (#3164)
Next entry from the verified coverage map on #3115.
`scopeDecayWipColumns` was uncovered — the existing case uses
`in-progress`, where the literal is correct, so blinding the resolver
left all 420 tests green.

## What the literal costs

A paused holder resting in a renamed wip lane is **never selected**. The
loop does not run, nothing is recorded, and its file scope decays with
nothing to rebound it — so its followers stay blocked behind a card that
is not coming back.

## The observable

The audit event. Reaching a no-action record proves the holder was
**selected by the lane filter**, which is the only thing this resolver
controls. Asserting on the rebound itself would have needed triple-proof
to succeed, dragging in state the resolver has nothing to do with.

**Measured:** 420 pass; blinding `scopeDecayWipColumns` fails exactly
this case.

## Two attempts thrown away first

Worth recording, because the remaining map entries are not uniform with
the ones already closed:

- **`reconcileInReviewBranchRebind`** — a git-free probe (workspace
task, rejected before any git runs) returned `{outcomes: [], repaired:
0}`. The loop never ran. I confirmed the `merge` trait does map to
`mergeOrchestration`, so the filter should have matched; something else
short-circuits and I could not establish what.
- **`recoverAgentsRunningOnInactiveTasks`** — the test passed, then
**both** resolvers stayed green when blinded. `agentLinkTerminalColumns`
never fires because the live card is caught by the wip∪review set first;
`agentParkedColumns` only feeds `evaluateParkedAgentTaskLink`, whose
result my fixture already forced true via a fresh run.

Both would have been green, plausible, and worthless. They were reverted
rather than adjusted until they passed — which is the failure this whole
effort exists to remove, and the one I committed myself in #3078.

## Map status

Closed: 9 resolvers across 7 PRs. **~18 remain.** The easy ones are
done; what is left needs real harness work — an `execAsync` git fixture,
and understanding how `evaluateParkedAgentTaskLink` weighs run-freshness
against lane. Budget for that rather than expecting the pattern that
closed the first nine.

## Verification

`self-healing.test.ts` **420 passed** · `pnpm test:gate` 161 + 13 + 487
+ 71 · lint — green.
2026-07-31 07:15:56 -07:00
gsxdsm
5d5a3ddd60 fix(core): live search excluded the archived id, not the board's archive lane (#3160)
The first **LANE** site from the archived triage (#3154), converted —
and it establishes that this family does **not** need the single 52-site
commit the parity gate's message implies.

## The bug

`liveSearchPredicate` builds the "not archived" half of every task
search. Keyed on the literal, a card filed away on a renamed board
stayed in **every live search result** — including the CREATE-time
near-duplicate check, which calls `searchTasks()`.

So creating a task could be rejected as a duplicate of one the operator
had **already archived**, with nothing on screen explaining why.

## Why this site and not its neighbours

The triage classifies all eight Drizzle `archived` sites. Two are
**STATE** markers that must never be resolved —
`cleanupArchivedTasksImpl` deletes directories,
`listSoftDeletedColumnDriftCandidates` would "repair" already-correct
rows. Both are marked at their sites in #3157.

This one asks about the board, so it must resolve. That distinction is
the entire product of the triage, and it is why this is a one-site PR
rather than a sweep.

## The parity gate is satisfied — and the reason generalises

That gate exists because converting one encoding while the others
compare the raw string makes them **disagree**.

This conversion is **additive**: it adds a resolved path and keeps the
literal as the documented fallback. The SQL encoding's literal count
does not move, and no encoding shifts relative to another. A caller
supplying no set gets **byte-identical SQL**.

So the family can be converted **incrementally** — one site at a time,
each keeping its fallback — rather than in one coordinated 52-site
commit. The gate's rule is about not letting the encodings *diverge*,
not about batching.

That was the last thing making this cluster look unapproachable, and it
turns out not to be true.

## Threading was two layers, and the root already had the answer

`reads.ts` resolves archived lanes for its cold-storage decision a few
hundred lines above; the same call now serves both search paths. My
earlier scoping note (#3147) guessed "one parameter each" — it is the
builder plus its two entry points, with the resolution already present
at the root.

## Measured

- 4 new cases; parity test still **2/2** (inventories unmoved — that is
the point).
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control** and both negatives green.
- The `includeArchived: true` negative earns its place: resolving lanes
must not start excluding them from a search that explicitly asked for
archived rows.
- The predicate walker needed **cycle detection** — Drizzle's SQL graph
is circular (column → table → columns) and my first version blew the
stack on the first assertion.
- core search / archive / cold-storage / reads suites — **5 files / 15
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — the literal remains as the fallback arm, by design. A
census drop here would mean the fallback had been removed, which is what
the parity gate is protecting against.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:13:01 -07:00
gsxdsm
71f459c2a5 fix(events): the last two live task:moved emitters carry the resolved lanes (#3135)
## The last two live emitters

#3109 attached lanes at `moves.ts`. #3120 attached them at the archive
and completion emits. Two live emitters were still sending `lanes:
undefined`:

```
task-update.ts:962        todo -> triage
update-task-deps.ts:502   emits to: task.column — the row's ACTUAL lane
```

A listener reads absence as "unknown" and falls back to
`resolveTaskParkedColumnsSync`, which returns the **default board** for
every task under PostgreSQL. So these two paths kept the pre-#3109
behaviour while the listener code reads as resolved at every site.

`update-task-deps.ts` is the sharper of the two: it emits `to:
task.column`, the row's real lane, so on a renamed board it sends e.g.
`"shipped"` to a listener comparing against `"done"`. The emitter had
already resolved the board and threw the answer away. Nothing errors;
the branch stops firing.

## Emitter coverage after this

```
LANES    moves.ts:1450               (#3109)
LANES    archive-lifecycle-2.ts:385  (#3120)
LANES    task-artifacts-ops.ts:578   (#3120)
LANES    task-update.ts:970          (this PR)
LANES    update-task-deps.ts:504     (this PR)
MISSING  lifecycle-ops.ts:668        deliberate — see below
MISSING  lifecycle-ops.ts:715        deliberate — see below
```

**Every emitter that can execute under the shipped backend now carries
lanes.**

## Flagged — the two I did NOT convert

`lifecycle-ops.ts:668` and `:715` stay lane-less on purpose. Both sit on
the polling-replica path that file already documents as
legacy-SQLite-only — it reaches `store.db`, which throws under
PostgreSQL — and the same note argues against spending a signature
change on dead code. I agreed rather than overrode it. If that path is
ever revived they must be attached, because absence resolves to the
default board rather than to nothing.

## Supersedes my own earlier PR

This replaces **#3119**, which carried four emitters. #3120 landed two
of them first, so I rebuilt against current `main` with only the
remainder rather than resolving a conflict into a half-redundant diff.
#3119 is closed with nothing lost.

## Census before / after

```
before:  COLUMN guards (the backlog):   17
after:   COLUMN guards (the backlog):   17
```

Unchanged — this converts no guards. It makes the resolved answer
*reach* guards that were already converted, which is the half that was
missing.

## Verification

Full `@fusion/core` suite **462 files / 4906 passed, 0 failed** ·
`test:gate` exit 0 · typecheck exit 0 · lifecycle-column census exit 0 ·
`pnpm lint` clean.

## Still open

`main` is **red on `check:inert-sync-lanes`** and nothing in CI runs it
— **#3127** fixes both halves. **#3122** restores 13 guards laundered
through `mergeParkedColumns`. **#3131** corrects a 44% under-report in
`--triage`.
2026-07-31 07:10:09 -07:00
gsxdsm
b9b7d14804 fix(core): a type that taught the wrong invariant — staleness signal column narrowed to legacy ids (#3159)
**Type-only. No runtime behaviour changes**, and I would rather say that
than let a green suite imply otherwise.

## The type described a guard that no longer exists

`TaskAgeStalenessSignal.column` was typed `"in-progress" | "in-review"`
and filled through a cast carrying this justification:

```ts
// The guard above proves `column` is one of these two legacy ids ... (#1403)
const activeColumn = task.column as "in-progress" | "in-review";
```

True when written. The guard now reads:

```ts
const wipColumn    = context.lifecycle?.wip    ?? "in-progress";
const reviewColumn = context.lifecycle?.review ?? "in-review";
if (task.column !== wipColumn && task.column !== reviewColumn) return undefined;
```

So on a renamed board it proves the column is `building` or `checking` —
and the cast asserted the **opposite** of what the guard established.
The runtime was always fine; the real id passed straight through.

## The damage is in what the type taught

A consumer writing `signal.column === "building"` got a **compile
error** saying the comparison was impossible. The type actively
instructed callers that `=== "in-progress"` is exhaustive — the exact
guard shape this program spends its time removing.

This is the second instance of the shape today. The first was
`dashboard/src/server.ts`:

```ts
moveTask(taskId: string, column: "todo", options?: …): Promise<unknown>;
```

which made the type system **reject** a resolved target (#3158). Neither
was a constraint anyone chose — both were inferred from a single legacy
call site and then hardened into an assertion about live data.

**A type narrowed to legacy ids is a lint against fixing the code**, and
it is invisible to every gate this program has: the census counts
comparisons, the move-target ratchet counts arguments, and neither looks
at type positions.

## The test is a characterization, and says so

No runtime test can differentiate a type-level fix — **`tsc` is what
differentiates it**. The added case pins a value that was already
correct, so a future narrowing has something to break against besides a
compile error nobody sees until they hit it. I have labelled it in the
file rather than presenting it as a regression test.

## Verification

| | result |
|---|---|
| `tsc` — core, engine, dashboard (app + src) | **0 errors** each |
| `task-age-staleness` | **17 passed** |
| all three staleness suites | **28 passed** |
| census `--strict` | exit 0 |

All three consumers of `.column` only display or compare it
(`taskAgeStalenessCopy.ts`, `TaskDetailModal`, a `TaskCard` memo
comparison), so nothing downstream narrows on the widened type.

## Scope note

I scanned for this class and the raw pattern is noisy — 192 candidates,
almost all object-literal **values** (`status: "archived"`), agent
roles, and unrelated `type: "done"` stream events. This one and
`server.ts` are the two I could confirm as genuine type-position
narrowings on a *task column*. I have not filed an issue for the class
because I cannot yet separate it from the noise reliably; if a cheap
discriminator turns up, it is worth a ratchet like the move-target one.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:06:48 -07:00
gsxdsm
44a67df65d test(engine): pin both dependency-lease resolvers on a renamed board (one case covered only half the conversion) (#3138)
Next two entries from the verified coverage map on #3115.
`reconcileDependencyBlockingLeases` had **both** of its resolvers
uncovered — blinding either `leaseWipColumns` or `leaseHoldColumns` back
to its legacy id left all 825 self-healing tests green, because every
fixture in that block uses `in-progress` / `todo`, where the literals
happen to be correct.

## What the literals cost

The holder scan matches no card **and** the dependency scan matches no
card. A stale file-scope lease blocking a real dependency is never
rebounded, so the dependent stays `overlapBlockedBy` behind a holder
that is not coming back. That is the deadlock this sweep exists to break
— silently not broken, no error, no log.

## Two cases, because one did not cover both — measured, not assumed

My first case (holder in a renamed wip lane, dependency marked
`overlapBlockedBy`) pinned `leaseWipColumns`. I then blinded
`leaseHoldColumns` against it and **it stayed green**.

The reason is in the control flow: the `overlapBlockedBy === holder.id`
branch short-circuits and `break`s **before** the hold membership is
consulted. So that fixture can never reach the guard `leaseHoldColumns`
feeds.

The second case drops the marker, leaving an unmarked dependency resting
in a renamed hold lane, which falls through to the
overlapping-hold-dependency branch.

| blinded | result |
|---|---|
| `leaseWipColumns` | **1 failed** |
| `leaseHoldColumns` | **1 failed** |

Before the second case, that table read `1 failed` / `still green`.
Checking each resolver separately is the only reason I noticed — a
single "the suite fails when reverted" would have looked like proof and
covered half the conversion.

## Remaining

23 uncovered resolvers on the map. Next by risk:
`reclaimStaleActiveBranches` (deletes branches) and
`reconcileInReviewBranchRebind` (rebinds branches of live cards), both
needing a git-shelling harness.

## Verification

`self-healing.test.ts` **415 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved recovery of stalled workflow tasks when dependency-blocking
leases become stale.
* Added support for workflows using customized task status lanes,
including marked and unmarked overlap blockers.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 06:58:27 -07:00
gsxdsm
7cdd3f3e87 test(engine): pin the last two worktree-metadata resolvers — completes the sweep with 3 of 3 uncovered (#3148)
Completes `reconcileTaskWorktreeMetadata` — the sweep with the most
uncovered resolvers on the #3115 map (3 of 3). #3132 pinned the wip
half; these are **terminal** and **review**.

Both were uncovered: blinding either back to its legacy ids left all 825
self-healing tests green, because no fixture in that block used a
renamed lane.

| resolver | what the literal cost |
|---|---|
| terminal | a finished card is not this sweep's business. Keyed on the
ids the skip never fired, so finished cards were reconciled on every
pass |
| review | the other half of the **FN-5256** liveness guard. Keyed on
the id it went silent — and this sweep nulls
`worktree`/`branch`/`sessionFile` on a live row |

## Blinded separately, not as a pair

Each resolver was blinded on its own and measured on its own. **#3138 is
exactly why**: there, one case pinned `leaseWipColumns` and left
`leaseHoldColumns` green, because the control flow short-circuited
before the second guard was ever reached.

A single revert that fails proves *one* resolver, not the conversion.
That is the finer-grained version of the lesson from #3078, where a
whole suite passing proved nothing at all.

| blinded | result |
|---|---|
| `worktreeReconcileTerminalColumns` | **1 failed** |
| `worktreeReconcileReviewColumns` | **1 failed** |

## Map progress

Closed: `archiveStaleDoneTasks` ×2,
`reconcileOrphanedPendingStepResults`, `recoverDriftedAgentTaskLinks`,
`reconcileDependencyBlockingLeases` ×2, `reclaimStaleActiveBranches`,
`reconcileTaskWorktreeMetadata` ×3. **20 uncovered resolvers remain** of
the original 26.

Next: `reconcileInReviewBranchRebind` (rebinds branches of live cards),
then `recoverAgentsRunningOnInactiveTasks` ×2.

## CI note

This will show red on Lint until **#3145** merges — main carries FNXC
stamps dated 2026-08-01 through 08-06 while UTC now is 07-31, so every
open PR inherits it. #3145 fixes it; this PR touches none of those
files.

## Verification

`self-healing.test.ts` **416 passed** · `pnpm test:gate` 161 + 487 + 13
+ 71 · lint — green locally.
2026-07-31 06:58:11 -07:00
gsxdsm
1136474a63 test(engine): pin the archive skip in branch reclaim — the first uncovered resolver whose failure deletes a branch (#3144)
Next entry from the verified coverage map on #3115 — and the first one
whose failure mode is **irreversible**.

## The gap

`reclaimArchivedColumns` was uncovered: blinding it back to the id
`archived` leaves all 825 self-healing tests green, because no fixture
in this suite puts a card in a renamed archive lane.

## Why it matters more than the other 22

That guard **skips** archived cards — their branches belong to archive
cleanup, not to branch reclaim. Keyed on the id, a card filed in a
renamed archive lane fails the skip, and this sweep reaches:

```
git branch -D "fusion/<id>"
```

Every other uncovered resolver I have pinned so far causes a wrong
lifecycle decision — a card not requeued, a lease not released, a
diagnostic not surfaced. All of those are recoverable from the task row.
**A deleted branch is not.**

## Measured

415 pass. Blinding `reclaimArchivedColumns` fails exactly this case, and
the assertion that fails is the one checking `git branch -D` was never
called — so the failure *is* the branch being deleted, not a proxy for
it.

## Progress on the map

Closed so far: `archiveStaleDoneTasks` ×2 (#3115),
`reconcileOrphanedPendingStepResults` (#3090),
`recoverDriftedAgentTaskLinks` (#3102), `reconcileTaskWorktreeMetadata`
wip (#3132), `reconcileDependencyBlockingLeases` ×2 (#3138), and this
one. **22 uncovered resolvers remain.**

Every sweep probed so far has been uncovered, and one
(`reconcileDependencyBlockingLeases`) was only half-covered by its own
first test — the branch short-circuited before the second resolver was
ever consulted. That is why I now blind each resolver separately rather
than trusting a single revert.

Next: `reconcileInReviewBranchRebind` (rebinds branches of live cards)
and the two remaining `reconcileTaskWorktreeMetadata` resolvers
(terminal, review).

## Verification

`self-healing.test.ts` **415 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.
2026-07-31 06:55:15 -07:00
gsxdsm
db71b4fffb docs(core): scope the archived three-encoding decision — it is not 52-or-nothing (#3147)
The `archived` family is the largest unclaimed cluster (52 sites) and
its blocker is that **nobody has scoped it**. This scopes it. It
converts nothing.

## The two options both read as enormous because 52 sites are counted as
one lump

They are not one lump. The sites answer **two different questions**:

| question | renameable? |
|---|---|
| **LANE** — "is this row resting in the board's archive lane?" | yes —
must resolve |
| **STATE** — "did Fusion archive this row?" (the marker `archiveTask`
writes) | **no** |

`async-maintenance.ts` already draws that line and marks its own site
DELIBERATE-LITERAL:

> `'archived'` is the STATE marker here, not a lane. This sweep collects
rows Fusion itself archived or soft-deleted; a card merely sitting in a
workflow's archived-TRAIT lane is live work and must not be collected.
Widening to the resolved archived set would pull real cards into a
cleanup pass.

**Converting that site would be a bug, not progress.**
`async-archive-lineage.ts`'s soft-delete path is the same shape —
`column = 'archived', deleted_at IS NOT NULL` is the storage state it
has just written.

So the first question is a **triage**, not a conversion: which of the 52
are lane questions? Nobody has answered it, which is exactly why the
cost reads as unbounded.

## Measured: the SQL half, which the existing note calls the hard part

8 Drizzle sites across 7 files.

**Four already have `store` in scope** — they could take a resolved set
today with no signature change:

- `branch-group-ops.ts` — `clearNearDuplicateReferencesToImpl(store,
...)`
- `branch-and-pr-entities.ts` —
`findRecentTasksByContentFingerprintImpl(store, ...)` (2 sites)
- `task-mutation-ops.ts` — `cleanupArchivedTasksImpl(store)`

**Four need one parameter each**, the same optional-lane-set shape used
throughout this program:

- `async-lifecycle.ts` — `liveLineageChildFilter(parentId, projectId?)`
- `async-search.ts` — `liveSearchPredicate(includeArchived, projectId?)`
- `async-self-healing.ts` — `listSoftDeletedColumnDriftCandidates(db,
...)`
- `store.ts` — the revert-lookup conditions (already holds
`this.asyncLayer`)

That is not *"threading a resolver into the persistence layer"*. It is
four call sites that already have what they need, plus four
one-parameter widenings — **before** any triage removes the STATE sites
from the count entirely.

## What I did not do, and why

The triage itself: a per-site judgement about what each guard *means*.
That belongs to whoever owns this gate, not to a passing fleet lane —
and getting it wrong in the STATE direction pulls live cards into a
cleanup sweep, which is the one failure mode here that destroys work
rather than hiding an affordance.

What was cheap and missing was the **shape** of the problem.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean.
- `check-fnxc-future-dates` is red from `main`'s own #3128 stamps —
**#3139** fixes that; this branch inherits and does not add to it.

## Census

No movement.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 06:55:03 -07:00
Phil Larson
920d68e10f fix(dashboard): expose column roles to browser bundle (#3151)
## Summary
- export the browser-safe `@fusion/core/column-roles` subpath
- keep Vite/Vitest aliases ahead of broad `@fusion/core` aliases
- restore production dashboard builds after task undo classification
adopted shared column-role helpers

## Test plan
- `node scripts/check-no-node-only-core-imports-in-dashboard.mjs`
- `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest
run app/utils/__tests__/taskRevert.test.ts --pool=threads
--maxWorkers=1`
- `pnpm --filter @fusion/core typecheck`
- `pnpm --filter @fusion/dashboard typecheck`
- `CI=true pnpm check:changesets`
- `pnpm build`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dashboard build compatibility for browser-based environments.
* Improved reliability when importing column role functionality across
supported application components.

* **Refactor**
* Made column role utilities available through a dedicated browser-safe
entry point.

* **Chores**
* Updated development and test configurations to consistently resolve
the new entry point.
* Documented the browser-safe module classification and recorded the
release patch.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 06:46:19 -07:00
gsxdsm
110d6fd150 docs(core): the archived LANE-vs-STATE triage, done — 8 SQL sites classified with evidence (#3154)
#3147 scoped this cluster and said the first question is *"which of
these are LANE questions and which are STATE markers?"* — and that
nobody had answered it. **This answers it** for the Drizzle half, per
site, by reading what each query is for.

I claimed it because it has sat unclaimed for many rounds and `--claims`
reports `AVAILABLE: 0 files / 0 guards` — this is the only real work
left in the area. Nothing is converted here.

## LANE (6) — must resolve a renamed archive lane

| site | evidence |
|---|---|
| `store.ts` revert lookup | `ne(archived)` + `ne(done)` picking
**live** revert candidates |
| `branch-group-ops.ts:82` | near-duplicate marker cleanup over **live**
rows |
| `branch-and-pr-entities.ts:438` | content-fingerprint duplicate guard,
gated on `!includeArchived` |
| `branch-and-pr-entities.ts:470` | recent **sibling** lookup |
| `async-lifecycle.ts:68` | `liveLineageChildFilter` — the name is the
classification |
| `async-search.ts:82` | `liveSearchPredicate(includeArchived)` — same |

Four already hold `store` / `this.asyncLayer`. The two predicate
builders need one optional parameter each — the shape used throughout
this program.

## STATE (2) — converting these would be a **bug**

**`task-mutation-ops.ts:1072`** — `cleanupArchivedTasksImpl` selects
`eq(column, "archived")` and then `rm`s each row's files. Widening it to
the resolved archived set would feed cards **merely resting in a board's
archive lane** into a filesystem delete.

This is the most destructive site in the family, and it **looks
identical to the LANE sites at a glance** — same column, same operator,
same file neighbourhood. That is the whole argument for triaging before
converting.

**`async-self-healing.ts:61`** — soft-deleted rows whose column
*drifted* from the archive marker (`isNotNull(deletedAt) && ne(column,
"archived")`). Resolving it would classify a soft-deleted row sitting in
a renamed archive lane as drift and "repair" it.

## The raw-SQL half is already partly triaged in place

`async-maintenance.ts` is marked DELIBERATE-LITERAL as a STATE marker,
and `async-archive-lineage.ts`'s soft-delete path writes `column =
'archived', deleted_at IS NOT NULL` as the storage state it has just set
— STATE by construction.

## What this changes about the decision

Roughly **three quarters LANE, one quarter STATE** — and the STATE sites
are the ones that destroy data if converted.

That is why "convert all three encodings" cannot be a sweep, and why the
raw count of 52 made it look larger than it is: there are fewer sites to
convert than the headline, and the ones that must **not** be touched are
the part worth being careful about.

## Not converted here, deliberately

The gate requires all three encodings to move together, so the
conversion is one coordinated change with its inventories updated in the
same commit. This supplies the classification that change needs without
pre-empting it — and without me making a 52-site coordinated change at
the tail of a long session, which is exactly when I have made my worst
calls today.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean. **No
census movement.**

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 06:43:18 -07:00
gsxdsm
a8dae03fdb fleet(dashboard): taskRevert 2 → 0 — the recorded blocker named the wrong variable (#3129)
The largest remaining census cluster. Deferred twice, with a blocker
that turns out to be false **in the same component where its
counter-example already lives**.

## What the earlier notes got right

`detailColumnFlags` describes the **modal's own task**, and the column
classified here belongs to a **neighbour**. Supplying it would answer
*"is this neighbour finished?"* with a different row's traits — wrong on
data, not merely stale on vocabulary. That reasoning stands and I kept
it.

An earlier pass also converted this, left the parameter unsupplied, and
**reverted it** — correctly. An unsupplied optional parameter is
strictly worse than the literal: the guard is gone, the census counts a
conversion, and the behaviour is the legacy fallback forever. That rule
is why the wiring ships in this same commit.

## What the conclusion got wrong

> "A correct conversion needs per-**neighbour** flags — which the modal
does not have and should not fetch mid-render."

`columnFlagsByTaskId` is a per-task map. It is **already a prop** of
`TaskDetailModal` (declared :367, destructured :727), and the call site
at :992 sits **below** that destructure.

And `TaskDetailModal` already uses it exactly this way, for the
near-duplicate canonical:

```ts
columnFlagsByTaskId?.get(nearDuplicateCanonical.id)
```

…under a note observing that *its* blocker had been *"asserted from the
shape of the problem rather than tested against what was in scope."*
Same assertion, one function over. So the supplier the earlier note went
looking for exists, is per-neighbour, and needs no fetch.

## What it fixes

This lookup skips **finished** candidates so a done/archived prior undo
attempt never renders as an active "Undo task" link. On a board that
renames those lanes it matched neither — a finished undo task kept
rendering as open, which is precisely the stale affordance the
function's own header says it exists to prevent.

## Census

| | before | after |
|---|---|---|
| `taskRevert.ts` | 2 | **0** |
| repo backlog | 17 | **15** |

## Measured

- 4 new cases; `taskRevert.test.ts` **11/11 pass**.
- **MUTATION**: restoring the literal pair fails the renamed case.
- **The negative is load-bearing.** The map is fail-soft, so a candidate
it does not cover must still be treated as **open**, not skipped. A
conversion that skipped unknown candidates would *hide live undo links*
— failing in the direction nobody reports.
- A **control** pins that an unwired caller (no flags at all) still
skips the legacy ids, so the optional parameter cannot regress default
boards.
- `TaskDetailModal` suites — **31 files / 664 tests pass**.
- `tsc --noEmit -p tsconfig.app.json` clean; census `--strict`,
`check-lane-wiring`, `check-fnxc-future-dates` clean.

## Pattern worth noting

This is the fourth deferral this session whose stated blocker had
dissolved or misidentified itself, and the second where the
counter-example was already in the same file. The common shape: a note
records *why* something is blocked, is accurate when written, and is
never re-checked — so the block outlives its cause. Re-reading them cost
minutes each and returned two real conversions.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:58:44 -07:00
gsxdsm
ada62a7c4a census: --claims shows which remaining files an open PR already holds (two duplicate claims today) (#3124)
The census says **where** the work is but not **who has it**, and
duplicate claims are now the dominant coordination cost of this phase.
This adds an opt-in `--claims` report mapping each remaining file to the
open PRs already touching it.

## The problem is measured, not suspected

- **`self-healing.ts` took three overlapping conversions** from
different lanes while one branch was open (#3049, #3075, #3078). Each
forced a full rebuild of #3094, and every conflict was the same shape:
*same guard, two spellings, different variable names*. That PR's body
asks, in as many words, for one lane to own the file.
- **`executor.ts` took two independent conversions today** — #3112 and
#3118 — same four literals, same payload-lanes fix, two branches. Two
workers each read the census, saw the top cluster, and started. Neither
could see the other; I only caught it because both appeared in one `gh
pr list`.

The census is what sends everyone to the same file, so the claim signal
belongs here rather than in a side channel nobody reads. `--triage`
(#3097) already measured the underlying fact — 53 of 88 guards sat
inside an open PR — one step short of being actionable.

## Measured on current main (29 guards)

```
  CLAIMED by an open PR: 6 files holding 15 guards
       6  packages/engine/src/self-healing.ts  ← #3121 #3116
       4  packages/engine/src/executor.ts  ← #3118 #3112
       2  packages/engine/src/auto-merge-finalization.ts  ← #3107
       1  packages/core/src/task-store/task-artifacts-ops.ts  ← #3120 #3119 #3091
       …
  UNCLAIMED: 12 files holding 14 guards — start here
       2  packages/dashboard/app/utils/taskRevert.ts
       2  packages/engine/src/scheduler.ts
       …
```

It independently reproduces **both** collisions I found by hand today,
which is the strongest evidence I can offer that it works: `executor.ts
← #3118 #3112` and `self-healing.ts ← #3121 #3116`.

It also answers the standing fleet instruction empirically. "Claim the
largest unclaimed cluster" currently resolves to **12 files holding 14
guards, none larger than 2** — and one of those two (`scheduler.ts`) is
in the SYNC-RESOLVED list, where conversion is inert. That is a
materially different picture from the headline `29`.

## Design decisions

**Report-only and fail-soft**, on the same terms as `--triage`: opt-in,
printed beside the totals, changes no count and no exit code. It shells
to `gh`, so it is unavailable offline, in CI without a token, and in
sandboxes — all of which print a notice and continue. A gate must not
depend on network state; this is a work-selection aid, not a gate.

**The fail-soft path is loud on purpose**, and it is the case I care
most about. A claim report that silently degrades to "nothing is
claimed" is *worse than no report*, because it actively sends the reader
into work another lane holds — the exact failure the flag exists to
prevent. So when `gh` cannot answer it prints `POSSIBLY CLAIMED` and
suppresses the start-here list entirely rather than rendering it empty.

**Heuristic, and says so.** A PR touching a file is not proof it
converts *that file's* guards — it may edit an unrelated function. It
over-reports rather than misses, which is the safe direction: a false
claim costs one comment asking, a missed one costs a rebuilt branch.

**One bulk `gh pr list` call**, not a request per PR — the per-PR shape
was too slow to become habitual, and a report nobody runs is not a fix.

## Verification

- `lifecycle-column-census.test.ts` — **42 passed** (was 40)
- Differential: disabling the flag gives **2 failed | 40 passed**. Both
new tests fail on the defect they were written for.
- `--strict` and `check-fnxc-future-dates` — exit 0
- Tests stub `gh` on PATH, so no network call and no dependency on the
live PR list. The fixture reads the census's **own current top file**
rather than a hardcoded path, so it cannot rot as the backlog shrinks
(same self-maintaining discipline as #3106).

## What this does not do

It does not reserve anything — there is no lock, and two workers who
both run it can still collide if they start simultaneously. It reports
what is already visible in the PR list, which is enough to catch the
every-case-so-far pattern of *starting work on a file someone has held
for hours*. A real reservation would need shared mutable state, and I
would not add that without an owner asking for it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:50:13 -07:00
gsxdsm
d09a856941 test(engine): pin the FN-5256 liveness guard on a renamed board (the sweep that clears a live task's worktree) (#3132)
Top item from the verified coverage map on #3115.
`reconcileTaskWorktreeMetadata` had **three** uncovered resolvers — the
most of any sweep in the file — and it is the one that nulls
`worktree`/`branch`/`sessionFile` on a live row.

## Why this sweep first

Its own header names FN-5256: the incident where clearing worktree
metadata yanked a checkout out from under a running shell. The guard
that prevents it is `scopeOverrideMergeActiveSafe`, and that guard is
exactly what the wip/review resolvers feed.

The existing guard test uses `column: "in-progress"` — **the literal**.
So blinding `worktreeReconcileWipColumns` back to `["in-progress"]`
leaves all 825 self-healing tests green. The guard is converted; nothing
in the suite could tell.

On a renamed board the pre-conversion form matched nothing,
`scopeOverrideMergeActiveSafe` became true for a card an executor was
actively running, and the sweep cleared its metadata.

## The case

The renamed twin of the existing FN-5256 test: a `scopeOverride` task
live in a **renamed wip lane** keeps its metadata. Same shape, same
assertions, different vocabulary — which is the whole point, since the
original passes either way.

**Measured:** 414 pass; blinding `worktreeReconcileWipColumns` to the
legacy id fails **exactly this test**.

## Remaining from the map

25 uncovered resolvers left. Next by risk: `reclaimStaleActiveBranches`
(deletes branches) and `reconcileInReviewBranchRebind` (rebinds branches
of live cards) — both need a git-shelling harness, so they are slower to
pin than this one was. Then the two `reconcileDependencyBlockingLeases`
resolvers.

I will keep working down that list. The map is on #3115 with verified
names; anyone can pick an entry and check it the same way — blind one
resolver, run `vitest run src/__tests__/self-healing`, and if it stays
green that conversion has nothing behind it.

## Verification

`self-healing.test.ts` **414 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.
2026-07-31 05:49:59 -07:00
gsxdsm
ce84aa48d0 test(self-healing): cover the renamed-board starved-refinement wake that main's conversion lacked (#3116)
**Rebased onto current `main`, and it shrank to one test.** Was
"self-healing consolidated (45 → 39)".

## What happened

**Every code change in this PR landed independently from other workers**
while it was open, and in each case theirs is equal or better. I took
theirs and dropped mine:

| My change | Landed on `main` as |
|---|---|
| pre-execution worktree seizure | `preExecLiveColumns` — same
"dangerous direction" reasoning |
| FN-5256 liveness cluster | `worktreeReconcileWipColumns` /
`worktreeReconcileReviewColumns` |
| agent-link membership | `agentLinkLiveColumns` /
`agentLinkTerminalColumns` |
| starved-refinement peer progress | `starvedWaitingColumns` — a project
union covering both duplicated sites |

Resolving the rebase by taking `main` left two orphaned declarations
(`activeOrQueuedColumns`, `holdPeerIds`) that nothing referenced. `tsc`
doesn't flag unused locals here, so I checked references by hand and
removed them rather than ship dead code that reads as converted.

## What's worth landing

**Their starved-refinement conversion has no renamed-board test — the
suite had zero.** This adds one.

A candidate resting in a renamed **intake** lane, with its peers in a
renamed **hold** lane, must still escalate. The two are deliberately
distinct columns so a wrong role set resolves no peers and escalates
nothing; a fixture where they coincide would pass either way.

The fake needed `listWorkflowDefinitions` — `starvedWaitingColumns` is a
**project union**, so per-task selection readers alone leave it
resolving nothing and the test would pass for the wrong reason. That
mismatch is how I found the gap: my original test failed against their
implementation.

## Verification

- Green against **their** code
- **Revert-proof against theirs:** restoring the literal fails it — 0
escalations against 1 expected
- 8 tests in the suite green

## Note for the fleet

This is the second PR of mine to shrink to a test on rebase (#3096 was
the first). Both times the duplicated work was real and mine was the
later arrival. The pattern is worth acting on at the coordination level,
not by me working faster.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:41:32 -07:00
gsxdsm
6483f9ce2b fix(scheduler): resolve task:updated / task:deleted lanes asynchronously (scheduler inert 5 → 0) (#3128)
The last inert guards in `scheduler.ts`. Independent of my other
branches.

## Inert-guard ratchet

| Scope | Before | After |
|---|---:|---:|
| `scheduler.ts` | 5 | **0** |
| total | 12 | **7** (triage.ts 8 → other worker; executor.ts 4 → #3112)
|

## The live bug

These read `resolveTaskParkedColumnsSync`, which answers with the
**default** workflow in production. On a renamed board the scheduler
**never woke** on unpause or planning-finish, and a **deleted blocker
never unblocked its dependents** — the card sat behind a task that no
longer existed.

## The criterion, restated because I got it wrong before

**What blocks a guard is whether its answer is consumed synchronously —
not whether the enclosing listener is declared sync.** I assumed the
latter earlier in this program and reverted for it.

All three fail that test: two only gate `schedule()`, which is itself
`async`, fire-and-forget and re-entrance-guarded; the third already sits
below an `await getSettings()`. The edge-trigger bookkeeping
(`planningTaskIds.delete`) **stays synchronous** on purpose — deferring
*that* would let a second update re-enter the branch.

## The union is load-bearing, not defensive

Post-U11 the default lineage has no `triage` column, so a **resolved**
answer returns `intake: "todo"` where the inert path fell back to
`"triage"`. Converting without unioning the legacy ids silently
**narrowed** the wake set and stopped waking cards in a legacy-named
lane — caught by *"schedules when planning clears in triage"*.

**A resolved conversion must be a superset of what it replaces, or it is
a behaviour change wearing a vocabulary change's clothes.** That's the
reusable lesson here.

## Tests

- Drained with the repo's existing **`flushAsyncHandlers`** helper —
written for exactly this fire-and-forget shape — rather than loosening
any assertion.
- **The characterization test flipped, as designed.**
`workflow-scheduler-parked-columns-live-e2e.pg.test.ts` asserted *"a
dependent in a RENAMED hold column is NEVER unblocked"*, with its author
noting: *"expected to flip to null the moment the resolver is fixed —
and that flip is the whole point of writing it down."* It flipped.
Inverted to a REGRESSION case so the assertion holds the fix rather than
the defect; it now matches its own CONTROL arm, which still guards
against a vacuous pass.

## Verification

- 21 scheduler suites — **361 green**, including the live PostgreSQL e2e
- **`pnpm test:gate` green**; eslint and `tsc` clean
- Changeset added; `check:changesets` passes

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:41:16 -07:00
gsxdsm
ad5172afd5 fix(engine): main is red on check:inert-sync-lanes — #3114's triage conversion is inert, revert the arm (#3126)
## `main` is red on `check:inert-sync-lanes` right now

```
inert-sync-lane: NEW inert conversions — a lane guard now reads a sync resolver
that always answers with the DEFAULT board.
  packages/engine/src/triage.ts: 7 -> 8
```

Verified on a clean `origin/main` checkout, not on my branch. #3114
converted this guard's third arm to `disposeLanes.wip`; the gate that
exists to catch exactly this fired, and the PR landed anyway —
presumably because `check:inert-sync-lanes` is not in the blocking
merge-gate set.

## The change did not change behaviour

`disposeLanes` comes from `resolvePlannerLanes`, which resolves through
`resolveTaskWorkflowIrSync` — inert under PostgreSQL for two independent
reasons (#3103). So `disposeLanes.wip` evaluates to `in-progress`: **the
same value as the literal it replaced.**

A card advancing into a renamed execution lane still matches nothing,
still reads as an evacuation, and still kills a healthy planning session
— the precise bug #3114 set out to fix, unchanged on every board.

So the arm goes back to the literal. The gate's own failure text rules
out the alternative:

> Do NOT re-record the baseline to clear this — that is the same false
green one layer up.

## #3114's analysis is kept — only the code reverts

Its behavioural description is **correct** and is the clearest statement
of this bug anywhere in the file. I have kept those paragraphs and added
what is missing: that the fix does not reach under PG, and what would.

Whoever supplies a lane answer that is not sync-resolved should make
this line read `disposeLanes.wip` and delete the note. The specification
is sitting right there for them.

## It also reconciles two contradictory notes, one of them mine

My #3108 flag said converting the third arm this way adds an inert
comparison and removes a census entry that is telling the truth. #3114
then converted it and added a note saying it fixes the bug. **Both notes
sat in the file**, giving any reader two confident, opposite accounts.
They are now one account with the evidence attached.

## Read this file's census count carefully

#3114 took it to **0** while the inert count went to **8**. The census's
own `--triage` output warns about exactly this shape:

> for a sync-resolved file, a count of 0 is the WORST case, not the best
— the file reads as fully converted

Reverting restores it to 1, which is the honest signal.

## Census

| | before | after |
|---|---|---|
| `triage.ts` | 0 | **1** |
| repo backlog | 26 | **27** |

**The number going up is the point.** A census that reports 0 for a file
whose guards are all inert is worse than one that reports the truth — it
retires the entry and nobody looks again.

## Measured

- `check-inert-sync-lane-conversions`: **exits 1 on `main`, 0 here** (8
→ 7).
- `src/__tests__/triage*` — **25 files / 374 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-fnxc-future-dates` clean.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:35:28 -07:00
gsxdsm
c38d784b88 fix(events): carry lanes on the archive and completion task:moved emits (#3120)
Follow-up to **#3109** (merged). Independent of my other branches.

## The gap

#3109 put resolved lanes on `task:moved` and wired the main move path in
`moves.ts`. **Two other emit paths still fired lane-less** — and to a
listener, a lane-less emit is not a *missing* answer, it is the
**legacy** answer, which on a renamed board is wrong.

Concretely: `archiveTaskBackendImpl` emits its own `task:moved`. The
executor's archive branch releases the task's active-session registry
entry, and **that entry is what blocks a successor task from acquiring
the same path**. Fixing the listener alone (#3112) leaves the leak
reachable *through this emitter*, because the listener still falls back
to the literal when the payload carries nothing.

That's the part worth noting for the pattern generally: **a
payload-carrying design is only as good as its emitters.** Converting
consumers without sweeping producers leaves a hole that looks fixed at
the call site.

## Scope

| Emit path | Action |
|---|---|
| `archive-lifecycle-2.ts:373` (`archiveTaskBackendImpl`) | carries
lanes |
| `task-artifacts-ops.ts:549` (`moveToDoneImpl`) | carries lanes |
| `lifecycle-ops.ts:715` | **untouched** — fires from a watcher callback
|
| `task-update.ts:962` | **untouched** — sync path |

Both converted sites are already `async` and already import
`resolveWorkflowIrForTask`, so this costs one IR read on a transition
that has just done database work. Fail-soft to `undefined`, matching
`moves.ts` — "unknown", never a wrong answer.

The two untouched ones each need their own look rather than a blanket
sweep; flagged, not guessed.

## Verification

- `@fusion/core` archive/lifecycle suites — **156 green**
- **`pnpm test:gate` green**; `tsc` clean
- Changeset added; `check:changesets` passes

## Census

**No change**, and that's expected — this converts emit *payloads*, not
comparison literals. The effect is that guards already converted in
#3109/#3112 receive a correct answer on these paths instead of a legacy
fallback.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Task move events now include accurate board lane information when
tasks are archived or marked complete.
* Improved handling for renamed workflows, ensuring task transitions use
the correct lane details.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:29:53 -07:00
gsxdsm
c5d5978a01 test(core): the emitter route does not generalise from task:moved to task:updated — measured (#3123)
#3109 solved the inert-guard class for `task:moved` by having the
**emitter** resolve lanes once and carry them on the payload. It is the
right fix there, and it retired several flags within a day — two of them
mine (#3118, #3121).

**The obvious next step is to do the same for `task:updated`**, which is
where every remaining sync-listener guard lives: `scheduler.ts`'s
mission-failure and PR-monitoring guards, and `triage.ts`'s
planning-evacuation guard. I went to do exactly that, and the cost
profile is opposite.

## Measured

| event | emit sites | files |
|---|---|---|
| `task:moved` | **7** | — |
| `task:updated` | **26** | 10 |

#3109 could argue its resolution away because a move *"is already async
and already post-commit, so the resolution costs one IR read on a
transition that has just done database work."*

`task:updated` fires on every **log append, comment, artifact write and
steering message**. `audit-ops.ts`'s logEntry **fast path** is one of
those emit sites — and that path exists specifically to avoid re-reading
the task. An IR read per emit there is a regression on the hottest write
path in the system.

## Partial coverage does not rescue it

The tempting narrower version is "add lanes only to the emit paths those
guards react to". `triage.ts` rules that out: its evacuation handler
reacts to **any** `task:updated` carrying a column, and explicitly
tolerates partial payloads (that is why it checks `typeof task.column
!== "string"`). It needs lanes on essentially every emit, or it keeps
its literal fallback regardless.

So those guards are **not one commit behind #3109**. Unblocking them
wants either a cached lane answer the emitter can attach for free, or
the sync reader this file already specifies.

## Pinned as a ratio, not prose

The new case asserts `task:updated` emits **more than twice**
`task:moved` — the shape of the argument rather than today's exact
numbers, which move with ordinary work. If the ratio ever converges, the
trade-off has genuinely changed and the note above it should be re-read.

That is deliberate: a comment stating "26 vs 7" would be wrong within a
week and would then argue for the opposite conclusion with full
confidence. This is the third time this session I have found a note that
was accurate when written and had quietly stopped being true.

## Measured

- 5 cases pass (4 pre-existing + 1 new).
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted.** This records why the cheap-looking
follow-on is not cheap, in the file that already owns this decision
(#3103).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:24:05 -07:00
gsxdsm
4b61170a51 fix(executor): read task:moved lanes from the payload (executor.ts 4 → 0) (#3112)
**Stacked on #3109** — merge that first; this is its first consumer.

## Census

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 47 | **43** |
| `executor.ts` | 4 | **0** |

`executor.ts` is off the census top-files list.

## Why these four could not be converted in place

This listener is synchronous and its branches **start execution**,
dispose worktrees and release sessions. An await ahead of them defers
the `execute()` dispatch itself. The sync IR resolver isn't an option
either — it answers with the default workflow under PostgreSQL, so a
guard written through it is inert.

Reading the lanes the emitter already resolved costs nothing and leaves
the prologue synchronous. This listener is the reason #3109 has the
shape it does.

## The archive branch is the one with teeth

`to === "archived"` matched nothing on a board with a renamed terminal
lane, so **archiving never released the task's active-session registry
entry** — and that entry is what blocks a **successor** task from
acquiring the same path. Not cosmetic: the next task wanting that path
fails to register.

## Verification

- **Revert-proof:** the new case drives a `shipped` terminal lane
(matching no legacy id) and asserts the release. Reverting the branch to
the literal leaves the entry held — `expected [Array(1)] to have a
length of 0`.
- 43 executor suites — **483 green**
- **`pnpm test:gate` green**; eslint clean

## Note on shape

Lanes are read as **single ids, not sets**, because each branch here is
a lane-identity test on one column — exactly what the literals were.
Widening to membership would change behaviour, not just vocabulary.
Fail-soft to the legacy ids when the emit path could not resolve,
matching every other consumer of this payload.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:15:38 -07:00
gsxdsm
218086bea2 fleet(engine): self-healing 6 → 1 — the board-stall counter, the last guard that needed a sync answer (#3121)
The last fan-out guard, and the one I explicitly said needed a
synchronous answer. #3109 made that answer available without an await,
so the flag comes off.

## Why this one was last

The other two guards in this listener gated work the listener **already
`void`s**, so they moved onto the async resolver in #3094. This one
increments in-memory state **in the handler's own tick**, so it
genuinely needed a synchronous answer.

The sync IR path was never that answer: `resolveTaskWorkflowIrSync`
cannot resolve a **custom** workflow at all — two independent blockers,
#3103 — which is why I wrote that conversion, measured it, and withdrew
it.

#3109's emitter-carried `lanes` removes the dilemma rather than trading
one horn for the other: reading them needs **no await**, so the
increment stays in the same tick *and* the guard becomes correct.

## What it fixes

On a renamed board this counter read **zero**. The board-stall watchdog
was blind to a board whose cards were moving out of implementation the
whole time — the signal it exists to raise was never raised.

## Census

| | before | after |
|---|---|---|
| `self-healing.ts` | 6 | **1** |
| repo backlog | 29 | **24** |

The remaining 1 is the log-dedup closure — a pre-existing flag whose
degraded answer costs a duplicate log line, not a lifecycle decision.

## Measured

- 3 new cases; `self-healing-completion-fanout.test.ts` **13/13 pass**.
- **MUTATION**: restoring the literal pair fails the renamed case.
- **The paired negative is the load-bearing one.** The guard means
*"left implementation for somewhere that is not implementation"*, so a
move **between two non-wip lanes** must not count. Without that case, a
conversion that counted every move would pass the positive and inflate
the watchdog's denominator — breaking it in the opposite direction,
which is harder to notice than a zero.
- A **fail-soft** case pins that an emit carrying no `lanes` still
counts on the legacy ids.
- **Asserted through the counter itself**, not a downstream alert. The
increment *is* what this guard decides; routing the assertion through
the watchdog would let an unrelated threshold change mask a regression
here.
- `src/__tests__/self-healing*` + `task-agent*` — **42 files / 848 tests
pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## On the withdrawal this reverses

#3094 withdrew a sync-IR conversion of this listener and recorded why,
precisely. That record is what made this cheap: I could tell in one read
that #3109 addressed the *specific* obstacle rather than a general
"async is hard". A flag that names its blocker exactly is a flag that
can be retired the day the blocker goes.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:15:14 -07:00
gsxdsm
56b5cdfeed test(notifications): cover the wedge-episode renamed-lane clear that main's conversion lacked (#3096)
**Rebased onto current `main`, and it shrank to a test.** Was "close the
wedge-episode race, then resolve its lanes."

## What happened

Another worker landed **both halves of this PR independently** while it
was open. Rebasing showed their versions are better, so I took theirs
and dropped mine:

- **The serialisation** — theirs is `enqueueWedgeHandling(taskId, run)`,
a general callback; mine was wedge-specific.
- **The conversion** — theirs is **project-union membership** over the
four roles; mine was first-match-per-role via `resolveLifecycleColumns`.
Membership is correct: more than one lane can fill a role on a renamed
board, and first-match silently ignores the rest.

My rebased branch initially compiled to a **duplicate
`wedgeHandlingChains` field and duplicate method** — caught by `tsc`,
removed. Nothing of my implementation survives, and it shouldn't.

## What's left is worth landing

Their conversion has **no renamed-board test**. This adds one.

A card recovering into a renamed hold lane must **clear** its episode.
Asserted through the *second* notification, because a stale active
episode also **refuses the next genuine wedge its claim** — so the
visible symptom is a real wedge going unannounced, not merely a stale
alert.

The fixture needed `listWorkflowDefinitions`:
`resolveProjectColumnsForRoles` unions across the project's workflows,
so the per-task selection readers alone leave it resolving nothing and
the test would pass for the wrong reason. That's how I found the
mismatch — my original test failed against their implementation.

## Verification

- Green as written against **their** implementation
- **Revert-proof against theirs:** restoring the four literals fails it
— 1 delivered, 2 expected
- 7 notification suites — **80 green**; `tsc` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed wedge notifications so they can trigger again after a task
recovers into a renamed workflow’s hold lane.

* **Tests**
* Added regression coverage confirming that recovered tasks correctly
clear their wedge state and support subsequent notifications.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:09:36 -07:00
gsxdsm
6050d6eb83 chore(engine): mark the auto-merge-finalization reviewed literals DELIBERATE (census 47→45) (#3107)
Fleet phase. `packages/engine/src/auto-merge-finalization.ts` was the
last census file with no branch, worktree, or open PR against it. Claim
published by pushing the branch before starting.

## Census before / after

| | total | this file |
|---|---|---|
| before | **47** | 2 |
| after | **45** | 0 |

`--strict` exits 0, baseline re-recorded. **Reclassification, not
conversion** — both lines are unchanged.

## Both sites were already reasoned, in a note that calls them
non-defects

- **Line 30** is the resolver's **degraded fallback arm**, inside
`catch`. The live arm two lines up calls `columnHasFlag(ir, columnId,
"complete")`. The literal is reached only when IR resolution throws,
where the legacy id is the only answer left — removing it would make a
failed resolve return nothing.
- **Line 99** picks an **error string**. The note above it works through
threading `isCompleteColumn` in and concludes the signature widening
costs more than the sharper diagnostic buys.

I did not revisit either judgement. The gap was mechanical: prose the
census cannot read, so both stayed in `byFile` as apparent debt for the
next pass to re-derive.

## This is the fourth, and it closes the set

With #3056, #3060, and #3063, **every census file that was unclaimed
during this phase has now been examined, and not one needed a
conversion.** Each site was a three-state fallback arm, or a site a
prior pass had already reviewed and kept.

The corollary is the finding I would most want carried forward: the
remaining count is not a work queue. A worker told to "claim the largest
cluster" reads the number, finds most of it already reasoned, and
reaches for whatever moves it — which is how three PRs converted guards
to a synchronous resolver that is inert under PostgreSQL.

One exception worth preserving: **`taskRevert.ts` should stay counted.**
I claimed, inspected, and released it without marking. Converting it
would classify a *neighbour* row using the modal task's flags — wrong on
data, not merely stale on vocabulary — and its note correctly calls the
entry **accurate debt** blocked on a per-neighbour flag map. Fallback
arms and dead paths → mark. Placeholders awaiting a capability → leave
counted.

## Verification

- `census --strict` exit 0; `tsc --noEmit` (engine) **0 errors**
- No dedicated test file for this module (`vitest` reports none), so no
suite to run — comment-only diff, no behaviour change

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:09:23 -07:00
gsxdsm
82329819f7 fleet: resolve the pre-archive unarchive target (census 69 → 68) (#3091)
## Census

| | column guards |
|---|---|
| before | **69** |
| after | **68** |

## How this was found

By finishing a triage I'd left incomplete. Of the 19 single-guard files,
I had actually examined six and flagged the rest partly on assumption —
so I went back and read them.

Five of the remaining ones turned out to be `archived` comparisons
**pinned by `archived-column-gate-parity.test.ts`** (`audit-ops`,
`task-id-integrity`, `mission-store`, `async-comments-attachments`, plus
`merge-queue-ops-2` for review). Converting any of those moves one of
three encodings that must move together.

**This one isn't pinned**, and that difference is the whole PR.

## What changed

```ts
if (!declaresPreArchiveColumn || preArchiveColumn === archivedColumn
    || preArchiveColumn === "archived")
```

Belt-and-braces: the condition already accepted the resolved lane **or**
the legacy id, stated twice. A set says it once, so the two halves can't
drift apart — the real risk with a duplicated condition, rather than the
census count.

## Why this isn't the split brain

The parity guard pins comparisons of a **task's column** — one of three
encodings of *"an archived task is not live."* This compares a **stored
`preArchiveColumn` value** against the board's archive lane: a different
question, about where to send a card on unarchive.

Verified rather than argued — that suite runs **green** here, and it
went **red** the last time I touched a pinned site (#3076, where I named
an arm and immediately reverted). It's a live check, not an assumption.

## Measured

| check | result |
|---|---|
| archive / artifact / unarchive suites | 7 files, **27 tests green** |
| `archived-column-gate-parity` | **2 passed** |
| five gates + strict census | green |
| core `tsc` | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:06:22 -07:00
gsxdsm
89b21e2906 fleet: triage's planning-evacuation check uses the resolved wip lane (census 45 → 44) (#3114)
## Census

| | column guards |
|---|---|
| before | **45** |
| after | **44** |

## What changed

```ts
if (task.column === disposeLanes.hold || task.column === disposeLanes.intake
    || task.column === "in-progress") return;
```

Two role questions and one id question on the same line.
`resolvePlannerLanes` is **already called immediately above**, and its
result carries `wip` — so this needs no new resolution and no new await.
The literal just stops being the odd one out among its neighbours.

## What it cost on a renamed board

This handler aborts a planning session when a card leaves the planner
lanes. `in-progress` is excluded because *a card advancing into
execution is not an evacuation* — that's stated in the note directly
above it.

Against the literal, that exclusion **never matched** on a board whose
execution lane is renamed. So a legitimate advance into execution read
as an evacuation and **killed a healthy planning session** — precisely
the case the comment says must not abort.

`wip` is optional by design (PR #2628: a missing role stays `undefined`
so callers refuse rather than invent a column). Undefined here means the
board declares no execution lane, so there's no advance-into-execution
to exclude and the comparison is correctly false.

## Not addressed, and pre-existing

This line resolves through `resolvePlannerLanes` — the **sync** twin,
which returns the default workflow's lanes under PostgreSQL. That
affects all three lanes on the line equally and predates this change:
the handler is `(task: Task) => {}` with no await available, so fixing
it needs the same emitter-side change as #3082.

Making the third lane consistent with the other two doesn't deepen that,
and it leaves **one** shape to fix there rather than two.

## Measured

| check | result |
|---|---|
| triage / evacuation / planner-lane suites | **453 tests green** |
| four gates + strict census | green |
| engine `tsc` | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:06:08 -07:00
gsxdsm
6eeeb43d4b test(engine): pin #3047's archive-sweep conversion — measured uncovered (825 tests passed against the reverted fix) (#3115)
Not a conversion — the fleet's conversions are landing faster than their
coverage, and this is the audit that shows which ones actually have any.

## Method

For each of today's fleet commits to `self-healing.ts`: revert that
single commit, re-run the file's suites, see whether anything fails. If
nothing fails, the conversion has no regression protection and the "N
tests passed" cited on its PR was measuring something else.

| commit | reverted | verdict |
|---|---|---|
| #3075 pause-abort recovery | suites **fail** | covered |
| **#3047 archiveStaleDoneTasks** | **825 tests all pass** |
**uncovered** |
| #3078 (mine) | 204 tests all passed | was uncovered — closed by #3090,
#3102 |

Every fixture in the `archiveStaleDoneTasks` describe block uses the id
`done`, where the literal is correct, so none of them could see the
conversion at all.

## What the literal cost

The sweep's dependent scan skips tasks in terminal lanes. On a renamed
board **nothing matched `done`/`archived`, so every task read as
active** — which means every archive candidate looked like it had active
dependents, and the sweep archived **nothing**. The board quietly stops
auto-archiving: no error, no log line, no failing test.

## The case covers both halves of #3047

- a stale card in a **renamed complete lane** is archived — the
`complete` role
- a card whose dependent is still live in the **renamed wip lane** is
**not** archived, and a dependent already in the **renamed archive
lane** does not count as live — the `terminal` role

That second assertion is the one that matters: it stops the fix from
degenerating into "archive everything", which is the failure mode a
one-sided test would miss.

## Measured

**413 pass** on current main; reverting #3047 fails **exactly this
test**.

## Remaining audit

I have now audited 3 of ~13 fleet conversions to this file this way. The
method is cheap (one revert, one 17s suite run) and I will keep working
through the rest unless someone else picks it up. #3049 could not be
auto-reverted — later commits overlap its hunks — so it needs a manual
read rather than a mechanical revert.

## Verification

`self-healing.test.ts` **413 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.
2026-07-31 04:59:44 -07:00
gsxdsm
ec2921b958 fleet(engine): self-healing 23 → 6 — async-reachable guards, plus 3 of 4 fan-out guards the sync path could not serve (#3094)
**Replaces #3093, which I am closing.** Third rebuild of this work.

## A coordination note first, because it is costing more than the code

`self-healing.ts` has had **three** overlapping conversions land from
other lanes while my branch was open — #3049, #3075, #3078. Every time,
replaying my commits produced conflicts that were all the same shape:
*same guard, two spellings, different variable names*. Each rebuild is a
full cycle spent on merge mechanics rather than on lanes.

I have rebuilt against `main`'s own census each time rather than argue
about whose spelling wins, and this PR contains only what `main` (23)
does not have. But if this file is going to keep receiving concurrent
fleet passes, one lane should own it — otherwise the next PR pays the
same tax again.

## Converted

| site | note |
|---|---|
| `isPhantomExecutorBinding` | caller resolved
`lanesOfReclaim(task.id).wip` **three lines above the call**, then
passed a task whose column the predicate compared against `in-progress`
|
| `isWorkspaceOwnerLive` | required `completeColumns` |
| `recoverPausedAbortFailures` **body** | #3075 converted this sweep's
*router* and left three body guards comparing ids |
| `reconcilePreExecutionWorktrees` | a four-id literal in a sweep that
**removes worktrees** |
| `recoverStarvedRefinementTriageTasks` | 2 peer counts that read zero,
so escalation never fired |
| `evaluateParkedAgentTaskLink` | the omitted `parkedColumns` argument |

All through the **async** `resolveProjectColumnsForRoles`, whose only
store read is `listWorkflowDefinitions()` — answerable under PostgreSQL.
That is what separates these from the inert kind.

**The half-converted sweep is the important one.** A router that
resolves correctly feeding a body that compares ids is worse than
converting neither: the route now fires on a renamed board and the body
then acts on the wrong lane. The `moveTask` **target** is the sharp end
— an undeclared target is rejected *except* under `recoveryRehome` with
a legacy id (`moves.ts:570`, the #1411 escape hatch), so a converted
route feeding the literal `"todo"` rehomes the card into a column its
workflow does not declare, which is the state other reconcilers exist to
repair.

**A real dropped-behaviour bug**: `evaluateParkedAgentTaskLink` was
called without `parkedColumns`, falling back to `LEGACY_PARKED_COLUMNS`.
A live durable agent linked to a card resting in a renamed hold lane
read as not-parked, so the safeguard preserving its task link never
applied.

## Withdrawn: the `task:moved` fan-out

I wrote the sync-IR conversion, measured it, removed it.
`getTaskWorkflowSelectionImpl` returns `undefined` **unconditionally**
under PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR and `columnsWithFlag` on it yields exactly the legacy
ids — inert on every board.

Worse than the literal, because **the literal is counted**. My own test
passed only because its store mock supplied a renamed IR: it pinned the
helper's shape, not production behaviour. The refutation is recorded in
place, and `check-inert-sync-lane-conversions` exits 0 on this branch.

## One question, one answer

An earlier pass of this work mapped the notification-attach guard onto a
wider `activeWork` set, and `self-healing-paused-abort-recovery >
"rehomes an in-progress pause-abort park back to todo"` caught it — an
in-progress park attached a transition notification it should not have.

The fix is not a narrower set. Both guards ask **one** question — *"is
the card already at the requeue target?"* — which the literal happened
to spell twice as `=== "todo"`. The target now resolves once, before the
write, and both read it. Deriving one question two ways is exactly how a
converted guard and an unconverted target drift apart.

## Census

| | before | after |
|---|---|---|
| `self-healing.ts` | 23 | **11** |
| repo backlog | 53 | **41** |

## Measured

- `src/__tests__/self-healing*` + `task-agent*` — **42 files / 836 tests
pass**
- `tsc --noEmit -p packages/engine` clean;
**`check-inert-sync-lane-conversions` exits 0**; census `--strict`,
`check-lane-wiring`, `check-fnxc-future-dates` clean

## The remaining 11, flagged not guessed

- **4** — the fan-out, withdrawn above; blocked on a sync-capable
selection reader.
- **1** — the log-dedup closure: pre-existing flag; it sits before the
lane prefetch it needs, and the degraded answer costs a duplicate log
line, not a lifecycle decision.
- **1** — the synthetic `{ column: "todo" }` for a *missing* task:
deliberate, and now correct rather than unconverted, because
`parkedColumns` is legacy-seeded.
- The rest are status/deliberate classifications the census counts but
that are not lane guards.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved workflow automation for boards using renamed or customized
workflow lanes.
* Fixed task completion fan-out, branch rebinding, recovery, and
stalled-task detection across custom lifecycle columns.
* Prevented completion actions from triggering when tasks move back to
the work-in-progress lane.
* Improved cleanup and pause recovery behavior for customized workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:56:49 -07:00
gsxdsm
50f089a14b test(core): archived is renameable today — remove a wrong reason for the cheap parity option (#3113)
Follow-on from #3110, where the archived-gate parity test blocked my
conversion and laid out two ways forward. This removes a **wrong
reason** for picking the cheap one.

## The cheap option looks like it has already been taken

`archived-column-gate-parity.test.ts` offers: convert all three
encodings, **or** *"declare `archived` a non-renameable system column
and mark the sites deliberate"*.

The second is far cheaper — and the codebase reads as if it is already
true:

> `"Globally archived; hidden from the board. RESTRICTED (built-in
only)."` — `trait-types.ts`

At a glance that says a custom board cannot have an archive lane of its
own, which would make all 52 sites correct by construction and the whole
family **deliberate rather than debt**. That is a very attractive
conclusion for anyone facing a 52-site conversion.

## It does not mean that

The restriction is over trait **registration**. `trait-registry.ts`
rejects a **non-builtin** (plugin-defined) trait that declares
`archived` or `complete` (R22). It says nothing about which **column**
may carry the built-in trait, and a custom workflow may put it on a
column with any id.

**Verified, not argued:**

| input | result |
|---|---|
| `columnsWithFlag(ir, "archived")` where the lane is named `filed` |
`["filed"]` |
| `resolveColumnFlags` on that column | `{ archived: true,
hiddenFromBoard: true }` |
| `columnsWithFlag(ir, "complete")` where the lane is named `shipped` |
`["shipped"]` |

So option two is a **capability removal** — it would silently break any
board that has already renamed its archive lane — not the documentation
of a constraint that already exists.

## What this does and does not do

It **does not** decide between the two options; that is a product call
with real blast radius either way. It removes the reason someone would
most plausibly reach for after reading that flag comment, and it does so
with an executable check rather than a claim.

The `complete` case is pinned alongside it, because both restricted
flags behave identically — anyone reaching the same conclusion about
`complete` would be wrong for the same reason.

The first case also doubles as a **tripwire**: if `columnsWithFlag(ir,
"archived")` ever returns `[]` or `["archived"]` for that fixture, the
cheap option has effectively been taken and the parity test's framing
needs revisiting. That is the signal the case exists to give.

## Measured

- 4 new cases pass.
- The parity test still passes with its corrected framing — it excludes
`__tests__` from its scan, so the added prose cannot move its own
inventory. I checked that before editing it rather than after.
- `src/__tests__/{archived,trait,log-entry}*` — **3 files / 26 tests
pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted.**

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:56:25 -07:00
gsxdsm
77d11a0a9a docs(core): resolveLifecycleColumns returns one id per role, not a set (#3111)
Comment-only. No behaviour change. `tsc` 0 errors, eslint clean, `census
--strict` and `check:fnxc-future-dates` exit 0.

## Why this is worth a PR

Three fleet PRs have independently read this function's fields as
**membership**:

| where | finding |
|---|---|
| `self-healing.ts` `hold` guard | **#3084**, `Major` — *"resolve `hold`
by membership, not first-per-role"* |
| `notification-service.ts` progressed-lane check | **#3096** — reads
`?.hold / ?.wip / ?.complete / ?.archived` |
| sync vs async review set | **#3088** — same shape, different symptom |

Each field is a **single id**: the first column carrying that trait. A
board may declare several columns with one role — two pre-review working
lanes, or a review lane plus a merge-blocked lane — and every one after
the first is invisible.

Used for membership, that reads as a working check while silently
ignoring lanes: a card resting in the second hold column classifies as
**not-held**, and the guard depending on it never fires. It is the quiet
direction of wrong, which is why it survives review three times.

**Three occurrences of one misreading is a property of the signature,
not three unlucky authors.** Fixing it at the call sites one at a time
leaves the next author to rediscover it, so this documents the contract
where the mistake is made — at the definition — and names the
alternatives:

- `columnsWithFlag(ir, role)` — every column with the role on one board
- `resolveProjectColumnsForRoles(store, roles)` — the union across a
project's workflows

Rule of thumb recorded in the doc: **a routing/move target wants this
function; a `.has(task.column)` test does not.**

## What this does not do

It does not fix the three call sites. #3084's is a live `Major` and
needs a real fix with mutation verification; #3096's needs its author to
resolve a two-implementation conflict first; #3088's is flagged and
open. This only stops the fourth occurrence.

A stronger version would rename the fields (`firstHoldColumn`) or return
branded single-id types so misuse fails to compile. That is a wider
change across every caller and belongs to whoever owns the helper's API
— worth considering once the current fleet PRs land, since doing it now
would conflict with all three.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:50:28 -07:00
gsxdsm
41cdcc741e fix(events): carry resolved lanes on task:moved so listener guards stop being inert (#3109)
Removes the **inert-guard class at its source** instead of one call site
at a time. Independent of my other branches.

## The problem

`task:moved` listeners run synchronously, so a listener needing a lane
answer had to resolve one synchronously — and
`resolveTaskWorkflowIrSync` returns the **default** workflow under
PostgreSQL, the shipped backend. Every such guard behaved exactly as the
literal it replaced, while the census scored it as converted.

**Resolving asynchronously inside the listener is not available**, and
that is measured rather than assumed. The scheduler's
`snapshotManager.invalidate` is asserted to run in the listener's
**synchronous prologue**; putting an await ahead of it produced **3
failures across 21 scheduler suites**.

## The fix

The emitter carries the answer, which removes the dilemma rather than
trading one horn for the other. `moves.ts` is already async and already
post-commit, so it resolves the moving task's lanes **once** and hands
them to every listener. The guard becomes correct **and** the prologue
stays synchronous.

This is the file's own recorded preferred fix — *"having the emitter
carry the resolved lanes on the event payload so no listener resolves at
all"* — now that the audit it was waiting on is done and came back as
**one** prologue-dependent consumer, not a class.

## Design choices

- **`lanes` is optional and fail-soft to `undefined`** — "unknown",
never "legacy". Some emit paths fire from sync contexts or a cached row
mid-teardown. Listeners keep their existing fallback, so those paths are
no better than before but **no worse**, and they become the exception
rather than the rule.
- **`mergeParkedColumns` overlays only fields the emitter actually
resolved**, so a partial payload cannot blank a lane back to a wrong
answer.
- **The sync resolver stays** as that fallback. Deleting it would strand
the emit paths that cannot resolve.

## Verification

- **Revert-proof and it pins the prologue:** the new case asserts
invalidation on a **renamed** hold lane with **no `waitFor`**. Ignoring
the payload gives **0 calls**.
- 21 scheduler suites — **361 green**
- self-healing + notification suites — **491 green**
- core moves + the `sync-workflow-ir-callsite-allowlist` ratchet — green
- **`pnpm test:gate` green** (71)
- Changeset added; `check:changesets` passes

## What it unblocks

`scheduler.ts`'s 10 allow-listed guards now resolve correctly for every
move that goes through `moves.ts` — the path real moves take. Those were
already absent from the backlog, so **the census number does not move**;
what changes is that they now do what the number claimed.
`executor.ts`'s 4 remaining sites can follow the same pattern in a
separate PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:47:09 -07:00
gsxdsm
4afb32ef98 test(core): cover the untested log-entry archive gate; correct a deferral that named the wrong blocker (#3110)
I converted `audit-ops.ts`'s archived gate, measured, and **backed it
out**. Both halves of that are the deliverable.

## The old deferral was stale on its own terms

It declined the conversion because *"the fix is the same one
`getLiveTaskColumn` needs"* and doing one of the pair would leave them
disagreeing.

But `getLiveTaskColumn` now **takes** a resolved `archivedColumns` set,
and both of its callers already pass `await resolveArchivedLanes(store)`
— including the sentinel path **twenty lines up in this same function**.
The pair it worried about was already half-converted, and this arm was
the half out of step. Converting it would have made them *agree*.

That is the third deferral I have found this session whose stated
blocker had dissolved. A deferral note records the blocker at the moment
it was written, and nothing re-checks it.

## The real blocker is one neither note named

`archived-column-gate-parity.test.ts` failed my conversion, and its
reasoning is correct and not obvious. This gate has **three encodings**:

1. TypeScript comparisons
2. Drizzle `eq`/`ne` predicates
3. raw SQL templates

Converting only the TypeScript arm makes them **diverge**: the gate
would call the row archived while the SQL side still returns it as live
— a log write rejected by its gate while its parent is listed as live.

Every builtin workflow names the column `archived`, so all three agree
*by accident* on every board we ship, and nothing except that parity
test can see the split.

Unblocking means converting all three together — the SQL sides need the
resolved id as a query-build value, including inside `for update`
transactions that receive no store today — or declaring `archived` a
non-renameable system column. That test lays out both options and owns
the inventory that has to move in the same commit. I am not doing it
here; it is a different change from a lane conversion.

## What ships

**The corrected note**, and **a test for a gate that had no coverage in
any form**.

The test asserts the legacy refusal and — the case that matters more —
that a **live lane is not refused**. A gate that refused everything
would satisfy a one-sided test and silently break every log write on the
board.

The renamed case is recorded as a **deliberate, explained omission**
rather than left as a silent hole, so the next reader knows it is a
decision.

## Measured

- 3 new cases pass; the parity gate passes.
- **MUTATION**, on the conversion before I reverted it: restoring the
literal failed the renamed case. The conversion *worked* — which is
exactly why the parity gate mattered. A working change can still be the
wrong change.
- The live-lane negative asserts **the gate did not fire**, not that the
call succeeded: past the gate the fast path performs a real Drizzle
write this fake layer cannot serve, so asserting success would drag a
database fixture into a test about a lane comparison, and asserting a
bare rejection would pass even if the gate *had* fired.
- `src/__tests__/{log-entry,archive,cold-storage,unarchive}*` — **7
files / 23 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted, deliberately.** The count stays where
it is because the gate is blocked, not because it is fine.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:46:58 -07:00
gsxdsm
15a664a8f5 docs(engine): flag executor's four task:moved literals — the obvious conversion is provably inert (#3104)
The largest unclaimed census cluster. **Nothing in this file said why
the sync-lane pass skipped it**, and that silence is the hazard: the
obvious next move is to convert these the way `scheduler.ts`'s ten were
converted, which would make them **inert rather than fixed**.

## The literals are genuinely wrong — this is not a "non-issue" flag

All four sit in one synchronous `task:moved` listener, and on a renamed
board:

- execution **never starts** on a move into the board's own wip lane;
- terminal session release **never runs** on a move into its archive
lane;
- both `from` guards never fire, so **in-flight work is not aborted**
when a card leaves implementation.

Nothing errors. The engine simply stops reacting.

## Why the obvious fix is inert — proved, not argued

`task:moved` is emitted synchronously, so an `await` here reorders this
handler against every other subscriber. That points at the sync IR path,
which cannot answer for a renamed board for **two independent reasons**
(`sync-workflow-ir-second-blocker.test.ts`, #3103):

1. `getTaskWorkflowSelectionImpl` returns `undefined` unconditionally
under PostgreSQL, so `resolveTaskWorkflowIrSync` always takes its
`!workflowId` branch.
2. Even **with** a selection, the custom-workflow branch loads its IR
through `store.db`, whose implementation is an **unconditional throw** —
so it falls into the catch and returns the default IR anyway.

**A renamed lane is a custom workflow, so (2) alone is decisive.** The
sync path can never serve this listener's case, whatever the selection
reader is fixed to do. That is the part the existing notes across this
repo miss, and it is why flagging beats attempting here.

`check-inert-sync-lane-conversions` already baselines **twenty** guards
in exactly that state in `scheduler.ts`. These four must not join them.

## Census

**Unchanged at 4, deliberately.**

Marking them DELIBERATE-LITERAL would buy a smaller number by asserting
the code is *fine*. It is not fine — it is *blocked*. Those are
different claims with different expiries, and the census should keep
pointing here until the block is lifted. An unconverted literal is
visible; an inert conversion leaves the backlog and takes the evidence
with it.

## Measured

- Comment-only change.
- `src/__tests__/executor*` — **84 files / 853 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean.

## Unblocking, for whoever takes it

Either an async listener contract — a behaviour change to handler
ordering, not a column conversion — or a sync reader that answers for
**custom** workflows *and* survives a writer on another node. All three
constraints are written up in `sync-workflow-ir-second-blocker.test.ts`
(#3103).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:40:42 -07:00
gsxdsm
f5926d3b54 docs(engine): flag triage's evacuation guard — it looks two-thirds converted and is fully literal (#3108)
Completes the sync-listener audit across the three files holding the
remaining blocked guards — `executor.ts` (#3104), `scheduler.ts`
(#3100), and this one. Triage is the most misleading of the three.

## The shape lies

The guard reads as **two resolved arms and one literal**:

> `task.column === disposeLanes.hold || task.column ===
disposeLanes.intake || task.column === "in-progress"`

So the obvious next move is to convert the third arm with the same
helper. That is wrong twice:

**1. The two "resolved" arms are not resolved.** `resolvePlannerLanes`
goes through `resolveTaskWorkflowIrSync`, which cannot answer for a
**custom** workflow — the sync selection reader returns `undefined`
unconditionally, *and* the custom-workflow IR read goes through
`store.db`, whose implementation is an unconditional throw (#3103). So
`disposeLanes.hold` / `.intake` are `todo` / `triage` on every board.
**All three arms are literal in effect.** Converting the third the same
way adds a third inert comparison and retires a census entry that is
currently telling the truth.

**2. The guard's answer is consumed synchronously** — the criterion I
had to correct in #3104. Below it, `pauseAborted.add`,
`session.dispose()` and `activeSessions.delete` mutate in-memory state
in this tick, and other paths read those maps. Contrast
`self-healing.ts`'s fan-out, where three of four guards only gated work
the listener already `void`s and so *were* convertible via the async
resolver (#3094).

## What it costs, and the obvious reading is backwards

An evacuation **into** a renamed destination still falls through and
disposes correctly — no bug there.

The failure is the other direction: on a board whose **hold or intake**
lane is renamed, arms 1 and 2 stop matching, so a card **sitting still
in its own planning lane** is treated as evacuated and its live triage
session is aborted mid-run.

I state it that way because "renamed board → guard misses → nothing
happens" is the pattern everywhere else in this program, and here it
inverts.

## Census

**Unchanged at 1**, deliberately. Blocked, and now documented as *fully
literal* rather than part-converted — which is the fact a future pass
needs in order not to make it worse.

## Measured

- Comment-only.
- `src/__tests__/triage*` — **25 files / 374 tests pass**.
- `tsc --noEmit -p packages/engine` clean; `check-fnxc-future-dates`
clean.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:40:21 -07:00
gsxdsm
0e4a559a0a test(census): make the tighten fixture self-maintaining instead of pinned to committed state (#3106)
## What broke, and why it will break again

The two cases in this block assert the CLI tightens an inflated
allowance **by exactly the inflation**. That arithmetic only held while
the *committed* baseline matched the tree — so it broke the moment a
fleet PR took `self-healing.ts` from 26 to 22 without re-recording. The
CLI correctly tightened to 22 while the fixture expected 26, and both
cases went red for a reason that had nothing to do with the code under
test.

#3101 fixed that instance by committing the number. **This fixes the
class.**

## Why it recurs

The census **exits 0 on a drop** — deliberately, so one worker's merge
can't redden the gate for everyone else. The cost is that the committed
baseline goes stale *silently*: every run rewrites the file, prints
`COMMIT IT`, and exits 0. This fixture is what eventually trips over it.

With a fleet actively converting the largest file (eight open PRs
against `self-healing.ts` as I write this), that's a recurring red, not
a one-off.

## The change

The fixture syncs its temp copy to the tree with `--strict
--update-baseline` **before** inflating. The assertion is then about the
CLI's behaviour rather than about what happens to be recorded on disk.

## Differential proof, both directions

Against an artificially staled baseline (22 → 26):

| | result |
|---|---|
| with this change | **40 passed** |
| without it | **2 failed / 38 passed** |

So the fixture now tolerates drift it previously broke on — and still
fails if the CLI stops tightening, which is the property it was written
to guard. That second half matters: a fixture made tolerant of
everything would be worse than the flake.

The CLI invocation is extracted to a `runCli` helper so the sync run and
the assertion run share one path. No behaviour rides on that extraction.

## Measured

| check | result |
|---|---|
| census suite, clean tree | 40/40 |
| census suite, staled baseline | 40/40 |
| `--strict` | exits 0, no residual drift |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:33:37 -07:00
gsxdsm
39e6891c93 chore(core): mark the dead sync-path lane literal DELIBERATE-LITERAL (census 104→102) (#3060)
Fleet phase. Claimed `packages/core/src/task-store/project-store-ops.ts`
— the largest census file with no branch, worktree, or open PR against
it. Claim published by pushing the branch **before** starting work.

## Census before / after

| | total | this file | deliberate |
|---|---|---|---|
| before | **104** | 2 | 130 |
| after | **102** | 0 | 130 |

`--strict` exits 0, baseline re-recorded in the same commit.
**Reclassification, not conversion** — the line is unchanged.

## The site was already audited today, in prose the tool cannot read

```
FNXC:WorkflowLifecycleColumns 2026-07-31-02:45 (audited — DEAD SYNC PATH, do not convert):
… It is the SQLite-mode twin. The live path is `dequeueMergeQueueOnColumnExitInTransaction`
… and it is ALREADY converted … This body reaches for `store.db.prepare`, which throws in
   PostgreSQL backend mode …
```

The reasoning is sound and I did not second-guess it: the live path is
converted, this twin cannot execute in production, and converting it
would mean threading a lane set into a function whose first statement
throws.

The problem is purely mechanical — **the note is prose, and the census
reads markers.** So the site stayed in `byFile` looking like unconverted
debt, and each fleet pass pays to re-derive the same conclusion. Adding
`DELIBERATE-LITERAL` moves it to `deliberateByFile`, where a
reviewed-and-kept literal belongs.

## This is the second one, which makes it a pattern

Same shape as #3056 (`async-mission-store-queries.ts`, fallback arms).
Across the files I have checked this phase — `agent-store`,
`github-tracking-state`, `planner-overseer`, `auto-merge-finalization`,
`async-mission-store-queries`, and this one — **every site was either a
fallback arm or an already-documented deliberate leave**, and
`agent-store.ts:236` carries its own "FLAGGED AND LEFT COUNTED" note
from today.

So the count is not a work queue, and the gap is not judgement —
previous passes reached the right answer. They recorded it where only a
human reader would find it. Two lines of marker per site closes that,
and the number then means "conversions owed", which is how every worker
reads it when picking a cluster.

## Verification

- `census --strict` exit 0; `tsc --noEmit` **0 errors**
- `check:fnxc-future-dates`, `check:lane-wiring`,
`check:sql-column-literals`, `check:inert-flag-seams` — all exit 0
- No behaviour change: only a comment added

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:30:00 -07:00
gsxdsm
0da19f7963 fix(core): a renamed archive lane was recorded as done in the eval corpus; flag the scheduler's two honest literals (#3100)
Two pieces, both about the same distinction: which literals are worth
**converting** and which are worth **naming**.

## Converted — the eval corpus was mislabelling renamed archive lanes

`collectDeterministicSignals` writes `column` as a two-value eval-record
field. Against the `archived` literal, a card resting in a renamed
archive lane was recorded as `"done"`.

No crash, no lifecycle decision — a **mislabelled row in the eval
corpus**, which is a dataset every later comparison reads. That is the
expensive kind of quiet: nothing fails, the numbers just drift.

The collector is sync and pure (no store, no workflow), so the lane
answer arrives as an optional parameter.
`HybridEvaluatorService.evaluateTask` is async and already holds an
optional store, which is where the resolution is paid; a store-less
evaluator degrades to the legacy literal rather than failing.

**Only the archived arm was ever wrong.** A renamed *complete* lane was,
and remains, recorded as `"done"` — which is correct. So only that
answer is resolved, and a third case pins that the widening did not turn
every renamed lane into `"archived"`.

## Flagged, not converted — the scheduler's two honest literals

These are the two `scheduler.ts` literals the sync-lane pass did not
take, and **nothing in the file said why**. That silence is the problem:
the obvious next move is to "finish the job" the way the other ten were
converted, and that would make them **inert, not fixed**.

`getTaskWorkflowSelectionImpl` returns `undefined` unconditionally under
PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR — proved in
`postgres/sync-workflow-ir-is-always-default.pg.test.ts`, and
`check-inert-sync-lane-conversions` already baselines **twenty** guards
in that state in this same file.

They stay literal and **counted**, which is the honest state. An
unconverted literal is visible to the census; an inert conversion leaves
the backlog and takes the evidence with it. The note names the real
blocker — a sync-capable workflow-selection reader — so the next pass
does not spend a cycle discovering this the way I did.

## Measured

- 3 new cases in `eval-signal-collector.test.ts` — file **5/5 pass**.
- **MUTATION**: restoring the `archived` literal fails the renamed case
and leaves **both** the legacy control and the renamed-complete negative
green. The negative matters here: the fix must not turn every renamed
lane into `"archived"`.
- core eval suites — **4 files / 20 tests**; engine scheduler +
evaluator — **14 files / 143 tests**.
- `tsc --noEmit` clean in both packages; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

Both files keep their counts, deliberately:

- `eval-signal-collector.ts` — the remaining entry is the new
parameter's documented default, which is the fallback doing its job.
- `scheduler.ts` — the two literals this PR deliberately leaves visible.

A census that fell here would mean the flags had been marked exempt,
which would assert the code is fine. It is not fine; it is blocked, and
those are different claims with different expiries.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:29:21 -07:00
gsxdsm
f1e96f7a17 test(engine): pin the agent-link-drift terminal check on a renamed board (an agent stayed linked to finished work) (#3102)
Second of the two uncovered sweeps I flagged when #3078 merged. Not a
conversion — the conversion is already on main
(`driftedTerminalColumns`, landed by another fleet PR). This is the
coverage it shipped without.

## What was unprotected

Every existing case in this file uses `done` or `archived`, where the
literal is correct. So the terminal check had **no renamed-board case at
all**, and the same measurement that caught #3078 applies: a green file
proves nothing about a conversion whose fixtures can't express the
failure.

What the literal cost on a renamed board: **a durable agent stayed
linked to a finished task forever.** A linked agent is not free to pick
up new work, so the drift this sweep exists to clear is exactly the
drift it stopped clearing.

## Measured, both directions

| case | result |
|---|---|
| agent linked to a task in a RENAMED complete lane is cleared | **fails
on revert** — `taskId` still `"FN-9"`, agent pinned to finished work |
| agent linked to a task still in a RENAMED wip lane keeps its link |
passes either way — the sweep must narrow, not widen |

14 pass on current main; reverting the terminal check to the id pair
fails exactly one.

## Note on the fleet

This sweep was converted by someone else's PR while I was writing the
test for it — I found out because my revert probe hit
`driftedTerminalColumns`, a name I did not write. That is the collision
pattern working in a *useful* direction for once: their conversion, my
coverage, no duplicated code.

It also means the two of us independently chose the same sweep from a
7-PR pileup on this file. Assigning files from the census list would
still be cheaper than discovering the overlap in a test harness.

## Verification

`self-healing-agent-link-drift` **14 passed** · `pnpm test:gate` 13 +
161 + 487 + 71 · lint — green.
2026-07-31 04:25:59 -07:00
gsxdsm
827386dde6 test(core): the sync IR path is blocked TWICE, not once — every note in the repo undercounts it (#3103)
Every remaining census cluster I could not convert — `executor.ts` (4),
`scheduler.ts` (2), `triage.ts` (1), and the four-guard fan-out I
withdrew from my own PR — is waiting on the same thing. So I went to
unblock it, and found the record is wrong.

## The repo says one blocker. There are two, plus a constraint

The call-site allow-list header, the live-PG proof, and a dozen FNXC
notes across engine and core — **several of which I wrote** — all say:
`resolveTaskWorkflowIrSync` is inert because the sync selection reader
returns `undefined`, and the fix is "a sync-capable workflow-selection
reader".

That understates the work by half, and the undercount is load-bearing:
it makes the unblock read like a caching job, so the next person ships a
selection cache and finds the rest at integration time.

### Blocker 2 — the IR read is dead too

`resolveTaskWorkflowIrSyncImpl` loads a **custom** workflow's IR through
`store.db.prepare("SELECT ir FROM workflows WHERE id = ?")`.

`TaskStore.db` is not "SQLite-only". Its implementation (`dbImpl`,
`task-id-integrity.ts`) is an **unconditional throw with no mode branch
at all**. That read always throws into the surrounding `catch`, which
always returns the default IR.

The consequence is precisely the one this program cares about:

| workflow kind | after a perfect selection reader |
|---|---|
| built-in | resolves — that branch never touches `store.db` |
| **custom** | **still the default IR, always** |

**A renamed lane is by definition a custom workflow.** So the sync path
cannot serve the renamed-board case *at all* until this second read is
replaced. Fixing the selection reader alone would produce a change that
looks like it works — on default boards.

### Blocker 3 — a node-local cache is unsafe here

Not a bug; a constraint that bounds the fix's shape. Multiple Fusion
nodes run their own engines against **one shared PostgreSQL**
(`docs/multi-project.md` → "Shared Postgres multi-node runbook").

A node-local synchronous cache of `task_workflow_selection` therefore
goes stale whenever *another node* rewrites a selection — and answers
with full confidence. That is **worse than today's default**, which is
at least uniformly wrong rather than intermittently wrong. Any sync
reader needs an invalidation story that survives a writer on a different
host.

## Why a test rather than a comment

A comment saying "db always throws" decays the moment someone adds a
mode branch, and the whole argument silently inverts — which is the same
decay mode this conversion program keeps hitting with allow-list entries
and stale notes.

The assertions are deliberately about `dbImpl`'s **source** rather than
a call. Calling it proves one construction path throws; the fix depends
on the stronger claim that **no mode returns a database**. Reintroducing
an `if`/`return` there fails the test, which is the correct outcome: the
premise really has changed and the file must be re-read.

## Measured

- 4 new cases pass; the allow-list file's own 7 still pass with its
corrected header.
- **MUTATION**: adding a mode branch to `dbImpl` fails the first case.
- An **anti-vacuity** case pins that the resolver is still live and
still allow-listed, so these source assertions cannot keep passing after
the concern is deleted.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean.

## Census

**No movement — this converts nothing.** It corrects the record about
what the remaining conversions are waiting on, and it corrects notes I
authored. I would rather spend a PR making the next attempt cheap than
leave a half-true blocker in place that costs someone a full cycle to
rediscover.

## What I did not do

I did not build the sync reader. With blockers 2 and 3 in view it is a
store-substrate change — a second read to replace, and an invalidation
story that survives a writer on another host — not a fleet conversion,
and starting it mid-sweep on a shared file would repeat the collision
pattern that has already cost this branch three rebuilds. It remains
unclaimed, and now it is fully specified.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:25:48 -07:00
gsxdsm
d448ab6951 fix(engine): the merge-refusal reason was classified by a column id, and it lands in run-audit (#3098)
Claimed `auto-merge-finalization.ts` — and this one is a **reversal of
an earlier audit in the same file**, which is the interesting part.

## The earlier note said "diagnostic only". It was wrong about the
consequence

`validateWorkflowDoneMergeProof` picks between two refusal reasons with
`task.column === "done"`. Both arms return `{ ok: false }`, so this
never changed which branch ran — and on that basis a prior pass recorded
it as *"REAL but DIAGNOSTIC-ONLY"* and declined it, reasoning that
widening a signature to improve an error string is a poor trade.

**The reason is not an error string.** It is written to run-audit
metadata alongside `previousColumn` — `merger-merge-lifecycle.test.ts`
asserts exactly that — and that row is what an operator reads to find
out why a merge was refused.

So on a board whose complete lane is not called `done`, a card resting
in that lane was refused with the generic `missing-merge-confirmation`:
the classification for a card that is **not in the complete lane at
all**. The audit trail recorded the opposite of what happened. A wrong
record is worse than a vague one, because it gets acted on.

## The trade was also cheaper than the note claimed

The function is **already async** and **already takes an options bag**.
`resolveFinalizationColumns`, two functions up in the same file,
**already builds this exact predicate** for its own guard.

Nothing new is resolved. The answer that existed is handed down instead
of being re-asked with an id — the half-conversion shape this program
keeps finding, here inside a single file, one line apart: the caller
guards on the resolved `isCompleteColumn(latest.column)`, then calls a
validator that re-asked the same question with the literal.

`isCompleteColumn` is **optional with the legacy literal as its
default** — the same default-to-legacy contract the lane-parameter
vocabulary uses elsewhere — and `check-lane-wiring` watches the
parameter, so the two call sites cannot silently stop passing it.

## Measured

- New `merge-proof-reason-renamed-complete-lane.test.ts` — **2 pass**.
- **MUTATION**: dropping the parameter fails the renamed case and leaves
the legacy **control** green. The control earns its place: a failure now
means *"renamed board"*, not *"the refusal stopped working"*.
- **Driven through `finalizeProvenAutoMergeTask`**, not by calling the
validator with the new argument. The contract under test is the
**wiring** — a test that passed the argument directly would assert my
own parameter works and prove nothing about the seam that was broken.
- **The audit row is asserted, not just the return value.** The return
value alone is not the contract that failed here.
- merger / auto-merge suites — **5 files / 159 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

`auto-merge-finalization.ts` stays at **2**, deliberately. Both
remaining entries are now documented **degraded-fallback arms** — the
resolver's `catch` and this parameter's default — which is the right
kind of literal rather than a missed conversion. Converting a fallback
to a resolution would defeat its purpose.

## A note on the FNXC gate

My first stamps were dated `2026-08-01` while local today is
`2026-07-31`. `check-fnxc-future-dates` caught it and I re-stamped.
Worth mentioning because it is the second time this session that a
date-only local-calendar comparison has caught a stamp written near
midnight — the gate is doing real work, not ceremony.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:22:33 -07:00
gsxdsm
a7b2a757fa fix(engine): serialise wedge handling per task, then convert the lane guards it was blocking (5 → 1) (#3087)
The largest unclaimed census cluster, and the one two earlier fleet
passes explicitly declined.

## The standing blocker, taken on

Both passes converted these four ids and reverted, each time after the
same test went red:

```
task-wedge-notification.test.ts > sends one actionable push and mailbox message per active terminal episode
  expected 2 calls, got 1
```

Their diagnosis was right and I have kept it: this branch **resolves** a
wedge episode, `handleTaskUpdated` starts it fire-and-forget from a
synchronous `(task) => void` listener, and **any** await introduced
before the resolve lets a re-wedge arriving close behind reach `claim`
while the previous episode is still active — `claimed: false`, second
operator notification silently dropped. Column resolution needs an
await, so the conversion could not be made safe from inside the branch.

Both notes named the fix and left it for "whoever owns the wedge episode
contract": *serialise wedge handling per task*. This PR does that, then
takes the conversion.

## 1. Serialisation

`enqueueWedgeHandling` chains handling per task id, so
resolve-then-claim keeps its order however many awaits either branch
acquires. Details that matter:

- **Keyed by task, not global** — different tasks stay concurrent, so
this is not a throughput regression on a busy board.
- **The map entry is dropped when its chain drains**, and only if no
later link was appended while it ran, so it does not grow with the task
table.
- **Links never reject.** `maybeNotifyTaskWedge` already owns its error
handling; a rejected link would poison every later notification for that
task.

## 2. The conversion it was blocking

The four ids are an enumeration of *"every lane except review"* — the
lanes whose occupancy proves a wedged card's lifecycle has visibly
resumed. On a renamed board none of them matched, so a recovered card's
episode never resolved. Two consequences, and the second is worse than
the first:

1. the operator keeps an open "needs operator action" alert for work
that has moved on;
2. an active episode **suppresses re-claim**, so the *next* genuine
wedge on that task is never delivered.

Membership over the four roles, legacy-seeded, so an unconverted board
resolves exactly the four ids it used to compare.

## Measured

**The acceptance test the earlier notes named is the gate on both
halves.** With the conversion and *without* the serialisation, "sends
one actionable push and mailbox message per active terminal episode"
fails exactly as they reported. With the serialisation, green. I
reproduced their finding rather than taking it on trust — it is the
evidence that the serialisation is load-bearing and not incidental
refactoring.

| | result |
|---|---|
| `task-wedge-notification.test.ts` | **15/15** (2 new) |
| notification suites | **11 files / 234 tests pass** |
| `tsc --noEmit -p packages/engine` | clean |
| census `--strict`, `check-lane-wiring`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` | clean |

**MUTATION**: restoring the four literals fails the renamed-recovery
case and leaves its paired negative green.

**A vacuity I caught and fixed, worth stating plainly.** My first
version of the renamed case recovered the card with `status: "queued"`.
`hasProgressed` is an OR whose other arm is *"status is a non-failed
string"* — so that arm answered true and the column comparison never
ran. The mutation did not fail it. The case now clears `status` and
`error` together, which makes column membership the only thing that can
resolve the episode, and the paired negative uses the identical shape so
only the lane differs.

## Census

| | before | after |
|---|---|---|
| `notification-service.ts` | 5 | **1** |
| repo backlog | 71 | **67** |

## The remaining 1, flagged not guessed

`isManualMergeHold` (`task.column !== "in-review"`) is sync, and so is
its only caller `classifyWorkflowTransitionNotification`, reached from
the same `handleTaskUpdated` listener. Converting it means making that
whole chain async — a change to notification *classification ordering*
against every other `task:updated` handler, which is a different
contract from the episode one this PR owns. The serialisation added here
does not cover it: it wraps wedge handling, not transition
classification. Threading a pre-resolved `LifecycleColumns` in as a
parameter is the likely fix, and it wants the same gate-placement
judgement applied deliberately rather than swept in behind this.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:10:03 -07:00
gsxdsm
701677a2e5 test(engine): pin #3078's executor-owned skip — it merged without coverage (204 tests passed against the reverted fix) (#3090)
#3078 merged its conversion of the orphaned-pending-step-results sweep
**before this test landed**, so that sweep is on main with no coverage.
This closes the gap.

## The gap was measured, not assumed

With all three of #3078's conversions reverted, **all 204 self-healing
tests still passed**. I had cited that number as verification when I
opened it. It was meaningless for that change: every existing test in
the file uses `in-review` / `in-progress`, where the literal is correct,
so none of them could see the defect.

This is the same "a green suite is not coverage" failure I flagged in
other PRs today — in my own work, twice. The only reason I caught it is
that I finally ran the revert check on myself.

## Two cases

- **An executor-owned card in a renamed wip lane is SKIPPED.** Against
the pre-#3078 sweep this fails: the sweep reaches a card an executor is
actively running and rewrites its `pending` step results to `failed` —
the one thing that file's header says it must never do. The liveness
triple does not cover it; those legs prove an *in-process* session, and
an executor on another node or between session handles is exactly what
the column skip is for.
- **A genuine orphan on that same renamed board is still recovered** —
the skip must narrow, not disable. Passes either way, deliberately.

## What it pins, precisely

The **invariant**, not a line. Reverting either single guard still
passes, because the page-snapshot check and the fresh-row re-read
protect independently. What fails is reverting the sweep's column
handling as a whole — which is the condition worth pinning, and matches
the project's "fix the invariant, not the repro" rule.

## Still uncovered, said plainly

#3078's other two sweeps — worktree-metadata liveness and agent-link
drift — have no dedicated case. The orphaned-step-results sweep got the
test first because it is the one that can corrupt a live executor's
state. The other two remain honest debt rather than implied coverage.

## Verification

`self-healing-orphaned-pending-step-results` **10 passed** on current
main · full self-healing suites 204 · `pnpm test:gate` 161 + 13 + 487 +
71 · lint — green.
2026-07-31 04:09:48 -07:00
gsxdsm
f7a7347e1b test(core): cover #3057's cold-storage conversion; audit two archived literals as dead sync (#3089)
**Replaces #3085, which I am closing.** #3057 landed the same
cold-storage conversion while that PR was open. Rather than argue about
which spelling wins, this keeps only what `main` does not have: the
coverage, and two audits.

## #3057 converted this and shipped no test

`listTasksImpl`'s `columnFilterIsArchive` replaced `columnFilter ===
"archived"`. Correct change — and the kind that needs a test more than
most, because of how it fails.

Archived rows do not live in `tasks`; `archiveTask` copies them into the
archive store and removes them. This decision is whether that second
store is read **at all**. Against the literal, a caller naming a renamed
archive lane — `listTasks({ column: "filed", includeArchived: true })`,
which is what an archive view does — got an empty page from the only API
that can reach those rows.

Note the shape: the **unfiltered** read (`!columnFilter`) was always
correct. It fails only for the caller that names the lane, so it
survives any board-level smoke test and presents as *"the archive is
empty"* rather than as a bug. That is precisely the class that regresses
quietly once the conversion that fixed it has nothing holding it.

Three cases, and each earns its place:

| case | what it stops |
|---|---|
| renamed archive lane | the regression itself |
| legacy `archived` id (**control**) | a future conversion that resolves
the renamed lane and *drops* the legacy seed — the seeding hazard in its
other direction |
| non-archive lane (**negative**) | the widening turning every filtered
board read into a second-store round-trip |

**MUTATION**: restoring `columnFilter === "archived"` fails **only** the
renamed case.

**A trap worth recording.** My first version left the LIVE read real
against a fake `layer.db`. It threw, the `.catch` swallowed it, and all
three cases passed the negative — *including the legacy control*, which
is what exposed it. A test whose subject is never reached looks
identical to one whose subject answered no. `readLiveTaskRows` is mocked
now, and the control is what made it detectable.

## Audited, not converted — two dead sync paths

Both would be real defects if they ran. Neither runs.

| site | why it is dead |
|---|---|
| `mission-store.ts` (feature-delete link check) | `getMissionStoreImpl`
returns the AsyncDataLayer-backed `AsyncMissionStore` under PostgreSQL;
the sync `MissionStore` reached via `this.db.prepare` is legacy SQLite
only |
| `lifecycle-ops.ts` (polling-replica archive emit) |
`checkForChangesImpl` opens with `store.db.getLastModified()` /
`store.db.prepare`, which throw in backend mode |

The second is the sharper one: **both** its guard and the `to:
"archived"` it emits are literals, so a polling replica on a renamed
board would emit a move to a column the board does not declare.

Recorded in place — the treatment `project-store-ops.ts`'s dequeue twin
already has — so the census entries are not mistaken for unconverted
debt, and whoever deletes the sync SQLite residue takes these with it.

**They stay COUNTED.** Marking them DELIBERATE-LITERAL would buy a
smaller number by asserting the code is *correct*. It is not correct; it
is unreachable. Those are different claims with different expiries, and
the census should keep pointing here until the code is gone.

## Census

No movement — by design. This PR adds coverage and audits; it converts
nothing that was not already converted on `main`.

## Measured

- 3 new cases pass; `src/__tests__/{cold-storage,archive,unarchive}*` —
**6 files / 21 tests pass**
- `tsc --noEmit -p packages/core` clean; census `--strict` clean

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:06:40 -07:00
gsxdsm
6949f22ef8 fleet: resolve the same-column handoff review target (census 84 → 83) (#3076)
## Census

| | column guards |
|---|---|
| before | **84** |
| after | **83** |

`moves.ts`: 2 → 1. One of its two sites converts; the other **must
not**, and that difference is the useful part of this PR.

## Converted — the move target at the same-column handoff

```ts
if (internal.fromHandoff && toColumn === "in-review")
```

Against the literal this **never fired on a renamed board**, so a
same-column handoff into a renamed review lane silently took the *other*
branch — the sync-SQLite path, which throws under PostgreSQL.

It now asks `moveReviewColumns`: the broad membership set
(`mergeOrchestration ∪ mergeBlocker ∪ humanReview`) already resolved
**three lines above** for the merge-queue pair. Same value, so this
branch cannot disagree with the enqueue/dequeue calls that receive it.

## Not converted — the archived fallback arm

I named it, and `archived-column-gate-parity.test.ts` went red on
**`TypeScript encoding changed`**.

That guard's argument holds: the archived gate is enforced in three
encodings, the SQL halves still compare the raw string, and moving the
TypeScript half alone is the split brain it exists to prevent. Restored
inline **with a note recording the measurement**, so the next person
doesn't retry it and rediscover the same red.

This is the second time that guard has stopped me this session. It's
doing exactly what it was built for.

## On the pre-existing red

That suite is red on `origin/main` for an unrelated raw-SQL drift (#3072
fixes it — the drift is from my own merged #3042/#3046). I verified this
branch produces the **identical** failure and no other, so it doesn't
compound it.

## Measured

| check | result |
|---|---|
| moves / handoff / merge-queue suites | green |
| four gates + strict census | green |
| core `tsc` | clean |
| parity suite | same single raw-SQL failure as `origin/main`, nothing
added |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:57:51 -07:00
gsxdsm
9e242ea294 fix(engine): backlog pressure called every dependency unfinished on a renamed board (#3081)
## The third lane question

This reporter had **three** lane questions. Two were resolved when the
file's query-blindness was fixed — hold and wip, both through
`resolveProjectColumnsForRoles`. The third sat one method down and was
never touched:

```ts
if (dependency.column !== "done") return false;
```

One board, two lane answers.

## What it cost

On a renamed board every dependency reads unfinished, so
`isRunnableCandidate` rejects every card that has one. The
backlog-pressure alert then names **only dependency-free cards** as the
runnable ones.

The failure mode is the quiet kind: the report still renders, the counts
are right, and the candidate list looks plausible. The operator is told
the queue is blocked on nothing in particular. No default-board test can
see it — which is exactly why the earlier conversion of this same file,
which fixed its reads, left this behind.

## Fix

`finishedColumns` (complete ∪ archived) resolved once by the async
caller alongside hold and wip, then passed into the sync predicate.

- **Required parameter, not optional-with-a-literal-default.** An
optional parameter leaves `done` in the file as a silent fallback and
the next caller gets pre-conversion behaviour by writing nothing.
- **Archived is included** because a dependency that has been archived
is finished too — and this reporter already reads with `includeArchived:
true` precisely so archived blockers resolve.
- **Async resolution.** `resolveProjectColumnsForRoles`' only store read
is `listWorkflowDefinitions()`, a project-wide async read that works
under PostgreSQL. That is the line between a real conversion and the
inert sync-IR kind (#3058), and the new test supplies its board through
that same reader so it exercises the production path.

## Census

| | before | after |
|---|---|---|
| `backlog-pressure-reporter.ts` | 1 | **0** |

## Measured

- One new case; file **11/11 pass**.
- **MUTATION**: restoring `dependency.column !== "done"` fails it.
- The case asserts **both directions in one test** — a dependency
resting in the board's own complete lane makes its card runnable, *and*
a dependency still in the hold lane still blocks it. Asserting only the
first would pass against a predicate that had simply stopped checking
dependencies.
- The file already had a `RENAMED_IR` scoped to its second describe;
mine is a distinct `RENAMED_DEPENDENCY_IR` with different lane names. I
hit the shadowing first and the test failed as `under-threshold` — worth
noting because a same-named fixture that silently resolves to the
*other* board is precisely how a renamed-lane test goes vacuous.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Flagged, not guessed

Adjacent census entries I looked at and deliberately left:

- **`executor.ts` (4)** — all inside a sync `task:moved` listener.
Converting via `resolveTaskWorkflowIrSync` would be inert for #3058's
reason, and making the listener async reorders it against every other
subscriber. Correctly out of scope, as #3048 judged.
- **`triage.ts:724`** — half-converted in the same shape:
`disposeLanes.hold`/`.intake` come from a sync resolver, so the resolved
arms are themselves inert and "finishing" the guard would add a third
inert comparison.
- **`auto-merge-finalization.ts` (2)** — one is the resolver's
documented degraded fallback (the live arm calls `columnHasFlag`), the
other is already recorded as a deferred signature-widening whose cost
exceeds the error string it sharpens.
- **`in-review-stall.ts:196`** — an explicitly marked DELIBERATE-LITERAL
no-metadata fallback.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:51:43 -07:00
gsxdsm
3c531d984c fix(engine): self-healing lane cluster round 2 — 38 → 26 (two sweeps could disturb live work) (#3078)
The largest census cluster became **unclaimed again** when #3055 closed
conflicting. I had closed my own #3050 an hour earlier expecting #3055
to land, so this re-applies the conversions #3047 and #3049 did not
cover.

**Re-applied from current main rather than rebasing the closed branch.**
The conversions are small; the conflict archaeology is what went wrong
last time — nine conflicts against #3049, several on variable names
identical to mine, and my mechanical fixup corrupted the file badly
enough that I aborted. Starting from main cost less than resolving that
and carries no risk of resurrecting a stale line.

## Census

| | before | after |
|---|---|---|
| `packages/engine/src/self-healing.ts` | **38** | **26** |
| repo-wide column guards | 84 | **72** |

## Three sweeps, existing role helpers only

| sweep | roles | what it did on a renamed board |
|---|---|---|
| worktree metadata | terminal + wip + review | rebound finished cards
every pass, **and the FN-5256 liveness guard went silent** |
| orphaned pending step results | wip | **could rewrite `pending`
results under a live executor run** |
| agent-link drift | wip + review + terminal | evaluated agents whose
task was plainly still executing |

Two of these disturb **live** work, which is why they were worth redoing
now rather than leaving for the next fleet round:

- The worktree-metadata sweep clears `worktree`/`branch` metadata. Its
liveness guard is the thing standing between that and a running shell
(FN-5256). Keyed on ids, it matched nothing on a renamed board. The
scope-override safety condition beside it now reads the **same resolved
sets**, so the two cannot disagree about which lanes are live —
previously they were two independent literal lists.
- The orphaned-step-results sweep's own header says it must never touch
an executor-owned row. The id-keyed skip made it do exactly that.
Resolved once per sweep, outside the paging loop, so a large board still
pays one resolve.

## Flagged, not guessed — the 26 that remain

Unchanged from my earlier audit and re-verified on this base:

- **Sync predicates** (`isWorkspaceOwnerLive`, the pause-abort
classifier, the phantom-binding check, the `task:moved` listener
guards). No store handle; converting means a signature change or making
a synchronous event listener async, which reorders handlers against a
synchronous emitter.
- **Already-converted fallbacks** — `own.length > 0 ? own.includes(...)
: task.column === "in-review"`. The resolved answer wins; the literal is
the documented no-metadata path.
- **The notification-route `fresh.column === "todo"` sites** — measured
previously: any `await` before the wedge resolve drops an operator
notification. Needs the wedge-episode contract, not a column pass.

## Verification

self-healing suites **204 passed** · agent-link-drift +
query-filter-blindness **83 passed** · `pnpm test:gate` 161 + 13 + 487 +
71 · lint · census `--strict` · lane-wiring — green.
2026-07-31 03:48:44 -07:00
gsxdsm
58791fac88 fix(self-healing): route pause-abort recovery on resolved columns (self-healing 38 → 34) (#3075)
First real cut into `self-healing.ts`, the last large cluster. Converts
the **pause-abort recovery router** — a coherent unit with one owner,
rather than a scattered pass.

## Census before/after

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 86 | **82** |
| `self-healing.ts` | 38 | **34** |

## The bug this was hiding

The router keyed on three literals — `in-review` twice (review progress,
manual merge hold) and `todo || in-progress` (active work). On a renamed
board **all three stop matching**, so a parked card in a renamed lane
falls through to `no-action` and is never recovered — silently, no log
line, and with every existing test still green because they all use
legacy ids.

## Conversion

Used the file's own resolvers. `resolveReviewColumnsFor` already
existed; added `resolveActiveWorkColumnsFor` as its sibling from the
same `columnsWithFlag` / `resolveLifecycleColumns` helpers — no new
vocabulary.

**ACTIVE WORK is hold + `countsTowardWip`, deliberately NOT the intake +
hold that the neighbouring `resolvePreWipColumns` returns.** The router
asks *"is this card mid-flight, so a requeue is right?"* — an intake
lane is not mid-flight; a WIP lane is. Reusing the intake-shaped helper
would have widened the requeue to triage rows and dropped in-progress
ones. Because the two sets **overlap on `hold`**, that mistake looks
correct in every legacy-id test. This is the trap worth knowing about
for the rest of the file: it has several resolvers, and picking the
nearest one is not the same as picking the right one.

**`columns` is required, not optional-with-a-fallback.** The router has
exactly two callers — the candidate filter and the post-re-read
re-verify — and they must agree. An optional parameter lets one resolve
and the other default, and that divergence surfaces as a sweep that
selects a card and then declines to act on it, writing nothing. Required
makes it a compile error. (Same reasoning as #3059; safe here because
both resolvers union the legacy ids internally, so "required" never
means callers invent a column set.)

**On the filter/re-verify hazard I flagged earlier:** both call sites
are the *same function*, so converting it once keeps them consistent by
construction — no split-CAS risk. Per-task resolution can't hoist out of
the loop but must not read an IR per row, so the marker test
(column-independent) stays a cheap sync prefilter and the IR is read
only for rows that pass it, over a shared cache. `parked` keeps its
exact former membership, so the log count still means what it said.

## Verification

- **Revert-proof:** reverting the three guards to literals fails 2 of 3
routing cases. The third exercises the new resolver directly, so it
cannot fail on revert — stated rather than counted as evidence.
- 4 self-healing suites, **437 tests green**
- `tsc --noEmit` clean; eslint clean
- Degraded-resolution case included, since the legacy-id union is what
keeps recovery alive on an unreadable workflow

## Still flagged in this file (not guessed)

34 remain. They are not one batch: the sweeps around
L2855/3297/8877/9018 depend on the unowned `listTasks({ column })`
decision, and L5330/5359 sit under the FN-5256 liveness guard whose own
comment says those columns can be live when the heuristic calls them
stale.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:45:52 -07:00