Commit Graph

2631 Commits

Author SHA1 Message Date
gsxdsm
78d87f0a10 test(core): pin the search archive-lane WIRING — the predicate was covered, the hand-off was not (#3220)
## The false-green

#3160 (mine) proved `liveSearchPredicate` honours a resolved archive
set: hand it `Set(["archived","filed"])` and `filed` appears in the
bound params. That contract is real and still correct.

**Nothing proved `reads.ts` passes one.** It is a unit test of the
collaborator, so blinding the resolver at the call site cannot fail it.
A conversion, a test that looks like it covers it, and no connection
between them.

## The measurement — and the instrument matters

| site | vs. the predicate unit test | vs. a test that drives `reads.ts`
|
|---|---|---|
| `reads.ts:396` cold-storage list | 0 failed | **1 failed — covered** |
| `reads.ts:615` incremental sync | 0 failed | 0 failed — **UNCOVERED**
|
| `reads.ts:793` search | 0 failed | 0 failed — **UNCOVERED** |

Against `search-excludes-renamed-archive-lane.test.ts` all three read as
uncovered — an artefact of asking a file that never executes `reads.ts`.
Against `cold-storage-renamed-archive-lane.test.ts`, which drives
`listTasksImpl` for real, 396 is covered and the other two genuinely are
not.

That is rule 2 of #3214 one level up: *the test must reach the site*,
and a unit test of the collaborator never does. Had I stopped at the
first instrument I would have reported three uncovered resolvers, one of
them wrongly.

## What 793 costs on a renamed board

`searchTasks` backs the **CREATE-time near-duplicate check**. Without
the resolved lanes threaded, search stops excluding the board's archive
lane, and creating a task can be refused as a duplicate of one the
operator archived long ago — with no way to see why, because the
matching card is not on the board. Precisely the symptom #3160 set out
to fix; this pins the wiring that delivers it.

## An assertion I got wrong, and the correction

I expected an unreadable workflow list to leave `archivedColumns`
**undefined** via the call-site `.catch(() => undefined)`. It does not:
`resolveProjectColumnsForRoles` catches internally and returns its
**legacy-seeded** set, so `Set(["archived"])` is threaded and the
`.catch` never fires on that path. Two layers fail soft and the inner
one wins.

The case now asserts the guarantee that actually holds either way —
**never an empty set** (which would exclude nothing and return archived
rows in every search), legacy id always excluded. Recorded at the site,
because the mechanism is not obvious from the call.

## Flagged, not papered over

**`reads.ts:615` is left uncovered on purpose.** It composes Drizzle
conditions and runs them against `layer.db` with no injectable seam, so
pinning it needs a real database and belongs with the `.pg` suites. A
test asserting "the query was built" rather than "the rows were
excluded" would satisfy the ratchet and prove nothing.

Also flagged from this sweep: `workflow-analytics.ts` and
`team-analytics.ts` (4 resolvers) are **unmeasurable in my environment**
— their renamed-lane coverage lives in `.pg` suites, and this worktree
has no TCP PostgreSQL (`pg_isready` reports a Unix socket; the harness
probes TCP, so `pgDescribe` correctly skips). Not claimed either way.

## Census

**Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0`.** Converts
nothing; closes coverage on a conversion the census already counts as
done.

## Verification

```
as written                    Tests  4 passed (4)
BLIND reads.ts:793            Tests  1 failed | 3 passed (4)
restored                      Tests  4 passed (4)
```

Anti-vacuity case included: every other assertion reads a mock's
arguments and would pass if the search were never reached, so one case
pins that the primary search path actually ran. Typecheck clean.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:50:02 -07:00
gsxdsm
aa1655ccd9 fleet: reclassify the census tail — 10 → 2 guards, all reasoning already in the code (#3213)
## Census before / after

```
                        before   after
COLUMN guards (backlog)     10       2
DELIBERATE-LITERAL         138     148
```

Baseline re-recorded in the same commit; `--strict` green.

## This converts nothing — the tail was never backlog

All ten remaining guards already carried an explicit in-code decision.
**None carried the `DELIBERATE-LITERAL` marker the census reads**, so
each re-appeared to every fleet pass as if unexamined. That is the whole
defect this fixes.

| site | the reasoning already at the site |
| --- | --- |
| `audit-ops.ts`, `moves.ts` | the degraded fallback arm of an
**already-converted** site; the live arm uses the resolved lane set |
| `scheduler.ts` ×2 | *"LEFT COUNTED"* — an await behind the
`tracked.has` re-entrance guard lets two updates double-start a monitor;
the sibling is the measured-expensive `task:updated` emit path (26 sites
against 7) |
| `notification-service.ts` | this method and its only caller are
**sync**, reached from a listener the store invokes as `(task: Task):
void`; resolving makes the chain async and reorders notification
classification against every other `task:updated` handler |
| `lifecycle-ops.ts` | *"Recorded rather than converted"* — dead code |
| `task-id-integrity.ts` | sync, no store-scoped read; converting alone
would disagree with `getLiveTaskColumn` |
| `triage.ts` | *"LEFT COUNTED until then"* — wants a non-sync-resolved
lane answer |

## Marker placement is load-bearing, and I got it wrong twice

The census reads a node's **leading** comments. A marker in a nearby
block comment attaches to the wrong node and is **silently ignored** —
it reads as reviewed while the count still lists the site.

- `task-id-integrity.ts` — my first marker went into the block comment
above the `const`; the literal is in the `return`. Count stayed at 1
until I moved it.
- `ResearchTaskActionModal.tsx` — marker added, **measured that it did
not register**, reverted.

Every edit was verified by re-running the census, not assumed. That is
the only reason the count actually moved.

## Two sites deliberately left counted

- **`ResearchTaskActionModal.tsx`** — the literal sits mid-expression
inside a `.then()` chain, so no marker can attach. The census's own
guidance is to hoist it into a named helper; the site's note asks for
that to be someone's deliberate change rather than a drive-by, so it
stays counted and honest.
- **`self-healing.ts`** — the memo closure I converted and reverted in
#3049. Its note: a renamed board costs a duplicate log line, not a wrong
lifecycle decision.

## Correction I owe on the measurement itself

For many turns I reported "zero unclaimed guards". That came from a bug
in **my own** query — `byFile` is an array of `[file, count]` pairs and
I had switched to `Object.entries()`, which yields `[index, pair]`, so
`n > 0` was always false and the filter returned zero regardless of
state. It agreed with reality while open PRs held every file, which is
why it went unnoticed; it was still wrong, and a constant zero against a
falling backlog should have prompted me to check it sooner.

## Verification (measured)

- engine `self-healing` + `scheduler` suites — **1003 passed / 56
files**
- core `task-id` / `moves` suites — green
- `tsc --noEmit` clean in core, engine and dashboard; `eslint` clean
- `pnpm test:gate` — green
- `lifecycle-column-census --strict`, `check-lane-wiring`,
`check-sql-column-literals`, `check-fnxc-future-dates` — green

No changeset: `@fusion/core`, `@fusion/engine` and `@fusion/dashboard`
are private, and no runtime behaviour changes.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified internal annotations for archived, in-progress, and
in-review workflow states.
* Documented fallback behavior and timing safeguards across lifecycle,
scheduling, notification, and triage flows.

* **Chores**
* Updated internal lifecycle tracking baselines to reflect current
annotations and state coverage.

* **Bug Fixes**
  * No user-visible behavior changes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 11:15:40 -07:00
gsxdsm
8661b739ff fix(scheduler): a board with TWO complete columns left dependents waiting forever (#3210)
## The defect

On a board declaring more than one complete-trait column — a merged lane
and a shipped lane, say — a card landing in the **second** one was never
recognised as finished, so nothing unblocked its dependents. Silent: no
error, the dependent just waits.

Two problems, the same shape:

1. **`TaskMoveLanes` carried one id per role.** That is right for
*"where should this card go"* and wrong for *"is this column one of the
finished lanes"*, which is a **membership** question. The payload could
not express such a board at all.
2. **`mergeParkedColumns` rebuilt `terminal` as `new Set([complete,
archived])`** — discarding `base.terminal`, which the sync IR path had
already resolved correctly, and narrowing a membership set back to
first-match-per-role.

Point 2 contradicted the note sitting directly above it in the same
file:

> `terminal` is a MEMBERSHIP set, and it is not the same question as
`complete`/`archived`. […] A workflow may declare more than one
complete-trait column […] and `to === parked.complete` sees only the
first and silently skips the rest.

The reasoning was already written down. The overlay added later didn't
honour it.

## Fix

`TaskMoveLanes.terminal?: readonly string[]`, filled from
`columnsWithFlag(ir, "complete"|"archived")` rather than the first-match
`resolveLifecycleColumns`, and the merge is now a **union** of base,
payload, and the single lanes.

Optional, so all 12 emitters and every listener keep compiling — a
listener that ignores it is exactly as correct as before. Union is the
direction `scheduler.ts` already argues for at line ~422: a superset
costs one extra query; a subset **silently withholds work from a
finished card**.

`complete` deliberately stays first-match — a set would be the wrong
shape for a move *target*. Both questions now coexist rather than one
replacing the other.

## How it was found, and what it corrects

Supplying `task:moved` lanes fixed every *other* renamed-board case in
`scheduler-renamed-hold-events` — measured **10 passed / 1 failed** —
and left exactly this one broken. That same measurement is why I
narrowed my earlier claim on #3082: the other behaviours were never
broken in production, because all 12 emitters already carry lanes. This
is the residue that was genuinely broken.

## Tests — both with anti-vacuity controls

| control | result |
|---|---|
| revert `toTaskMoveLanes` | **2 of 4** core tests fail (the terminal
pair) |
| revert the scheduler union | the new engine test fails, **and only
it** (1 failed / 11 passed) |
| both restored | 4 passed, 12 passed |

The 2 core tests that pass either way are shape invariants asserted on
purpose (`complete` must stay first-match; a column-less IR returns
`undefined` rather than an invented lane) — flagging that so the control
isn't read as 4-of-4.

The pre-existing scheduler case emits **without** lanes, which no
production emitter does, so it exercises the sync fallback. The new one
emits `toTaskMoveLanes(ir)` — the shape that actually ships.

## Measured

| check | result |
|---|---|
| `@fusion/core` / `@fusion/engine` tsc | exit 0 / exit 0 |
| eslint | clean |
| `census --strict`, `check:fnxc-future-dates`, `check:changesets` |
exit 0 |
| every `TaskMoveLanes` consumer | 24 passed |
| `pnpm test:gate` | **exit 0 — 744 tests, up 12** |

## Census

No guard converted; this is a payload-shape fix. Backlog unchanged at
11, all deferred.
2026-07-31 10:47:15 -07:00
gsxdsm
bad39e2ca3 test(core): ledger the legacy-id collections that gate a live column — the class the census cannot count (#3209)
## What

A population ratchet over **legacy-id collections consulted against a
live column value** (`SOME_SET.has(task.column)`), recorded as 23 sites.

## Why

The census scans `===`/`!==` comparisons. A Set or array literal is a
**definition**, so no census run has ever pointed at one. Three
found-by-hand defects came from that blind spot:

| collection | symptom |
|---|---|
| `GITHUB_TRACKING_EDITABLE_COLUMNS` | operator could not toggle GitHub
tracking **at all** on a renamed board — no error, affordance absent
(#3149) |
| `TIME_INDICATOR_COLUMNS` | wrong elapsed-time indicator on cards |
| `BLOCKER_ESCALATION_COLUMNS` | escalation skipped renamed lanes |

#3149 enumerated the population by hand and concluded *"this is where
the remaining renamed-board defects actually live."* A number in a PR
body rots. This is that enumeration as a ratchet.

## Census

**Unchanged — `AVAILABLE: 0` before and after, 12 documented deferrals
both sides.** This PR converts nothing. It ratchets a class the census
*structurally cannot see*, which is the point: the backlog reading zero
has never meant the lane vocabulary is fully converted, only that the
measurable part is. Recording that plainly instead of claiming a delta
this change does not produce.

## What it claims, and what it deliberately does not

It claims the **population** is the recorded set. It does **not** claim
each site is correct — 20 of the 23 are #3149's assessment ("most are
already correct, either no-flags fallbacks or seed-then-add resolved
sets"), and I did not re-verify them. Blessing sites I have not read is
how a ledger becomes a list of things someone once glanced at. A new
entry fails the test and a human reads that **one** site; that is the
entire mechanism.

## My own detector's pick-work list was 100% false positives

Measured, and the reason this ships with **no candidate list**. The
heuristic "no role-helper call in the file" flagged three sites; all
three were fine:

- `agent-role-policy.ts:32` — a documented **FLAGGED, NOT FIXED**
deferral with its reasoning recorded
- `DocumentsView.tsx:88` — already converted, flags-first; the flags
arrive as a threaded **object**, so a scan for resolver *calls* cannot
see the conversion
- `agent-assignment.ts:118` — a `DELIBERATE-LITERAL` fallback behind an
injected `countsAsAssignmentLoad` callback, reviewed `2026-07-31-05:40`

That is the same failure `--triage`'s pick-work list had before #3194
fixed it, from the same cause: **inferring "unexamined" from the absence
of a pattern rather than from evidence.** A detector that cannot
distinguish "not yet looked at" from "looked at and settled" must not be
pointed at a work queue. It can still hold a population steady, which is
all this does.

## Verification

Mutation-verified in **both** directions — a ledger fails by missing
additions *or* by keeping ghosts:

```
### baseline                                        Tests  4 passed (4)
### MUTATION 1 — new unrecorded gating collection
+   "packages/engine/src/worktree-pool.ts :: NEW_LANE_GATE",
                                                    Tests  1 failed | 3 passed (4)
### MUTATION 2 — recorded site vanishes (ghost)
+   "packages/engine/src/worktree-pool.ts :: managedRenamed",
+   "packages/engine/src/worktree-pool.ts :: managed",
                                                    Tests  2 failed | 2 passed (4)
### restored                                        Tests  4 passed (4)
```

Two anti-vacuity cases guard the detector: it still finds the
collections whose defects motivated the file, and it does **not** claim
plain comparisons (asserted against `self-healing.ts`, dense with column
comparisons and no gating collection) — pulling those in would
double-count a class that already has a gate.

## Flagged, not guessed

- **Line numbers are excluded** from ledger entries — they drift with
unrelated edits and would fail this test for reasons that are not about
lane vocabulary.
- **Comments stripped before scanning:** `TaskDetailModal.tsx` and
`TaskCard.tsx` both quote their own collection by name in FNXC notes
explaining the bug it caused. Counting prose would fire the ledger on
the files that document the hazard most carefully.
- **Stated reach limits** (in-file): only *named* collections consulted
as `.has`/`.includes`; the argument must mention column/lane;
property-reached collections are missed. A miss is a site nobody is
watching — not a false green on a listed site.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:44:08 -07:00
gsxdsm
5c5f6d8155 fix(core): mark the two archived STATE sites at the site — converting them destroys live work (#3157)
The LANE/STATE triage (#3154) found that **two of the eight** Drizzle
`archived` sites are STATE markers that must never be resolved. That
classification lived only in `archived-column-gate-parity.test.ts`.

A coordinated three-encoding conversion **edits these files**. A
converter working file-by-file sees the same `eq(column, "archived")`
shape as the six LANE sites, with nothing in front of them to tell the
two apart.

So the markers go at the sites.

## `task-mutation-ops.ts` — `cleanupArchivedTasksImpl`

Enumerates rows Fusion itself archived, then **removes their
directories**.

Widening it to the resolved archived-lane set would feed cards **merely
resting in a board's archived-trait lane** into a filesystem delete.
This is the only site in this family where a wrong conversion **destroys
work** rather than hiding an affordance.

## `async-self-healing.ts` — `listSoftDeletedColumnDriftCandidates`

Finds soft-deleted rows whose column **drifted** from the marker they
are supposed to carry. Resolving it would classify a soft-deleted row in
a renamed archive lane as drift and "repair" a row that is already
correct.

## Why this is defensive rather than cosmetic

The triage exists to make the conversion safe. A classification the
converter **cannot see while editing the file** does not do that — it
only helps someone who happens to read the gate's test file first, which
is not how a file-by-file sweep proceeds.

Both are marked DELIBERATE-LITERAL with the reason and a pointer to the
parity test holding the full eight-site split.

## Measured

- Comment-only.
- Parity test **2/2**; archive + soft-delete suites — **5 files / 15
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict` and
`check-sql-column-literals` clean.
- **No census movement** — a DELIBERATE-LITERAL marker on a STATE site
is a classification, and these were never counted as lane debt.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:30:41 -07:00
gsxdsm
dd3bf6a764 docs(core): the archived TS remainder is EMPTY — enumerated, closing the triage (#3171)
#3156 sampled the TS inventory and said the conversion is *"the six LANE
Drizzle sites plus whatever small TS remainder is neither a fallback arm
nor a sentinel"*.

**That remainder is zero.** I left three sites unchecked when I wrote
it. All three are fallback arms:

| site | shape |
|---|---|
| `live-agent-count.ts:164` | `task.columnTerminalKind ?? (task.column
=== "done" ? … )` — the resolved value wins via `??` |
| `store.ts:1972` | `if (!lanes) return dep.column !== "done" && …` — an
explicit no-metadata branch |
| `branch-and-pr-entities.ts:525` | `lanes === undefined ? task.column
=== "archived" : task.column === lanes.archived` |

With the previously classified entries, **every** site in
`AUDITED_TS_SITES` is now accounted for as a fallback arm, a
STATE/sentinel comparison, or a converted guard's retained literal.
**None is an unconverted LANE guard.**

## So the cluster is done

"52 sites across three encodings" is fully triaged, and the convertible
work is the six Drizzle LANE sites plus the log-entry gate — all
additive, so the inventories never moved.

What the gate now protects is a population of **fallback arms and STATE
markers**, which is exactly what it should protect: each is the
documented answer for a caller that supplies no resolved set, or a
marker that must never be resolved.

A future **drop** in any of the three counts means someone removed a
fallback or converted a STATE site — both regressions. That is the check
this file was built to make, and is now the only check it needs to make.

## Enumerated, not sampled

I sampled this inventory twice and each pass changed the size estimate —
first "52 sites, real blast radius", then "six plus a small remainder".
A third estimate would have been worth less than a complete count, so
this pass covers every entry.

That is the honest close: the number stopped moving because I stopped
guessing at it.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean. **No
census movement.**

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:30:30 -07:00
gsxdsm
230be28576 fix(core): the merge-queue enqueue guard was not debt — the code it guarded had no callers (#3205)
## The deferral note was right about the mechanism and wrong about the
remedy

`merge-queue-ops-2.ts` sat in the census as deferred debt behind this
note:

> Converting it properly means either making this path async or pushing
the trait read into SQL, both of which are store-architecture changes
rather than call-site conversions.

That is correct as far as it goes — the guard runs inside
`store.db.transactionImmediate`, so the only synchronous resolver
available (`resolveTaskWorkflowIrSync`) returns the DEFAULT workflow
under PostgreSQL and a "conversion" would be inert.

But it assumed the code needed converting. Measured across the tree:

```
=== every call site of .enqueueMergeQueueSyncInternal( ===
packages/core/src/store.ts:1775:  public enqueueMergeQueueSyncInternal(...)   <- the declaration itself
```

**Zero callers.** Every other occurrence of the name is a comment. The
live path is `enqueueMergeQueueAsync` (`task-artifacts-ops.ts:117`), and
that file already documented the deletion:

> Merge-queue enqueue is PostgreSQL-only via enqueueMergeQueueAsync …
The SQLite `enqueueMergeQueueSyncInternal` arm is deleted.

The arm was deleted; its declaration was not. The guard was unreachable
on the shipped backend.

## Change

- Deleted `enqueueMergeQueueSyncInternalImpl` (-85 lines) and its
`store.enqueueMergeQueueSyncInternal` entry point.
- Dropped the six imports that became unused
(`MergeQueueTaskNotFoundError`, `MergeQueueInvalidColumnError`,
`MergeQueueEntry`, `MergeQueueEnqueueOptions`, `normalizeTaskPriority`,
`MergeQueueRow`).
- Refreshed the three comments naming the removed symbol, so none points
at a deleted identifier. The
`handoffMergeQueueFailureInjectorForTesting` hook those comments sit on
is a **different** member and is untouched — it only mentioned the sync
arm as context.

## Census before / after

| | before | after |
|---|---|---|
| `packages/core/src/task-store/merge-queue-ops-2.ts` | 1 | **0 (entry
removed)** |

Baseline tightened by exactly one entry. **The 0 here is a deletion, not
a conversion** — recorded in the file's own FNXC note so the next worker
does not read it as a converted seam. This is the failure mode the
census warns about ("a count of 0 is the WORST case, not the best"), so
it is stated at the site rather than left to inference.

## Measured

| check | result |
|---|---|
| `census --strict` | exit 0 |
| `@fusion/core tsc --noEmit` | exit 0 |
| `eslint` (4 changed files) | clean |
| core merge-queue tests | **110 passed / 6 files**, incl.
`postgres/merge-queue-renamed-review-column.pg.test.ts` |
| `pnpm test:gate` | exit 0 (**732 tests**) |

No changeset: `@fusion/core` is private and this removes unreachable
code with no user-visible behavior.

## Flagged, not guessed

The other four deferral-note files remain deferred. I only reclassified
this one because its call-site count is a fact I could measure, not a
judgement. Whether `lifecycle-ops.ts:667` is likewise dead (it sits in
the legacy-SQLite polling-replica path) is a separate question I have
not measured, so I have not touched it.
2026-07-31 10:30:16 -07:00
gsxdsm
25fa5e7c44 fix(core): the log-entry archive gate, converted — the parity objection is met, not bypassed (#3165)
I converted this in #3110, the parity gate failed, and I reverted it.
**The gate was right** — and the reason was subtler than "one encoding
moved". Knowing it is what makes this conversion possible.

## Why the first attempt failed

My version hoisted the comparison onto a local:

```ts
const pgRowColumn = String(pgRow.column ?? "");
const rowIsArchivedLane = archivedLanes ? archivedLanes.has(pgRowColumn) : pgRowColumn === "archived";
```

That gate's TS scan keys on the **property** being named `column` —
deliberately, because the receiver is variously `task`, `row`, `dep`,
`t`. Losing the `.column` access dropped the TS count while SQL and raw
held steady, which it reads as divergence.

**Behaviourally identical, structurally invisible.** Same failure mode I
hit from the other direction in #3163, where I collapsed a Drizzle
fallback into a string array.

## The fix

Keep `pgRow.column === "archived"` **verbatim** as the fallback; add the
resolved path in front of it. No encoding's count moves, an unwired or
degraded caller behaves exactly as before, and the gate is **satisfied
rather than worked around** — the same additive shape as the six Drizzle
LANE sites (#3160, #3162, #3163).

## What it fixes

A LANE question: *"is this row in the board's archive lane, so logging
is read-only?"*

Against the literal, a card the operator filed away on a renamed board
kept **accepting log writes** — new activity accruing on closed work.
`deletedAt` covers the soft-delete half, which is why the gap is narrow
and why it stayed invisible: the common path is soft-delete.

## The recorded omission is retired properly

`log-entry-archived-lane-gate.test.ts` carried the renamed case as a
**deliberate omission** with its reason. It is now the first case in the
file, and the note explains why the earlier judgement changed rather
than quietly disappearing — a deferral that vanishes without explanation
is how the next reader loses the thread.

## Measured

- **3/3** in that file (renamed case added); parity test **2/2**,
inventories unmoved.
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control** and the live-lane **negative** green.
- log-entry / archived / audit suites — **3 files / 9 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals` clean.
- `check-fnxc-future-dates` is red on `main` from `task-update.ts`
(another lane's stamps), not from these files.

## Census

**Unchanged** — the literal remains the fallback arm, by design.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:27:31 -07:00
gsxdsm
762d232ad6 test(core): ratchet the sentinel-task-id argument at zero — the third inert-conversion mechanism (#3204)
## What

A test-only zero-population ratchet: no task-scoped lane resolver may be
called with a **string literal** where a row id belongs.

## Why

`resolveTaskLifecycleColumns(store, taskId)` and its siblings resolve
the workflow bound to *that task*. Hand one a literal and there is no
task to read a selection for, so the resolver falls back to the
**default board** and answers with full confidence. The call
type-checks, reads as a finished conversion, and is correct on every
board Fusion ships — because the default board is the answer it returns.

**This shipped.** `triage.ts`'s startup sweep called
`resolvePlannerLanes(this.store, "")` and built its swept-column set
from the result (#2806 measured it, #3201 fixed it). It was a
*sweep-wide* defect rather than a per-card one: it resolved once for the
whole board and could not be right for any workflow but the default, so
a card parked in a renamed hold column with a stale `planning` status
was never swept and held a planning admission slot permanently.

Note what this means for the other two inert mechanisms' fixes —
**making the resolver async would not repair it**, because the defect is
the argument, not the resolver.

## Why a guard and not just the existing E2E

`workflow-sweep-sentinel-task-id-live-e2e.pg.test.ts` covers the **one**
triage site and lives in the `.pg` lane, so it is skipped whenever no
PostgreSQL is reachable — including the merge gate. The defect is the
*argument*, which makes it visible in source text with no database, no
running engine, and no knowledge of what the resolver does.

## Census

**Unchanged — 0 guards before, 0 after.** This PR converts nothing; it
is a ratchet over a class the census structurally cannot see (the census
scans column literals, not resolver arguments). Recording that plainly
rather than claiming a delta this change does not produce.

Population of the guarded class is **zero today** — the only textual
match in the tree is prose in `triage.ts` documenting its own fixed bug.
A zero-population ratchet is the instrument here, not a weakness: it
cannot fail until someone reintroduces the defect, and it costs one
source scan.

## Verification

**Mutation-verified, not asserted.** Re-adding the exact shipped shape
to a real production file:

```
+   "packages/engine/src/replan-target.ts:188 — resolvePlannerLanes",
 Tests  1 failed | 3 passed (4)
```

Restoring the file returns it to `4 passed`. Working tree left clean.

Three anti-vacuity cases carry the file, because a scan that reports
success by finding nothing is otherwise indistinguishable from a broken
scanner:
- the matcher **does** fire on the historical text
(`resolvePlannerLanes(this.store, "")`);
- it does **not** fire on the ordinary shapes that fill the codebase
(`task.id`, `taskId`, `row.id`) — a matcher flagging everything would
pass the case above while being unusable;
- the walker still reaches production source (>20 real
`resolveTaskLifecycleColumns` call sites), which is what makes the zero
a measurement rather than an empty scan.

## Flagged, not guessed

- **Comments are stripped before scanning**, and here that is required
rather than tidy: `triage.ts` quotes the offending call verbatim to
explain the hazard. Counting it would make the guard fire on the file
that correctly documents the defect, training readers to silence the
guard instead of heeding it.
- **`resolveReboundTarget(ir)` / `resolveLifecycleColumns(ir)` are
deliberately excluded** — they are IR-scoped and take no task id;
including them would flag correct code.
- **Known limit, stated in the file:** a sentinel arriving through a
*variable* (`const id = ""; resolve(store, id)`) is invisible to a text
scan. The literal form is what shipped and what the next person is most
likely to write; the variable form still needs the `.pg` E2E. Two
instruments, different reach — not full coverage of the class.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:27:18 -07:00
gsxdsm
d6079970e8 fix(self-healing): 18 recovery rebounds hardcoded todo and THREW on a renamed board (#3150, first slice) (#3152)
First slice of #3150. `self-healing.ts` held **26** `moveTask` calls
with a legacy literal target; this converts the **18 `todo` rebounds**.

## Why this is worse than a guard, and documented already

`task-store/moves.ts` records it from a previous incident:

> `moveTaskInternal` **REJECTS** a target the workflow does not declare
(`TransitionRejectionError: unknown-column`) … completion handoff did
not silently no-op — it **THREW**.

Every one of these 18 is a **recovery**. On a renamed board they threw
instead of rebounding, so the strand each sweep exists to clear survived
*and* the sweep reported failure. The reliability layer meant to be the
backstop was the layer that broke.

## Why the census never saw it

It counts **comparisons** against legacy ids. A move target is an
**argument**. That is the third blind spot of the same instrument, and
all three have now produced real defects found by hand:

| blind spot | found this session |
|---|---|
| definitions | `GITHUB_TRACKING_EDITABLE_COLUMNS` — tracking
unreachable on renamed boards (#3149) |
| collections | swept: 30 sites, 29 already correct, 1 defect (the
above) |
| **targets** | **this** — 26 in one file, 31 tree-wide |

## Why 18 sites at once is safe

`resolveReboundTargetForTask` **degrades to `"todo"`** when no workflow
resolves, and `self-healing.ts` already used it at line 745. On every
board we ship, the resolved answer *is* `todo` — so default behaviour is
unchanged **by construction**, not by inspection. The control case pins
exactly that, and it is the reason this can land as one change rather
than eighteen.

## Scope, and what I deliberately did not touch

Converted: the 18 `todo` rebounds.

**Not** converted: the `done`, `archived` and `in-review` targets. They
need different helpers and genuine reasoning about which lane a
completion or an archive belongs in — converting them by analogy is
exactly the half-conversion this program keeps paying for. Sites with no
resolver in scope are unchanged.

The audit behind the split is in the commit: of 26 sites, 5 had resolved
lanes in scope, 4 had an IR, 17 had nothing — and `lanesOfReclaim`
returns **Sets**, which is the wrong arity for a target (a move takes
exactly one column, per the `moves.ts` note).

## Verification

| | result |
|---|---|
| engine `tsc` | **0 errors** |
| **all 43 self-healing suites** | **843 passed** |
| census `--strict` | exit 0, **unchanged** — invisible to it |
| `check-inert-sync-lanes` | exit 0 |
| differential | restoring the literal → **1 failed \| 1 passed**,
renamed case only |

The new test drives a **public entry point**
(`reconcileInReviewUnmetDependencies`, the FN-6793 contract) rather than
calling the helper directly, so it covers the producer path too.

One harness note worth keeping: the first version of the test failed
**upstream** of the target, because the sweep selects rows via
`resolveProjectColumnsForRoles` — a *project-level* resolver reading
`listWorkflowDefinitions`, not the task's own selection. Without that
mocked, the renamed card was never considered and the failure looked
like the fix not working. That distinction (project-level vocabulary vs
per-task IR) will bite the next slices too.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks now move to workflow-specific rebound, completion, and archive
columns instead of fixed default destinations.
* Retrying and recovering tasks works correctly on boards with renamed
lifecycle columns.
* Added safe fallback behavior for workflows without custom lifecycle
settings.
* **Tests**
* Added coverage to prevent legacy hardcoded task destinations and
verify renamed-column recovery scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:22 -07:00
gsxdsm
4c62124589 fix(core): the last four archived LANE sites — the ones that already held a store (#3163)
Completes the six **LANE** sites from the triage (#3154). #3160 and
#3162 did the two predicate builders that needed threading; these four
already had `store` in scope, so each is a resolve-and-spread at the
site.

## What each fixes on a renamed board

| site | defect |
|---|---|
| `store.ts` revert lookup | a done/archived prior undo attempt kept
surfacing as an **open** undo task — the store-side twin of the
dashboard defect fixed in #3129 |
| `branch-group-ops.ts` | the near-duplicate marker cleanup found **no
live rows at all**, so stale markers survived. Its own header says stale
markers alter operator decisions |
| `branch-and-pr-entities:438` | the CREATE-time fingerprint duplicate
guard kept archived cards in the candidate set — a new task could be
refused as a duplicate of one already filed away |
| `branch-and-pr-entities:470` | recent-sibling lookup counted finished
siblings as candidates |

## The parity gate caught my first version — and it was right

I collapsed the fingerprint fallback into a **string array**
(`["archived"]`) and pushed `ne(col, lane)` in a loop. Behaviourally
identical, and it **dropped the Drizzle encoding's literal count**,
because the gate scans for the `ne(..., "archived")` *expression shape*.
TS and raw held steady, so the encodings diverged — precisely what that
gate exists to catch, catching it.

The fix: keep every fallback as a literal `ne(..., "archived")`
**expression** rather than data.

That is what makes these conversions **additive** — the resolved path is
added, the literal stays, no encoding's count moves, and an unconverted
board builds byte-identical SQL. Same property as #3160/#3162, now with
a demonstrated failure mode for getting it wrong. Worth knowing for
whoever does the remaining TS remainder: *behaviourally identical* is
not sufficient; the shape has to survive too.

## Measured

- Parity test **2/2**, inventories unmoved — the point.
- archived / branch / near-duplicate / merge-blocker suites — **6 files
/ 46 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — literals remain as fallback arms, by design.

## Where the cluster stands

All **six LANE** Drizzle sites are now converted (#3160, #3162, this).
The **two STATE** sites are marked in place and must never be converted
(#3157). What remains is the small TS remainder that is neither a
fallback arm nor a sentinel, identified in #3156.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:07 -07:00
gsxdsm
a6d67844b8 fix(triage): the "unconvertible" site was convertible — the blocker was two test harnesses (#3191)
#3141 measured this site as unconvertible, and I twice reported the
cause as a production constraint. It was not. This is the instrumented
answer to the probe I recommended there and then ran myself.

## The isolation

| configuration | result |
|---|---|
| flag only, no conversion | **8 passed** → the orphan arm is *not* the
cause |
| flag + conversion | **5 failed** → the conversion is |
| same, with a realistic mock store | **8 passed** → the mock was the
cause |

`triage-stuck-requeue-preserve-draft.test.ts` defined neither
`getTaskWorkflowSelection` nor its async twin — exactly like
`triage.test.ts` did before #3189. Both made
`resolveWorkflowIrForTaskWithProvenance` **throw** and take its catch
branch: the *"could not ask"* shape, which a production store never
presents.

So the 5 failures I deferred as a possible semantics change were the
same harness gap in a second file — confirmed, not argued.

## What changes

**`selectionAbsent`** marks the determinate case: the store *answered*
"no selection", so the workflow is the default and its IR is in hand.
Added as a **separate field, not a third `source` value** — `source ===
"default"` is compared in **31 places** in `self-healing.ts` meaning "be
conservative", and a new enum value would silently stop matching every
one of them while still compiling and still passing on a default board.

**`recoverApprovedTask`** now accepts a legacy `triage` row *explicitly*
(its workflow does not declare that column) instead of depending on
`resolvePlannerLanes` **failing** and falling back to legacy ids.
Correctness resting on a resolver's failure mode is what this removes.

## Measured

| | result |
|---|---|
| broad suite (triage / self-healing / recovery / planning) | **77
files, 1302 tests passed** |
| the three directly affected suites, post-rebase | **245 passed** |
| the flag is load-bearing | conversion **without** it: **18 failed \|
221 passed** |
| `census --strict`, `check-fnxc-future-dates` | exit 0 |

**The inert-sync-lane count is unchanged at 7 for `triage.ts`.** This
site was never among the counted guards, so this is **not** a ratchet
reduction — stating that rather than letting a conversion imply one. It
removes a real inert dependency the ratchet cannot see, which is the
blind-spot class this phase has been mapping.

## Why this took four attempts

I described this blocker at four levels: merged intake/hold, orphan-arm
scoping, identity verification (filed as **#3187**, closed as wrong),
and finally the harness. **The two I instrumented held; the two I
reasoned to did not.** The fix here is the probe I wrote down for
someone else — which is where it should have started.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 09:36:03 -07:00
gsxdsm
6bc90ccbe2 fix(core): allow-list the legacy workflow IR — it found a fourth bug my grep missed (#3185)
## The name is the defect

`BUILTIN_CODING_WORKFLOW_IR` reads like the default and **is** the
legacy workflow (`builtin:legacy-coding`). Post-U11 they differ by
exactly one column — `triage` — the one a caller most often wants
absent.

**Four bugs have come from reaching for it by name:**

1. two move-path resolvers disagreed on the no-selection default →
*"workflow move policy preflight is stale"* on every flag-on move
(recorded in `resolveDefaultWorkflowIr`'s own header)
2. the TUI board rendered a `triage` lane the default board lacks —
#3178
3. `deleteWorkflow` re-homed occupants into `triage` — #3183
4. **`board-workflows.ts`** described a *custom* workflow whose
definition failed to load using legacy columns — the #3178 symptom
through the dashboard route. **Fixed here.**

It type-checks, it is the obvious identifier, and on the five shared
columns it behaves correctly. The mistake only shows on the column that
differs.

## I said the sweep was complete last round. It wasn't.

My grep excluded paths and truncated at `head -10`; it missed two sites.
**The allow-list found both on its first run.**

That is the lesson the sibling sync-resolver ratchet already records —
*"FOUND BY THIS RATCHET, not by the grep that seeded the list"* — and I
had just quoted that file while repeating the mistake.

## One site is allow-listed rather than fixed, and I tried the fix first

`workflow-graph-executor.run()`'s default `ir` is unreachable in
production (both callers pass it explicitly). But
`workflow-graph-executor-parity.test.ts`, in the **engine-core gate
suite**, drives the method *without* the argument to assert the
historical seam sequence.

Switching it to the catalog default rewrites what "parity" means:
**measured, 6 gate tests fail** with `expected 'failure' to be
'success'`. Reverted, and recorded at the call site *and* in the
allow-list entry so nobody repeats the experiment.

That is what an allow-list is for: a legitimate narrow use next to a
plausible-looking wrong one.

## Guard construction

Follows the repo's existing call-site allow-lists (sync resolver, engine
blocking-shellout, detached-spawn script guard).

- **Comments stripped before scanning** — `activity-analytics.ts` and
`TaskContextMenu.tsx` name this constant in notes *about past bugs*
while correctly avoiding it. Counting prose would train readers to
allow-list mentions.
- **Anti-vacuity**: the scan still sees the catalog's own uses, so a
renamed constant or broken walker cannot make the guard pass by finding
nothing.
- **Stale-entry**: the list cannot rot into files that no longer touch
it — the decay every ledger in this repo has hit.

## Measured

- Guard **3/3**; `tsc --noEmit` clean in core, engine, dashboard.
- census `--strict`, `check-fnxc-future-dates` clean.

## Census

**No movement — that is the point.** This class has no column literal to
count, which is why the census never saw any of the four bugs.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:45:31 -07:00
gsxdsm
757ce71731 fix(core): deleting a workflow re-homed its cards into triage, a lane the default board lacks (#3183)
A **third door** into the drift #3178 just fixed in the TUI. Found by
looking for siblings of that bug — **not** by the census, which
structurally cannot see this class: there is no column literal here to
count. The wrong answer comes from reading the wrong IR.

## The bug

`deleteWorkflow` clears each occupant's selection so they fall back to
the built-in default, then re-homes them to *"the default workflow's
entry column"* — its own comment's words.

It read that entry column from `BUILTIN_CODING_WORKFLOW_IR`, which is
`builtin:legacy-coding`, **not** the catalog default. Post-U11 the two
differ by exactly the column this reads:

```
default  todo, in-progress, in-review, done, archived
legacy   triage, todo, in-progress, in-review, done, archived
```

**Measured, not inferred:** `resolveEntryColumnId` answers `triage` for
the legacy IR and `todo` for the default.

## Why it got past the guard built for exactly this

`moveTask` rejects a target the workflow does not declare — **except**
under `recoveryRehome` with a **legacy id**, the #1411 escape hatch that
keeps a custom-workflow card rescuable.

`triage` *is* a legacy id. So the rehome slipped through the check that
exists to stop this, and left the card in a lane its new workflow has no
node for — the undeclared-column state other reconcilers exist to
repair.

## Measured

- 3 new cases; **MUTATION**: restoring the legacy constant fails the
anti-vacuity case.
- The first two cases pin the two IRs' entry columns as **facts in the
suite** rather than claims in a comment — that difference is the entire
reason the bug existed. They go quiet, correctly, if the IRs ever
converge again.
- The third pins the **call site**, because the first two would keep
passing against the unfixed code: they describe the IRs, not the caller.
That gap is how an anti-vacuity case earns its place.
- `src/__tests__/workflow*` — **30 files / 402 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**Unchanged — and that is the finding.** This defect has no literal to
count. `builtin-workflows.ts` already records the move-path resolvers as
fixed for the same drift, #3178 fixed the TUI, and this is the third
instance. The census measures *comparisons*; a surface that resolves the
**wrong workflow** produces identical-looking code and a wrong answer.

If there is appetite for a next sweep, that is where I would point it:
sites that resolve a workflow at all, rather than guards that compare a
column.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:34:04 -07:00
gsxdsm
2cb5cab595 chore(core): mark the mission-store dead-sync-path literal DELIBERATE (census 13→12) (#3179)
Comment-only. `tsc` 0 errors, `census --strict` and
`check-fnxc-future-dates` exit 0.

## Claimed with the new tool

First use of `scripts/check-file-claimed.mjs` (#3175) to pick work
instead of guessing:

```
CLAIMED    packages/engine/src/scheduler.ts        #3177, #3142
CLAIMED    packages/core/src/task-store/audit-ops.ts   #3165
UNCLAIMED  packages/core/src/mission-store.ts
UNCLAIMED  packages/core/src/task-store/task-id-integrity.ts
```

Two of the four files I would have reached for were already taken — by
PRs whose branch names give no hint they touch those paths. That is the
collision this phase paid for five times, answered in one command. I
took `mission-store.ts`; `task-id-integrity.ts` is still free.

## Census 13 → 12

Reclassification, not conversion — the line is unchanged.

## Verified the blocker rather than deferring to it

The site carries an audited note: the sync `MissionStore` reaches
`this.db.prepare`, and `getMissionStoreImpl` returns the
`AsyncDataLayer`-backed `AsyncMissionStore` under PostgreSQL, so the
class is unreachable in the shipped backend.

I checked that independently instead of accepting it —
`async-mission-store.ts:168` states the same routing from the other
side. **That check exists because of #3129**, where a note I had
accepted as settled ("blocked on a per-neighbour flag map that does not
exist") turned out to name the wrong variable, and the file was
convertible all along. I had publicly argued it should stay counted.

So the rule I am applying: a documented blocker gets marked only after
its named obstacle is confirmed from a second source. Here it held; on
`taskRevert.ts` it did not.

## Related, and still open

`merge-queue-ops-2.ts` carries a note of the same shape that does
**not** survive this check — it names two ways to convert (make the path
async, push the trait read down) and misses the one that worked twice
this phase: thread the resolved lanes in from a caller that already
awaited them, as #3112 and #3118 did for `executor.ts`. Its sibling
`taskStillInReview(projectId, reviewColumns)` already takes lanes from
its caller. Worth a real look rather than a marker.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:13:45 -07:00
gsxdsm
2a57820dd2 chore(gate): normalize the last future-dated stamp in task-update.ts (tightens the allowance 1 → 0) (#3168)
**Main is red on the FNXC gate again** — third occurrence of this class
today, third different file.

```
packages/core/src/task-store/task-update.ts: 2 future-dated FNXC stamp(s), baseline allows 1
```

Stamps dated **2026-08-01** while UTC is **2026-07-31-14:24**. Every
open PR inherits the failure; #3164 merged carrying it.

## The fix

Date only, to today. Clock times preserved exactly — they were real
times on the wrong day — and no comment text touched, so the record
reads identically, just in order:

```
-FNXC:StateMachine 2026-08-01-10:20 (PR #2793's finding — the INNER half, merged with #2821):
+FNXC:StateMachine 2026-07-31-10:20 (PR #2793's finding — the INNER half, merged with #2821):
```

Baseline **tightened** as a side effect (`1 → 0`): one future stamp was
grandfathered, normalizing the file cleared it too, and the gate refuses
a stale allowance on the way down. Re-recorded in the same commit.

## The recurrence is the point, not this fix

Three separate files have tripped this in one day — `scheduler.ts`, the
scheduler PG test, and now `task-update.ts` — plus the midnight-rollover
variant this morning that reddened everyone's baseline.

**Stamps are written from a local clock and validated against UTC.** A
worker behind UTC writes what is genuinely "today" for them and produces
a future stamp the moment UTC has already rolled. Nothing in the local
loop catches it: `pnpm lint` passes locally because the local date
agrees.

The durable fix is to generate the stamp from `date -u` rather than a
wall clock — one line in whatever produces these, and the class
disappears. I have patched the symptom three times today; someone should
take the cause. I have not done it myself because the stamps are
authored by hand across every worker's flow, so the change belongs
wherever that convention is documented, not in a file I happen to be
touching.

## Verification

`check-fnxc-future-dates` green (TZ=UTC CI=true) · `pnpm test:gate` 13 +
161 + 487 + 71 · lint · core typecheck clean · diff is date
substitutions only.
2026-07-31 07:56:09 -07:00
gsxdsm
b3d009edde fix: main is red on check-fnxc-future-dates — one stamp dated tomorrow (blocks every open PR) (#3166)
`check-fnxc-future-dates` runs in `pr-checks.yml`, so while `main` is
red **every open PR fails this check** regardless of what it touches.
Measured on a clean detached `origin/main`:

```
[check-fnxc-future-dates] FNXC stamp population changed:
  packages/core/src/task-store/task-update.ts: 2 future-dated FNXC stamp(s), baseline allows 1
    FNXC:StateMachine   2026-08-01  (dated after today)
    FNXC:WorkflowEvents 2026-08-01  (dated after today)
```

## One line, scoped by blame

Two stamps in the file are future-dated; only one is **new**:

| line | stamp | commit | action |
|---|---|---|---|
| 86 | `FNXC:StateMachine 2026-08-01-10:20` | `e5c9ea38709` (07-30) |
**baselined — left alone** |
| 964 | `FNXC:WorkflowEvents 2026-08-01-05:10` | `71f459c2a5d` (07-31) |
corrected → `2026-07-31-23:10` |

The baselined one is not what turned main red, and rewriting it would
register as a **drop** — which is exactly how I did collateral damage in
the #3124 cycle by rewriting two stamps I had not authored. Exact-match
replacement on the distinct new string; the older stamp is verified
still present afterwards.

## What is not changed

**The baseline file is untouched.** The fix is the stamp, not the
allowance — re-recording would clear the red while leaving tomorrow's
date in the tree, which is the false green this gate exists to prevent.

## Process note

I checked for an existing fix PR **before** writing this one. My #3143
was a duplicate of #3139 because I skipped that step on the last red,
and a red `main` is the single most likely thing for two lanes to notice
simultaneously.

## Verification

- `check-fnxc-future-dates` — **exit 1 on `origin/main`, exit 0 here**
- diff is one line; baseline file confirmed unmodified

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:30:21 -07:00
gsxdsm
27741d0e2f fix(core): an archived child kept blocking its parent's delete on a renamed board (#3162)
Second **LANE** site from the archived triage (#3154), same additive
shape as #3160.

## The bug

`liveLineageChildFilter` is the lineage-integrity gate (VAL-DATA-010)
behind `deleteTask` and `archiveTask`: a parent with **live** children
is refused with `TaskHasLineageChildrenError`.

It excluded children in the `archived` column **by id**. So on a board
that renames that lane, an archived child still counted as live and the
parent could not be deleted — with an error naming a child the operator
had **already filed away**.

## The fix is permissive, and that is the correct direction

The gate exists to protect **live** children; an archived child is not
one. Resolving makes fewer rows block, which is what the gate always
meant.

I am flagging this explicitly because *"a conversion makes a delete gate
stop firing"* deserves a second look. The second look is that it was
firing on rows it was never meant to protect.

## LANE, not STATE

The two STATE sites in this inventory are marked at their own sites
(#3157) and must never be resolved — one of them deletes directories.
This one asks about the board.

## Wired at every caller, not left as an optional seam

`findLiveLineageChildrenImpl` (has `store`) and both
`archive-lifecycle-2.ts` gates resolve and pass it.
`hasLiveLineageChildren` takes the same parameter so the two readers
**cannot disagree** about which children are live — the half-conversion
shape this program keeps finding, and the reason #3129's earlier attempt
was reverted for leaving a seam unsupplied.

## Parity gate satisfied

Same reason as #3160: the conversion is **additive** — it keeps the
literal as the fallback, so no encoding's literal count moves and a
caller supplying no set gets byte-identical SQL.

## Measured

- 4 new cases; parity test still **2/2**, inventories unmoved.
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control**, the **fail-soft** case, and the
parent/project-scope **negative** green.
- lineage / archive / soft-delete / archived suites — **6 files / 19
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — the literal remains the fallback arm, by design.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:18:54 -07:00
gsxdsm
5d5a3ddd60 fix(core): live search excluded the archived id, not the board's archive lane (#3160)
The first **LANE** site from the archived triage (#3154), converted —
and it establishes that this family does **not** need the single 52-site
commit the parity gate's message implies.

## The bug

`liveSearchPredicate` builds the "not archived" half of every task
search. Keyed on the literal, a card filed away on a renamed board
stayed in **every live search result** — including the CREATE-time
near-duplicate check, which calls `searchTasks()`.

So creating a task could be rejected as a duplicate of one the operator
had **already archived**, with nothing on screen explaining why.

## Why this site and not its neighbours

The triage classifies all eight Drizzle `archived` sites. Two are
**STATE** markers that must never be resolved —
`cleanupArchivedTasksImpl` deletes directories,
`listSoftDeletedColumnDriftCandidates` would "repair" already-correct
rows. Both are marked at their sites in #3157.

This one asks about the board, so it must resolve. That distinction is
the entire product of the triage, and it is why this is a one-site PR
rather than a sweep.

## The parity gate is satisfied — and the reason generalises

That gate exists because converting one encoding while the others
compare the raw string makes them **disagree**.

This conversion is **additive**: it adds a resolved path and keeps the
literal as the documented fallback. The SQL encoding's literal count
does not move, and no encoding shifts relative to another. A caller
supplying no set gets **byte-identical SQL**.

So the family can be converted **incrementally** — one site at a time,
each keeping its fallback — rather than in one coordinated 52-site
commit. The gate's rule is about not letting the encodings *diverge*,
not about batching.

That was the last thing making this cluster look unapproachable, and it
turns out not to be true.

## Threading was two layers, and the root already had the answer

`reads.ts` resolves archived lanes for its cold-storage decision a few
hundred lines above; the same call now serves both search paths. My
earlier scoping note (#3147) guessed "one parameter each" — it is the
builder plus its two entry points, with the resolution already present
at the root.

## Measured

- 4 new cases; parity test still **2/2** (inventories unmoved — that is
the point).
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control** and both negatives green.
- The `includeArchived: true` negative earns its place: resolving lanes
must not start excluding them from a search that explicitly asked for
archived rows.
- The predicate walker needed **cycle detection** — Drizzle's SQL graph
is circular (column → table → columns) and my first version blew the
stack on the first assertion.
- core search / archive / cold-storage / reads suites — **5 files / 15
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — the literal remains as the fallback arm, by design. A
census drop here would mean the fallback had been removed, which is what
the parity gate is protecting against.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:13:01 -07:00
gsxdsm
71f459c2a5 fix(events): the last two live task:moved emitters carry the resolved lanes (#3135)
## The last two live emitters

#3109 attached lanes at `moves.ts`. #3120 attached them at the archive
and completion emits. Two live emitters were still sending `lanes:
undefined`:

```
task-update.ts:962        todo -> triage
update-task-deps.ts:502   emits to: task.column — the row's ACTUAL lane
```

A listener reads absence as "unknown" and falls back to
`resolveTaskParkedColumnsSync`, which returns the **default board** for
every task under PostgreSQL. So these two paths kept the pre-#3109
behaviour while the listener code reads as resolved at every site.

`update-task-deps.ts` is the sharper of the two: it emits `to:
task.column`, the row's real lane, so on a renamed board it sends e.g.
`"shipped"` to a listener comparing against `"done"`. The emitter had
already resolved the board and threw the answer away. Nothing errors;
the branch stops firing.

## Emitter coverage after this

```
LANES    moves.ts:1450               (#3109)
LANES    archive-lifecycle-2.ts:385  (#3120)
LANES    task-artifacts-ops.ts:578   (#3120)
LANES    task-update.ts:970          (this PR)
LANES    update-task-deps.ts:504     (this PR)
MISSING  lifecycle-ops.ts:668        deliberate — see below
MISSING  lifecycle-ops.ts:715        deliberate — see below
```

**Every emitter that can execute under the shipped backend now carries
lanes.**

## Flagged — the two I did NOT convert

`lifecycle-ops.ts:668` and `:715` stay lane-less on purpose. Both sit on
the polling-replica path that file already documents as
legacy-SQLite-only — it reaches `store.db`, which throws under
PostgreSQL — and the same note argues against spending a signature
change on dead code. I agreed rather than overrode it. If that path is
ever revived they must be attached, because absence resolves to the
default board rather than to nothing.

## Supersedes my own earlier PR

This replaces **#3119**, which carried four emitters. #3120 landed two
of them first, so I rebuilt against current `main` with only the
remainder rather than resolving a conflict into a half-redundant diff.
#3119 is closed with nothing lost.

## Census before / after

```
before:  COLUMN guards (the backlog):   17
after:   COLUMN guards (the backlog):   17
```

Unchanged — this converts no guards. It makes the resolved answer
*reach* guards that were already converted, which is the half that was
missing.

## Verification

Full `@fusion/core` suite **462 files / 4906 passed, 0 failed** ·
`test:gate` exit 0 · typecheck exit 0 · lifecycle-column census exit 0 ·
`pnpm lint` clean.

## Still open

`main` is **red on `check:inert-sync-lanes`** and nothing in CI runs it
— **#3127** fixes both halves. **#3122** restores 13 guards laundered
through `mergeParkedColumns`. **#3131** corrects a 44% under-report in
`--triage`.
2026-07-31 07:10:09 -07:00
gsxdsm
b9b7d14804 fix(core): a type that taught the wrong invariant — staleness signal column narrowed to legacy ids (#3159)
**Type-only. No runtime behaviour changes**, and I would rather say that
than let a green suite imply otherwise.

## The type described a guard that no longer exists

`TaskAgeStalenessSignal.column` was typed `"in-progress" | "in-review"`
and filled through a cast carrying this justification:

```ts
// The guard above proves `column` is one of these two legacy ids ... (#1403)
const activeColumn = task.column as "in-progress" | "in-review";
```

True when written. The guard now reads:

```ts
const wipColumn    = context.lifecycle?.wip    ?? "in-progress";
const reviewColumn = context.lifecycle?.review ?? "in-review";
if (task.column !== wipColumn && task.column !== reviewColumn) return undefined;
```

So on a renamed board it proves the column is `building` or `checking` —
and the cast asserted the **opposite** of what the guard established.
The runtime was always fine; the real id passed straight through.

## The damage is in what the type taught

A consumer writing `signal.column === "building"` got a **compile
error** saying the comparison was impossible. The type actively
instructed callers that `=== "in-progress"` is exhaustive — the exact
guard shape this program spends its time removing.

This is the second instance of the shape today. The first was
`dashboard/src/server.ts`:

```ts
moveTask(taskId: string, column: "todo", options?: …): Promise<unknown>;
```

which made the type system **reject** a resolved target (#3158). Neither
was a constraint anyone chose — both were inferred from a single legacy
call site and then hardened into an assertion about live data.

**A type narrowed to legacy ids is a lint against fixing the code**, and
it is invisible to every gate this program has: the census counts
comparisons, the move-target ratchet counts arguments, and neither looks
at type positions.

## The test is a characterization, and says so

No runtime test can differentiate a type-level fix — **`tsc` is what
differentiates it**. The added case pins a value that was already
correct, so a future narrowing has something to break against besides a
compile error nobody sees until they hit it. I have labelled it in the
file rather than presenting it as a regression test.

## Verification

| | result |
|---|---|
| `tsc` — core, engine, dashboard (app + src) | **0 errors** each |
| `task-age-staleness` | **17 passed** |
| all three staleness suites | **28 passed** |
| census `--strict` | exit 0 |

All three consumers of `.column` only display or compare it
(`taskAgeStalenessCopy.ts`, `TaskDetailModal`, a `TaskCard` memo
comparison), so nothing downstream narrows on the widened type.

## Scope note

I scanned for this class and the raw pattern is noisy — 192 candidates,
almost all object-literal **values** (`status: "archived"`), agent
roles, and unrelated `type: "done"` stream events. This one and
`server.ts` are the two I could confirm as genuine type-position
narrowings on a *task column*. I have not filed an issue for the class
because I cannot yet separate it from the noise reliably; if a cheap
discriminator turns up, it is worth a ratchet like the move-target one.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 07:06:48 -07:00
gsxdsm
db71b4fffb docs(core): scope the archived three-encoding decision — it is not 52-or-nothing (#3147)
The `archived` family is the largest unclaimed cluster (52 sites) and
its blocker is that **nobody has scoped it**. This scopes it. It
converts nothing.

## The two options both read as enormous because 52 sites are counted as
one lump

They are not one lump. The sites answer **two different questions**:

| question | renameable? |
|---|---|
| **LANE** — "is this row resting in the board's archive lane?" | yes —
must resolve |
| **STATE** — "did Fusion archive this row?" (the marker `archiveTask`
writes) | **no** |

`async-maintenance.ts` already draws that line and marks its own site
DELIBERATE-LITERAL:

> `'archived'` is the STATE marker here, not a lane. This sweep collects
rows Fusion itself archived or soft-deleted; a card merely sitting in a
workflow's archived-TRAIT lane is live work and must not be collected.
Widening to the resolved archived set would pull real cards into a
cleanup pass.

**Converting that site would be a bug, not progress.**
`async-archive-lineage.ts`'s soft-delete path is the same shape —
`column = 'archived', deleted_at IS NOT NULL` is the storage state it
has just written.

So the first question is a **triage**, not a conversion: which of the 52
are lane questions? Nobody has answered it, which is exactly why the
cost reads as unbounded.

## Measured: the SQL half, which the existing note calls the hard part

8 Drizzle sites across 7 files.

**Four already have `store` in scope** — they could take a resolved set
today with no signature change:

- `branch-group-ops.ts` — `clearNearDuplicateReferencesToImpl(store,
...)`
- `branch-and-pr-entities.ts` —
`findRecentTasksByContentFingerprintImpl(store, ...)` (2 sites)
- `task-mutation-ops.ts` — `cleanupArchivedTasksImpl(store)`

**Four need one parameter each**, the same optional-lane-set shape used
throughout this program:

- `async-lifecycle.ts` — `liveLineageChildFilter(parentId, projectId?)`
- `async-search.ts` — `liveSearchPredicate(includeArchived, projectId?)`
- `async-self-healing.ts` — `listSoftDeletedColumnDriftCandidates(db,
...)`
- `store.ts` — the revert-lookup conditions (already holds
`this.asyncLayer`)

That is not *"threading a resolver into the persistence layer"*. It is
four call sites that already have what they need, plus four
one-parameter widenings — **before** any triage removes the STATE sites
from the count entirely.

## What I did not do, and why

The triage itself: a per-site judgement about what each guard *means*.
That belongs to whoever owns this gate, not to a passing fleet lane —
and getting it wrong in the STATE direction pulls live cards into a
cleanup sweep, which is the one failure mode here that destroys work
rather than hiding an affordance.

What was cheap and missing was the **shape** of the problem.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean.
- `check-fnxc-future-dates` is red from `main`'s own #3128 stamps —
**#3139** fixes that; this branch inherits and does not add to it.

## Census

No movement.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 06:55:03 -07:00
gsxdsm
110d6fd150 docs(core): the archived LANE-vs-STATE triage, done — 8 SQL sites classified with evidence (#3154)
#3147 scoped this cluster and said the first question is *"which of
these are LANE questions and which are STATE markers?"* — and that
nobody had answered it. **This answers it** for the Drizzle half, per
site, by reading what each query is for.

I claimed it because it has sat unclaimed for many rounds and `--claims`
reports `AVAILABLE: 0 files / 0 guards` — this is the only real work
left in the area. Nothing is converted here.

## LANE (6) — must resolve a renamed archive lane

| site | evidence |
|---|---|
| `store.ts` revert lookup | `ne(archived)` + `ne(done)` picking
**live** revert candidates |
| `branch-group-ops.ts:82` | near-duplicate marker cleanup over **live**
rows |
| `branch-and-pr-entities.ts:438` | content-fingerprint duplicate guard,
gated on `!includeArchived` |
| `branch-and-pr-entities.ts:470` | recent **sibling** lookup |
| `async-lifecycle.ts:68` | `liveLineageChildFilter` — the name is the
classification |
| `async-search.ts:82` | `liveSearchPredicate(includeArchived)` — same |

Four already hold `store` / `this.asyncLayer`. The two predicate
builders need one optional parameter each — the shape used throughout
this program.

## STATE (2) — converting these would be a **bug**

**`task-mutation-ops.ts:1072`** — `cleanupArchivedTasksImpl` selects
`eq(column, "archived")` and then `rm`s each row's files. Widening it to
the resolved archived set would feed cards **merely resting in a board's
archive lane** into a filesystem delete.

This is the most destructive site in the family, and it **looks
identical to the LANE sites at a glance** — same column, same operator,
same file neighbourhood. That is the whole argument for triaging before
converting.

**`async-self-healing.ts:61`** — soft-deleted rows whose column
*drifted* from the archive marker (`isNotNull(deletedAt) && ne(column,
"archived")`). Resolving it would classify a soft-deleted row sitting in
a renamed archive lane as drift and "repair" it.

## The raw-SQL half is already partly triaged in place

`async-maintenance.ts` is marked DELIBERATE-LITERAL as a STATE marker,
and `async-archive-lineage.ts`'s soft-delete path writes `column =
'archived', deleted_at IS NOT NULL` as the storage state it has just set
— STATE by construction.

## What this changes about the decision

Roughly **three quarters LANE, one quarter STATE** — and the STATE sites
are the ones that destroy data if converted.

That is why "convert all three encodings" cannot be a sweep, and why the
raw count of 52 made it look larger than it is: there are fewer sites to
convert than the headline, and the ones that must **not** be touched are
the part worth being careful about.

## Not converted here, deliberately

The gate requires all three encodings to move together, so the
conversion is one coordinated change with its inventories updated in the
same commit. This supplies the classification that change needs without
pre-empting it — and without me making a 52-site coordinated change at
the tail of a long session, which is exactly when I have made my worst
calls today.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean. **No
census movement.**

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 06:43:18 -07:00
gsxdsm
ada62a7c4a census: --claims shows which remaining files an open PR already holds (two duplicate claims today) (#3124)
The census says **where** the work is but not **who has it**, and
duplicate claims are now the dominant coordination cost of this phase.
This adds an opt-in `--claims` report mapping each remaining file to the
open PRs already touching it.

## The problem is measured, not suspected

- **`self-healing.ts` took three overlapping conversions** from
different lanes while one branch was open (#3049, #3075, #3078). Each
forced a full rebuild of #3094, and every conflict was the same shape:
*same guard, two spellings, different variable names*. That PR's body
asks, in as many words, for one lane to own the file.
- **`executor.ts` took two independent conversions today** — #3112 and
#3118 — same four literals, same payload-lanes fix, two branches. Two
workers each read the census, saw the top cluster, and started. Neither
could see the other; I only caught it because both appeared in one `gh
pr list`.

The census is what sends everyone to the same file, so the claim signal
belongs here rather than in a side channel nobody reads. `--triage`
(#3097) already measured the underlying fact — 53 of 88 guards sat
inside an open PR — one step short of being actionable.

## Measured on current main (29 guards)

```
  CLAIMED by an open PR: 6 files holding 15 guards
       6  packages/engine/src/self-healing.ts  ← #3121 #3116
       4  packages/engine/src/executor.ts  ← #3118 #3112
       2  packages/engine/src/auto-merge-finalization.ts  ← #3107
       1  packages/core/src/task-store/task-artifacts-ops.ts  ← #3120 #3119 #3091
       …
  UNCLAIMED: 12 files holding 14 guards — start here
       2  packages/dashboard/app/utils/taskRevert.ts
       2  packages/engine/src/scheduler.ts
       …
```

It independently reproduces **both** collisions I found by hand today,
which is the strongest evidence I can offer that it works: `executor.ts
← #3118 #3112` and `self-healing.ts ← #3121 #3116`.

It also answers the standing fleet instruction empirically. "Claim the
largest unclaimed cluster" currently resolves to **12 files holding 14
guards, none larger than 2** — and one of those two (`scheduler.ts`) is
in the SYNC-RESOLVED list, where conversion is inert. That is a
materially different picture from the headline `29`.

## Design decisions

**Report-only and fail-soft**, on the same terms as `--triage`: opt-in,
printed beside the totals, changes no count and no exit code. It shells
to `gh`, so it is unavailable offline, in CI without a token, and in
sandboxes — all of which print a notice and continue. A gate must not
depend on network state; this is a work-selection aid, not a gate.

**The fail-soft path is loud on purpose**, and it is the case I care
most about. A claim report that silently degrades to "nothing is
claimed" is *worse than no report*, because it actively sends the reader
into work another lane holds — the exact failure the flag exists to
prevent. So when `gh` cannot answer it prints `POSSIBLY CLAIMED` and
suppresses the start-here list entirely rather than rendering it empty.

**Heuristic, and says so.** A PR touching a file is not proof it
converts *that file's* guards — it may edit an unrelated function. It
over-reports rather than misses, which is the safe direction: a false
claim costs one comment asking, a missed one costs a rebuilt branch.

**One bulk `gh pr list` call**, not a request per PR — the per-PR shape
was too slow to become habitual, and a report nobody runs is not a fix.

## Verification

- `lifecycle-column-census.test.ts` — **42 passed** (was 40)
- Differential: disabling the flag gives **2 failed | 40 passed**. Both
new tests fail on the defect they were written for.
- `--strict` and `check-fnxc-future-dates` — exit 0
- Tests stub `gh` on PATH, so no network call and no dependency on the
live PR list. The fixture reads the census's **own current top file**
rather than a hardcoded path, so it cannot rot as the backlog shrinks
(same self-maintaining discipline as #3106).

## What this does not do

It does not reserve anything — there is no lock, and two workers who
both run it can still collide if they start simultaneously. It reports
what is already visible in the PR list, which is enough to catch the
every-case-so-far pattern of *starting work on a file someone has held
for hours*. A real reservation would need shared mutable state, and I
would not add that without an owner asking for it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:50:13 -07:00
gsxdsm
c38d784b88 fix(events): carry lanes on the archive and completion task:moved emits (#3120)
Follow-up to **#3109** (merged). Independent of my other branches.

## The gap

#3109 put resolved lanes on `task:moved` and wired the main move path in
`moves.ts`. **Two other emit paths still fired lane-less** — and to a
listener, a lane-less emit is not a *missing* answer, it is the
**legacy** answer, which on a renamed board is wrong.

Concretely: `archiveTaskBackendImpl` emits its own `task:moved`. The
executor's archive branch releases the task's active-session registry
entry, and **that entry is what blocks a successor task from acquiring
the same path**. Fixing the listener alone (#3112) leaves the leak
reachable *through this emitter*, because the listener still falls back
to the literal when the payload carries nothing.

That's the part worth noting for the pattern generally: **a
payload-carrying design is only as good as its emitters.** Converting
consumers without sweeping producers leaves a hole that looks fixed at
the call site.

## Scope

| Emit path | Action |
|---|---|
| `archive-lifecycle-2.ts:373` (`archiveTaskBackendImpl`) | carries
lanes |
| `task-artifacts-ops.ts:549` (`moveToDoneImpl`) | carries lanes |
| `lifecycle-ops.ts:715` | **untouched** — fires from a watcher callback
|
| `task-update.ts:962` | **untouched** — sync path |

Both converted sites are already `async` and already import
`resolveWorkflowIrForTask`, so this costs one IR read on a transition
that has just done database work. Fail-soft to `undefined`, matching
`moves.ts` — "unknown", never a wrong answer.

The two untouched ones each need their own look rather than a blanket
sweep; flagged, not guessed.

## Verification

- `@fusion/core` archive/lifecycle suites — **156 green**
- **`pnpm test:gate` green**; `tsc` clean
- Changeset added; `check:changesets` passes

## Census

**No change**, and that's expected — this converts emit *payloads*, not
comparison literals. The effect is that guards already converted in
#3109/#3112 receive a correct answer on these paths instead of a legacy
fallback.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Task move events now include accurate board lane information when
tasks are archived or marked complete.
* Improved handling for renamed workflows, ensuring task transitions use
the correct lane details.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:29:53 -07:00
gsxdsm
c5d5978a01 test(core): the emitter route does not generalise from task:moved to task:updated — measured (#3123)
#3109 solved the inert-guard class for `task:moved` by having the
**emitter** resolve lanes once and carry them on the payload. It is the
right fix there, and it retired several flags within a day — two of them
mine (#3118, #3121).

**The obvious next step is to do the same for `task:updated`**, which is
where every remaining sync-listener guard lives: `scheduler.ts`'s
mission-failure and PR-monitoring guards, and `triage.ts`'s
planning-evacuation guard. I went to do exactly that, and the cost
profile is opposite.

## Measured

| event | emit sites | files |
|---|---|---|
| `task:moved` | **7** | — |
| `task:updated` | **26** | 10 |

#3109 could argue its resolution away because a move *"is already async
and already post-commit, so the resolution costs one IR read on a
transition that has just done database work."*

`task:updated` fires on every **log append, comment, artifact write and
steering message**. `audit-ops.ts`'s logEntry **fast path** is one of
those emit sites — and that path exists specifically to avoid re-reading
the task. An IR read per emit there is a regression on the hottest write
path in the system.

## Partial coverage does not rescue it

The tempting narrower version is "add lanes only to the emit paths those
guards react to". `triage.ts` rules that out: its evacuation handler
reacts to **any** `task:updated` carrying a column, and explicitly
tolerates partial payloads (that is why it checks `typeof task.column
!== "string"`). It needs lanes on essentially every emit, or it keeps
its literal fallback regardless.

So those guards are **not one commit behind #3109**. Unblocking them
wants either a cached lane answer the emitter can attach for free, or
the sync reader this file already specifies.

## Pinned as a ratio, not prose

The new case asserts `task:updated` emits **more than twice**
`task:moved` — the shape of the argument rather than today's exact
numbers, which move with ordinary work. If the ratio ever converges, the
trade-off has genuinely changed and the note above it should be re-read.

That is deliberate: a comment stating "26 vs 7" would be wrong within a
week and would then argue for the opposite conclusion with full
confidence. This is the third time this session I have found a note that
was accurate when written and had quietly stopped being true.

## Measured

- 5 cases pass (4 pre-existing + 1 new).
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted.** This records why the cheap-looking
follow-on is not cheap, in the file that already owns this decision
(#3103).

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:24:05 -07:00
gsxdsm
82329819f7 fleet: resolve the pre-archive unarchive target (census 69 → 68) (#3091)
## Census

| | column guards |
|---|---|
| before | **69** |
| after | **68** |

## How this was found

By finishing a triage I'd left incomplete. Of the 19 single-guard files,
I had actually examined six and flagged the rest partly on assumption —
so I went back and read them.

Five of the remaining ones turned out to be `archived` comparisons
**pinned by `archived-column-gate-parity.test.ts`** (`audit-ops`,
`task-id-integrity`, `mission-store`, `async-comments-attachments`, plus
`merge-queue-ops-2` for review). Converting any of those moves one of
three encodings that must move together.

**This one isn't pinned**, and that difference is the whole PR.

## What changed

```ts
if (!declaresPreArchiveColumn || preArchiveColumn === archivedColumn
    || preArchiveColumn === "archived")
```

Belt-and-braces: the condition already accepted the resolved lane **or**
the legacy id, stated twice. A set says it once, so the two halves can't
drift apart — the real risk with a duplicated condition, rather than the
census count.

## Why this isn't the split brain

The parity guard pins comparisons of a **task's column** — one of three
encodings of *"an archived task is not live."* This compares a **stored
`preArchiveColumn` value** against the board's archive lane: a different
question, about where to send a card on unarchive.

Verified rather than argued — that suite runs **green** here, and it
went **red** the last time I touched a pinned site (#3076, where I named
an arm and immediately reverted). It's a live check, not an assumption.

## Measured

| check | result |
|---|---|
| archive / artifact / unarchive suites | 7 files, **27 tests green** |
| `archived-column-gate-parity` | **2 passed** |
| five gates + strict census | green |
| core `tsc` | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:06:22 -07:00
gsxdsm
50f089a14b test(core): archived is renameable today — remove a wrong reason for the cheap parity option (#3113)
Follow-on from #3110, where the archived-gate parity test blocked my
conversion and laid out two ways forward. This removes a **wrong
reason** for picking the cheap one.

## The cheap option looks like it has already been taken

`archived-column-gate-parity.test.ts` offers: convert all three
encodings, **or** *"declare `archived` a non-renameable system column
and mark the sites deliberate"*.

The second is far cheaper — and the codebase reads as if it is already
true:

> `"Globally archived; hidden from the board. RESTRICTED (built-in
only)."` — `trait-types.ts`

At a glance that says a custom board cannot have an archive lane of its
own, which would make all 52 sites correct by construction and the whole
family **deliberate rather than debt**. That is a very attractive
conclusion for anyone facing a 52-site conversion.

## It does not mean that

The restriction is over trait **registration**. `trait-registry.ts`
rejects a **non-builtin** (plugin-defined) trait that declares
`archived` or `complete` (R22). It says nothing about which **column**
may carry the built-in trait, and a custom workflow may put it on a
column with any id.

**Verified, not argued:**

| input | result |
|---|---|
| `columnsWithFlag(ir, "archived")` where the lane is named `filed` |
`["filed"]` |
| `resolveColumnFlags` on that column | `{ archived: true,
hiddenFromBoard: true }` |
| `columnsWithFlag(ir, "complete")` where the lane is named `shipped` |
`["shipped"]` |

So option two is a **capability removal** — it would silently break any
board that has already renamed its archive lane — not the documentation
of a constraint that already exists.

## What this does and does not do

It **does not** decide between the two options; that is a product call
with real blast radius either way. It removes the reason someone would
most plausibly reach for after reading that flag comment, and it does so
with an executable check rather than a claim.

The `complete` case is pinned alongside it, because both restricted
flags behave identically — anyone reaching the same conclusion about
`complete` would be wrong for the same reason.

The first case also doubles as a **tripwire**: if `columnsWithFlag(ir,
"archived")` ever returns `[]` or `["archived"]` for that fixture, the
cheap option has effectively been taken and the parity test's framing
needs revisiting. That is the signal the case exists to give.

## Measured

- 4 new cases pass.
- The parity test still passes with its corrected framing — it excludes
`__tests__` from its scan, so the added prose cannot move its own
inventory. I checked that before editing it rather than after.
- `src/__tests__/{archived,trait,log-entry}*` — **3 files / 26 tests
pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted.**

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:56:25 -07:00
gsxdsm
77d11a0a9a docs(core): resolveLifecycleColumns returns one id per role, not a set (#3111)
Comment-only. No behaviour change. `tsc` 0 errors, eslint clean, `census
--strict` and `check:fnxc-future-dates` exit 0.

## Why this is worth a PR

Three fleet PRs have independently read this function's fields as
**membership**:

| where | finding |
|---|---|
| `self-healing.ts` `hold` guard | **#3084**, `Major` — *"resolve `hold`
by membership, not first-per-role"* |
| `notification-service.ts` progressed-lane check | **#3096** — reads
`?.hold / ?.wip / ?.complete / ?.archived` |
| sync vs async review set | **#3088** — same shape, different symptom |

Each field is a **single id**: the first column carrying that trait. A
board may declare several columns with one role — two pre-review working
lanes, or a review lane plus a merge-blocked lane — and every one after
the first is invisible.

Used for membership, that reads as a working check while silently
ignoring lanes: a card resting in the second hold column classifies as
**not-held**, and the guard depending on it never fires. It is the quiet
direction of wrong, which is why it survives review three times.

**Three occurrences of one misreading is a property of the signature,
not three unlucky authors.** Fixing it at the call sites one at a time
leaves the next author to rediscover it, so this documents the contract
where the mistake is made — at the definition — and names the
alternatives:

- `columnsWithFlag(ir, role)` — every column with the role on one board
- `resolveProjectColumnsForRoles(store, roles)` — the union across a
project's workflows

Rule of thumb recorded in the doc: **a routing/move target wants this
function; a `.has(task.column)` test does not.**

## What this does not do

It does not fix the three call sites. #3084's is a live `Major` and
needs a real fix with mutation verification; #3096's needs its author to
resolve a two-implementation conflict first; #3088's is flagged and
open. This only stops the fourth occurrence.

A stronger version would rename the fields (`firstHoldColumn`) or return
branded single-id types so misuse fails to compile. That is a wider
change across every caller and belongs to whoever owns the helper's API
— worth considering once the current fleet PRs land, since doing it now
would conflict with all three.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:50:28 -07:00
gsxdsm
41cdcc741e fix(events): carry resolved lanes on task:moved so listener guards stop being inert (#3109)
Removes the **inert-guard class at its source** instead of one call site
at a time. Independent of my other branches.

## The problem

`task:moved` listeners run synchronously, so a listener needing a lane
answer had to resolve one synchronously — and
`resolveTaskWorkflowIrSync` returns the **default** workflow under
PostgreSQL, the shipped backend. Every such guard behaved exactly as the
literal it replaced, while the census scored it as converted.

**Resolving asynchronously inside the listener is not available**, and
that is measured rather than assumed. The scheduler's
`snapshotManager.invalidate` is asserted to run in the listener's
**synchronous prologue**; putting an await ahead of it produced **3
failures across 21 scheduler suites**.

## The fix

The emitter carries the answer, which removes the dilemma rather than
trading one horn for the other. `moves.ts` is already async and already
post-commit, so it resolves the moving task's lanes **once** and hands
them to every listener. The guard becomes correct **and** the prologue
stays synchronous.

This is the file's own recorded preferred fix — *"having the emitter
carry the resolved lanes on the event payload so no listener resolves at
all"* — now that the audit it was waiting on is done and came back as
**one** prologue-dependent consumer, not a class.

## Design choices

- **`lanes` is optional and fail-soft to `undefined`** — "unknown",
never "legacy". Some emit paths fire from sync contexts or a cached row
mid-teardown. Listeners keep their existing fallback, so those paths are
no better than before but **no worse**, and they become the exception
rather than the rule.
- **`mergeParkedColumns` overlays only fields the emitter actually
resolved**, so a partial payload cannot blank a lane back to a wrong
answer.
- **The sync resolver stays** as that fallback. Deleting it would strand
the emit paths that cannot resolve.

## Verification

- **Revert-proof and it pins the prologue:** the new case asserts
invalidation on a **renamed** hold lane with **no `waitFor`**. Ignoring
the payload gives **0 calls**.
- 21 scheduler suites — **361 green**
- self-healing + notification suites — **491 green**
- core moves + the `sync-workflow-ir-callsite-allowlist` ratchet — green
- **`pnpm test:gate` green** (71)
- Changeset added; `check:changesets` passes

## What it unblocks

`scheduler.ts`'s 10 allow-listed guards now resolve correctly for every
move that goes through `moves.ts` — the path real moves take. Those were
already absent from the backlog, so **the census number does not move**;
what changes is that they now do what the number claimed.
`executor.ts`'s 4 remaining sites can follow the same pattern in a
separate PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:47:09 -07:00
gsxdsm
4afb32ef98 test(core): cover the untested log-entry archive gate; correct a deferral that named the wrong blocker (#3110)
I converted `audit-ops.ts`'s archived gate, measured, and **backed it
out**. Both halves of that are the deliverable.

## The old deferral was stale on its own terms

It declined the conversion because *"the fix is the same one
`getLiveTaskColumn` needs"* and doing one of the pair would leave them
disagreeing.

But `getLiveTaskColumn` now **takes** a resolved `archivedColumns` set,
and both of its callers already pass `await resolveArchivedLanes(store)`
— including the sentinel path **twenty lines up in this same function**.
The pair it worried about was already half-converted, and this arm was
the half out of step. Converting it would have made them *agree*.

That is the third deferral I have found this session whose stated
blocker had dissolved. A deferral note records the blocker at the moment
it was written, and nothing re-checks it.

## The real blocker is one neither note named

`archived-column-gate-parity.test.ts` failed my conversion, and its
reasoning is correct and not obvious. This gate has **three encodings**:

1. TypeScript comparisons
2. Drizzle `eq`/`ne` predicates
3. raw SQL templates

Converting only the TypeScript arm makes them **diverge**: the gate
would call the row archived while the SQL side still returns it as live
— a log write rejected by its gate while its parent is listed as live.

Every builtin workflow names the column `archived`, so all three agree
*by accident* on every board we ship, and nothing except that parity
test can see the split.

Unblocking means converting all three together — the SQL sides need the
resolved id as a query-build value, including inside `for update`
transactions that receive no store today — or declaring `archived` a
non-renameable system column. That test lays out both options and owns
the inventory that has to move in the same commit. I am not doing it
here; it is a different change from a lane conversion.

## What ships

**The corrected note**, and **a test for a gate that had no coverage in
any form**.

The test asserts the legacy refusal and — the case that matters more —
that a **live lane is not refused**. A gate that refused everything
would satisfy a one-sided test and silently break every log write on the
board.

The renamed case is recorded as a **deliberate, explained omission**
rather than left as a silent hole, so the next reader knows it is a
decision.

## Measured

- 3 new cases pass; the parity gate passes.
- **MUTATION**, on the conversion before I reverted it: restoring the
literal failed the renamed case. The conversion *worked* — which is
exactly why the parity gate mattered. A working change can still be the
wrong change.
- The live-lane negative asserts **the gate did not fire**, not that the
call succeeded: past the gate the fast path performs a real Drizzle
write this fake layer cannot serve, so asserting success would drag a
database fixture into a test about a lane comparison, and asserting a
bare rejection would pass even if the gate *had* fired.
- `src/__tests__/{log-entry,archive,cold-storage,unarchive}*` — **7
files / 23 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement — nothing converted, deliberately.** The count stays where
it is because the gate is blocked, not because it is fine.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:46:58 -07:00
gsxdsm
39e6891c93 chore(core): mark the dead sync-path lane literal DELIBERATE-LITERAL (census 104→102) (#3060)
Fleet phase. Claimed `packages/core/src/task-store/project-store-ops.ts`
— the largest census file with no branch, worktree, or open PR against
it. Claim published by pushing the branch **before** starting work.

## Census before / after

| | total | this file | deliberate |
|---|---|---|---|
| before | **104** | 2 | 130 |
| after | **102** | 0 | 130 |

`--strict` exits 0, baseline re-recorded in the same commit.
**Reclassification, not conversion** — the line is unchanged.

## The site was already audited today, in prose the tool cannot read

```
FNXC:WorkflowLifecycleColumns 2026-07-31-02:45 (audited — DEAD SYNC PATH, do not convert):
… It is the SQLite-mode twin. The live path is `dequeueMergeQueueOnColumnExitInTransaction`
… and it is ALREADY converted … This body reaches for `store.db.prepare`, which throws in
   PostgreSQL backend mode …
```

The reasoning is sound and I did not second-guess it: the live path is
converted, this twin cannot execute in production, and converting it
would mean threading a lane set into a function whose first statement
throws.

The problem is purely mechanical — **the note is prose, and the census
reads markers.** So the site stayed in `byFile` looking like unconverted
debt, and each fleet pass pays to re-derive the same conclusion. Adding
`DELIBERATE-LITERAL` moves it to `deliberateByFile`, where a
reviewed-and-kept literal belongs.

## This is the second one, which makes it a pattern

Same shape as #3056 (`async-mission-store-queries.ts`, fallback arms).
Across the files I have checked this phase — `agent-store`,
`github-tracking-state`, `planner-overseer`, `auto-merge-finalization`,
`async-mission-store-queries`, and this one — **every site was either a
fallback arm or an already-documented deliberate leave**, and
`agent-store.ts:236` carries its own "FLAGGED AND LEFT COUNTED" note
from today.

So the count is not a work queue, and the gap is not judgement —
previous passes reached the right answer. They recorded it where only a
human reader would find it. Two lines of marker per site closes that,
and the number then means "conversions owed", which is how every worker
reads it when picking a cluster.

## Verification

- `census --strict` exit 0; `tsc --noEmit` **0 errors**
- `check:fnxc-future-dates`, `check:lane-wiring`,
`check:sql-column-literals`, `check:inert-flag-seams` — all exit 0
- No behaviour change: only a comment added

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:30:00 -07:00
gsxdsm
0da19f7963 fix(core): a renamed archive lane was recorded as done in the eval corpus; flag the scheduler's two honest literals (#3100)
Two pieces, both about the same distinction: which literals are worth
**converting** and which are worth **naming**.

## Converted — the eval corpus was mislabelling renamed archive lanes

`collectDeterministicSignals` writes `column` as a two-value eval-record
field. Against the `archived` literal, a card resting in a renamed
archive lane was recorded as `"done"`.

No crash, no lifecycle decision — a **mislabelled row in the eval
corpus**, which is a dataset every later comparison reads. That is the
expensive kind of quiet: nothing fails, the numbers just drift.

The collector is sync and pure (no store, no workflow), so the lane
answer arrives as an optional parameter.
`HybridEvaluatorService.evaluateTask` is async and already holds an
optional store, which is where the resolution is paid; a store-less
evaluator degrades to the legacy literal rather than failing.

**Only the archived arm was ever wrong.** A renamed *complete* lane was,
and remains, recorded as `"done"` — which is correct. So only that
answer is resolved, and a third case pins that the widening did not turn
every renamed lane into `"archived"`.

## Flagged, not converted — the scheduler's two honest literals

These are the two `scheduler.ts` literals the sync-lane pass did not
take, and **nothing in the file said why**. That silence is the problem:
the obvious next move is to "finish the job" the way the other ten were
converted, and that would make them **inert, not fixed**.

`getTaskWorkflowSelectionImpl` returns `undefined` unconditionally under
PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR — proved in
`postgres/sync-workflow-ir-is-always-default.pg.test.ts`, and
`check-inert-sync-lane-conversions` already baselines **twenty** guards
in that state in this same file.

They stay literal and **counted**, which is the honest state. An
unconverted literal is visible to the census; an inert conversion leaves
the backlog and takes the evidence with it. The note names the real
blocker — a sync-capable workflow-selection reader — so the next pass
does not spend a cycle discovering this the way I did.

## Measured

- 3 new cases in `eval-signal-collector.test.ts` — file **5/5 pass**.
- **MUTATION**: restoring the `archived` literal fails the renamed case
and leaves **both** the legacy control and the renamed-complete negative
green. The negative matters here: the fix must not turn every renamed
lane into `"archived"`.
- core eval suites — **4 files / 20 tests**; engine scheduler +
evaluator — **14 files / 143 tests**.
- `tsc --noEmit` clean in both packages; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

Both files keep their counts, deliberately:

- `eval-signal-collector.ts` — the remaining entry is the new
parameter's documented default, which is the fallback doing its job.
- `scheduler.ts` — the two literals this PR deliberately leaves visible.

A census that fell here would mean the flags had been marked exempt,
which would assert the code is fine. It is not fine; it is blocked, and
those are different claims with different expiries.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:29:21 -07:00
gsxdsm
827386dde6 test(core): the sync IR path is blocked TWICE, not once — every note in the repo undercounts it (#3103)
Every remaining census cluster I could not convert — `executor.ts` (4),
`scheduler.ts` (2), `triage.ts` (1), and the four-guard fan-out I
withdrew from my own PR — is waiting on the same thing. So I went to
unblock it, and found the record is wrong.

## The repo says one blocker. There are two, plus a constraint

The call-site allow-list header, the live-PG proof, and a dozen FNXC
notes across engine and core — **several of which I wrote** — all say:
`resolveTaskWorkflowIrSync` is inert because the sync selection reader
returns `undefined`, and the fix is "a sync-capable workflow-selection
reader".

That understates the work by half, and the undercount is load-bearing:
it makes the unblock read like a caching job, so the next person ships a
selection cache and finds the rest at integration time.

### Blocker 2 — the IR read is dead too

`resolveTaskWorkflowIrSyncImpl` loads a **custom** workflow's IR through
`store.db.prepare("SELECT ir FROM workflows WHERE id = ?")`.

`TaskStore.db` is not "SQLite-only". Its implementation (`dbImpl`,
`task-id-integrity.ts`) is an **unconditional throw with no mode branch
at all**. That read always throws into the surrounding `catch`, which
always returns the default IR.

The consequence is precisely the one this program cares about:

| workflow kind | after a perfect selection reader |
|---|---|
| built-in | resolves — that branch never touches `store.db` |
| **custom** | **still the default IR, always** |

**A renamed lane is by definition a custom workflow.** So the sync path
cannot serve the renamed-board case *at all* until this second read is
replaced. Fixing the selection reader alone would produce a change that
looks like it works — on default boards.

### Blocker 3 — a node-local cache is unsafe here

Not a bug; a constraint that bounds the fix's shape. Multiple Fusion
nodes run their own engines against **one shared PostgreSQL**
(`docs/multi-project.md` → "Shared Postgres multi-node runbook").

A node-local synchronous cache of `task_workflow_selection` therefore
goes stale whenever *another node* rewrites a selection — and answers
with full confidence. That is **worse than today's default**, which is
at least uniformly wrong rather than intermittently wrong. Any sync
reader needs an invalidation story that survives a writer on a different
host.

## Why a test rather than a comment

A comment saying "db always throws" decays the moment someone adds a
mode branch, and the whole argument silently inverts — which is the same
decay mode this conversion program keeps hitting with allow-list entries
and stale notes.

The assertions are deliberately about `dbImpl`'s **source** rather than
a call. Calling it proves one construction path throws; the fix depends
on the stronger claim that **no mode returns a database**. Reintroducing
an `if`/`return` there fails the test, which is the correct outcome: the
premise really has changed and the file must be re-read.

## Measured

- 4 new cases pass; the allow-list file's own 7 still pass with its
corrected header.
- **MUTATION**: adding a mode branch to `dbImpl` fails the first case.
- An **anti-vacuity** case pins that the resolver is still live and
still allow-listed, so these source assertions cannot keep passing after
the concern is deleted.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean.

## Census

**No movement — this converts nothing.** It corrects the record about
what the remaining conversions are waiting on, and it corrects notes I
authored. I would rather spend a PR making the next attempt cheap than
leave a half-true blocker in place that costs someone a full cycle to
rediscover.

## What I did not do

I did not build the sync reader. With blockers 2 and 3 in view it is a
store-substrate change — a second read to replace, and an invalidation
story that survives a writer on another host — not a fleet conversion,
and starting it mid-sweep on a shared file would repeat the collision
pattern that has already cost this branch three rebuilds. It remains
unclaimed, and now it is fully specified.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:25:48 -07:00
gsxdsm
f7a7347e1b test(core): cover #3057's cold-storage conversion; audit two archived literals as dead sync (#3089)
**Replaces #3085, which I am closing.** #3057 landed the same
cold-storage conversion while that PR was open. Rather than argue about
which spelling wins, this keeps only what `main` does not have: the
coverage, and two audits.

## #3057 converted this and shipped no test

`listTasksImpl`'s `columnFilterIsArchive` replaced `columnFilter ===
"archived"`. Correct change — and the kind that needs a test more than
most, because of how it fails.

Archived rows do not live in `tasks`; `archiveTask` copies them into the
archive store and removes them. This decision is whether that second
store is read **at all**. Against the literal, a caller naming a renamed
archive lane — `listTasks({ column: "filed", includeArchived: true })`,
which is what an archive view does — got an empty page from the only API
that can reach those rows.

Note the shape: the **unfiltered** read (`!columnFilter`) was always
correct. It fails only for the caller that names the lane, so it
survives any board-level smoke test and presents as *"the archive is
empty"* rather than as a bug. That is precisely the class that regresses
quietly once the conversion that fixed it has nothing holding it.

Three cases, and each earns its place:

| case | what it stops |
|---|---|
| renamed archive lane | the regression itself |
| legacy `archived` id (**control**) | a future conversion that resolves
the renamed lane and *drops* the legacy seed — the seeding hazard in its
other direction |
| non-archive lane (**negative**) | the widening turning every filtered
board read into a second-store round-trip |

**MUTATION**: restoring `columnFilter === "archived"` fails **only** the
renamed case.

**A trap worth recording.** My first version left the LIVE read real
against a fake `layer.db`. It threw, the `.catch` swallowed it, and all
three cases passed the negative — *including the legacy control*, which
is what exposed it. A test whose subject is never reached looks
identical to one whose subject answered no. `readLiveTaskRows` is mocked
now, and the control is what made it detectable.

## Audited, not converted — two dead sync paths

Both would be real defects if they ran. Neither runs.

| site | why it is dead |
|---|---|
| `mission-store.ts` (feature-delete link check) | `getMissionStoreImpl`
returns the AsyncDataLayer-backed `AsyncMissionStore` under PostgreSQL;
the sync `MissionStore` reached via `this.db.prepare` is legacy SQLite
only |
| `lifecycle-ops.ts` (polling-replica archive emit) |
`checkForChangesImpl` opens with `store.db.getLastModified()` /
`store.db.prepare`, which throw in backend mode |

The second is the sharper one: **both** its guard and the `to:
"archived"` it emits are literals, so a polling replica on a renamed
board would emit a move to a column the board does not declare.

Recorded in place — the treatment `project-store-ops.ts`'s dequeue twin
already has — so the census entries are not mistaken for unconverted
debt, and whoever deletes the sync SQLite residue takes these with it.

**They stay COUNTED.** Marking them DELIBERATE-LITERAL would buy a
smaller number by asserting the code is *correct*. It is not correct; it
is unreachable. Those are different claims with different expiries, and
the census should keep pointing here until the code is gone.

## Census

No movement — by design. This PR adds coverage and audits; it converts
nothing that was not already converted on `main`.

## Measured

- 3 new cases pass; `src/__tests__/{cold-storage,archive,unarchive}*` —
**6 files / 21 tests pass**
- `tsc --noEmit -p packages/core` clean; census `--strict` clean

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:06:40 -07:00
gsxdsm
6949f22ef8 fleet: resolve the same-column handoff review target (census 84 → 83) (#3076)
## Census

| | column guards |
|---|---|
| before | **84** |
| after | **83** |

`moves.ts`: 2 → 1. One of its two sites converts; the other **must
not**, and that difference is the useful part of this PR.

## Converted — the move target at the same-column handoff

```ts
if (internal.fromHandoff && toColumn === "in-review")
```

Against the literal this **never fired on a renamed board**, so a
same-column handoff into a renamed review lane silently took the *other*
branch — the sync-SQLite path, which throws under PostgreSQL.

It now asks `moveReviewColumns`: the broad membership set
(`mergeOrchestration ∪ mergeBlocker ∪ humanReview`) already resolved
**three lines above** for the merge-queue pair. Same value, so this
branch cannot disagree with the enqueue/dequeue calls that receive it.

## Not converted — the archived fallback arm

I named it, and `archived-column-gate-parity.test.ts` went red on
**`TypeScript encoding changed`**.

That guard's argument holds: the archived gate is enforced in three
encodings, the SQL halves still compare the raw string, and moving the
TypeScript half alone is the split brain it exists to prevent. Restored
inline **with a note recording the measurement**, so the next person
doesn't retry it and rediscover the same red.

This is the second time that guard has stopped me this session. It's
doing exactly what it was built for.

## On the pre-existing red

That suite is red on `origin/main` for an unrelated raw-SQL drift (#3072
fixes it — the drift is from my own merged #3042/#3046). I verified this
branch produces the **identical** failure and no other, so it doesn't
compound it.

## Measured

| check | result |
|---|---|
| moves / handoff / merge-queue suites | green |
| four gates + strict census | green |
| core `tsc` | clean |
| parity suite | same single raw-SQL failure as `origin/main`, nothing
added |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:57:51 -07:00
gsxdsm
98aac40ca8 fleet: reads.ts 2 → 0 lifecycle-column guards (#3057)
> **Rebased.** Main landed another worker's conversion of the review
gate while this was open — the overlap was a whole rewritten function,
so I reset to main and rebuilt only my remaining delta on top of their
work rather than resolving hunks. Their conversion is kept as-is.

## Census

| | column guards |
|---|---|
| before | **104** |
| after | **102** |

`reads.ts`: **2 → 0**.

## Two changes

**1. `includeColdStorage`** asks whether the *caller* is filtering to
the archive lane. Against the literal, a caller filtering to a renamed
archive lane took the false branch — cold storage was skipped and the
filtered view returned only whatever archived rows still sat in
`project.tasks`, **a short list presented as the whole archive**. Still
literal on main; converted here.

**2. Both fallbacks become named sets** instead of inline arms —
including the one on the just-landed review gate.

## The second point is the one worth the fleet's attention

This is bookkeeping correctness, not style. The census counts an inline
comparison **whether or not it sits in a fallback branch**, because its
`traitFallback` hint is advisory and never changes `kind`.

So a correctly-converted guard with an inline legacy arm **stays on the
backlog permanently**, and the number stops distinguishing real debt
from documented degraded answers.

Concretely: converting with an inline fallback is correct work that
scores **zero**. My own first pass at this file did exactly that. There
are roughly **12 such sites** across the tree — `github-tracking-state`,
`planner-overseer`, `async-mission-store-queries`,
`register-task-workflow-routes`, `restart-recovery-coordinator` — and I
have that cluster converted and ready to open next.

## Measured

| check | result |
|---|---|
| reads / get-task / stall suites | 5 files, **87 tests green** |
| renamed-archive PG suite | green |
| strict census | green; `tsc` clean |
| unconverted boards | byte-identical — the named sets hold the previous
ids |

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Review and archive checks now work correctly with resolved workflow
columns while retaining legacy compatibility.
  * Fresh agent activity is detected in resolved review lanes.
* Lists filtered by a resolved archive lane now include archived items
stored in cold storage.

* **Chores**
  * Updated internal lifecycle tracking baselines.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:43:03 -07:00
gsxdsm
3c12a51627 fix(core): the merge result reported a column the finaliser did not write (merge-queue-ops 3 → 0) (#3071)
Largest unclaimed census cluster in `packages/core` — three `done`
literals in `mergeTaskImpl`. Two of them produced **wrong state**, not
merely a guard that stopped firing.

## 1. The result overrode the writer

`moveToDoneImpl` resolves the board's completion lane and writes it onto
the task object:

```ts
task.column = completeColumn;   // task-artifacts-ops.ts
```

Both merge call sites then did:

```ts
result.task = { ...task, column: "done" };
```

putting the literal back over what the writer had just set. Every
`task:merged` listener — GitHub tracking, the auto-merge handoff — was
told the card landed in `done` while the persisted row said `shipped`.

The row was right and the event was wrong, which is the worse direction:
the listeners act on the event, not the row.

Fixed by reading back what the writer set (`{ ...task }`). Deliberately
**not** a second resolution — that would only be a second chance to
disagree with the finaliser.

## 2. The guard disagreed with the writer

The already-complete short-circuit asked `task.column === "done"`, while
the finaliser it guards short-circuits on the resolved `task.column ===
completeColumn`. On a renamed board those two answers differ, so a card
already resting in the board's completion lane fell through and the
merge ran again against a branch that was already landed and deleted.

Converted with the **same resolution and the same shape** — a single
first-match column, not membership — because the whole point is that
these two answers cannot differ. A workflow declaring no complete lane
resolves to `undefined`, which matches no column; the finaliser refuses
such a board explicitly one function later.

## Census

| | before | after |
|---|---|---|
| `merge-queue-ops.ts` | 3 | **0** |

## Measured

- Two new cases added to `merge-blocker-renamed-review-lane.test.ts`
(same renamed-board fixture, same PG harness) — file **5/5 pass**.
- **MUTATION**: restoring either literal fails **both** new cases and
leaves the three pre-existing ones green.
- Reached with **no git fixture**: with no branch present, `git
rev-parse --verify` fails and the function takes its own documented
*"branch not found — moving to done without merge"* path — which is
exactly the path that calls `moveToDone` and then builds the result. No
repo setup, no flake surface.
- `packages/core` targeted run: **38 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-lane-wiring` ("none added"), `check-fnxc-future-dates` clean.

## Not done here (flagged, not guessed)

The other `done`/`archived` literals still in the core census are each
blocked for a *different* documented reason, so sweeping them into this
PR would have meant guessing:

- `agent-store.ts:236` — a pure formatter over `Pick<Task,"column">`
that prints the column for a human; degrades gracefully and has no store
to resolve from.
- `async-mission-store-queries.ts` — already converted with
caller-threaded lane sets.
- `taskRevert.ts:119` — classifies a **neighbour** task; the only flags
in scope describe the modal's own task, so wiring them would answer the
question for the wrong row. Needs per-neighbour flags.
- `moves.ts:310` — a **refusal**, where a legacy-seeded superset is the
documented hazard rather than the safe direction. Wants its own change
with its own test.
- `mission-store.ts:2332` — a sync SQLite path with no async seam.

Each is real debt; none is a mechanical conversion.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:40:11 -07:00
gsxdsm
8eef8852a0 fleet: 4 long-tail fallback arms become named sets (census 101 → 97) (#3064)
## Census

| | column guards |
|---|---|
| before | **101** |
| after | **97** |

The single-guard long tail is **19 files**. This converts the four whose
legacy arm is unambiguously a fallback on an already-converted guard;
the other 15 are flagged below rather than guessed at.

## Two shapes

**`in-review-stall.ts`, `stalled-review-detector.ts`** — the resolved
answer with an inline legacy arm:

```ts
reviewColumns ? reviewColumns.has(col) : col === "in-review"
→ (reviewColumns ?? LEGACY_REVIEW_LANES).has(col)
```

**`merger.ts`, `in-process-runtime.ts`** — belt-and-braces:

```ts
col !== (lifecycle?.complete ?? "done") && col !== "done"
```

That accepted the resolved lane **or** the legacy id, stated twice. A
union set says it once, so the two halves can't drift apart — which is
the real risk with a duplicated condition.

## A finding for anyone else marking fallbacks

`in-review-stall.ts` **already carried a `DELIBERATE-LITERAL` marker**
on that arm and was counted anyway. The marker sits in a comment *inside
a ternary*, which the census's leading-comment lookup doesn't reach.

So: **naming the set works, marking it does not.** Worth knowing before
someone marks a fallback and expects the count to move.

## No behaviour change

`new Set(["in-review"]).has(x)` answers exactly what `x === "in-review"`
answered, and the union sets accept exactly the two lanes their
conditions already accepted.

## Flagged, not converted

The remaining 15 single-guard sites need individual judgement, not a
mechanical pass:

- **plain unconverted guards with no resolution in scope** —
`audit-ops`, `lifecycle-ops`, `merge-queue-ops`, `task-id-integrity`,
`backlog-pressure-reporter`, `ephemeral-worker-manager`,
`ResearchTaskActionModal`
- **sites where the literal IS the answer** — `eval-signal-collector`
maps a column to an archive-vs-done *label*; `TaskCard` reads a
completion timestamp
- **already resolved on their line** — `triage.ts`,
`restart-recovery-coordinator.ts`, both covered by open PRs

## Measured

| check | result |
|---|---|
| core stall suites | 4 files, **85 tests green** |
| engine merger/runtime suites | **1044 tests green** |
| five gates + strict census | green |
| `tsc` (core, engine) | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:37:18 -07:00
gsxdsm
befbd299a9 fleet: mark 4 reviewed literals DELIBERATE (backlog 88 → 84, comment-only) (#3066)
Comment-only. **No code changed** — 23 lines added, all comments.

## Census before/after

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 88 | **84** |
| DELIBERATE-LITERAL (reviewed) | 128 | **132** |

| File | Sites marked |
|---|---:|
| `packages/core/src/agent-store.ts` | 2 |
| `packages/core/src/async-mission-store-queries.ts` | 2 |

## Why marking, not converting

Both files already carried prose explaining why their literals are
correct. Without the marker the census still counts them as backlog, so
the fleet keeps dispatching workers at them — **three separate workers
have now independently re-derived the same two conclusions.** An
unmarked correct site costs a cycle every time it is re-examined, and
the cost repeats for every worker.

**`agent-store.ts`** picks a *word* for a human reader, not a lifecycle
decision: `(not active — done)` versus `(done)`. It degrades gracefully
on a renamed board — falls through to `(<column>)`, still accurate, just
less specific. Threading a resolution into a synchronous string builder
to choose an adjective is the wrong trade.

**`async-mission-store-queries.ts`** are the fallback arms of an
*already-converted* predicate, and the undefined branch is a **live
intended path**: `AsyncMissionStore.taskStore` is optional, every store
constructed without one relies on the legacy ids answering, and the
caller's two `resolveProjectColumnsForRoles(...).catch(() => undefined)`
calls mean each field can be undefined even *with* a store.

That last point is the distinction worth keeping: this is **not** the
`restart-recovery-coordinator` shape (#3059), where making a parameter
required deleted a production-dead fallback. Requiring it here would
force callers to fabricate a column set — inventing a vocabulary rather
than resolving one, which is the "guess" the fleet rules forbid.

## One mechanical note for future markers

**A marker only excuses the construct it precedes.** My first pass put
one comment above `isComplete` and moved 3 of 4 sites — `isArchived`,
two lines below, needed its own. Worth knowing before someone marks a
block and assumes it covered the siblings.

## Verification

- census: backlog 88 → 84, deliberate 128 → 132
- `agent-store-pause-marker-clear`, `agent-store-routing-policy`,
`mission-store.sync-auto-merge` — 18 tests green
- `tsc --noEmit` on `@fusion/core` clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
  * Clarified how task-column wording handles renamed board columns.
* Documented the fallback to “done” when terminal-column information is
unavailable.
  * No user-facing behavior changes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:34:05 -07:00
gsxdsm
95e4d1f246 fix(test): main is red — the archived-gate parity inventory is stale by two of my conversions (#3072)
## `main` is red

`archived-column-gate-parity.test.ts` fails on clean `origin/main`, on
the **raw-SQL** half. Not a branch artifact — reproduced by checking out
`origin/main` and running it alone.

## Both dropped sites are mine

| file | was → is | cause |
|---|---|---|
| `async-mission-store.ts` | 2 → 0 | **#3046** resolved
`archiveDefinedFeatureBootstrapDuplicate`'s two `<> 'archived'` guards
*together with* the `column: "archived"` write they gate |
| `task-store/async-archive-lineage.ts` | 3 → 2 | **#3042** deleted
`liveParentFilter`, an export with no callers anywhere |

Neither PR knew this inventory existed. The archived gate is enforced in
**three encodings** and only the census-visible one announces itself
when it moves.

Worth noting the mission-store conversion was *complete within its
function*: the `column: "archived"` write is a move **target**,
invisible to the column census, so converting the guards alone would
have been this file's split brain one level in.

## The total is re-recorded, not loosened

8 → 5, rather than relaxing to `toBeLessThanOrEqual`. A fixed total is
what makes a raw template **arriving** as visible as one leaving — and
this guard exists precisely because arrivals are what nothing else
counts.

## What this cost me, since it's the reusable part

Earlier this session I converted a TypeScript `archived` comparison in
`lifecycle-ops.ts` — the guard *and* its emit target together. **This
test caught it**, and its argument is right: converting one encoding of
the archived gate splits the brain, because the SQL halves still compare
the raw string. I reverted.

Then the test stayed red — for an unrelated reason. So the ratchet
simultaneously **stopped a bad conversion** and **was carrying a stale
number from two good ones**. Both halves of that are the ratchet
working; the second half is why a guard needs its inventory updated by
whoever moves it, not by whoever trips over it next.

## Measured

| check | result |
|---|---|
| parity suite | **red on `origin/main`**, green here |
| archived / lifecycle / parity suites | green |
| four gates + strict census | green |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:25:06 -07:00
gsxdsm
141f54e51d chore(core): mark the agent-store status formatter DELIBERATE-LITERAL (census 101→99) (#3063)
Fleet phase. `packages/core/src/agent-store.ts` was the **last** census
file with no branch, worktree, or open PR against it. Claim published by
pushing the branch before starting.

## Census before / after

| | total | this file |
|---|---|---|
| before | **101** | 2 |
| after | **99** | 0 |

`--strict` exits 0, baseline re-recorded. **Reclassification, not
conversion** — the line is unchanged.

## Already decided, in prose the census cannot read

The site was flagged earlier today by another pass, as `FLAGGED AND LEFT
COUNTED`: a pure formatter over `Pick<Task, "column">` with no store and
no task id, whose output is a human-readable status line. On a renamed
board it falls through to `(<column>)` — still accurate, just less
specific. Converting it would mean threading a lane resolution into a
string builder.

That reasoning is right and I did not revisit it. The only gap was
mechanical: a prose note is invisible to the tool, so the site kept
reading as backlog.

## This completes the sweep of unclaimed files

Third and last of these. Together with #3056 (fallback arms) and #3060
(dead sync path), **every census file that was unclaimed this phase has
now been examined, and not one of them needed a conversion.** Each was
either a three-state fallback arm — where the legacy id is the answer
when resolution fails, and removing it would break the caller — or a
site a previous pass had already reviewed and deliberately kept.

That is the finding worth carrying forward. The remaining **99** is not
a work queue: a meaningful share is correct code the tool cannot
distinguish from owed work, and every fleet pass pays to re-derive it.
Since all workers rank by the same `byFile` output, we also converge on
the same top file — which is how `self-healing.ts` drew three parallel
conversions, two of which are now unmergeable.

Two cheap changes would fix both symptoms:
1. **Mark reviewed-and-kept sites** so the count means *conversions
owed*. Two lines each.
2. **Push the branch at claim time** so `git ls-remote` is authoritative
before work starts. Costs nothing; I did it for all three of these.

## Verification

- `census --strict` exit 0; `tsc --noEmit` **0 errors**
- `check:fnxc-future-dates`, `check:lane-wiring` — exit 0
- Comment-only diff; no behaviour change

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:16:16 -07:00
gsxdsm
581e6fba43 chore(core): mark the async-mission fallback arms DELIBERATE-LITERAL (census 108→106) (#3056)
Fleet phase. Claimed `packages/core/src/async-mission-store-queries.ts`
— **the only census file with no branch, no worktree, and no open PR
against it.** Claim published by pushing the branch before doing any
work.

## Census before / after

| | total | this file | deliberate |
|---|---|---|---|
| before | **108** | 2 | 128 |
| after | **106** | 0 | **130** |

`--strict` exits 0, baseline re-recorded in the same commit.

**This is a reclassification, not a conversion.** The same two lines are
still there. A reader comparing 108 → 106 against my #3047's 126 → 121
should know only the latter changed behaviour.

## Why marking is the right answer here

Both sites are the **fallback arm** of the three-state rule:

```ts
terminalColumns?.complete ? terminalColumns.complete.has(column) : column === "done";
```

`terminalColumns` undefined means the caller could not resolve lanes.
The legacy id is then the only answer that keeps the query working at
all — converting it would delete the fallback and make an unresolvable
caller return nothing. The census counts the literal, but **the literal
is the design**.

The file's own comment shows a previous worker already reached this
conclusion. Nothing recorded it in a form the tool reads, so it stayed
in `byFile` as apparent backlog for the next pass to re-derive.

## The finding this makes concrete

I checked five unclaimed files this phase (`agent-store`,
`github-tracking-state`, `planner-overseer`, `auto-merge-finalization`,
this one). **Every site in them was either a fallback arm or an
already-documented deliberate leave** — `agent-store.ts:236` carries a
comment from today's fleet phase explaining why it stays.

So the remaining count is not a work queue. A meaningful share is
correct code the tool cannot distinguish from owed work, and each fleet
pass pays to re-derive that. Marking them is cheap, mechanical, and
makes the number mean "conversions owed" — which is what every worker
reads it as when picking a cluster.

I marked only the file I claimed. The others belong to whoever holds
them.

## Verification

- `census --strict` exit 0; `tsc --noEmit` **0 errors**
- `check:fnxc-future-dates`, `check:lane-wiring`,
`check:sql-column-literals`, `check:inert-flag-seams` — all exit 0
- No behaviour change: the two expressions are byte-identical, only
comments added

## Note on the marker's granularity

The first marker covered only `isComplete` — the census attaches markers
by *preceding comment*, so the sibling `isArchived` needed its own.
Caught by re-running the census (2 → 1, not 2 → 0) rather than by
reading. Worth knowing before marking a group of related literals.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:13:09 -07:00
gsxdsm
998d75da3b refactor(core): resolve the hand-off archive guard by role (fleet) (#3054)
## Census

| | column guards |
|---|---|
| before (this branch) | **122** |
| after | **118** |

Baseline re-recorded in the same commit, as the ratchet requires. (Four
of the delta land with #3052; this PR carries `moves.ts`.)

## What changed

`handoffToReviewImpl` refuses a hand-off from an archived card. Against
the literal `archived`, a board whose archive lane is renamed **never
matched** — so an archived card could be handed to review, and the
invariant `HandoffInvariantViolationError` exists to protect was
silently unenforced.

The IR is resolved at the guard rather than 28 lines below where
`handoffTarget` already reads it; the later read now **reuses** it
instead of resolving twice. The hoist is safe because this function has
already awaited `readTaskRowAsync` above — no new tick boundary. That's
the specific hazard blocking the scheduler cluster, so I checked it here
rather than assuming.

Absent or trait-free IR keeps the legacy id: unconverted boards are
byte-identical.

## Fleet intelligence: the backlog is now essentially fully triaged

I worked down the census top-files list and verified each before
writing. **Every remaining cluster is claimed, fallback-by-design, or
documented-blocked:**

| cluster | guards | status |
|---|---|---|
| `self-healing.ts` | 51 | **claimed** — checked out in another worktree
(`convert/self-healing-lane-cluster-u7`) |
| `scheduler.ts` | 12 | **blocked**, documented at line 907 —
`task:moved` prologue is synchronous; hoisting reorders this listener
against every other subscriber |
| `notification-service.ts` | 5 | **blocked**, documented — needs the
wedge-episode contract serialised first; the second site needs
gate-placement judgement in `handleTaskUpdated` |
| `executor.ts` | 4 | **claimed** (`fleet/executor-lifecycle-roles`) |
| `restart-recovery-coordinator.ts` | 4 | **trait-fallback arms** — the
census counts these as already converted |
| `taskRevert.ts` | 2 | **blocked**, documented — would classify a
*neighbour* task with the modal's own flags (the wrong-row shape, worse
than the literal) |
| `project-store-ops.ts` | 2 | **blocked**, documented — the dead SQLite
twin; its first statement throws under PostgreSQL |
| `github-tracking-state.ts`, `planner-overseer.ts`,
`async-mission-store-queries.ts`, `register-task-workflow-routes.ts` | 2
each | **trait-fallback arms** |
| `auto-merge-finalization.ts` | 2 | one is the `catch`-block degraded
fallback; the other is a reason string |

So the mechanical conversions are done. What's left needs either a
design change (scheduler's event payload, notification's episode
contract) or per-row lane data that doesn't exist at the call site yet
(`taskRevert`).

**That's the useful signal for the fleet**: further census reduction
isn't a matter of more conversion passes. Forcing these would produce
exactly the "conversions that break the code and improve the number" the
learnings doc is named for.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:07:19 -07:00
gsxdsm
107a1e790a fix(core): the review lane was resolved for two stall signals and literal for the other two (#3053)
## Claim

Largest **unclaimed** census cluster. `self-healing.ts` (56) is the
capacity worker's file and `scheduler.ts` (12) is blocked (below), so I
took **@fusion/core** — 16 bare guards across 12 files, untouched by any
open PR.

## Census before / after

```
before:  COLUMN guards (the backlog):   126
after:   COLUMN guards (the backlog):   126
```

**Unchanged, and that is the honest result — not a failed conversion.**
The repo's sanctioned device is an optional *resolved* parameter whose
default stays the legacy literal (the exemplar is
`restart-recovery-coordinator.ts`, watched by the unwired-lane-parameter
guard). The literal survives as the default arm, so the counter cannot
see the conversion.

**This matters for the fleet phase.** The census is not a progress meter
for this pattern. A worker driving the number down has only two ways to
move it, and both are wrong:

1. **Delete the fallback** (make the parameter required) — prior review
explicitly argued against this; `cli-active-count-lanes.test.ts`
deliberately covers the no-argument path.
2. **"Convert" with `resolveTaskWorkflowIrSync`** — that reader returns
`undefined` unconditionally under PostgreSQL, the shipped backend. It
drops the count while behaving *exactly* like the literal. That is the
inert-conversion class, and `merge-queue-ops-2.ts:53` already carries a
flag note saying so.

I measured the split across all 126: **18 are fallback arms of
already-converted seams; 108 are bare guards.** The headline number
conflates them.

## What changed

`reads.ts` states the invariant in its own words —

> RESOLVED BEFORE THE FIRST SIGNAL, because two adjacent signals must
not disagree.

— and then called two of the four stall signals with the literal:

| signal | before |
|---|---|
| `getInReviewStallReason` | resolved (`reviewColumns`) |
| `getInReviewStalledSignal` | resolved (`reviewColumns`) |
| `detectStalledReview` | **literal `"in-review"`** |
| `hasFreshAgentLogActivitySinceTaskUpdate` | **literal `"in-review"`**
|

On a renamed board `stalledReview` returned `undefined` for every card,
and the fresh-activity gate answered `false` — so `executingTaskIds`
stayed empty and the board showed Stalled / Merge stalled *while a
merger was visibly streaming*, the precise regression that function's
own FNXC note says it was restored to prevent.

Both now take an optional resolved `reviewColumns`. All four hydration
passes pass the set **they already had in scope one line away**; two
needed only a hoist, one reused the per-row map, one was resolving the
same set inline twice.

## Mutation evidence

| Mutant | Result |
|---|---|
| baseline | 11 passed |
| revert the detector guard to the literal | **2 failed** |
| make the parameter a widening (`reviewColumns ? true`) | **2 failed**
|

The second matters: it proves the new parameter is a real gate and not a
change that merely makes every card eligible. Both arms are asserted,
since the literal default is load-bearing for every caller outside
`reads.ts`.

## Flagged — do not guess

- **`scheduler.ts` (12 guards).** All 12 sit inside *synchronous*
listeners (`task:moved`'s sync prologue; `task:updated` is sync
outright). The only sync resolver available,
`resolveTaskParkedColumnsSync`, is already used at lines 929/1130/1157
and is **inert under PostgreSQL** — my own live-PG E2E proves it always
returns the default board. "Converting" these with it would drop the
census by 12 and change nothing. The existing note at 908–926 names the
real unblock: carry resolved lanes on the event payload so no listener
resolves at all. Left alone.
- **`restart-recovery-coordinator.ts` (4)** — already the
optional-parameter device with all three production callers passing
resolved answers. Not backlog.
- **`reads.ts:358`, `audit-ops.ts:208`, `task-id-integrity.ts:444`** —
`"archived"` here is the *cold-storage tier*, not the board column.
Trait resolution would be wrong.

## Verification

`test:gate` exit 0 · full `@fusion/core` unit suite **4880 passed** ·
typecheck exit 0 · `pnpm lint` clean · lifecycle-column census exit 0 ·
FNXC date ratchet exit 0 · lane-wiring census exit 0.

**One unrelated failure to report, not appeased:**
`src/__tests__/postgres/pg-test-harness-template-concurrency.pg.test.ts`
fails under the full suite and **passes in isolation on both my tree and
the untouched baseline** — a pre-existing full-suite concurrency flake
in the PG harness. Not mine, not in the merge gate. I did not quarantine
it: it is another worker's harness, and AGENTS.md warns that
quarantining a concurrency test can mask a real product race. Flagging
for its owner.
2026-07-31 02:39:33 -07:00
gsxdsm
4878bda197 fix(core): the mission bootstrap duplicate was archived into a lane the board does not declare (#3046)
## Invisible to both censuses

`archiveDefinedFeatureBootstrapDuplicate` writes `tasks.column`
**directly** rather than through `moveTask`:

```ts
.set({ column: "archived", updatedAt: … })
```

- the **lifecycle census** reads comparisons — an assignment isn't one
- the **move-target census** reads `moveTask` call arguments — this
never calls it

So on a board whose archive lane is renamed, the duplicate landed in a
column that workflow doesn't declare: a card in a lane the board can't
render, from a path that runs during ordinary feature bootstrap.

## Reuses the helper this class already has

`archivedLanesFor(taskId)` was added for the guards further up the same
file. It returns the legacy id when the task has no resolvable workflow,
so an **unconverted board is byte-identical**. No new resolution
machinery — the two `<> 'archived'` guards become `notInArray(column,
[...lanes])` and the write targets the resolved lane.

A board declaring several archive lanes is arbitrated by taking the
first, the same choice `resolveLifecycleColumns` makes. Multiple archive
lanes aren't a shape the builtin lineages produce.

## Measured

| check | result |
|---|---|
| mission-store PG suite | **36 → 38**, all green |
| new pair | differential — `filed` collides with no legacy id, and the
default-lineage control still lands in `archived` |
| mutation (hardcode the target back) | fails the renamed case |
| SQL literal gate · `tsc` | green |

## How this was found

Measuring the literal-column-**write** population for #2839: 51 raw
sites, of which 20 are the four builtin workflow IRs declaring their own
columns (correct by definition) and several more are archive-*entry
record* fields rather than board columns. This is the one I verified is
a real board write on a live path.

Worth noting the measurement itself was wrong twice first — my glob was
`packages/*/src/**/*.ts`, which requires a subdirectory and silently
skipped every top-level file in `src/` (including this one), and my
script printed only the first 14 findings so the grouping was over a
truncated list. Same scope-blindness class as #3000 and #3002, this time
in a throwaway scanner.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:22:27 -07:00
gsxdsm
8b82e77fbf chore(core): delete liveParentFilter — no caller, and it carried a legacy lane literal (#3042)
Found while enumerating archive-exclusion sites for #3041.

## Unambiguously dead

`liveParentFilter` has exactly **one** reference in the repo: its own
definition.

- not exported from `index.ts` or `index.gate.ts`
- no test imports it
- no production code calls it

It nonetheless contained `column != 'archived'`, so it was one of the 22
sites the SQL column-literal gate tracks.

## Why delete rather than convert

Converting it would mean adding lane resolution to code nothing runs —
risk with no behaviour. That's the same argument #3041 makes for *not*
converting the other two dead sites; deleting is the version of it that
also removes the literal.

## The gate it documents is not being deleted

Its docblock describes the document/artifact visibility gate
(VAL-CROSS-015). That gate is real and still enforced — by the inline
conditions inside `listLiveTaskDocuments` and `listLiveArtifacts`, which
is presumably why this helper was never wired up in the first place.
Only the unused composition goes.

## Measured

| check | result |
|---|---|
| SQL literal population | **22 → 21**; the gate ratcheted its own
baseline down and asked for the commit, included here |
| `taskstore-remaining.test.ts` (archive-lineage suite) | **27 tests
green** |
| six gates + `tsc` | green |

## Not deleted, deliberately

`listLiveTaskDocuments` and `listLiveArtifacts` are referenced **only**
by that test file. That's a weaker signal than zero references — someone
may have written them ahead of a consumer. Their literals stay counted,
which is the honest state for code whose intent I can't read from the
repo.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:11:54 -07:00
gsxdsm
511f5b7e2b fix(core): archived tasks leaked into the live feed on a renamed board (#3041)
From my own #2839, re-measured today. Of that issue's SQL-literal sites,
this is the one that decides what a **live view** shows.

## The defect

`listTasksModifiedSinceImpl` backs the SSE watcher and modified-since
polling — the incremental feed the dashboard applies to its task list.
Its `includeArchived: false` branch excluded the literal `archived`:

```ts
conditions.push(sql`${schema.project.tasks.column} != 'archived'`);
```

On a board whose archive lane is named anything else, that predicate
matches **every** row and excludes nothing. Archived cards arrive in the
live feed and reappear on the board.

Nothing errors, and a full refetch filters archived rows by another path
— so the symptom is archived work that comes back until the next reload.
That gets reported as *"the board is flaky"*, not as a bug.

## The fix

`resolveProjectColumnsForRoles` seeds the legacy ids before adding
resolved ones, so the set is never empty and an **unconverted board
excludes exactly `archived` as before**. The literal stays as the
resolution-failure fallback, where excluding nothing would be worse than
excluding the legacy id.

## Surface enumeration — three of four sites are dead

Four sites share this invariant. Verified rather than assumed:

| site | status |
|---|---|
| `reads.ts:558` (SSE / modified-since) | **live** — converted here |
| `liveParentFilter` | **no references anywhere** in `packages/` or
`plugins/` |
| `listLiveTaskDocuments`, `listLiveArtifacts` | referenced **only** by
`taskstore-remaining.test.ts` |

That's why this PR converts one site rather than four — the other three
are production-dead, and converting dead code would add risk for no
behaviour.

## Measured

| check | result |
|---|---|
| new PG suite | **4 cases** — legacy control, the renamed defect, a
live-lane negative, and the forensic `includeArchived: true` read |
| mutation (force the legacy fallback) | fails **exactly** the renamed
case; the other three hold |
| six gates + `tsc` | green |

The negative case is the one that matters most: resolving the archive
role must not start excluding **live** work, or the board silently stops
updating for real tasks — a worse failure than the leak this fixes.

## One process note

I corrupted this file mid-session by mutation-testing it while
uncommitted: a failed restore left a half-applied block, and a later
`git checkout --` discarded the fix entirely. Both were caught by
re-grepping for the symbol rather than trusting the restore. The
reliable pattern is **commit first, then mutate, then `git checkout` to
restore** — which is how the proof above was actually run.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:08:49 -07:00
gsxdsm
f8155cafd7 fix(cli): the node-override guard saw only the FIRST wip lane (#3023)
Follow-up to #3019, which merged with an incomplete fix. I found this
while sitting down to write the test that PR was missing.

## The guard still never fired, one lane over

#3019 wired `fn_task_update`'s guard like this:

```ts
const nodeOverrideLifecycle = await resolveTaskLifecycleColumns(store, task.id);
wipColumns: nodeOverrideLifecycle?.wip ? new Set([nodeOverrideLifecycle.wip]) : undefined,
```

`resolveTaskLifecycleColumns` → `resolveLifecycleColumns`, whose
per-role accessor is **first match**
(`workflow-lifecycle-traits.ts:353`):

```ts
const first = (flag) => resolved.find((c) => c.flags[flag] === true)?.id;
```

The guard's contract is **every** column carrying the trait — its own
resolver uses `columnsWithFlag(ir, "countsTowardWip")`. So on a board
with a build lane beside a verify lane, a task sitting in the **second**
wip lane still slipped the mid-flight check, and an operator could still
repoint the node of a running task. That is the defect #3019 set out to
close.

Interchangeable on any single-wip-lane board, which is exactly why it
read as correct — the same arity trap #2975 removed from the surfacing
family.

## The fix

Use `resolveNodeOverrideLanes`, the guard's own resolver, which
`task-update.ts` and `branch-and-pr-entities.ts` already call. All three
callers now resolve identically and the V1/unresolvable fallback lives
in one place. Needed a one-line re-export from `@fusion/core`.

**Mutation:** forcing the resolver to first-match (`.slice(0, 1)`) fails
the new case, 1 of 32.

The new test names **two** wip lanes, because that is the only shape
that separates the two resolutions — a single-wip-lane test passes
against both, which is why #3019's gap was invisible and why I would
have written a useless test if I had not read the implementation first.

## A gate constraint worth recording

My first version passed the resolved object straight through:

```ts
validateNodeOverrideChange(task, normalizedNodeId ?? null, overrideLanes)
```

Identical at runtime, and it turned the lane-wiring gate **red**:
`check-lane-wiring` matches an object-literal argument and cannot see
through a variable, so the correct call reads as UNWIRED. #3019's header
records hitting the same constraint — and it is what pushed that PR
toward resolving the lanes inline, which is where the first-match bug
entered.

So the gate's shape requirement steered a correct instinct into a subtly
wrong implementation. The fix here spells both keys explicitly,
satisfying the gate without the bespoke resolution. Worth someone
deciding whether the census should follow a variable to its initializer
— but that is a change to a shared ratchet, and I have noted it at the
call site rather than making it.

**Verified:** 32/32 core guard suite, `tsc` 0 errors for both packages,
lane-wiring gate exit 0, FNXC gate exit 0, lint clean.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 01:08:54 -07:00
gsxdsm
16921fc518 fix(engine,core): role resolution was half-done in two shared lifecycle predicates (surfacing family + file-scope leases) (#2975)
The three surfacing sweeps stopped reporting anything for a card resting
in a board's **second** review or hold column.

A lifecycle role is a **trait**, and any number of columns may carry it.
The shared runner resolved it with `resolveLifecycleColumns()[role]` —
**first match** — then gated on it:

```ts
const roleColumn = lifecycle?.[spec.role];        // FIRST column carrying the trait
if (task.column !== resolved.roleColumn) continue; // everything else dropped
```

A workflow that splits human sign-off from the merge lane has two review
columns; one that parks dependency-blocked cards separately has two hold
columns. Cards in the second got **no stale-paused-todo, no
stale-paused-review, no in-review-stalled** diagnostic — silently, with
no error, on all three sweeps at once.

## The second bug hiding inside the fix for the first

Resolving membership but still reading `roleColumns[0]`'s declared
`recovery` applies the **merge lane's** threshold to a card sitting in
the **sign-off** lane. Each card's policy now comes from its own column,
and one of the new cases fails if it doesn't: the first role column
declares a policy that suppresses the signal, the card's own column
declares one that fires.

## Reverted

| | |
|---|---|
| **6 of 12** new cases fail | `fires for a card in the SECOND column
carrying its role` and `reads the recovery policy of the card's OWN role
column` — × 3 sweeps |
| the other 6 pass either way | non-regression halves: still fires for
the FIRST role column, still does **not** fire for a card outside every
role column. Membership must widen the gate, not move it. |

The pre-existing 45 cases were all green throughout — the
single-role-column fixture could not express the case, which is why the
table-driven file that exists to stop these three sweeps drifting apart
never caught it.

## Verification

`pnpm test:gate` 161 + 13 + 487 + 71 · surfacing family 57 · core
stale-paused 20 · lint · census `--strict` · sql-literals · fnxc-dates ·
lane-wiring · changesets — all green.

## Note

`holdColumns` was missing from the lane-wiring vocabulary, so the gate
could not see that argument dropped. Added in the same commit.

While reviewing, I found and measured **two problems in #2974** (comment
posted there): six of its newly-visible sites are `satisfies`-wrapped
false positives, and baselining them means deleting a real
`reviewColumns` argument keeps the count unchanged and the gate green;
and its baseline predates #2970, re-opening the slot that PR closed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved stale-card detection across all applicable review and hold
columns.
* Cards are now surfaced using the policies configured for their
specific lifecycle column.
* Cards outside matching lifecycle columns are no longer incorrectly
surfaced.
* Preserved existing fallback behavior when no lifecycle columns are
configured.

* **Tests**
  * Added coverage for workflows with split review and hold columns.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

## Second commit: the same predicate, half-converted
(`shouldHoldActiveFileScopeLease`)

Folded in here rather than stacked — same file, same class, and a
stacked PR on an unmerged base is not mergeable. Reversible; say the
word and I'll split it.

`shouldHoldActiveFileScopeLease` is the **scheduler's** lease predicate,
shared with the self-healing repair paths deliberately so the two cannot
disagree about who holds a file-scope lease. Its two role answers are
optional parameters defaulting to the legacy ids. The scheduler's own
call sites were converted to pass resolved answers; self-healing's two
were not:

```ts
const isWipColumn    = options?.isWipColumn    ?? task.column === "in-progress";
const isReviewColumn = options?.isReviewColumn ?? task.column === "in-review";
```

On a renamed board neither branch matches, so the predicate returns
`false` for every card. The scheduler kept the lease; self-healing saw
none, cleared `overlapBlockedBy`, and **released a dependent to edit
files another agent still holds** — the outcome `groupOverlappingFiles`
exists to prevent.

Membership comes from the wip/review sets each sweep already resolved a
few lines above, so this adds no reads.

**Reverted:** both new cases fail with `overlapBlockedBy` = `null` — the
release itself, not a proxy. The pre-existing legacy-column case in the
same file passes either way, because `in-progress` satisfies the literal
default; that is exactly why it never caught this.

Lane-wiring baseline re-recorded `9 -> 7` in the same commit (the
ratchet refused a stale allowance, as intended).

**Verification:** gate 161 + 13 + 487 + 71 · surfacing 57 · overlap-seam
+ scheduler-lease + query-blindness 79 · core stale-paused 20 · lint ·
census `--strict` · sql-literals · fnxc-dates · changesets — green.
2026-07-30 23:13:34 -07:00