Commit Graph

12958 Commits

Author SHA1 Message Date
gsxdsm
851369a480 docs(solutions): record "silence is not success" (#3241)
## What

A new `docs/solutions` note recording a failure that hit **three
different tools in one session**, each time reading as a pass. Docs
only.

## The three costumes

| what happened | looked like | was |
|---|---|---|
| `git stash --keep-index` swept the new test file out of the tree | "45
passed" | the pre-existing count; the new test never ran |
| a blinding script hit an unmapped role and `sys.exit(2)` **with no
message**; `&&` skipped the check, `;` let the run proceed | "375/375
green under blinding" | nothing blinded — run was against unmodified
source |
| a gate piped to `tail -1`, printing a blank line | "gate ran, no
complaints" | exit code 1; the FNXC stamp check had failed, and **CI
caught it in #3238** |

## Why it deserves its own note

**A passing run and a run that never happened produce the same evidence:
no failure text.** Every other bug announces itself; this one is defined
by the absence of an announcement. The instinct that catches ordinary
bugs — *"nothing looks wrong"* — is precisely the instinct that
certifies this one.

It gets worse under automation, where output is piped and skimmed. `|
tail -1`, `| grep "Tests"`, `>/dev/null 2>&1` all discard the part that
would have said `No test files found` or `command not found`.

## The five rules, each paid for above

1. **Assert the exit code before any pipe.** A pipeline's status is the
*last* stage's — `cmd | tail -1` reports `tail`'s success, never
`cmd`'s.
2. **Confirm the run did the work.** "Test Files 1 passed" when you
expected 16 is a finding, not a pass.
3. **A tool that can no-op must say what it did** — print the
substitution and location, fail loudly where it cannot act.
4. **Verify the mutation, not the tool's promise** — `git diff --stat`,
not the exit code.
5. **Break the guard on purpose once** and watch it fail. A guard never
observed failing has not been shown to work — the standard this repo
already applies to product ratchets, turned on your own verification.

## The uncomfortable part, kept in

The third instance was a rule **I added to AGENTS.md myself in #3174**,
broken for the second time. I ran the gate. I read `tail -1`. I moved
on.

Writing a rule down does not make you follow it. The only reason it was
caught is that **CI read the output when I did not** — an argument for
the gate existing, not for me having been careful.

Cross-linked from the resolver-audit note, whose every wrong reading
came from a run that never happened rather than from the blinding
itself. That connection is the point: I spent this session auditing a
program whose subject is defects hiding behind green results, and
reproduced the same class three times in my own tooling.

```
lint clean; fnxc-future-dates: none added (exit code checked before piping this time)
```


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added workflow guidance explaining why silent or seemingly successful
output does not confirm that a test, script, or validation gate ran.
* Documented verification practices including checking exit codes, work
counts, no-op detection, post-run changes, and intentional failure
checks.
* Added a case study highlighting how filtered output can conceal
verification failures.
* Added cross-references connecting resolver interpretation, test
execution, and conversion coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 13:47:08 -07:00
gsxdsm
efd8454b6c fix(planning): the add-comment trigger sat below the fold — the sheet outgrew its floating host (#3242)
Fixes the red `main`: `planning-browser-e2e.test.ts > places the sole
contextual comment trigger by viewport in embedded and modal Planning`.

**Unclaimed and not a flake.** `check-file-claimed` reported UNCLAIMED,
it is not in the quarantine ledger, it reproduced locally and
deterministically, and it failed identically across three consecutive
Full Suite runs. Quarantine would have been the wrong instrument — that
rule is for flakes, and appeasing a consistent failure buries a real
regression.

## Root cause

`PlanningModeModal.css` sizes dialog Planning as a **full-viewport
sheet**:

```css
.planning-modal:not(.planning-modal--embedded) {
  height: 100dvh;
  min-height: 100dvh;   /* ← beats max-height: 100% */
  max-height: 100%;
}
```

That was correct until the modal branch moved **inside
`FloatingWindow`** (`FNXC:ModalTouchGeometry 2026-07-26-14:10`). The
floating host's body is shorter than the viewport — it sits below a
title bar — so the rule now asks the sheet to be *taller than the box
containing it*. `min-height` wins over `max-height`, so the sheet cannot
shrink to its host and overflows.

Measured by walking the ancestor chain at 768×900:

```
BUTTON.btn                 top=928 h=36        ← 28px past the fold
DIV.planning-actions       top=919 h=101
DIV.modal                  h=900               ← forced to full viewport height
DIV.floating-window__body  h=763  sh=900       ← host is 763 tall, content is 900
DIV.floating-window        h=765
```

The "Add comment to selection" control needed a scroll to reach —
exactly what the placement case exists to prevent.

## The fix

A scoped override under `.floating-window`, rather than editing the
sheet rule, so Planning rendered **outside** a floating host keeps its
full-viewport sizing:

```css
.floating-window .planning-modal:not(.planning-modal--embedded) {
  height: 100%;
  min-height: 0;
}
```

## Surface enumeration

Embedded Planning was **never affected** — it is excluded from the sheet
rule, and all four embedded viewports passed throughout. The failure was
modal-only, at every modal viewport (768, 769, 1024, 1280 — it fails
fast at the first).

**No new test.** The existing placement case already asserts this
invariant across **4 viewports × 2 presentations = 8 combinations**,
which is the surface enumeration for this affordance. It was red; it is
now green. Adding a narrower repro-only test would be the anti-pattern
the Fix-the-Invariant rule names.

## A disproven hypothesis, recorded

`min-height: 0` on `.planning-plan-review > .planning-plan-pane` — the
canonical flex-overflow fix, and a pattern used 10+ times in this very
file — **does not fix it**. Measured, not assumed. The overflow is one
level up, at the sheet/host boundary. Noted so the next reader does not
repeat the experiment.

## Verification

```
fix applied      Tests  5 passed (5)
fix reverted     Tests  1 failed | 4 passed (5)     ← the test genuinely holds this fix
fix restored     Tests  5 passed (5)
```

Neighbours green: **57 tests across 8 suites** (mobile
footer/bottom-space/pan-containment, terminal keyboard layout,
task-detail tablet width, mission planning modals mobile, mobile
planning input font size, task-detail floating geometry) plus **9**
planning e2e.

Changeset included (`patch`, category `fix`) — this is user-visible
dashboard behaviour.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 13:44:25 -07:00
gsxdsm
f31a716a2a FN-8629: prevent false Grok usage percentages
Prevent omitted Grok billing percentages from being displayed as fully consumed credits.

- Require a finite CLI-supplied credit usage percentage before creating a billing window.
- Cover omitted, zero, invalid, and non-weekly Grok billing responses.
- Add a patch changeset for the corrected usage display.

Files changed:
 .changeset/fn-8629-grok-usage-percent.md       |  7 +++
 packages/dashboard/src/__tests__/usage.test.ts | 71 ++++++++++++++++++++++++--
 packages/dashboard/src/usage.ts                | 13 ++---
 3 files changed, 77 insertions(+), 14 deletions(-)

Fusion-Task-Id: FN-8629

Fusion-Task-Lineage: b5b7c83b-e34f-43d9-a31a-d1fd769c4eb8

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-31 13:43:28 -07:00
gsxdsm
4bcaddafc5 docs: index workflow-owned-lifecycle-closing-verification.md in Audit Reports
Found during routine docs orphan scan. The file (added 2026-07-30,
commit 13bf7e001d) was not linked from the README.md index. It is a
verification runbook and recorded-pass history for the workflow-owned
lifecycle cutover programme.
2026-07-31 13:43:27 -07:00
gsxdsm
56e16d9dea test(cli): pin the board glyph's terminal-lane resolve (extract seam + pin) (#3238)
## What

Pins the CLI board glyph's terminal-lane resolve — **the last flagged
site in the repo-wide resolver audit.**

Two commits: a behaviour-preserving extraction, then the test.

## I was wrong to flag this as unpinnable

In #3236 I recorded this site as not pinnable, reasoning that
*"extracting a pure helper and testing it would look like coverage and
would not be."*

That is true of a helper that **receives** the lane set — such a test
passes with the resolve blinded, which is exactly the `reads.ts` trap
the audit note records. It is **not** true of one that **resolves** it.
Building `resolveReliabilityLanes` in #3237 made the distinction
obvious: the seam has to contain the resolve, and then blinding fails a
test of it.

So the flag was too broad, and correcting it closes the site rather than
leaving a permanent excuse. That is the same failure mode I corrected in
someone else's note earlier today — a caution that hardens into a reason
not to look.

## Measured

```
converted: Tests 5 passed (5)
blinded:   Tests 2 failed | 3 passed (5)
```

The two failures are the **renamed complete** and **renamed archive**
lanes. The three survivors are the default-vocabulary control, the
active-lane negative, and the degrade path — all of which should
survive.

```
task-list-board-columns + bin: 82 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```

## Why the sibling file did not cover it

`task-list-board-columns.test.ts` pins `boardColumnsForDisplay`, which
decides **which** lanes print. That function takes no lane set, so it
cannot fail when this resolve is blinded — and its own header says so
honestly. Two tests about the same command, one of which cannot see the
other's bug.

## What breaks without the conversion

On a board whose complete lane is `shipped`, a finished lane renders `●`
— the same glyph as active work. The board says work is in flight when
it shipped. Cosmetic next to the blank-board bug this area already
fixed, but wrong in the direction an operator reads at a glance.

## Also pinned

Two contracts the surrounding comments assert but nothing tested:

- **Cards come from the TASKS, not a resolved IR** — a card must never
depend on resolution succeeding to be *visible*. Asserted with an
unreadable workflow list.
- **A failed resolve degrades to the legacy pair**, with an unresolved
custom lane rendering as active — the documented fail-open direction.

Plus the paired negative: an ACTIVE lane keeps the active glyph under
both vocabularies, so widening the terminal set cannot mark the whole
board finished.

## Audit complete

Every `resolveProjectColumnsForRoles` call site in the repository —
`engine`, `core`, `dashboard`, `cli` — has now been blinded
individually, and every uncovered one is either pinned or has a recorded
reason it cannot be. Nothing is left flagged.
2026-07-31 13:34:03 -07:00
gsxdsm
0698ce6f9c test(dashboard): restore the missing api mock export in ResearchView tests (#3239)
## What

Partial fix for **red main**. Test-only.

`ResearchView.test.tsx` has **4 failing tests on main**; 2 fail with:

```
No "fetchBoardWorkflows" export is defined on the "../../api" mock
```

The `vi.mock("../../api")` factory **replaces the whole module**, so
every import anywhere in the rendered tree must appear in it.
`fetchBoardWorkflows` reached this file *indirectly* — the task modals
ResearchView opens import it — so adding that export to product code
broke four cases that have nothing to do with board workflows.

Stubbed with the flag-OFF payload the server sends when multi-lane
boards are disabled, which is the shape these cases already assume.

## Measured

```
before: Tests 4 failed | 23 passed (27)
after:  Tests 2 failed | 25 passed (27)
lint clean; fnxc-future-dates: none added
```

## The remaining 2 are a different cause and are NOT fixed here

They fail with `Number of calls: 0` — the enrich-task and create-task
actions never fire. That is a UI-wiring question, not mock completeness.

**I checked that my stub is not responsible**, rather than assuming:
re-running with `flagEnabled: true` and a populated workflow list
produces the *same* 2 failures, so the payload shape does not gate those
affordances. Left for whoever owns that surface.

## How this was found

While establishing a clean baseline for the resolver audit. That sweep
also reported `lazy-loaded-views-docs.test.ts` red — **it now passes**,
fixed by another worker between my measurement and this PR, which is why
the count here is 3 files rather than the 4 I reported in #3236.

Still red on main, untouched by this PR:
- `src/__tests__/planning-browser-e2e.test.ts` — `expected {
totalButtons: 1, …(7) } to match object { totalButtons: 1, …(6) }`; an
assertion shape gained a field.
- `src/__tests__/register-model-routes-kimi-k3-supplemental.test.ts` —
`Test timed out in 15000ms`. Per the standing rule a timeout with no
corresponding bug in the change is a **quarantine candidate**, not
something to appease with a longer timeout; I am not quarantining it
unilaterally since it is not my subsystem, but flagging it as the shape
that rule describes.
2026-07-31 13:33:51 -07:00
gsxdsm
05f09c29f8 docs(solutions): complete the repo-wide resolver audit; correct a superseded note (#3236)
## What

Completes the repo-wide resolver audit and **corrects a note of mine
that had gone stale**. Docs only.

Every `resolveProjectColumnsForRoles` call site in the repository has
now been blinded individually.

## Final results

| package | sites | outcome |
|---|---|---|
| `engine` | 10 files | scheduler, triage, evaluator uncovered → pinned;
executor, restart-recovery, notification already covered; self-healing
21 pinned / 1 inert |
| `core` | 14 | 9 covered, **5 uncovered → all 5 pinned** (#3225, #3227,
#3233, #3234, #3235) |
| `dashboard` | 4 | `register-task-workflow-routes.ts:1268` covered;
`server.ts` ×3 flagged |
| `cli` | 1 | flagged |

## The correction

A note recorded `workflow-analytics.ts` and `team-analytics.ts` — 4
resolvers — as **unmeasurable**, because `pgDescribe` probes TCP and the
`.pg` suites skip without it.

The caution is real and stays: a skipped suite reads exactly like a
passing one. But on an environment where those suites **do** run, all 4
were measured, and `team-analytics.ts` turned out to have a half-covered
pair — `completeLanes` covered, **`activeLanes` not** — in a file named
`team-analytics-renamed-lanes`. That is now pinned (#3227, merged).

Left standing, the note converts a real finding into a **permanent
excuse for not looking**. It now says: confirm the suite actually skips
*here* before recording a site as unmeasurable for environment reasons.

## A fourth measurement failure mode — the opposite direction

The three already recorded all produce false *uncovered*. This one
produces false *covered*:

**A COVERED verdict needs a baseline.** The dashboard sweep reported 5
failing files under the global blind. **4 of them fail on clean `main`**
and have nothing to do with lanes — a docs-inventory test and a
model-routes test among them. Read as-is, that is four resolvers falsely
credited as covered. Only
`register-task-workflow-routes.awaiting-planning.test.ts` passes clean
and fails blinded, so it is the sole real detector.

Second time today a baseline changed a conclusion (the first found a
genuine red on main, #3229).

## Why 4 sites are flagged rather than pinned

- **`server.ts:1922/1923/1938`** — inside the `/api/health/reliability`
route closure. No route-level test exists, and the only way in is
booting `createServer(store)` behind a mock-the-world shell, which the
slow-test rule forbids. The alternative is a refactor to expose a seam —
its own commit, since moving code and changing behaviour do not ride
together. (A note already in this doc reached the same conclusion
independently; this confirms it by measurement.)
- **`cli/commands/task.ts:660`** — worth its own warning. Extracting a
pure helper and testing it **would look like coverage and would not
be**: blinding the resolver leaves such a test green, because the helper
*receives* the lane set rather than resolving it. The uncovered thing is
the resolve call, not the decision it feeds. Its sibling test file
already records the same limit honestly for `boardColumnsForDisplay`.

## Reported, not fixed: 4 pre-existing red dashboard files on main

`lazy-loaded-views-docs.test.ts` (AGENTS lazy-view inventory drifted —
24 actual vs 18 documented), `ResearchView.test.tsx`,
`planning-browser-e2e.test.ts`,
`register-model-routes-kimi-k3-supplemental.test.ts` — 7 failing tests,
all in the non-blocking suite.

I am not fixing them here: the lazy-views inventory is a curated list
other workers are actively adding to, and rewriting it mid-flight would
collide. Flagging so it is visible rather than silently absorbed into my
blind's noise.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated workflow guidance to require baseline comparisons and
verification that all relevant tests run.
* Added safeguards for detecting ineffective changes and distinguishing
pre-existing failures.
* Expanded PostgreSQL audit documentation with measured coverage
results, including uncovered resolver paths.
* Recorded completed coverage sweeps across core, dashboard, and CLI
areas, including pinned and non-pinnable sites.
* Clarified limitations when testing extracted decision helpers instead
of resolver calls.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 13:23:31 -07:00
gsxdsm
476c5c360c test(dashboard): pin the Reliability endpoint's three lane reads (extract seam + pin) (#3237)
## What

Pins the **Reliability endpoint's three lane reads** — the last
uncovered resolver cluster the repo-wide audit found.

Two commits, deliberately separate:
1. **refactor** — extract the three resolves behind
`resolveReliabilityLanes(store)`. Behaviour-preserving, no test changes.
2. **test** — pin all three through that seam.

## Why a seam was needed

The three resolves lived inline in the `/api/health/reliability` route
closure. Blinding any of them left the **entire dashboard suite green —
21,582 tests** — and the only way to reach them was booting
`createServer` behind a mock-the-world shell the slow-test rule forbids.

**And the obvious test would not have helped.**
`reliability-metrics.test.ts` exercises `countEntriesInto`,
`countBouncesOut` and `inReviewDurationMetrics` with lane sets **passed
in by hand**. That proves the collaborators honour a resolved set; it
says nothing about whether the caller passes one. *A unit test of the
collaborator can never fail when the caller's resolve is blinded* — the
same trap the audit note records for `reads.ts`, where a suite written
for the exact conversion still could not see it.

The seam is the caller. It resolves, so blinding a resolve fails a test
of it.

## Measured — each blind fails exactly its own case

| blinded | fails |
|---|---|
| `REVIEW_ROLES` | "resolves the board's OWN review lane" |
| `["countsTowardWip"]` | "resolves the board's OWN wip lane" |
| `["complete"]` | "resolves the board's OWN complete lane" |

```
converted: Tests 6 passed (6)
each blind: Tests 1 failed  (its own case only)
reliability-metrics.test.ts + this file: 28 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```

That isolation is the point: **three resolves in one function invite a
copy-paste that hands the same set to all three**, and every positive
assertion would still pass. There is a paired negative asserting each
renamed lane appears in *its* bucket and nowhere else — without it the
duration metric could silently measure review → review.

Also pinned: the degrade path. An unreadable workflow list must not fail
the endpoint, so the legacy ids still answer.

## What breaks without the conversion

On a board that renames either lane, every underlying query returns `{}`
— so `tasksEnteredInReview` and `tasksBouncedToInProgress` are zero for
every day, and `inReviewFailureRate7d` divides one zero by another and
reports a **healthy** rate. It produces a NUMBER, not an error, and the
number says everything is fine. An operator reading 0% review failures
beside a populated audit list has no reason to suspect the metric is
blind.

## The one observable difference in the refactor, stated not buried

The complete-lane read moves from *after* the counting `Promise.all`
into the same phase as the review/wip pair. These are pure reads of
workflow definitions — no writes, no ordering dependency — so the
resolved values are identical; only the concurrency shape changes (three
parallel reads instead of two-then-one). Flagging it because
"behaviour-preserving" should be a claim someone can check, not an
assertion.

## Audit status

With this, **3 of the 4 flagged sites are closed**. Remaining:
`cli/commands/task.ts:660`, where the glyph decision is inline in
`runTaskList` and the same seam argument applies — but its sibling test
file already documents that driving that function needs the forbidden
shell, and extracting a helper there would produce a test that *looks*
like coverage while leaving the resolve unpinned. Left flagged rather
than faked.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Reliability health metrics now recognize configured review,
work-in-progress, and completion lanes, including renamed workflow
lanes.

- **Bug Fixes**
- Improved fallback behavior when workflow definitions are unavailable,
preserving compatibility with legacy lane configurations.
- Ensured lane resolution remains isolated by role for more accurate
reliability metrics.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 13:23:20 -07:00
gsxdsm
f0a13745b2 fix(test): lazy-views doc parser stops at any heading — six phantom views came from the H2 that follows
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 13:09:55 -07:00
gsxdsm
fb8d37e30d test(core): pin the last two uncovered lane reads (mission archive, lineage gate) (#3235)
## What

Pins the **last two uncovered lane reads** in `packages/core`.
Test-only. This closes the per-site core audit.

| site | what it decides |
|---|---|
| `async-mission-store.ts:1179` | is an ARCHIVED card valid terminal
evidence for mission repair? |
| `task-id-integrity.ts:502` | does an archived child still count as a
LIVE lineage child? |

## Measured

```
mission-store:  39 passed clean;  1 failed | 38 passed blinded
lineage:         3 passed clean;  1 failed |  2 passed blinded
lint clean; fnxc-future-dates: none added; census unchanged
```

Both blinds confirmed applied with `git diff --stat` before each run.

## The third adjacent-pair split

`:1179` is the **archived** half of a pair whose **complete** half
(`:1178`, *one line above*) was already covered by a test in the same
file, written for exactly this concern. Terminal evidence is "done OR
supported archived state," so an archived card is equally valid repair
evidence — but on a board whose archive lane is `vaulted` the archived
half could not see it, and reconciliation threw `TASK_NOT_TERMINAL` for
a card that was genuinely filed away. Same refusal the covered case
fixed, reached through the other door.

That is now the third confirmed instance in core (after `team-analytics`
in #3227 and the scheduler pair earlier). **Being adjacent to a covered
resolver is not coverage**, and it is the most reliable place to look.

## What breaks without the lineage read

An archived child is filed away, not live, so it must not hold the
delete gate shut. Renamed, it still counted as live and
`TaskHasLineageChildrenError` blocked the parent's delete **forever** —
the operator archived the child *precisely* to clear the way, and the
gate could not see that they had.

## A fixture detail I got wrong first

My first mission fixture created a live card in a `vaulted` column and
failed with `deleted or archived without a valid retained tombstone and
archive snapshot` — nothing to do with the lane read.

The `archived` verdict requires **all three** of `deletedAt !== null`,
an archive-snapshot row, and `isArchived(column)`. A live card merely
sitting in an archive-trait column is `invalid-deleted`, not `archived`.
The test now archives for real and *then* renames the recorded lane,
which isolates the third condition — the only one under test. Recorded
in the file so the next person does not re-derive it.

## Paired positives

Both files pin the complement: a WORKING child still counts as live.
Recognising the renamed archive lane must not degrade into "no child is
ever live" — that would silently **disable** the lineage gate and let a
parent be deleted out from under real descendants, which is worse than
the bug being fixed.

## Core audit complete

**14 sites blinded individually: 9 already covered, 5 uncovered, all 5
now pinned** (#3233, #3234, this PR).

Every `resolveProjectColumnsForRoles` call site in `packages/engine` and
`packages/core` has now been blinded. Remaining unaudited: `dashboard`
(2 files) and `cli` (1) — I claim nothing about those.
2026-07-31 13:00:22 -07:00
gsxdsm
623581837a fix(engine): mock provider sends 0-based steps — test mode full-task runs complete again (#3231)
Found by a live browser E2E of the coding workflow in test mode: every
scripted full-task run failed at `steps#0:step-execute` with `Step 4 out
of range (task has 4 steps)`, rebounding through recovery forever.

**Root cause:** `fn_task_update.step` has been **0-based since FN-6607**
(executor.ts FNXC:StepNumbering — the old `step - 1` conversion made
Step 0 impossible to mark). `mock-provider.ts` still sent `index + 1`,
so test mode marked steps 1..N instead of 0..N-1: Step 0 (Preflight)
never completed and step N threw out-of-range. Test mode's full-task
path has been broken since June.

**Also fixes the test that pinned the bug:** `mock-provider.test.ts`
expected `{ step: 1 }` for a fixture whose first unfinished step is
index 0 — the expectation encoded the 1-based off-by-one.

Verified: 12/12 mock-provider tests; the live E2E instance completes the
task after this patch (see follow-up screenshot in the session).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:54:58 -07:00
gsxdsm
7f3acf8929 test(core): pin both create-time duplicate guards' lane exclusions (#3234)
## What

Pins **both** create-time duplicate guards in
`branch-and-pr-entities.ts`. Test-only.

| site | method | excludes |
|---|---|---|
| `:445` | `findRecentTasksByContentFingerprint` | ARCHIVED (unless
`includeArchived`) |
| `:484` | `findRecentTasksBySourceParentTaskId` | COMPLETE and ARCHIVED
|

Blinding either back to its literals left the entire 16-file
lane-detector set green. **No test in `packages/core` reaches either
method.**

## Measured

```
converted:     Tests 8 passed (8)
blinded :445   Tests 1 failed | 7 passed (8)    <- only the fingerprint case
blinded :484   Tests 2 failed | 6 passed (8)    <- only the sibling cases
lint clean; fnxc-future-dates: none added; census unchanged
```

**Each blind fails exactly its own cases.** That matters: it proves the
two resolvers are pinned *independently*, rather than one broad test
appearing to cover both. Blinding `:445` leaves every sibling case green
and vice versa — so neither is riding on the other's coverage.

## They fail in opposite directions

This is why both belong in one file:

- **Fingerprint guard** — a renamed board leaves archived cards in the
candidate set, so filing a new task is **refused as a duplicate** of one
the operator already archived. The create is blocked and the thing
blocking it is invisible.
- **Sibling guard** — a renamed board leaves finished siblings in the
"recent live siblings" set, so completed work keeps counting as active.

One over-includes into a *refusal*, the other over-includes into
*phantom activity*. Neither raises an error.

## Positives pinned too

A LIVE fingerprint match is still a duplicate candidate;
`includeArchived: true` opts the renamed archived lane back in; a
WORKING sibling is still live. Excluding the finished lanes must not
degrade into excluding everything, or the guards stop guarding — the
failure mode a lane-widening change invites.

## A fixture detail that would have made this vacuous

Both queries cut off at `Date.now() - windowMs`, with `windowMs` capped
at 24h. The sibling harness I copied from seeds a **fixed past
timestamp**, which falls outside that window — every case would then
pass on an empty result, including the ones that are supposed to fail
under blinding. Fixtures are seeded at current time instead, and the
reason is recorded in the file so nobody "tidies" it back to a frozen
date.

## Progress

3 of the 5 uncovered core sites are now pinned (`store.ts:1135` in
#3233, these two here). Remaining and unclaimed:
`async-mission-store.ts:1179` and `task-id-integrity.ts:502`.
2026-07-31 12:52:36 -07:00
gsxdsm
5f6f39e115 fix(census): the scan root and the READ root could disagree, so an injected file list ENOENTs (#3230)
Picks up the bug **@#3228's author diagnosed and deliberately left
documented** rather than guessing at it mid-revision. Their diagnosis
was correct; the bug is mine, from extracting `triageFindings` in #3207
without considering an injected file list.

## The defect

`REPO_ROOT` came from the **script's** location; the file list comes
from `git ls-files` in the **CWD**. Identical in production and nowhere
else. Override the list — which a synthetic-tree fixture must do — and
every path is *listed* against the fixture but *read* against the repo:
`ENOENT` on every read.

## There were THREE read roots, not one

That is why a partial fix still ENOENTs, and I hit it myself: I fixed
`REPO_ROOT`, re-ran, and still got `ENOENT: open 'pkg/src/a.ts'`. The
scanners read the path **as given**:

| consumer | read root before |
|---|---|
| `triageFindings` | `join(REPO_ROOT, …)` |
| sync-resolver probe | `join(REPO_ROOT, …)` |
| `censusFiles` (AST) | path as given → CWD |
| `censusFilesText` | path as given → CWD |

All four now go through a single `readCensusFile`, so a listed path and
a read path cannot diverge again. **No lib change needed** — both
scanners already accept an injectable reader, which is the seam that
made this a small fix.

## What it unblocks

The two ratchet cases #3228 records as *permanently* vacuous at zero
backlog. On a three-file synthetic tree the full cycle is constructible
again:

```
inflated baseline  -> exit 0   TIGHTENED
deflated baseline  -> exit 1   ROSE
```

Neither is constructible against a real tree with nothing left to count
— which is exactly why that coverage was lost when the backlog hit zero,
and why `-1` (my #3218) and skip-at-zero (#3226) were both workarounds
for a missing seam rather than fixes.

## Measured

| | result |
|---|---|
| production scan | **unchanged** — 1961 files, 0 guards, `BACKLOG ZERO`
|
| `--strict` / `--json` / `--compare` | 0 / 0 / **0** (AST and text
classifiers still agree) |
| synthetic fixture | 3 files, **1 backlog / 1 deliberate** — the
numbers #3228 predicted |
| census suite | 53 passed |
| eslint / `check-fnxc-future-dates` / `pnpm test:gate` | clean / 0 / 0
(744 tests) |

`--compare` is the one I would look at first as a reviewer: it runs both
classifiers over the same list and fails if they disagree, so it catches
a reader change that silently alters what either one sees.

## Scope

Seams only — `FUSION_CENSUS_FILE_ROOT` and `FUSION_CENSUS_FILE_LIST`,
both required together (a root with no list still scans the real tree; a
list with no root still reads from it). Production sets neither.

I have **not** rewritten the two vacuous cases. That is #3228's work,
they already have two of six green, and duplicating it is how this fleet
loses PRs to collisions. This just removes the blocker.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved repository analysis reliability when run against configured
file sets or alternate repository locations.
* Prevented analysis from unintentionally reading unrelated files
outside the selected repository context.
  * Existing production behavior remains unchanged.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 12:49:54 -07:00
gsxdsm
3f06d7201b test(engine): assert the evaluator's exact archived-lane set, not toContain (#3232)
## What

Follow-up to the #3224 review comment *"reject legacy archived
identifiers for renamed workflows."* Test-only.

That comment had two halves. **The half it got wrong** is already
answered on main: asserting an exact single-column set `["vaulted"]`
*fails*, because `resolveProjectColumnsForRoles` unions
`LEGACY_COLUMN_IDS_BY_ROLE` in as a documented floor — so a row whose
workflow cannot be resolved still classifies. Pinning `["vaulted"]`
would encode the opposite of the design.

**The half it got right was never addressed.** `toContain` also passes
when the set grows a lane nobody intended, and an over-broad archived
set silently classifies *live* rows as archived. So the reviewer's worry
was legitimate even though the proposed fix was not.

The exact set is assertable — it just is not the one the review
proposed:

| board | resolved set |
|---|---|
| default | `["archived"]` |
| renamed | `["archived", "vaulted"]` — legacy floor + the board's own
lane |

`[...].sort()` in the helper makes ordering stable, so these pin the
resolver's whole answer rather than a substring of it.

## Proven to catch what `toContain` missed

Giving the fixture a second `archived`-trait column fails both new
assertions:

```
AssertionError: expected [ 'archived', 'cold-store' ] to deeply equal [ 'archived' ]
AssertionError: expected [ 'archived', 'cold-store', 'vaulted' ] to deeply equal [ 'archived', 'vaulted' ]
Tests  2 failed | 1 passed (3)
```

The previous `toContain` assertions pass unchanged against that same
spurious lane. That is the whole justification for this PR — without the
injection test it would be a stylistic preference.

```
clean:            Tests 3 passed (3)
spurious lane:    Tests 2 failed | 1 passed (3)
lint clean
```

## Note

I said in the review thread I would tighten this, so this closes that
loop. I also checked before editing whether another worker had already
done it — main already carries the FNXC tag and the legacy-floor
explanation from the same review round, so this PR adds only the part
still missing rather than redoing settled work.
2026-07-31 12:47:08 -07:00
gsxdsm
2780a8ae7b test(core): pin the open-undo query's finished-lane exclusion (#3233)
## What

Pins the **open-undo query's finished-lane exclusion** in
`packages/core/src/store.ts`. Test-only.

`findOpenRevertTaskForSource` answers *"is there an OPEN undo task for
this source?"* — the question behind the dashboard's Undo affordance. It
answers by **excluding the finished lanes**, so a prior undo that
already landed does not keep rendering as open.

Blinding that exclusion back to `ne(column,"archived"),
ne(column,"done")` left the entire 16-file lane-detector set green. **No
test in `packages/core` reaches this method at all.** The dashboard-side
twin (`taskRevert.ts`, #3129) is tested; the store-side query behind it
was not.

## Measured

| | default (control) | renamed complete | renamed archived | working
lane |
|---|---|---|---|---|
| converted | pass | pass | pass | pass |
| blinded to `["done","archived"]` | pass | **FAIL** | **FAIL** | pass |

```
converted: Test Files 1 passed (1) / Tests 4 passed (4)
blinded:   Test Files 1 failed (1) / Tests 2 failed | 2 passed (4)
lint clean; fnxc-future-dates: none added; census unchanged
```

Blind confirmed applied with `git diff --stat` before the run.

## What breaks without it

On a board whose complete lane is `shipped`, neither literal matches, so
a **done** undo task is never excluded and the query keeps returning it.
The card shows an undo already in flight *forever*, and the real
affordance is unreachable. Nothing errors — the button is just
permanently wrong, which is why it went unnoticed.

## Includes the paired positive

An undo still in a **working** lane IS reported as open. Excluding the
finished lanes must not degrade into excluding everything, or the
affordance breaks in the other direction and no undo is ever reported in
flight. Both new failing cases are renamed-lane cases; both survivors
are cases that should survive.

## Where this came from

Per-site blinding of all 14 remaining `resolveProjectColumnsForRoles`
call sites in `core`, run against a 16-file detector set. **9 covered, 5
uncovered:**

| site | verdict |
|---|---|
| `store.ts:1135` | **uncovered** → pinned here |
| `async-mission-store.ts:1179` (archived) | **uncovered** — its
neighbour `:1178` (complete) is covered |
| `branch-and-pr-entities.ts:445` | **uncovered** |
| `branch-and-pr-entities.ts:484` | **uncovered** |
| `task-id-integrity.ts:502` | **uncovered** |
| `reads.ts` ×3, analytics ×3, `eval-automation`, `task-artifacts-ops`,
`async-mission-store:1178` | covered |

The first run of that probe was **invalid and I nearly published it**:
it reported all 14 sites "COVERED" with *zero failing tests*. zsh does
not word-split unquoted parameter expansions, so `vitest run $DET`
passed 16 paths as one argument and vitest exited 1 with "No test files
found" — which my script read as a failing test. The re-run treats that
string as `INVALID` rather than a result. Third time this session a
wrong reading came from test *selection* rather than from blinding.

## Flagged, not guessed

The four remaining uncovered sites are named above rather than quietly
left; `async-mission-store` shows the same adjacent-pair split as
`team-analytics` in #3227, which is now the third confirmed instance of
that shape.
2026-07-31 12:46:56 -07:00
gsxdsm
dd09e57511 test(core): fix red main — assert the delete re-home against the resolver, not "triage" (#3229)
## What

**Fixes a red main.**
`workflow-reconciliation-production-shape.pg.test.ts` has been failing
with `expected 'todo' to be 'triage'`. Test-only.

Found while establishing a clean baseline for an unrelated coverage
audit — my tree was clean at `origin/main` (`76c73238a0`), so this is
not something I introduced. It is in the non-blocking suite, which is
why it has stayed red.

## It is not a regression — the test was the stale half

The delete path was deliberately fixed to re-home occupants using
`resolveEntryColumnId(resolveDefaultWorkflowIr())` instead of
`BUILTIN_CODING_WORKFLOW_IR`. This assertion was not updated with it.

The two IRs are **not the same board**:

| IR | entry column |
|---|---|
| `BUILTIN_CODING_WORKFLOW_IR` (`builtin:legacy-coding`) | `triage` |
| `resolveDefaultWorkflowIr()` (the catalog default) | `todo` |

Re-homing into `triage` put cards in a column the default board never
declares. It slipped past `moveTask`'s undeclared-target guard **only
because `triage` is a legacy id** and the recovery-rehome path exempts
those — so the guard that exists to stop exactly this could not see it.

So `todo` is the correct behaviour and the literal `"triage"` was what
needed fixing.

## Why it asserts a resolver rather than `"todo"`

Swapping one hardcoded id for another would be the identical trap one
rename later — the same class of defect this whole program exists to
remove. The expectation now derives from **the same two functions the
product path calls**, so it cannot drift out of sync with them again.

I also added the complement: the card must genuinely have **left** the
vanished column, not merely match whatever a resolver returns. Without
it, a resolver that started returning `custom-hold` would pass.

## Proven not appeasement

Reverting the product line to the legacy IR — the original defect —
fails this test:

```
AssertionError: expected 'triage' to be 'todo'
Test Files  1 failed (1) / Tests  1 failed | 6 passed (7)
```

That is the check that matters for a test edit that turns a red green.
It fails on the defect it describes.

## Measured

```
before: Tests 1 failed | 6 passed (7)
after:  Tests 7 passed (7)
16-file detector set: 181 passed (16 files)   [was 1 failed | 180 passed]
lint clean
```

## Note on the reading

I got this wrong twice before getting it right, and the record is worth
having. My first read was "the test is stale, `triage` was merged away."
My second was "the builtin IR still declares `triage`, so the
*behaviour* is the defect" — which the IR file superficially supports.
Only the third reading, of the FNXC note at the fix site, showed the
file I was reading is the **legacy** IR and not the default one. Two of
those three readings would have produced a confidently wrong PR; the
deciding evidence was the comment the fixing author left at the call
site, which is a good argument for writing them.
2026-07-31 12:31:04 -07:00
gsxdsm
01ab2400d0 docs(learnings): blinding measures the instrument you picked — rule 5, and where the measurement cannot be taken (#3222)
Extends `blind-the-resolver-to-find-uncovered-conversions.md` rather
than forking a second doc on the same technique.

## Rule 5: blinding measures the instrument you picked, not the site

A suite that never reaches the blinded site reports `0 failed` for the
same reason a covered one does. The outputs are identical. This produced
a **wrong answer twice in one sweep**, both times reading as a finding:

| blinded | suite run | said | actually |
|---|---|---|---|
| `reads.ts` ×3 | `search-excludes-renamed-archive-lane.test.ts` | 3
uncovered | that file unit-tests `liveSearchPredicate` and never runs
`reads.ts`; against `cold-storage-renamed-archive-lane.test.ts` one of
the three is covered |
| `server.ts` ×3 | `reliability-metrics.test.ts` | 3 uncovered | that
file imports `../reliability-metrics`; nothing executes the route at all
|

The `reads.ts` case is the one to remember, because **the misleading
suite was written for that exact conversion**. It proves the
collaborator honours a resolved set — which says nothing about whether
the caller passes one, and can never fail when the call site is blinded.
That gap shipped as a real hole and was closed in #3220.

Doc adds the cheap guard: make the blinded edit obviously fatal (`throw
new Error("x")`) and re-run. Still green means the suite does not reach
the site and the measurement is void.

## Where the measurement cannot be taken

Per #3212's stance that recording *why* something cannot be pinned is a
result, three groups are written down so nobody re-derives them:

- **No TCP PostgreSQL** — `workflow-analytics.ts` / `team-analytics.ts`
(4 resolvers) keep renamed-lane coverage in `.pg` suites. `pgDescribe`
probes **TCP**; `pg_isready` succeeding on a **Unix socket** is not the
same thing. I made exactly this mistake and reported PG as reachable one
round before correcting it — mistaking the two turns 4 skipped suites
into 4 false "uncovered" readings.
- **No injectable seam** — `reads.ts`'s incremental-sync scan composes
Drizzle conditions against `layer.db`. A test there asserts the query
built, not the rows excluded: green, and blind to the bug.
- **Logic inside a route closure** — `server.ts`'s three resolvers sit
in the `/api/health/reliability` handler, which has no route-level test.
The only harness in that package is a mock-the-world shell the slow-test
rule forbids; the alternative is a refactor to expose a seam, which is
its own commit.

## Census

**Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0`.** Documentation
only.

Gates verified green (`check-fnxc-future-dates`,
`lifecycle-column-census --strict`).

No changeset: internal docs, per AGENTS.md.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added guidance for verifying the test instrument used during blinding.
  * Documented fatal-edit reachability checks.
* Added troubleshooting guidance for situations where resolver coverage
cannot be measured.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 12:28:20 -07:00
gsxdsm
d4c25384ae test(engine): record what the zero-backlog early return stops testing (#3228)
## A green case that stopped testing what its name says

#3226 fixed a red `main` correctly: `fileWithGuards()` now returns
`null` at zero backlog, and with nothing to inflate there is no rise to
manufacture. Asserting `totals.column === 0` and returning is the honest
response.

What went unrecorded is the cost. At zero, these two cases:

- *"exits 0 and REWRITES the baseline under `--update-baseline`, even
when the count rose"*
- *"exits 1 and LEAVES the baseline alone on a rise without
`--update-baseline`"*

no longer exercise the CLI's ordering or exit codes. They assert the
backlog is empty and return. **If the write-before-exit ordering
regressed — the exact bug those cases were written for — both would
still pass.**

That matters more than it would elsewhere, because **zero is not a state
to wait out.** It is this program's terminal state: the backlog went 126
→ 0 and is meant to stay there. So the vacuity is permanent, not
transitional.

This file already legislates against precisely this, two hundred lines
down:

> `/* Anti-vacuity: an empty exclusion list would make the assertion
below trivially true. */`

## What this PR does

Adds a comment on `fileWithGuards()` recording (a) which cases go
vacuous at zero and why, (b) that zero is terminal so it will not
resolve itself, and (c) the durable fix.

**Comment only. No behaviour change — suite stays 53/53.**

## The durable fix, recorded rather than done

Point the scan at a synthetic tree so the fixture stops being a function
of the real backlog — the same seam `FUSION_CENSUS_BASELINE_PATH`
already provides for the baseline, applied to the file list.

It needs one CLI correction to work, and that is a genuine bug in my own
code regardless of this suite: `triageFindings` and the sync-resolver
check read files via `join(REPO_ROOT, f.file)`, where `REPO_ROOT` is
derived from the **script's** location. An overridden file list
therefore changes which paths are *listed* without changing where they
are *read from*, and every read misses with `ENOENT`.

I verified that approach works (a three-file fixture yields a stable `1
backlog / 1 deliberate / 1 sync-resolved`) and got two of the six
failing cases green with it, then stopped rather than keep guessing in a
file being actively revised. Left as a comment so whoever takes it does
not re-derive the diagnosis.

## Why this is worth a PR at all

This program's recurring failure is instruments that report green while
measuring nothing — an inert conversion the census scored as a win, a
ratchet wired to nothing, a gate that could not fail. A test asserting
`0 === 0` under a name promising ordering coverage is the same shape at
the test layer. The suite cannot be fixed in this PR without re-opening
work someone else owns, but it can at least stop being silent about it.

## Census before / after

```
before:  COLUMN guards (the backlog):   0
after:   COLUMN guards (the backlog):   0
```

## Verification

`test:gate` exit 0 · `lifecycle-column-census.test.ts` **53 passed** ·
`fnxc-future-dates`, `lifecycle-columns`, `inert-sync-lanes`,
`quarantine-ledger`, `inert-flag-seams`, `lane-wiring`,
`sql-column-literals` — all exit 0.
2026-07-31 12:28:09 -07:00
gsxdsm
6646c1b95d docs: regenerate synced skill tool tables (unblocks all four full-suite shards)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:27:11 -07:00
gsxdsm
32edd1421a test(core): pin team analytics' in-flight lane read — the other half of the pair (#3227)
## What

Pins `aggregateTeamAnalytics`' **in-flight lane read** (`activeLanes`) —
the unpinned half of an adjacent resolver pair. Test-only, no product
change.

`completeLanes` and `activeLanes` are declared **two lines apart**.
Every existing case in this file asserts only `totals.tasksCompleted`,
so the in-flight query `activeLanes` feeds was never observed.

Measured on main:

| blinded resolver | result |
|---|---|
| `completeLanes` → `["done"]` | **FAILS** the file (1 failed / 3
passed) — pinned |
| `activeLanes` → `["in-progress","in-review"]` | **entirely GREEN** (4
passed) — unpinned |

One resolver held, its neighbour not, in a file named
`team-analytics-renamed-lanes`. This is the half-covered-pair shape the
program keeps finding; being *next to* a covered resolver is not
coverage.

## How I found it

Rather than blind core's 17 files one at a time, I made
`resolveProjectColumnsForRoles` itself return legacy-only — its own
documented degrade path — which blinds **all 116 call sites in one
edit**. The full core suite then reported **24 failures across 16
files** out of 4,987 tests, which maps the covered areas in a single
run: the analytics renamed-lane pg tests, the archived-lane family,
eval-automation, mission-store, and the resolver's own tests.

That global probe finds *areas* that are covered, not *resolvers* — so
the pairs still needed individual blinding, which is what surfaced this
one. `workflow-analytics.ts` has the identical two-resolver shape and
**both halves are covered**; the gap is specific to `team-analytics.ts`.

## Two things have to be right, and both are now asserted

1. **The SQL must ASK for the board's real wip lane** — `activeLanes`.
2. **`buildTeamAnalytics` must RECOGNISE the row it gets back.** It
classifies via `isWipColumnRole(query.columnFlagsByName?.get(name),
name)`, which **without flags falls back to `name === "in-progress"`**
and drops a renamed lane it already fetched.

So supplying `columnFlagsByName` is part of the caller contract, not
test scaffolding: **widening the query alone would still report zero.**
A test that only widened the first half would pass while the feature
stayed broken.

## Measured

```
converted:            Test Files 1 passed (1) / Tests 7 passed (7)
blinded activeLanes:  Test Files 1 failed (1) / Tests 2 failed | 5 passed (7)
lint clean; fnxc-future-dates: none added; census unchanged
```

Blind confirmed applied with `git diff --stat` before each run.

## What breaks without it

A per-agent `tasksInProgress: 0` sitting beside a nonzero completed
count and real token spend — an agent that looks idle while it is
working. Same wrong-but-plausible shape this file's own header
describes: nothing errors, and a plausible-looking number is the least
likely defect for anyone to file.

## Flagged, not guessed

- `packages/core` is not my package; this is additive tests only. I
raised the same note on #3225.
- The global probe shows core has **substantial** renamed-lane coverage
— it is not the uniformly-unpinned surface I implied when I first
reported 17 unaudited files. Correcting that here rather than leaving
the stronger claim standing.
- Still unblinded individually: the resolver pairs in
`async-mission-store.ts` (1178/1179) and the archived reads in
`task-store/reads.ts` (396/615/793). Their *files* fail under the global
blind, so something covers each area — but that is not per-resolver
evidence, and I am not claiming it is.
2026-07-31 12:25:26 -07:00
gsxdsm
12c4ab5a6e test(engine): pin the evaluator's archived-lane read — the service had no test at all (#3224)
## What

Pins the **evaluator's archived-lane read**. Test-only — no product
change.

`HybridEvaluatorService.evaluateTask` resolves the board's archived
lanes and hands them to `collectDeterministicSignals`, which decides
which of a task's related rows count as archived when scoring a run.

**The service had no test anywhere in the repo.** Four test files import
the module; none construct or exercise it. So this conversion was
unobservable for the simplest possible reason — *nothing ran the code*.
That is a different failure from the ones this audit has been finding
(harnesses that run the code but cannot see the difference), and worth
distinguishing: no amount of fixture care helps when the entry point is
never called.

## Measured

| | default (control) | renamed | differential |
|---|---|---|---|
| converted | pass | pass | pass |
| blinded to `["archived"]` | pass | **FAIL** | **FAIL** |

```
converted: Test Files 1 passed (1) / Tests 3 passed (3)
blinded:   Test Files 1 failed (1) / Tests 2 failed | 1 passed (3)
the 4 files importing evaluator.ts, plus this one: 5 files/76 tests, all green
lint clean; fnxc-future-dates: none added; census unchanged
```

Per the rule I documented in #3223, the blind was confirmed applied with
`git diff --stat` **before** the run rather than trusting the tool's
exit code.

## What breaks without it

On a board whose archived lane is `vaulted`, the evaluator hands the
collector the legacy `{archived}` set. Rows resting in `vaulted` are not
recognised as archived, and the deterministic half of every evaluation
score is computed from a wrong picture of the task's history. **Nothing
errors, the run completes, the number is just wrong** — which is why it
survived unnoticed.

## Pinned without faking a provider response

The assertion is about what the collector *receives*, which is decided
before any model call. `collectDeterministicSignals` is mocked to record
its arguments and throw a sentinel; the test asserts the resolved lane
set and stops.

This is deliberate over the obvious alternative of feeding `runPrompt` a
canned AI payload: `deps.runPrompt` is injectable so either approach is
offline, but a canned payload has to satisfy `parseAiResponse` and every
`EVAL_SCORE_CATEGORIES` entry, and would silently rot into a maintenance
burden on a test whose subject is one `Set`. Reversible if someone later
wants full end-to-end evaluator coverage — that is a different test, not
this one.

## Completes the engine audit

With this, every `resolveProjectColumnsForRoles` call site in
`packages/engine` has been blinded:

| file | resolvers | result |
|---|---|---|
| `self-healing.ts` | 64 | 21 pinned, 1 recorded inert by construction,
remainder mapped |
| `executor.ts` | 2 | both already covered |
| `scheduler.ts` | 1 | uncovered → pinned (#3219, merged) |
| `triage.ts` | 1 | uncovered → pinned (#3221) |
| `restart-recovery-coordinator.ts` | 1 | already covered |
| `notification-service.ts` | 1 | already covered |
| `evaluator.ts` | 1 | uncovered → pinned (this PR) |

`project-engine.ts:5154` takes `roles` as a **parameter**, so it is a
generic wrapper with no fixed role set to blind — flagged rather than
guessed at; its callers are where the question belongs.

**`packages/core`'s 17 files remain entirely unaudited** and I am
claiming nothing about them.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added regression coverage to verify reliable resolution of archived
workflow lanes.
* Covered both the default archived-lane name and custom renamed
configurations.
  * Confirmed compatibility with legacy archived-lane naming behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 12:20:08 -07:00
gsxdsm
76c73238a0 test(core): pin the engine-downtime shift's wip read (428 tests could not see it) (#3225)
## What

Pins the **engine-downtime timing shift's wip read** in
`packages/core/src/store.ts`. Test-only — no product change. First
audited site in `core`.

`reconcileActiveTimingForEngineDowntime` (FN-7011/FN-7975) excludes
proven stopped-engine wall-clock from a card's active time. It finds the
cards to fix by querying the board's wip lane.

**Blinding that read back to `["in-progress"]` left every test that
touches the sweep green — 4 in this file plus 424 in the two engine
files that exercise it, 428 in total.**

## Why 428 tests were blind to it

The existing store double is 10 lines and contains **both** documented
anti-patterns, either one sufficient on its own:

1. **`listTasks: vi.fn(async () => tasks)` ignores its `column`
argument** — it returns the same rows whichever lane is requested. A
fake that ignores its own filter cannot see a filter bug, which is
exactly the bug this resolver exists to fix.
2. **No `listWorkflowDefinitions`** — `resolveProjectColumnsForRoles`
then returns the legacy ids and nothing else (an intentional degrade in
`project-lane-vocabulary.ts` so an unreadable workflow list cannot fail
a sweep). The resolved set and the literal set were *equal by
construction*.

The new double fixes both and changes nothing else. **The existing cases
keep the original double on purpose:** they are about heartbeat and
threshold arithmetic, not lanes, and rewriting them would put unrelated
churn in the same commit.

## Measured

| | default (control) | renamed | differential | non-wip card |
|---|---|---|---|---|
| converted | pass | pass | pass | pass |
| blinded to `["in-progress"]` | pass | **FAIL** | **FAIL** | pass |

```
converted: Test Files 1 passed (1) / Tests 8 passed (8)
blinded:   Test Files 1 failed (1) / Tests 2 failed | 6 passed (8)
engine neighbours (project-engine-unpause-active-timing + self-healing): 424 tests, green and unchanged
lint clean; fnxc-future-dates: none added; census unchanged
```

Blind confirmed applied with `git diff --stat` before each run, not
inferred from the tool's exit code.

## What breaks without it

On a board whose wip lane is `building`, the sweep queries
`in-progress`, finds **no tasks**, and shifts no anchor. Every card
silently absorbs the stopped-engine wall-clock the sweep exists to
exclude. The reported active time is simply wrong and nothing fails to
signal it — the same silent-wrong-number shape as the evaluator defect
in #3224.

## Also covers the complement

A held card *outside* the wip lane is **not** shifted. Widening a lane
read is the kind of change that can quietly turn a targeted sweep into a
board-wide rewrite; a card in `todo` has no stopped-engine time to
exclude, and there is now a case saying so.

## Scope note

`packages/core` is not my package. This is an additive test file with no
product change, so collision risk is low, but I am flagging it rather
than assuming: **16 of core's 17 files with resolver call sites remain
unaudited** and I claim nothing about them. The audit method and its
failure modes are documented in #3223 if core's owner wants to continue
it.
2026-07-31 12:14:40 -07:00
Phil Larson
3c7dd9a803 test(engine): keep census regressions valid at zero backlog (#3226)
## Summary
- keeps lifecycle-census end-to-end fixtures valid after the conversion
backlog reaches zero
- treats zero backlog as a real protected end state instead of requiring
a remaining guard or claim target
- preserves nonzero rise/claim assertions when guards remain

## Test plan
- `corepack pnpm --filter @fusion/engine exec vitest run
src/__tests__/lifecycle-column-census.test.ts --silent=passed-only
--reporter=dot --project=engine-default`
- `corepack pnpm --filter @fusion/engine typecheck`
- verified the same suite at zero backlog on the aggregate runtime
(53/53)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Improved lifecycle validation to handle empty backlogs without errors.
  * Added coverage for zero-item results and changing file sets.
* Enhanced baseline and trend checks to report completed backlogs
consistently.
* Improved resilience when expected result entries are unavailable,
ensuring validation completes cleanly with accurate zero counts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 12:14:28 -07:00
gsxdsm
206ff11874 docs(solutions): record the blinding audit's own failure modes (#3223)
## What

Extends
`docs/solutions/workflow-learnings/blind-the-resolver-to-find-uncovered-conversions.md`
with what this session's audit work paid for. Docs only — no code, no
changeset (internal doc).

## The main addition: the audit's own failure modes

**Every wrong reading this method has produced came from test
*selection*, not from the blind.** Three in one session, each of which
reads exactly like coverage:

| what I ran | why it lied |
|---|---|
| `vitest run src/__tests__ -t "executor"` | `-t` filters test
**names**, not files. Reported two `executor.ts` resolvers uncovered;
**both are covered.** |
| `blind3.py <file> <var>` with an unmapped role | exited non-zero
**silently**; `&&` skipped the check and `;` let vitest run against
**unmodified source**. Reported "375/375 green under blinding" with
nothing blinded. |
| `vitest run src/__tests__/notification` | missed
`src/notification/__tests__/` — a nested `__tests__` the glob never
reached. Reported covered code as uncovered. |

The rule that follows: an UNCOVERED verdict is a claim about the whole
tree and needs the whole tree's tests. Confirm the blind actually
modified the file with `git diff --stat` — *not* the tool's exit code —
and that the run included every file importing the module.

I am documenting my own instrument failing the standard I have been
applying to product guards all phase: *a guard that reports success
without checking anything is worse than no guard.* Mine reported success
without checking anything. It now echoes what it substituted and where,
and fails loudly on an unmapped role or missing variable; I self-tested
both directions before trusting any number in #3219 and #3221.

## Rule 5: the resolver must be able to answer differently in the
harness

`resolveProjectColumnsForRoles` returns **legacy ids and nothing else**
when the store has no `listWorkflowDefinitions` — an intentional degrade
so an unreadable workflow list cannot fail a sweep. A harness omitting
it makes the resolved set and the literal set **equal by construction**,
so the conversion is unobservable however good the assertion is.

This is not a test bug. It is correct production behaviour that erases
the difference the test is trying to measure — and it alone left both
the `scheduler.ts` and `triage.ts` conversions unpinnable.

## A correction to my own earlier rule

I had "seed-then-union sites hide defects" too broad. Such a site hides
a defect **only while every lane you assert on is already in the seed**.
On a renamed board the resolver is the sole contributor of the renamed
lane, so the legacy blind is *not* a no-op — I predicted it would be and
it failed. Also: expand roles to legacy ids **per role** from
`LEGACY_COLUMN_IDS_BY_ROLE`; `intake` is `["todo","triage"]`, not
`["triage"]`, and a stricter-than-real blind manufactures failures that
read as coverage.

## Inventory, so the gap is legible

**116 non-test call sites across 30 files** — core 17, engine 10,
dashboard 2, cli 1. Audited so far, all in engine: `self-healing.ts` (64
mapped / 21 pinned / 1 inert by construction), `executor.ts` (2,
covered), `scheduler.ts` (uncovered → pinned in #3219), `triage.ts`
(uncovered → pinned in #3221), `restart-recovery-coordinator.ts`
(covered), `notification-service.ts` (covered).

**`packages/core`'s 17 files are entirely unaudited.** Stated as a gap
rather than left implied, so nobody reads engine's coverage as a
repo-wide clean bill.

## Flagged, not guessed

- `evaluator.ts`'s archived read is uncovered — **no test file imports
that module at all.** Left unpinned deliberately: it is a thin
pass-through into `collectDeterministicSignals`, which is testable
directly, and it affects eval signal quality rather than task lifecycle.
Recorded in the doc rather than silently skipped.
- I did not audit core; it is outside my package and I am not claiming
anything about it either way.
2026-07-31 12:08:59 -07:00
gsxdsm
719cf281fd test(engine): pin the startup stale-planning sweep's planner-lane read (#3221)
## What

Pins the **startup stale-planning sweep's** planner-lane read in
`triage.ts`. Test-only — no product change.

A card can hold `status: "planning"` while triage specifies it in place.
A crash or restart before planning completes leaves that status set, and
a startup sweep clears it. If the sweep misses the card, it occupies a
planning admission slot **permanently** and new triage work is never
admitted.

The lane read was converted from `resolvePlannerLanes(store, "")` —
called with an **empty task id**, so it could never resolve a task and
always answered with the default board — to a project-level
`resolveProjectColumnsForRoles(store, ["intake", "hold"])`.

**No test could observe that conversion.** All 25 existing triage test
files stayed green (375/375) with the resolver neutered — under *both*
blinds tried.

## Measured

| | default (control) | legacy ids | renamed | differential |
|---|---|---|---|---|
| converted | pass | pass | pass | pass |
| blinded to empty set | pass | pass | **FAIL** | **FAIL** |
| blinded to legacy pair | pass | pass | **FAIL** | **FAIL** |

```
converted:            Tests 4 passed (4)
blinded to empty set: Tests 2 failed | 2 passed (4)
blinded to legacy:    Tests 2 failed | 2 passed (4)
existing 25 files:    green under BOTH blinds  (only this new file fails)
triage suite: 25 files/375 tests -> 26 files/379 tests, all green
lint clean; fnxc-future-dates: none added; census unchanged
```

## A prediction I got wrong, and what it changes

The site is **seed-then-union**:

```ts
const sweepColumns = [...new Set(["triage", "todo", ...projectPlannerColumns])];
```

I expected blinding the resolver to its legacy pair `["triage","todo"]`
to be a **no-op by construction** — the seed already contains both. It
is not. That reasoning holds only for a *default* board; on a renamed
board the resolver is the sole contributor of `drafting`, so the legacy
blind drops it and the renamed cases fail.

The corrected rule, which is narrower than the one I was carrying: **a
seed-then-union site hides a defect only while every lane you assert on
is already in the seed.** Assert on a lane that is not, and the union
stops protecting it. The empty-set blind remains the stricter of the two
because it also models a resolver returning nothing at all. Both are
recorded in the test file so the next person does not re-derive it.

## A tooling failure worth naming

My first triage measurement reported **375/375 green under blinding** —
and was a lie. The blinding script had no mapping for
`["intake","hold"]` and exited `2` **silently**; the `&&`
short-circuited the check while a `;` let vitest run against
**unmodified source**. A no-op blind produces a green run that is
indistinguishable from real coverage.

The script now prints what it substituted and where, fails loudly on an
unmapped role or a missing variable, and I self-tested both directions
(bogus var → rc 2 with a message; real var → rc 0 with the substitution
echoed) before trusting any number above. This is the same standard I
have been applying to other people's guards — *a guard that reports
success without checking anything is worse than no guard* — and my own
instrument failed it.

## Also pinned

The **legacy half** of the union. `triage`/`todo` stay in the sweep even
on a board whose workflow declares neither, because pre-U11 and Coding
(Ideas) rows can rest there. Dropping them in favour of the resolved
lanes alone would strand exactly those rows, so there is now a case
asserting it.

## Flagged, not guessed

- The FNXC stamp at `triage.ts:987` reads `2026-07-31-23:59`, hours
ahead of the `date -u` clock. It is one of the 181 grandfathered stamps
so the gate is green; I left it rather than widen this PR.
2026-07-31 12:01:00 -07:00
gsxdsm
9ab9822b8d test(engine): pin the deleted-blocker WIP-lane read, which no test could see (#3219)
## What

Pins the deleted-blocker sweep's **WIP-lane read** in `scheduler.ts`.
Test-only — no product change.

When a task is soft-deleted, a `task:deleted` listener clears
`blockedBy` on every dependent so the work can be scheduled again. It
reads two lanes to find them: hold and WIP. The WIP read was already
converted to `resolveProjectColumnsForRoles(store,
["countsTowardWip"])`.

**Blinding that resolver back to the literal `["in-progress"]` left all
14 existing scheduler test files green — 145/145.** The conversion was
load-bearing and unpinned.

## Why nothing could see it

Two independent harness properties, **either one sufficient** to hide
the defect:

1. **`resolveProjectColumnsForRoles` returns legacy ids and nothing else
when the store has no `listWorkflowDefinitions`** — an intentional
degrade in `project-lane-vocabulary.ts` so an unreadable workflow list
cannot fail a sweep. The shared scheduler harness does not define it, so
*the resolved set and the literal set were equal by construction* in
every existing test.
2. **`listTasks` was mocked as `vi.fn(async () => tasks)`**, ignoring
its `column` filter. A mock that returns every task regardless of the
lane asked for cannot detect a wrong lane.

This is worth naming because #1 is not a test bug — it is correct
production behaviour that happens to erase the difference a test is
trying to measure. A harness can satisfy a conversion's *shape* while
making its *effect* unobservable.

There is a test file named
`scheduler-renamed-dependency-and-review-lanes.test.ts` covering
dependency satisfaction, file-scope leases, base-branch stacking, PR
hydration and mission completion on a renamed board. It does not reach
this sweep, and its harness has both properties above.

## What breaks without the conversion

On a board whose WIP lane is `building`, the literal read asks for
`in-progress`, finds nothing, and the in-flight dependent is never
reconciled. It keeps `blockedBy` pointing at a task that no longer
exists — **permanently**, because the blocker can never be completed or
re-deleted to trigger another sweep. Work stops with nothing to rescue
it.

## Measured

| | default (control) | renamed | differential |
|---|---|---|---|
| converted | pass | pass | pass |
| blinded to `["in-progress"]` | pass | **FAIL** | **FAIL** |

```
converted: Test Files 1 passed (1) / Tests 3 passed (3)
blinded:   Test Files 1 failed (1) / Tests 2 failed | 1 passed (3)
scheduler suite: 14 files/145 tests -> 15 files/148 tests, all green
lint: clean   fnxc-future-dates: none added
```

The default-vocabulary control passes in **both** columns by design: it
is what keeps a generic break in this path from hiding in the renamed
assertion.

## Census

Unchanged — **2 / 148 deliberate**. `scheduler.ts` was already at 0.
This PR adds coverage, not conversions; the census counts comparisons
and would not have moved either way, which is the same blind spot that
let #3078 merge green.

## Flagged, not guessed

- The `resolveTaskParkedColumns` **hold** read one line above
(`scheduler.ts:1402`) is the other half of this pair. My scenario places
its dependent in the WIP lane, so it does not exercise the hold read and
I am not claiming it is pinned.
- The FNXC stamp at `scheduler.ts:1404` reads `2026-08-01-05:00`, which
is future-dated. It is one of the 181 grandfathered stamps, so the gate
is green and I left it alone rather than widen this PR.
2026-07-31 11:55:43 -07:00
gsxdsm
3146a745bf test(engine): cover two reporter resolvers that no test could tell from the literal (#3217)
## What

Applied #3214's blinding procedure **outside `self-healing.ts`**, where
that measurement has never been run. Two of the five resolvers across
the two reporters were uncovered; this covers both.

## The measurement

One resolver at a time, blinded back to its legacy ids, against each
file's existing suite:

| site | blinded to | result |
|---|---|---|
| `backlog-pressure-reporter.ts:87` hold | `["todo"]` | 2 failed —
covered |
| **`backlog-pressure-reporter.ts:88` wip** | `["in-progress"]` | **0
failed of 11 — UNCOVERED** |
| `backlog-pressure-reporter.ts:89` terminal | `["done","archived"]` | 1
failed — covered |
| `stale-task-reporter.ts:59` wip | `["in-progress"]` | 1 failed —
covered |
| **`stale-task-reporter.ts:60` review** | `["in-review"]` | **0 failed
of 7 — UNCOVERED** |

Both uncovered resolvers sit in a `Promise.all` **beside one that is
covered**, so each sweep reads as converted while half of it was held by
nothing. That is rule 1 in the doc — coverage is per-resolver, not
per-sweep — and it is why the census cannot answer this: a syntactic
scan sees five resolved sites and five is what it counts.

`stale-task-reporter.ts` is the sharper case. Its describe block
**already declared `signoff` in the fixture IR** and no case ever put a
card there, so the review resolver was decorative.

## What they cost on a renamed board

- **wip** feeds `inProgressCount`, the *denominator* of `ratio =
todoCount / max(inProgressCount, 1)`. Against the literal, busy work in
a renamed lane counts as **zero**, the ratio inflates, and the
backlog-pressure alert fires on a queue that is draining normally — the
operator is paged that the board is jammed while agents work through it.
- **review** decides which rows the staleness read *fetches at all*. A
review stalled for days in a renamed lane is never queried and never
surfaced — precisely the condition this reporter exists to report.

## Following the four rules

**Rule 2 — the fixture reaches the guarded branch.** 12 hold cards over
2 wip cards is a ratio of 6, *under* the default threshold of 10, so the
correct answer is "no alert"; blinding collapses the denominator to 1,
the ratio becomes 12, and it alerts. A fixture whose ratio cleared the
threshold either way would exercise the sweep and never touch the line
under test.

**Rule 3 — assert the path-specific side effect.** `upsertInsight` not
called, and `logEntry` called with `column=signoff`. Asserting `alerted
=== false` alone would also pass if the run bailed for an unrelated
reason — missing insight store, cooldown, too few candidates — none of
which involve the wip lane.

**Rule 4 — the store fake honours `options.column`.** Both harnesses
already did; reused rather than replaced.

Each new case is paired with a negative so it cannot pass vacuously: the
"does not alert" case is backed by a *same renamed board still alerts
when in-progress work really is thin* case, so a reporter broken into
never firing fails.

## Census

**Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0` before and
after.** This converts nothing. It closes coverage on conversions the
census already counts as done, which is the gap #3214 names: *"the
census counts comparisons; it cannot tell a working conversion from one
a later merge silently reverted."*

## Verification

Blind-verified in both directions — blinding each resolver fails
**exactly** the new case and nothing else:

```
backlog-pressure  BLIND wip     -> 1 failed | 12 passed (13)   restored: 13 passed
stale-task        BLIND review  -> 1 failed |  7 passed  (8)   restored:  8 passed
combined                                                        21 passed (2 files)
```

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:50:25 -07:00
gsxdsm
cfcbba6f81 fix(census): 4 RED ratchet tests on main, and the report said nothing at zero (#3218)
Two problems, both caused by the backlog actually shrinking.

## 1. Four failing tests on main

**Pre-existing, not introduced here** — running this file on clean
`origin/main` gives `49 passed / 4 failed` with identical messages. I
checked that before touching anything, because the failures surfaced
while I was editing the same file.

The ratchet cases build their fixture like this:

```ts
Object.entries(baseline.byFile).find(([, c]) => c > 1)   // needs a file with MORE THAN ONE guard
```

After the tail reclassification no such entry exists. `find` returns
undefined → `byFile[undefined] = NaN` → the baseline is corrupt → every
case fails with `expected … to contain 'TIGHTENED'`, a message that
points squarely at the CLI when the **fixture** is at fault. That
misdirection is why this sat red.

The ratchet doesn't care *which* file it tightens, only that an
allowance exceeds the measured count. So `inflate` now takes any entry,
and synthesises one against a real scanned file when the backlog is
empty.

`deflate` is the harder half: a RISE needs an allowance **below** the
real count, and once every measured count is 0 the only value below is
negative. The empty case uses `-1`. That is not a realistic baseline
value and the comment says so — it is the sole way to exercise the
`measured > allowed` comparison against a tree with nothing left to
count, which is the tree this suite now runs on.

Same class as the unbounded-slice rot in #3207: **census self-tests
coupled to the size of a shrinking backlog.** That is now twice, so it
is a pattern rather than an accident.

## 2. The report went silent at the finish line

The verdict was two inline branches and neither fired at zero —
`CONVERSION QUEUE EMPTY` required `totals.column > 0`. So the one state
the entire fleet phase was working toward printed **nothing**, which
reads as a broken scan rather than the protected end state.

Extracted to a pure `describeBacklogState({ columnGuards,
unexaminedGuards })` returning lines, so the caller stays a dumb
printer:

```
BACKLOG ZERO: no lifecycle-column guard remains.
This is the protected end state, not an empty scan — `--strict` fails on any RISE, so a new
guard cannot land silently. Use the role helpers (resolveLifecycleColumns / columnHasRole).
```

Pure **specifically** so the zero state is testable before the tree
reaches zero. While it was inline, only the *current* backlog state was
observable — and a message nobody can test before they need it is the
one that is wrong when they do.

## Evidence

| check | result |
|---|---|
| census test file | **53 passed** (was 49 passed / 4 failed) |
| behaviour on today's tree | **unchanged** — identical `CONVERSION
QUEUE EMPTY` block |
| empty-baseline probe | exits 1, `column-guard count ROSE` |
| forced zero verdict | prints `BACKLOG ZERO … not an empty scan` |
| `--strict` / `check-fnxc-future-dates` / eslint | 0 / 0 / clean |
| `pnpm test:gate` | exit 0 (744 tests) |

Four new tests pin all three states, including that the unexamined
branch must **not** claim the queue is empty while real work is
outstanding.

## Census

No guard converted — this is tooling and test repair. Backlog unchanged
at 1, which #3215 takes to 0.
2026-07-31 11:50:14 -07:00
gsxdsm
78d87f0a10 test(core): pin the search archive-lane WIRING — the predicate was covered, the hand-off was not (#3220)
## The false-green

#3160 (mine) proved `liveSearchPredicate` honours a resolved archive
set: hand it `Set(["archived","filed"])` and `filed` appears in the
bound params. That contract is real and still correct.

**Nothing proved `reads.ts` passes one.** It is a unit test of the
collaborator, so blinding the resolver at the call site cannot fail it.
A conversion, a test that looks like it covers it, and no connection
between them.

## The measurement — and the instrument matters

| site | vs. the predicate unit test | vs. a test that drives `reads.ts`
|
|---|---|---|
| `reads.ts:396` cold-storage list | 0 failed | **1 failed — covered** |
| `reads.ts:615` incremental sync | 0 failed | 0 failed — **UNCOVERED**
|
| `reads.ts:793` search | 0 failed | 0 failed — **UNCOVERED** |

Against `search-excludes-renamed-archive-lane.test.ts` all three read as
uncovered — an artefact of asking a file that never executes `reads.ts`.
Against `cold-storage-renamed-archive-lane.test.ts`, which drives
`listTasksImpl` for real, 396 is covered and the other two genuinely are
not.

That is rule 2 of #3214 one level up: *the test must reach the site*,
and a unit test of the collaborator never does. Had I stopped at the
first instrument I would have reported three uncovered resolvers, one of
them wrongly.

## What 793 costs on a renamed board

`searchTasks` backs the **CREATE-time near-duplicate check**. Without
the resolved lanes threaded, search stops excluding the board's archive
lane, and creating a task can be refused as a duplicate of one the
operator archived long ago — with no way to see why, because the
matching card is not on the board. Precisely the symptom #3160 set out
to fix; this pins the wiring that delivers it.

## An assertion I got wrong, and the correction

I expected an unreadable workflow list to leave `archivedColumns`
**undefined** via the call-site `.catch(() => undefined)`. It does not:
`resolveProjectColumnsForRoles` catches internally and returns its
**legacy-seeded** set, so `Set(["archived"])` is threaded and the
`.catch` never fires on that path. Two layers fail soft and the inner
one wins.

The case now asserts the guarantee that actually holds either way —
**never an empty set** (which would exclude nothing and return archived
rows in every search), legacy id always excluded. Recorded at the site,
because the mechanism is not obvious from the call.

## Flagged, not papered over

**`reads.ts:615` is left uncovered on purpose.** It composes Drizzle
conditions and runs them against `layer.db` with no injectable seam, so
pinning it needs a real database and belongs with the `.pg` suites. A
test asserting "the query was built" rather than "the rows were
excluded" would satisfy the ratchet and prove nothing.

Also flagged from this sweep: `workflow-analytics.ts` and
`team-analytics.ts` (4 resolvers) are **unmeasurable in my environment**
— their renamed-lane coverage lives in `.pg` suites, and this worktree
has no TCP PostgreSQL (`pg_isready` reports a Unix socket; the harness
probes TCP, so `pgDescribe` correctly skips). Not claimed either way.

## Census

**Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0`.** Converts
nothing; closes coverage on a conversion the census already counts as
done.

## Verification

```
as written                    Tests  4 passed (4)
BLIND reads.ts:793            Tests  1 failed | 3 passed (4)
restored                      Tests  4 passed (4)
```

Anti-vacuity case included: every other assertion reads a mock's
arguments and would pass if the search were never reached, so one case
pins that the primary search path actually ran. Typecheck clean.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 11:50:02 -07:00
gsxdsm
0bdc9bf4fb fix(dashboard): archived tasks stayed in the research picker on a renamed board (#3215)
## The defect

The enrich-mode task picker filtered with `task.column !== "archived"`.
On a board whose archive lane is renamed, that matched nothing — so
filed-away tasks stayed in the picker and an operator could attach
research findings to work they had deliberately archived.

## Census before / after

| | before | after |
|---|---|---|
| COLUMN guards (backlog) | 10 | **9** |
| `ResearchTaskActionModal.tsx` | 1 | **0 — converted** |

Baseline re-recorded in the same commit; `--strict` green.

## This site was declined twice, and I wrote the second wrong estimate

#3213 left it counted, correctly, on the note that was here — which was
mine. Both prior cost estimates were wrong, so this corrects my own
work:

1. **"Needs a data-fetch change"** — reasoned about
`columnFlagsByTaskId`, a per-**task** map built from board-resident
rows. Right that such a map can't help (archived rows are exactly what a
board map omits), but this guard asks a per-**column** question, so it
never needed one.
2. **"Needs prop threading, MainContent → ResearchView → here"** — right
that the answer is column-keyed, wrong about where it lives. `ListView`
builds `columnFlagsById` *inline*, which made it look like the owner.
The data is `useBoardWorkflows`, a hook already called from `App`,
`Board`, and `HeaderWorkflowSwitcherSlot`.

**Measured cost: one file.** The modal already takes `projectId`, and
`ResearchView` renders it only when a finding is open (`open` hardcoded
beside `if (!finding) return null`) — so the hook cannot fetch for a
closed modal, which was the one real objection to calling it here.

Union across workflows keyed by column id, first declaration wins — the
same convention `ListView` uses, so the two cannot disagree about a
shared id. `isArchivedColumnRole` fail-softs to the legacy id when a
column has no flags, so an unresolved workflow behaves exactly as the
literal did.

## Tests — the invariant, not the repro

Per the surface-enumeration rule, four cases: renamed archive lane,
legacy id, unresolved workflow (fail-soft), and a second workflow's
archive lane through the cross-workflow union. A repro-only test would
pass on the legacy board and prove nothing about the case the guard
exists for.

**Anti-vacuity control:**

| | renamed lane | union | legacy id | fail-soft |
|---|---|---|---|---|
| pre-fix literal | **FAIL** | **FAIL** | pass | pass |
| converted | pass | pass | pass | pass |

The legacy and fail-soft cases hold in both directions **on purpose** —
they pin that this conversion did not change the pre-resolution answer.
Flagging that so 4/4 isn't read as four independent proofs.

## Measured

| check | result |
|---|---|
| `census --strict` / `check-fnxc-future-dates` | exit 0 / exit 0 |
| `eslint` | clean |
| `tsc -p tsconfig.app.json` (the config that actually covers `app/`) |
exit 0 |
| new tests | 4/4 |
| `pnpm test:gate` | exit 0 (744 tests) |

## Note on process

My first attempt at the control silently did nothing — the revert script
threw a `SyntaxError`, so the "pre-fix" run was the fixed code and
reported 4/4. Caught it because the error printed. The table above is
from the re-run.
2026-07-31 11:42:15 -07:00
gsxdsm
c66b434b7b fix(self-healing): a renamed hold lane re-logged the same overlap blocker on every sweep (#3216)
## The defect

`clearStaleBlockedBy` keeps a per-task memo of which overlap blocker it
already logged, so a sweep running every few seconds doesn't repeat the
same line forever. The memo was retained only while the card sat in a
column matching the literal `todo` — so on a renamed board it was
dropped on **every** sweep and `still blocked by file scope overlap with
<id>` was re-logged each time.

## Census before / after

| | before | after |
|---|---|---|
| COLUMN guards (backlog) | 9 | **8** |
| `packages/engine/src/self-healing.ts` | 1 | **0 — converted** |

Baseline re-recorded in the same commit; `--strict` green. (Counts
follow #3215, which took 10 → 9.)

## The stated blocker was not real

The note here declined the conversion because the lane prefetch is keyed
on `candidates`, *"which this closure helps build"*. Measured — it does
not:

```
6033|  for (const task of blockedTasks) candidates.set(task.id, task);
6034|  for (const task of queuedDependencyTasks) candidates.set(task.id, task);
6036|  for (const [taskId, lastLoggedBlockerId] of this.preservedQueuedOverlapLogged) {   <- only CLEARS memos
```

`candidates` is fully populated two statements earlier, and this loop
only clears memo entries. So the prefetch was hoistable; it now sits
above the loop. That is a pure move of a read-only computation with no
conditional between the two positions.

Reaching the lane clause already proves the id is a candidate —
`!candidates.has(taskId)` is the first arm of the same `||` chain, so
short-circuit means the lane question is only asked for ids the prefetch
covered (`referencedIds.add(task.id)` runs for every candidate).
`lanesOf` still falls back to the legacy set, so an unresolvable
workflow answers exactly as the literal did.

This is the second inherited "too expensive" estimate to fail on
inspection this session (see #3215). Both were written in good faith and
both were checkable in a few minutes.

## One thing typecheck caught that review would not have

`memoTask?.column !== "todo"` was **also** the undefined check, and tsc
narrowed the later clauses on it. Replacing it without that arm compiled
clean to the eye but broke narrowing — `TS18048: 'memoTask' is possibly
'undefined'` on the next line. `|| !memoTask` is now explicit rather
than implied.

## Evidence

The test drives the sweep **twice**, because a single pass cannot
observe a dedup memo at all.

| | pre-fix literal | converted |
|---|---|---|
| `still blocked by file scope overlap` log lines | **2 — FAILS** | **1
— passes** |

Failure message against the pre-fix code: `expected [ [ 'FN-DEPENDENT',
…(1) ], …(1) ] to have a length of 1 but got 2`.

Worth correcting the record: the note called the cost *"a duplicate log
line, not a wrong lifecycle decision"*. The lifecycle half is right —
but it is a duplicate on **every sweep**, so it is recurring log spam,
not a one-off. That is a bigger cost than the note implies, though still
not a correctness bug.

| check | result |
|---|---|
| `census --strict` / `check-fnxc-future-dates` | exit 0 / exit 0 |
| `eslint` / engine `tsc --noEmit` | clean / exit 0 |
| self-healing + overlap suites | 15 / 21 / 6 passed |
| `pnpm test:gate` | exit 0 (744 tests) |

Reused the existing `RENAMED_BOARD_IR` harness in that file rather than
building a new one.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Improved cleanup of stale workflow blockers, including renamed
workflow lanes.
  - Prevented duplicate overlap warnings during repeated cleanup.
- More reliably preserves valid queued overlaps while ignoring missing
or inactive tasks.

- **Tests**
- Added regression coverage for repeated stale-blocker cleanup and
duplicate warning prevention.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 11:29:25 -07:00
gsxdsm
2868eb4797 docs(learnings): blind the resolver to find uncovered conversions (#3214)
Sibling to #3203 (`a-falling-count-is-not-evidence`), which records that
a metric moving is not proof the system moved. **This is the positive
procedure**: how to find out whether a landed conversion is held by
anything, and how to write a test that holds it.

## The measurement it is written from

Of **64 resolved lane sets in `self-healing.ts`, 26 had no test that
could distinguish them from the literal they replaced** — including
three conversions I shipped that same day, and two halves of sweeps I
had already recorded as covered.

## The procedure

```
- const reviewColumns = await resolveProjectColumnsForRoles(this.store, REVIEW_ROLES);
+ const reviewColumns = new Set<string>(["in-review"]);
```

Suite fails → covered. Suite passes → nothing in the tree can tell the
conversion from the literal. One resolver, one 17-second run — cheaper
than writing the conversion was.

## Why the census cannot answer this

| instrument | question |
|---|---|
| census / lane-wiring ratchet | is this site written in the resolved
vocabulary? |
| blinding | does anything break if it stops being? |

Neither substitutes for the other. A conversion merged with 204 green
tests behind it and zero able to see it.

## Four rules, each paid for by a test that proved nothing

1. **Blind each resolver separately** — coverage is per-resolver, not
per-sweep. Twice a sweep recorded as done was half-done, because control
flow short-circuited before the second guard.
2. **The fixture must reach the branch the resolver gates.** A card in a
renamed *wip* lane cannot exercise a *terminal* skip — it is caught by
the wip∪review set first.
3. **Assert a path-specific side effect, never a return value.**
`outcome === "reclaimed"` is reachable without the guarded branch.
4. **A store fake must honour `options.column`.** Flat and call-order
stubs answer identically whatever column is requested — a fake that
ignores its own filter cannot see a filter bug.

## The two shapes a ratchet cannot distinguish

- **resolved gate, literal branch** — reads as *unwired*, was a live
defect (#3208: a working agent lost its task link)
- **passed-but-unread** — reads as *wired*, is dead code (#3212)

A ratchet counting call sites scores the first as debt and the second as
done. Both wrong.

## Why a doc and not more PR comments

Everything above currently lives in ~20 PR descriptions. The next person
to touch a lane conversion will not read those. `docs/solutions/` is
where this project already keeps the things it learned the expensive
way, and the frontmatter (`applies_when: deciding whether a lane
conversion is actually protected by a test`) is what makes it findable.

## Verification

`pnpm test:gate` 13 + 161 + 499 + 71 · lint · fnxc-dates (TZ=UTC) ·
`self-healing-docs` 2 passed. Docs only; no changeset, per the AGENTS.md
rule for internal docs.
2026-07-31 11:18:23 -07:00
gsxdsm
aa1655ccd9 fleet: reclassify the census tail — 10 → 2 guards, all reasoning already in the code (#3213)
## Census before / after

```
                        before   after
COLUMN guards (backlog)     10       2
DELIBERATE-LITERAL         138     148
```

Baseline re-recorded in the same commit; `--strict` green.

## This converts nothing — the tail was never backlog

All ten remaining guards already carried an explicit in-code decision.
**None carried the `DELIBERATE-LITERAL` marker the census reads**, so
each re-appeared to every fleet pass as if unexamined. That is the whole
defect this fixes.

| site | the reasoning already at the site |
| --- | --- |
| `audit-ops.ts`, `moves.ts` | the degraded fallback arm of an
**already-converted** site; the live arm uses the resolved lane set |
| `scheduler.ts` ×2 | *"LEFT COUNTED"* — an await behind the
`tracked.has` re-entrance guard lets two updates double-start a monitor;
the sibling is the measured-expensive `task:updated` emit path (26 sites
against 7) |
| `notification-service.ts` | this method and its only caller are
**sync**, reached from a listener the store invokes as `(task: Task):
void`; resolving makes the chain async and reorders notification
classification against every other `task:updated` handler |
| `lifecycle-ops.ts` | *"Recorded rather than converted"* — dead code |
| `task-id-integrity.ts` | sync, no store-scoped read; converting alone
would disagree with `getLiveTaskColumn` |
| `triage.ts` | *"LEFT COUNTED until then"* — wants a non-sync-resolved
lane answer |

## Marker placement is load-bearing, and I got it wrong twice

The census reads a node's **leading** comments. A marker in a nearby
block comment attaches to the wrong node and is **silently ignored** —
it reads as reviewed while the count still lists the site.

- `task-id-integrity.ts` — my first marker went into the block comment
above the `const`; the literal is in the `return`. Count stayed at 1
until I moved it.
- `ResearchTaskActionModal.tsx` — marker added, **measured that it did
not register**, reverted.

Every edit was verified by re-running the census, not assumed. That is
the only reason the count actually moved.

## Two sites deliberately left counted

- **`ResearchTaskActionModal.tsx`** — the literal sits mid-expression
inside a `.then()` chain, so no marker can attach. The census's own
guidance is to hoist it into a named helper; the site's note asks for
that to be someone's deliberate change rather than a drive-by, so it
stays counted and honest.
- **`self-healing.ts`** — the memo closure I converted and reverted in
#3049. Its note: a renamed board costs a duplicate log line, not a wrong
lifecycle decision.

## Correction I owe on the measurement itself

For many turns I reported "zero unclaimed guards". That came from a bug
in **my own** query — `byFile` is an array of `[file, count]` pairs and
I had switched to `Object.entries()`, which yields `[index, pair]`, so
`n > 0` was always false and the filter returned zero regardless of
state. It agreed with reality while open PRs held every file, which is
why it went unnoticed; it was still wrong, and a constant zero against a
falling backlog should have prompted me to check it sooner.

## Verification (measured)

- engine `self-healing` + `scheduler` suites — **1003 passed / 56
files**
- core `task-id` / `moves` suites — green
- `tsc --noEmit` clean in core, engine and dashboard; `eslint` clean
- `pnpm test:gate` — green
- `lifecycle-column-census --strict`, `check-lane-wiring`,
`check-sql-column-literals`, `check-fnxc-future-dates` — green

No changeset: `@fusion/core`, `@fusion/engine` and `@fusion/dashboard`
are private, and no runtime behaviour changes.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified internal annotations for archived, in-progress, and
in-review workflow states.
* Documented fallback behavior and timing safeguards across lifecycle,
scheduling, notification, and triage flows.

* **Chores**
* Updated internal lifecycle tracking baselines to reflect current
annotations and state coverage.

* **Bug Fixes**
  * No user-visible behavior changes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 11:15:40 -07:00
gsxdsm
f94ff68391 docs(engine): record that agentParkedColumns is inert at this call site (passed-but-unread) (#3212)
Closes the last entry I could not pin on the #3115 coverage map — by
establishing **why** it cannot be pinned, rather than leaving it open or
forcing a green test around it.

## `agentParkedColumns` is inert at this call site

It influences exactly one output — `shouldPreserveParkedLink` — and
`recoverAgentsRunningOnInactiveTasks` **never reads it**. The gate is:

```ts
if (proof.hasFreshRun || proof.hasActiveExecution) continue;
```

Neither depends on the lane. So passing a resolved set changes nothing
today, and blinding it back to the legacy ids leaves every test green
**because there is no behaviour to observe**. That is not a coverage gap
— there is nothing there to cover.

## Kept, not deleted

- Removing it makes this call site read as **unwired** to the
lane-wiring ratchet, inviting the next worker to "fix" it by re-adding
exactly this.
- If the gate ever adopts `shouldPreserveParkedLink` — the
parked-specific semantics the sibling sweep uses
(`recoverDriftedAgentTaskLinks`, wired in #3208 an hour ago) — the
resolved set is already correct here.

## The shape worth naming

**Passed-but-unread** is the mirror of the **resolved-gate,
literal-branch** defect #3208 fixed. Both read as converted while
deciding nothing — and only one of them is a bug.

A ratchet that counts call sites cannot tell them apart: #3208's site
looked *unwired* and was a live defect; this one looks *wired* and is
dead code. That is why the distinction belongs in a comment at the site
rather than in a baseline number.

## Map status

**21 of 26 pinned**, 1 established as unpinnable-by-construction, 4
remaining with obstacles recorded:
- `reclaimHoldColumns` / `reclaimReviewColumns` — both audit paths emit
the same `branch:auto-reclaim` type, differing only by a `trigger`
string the branch-level scan also produces.
- `completedHoldColumns`, `wsDoneColumns`, `doneMetaColumns` and others
have since gone green from other workers' PRs.

## Verification

`pnpm test:gate` 13 + 161 + 499 + 71 · lint · census `--strict` ·
fnxc-dates (TZ=UTC) — green. Comment-only change.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added an explanatory note clarifying parked-column handling and its
connection to future parked-link preservation behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 11:05:17 -07:00
gsxdsm
c027a72d23 fix(engine): a live agent lost its task link on a renamed hold lane (found while testing, not converting) (#3208)
**A defect, not a coverage gap** — found while trying to pin
`agentParkedColumns` from the #3115 map.

`recoverDriftedAgentTaskLinks` enters its preservation branch on a
**resolved** question (`isPreWipColumn`) and then decided it on a
**literal** one: `evaluateParkedAgentTaskLink` was called without
`parkedColumns`, so parked-ness fell back to `todo`/`triage`. The
sibling sweep passes the resolved set; this call site did not.

On a renamed board: the card is pre-wip, `isParkedTaskColumn` says no,
`shouldPreserveParkedLink` is false, and **an agent with a fresh
heartbeat run has its task link cleared — while it is working.**

`task-agent-sync.ts` predicted this in writing when the parameter was
introduced:

> *"turning a stale-link bug into a dropped-link bug, since the card
would be treated as unparked and its live agent link cleared"*

That is what an unpassed optional lane parameter costs — the same
missed-pair shape as #2956, #2963 and #3186.

## The test needed two fixture corrections, both caught by failing

- the renamed IR had **no hold column**, so no card could be pre-wip at
all;
- the **per-task selection readers** were missing, so `isPreWipColumn`
resolved the default IR and the branch was never entered.

Either alone made the case pass while exercising nothing. Third time
today a fixture passed for a reason unrelated to the resolver — a
pattern, not an anecdote.

## Measured

15 pass; removing `parkedColumns` from the call fails exactly this case.

## Note on how it was found

I had discarded a probe at this sweep earlier for failing to
discriminate. Coming back with the obstacle understood — the fixture
must reach the branch the resolver gates — turned a coverage miss into a
defect find. The five discards this session were not wasted; three of
them named the obstacle that made a later attempt work.

## Verification

`self-healing-agent-link-drift` **15 passed** · `pnpm test:gate` 161 +
13 + 499 + 71 · lint · lane-wiring — green.
2026-07-31 10:49:57 -07:00
gsxdsm
8661b739ff fix(scheduler): a board with TWO complete columns left dependents waiting forever (#3210)
## The defect

On a board declaring more than one complete-trait column — a merged lane
and a shipped lane, say — a card landing in the **second** one was never
recognised as finished, so nothing unblocked its dependents. Silent: no
error, the dependent just waits.

Two problems, the same shape:

1. **`TaskMoveLanes` carried one id per role.** That is right for
*"where should this card go"* and wrong for *"is this column one of the
finished lanes"*, which is a **membership** question. The payload could
not express such a board at all.
2. **`mergeParkedColumns` rebuilt `terminal` as `new Set([complete,
archived])`** — discarding `base.terminal`, which the sync IR path had
already resolved correctly, and narrowing a membership set back to
first-match-per-role.

Point 2 contradicted the note sitting directly above it in the same
file:

> `terminal` is a MEMBERSHIP set, and it is not the same question as
`complete`/`archived`. […] A workflow may declare more than one
complete-trait column […] and `to === parked.complete` sees only the
first and silently skips the rest.

The reasoning was already written down. The overlay added later didn't
honour it.

## Fix

`TaskMoveLanes.terminal?: readonly string[]`, filled from
`columnsWithFlag(ir, "complete"|"archived")` rather than the first-match
`resolveLifecycleColumns`, and the merge is now a **union** of base,
payload, and the single lanes.

Optional, so all 12 emitters and every listener keep compiling — a
listener that ignores it is exactly as correct as before. Union is the
direction `scheduler.ts` already argues for at line ~422: a superset
costs one extra query; a subset **silently withholds work from a
finished card**.

`complete` deliberately stays first-match — a set would be the wrong
shape for a move *target*. Both questions now coexist rather than one
replacing the other.

## How it was found, and what it corrects

Supplying `task:moved` lanes fixed every *other* renamed-board case in
`scheduler-renamed-hold-events` — measured **10 passed / 1 failed** —
and left exactly this one broken. That same measurement is why I
narrowed my earlier claim on #3082: the other behaviours were never
broken in production, because all 12 emitters already carry lanes. This
is the residue that was genuinely broken.

## Tests — both with anti-vacuity controls

| control | result |
|---|---|
| revert `toTaskMoveLanes` | **2 of 4** core tests fail (the terminal
pair) |
| revert the scheduler union | the new engine test fails, **and only
it** (1 failed / 11 passed) |
| both restored | 4 passed, 12 passed |

The 2 core tests that pass either way are shape invariants asserted on
purpose (`complete` must stay first-match; a column-less IR returns
`undefined` rather than an invented lane) — flagging that so the control
isn't read as 4-of-4.

The pre-existing scheduler case emits **without** lanes, which no
production emitter does, so it exercises the sync fallback. The new one
emits `toTaskMoveLanes(ir)` — the shape that actually ships.

## Measured

| check | result |
|---|---|
| `@fusion/core` / `@fusion/engine` tsc | exit 0 / exit 0 |
| eslint | clean |
| `census --strict`, `check:fnxc-future-dates`, `check:changesets` |
exit 0 |
| every `TaskMoveLanes` consumer | 24 passed |
| `pnpm test:gate` | **exit 0 — 744 tests, up 12** |

## Census

No guard converted; this is a payload-shape fix. Backlog unchanged at
11, all deferred.
2026-07-31 10:47:15 -07:00
gsxdsm
8d393422ac chore(fnxc): tighten the future-dates baseline — merge-queue-ops-2 4 -> 3 (#3211)
One-line baseline tightening, produced by the gate's own auto-tighten
path.

`check-fnxc-future-dates` deliberately auto-tightens rather than failing
on a drop, because its population moves with the calendar and a drop has
**no author** — the counterpart asymmetry to
`check-inert-sync-lane-conversions`, where a drop *does* have an author
and must fail. Any gate run regenerates this; `main`'s committed
baseline had simply not caught up.

**Why this isn't churn:** left loose, the baseline permits 4 future
stamps in a file that now has 3. That slack silently absorbs one genuine
future-dated stamp — precisely the failure this gate exists to catch,
and one the fleet hit four times in a single day (`scheduler.ts`, a
scheduler PG test, `task-update.ts` twice by different authors), each a
real time on the wrong day that passed locally and reddened `main` for
everyone else.

Verified: both `check-fnxc-future-dates` and
`check-inert-sync-lane-conversions` green on the tightened baseline.

No changeset: tooling baseline, no published-package surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:47:03 -07:00
gsxdsm
e9a57ca8ba docs(learnings): a falling count is not evidence that anything changed (#3203)
## The metric counterpart to #3200

#3200 (merged) records the **shapes** an inert conversion takes, and its
grammatical tell is the portable one: *if a claim can be written without
running anything, it has not been tested.* I offered this material there
and said I would write it as a sibling rather than bloat that doc; it
merged without it, so here it is.

That doc is about **claims**. This one is about **numbers**.

## The tell

**A count that falls is not evidence that anything changed.** Every gate
here reports a number, and a number goes down three ways — work
happened, the code got denser and the scan stopped matching, or someone
lowered the allowance. Only the first is progress, and from inside the
check all three look identical.

Four times in one phase:

| what moved | what actually happened |
|---|---|
| census 12 → 2 (`scheduler.ts`, #3051) | ten guards routed through
`resolveTaskWorkflowIrSync`, which answers with the DEFAULT board under
PostgreSQL. Byte-identical. Refuted in #3058 |
| census 45 → 44 (`triage.ts`, #3114) | converted the exact arm #3108
flagged hours earlier. #3126 reverted it — three PRs for one line |
| ratchet 20 → 15 | not a conversion: #3065 rewrote `a === x \|\| a ===
y` as `set.has(a)`. It printed *"total fell — re-record"*, which would
have **permanently retired live guards** |
| ratchet 22 → 9 | `scheduler.ts` reported **0** while 13 guards still
fell back to the default board |

Rows three and four are the dangerous shape: **the gate went quiet
exactly when someone improved the code**, and the remedy it suggested
was to lower the allowance.

## Also recorded

- **One defect, four spellings** (#3062, #3068, #3079, #3181) — each fix
correct about the shape in front of it and blind to a respelling. The
lesson is not "write a better regex": enumerating consuming syntax is a
losing game, and the durable form keys on the *source*.
- **An instrument that runs nowhere and one that cannot fail are the
same defect.** `check:inert-sync-lanes` was invoked by nothing for six
PRs; `check:quarantine-ledger` ran nowhere *and* omitted `--strict`, so
wiring it alone would have been theatre. Includes the mechanical audit
that finds both.
- **Base drift makes branch numbers incomparable** — three false alarms,
one of them mine, from comparing against a remembered figure. The
procedure that works is extracting both scripts and running them against
one tree; that is how #3169 and #3181 were shown additive (22 = 13 + 7 +
2), which decided merge order and collapsed one into six lines inside
the other.
- **A pick-work list at 100% false positives**, because under-reporting
deferrals is the direction that manufactures the #3108 → #3114
collision.

## Every claim is a measurement

No mechanism here is derived from reading. Nine PRs cited, each the one
that produced or refuted the finding — including the ones where I was
wrong: a stale number I mistook for a regression, and two future-dated
stamps of my own that the full ratchet set caught before they shipped
(one earlier one it did not, and that broke `main`).

## Census before / after

```
before:  COLUMN guards (the backlog):   12
after:   COLUMN guards (the backlog):   12
```

Docs only.

## Verification

`test:gate` exit 0 · `fnxc-future-dates`, `lifecycle-columns`,
`inert-sync-lanes`, `quarantine-ledger`, `inert-flag-seams`,
`lane-wiring`, `sql-column-literals` — all exit 0 · `pnpm lint` clean.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added guidance for interpreting workflow metrics and avoiding
misleading conclusions from declining counts.
* Documented detection blind spots, branch comparison issues, false
positives, and validation procedures.
* Included a practical checklist for reviewing metrics, quality gates,
comparisons, and potential conversions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:44:21 -07:00
gsxdsm
bad39e2ca3 test(core): ledger the legacy-id collections that gate a live column — the class the census cannot count (#3209)
## What

A population ratchet over **legacy-id collections consulted against a
live column value** (`SOME_SET.has(task.column)`), recorded as 23 sites.

## Why

The census scans `===`/`!==` comparisons. A Set or array literal is a
**definition**, so no census run has ever pointed at one. Three
found-by-hand defects came from that blind spot:

| collection | symptom |
|---|---|
| `GITHUB_TRACKING_EDITABLE_COLUMNS` | operator could not toggle GitHub
tracking **at all** on a renamed board — no error, affordance absent
(#3149) |
| `TIME_INDICATOR_COLUMNS` | wrong elapsed-time indicator on cards |
| `BLOCKER_ESCALATION_COLUMNS` | escalation skipped renamed lanes |

#3149 enumerated the population by hand and concluded *"this is where
the remaining renamed-board defects actually live."* A number in a PR
body rots. This is that enumeration as a ratchet.

## Census

**Unchanged — `AVAILABLE: 0` before and after, 12 documented deferrals
both sides.** This PR converts nothing. It ratchets a class the census
*structurally cannot see*, which is the point: the backlog reading zero
has never meant the lane vocabulary is fully converted, only that the
measurable part is. Recording that plainly instead of claiming a delta
this change does not produce.

## What it claims, and what it deliberately does not

It claims the **population** is the recorded set. It does **not** claim
each site is correct — 20 of the 23 are #3149's assessment ("most are
already correct, either no-flags fallbacks or seed-then-add resolved
sets"), and I did not re-verify them. Blessing sites I have not read is
how a ledger becomes a list of things someone once glanced at. A new
entry fails the test and a human reads that **one** site; that is the
entire mechanism.

## My own detector's pick-work list was 100% false positives

Measured, and the reason this ships with **no candidate list**. The
heuristic "no role-helper call in the file" flagged three sites; all
three were fine:

- `agent-role-policy.ts:32` — a documented **FLAGGED, NOT FIXED**
deferral with its reasoning recorded
- `DocumentsView.tsx:88` — already converted, flags-first; the flags
arrive as a threaded **object**, so a scan for resolver *calls* cannot
see the conversion
- `agent-assignment.ts:118` — a `DELIBERATE-LITERAL` fallback behind an
injected `countsAsAssignmentLoad` callback, reviewed `2026-07-31-05:40`

That is the same failure `--triage`'s pick-work list had before #3194
fixed it, from the same cause: **inferring "unexamined" from the absence
of a pattern rather than from evidence.** A detector that cannot
distinguish "not yet looked at" from "looked at and settled" must not be
pointed at a work queue. It can still hold a population steady, which is
all this does.

## Verification

Mutation-verified in **both** directions — a ledger fails by missing
additions *or* by keeping ghosts:

```
### baseline                                        Tests  4 passed (4)
### MUTATION 1 — new unrecorded gating collection
+   "packages/engine/src/worktree-pool.ts :: NEW_LANE_GATE",
                                                    Tests  1 failed | 3 passed (4)
### MUTATION 2 — recorded site vanishes (ghost)
+   "packages/engine/src/worktree-pool.ts :: managedRenamed",
+   "packages/engine/src/worktree-pool.ts :: managed",
                                                    Tests  2 failed | 2 passed (4)
### restored                                        Tests  4 passed (4)
```

Two anti-vacuity cases guard the detector: it still finds the
collections whose defects motivated the file, and it does **not** claim
plain comparisons (asserted against `self-healing.ts`, dense with column
comparisons and no gating collection) — pulling those in would
double-count a class that already has a gate.

## Flagged, not guessed

- **Line numbers are excluded** from ledger entries — they drift with
unrelated edits and would fail this test for reasons that are not about
lane vocabulary.
- **Comments stripped before scanning:** `TaskDetailModal.tsx` and
`TaskCard.tsx` both quote their own collection by name in FNXC notes
explaining the bug it caused. Counting prose would fire the ledger on
the files that document the hazard most carefully.
- **Stated reach limits** (in-file): only *named* collections consulted
as `.has`/`.includes`; the argument must mention column/lane;
property-reached collections are missed. A miss is a site nobody is
watching — not a false green on a listed site.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:44:08 -07:00
gsxdsm
215f09d88f fix(census): the bare command could not say the conversion queue is EMPTY — and a test fix for main (#3207)
## Why this exists

The fleet instruction is *"claim the largest unclaimed census file
cluster (`node scripts/lifecycle-column-census.mjs`)"*. That command
cannot answer it. The availability verdict lived **only** behind
`--claims`, which shells to `gh`:

```
line 342:  if (claims && !json) {
```

So a worker following the instruction literally sees per-file counts,
reads a nonzero backlog as a work queue, and picks a file whose guard is
already documented as deferred. Counts alone cannot separate *work left*
from *debt left*.

**Measured cost:** the queue reached **zero unexamined guards** while
dispatch continued. I re-audited the last three candidates —
`merge-queue-ops-2`, `lifecycle-ops`, `notification-service` — and all
three were already documented. Only one was reclassifiable, and by
**deletion** rather than conversion (#3205).

## What the bare command prints now

```
  COLUMN guards (the backlog):   11

  CONVERSION QUEUE EMPTY: all 11 remaining column guard(s) carry a documented deferral note.
  There is no unexamined guard to claim. A nonzero backlog above is DEBT, not a work queue.
  Re-read the note at a site before converting it; run --claims to also check open-PR ownership.
```

Or, when work does exist: `N unexamined guard(s) remain (no deferral
note) — run --triage to list them by file.`

**Local signals only**, so it is honest offline. It reports what it can
prove — no *unexamined* guard remains — and explicitly does **not**
claim the files are unclaimed, because only `--claims` sees open PRs. No
count, no exit code, `--strict`/`--json` untouched.

## Three commits, deliberately separated

1. **`refactor`** — move `FLAG_MARKERS` + the 40-line window into the
lib as `hasDeferralNote()`, verbatim. It was a private const plus an
inline `.slice()` in the CLI, so the rule deciding where the fleet is
sent had **no test in either direction**. Proven identical on the real
tree: `11 documented / 0 unexamined` before and after.
2. **`feat`** — the verdict + 6 tests.
3. **`fix`** — an unrelated pre-existing failure (below).

## The test fix — this one is turning main red

`attributes a remaining file to the open PR that touches it` asserted
over `out.slice(out.indexOf("UNCLAIMED:"))`, which runs to **end of
output** and so also covers the `SYNC-RESOLVED` section printed
afterward. That section legitimately lists `scheduler.ts`.

Latent until `topRemainingFile()` returned `scheduler.ts` — which
happened as the backlog shrank, **a state every conversion moves
toward**. Confirmed pre-existing: clean `origin/main` runs `42 passed /
1 failed` with the identical message.

## Evidence

| check | result |
|---|---|
| `hasDeferralNote` tests | both directions, boundary exact at 40 above
/ not below, 5 real phrasings |
| verdict control (by hand) | one tracked undocumented guard → **11 →
12**, verdict flips to `1 unexamined`; removed → restored |
| test-fix anti-vacuity | claim split broken → **FAILS**; restored →
passes |
| census file | **49 passed** (was 42 passed / 1 failed) |
| `census --strict` / `check:fnxc-future-dates` | exit 0 / exit 0 |
| `pnpm test:gate` | **exit 0** (732 tests) |

The verdict control was **invalid on the first attempt** — my probe file
was untracked and `git ls-files` never scanned it, so the verdict did
not flip and nothing was proven. Recording that because a control that
silently proves nothing is the exact failure this PR is about.

## Census before / after

No guard converted here; this is tooling. Backlog unchanged at 11, all
deferred.
2026-07-31 10:38:48 -07:00
gsxdsm
23403e1426 test(engine): pin the agent sweep's terminal skip (21st resolver, after a corrected fixture) (#3206)
`agentLinkTerminalColumns` was uncovered on the #3115 map. The existing
case uses `todo` and `in-progress`, so the terminal skip is never the
deciding branch.

## The corrected fixture is the lesson

An earlier attempt of mine put the card in a renamed **wip** lane and
stayed green when blinded — **correctly**. Such a card is caught by the
wip∪review set first, so the terminal resolver never decides anything.
The card has to rest in a renamed **complete** lane for this guard to be
the one that matters.

That is the same class as my two discards on the branch-conflict sweeps:
**the fixture has to reach the branch the resolver gates.** A test can
exercise the sweep, pass, and still never touch the line under test.

## What the literal costs

A finished task's agent is not skipped, so the sweep **unlinks an agent
from a task that completed normally** — churn on a row that needed no
repair, and a lost link if that agent was about to be reused.

## Observable

Asserts `syncExecutionTaskLink` — the action the guard prevents — rather
than a return value, per the rule from #3202.

## Measured

421 pass; blinding `agentLinkTerminalColumns` fails exactly this case.

**21 of 26 pinned** across 20 merged PRs.

## Still open, with the obstacle recorded

`reclaimHoldColumns` / `reclaimReviewColumns` resist: both audit paths
in that sweep emit the same `branch:auto-reclaim` type, differing only
by a `trigger` string the branch-level scan also produces, so no
observable I found isolates the bucket resolvers from the branch scan.
`agentParkedColumns` needs a fixture where parked-ness changes the
outcome — mine forced the proof true via a fresh run.

## Verification

`self-healing.test.ts` **421 passed** · `pnpm test:gate` full pass ·
lint — green.
2026-07-31 10:33:25 -07:00
gsxdsm
5c5f6d8155 fix(core): mark the two archived STATE sites at the site — converting them destroys live work (#3157)
The LANE/STATE triage (#3154) found that **two of the eight** Drizzle
`archived` sites are STATE markers that must never be resolved. That
classification lived only in `archived-column-gate-parity.test.ts`.

A coordinated three-encoding conversion **edits these files**. A
converter working file-by-file sees the same `eq(column, "archived")`
shape as the six LANE sites, with nothing in front of them to tell the
two apart.

So the markers go at the sites.

## `task-mutation-ops.ts` — `cleanupArchivedTasksImpl`

Enumerates rows Fusion itself archived, then **removes their
directories**.

Widening it to the resolved archived-lane set would feed cards **merely
resting in a board's archived-trait lane** into a filesystem delete.
This is the only site in this family where a wrong conversion **destroys
work** rather than hiding an affordance.

## `async-self-healing.ts` — `listSoftDeletedColumnDriftCandidates`

Finds soft-deleted rows whose column **drifted** from the marker they
are supposed to carry. Resolving it would classify a soft-deleted row in
a renamed archive lane as drift and "repair" a row that is already
correct.

## Why this is defensive rather than cosmetic

The triage exists to make the conversion safe. A classification the
converter **cannot see while editing the file** does not do that — it
only helps someone who happens to read the gate's test file first, which
is not how a file-by-file sweep proceeds.

Both are marked DELIBERATE-LITERAL with the reason and a pointer to the
parity test holding the full eight-site split.

## Measured

- Comment-only.
- Parity test **2/2**; archive + soft-delete suites — **5 files / 15
tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict` and
`check-sql-column-literals` clean.
- **No census movement** — a DELIBERATE-LITERAL marker on a STATE site
is a classification, and these were never counted as lane debt.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:30:41 -07:00
gsxdsm
dd3bf6a764 docs(core): the archived TS remainder is EMPTY — enumerated, closing the triage (#3171)
#3156 sampled the TS inventory and said the conversion is *"the six LANE
Drizzle sites plus whatever small TS remainder is neither a fallback arm
nor a sentinel"*.

**That remainder is zero.** I left three sites unchecked when I wrote
it. All three are fallback arms:

| site | shape |
|---|---|
| `live-agent-count.ts:164` | `task.columnTerminalKind ?? (task.column
=== "done" ? … )` — the resolved value wins via `??` |
| `store.ts:1972` | `if (!lanes) return dep.column !== "done" && …` — an
explicit no-metadata branch |
| `branch-and-pr-entities.ts:525` | `lanes === undefined ? task.column
=== "archived" : task.column === lanes.archived` |

With the previously classified entries, **every** site in
`AUDITED_TS_SITES` is now accounted for as a fallback arm, a
STATE/sentinel comparison, or a converted guard's retained literal.
**None is an unconverted LANE guard.**

## So the cluster is done

"52 sites across three encodings" is fully triaged, and the convertible
work is the six Drizzle LANE sites plus the log-entry gate — all
additive, so the inventories never moved.

What the gate now protects is a population of **fallback arms and STATE
markers**, which is exactly what it should protect: each is the
documented answer for a caller that supplies no resolved set, or a
marker that must never be resolved.

A future **drop** in any of the three counts means someone removed a
fallback or converted a STATE site — both regressions. That is the check
this file was built to make, and is now the only check it needs to make.

## Enumerated, not sampled

I sampled this inventory twice and each pass changed the size estimate —
first "52 sites, real blast radius", then "six plus a small remainder".
A third estimate would have been worth less than a complete count, so
this pass covers every entry.

That is the honest close: the number stopped moving because I stopped
guessing at it.

## Measured

- Comment-only; parity test **2/2**.
- `tsc --noEmit -p packages/core` clean; census `--strict` clean. **No
census movement.**

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:30:30 -07:00
gsxdsm
230be28576 fix(core): the merge-queue enqueue guard was not debt — the code it guarded had no callers (#3205)
## The deferral note was right about the mechanism and wrong about the
remedy

`merge-queue-ops-2.ts` sat in the census as deferred debt behind this
note:

> Converting it properly means either making this path async or pushing
the trait read into SQL, both of which are store-architecture changes
rather than call-site conversions.

That is correct as far as it goes — the guard runs inside
`store.db.transactionImmediate`, so the only synchronous resolver
available (`resolveTaskWorkflowIrSync`) returns the DEFAULT workflow
under PostgreSQL and a "conversion" would be inert.

But it assumed the code needed converting. Measured across the tree:

```
=== every call site of .enqueueMergeQueueSyncInternal( ===
packages/core/src/store.ts:1775:  public enqueueMergeQueueSyncInternal(...)   <- the declaration itself
```

**Zero callers.** Every other occurrence of the name is a comment. The
live path is `enqueueMergeQueueAsync` (`task-artifacts-ops.ts:117`), and
that file already documented the deletion:

> Merge-queue enqueue is PostgreSQL-only via enqueueMergeQueueAsync …
The SQLite `enqueueMergeQueueSyncInternal` arm is deleted.

The arm was deleted; its declaration was not. The guard was unreachable
on the shipped backend.

## Change

- Deleted `enqueueMergeQueueSyncInternalImpl` (-85 lines) and its
`store.enqueueMergeQueueSyncInternal` entry point.
- Dropped the six imports that became unused
(`MergeQueueTaskNotFoundError`, `MergeQueueInvalidColumnError`,
`MergeQueueEntry`, `MergeQueueEnqueueOptions`, `normalizeTaskPriority`,
`MergeQueueRow`).
- Refreshed the three comments naming the removed symbol, so none points
at a deleted identifier. The
`handoffMergeQueueFailureInjectorForTesting` hook those comments sit on
is a **different** member and is untouched — it only mentioned the sync
arm as context.

## Census before / after

| | before | after |
|---|---|---|
| `packages/core/src/task-store/merge-queue-ops-2.ts` | 1 | **0 (entry
removed)** |

Baseline tightened by exactly one entry. **The 0 here is a deletion, not
a conversion** — recorded in the file's own FNXC note so the next worker
does not read it as a converted seam. This is the failure mode the
census warns about ("a count of 0 is the WORST case, not the best"), so
it is stated at the site rather than left to inference.

## Measured

| check | result |
|---|---|
| `census --strict` | exit 0 |
| `@fusion/core tsc --noEmit` | exit 0 |
| `eslint` (4 changed files) | clean |
| core merge-queue tests | **110 passed / 6 files**, incl.
`postgres/merge-queue-renamed-review-column.pg.test.ts` |
| `pnpm test:gate` | exit 0 (**732 tests**) |

No changeset: `@fusion/core` is private and this removes unreachable
code with no user-visible behavior.

## Flagged, not guessed

The other four deferral-note files remain deferred. I only reclassified
this one because its call-site count is a fact I could measure, not a
judgement. Whether `lifecycle-ops.ts:667` is likewise dead (it sits in
the legacy-SQLite polling-replica path) is a separate question I have
not measured, so I have not touched it.
2026-07-31 10:30:16 -07:00
gsxdsm
25fa5e7c44 fix(core): the log-entry archive gate, converted — the parity objection is met, not bypassed (#3165)
I converted this in #3110, the parity gate failed, and I reverted it.
**The gate was right** — and the reason was subtler than "one encoding
moved". Knowing it is what makes this conversion possible.

## Why the first attempt failed

My version hoisted the comparison onto a local:

```ts
const pgRowColumn = String(pgRow.column ?? "");
const rowIsArchivedLane = archivedLanes ? archivedLanes.has(pgRowColumn) : pgRowColumn === "archived";
```

That gate's TS scan keys on the **property** being named `column` —
deliberately, because the receiver is variously `task`, `row`, `dep`,
`t`. Losing the `.column` access dropped the TS count while SQL and raw
held steady, which it reads as divergence.

**Behaviourally identical, structurally invisible.** Same failure mode I
hit from the other direction in #3163, where I collapsed a Drizzle
fallback into a string array.

## The fix

Keep `pgRow.column === "archived"` **verbatim** as the fallback; add the
resolved path in front of it. No encoding's count moves, an unwired or
degraded caller behaves exactly as before, and the gate is **satisfied
rather than worked around** — the same additive shape as the six Drizzle
LANE sites (#3160, #3162, #3163).

## What it fixes

A LANE question: *"is this row in the board's archive lane, so logging
is read-only?"*

Against the literal, a card the operator filed away on a renamed board
kept **accepting log writes** — new activity accruing on closed work.
`deletedAt` covers the soft-delete half, which is why the gap is narrow
and why it stayed invisible: the common path is soft-delete.

## The recorded omission is retired properly

`log-entry-archived-lane-gate.test.ts` carried the renamed case as a
**deliberate omission** with its reason. It is now the first case in the
file, and the note explains why the earlier judgement changed rather
than quietly disappearing — a deferral that vanishes without explanation
is how the next reader loses the thread.

## Measured

- **3/3** in that file (renamed case added); parity test **2/2**,
inventories unmoved.
- **MUTATION**: dropping the resolved branch fails the renamed case and
leaves the legacy **control** and the live-lane **negative** green.
- log-entry / archived / audit suites — **3 files / 9 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals` clean.
- `check-fnxc-future-dates` is red on `main` from `task-update.ts`
(another lane's stamps), not from these files.

## Census

**Unchanged** — the literal remains the fallback arm, by design.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:27:31 -07:00
gsxdsm
762d232ad6 test(core): ratchet the sentinel-task-id argument at zero — the third inert-conversion mechanism (#3204)
## What

A test-only zero-population ratchet: no task-scoped lane resolver may be
called with a **string literal** where a row id belongs.

## Why

`resolveTaskLifecycleColumns(store, taskId)` and its siblings resolve
the workflow bound to *that task*. Hand one a literal and there is no
task to read a selection for, so the resolver falls back to the
**default board** and answers with full confidence. The call
type-checks, reads as a finished conversion, and is correct on every
board Fusion ships — because the default board is the answer it returns.

**This shipped.** `triage.ts`'s startup sweep called
`resolvePlannerLanes(this.store, "")` and built its swept-column set
from the result (#2806 measured it, #3201 fixed it). It was a
*sweep-wide* defect rather than a per-card one: it resolved once for the
whole board and could not be right for any workflow but the default, so
a card parked in a renamed hold column with a stale `planning` status
was never swept and held a planning admission slot permanently.

Note what this means for the other two inert mechanisms' fixes —
**making the resolver async would not repair it**, because the defect is
the argument, not the resolver.

## Why a guard and not just the existing E2E

`workflow-sweep-sentinel-task-id-live-e2e.pg.test.ts` covers the **one**
triage site and lives in the `.pg` lane, so it is skipped whenever no
PostgreSQL is reachable — including the merge gate. The defect is the
*argument*, which makes it visible in source text with no database, no
running engine, and no knowledge of what the resolver does.

## Census

**Unchanged — 0 guards before, 0 after.** This PR converts nothing; it
is a ratchet over a class the census structurally cannot see (the census
scans column literals, not resolver arguments). Recording that plainly
rather than claiming a delta this change does not produce.

Population of the guarded class is **zero today** — the only textual
match in the tree is prose in `triage.ts` documenting its own fixed bug.
A zero-population ratchet is the instrument here, not a weakness: it
cannot fail until someone reintroduces the defect, and it costs one
source scan.

## Verification

**Mutation-verified, not asserted.** Re-adding the exact shipped shape
to a real production file:

```
+   "packages/engine/src/replan-target.ts:188 — resolvePlannerLanes",
 Tests  1 failed | 3 passed (4)
```

Restoring the file returns it to `4 passed`. Working tree left clean.

Three anti-vacuity cases carry the file, because a scan that reports
success by finding nothing is otherwise indistinguishable from a broken
scanner:
- the matcher **does** fire on the historical text
(`resolvePlannerLanes(this.store, "")`);
- it does **not** fire on the ordinary shapes that fill the codebase
(`task.id`, `taskId`, `row.id`) — a matcher flagging everything would
pass the case above while being unusable;
- the walker still reaches production source (>20 real
`resolveTaskLifecycleColumns` call sites), which is what makes the zero
a measurement rather than an empty scan.

## Flagged, not guessed

- **Comments are stripped before scanning**, and here that is required
rather than tidy: `triage.ts` quotes the offending call verbatim to
explain the hazard. Counting it would make the guard fire on the file
that correctly documents the defect, training readers to silence the
guard instead of heeding it.
- **`resolveReboundTarget(ir)` / `resolveLifecycleColumns(ir)` are
deliberately excluded** — they are IR-scoped and take no task id;
including them would flag correct code.
- **Known limit, stated in the file:** a sentinel arriving through a
*variable* (`const id = ""; resolve(store, id)`) is invisible to a text
scan. The literal form is what shipped and what the next person is most
likely to write; the variable form still needs the `.pg` E2E. Two
instruments, different reach — not full coverage of the class.

No changeset: test-only, behavior-preserving, no published-package
surface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:27:18 -07:00
gsxdsm
40e64468d2 fix(dashboard): GitHub tracking was unreachable on a renamed board — a defect class the census cannot see (#3149)
The census backlog is verified-exhausted (12 guards, every blocker
re-checked in #3082). This is from the class **the census structurally
cannot count**, and it is a real capability loss.

## The defect

```ts
const GITHUB_TRACKING_EDITABLE_COLUMNS: Set<ColumnId> =
  new Set<ColumnId>(["triage", "todo", "in-progress", "in-review", "ideas"]);

function canTaskEditGithubTracking(column, workflowId) {
  return GITHUB_TRACKING_EDITABLE_COLUMNS.has(column) || workflowId === CODING_IDEAS_WORKFLOW_ID;
}
```

No resolved branch, no flags fallback. On a board whose lanes are
renamed this matched **nothing**, so the helper returned `false` for
every task and `showGithubTrackingSection` hid the section outright.
**The operator could not turn GitHub tracking on or off** — no error, no
explanation, the affordance simply absent. The only thing keeping it
reachable was the unrelated `builtin:coding-ideas` escape hatch on the
right-hand side.

## Why no gate saw it, and why this class matters now

The census counts **comparisons** against legacy ids. This is a **Set
literal — a definition** — consulted with `.has()`. Nothing in the
backlog ever pointed here. It is the same blind spot that hid
`TIME_INDICATOR_COLUMNS` and `BLOCKER_ESCALATION_COLUMNS`, both of which
were also found by hand rather than by any gate.

I found it by scanning for legacy-id **collections that gate a live
column value**, rather than for comparisons: **19 such sites** across
the tree. Most are already correct — either `if (!flags) return
LEGACY_…has(column)` fallbacks, or seed-then-add resolved sets
(`agent-reflection.ts`, `ephemeral-worker-manager.ts`). This one had
neither.

With the comparison backlog at 12 and every remaining entry blocked or
documented, **this is where the remaining renamed-board defects actually
live.**

## The fix

The set's meaning is "not finished" — every lane except complete and
archived — which is what the roles now express:

```ts
if (workflowId === CODING_IDEAS_WORKFLOW_ID) return true;
if (!columnFlags) return GITHUB_TRACKING_EDITABLE_COLUMNS.has(column);   // unchanged pre-fetch
return !isCompleteColumnRole(columnFlags, column) && !isArchivedColumnRole(columnFlags, column);
```

The caller passes `detailColumnFlags` — the **task-identity-guarded**
value. `workflowMoveMetadata` outlives a task switch, and this file's
own `2026-07-30-17:30` note records **six** review findings from
consumers that read around that guard. Passing the unguarded value would
answer about the previous card's workflow: worse than the legacy
fallback, because it is confidently wrong rather than merely stale.

## Verification

| | result |
|---|---|
| suite | **3 passed** |
| mutation (restore the literal) | **1 failed \| 2 passed** — the
renamed-WIP case only |
| dashboard `tsc -p tsconfig.app.json` | **0 errors** |
| census `--strict` | exit 0, **unchanged** — this class is invisible to
it |

The test drives the **production path** (`fetchBoardWorkflows` →
`resolveTaskWorkflowMetadata` → `currentColumnFlags`) rather than
injecting flags as props, so it covers the producer as well as the
consumer. `building` and `shipped` collide with no legacy id, so a
surviving `.has(column)` cannot pass by luck; the `todo` control pins
that the default vocabulary is unaffected, and the renamed-COMPLETE
negative pins that the fix does not hand editability to a finished card.

Note: `tsconfig.test-check.json` fails on `main` as well — pre-existing,
and **zero** of its errors come from this branch's files.

## Suggested follow-up

The remaining 17 collection sites deserve the same pass, and the scan
that found this should probably become a gate — a census that counts
comparisons will keep reporting zero while this class quietly grows. I
have not built that here because the existing gates already need
`#3136`'s attention first, and adding a sixth advisory check that nobody
blocks on would repeat the pattern this session keeps running into.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:34 -07:00
gsxdsm
d6079970e8 fix(self-healing): 18 recovery rebounds hardcoded todo and THREW on a renamed board (#3150, first slice) (#3152)
First slice of #3150. `self-healing.ts` held **26** `moveTask` calls
with a legacy literal target; this converts the **18 `todo` rebounds**.

## Why this is worse than a guard, and documented already

`task-store/moves.ts` records it from a previous incident:

> `moveTaskInternal` **REJECTS** a target the workflow does not declare
(`TransitionRejectionError: unknown-column`) … completion handoff did
not silently no-op — it **THREW**.

Every one of these 18 is a **recovery**. On a renamed board they threw
instead of rebounding, so the strand each sweep exists to clear survived
*and* the sweep reported failure. The reliability layer meant to be the
backstop was the layer that broke.

## Why the census never saw it

It counts **comparisons** against legacy ids. A move target is an
**argument**. That is the third blind spot of the same instrument, and
all three have now produced real defects found by hand:

| blind spot | found this session |
|---|---|
| definitions | `GITHUB_TRACKING_EDITABLE_COLUMNS` — tracking
unreachable on renamed boards (#3149) |
| collections | swept: 30 sites, 29 already correct, 1 defect (the
above) |
| **targets** | **this** — 26 in one file, 31 tree-wide |

## Why 18 sites at once is safe

`resolveReboundTargetForTask` **degrades to `"todo"`** when no workflow
resolves, and `self-healing.ts` already used it at line 745. On every
board we ship, the resolved answer *is* `todo` — so default behaviour is
unchanged **by construction**, not by inspection. The control case pins
exactly that, and it is the reason this can land as one change rather
than eighteen.

## Scope, and what I deliberately did not touch

Converted: the 18 `todo` rebounds.

**Not** converted: the `done`, `archived` and `in-review` targets. They
need different helpers and genuine reasoning about which lane a
completion or an archive belongs in — converting them by analogy is
exactly the half-conversion this program keeps paying for. Sites with no
resolver in scope are unchanged.

The audit behind the split is in the commit: of 26 sites, 5 had resolved
lanes in scope, 4 had an IR, 17 had nothing — and `lanesOfReclaim`
returns **Sets**, which is the wrong arity for a target (a move takes
exactly one column, per the `moves.ts` note).

## Verification

| | result |
|---|---|
| engine `tsc` | **0 errors** |
| **all 43 self-healing suites** | **843 passed** |
| census `--strict` | exit 0, **unchanged** — invisible to it |
| `check-inert-sync-lanes` | exit 0 |
| differential | restoring the literal → **1 failed \| 1 passed**,
renamed case only |

The new test drives a **public entry point**
(`reconcileInReviewUnmetDependencies`, the FN-6793 contract) rather than
calling the helper directly, so it covers the producer path too.

One harness note worth keeping: the first version of the test failed
**upstream** of the target, because the sweep selects rows via
`resolveProjectColumnsForRoles` — a *project-level* resolver reading
`listWorkflowDefinitions`, not the task's own selection. Without that
mocked, the renamed card was never considered and the failure looked
like the fix not working. That distinction (project-level vocabulary vs
per-task IR) will bite the next slices too.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks now move to workflow-specific rebound, completion, and archive
columns instead of fixed default destinations.
* Retrying and recovering tasks works correctly on boards with renamed
lifecycle columns.
* Added safe fallback behavior for workflows without custom lifecycle
settings.
* **Tests**
* Added coverage to prevent legacy hardcoded task destinations and
verify renamed-column recovery scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:22 -07:00
gsxdsm
4c62124589 fix(core): the last four archived LANE sites — the ones that already held a store (#3163)
Completes the six **LANE** sites from the triage (#3154). #3160 and
#3162 did the two predicate builders that needed threading; these four
already had `store` in scope, so each is a resolve-and-spread at the
site.

## What each fixes on a renamed board

| site | defect |
|---|---|
| `store.ts` revert lookup | a done/archived prior undo attempt kept
surfacing as an **open** undo task — the store-side twin of the
dashboard defect fixed in #3129 |
| `branch-group-ops.ts` | the near-duplicate marker cleanup found **no
live rows at all**, so stale markers survived. Its own header says stale
markers alter operator decisions |
| `branch-and-pr-entities:438` | the CREATE-time fingerprint duplicate
guard kept archived cards in the candidate set — a new task could be
refused as a duplicate of one already filed away |
| `branch-and-pr-entities:470` | recent-sibling lookup counted finished
siblings as candidates |

## The parity gate caught my first version — and it was right

I collapsed the fingerprint fallback into a **string array**
(`["archived"]`) and pushed `ne(col, lane)` in a loop. Behaviourally
identical, and it **dropped the Drizzle encoding's literal count**,
because the gate scans for the `ne(..., "archived")` *expression shape*.
TS and raw held steady, so the encodings diverged — precisely what that
gate exists to catch, catching it.

The fix: keep every fallback as a literal `ne(..., "archived")`
**expression** rather than data.

That is what makes these conversions **additive** — the resolved path is
added, the literal stays, no encoding's count moves, and an unconverted
board builds byte-identical SQL. Same property as #3160/#3162, now with
a demonstrated failure mode for getting it wrong. Worth knowing for
whoever does the remaining TS remainder: *behaviourally identical* is
not sufficient; the shape has to survive too.

## Measured

- Parity test **2/2**, inventories unmoved — the point.
- archived / branch / near-duplicate / merge-blocker suites — **6 files
/ 46 tests pass**.
- `tsc --noEmit -p packages/core` clean; census `--strict`,
`check-sql-column-literals`, `check-fnxc-future-dates` clean.

## Census

**Unchanged** — literals remain as fallback arms, by design.

## Where the cluster stands

All **six LANE** Drizzle sites are now converted (#3160, #3162, this).
The **two STATE** sites are marked in place and must never be converted
(#3157). What remains is the small TS remainder that is neither a
fallback arm nor a sentinel, identified in #3156.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:24:07 -07:00
gsxdsm
e1a01cb252 docs(dashboard): the research-modal archive filter is prop threading, not a data-fetch change (#3134)
Correcting the **shape** of the last census entry in this file. The
existing note sizes it as needing new data. It does not, and the
difference changes who can pick it up and for how much.

## What the note gets right

A per-**task** map (`columnFlagsByTaskId`) cannot help here. It is built
from board-resident rows, and the rows this filter cares about are
**archived** ones — exactly what a board map omits. Threading that map
would look converted and leave the case it exists for unresolved. That
reasoning is correct and I kept it.

## What it misses

This guard does not ask a per-task question.

*"Is `task.column` an archive lane"* is a question about a **column**,
and the answer lives in the workflow definition — a lane exists there
whether or not any row currently sits in it. **Archived rows being
absent from the board is irrelevant to a column-keyed answer.**

That map already exists on the board:

- `ListView.tsx:756` derives `columnFlagsById` (`ColumnId -> flags`)
from its workflow columns
- `useExecutorStats` takes the same shape

## So the real cost

`MainContent → ResearchView → this modal`, plus sourcing the column map
where MainContent renders ResearchView (it holds none today).

Three layers for one guard is a real cost and a fair thing to decline —
but it is a **different decision** from *"needs new data"*, and the two
have very different prices. The note as written would have the next
reader believe an API change is required.

## Left counted and unconverted, deliberately

A three-component prop chain wants to be someone's considered change,
not a drive-by on the last entry in a file — especially from me, at the
end of a long session where two rushed changes already went wrong.
Recording the corrected shape is the part that was cheap and wrong to
leave.

## Measured

- Comment-only.
- `ResearchView.test.tsx` — **27 tests pass**.
- `tsc --noEmit -p tsconfig.app.json` clean; census `--strict`,
`check-fnxc-future-dates` clean.

## Census

**No movement.** The entry stays, with an accurate price on it.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 10:18:41 -07:00