Last loose thread from the 2026-07-31 stamp-repointing wave. Two comment
lines.
```
line 1027 ConcurrencyAdmission 2026-07-31-09:00 -> 2026-07-21-22:30 (eef5eb751e)
line 1471 WorkflowLifecycleColumns 2026-07-31-05:00 -> 2026-07-30-20:55 (109204c590)
```
## Not a revert — neither value was ever right
| stamp | originally | after #3280 | authoring commit (UTC) |
|---|---|---|---|
| `ConcurrencyAdmission` | `2026-08-06-09:00` (16 days ahead) |
`2026-07-31-09:00` (10 days late) | **2026-07-21 22:30** |
| `WorkflowLifecycleColumns` | `2026-08-01-05:00` (~1.5 days ahead) |
`2026-07-31-05:00` (~1 day late) | **2026-07-30 20:55** |
Both were written **ahead of their own commits** to begin with. Three
lanes then repointed stamps to turn `main` green (#3261, #3269, #3280),
moving the **date** back a day while keeping the clock time — which
converts an hours-off stamp into a days-off one in the opposite
direction. #3282 reverted the batch it owned; these two were outside its
scope.
So restoring the originals would be wrong too. The defensible values are
the authoring commits' UTC timestamps, per the `date -u` rule #3281
settled.
## Why now
`scheduler.ts` is a hot file. This survived two successive claimants — I
flagged it on #3262 and again on #3288 rather than opening a conflicting
PR, and said I'd take it once the file was unclaimed.
`check-file-claimed` now reports UNCLAIMED, so here it is.
## Scope
The gate is **green either way** — #3277 fixed the comparison, so
nothing is blocked by this. It is purely about the FNXC trail recording
when the work actually happened, which is the only reason the trail
exists. A stamp that satisfies a check while misstating the date by ten
days is worse than no stamp.
## Verification
```
check-fnxc-future-dates green
check-inert-sync-lane-conversions green
check-lane-wiring green
check-sql-column-literals green
census --strict green
```
Diff is two comment lines — no executable change. No changeset.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## What
Pins the **worktree-capacity arithmetic** — the gap #3262 measured,
named, and explicitly left for someone to claim. Two commits: a
behaviour-preserving seam extraction, then the test.
#3262's own scope note:
> blinding this predicate to `false` leaves all 22 scheduler suites
green (365 tests). The capacity logic it feeds has no behavioural
coverage at all.
## Two live defects, opposite directions
Both came out of these few lines:
- **UNDER-COUNT admits work over the cap.** `maxWorktrees=4`, four
planning sessions each holding a worktree, and a replan dispatch
admitted as the **fifth** — the ledger counted WIP cards only and never
learned to count planners.
- **OVER-COUNT self-deadlocks.** A planned Ready card *retains* its
planning worktree for execution reuse, so counting it as a holder blocks
its own release: `2 wip + 3 idle-held = 5/4`, and the first unpause
released 2 of 4 slots' worth of work.
**Both are pinned, and the asymmetry is why.** Under-counting breaks the
cap and lets real work over it; over-counting only starves dispatch. A
test covering the "safe" direction alone would leave the expensive one
open.
## Mutation-tested — all four caught
| mutation | result |
|---|---|
| drop the terminal exclusion | 1 failed / 7 passed |
| count WIP cards twice | 2 failed / 6 passed |
| count cards holding no worktree | 1 failed / 7 passed |
| drop the self-slot subtraction | 1 failed / 7 passed |
```
clean: 8 passed (8)
scheduler suite: 16 files / 151 tests passed (behaviour preserved by the extraction)
typecheck, lint: clean
```
## Scope, stated rather than implied
The terminal predicate is **injected**, not resolved here. Which lanes
are terminal is #3262's test; resolving it in this file would make it
fail for that reason instead of this one. This pins the **set
arithmetic** — who is excluded, and how the total is formed.
Still not covered, and I am not claiming otherwise: the *stateful* half
of the ledger — the `+= 1` on dispatch and the `Math.max(0, … - 1)` on
failure inside `schedule()`'s loop. Extracting that would mean
restructuring dispatch itself, which is a different change from this
one.
## Process note
I claimed this on #3262 **before** starting rather than after, because
`scheduler.ts` is the hottest file in the tree and I produced three
duplicate PRs earlier tonight by picking up small shared-surface work
someone else already had in flight. Announcing first cost one comment;
the duplicates cost three PRs and two closes.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved scheduler capacity calculations for tasks that retain
existing worktrees.
* Prevented WIP tasks from being counted twice.
* Excluded completed and worktree-less tasks from reserved capacity.
* Corrected candidate capacity calculations when no worktree capacity is
reserved.
* **Tests**
* Added coverage for worktree reservation totals and candidate reuse
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
#3286 fixed a real user-facing regression and **shipped no test**, so
nothing stops it returning. The regression was mine.
## The bug
#3215 (mine) added `isArchivedColumn` to an effect's dependency list to
keep the task fetch honest. That effect **also owned four `setState`
calls**, and `isArchivedColumn` is a `useMemo` over
`useBoardWorkflows()` — which revalidates asynchronously.
Every revalidation re-ran the reset over whatever the operator had
typed. A title entered before the workflows settled silently reverted to
`Research: <heading>`, and the task was created with a title nobody
wrote.
## Why my own four tests could not see it
Every existing case in this file asserts the **filtered task list** —
render, await `fetchTasks`, read the datalist. **None types into the
form.**
I tested what I added and not what I touched. That is why the regression
belongs in this file rather than a new one: the gap is this file's.
## The case
Renders with `boardWorkflows: null` — the state when an operator opens
the modal and starts typing — types a title with per-character
`userEvent`, then rerenders with a resolved workflow set (a **new object
identity**, which is the entire mechanism) and asserts the typed text
survived.
Two details that each cost a cycle, recorded at the site:
- **`fetchTasks` is not awaited.** It runs only in enrich mode, while
the title field exists only in create mode — so the reset effect, not
the fetch, is under test. My first version waited on it and failed for
the wrong reason.
- **`userEvent.type`, not `fireEvent.change`.** The documented failure
is state overwritten between renders; a single synthetic change event
can land after the reset and mask it.
## Measured both directions, on main `6834ba35bd`
| state | result |
|---|---|
| fixed main | **5 passed** |
| dependency re-added to the reset effect (my bug) | **1 failed / 4
passed** — and only that case |
The second row is the point: it fails on precisely the mutation that
recreates the defect, and leaves the four archived-lane cases green — so
it pins the regression without duplicating what is already covered.
`eslint` clean, `check-fnxc-future-dates` 0. Test-only.
## Note on provenance
Getting this measurement took three attempts: a `git checkout` of the PR
branch silently failed (stderr suppressed), so I twice ran against the
wrong tree and nearly concluded the fix did not work. HEAD and
dirty-count are printed beside every number above for that reason.
Found by chasing the deterministic half of #3264 (dashboard red on
`main`). **The tests were right; the product is broken.**
## The bug
Open **Create Task** from a research finding, type a title before the
board workflows settle, and the field silently reverts to the derived
default `Research: <heading>`. The task is then created with a title the
operator did not write. `description`, `priority` and `taskId` reset the
same way.
`ResearchTaskActionModal` reset those four fields in the same effect
that fetched the task list, and that effect's dependency list carried
`isArchivedColumn`:
```ts
const isArchivedColumn = useMemo(() => { … }, [boardWorkflows]); // useBoardWorkflows() — async
useEffect(() => {
setTitle(`Research: ${finding.heading || run.title}`); // ← re-runs on every revalidation
…
}, [open, mode, projectId, finding.heading, preview, run.title, isArchivedColumn]);
```
`useBoardWorkflows` resolves and revalidates asynchronously, so the
memo's identity changes and the reset re-runs over whatever the operator
has typed.
Introduced by #3215, which correctly added the archived-column filter
but hung its dependency on an effect that also owns form state. Same
class as the documented
`docs/solutions/ui-bugs/skill-autocomplete-highlight-reset-on-swr-revalidation.md`.
## The fix
Split into two effects: the reset depends only on what it derives from;
the fetch keeps `isArchivedColumn`. No behaviour change to the archived
filter — #3215's guard is untouched.
## Verification, both directions
The three standing `ResearchView` tests fail without this and pass with
it:
```
isArchivedColumn back on the reset effect: 3 failed | 24 passed (27)
as committed: 27 passed (27)
```
## What I tried and removed, because it matters
I wrote a dedicated invariant test (per "fix the invariant, not the
repro") asserting that *all* typed fields survive a revalidation. **I
deleted it, because it did not work.**
- First draft used `mockImplementationOnce` to defer
`fetchBoardWorkflows`. `ResearchView` resolves board workflows on mount,
so that once-implementation was consumed before the modal opened.
Reverting the product fix left the test **green** — it proved nothing.
- Second draft deferred *every* call. `beforeEach` uses
`vi.clearAllMocks()`, which clears calls but **not implementations**, so
the deferral leaked into later tests and left `fetchBoardWorkflows`
permanently pending — masking two of the three genuine failures. The
revert then showed `1 failed` instead of `3`, i.e. my test was hiding
real bugs.
Rather than ship a regression test that cannot regress, I removed it.
The three existing tests already fail without the fix, which is real
coverage; a broader invariant test needs a modal-level harness that
resets implementations between cases, and that is worth doing properly
rather than badly here.
## Scope
Also in #3264: `TaskCard.badge-wrap` (1 deterministic failure, unrelated
— CSS/layout), and `useChat` / `WorkflowNodeEditor` /
`PlanningModeModal`, which pass standalone and are cross-file
contamination, not product bugs. Untouched here; the issue has the
per-file matrix.
No changeset — `@fusion/dashboard` is private.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Preserved form entries during workflow revalidation in the research
task modal.
* Limited task selection to active workflow columns when enriching
findings.
* Prevented outdated task results from replacing newer selections.
* Improved loading and task-list behavior when source findings or modal
state changes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**Rebased. The census claim in my original title was overtaken — this is
now a correction, not a conversion.**
## Census: 0 before, 0 after
#3261 got there first, by **recording** the fallback rather than
converting it. Its `DELIBERATE` reasoning is correct and I kept it
verbatim.
## What this corrects
That note says:
> Recorded rather than converted because **there is nothing to convert
TO**.
There is. **`isTerminalColumnRole` in core is this predicate, term for
term** — verified against `column-roles.ts` rather than assumed:
| | hand-rolled | `isTerminalColumnRole` |
|---|---|---|
| flags present | `flags.complete === true \|\| flags.archived === true`
| same, via the two role helpers |
| flags undefined | `columnId === "done" \|\| columnId === "archived"` |
same, via `LEGACY_COMPLETE/ARCHIVED_COLUMN_ID` |
And the helper's own doc names this exact case — it exists *"because the
pattern `column !== \"done\" && column !== \"archived\"` is the single
most repeated shape in the backlog"* and *"keeps callers from
re-deriving it and from accidentally dropping one half."*
**The rest of #3261's argument stands and is preserved.** The
undefined-flags arm is a **live** path, and treating an unreadable
workflow as non-terminal would count a finished card's retained worktree
against live capacity. That reasoning is about the *fallback's
existence*, not about *where the predicate lives* — and the shared
helper carries the identical fallback.
Second time in this file: `isWipColumnTask` two lines up records that it
was itself once *"a hand-rolled copy of `isWipColumnRole`"*. That's an
argument for the helper being easy to miss, not for anyone being
careless.
## Coverage, stated rather than implied
**Blinding this predicate to `false` leaves all 22 scheduler suites
green (365 tests)** — the capacity logic it feeds has no behavioural
coverage at all.
The added test pins the **lane vocabulary** (both renamed terminal
lanes, the non-terminal lanes, the legacy fallback). It does **not** pin
the capacity arithmetic, which stays unguarded and belongs to that
gate's owner. Under-counting is the dangerous direction: the commit
adding the gate reports `maxWorktrees=4` with **a fifth worktree
admitted**.
## Verification
- 23 scheduler suites — **368 green**; `tsc` clean
- Census 0 → 0; DELIBERATE count unchanged at 148; inert ratchet green
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#3279. Undoes the damage my #3261 did, now that #3277 has landed
and made it safe.
## What went wrong
#3277 established that last night's "future-dated" stamps were
**correct** — the author's local date in a UTC+1 container, written
minutes before their commits. The gate compared against the runner's
local calendar (PDT) and called them tomorrow.
I diagnosed it as author error and repointed seven stamps to turn main
green. The values I wrote were **neither the author's local time nor
UTC** — invented times chosen to satisfy a broken check. The FNXC record
is this project's why-does-this-exist trail, so those stamps misstated
when the work happened.
## Restored verbatim
| file | mine (wrong) | restored |
|---|---|---|
| `workflow-column-boundary-capacity.test.ts` | `22:30` |
`2026-08-01-00:30` |
| `runtimes/in-process-runtime.ts` | `22:20` | `2026-08-01-00:20` |
| `scheduler.ts` (`MissionReconciliation`) | `22:00` |
`2026-08-01-00:00` |
| `workflow-column-boundary-hooks.ts` | `22:20` | `2026-08-01-00:20` |
| `workflow-column-boundary.ts` ×2 | `22:20` | `2026-08-01-00:20` |
| `workflow-graph-task-runner.ts` | `22:20` | `2026-08-01-00:20` |
## The check that mattered
Sequencing was deliberate — #3277 had to land first or this would have
re-reddened main. The real question is whether the gate now accepts the
**originals**, measured across the rollover boundary at local
`2026-07-31 17:23 PDT` / UTC `2026-08-01 00:23`:
```
America/Los_Angeles exit 0 Europe/Paris exit 0
UTC exit 0 Asia/Tokyo exit 0
```
`123 known future-dated stamp(s), none added`. **No baseline change
needed** — #3278's pruning already re-recorded `scheduler.ts`, and these
are known stamps rather than new ones.
Stamps only: `git diff` shows **zero** non-FNXC lines, 7 insertions / 7
deletions across 6 files. `census --strict` 0, `pnpm test:gate` 0.
## The part worth keeping
I argued against exactly this on #3263 — *"it rewrites stamps whose
authors are not us"* — and then did it myself six lines later, because I
was confident about a cause I had not checked. The commits' timestamps
were available the entire time; I read the runner's clock and never
asked what timezone the **author** was in.
Four of last night's seven PRs were fixing something that was not
broken. This is the cleanup for my share of that.
**`check-fnxc-future-dates` exits 1 on `origin/main`.**
```
packages/engine/src/scheduler.ts: 3 future-dated stamps, baseline allows 2
FNXC:ConcurrencyAdmission 2026-08-06-09:00 (six days out)
FNXC:WorkflowLifecycleColumns 2026-08-01-05:00
FNXC:WorkflowScheduling 2026-08-01-01:05
```
All three repointed to `2026-07-31`, times preserved. Gate now exits 0.
## This red has outlived three owners
#3270, #3272 and #3274 were each opened against it and each **closed
without merging**. Main has been red on this gate for hours while three
fixes came and went.
Claimed with `check-file-claimed.mjs` before starting — only #3262
touches `scheduler.ts`, and it is a terminal-role refactor rather than a
stamp fix, so this was genuinely unowned.
## Why this keeps recurring
Seven incidents in roughly two hours. The mechanism, in one line: **the
date check runs only in CI** (`pr-checks.yml:66`, no pre-commit or
pre-push hook), so every PR is validated against main's baseline *at its
own CI time* and cannot see a concurrent or later change. Two PRs
stamping the same file both pass, then compose into a red main. One case
(#3273) was a stale branch **reverting** an already-merged fix.
Patching instances has not converged — this PR is the eighth attempt at
the same class. Two structural options, neither of which I am landing
unilaterally since the second changes the gate's contract:
- run the date check at **author time** (pre-push); it needs no baseline
for "is this date in the future", so it cannot be raced
- make the date rule **baseline-free** — a future-dated stamp is always
wrong, unlike a lifecycle literal that may be a deliberate fallback
`2026-08-06` being six days out also suggests these are not off-by-one
timezone slips but stamps written from an intended future date.
## Verification
- `check-fnxc-future-dates` — **exit 0** (was exit 1 on main)
- `scheduler` suites — **148 pass**
- `tsc --noEmit` (engine) — 0 errors
- comment-only diff, no behaviour change
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lifecycle-column-census test file spawned the census CLI ~14 times, each
parsing every tracked source file to a TypeScript AST (~2s for ~1960 files).
Many spawns were byte-identical, deterministic, read-only real-repo scans:
the --json census 4x, the plain report 2x, plus a repeated identical
--update-baseline tree-sync across the ratchet cases. Memoize each distinct
read-only spawn's output (keyed by argv) and reuse the synced baseline JSON,
collapsing duplicate full-AST scans without changing any assertion.
File wall-time: 30.5s -> 18.6s (-39%). 57/57 tests still pass.
Fusion-Task-Id: FN-slow-test-census
**`check-fnxc-future-dates` exits 1 on `origin/main`.** This is the last
stamp causing it.
```
packages/core/src/task-store/lifecycle-ops.ts: 1 future-dated FNXC stamp, baseline allows 0
FNXC:Diagnostics 2026-08-01-00:50 (today is 2026-07-31)
```
Corrected to `2026-07-31-00:50`. One character.
`check-fnxc-future-dates` now exits 0; `tsc --noEmit` clean.
## Why this was left behind
Four PRs converged on this red main — #3262, #3263, #3265, and my own
#3266 (closed as superseded). Between them they covered the census rise
and the boundary-work stamps. **None touched `lifecycle-ops.ts`**, so
the gate stayed red after the others landed.
That is the predictable failure of parallel work on one symptom:
everyone fixes the part they saw first, and the residue survives because
each author checked "is main green now?" against their own branch rather
than against main.
## I claimed before working this time
```
node scripts/check-file-claimed.mjs packages/core/src/task-store/lifecycle-ops.ts
→ UNCLAIMED
```
Then pushed the branch before editing. I did the opposite on #3266 —
built it, then discovered #3265 already covered it — which was the sixth
duplication of the phase and my third. The tool answers in one command;
the discipline is running it *first*.
## Verification
- `check-fnxc-future-dates` — **exit 0** (was exit 1 on main)
- `census --strict` — exit 0 (already green; #3265's marker landed)
- `tsc --noEmit` (core) — 0 errors
- one-character diff, no behaviour change
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## main is red, and this is the second half of it
Running the full `engine-default` project on `origin/main` — 754 files,
10,538 tests — returns **3 files / 4 tests failing**. One is the
parked-seam audit counter, fixed in #3258. The other three are here.
All three assert that lifecycle debt **still exists**. It does not:
| assertion | expected | actual |
| --- | --- | --- |
| `finds >10 literal move targets, so the census is not vacuous` | > 10
| **0** |
| `both recoveryRehome groups are non-empty` | > 0 each | **0 / 0** |
| `still says ROSE when guards genuinely grew` | exit 1 | **exit 0**
(mutation was a no-op) |
Nothing regressed. The conversion program drove the engine's literal
move-target population to zero and the census baseline to zero entries.
**Each guard's premise was "the debt still exists", so each expired the
moment the work succeeded — and expired by failing, which reads as a
regression in the very thing it was guarding.**
The third is the sharpest: it manufactured a rise by finding a baseline
entry with more than one guard and zeroing it. With `byFile` empty there
was nothing to find, so it mutated nothing, the census correctly passed,
and the test asserted exit 1 against a correct pass.
## Repointed, not deleted
A vacuity guard must not depend on real debt existing. Two halves:
- **Vacuity now runs against a synthetic fixture** — a small in-memory
source with three `moveTask` literals (one with `recoveryRehome`, one
targeting an undeclared column). The collector is exercised forever
regardless of how much real debt remains. This is what keeps the rest
honest: at a real population of 0 the zero-assertions are trivially
true, and **only the fixture proves they would still fire**.
- **The real-tree assertions now assert zero**, so a reintroduced
literal move target fails them. Same guarantee as before, pointed at the
state the tree is actually in.
The ROSE case builds its own one-file tree through the
`FUSION_CENSUS_FILE_ROOT`/`FILE_LIST` seam from #3230 rather than
borrowing a baseline entry that no longer exists — a rise
**constructed** instead of borrowed. It also now asserts the failure is
*not* the reclassification wording, which is the distinction that file
exists to protect.
## Verification
| mutation | expected | result |
| --- | --- | --- |
| reintroduce a literal `moveTask(id, "in-review")` in the engine | fail
| exit 1 ✅ |
| blind the collector to return `[]` | fail | exit 1 ✅ |
| remove the `ROSE` wording from the census script | fail | exit 1 ✅ |
| clean tree | pass | exit 0 ✅ |
60 tests green across the three census suites. All eight ratchets exit
0. Test-only; no changeset.
**One honest note on my own method.** My first `ROSE` mutation replaced
1 of the 2 occurrences in the script and the test stayed green — which
looks exactly like a dead assertion. It was an ineffective mutation, not
a dead test; the manual run still printed `ROSE` from the other
occurrence. Re-run against both, it failed. A mutation that does not
actually change behaviour proves nothing, and it is worth checking that
the mutation landed before concluding the test is dead — the same trap
as reading a report-only ratchet's exit 0 as a pass.
## Not in scope
`defaultColumnIds()` has a pre-existing type error (`Property 'columns'
does not exist on WorkflowIrV1` — union narrowing). Untouched by this PR
and the test executes fine; flagging rather than fixing, since it is
unrelated to the red.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved census analysis reliability when evaluating guard-count
increases and reclassification messages.
* Improved detection and classification of literal move targets and
declared columns.
* Enhanced file path reporting for inputs outside the primary source
directory.
* **Tests**
* Added isolated test scenarios using temporary census data and
fixtures.
* Strengthened validation of recovery, plain, and declared-column
classifications.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
`9094d1640e` (globalPause gates every graph node entry) reddened **two**
lifecycle gates on main. Both are fixed here, in separate commits.
## 1. The census ratchet went 0 → 2
`isTerminalColumnTask` in `scheduler.ts`:
```ts
const flags = columnFlagsForTask(task);
if (flags) return flags.complete === true || flags.archived === true;
return task.column === "done" || task.column === "archived"; // ← counted
```
**The code is correct.** It resolves traits first and falls back only
when the workflow is unreadable. The census counts fallback literals on
purpose — *"a fallback literal is still a literal and should go when the
trait path becomes unconditional"* — and reports them beside the backlog
as already-converted. Its own remedy for a legitimate one is a
`DELIBERATE-LITERAL` marker at the site.
Recorded rather than converted because **there is nothing to convert
to**: a task whose workflow cannot be read has no resolved lane, and
treating it as non-terminal would count a finished card's retained
worktree against live capacity — the opposite of what the surrounding
fix does.
Marker sits in the declaration's **leading** comments; an inline one
attaches to the wrong node and is silently ignored, which cost a
miscount once before. Baseline re-recorded in the same commit, since the
census tracks deliberate counts and reports a marker addition as
`RECLASSIFIED`.
## 2. The stamp gate was red as well
Six files stamped `2026-08-01-00:2x` while UTC was `2026-07-31`:
```
workflow-column-boundary.ts 2 workflow-graph-task-runner.ts 1
workflow-column-boundary-hooks.ts 1 in-process-runtime.ts 5 (allows 4)
workflow-column-boundary-capacity.test 1
```
This checkout is UTC-7, so "just after midnight local" is tomorrow in
UTC — the case AGENTS.md documents, which passes `pnpm lint` locally
*because* the local clock agrees with what was written. Second
occurrence today; I fixed the same shape on #3208 for another worker.
Repointed to `2026-07-31-22:2x`, preserving relative order. **Zero
non-comment lines changed** — 8 lines across 6 files, verified by
diffing out FNXC lines.
## Measured
| check | before | after |
|---|---|---|
| `census --strict` | **1** | **0** |
| backlog | **2** | **0** (DELIBERATE-LITERAL 148 → 150) |
| `check-fnxc-future-dates` | **1** | **0** |
| `pnpm test:gate` | 0 | 0 |
| `census-reclassification-message` | 2 failed | **1 failed** |
That last row is deliberate: the remaining failure is the
expired-premise case #3260 fixes, and I have not touched it. The
capacity test from `9094d1640e` still passes 9/9.
## Why this landed at all
Both gates run in `pr-checks.yml`, so a PR carrying either would have
gone red. Worth someone checking how it merged — a stale merge base
would explain it, and if so the same hole is open for the next merge.
Follow-through on the recommendation I made reviewing #3256: **a gate
needs a test for its file discovery, not only its matcher.**
## The gap
This suite pinned the matcher and never the scan. Every case either
feeds the classifier a source string or drives the CLI against the real
tree — so **the file list could return empty and all 53 tests would
still pass.**
Not hypothetical. `git ls-files` lists tracked files only, so a new file
with a plain `task.column === "in-review"` scored 0 until staged
(#3254). The identical bug then turned up in the move-target ratchet
**behind its own 12 matcher tests** (#3256) — I wrote those 12
specifically to stop that gate regressing, and they could not see it,
because they import the matcher and never run a scan.
## Four cases, on a synthetic tree
Driven through `FUSION_CENSUS_FILE_ROOT` + `FUSION_CENSUS_FILE_LIST`, so
discovery is testable without creating files inside a checkout the
operator writes to concurrently:
- a guard in a scanned file **reaches the classifier** and is counted
- `--strict` fails **for the right reason** (message names the file; not
an ENOENT fail-closed)
- **every** listed file is counted, not just the first
- files are read from the **scan root**, so a listed path and a read
path cannot diverge
## The second case earns its wording
Its first version asserted only `code === 1` — and **passed while
discovery was broken.** With the injected list ignored, paths come from
the real repo while reads resolve against the fixture root, every read
misses, and the gate fails closed with exit 1. Right code, unrelated
cause.
A test that cannot tell *"found a guard"* from *"could not read
anything"* is not testing the ratchet. Asserting the message is what
separates them.
I found that only by checking which cases the control actually failed —
3 of 4, not 4 of 4. Had I stopped at "the control fails, ship it", I
would have added a test that passes for the wrong reason to a suite
whose whole purpose is catching tests that pass for the wrong reason.
## Measured
| check | result |
|---|---|
| suite | **57 passed** (53 + 4) |
| anti-vacuity: `injectedList` forced undefined | **all 4 fail** (3/4
before strengthening case 2) |
| restored | 57/57 |
| `census --strict` / `check-fnxc-future-dates` | 0 / 0 |
Tests only — no gate or product change. The same four assertions port
directly to the other lifecycle gates once each grows the fixture seam;
the move-target ratchet is the obvious next one, and its `.mjs` is
currently claimed by #3256.
## main is red
`workflow-optional-role-param-caller-audit-live-e2e.pg.test.ts` fails on
`origin/main` in `engine-default`:
```
AssertionError: expected 4 to be 2
expect(parkedConverted.length).toBe(2);
```
Found by running the whole live-E2E corpus rather than trusting that it
passes — 27 files, 159 tests, 1 red. Not in the merge gate, so it has
been sitting there. The alarm fired **downward**, exactly as that file
was written to: two more call sites started passing `parkedColumns`, and
a counter fails when someone *closes* a gap as well as when someone
widens it.
## Why I did not just write 4
Re-recording it at 4 would have laundered an inert conversion through
the audit written to catch inert conversions. Walking all six sites
before touching the number:
| site | `parkedColumns` provenance | verdict |
| --- | --- | --- |
| `agent-heartbeat.ts:1267` | — | unconverted |
| `agent-heartbeat.ts:3796` | — | unconverted |
| `self-healing.ts:13184` | `await resolveProjectColumnsForRoles`
(:13170) | async-resolved |
| `self-healing.ts:13294` | `await resolveProjectColumnsForRoles`
(:13293) | async-resolved |
| `task-agent-sync.ts:243` | `await resolveLinkSyncColumnRoles` (:225) |
async-resolved |
| `scheduler.ts:1798` | `resolveTaskParkedColumnsSync` (:1797) | **SYNC
— INERT** |
`scheduler.ts:1798` resolves through `resolveTaskParkedColumnsSync` →
`getTaskWorkflowSelectionImpl`, which is `undefined` for every task
under PostgreSQL. The resolver then takes its `!workflowId` branch and
returns the **default builtin IR** — not `undefined` falling through to
a legacy arm, but a real IR resolving real traits, with full confidence.
It answers `hold`/`intake` as `todo`/`triage` on every board, exactly as
the literal did. Driven proof:
`workflow-scheduler-sync-role-conversion-inert-live-e2e.pg.test.ts`.
**The shape count (4) and the live count (3) are different numbers, and
only the second is about behaviour.** Both are now asserted, plus the
sync-resolved site by name so it cannot quietly become "just one of the
four".
## A mistake worth recording, because the test caught it
My first draft keyed on `await` appearing inside the call window. The
argument is nearly always a variable (`[...driftedParkedColumns]`,
`roles.parked`) and the `await` lives in that variable's **assignment**,
several lines above. That draft classified all four sites as inert — and
it **would have passed** had I written the expected number to match what
it measured. It failed only because I asserted 2 live from reading the
source first, and the mismatch exposed the detector.
Classification is now by provenance: take the root identifier, find
where the file assigns it, ask whether *that* is awaited. The same bug
recurred in my named-site check and failed the same way.
That is the whole hazard of this program in miniature — a source-text
audit that measures nothing looks exactly like one that measures
everything, and it is the *number you expected* that catches it, not the
green.
## Not asserted, deliberately
`self-healing.ts:13184` is inert for an unrelated reason: its gate is
`hasFreshRun || hasActiveExecution` and never reads
`shouldPreserveParkedLink`, so its correctly-resolved set decides
nothing today. Its own FNXC note says so. Resolution path is
mechanically checkable; "the gate never reads the answer" is not, and
asserting it on a string match would produce a number nobody could
maintain. Recorded as prose in the header.
I also did not touch `scheduler.ts`. It is a real defect, not a
deferral, but converting it is not this file's job — it is named in the
test so the next person converting it is sent here to move it from the
inert list to the live one.
## Verification
Mutation-verified — a passing audit proves nothing until it has been
seen to fail:
| mutation | expected | result |
| --- | --- | --- |
| convert the scheduler site to the async resolver | fail (3/1 → 4/0) |
exit 1 ✅ |
| drop `parkedColumns` from a converted site | fail (shape 4 → 3) | exit
1 ✅ |
| add a new unconverted caller | fail (calls 6 → 7) | exit 1 ✅ |
| clean tree | pass | exit 0 ✅ |
Full 27-file live-E2E corpus green (159 tests). Ratchets exit 0.
Test-only; no changeset.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Expanded end-to-end audit coverage for workflows with optional role
parameters.
* Improved validation of caller resolution paths, including asynchronous
and synchronous scheduling scenarios.
* Updated expectations to reflect all supported conversion paths and
strengthened verification of scheduler behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
The structural half of #2784. I re-measured its 123 failures on current
main and **all four reported lanes are green** (143 / 2010 / 1961 / 5149
/ 2103 passing). This fixes the reason nobody saw them.
## The mechanism
`pnpm --filter @fusion/dashboard test` sets `stopScheduling = true` on
the first failing lane, so the rest never run — and there was **no flag
to ask for a full pass**. The report said:
```
[dashboard-quality] skipped 9 lane(s) after first failure
```
Nine lanes with **unknown** status and nine **passing** lanes produce
the same absence of failure text. That is how 123 failures accumulated
behind one red lane, and it is why the original issue could only be
written by running all twelve lanes by hand.
## What changes, and what deliberately does not
Fail-fast stays the **default** — fast feedback on a broken lane is
right, and changing it would slow everyone for a rare case.
- `--all` (alias `--no-fail-fast`) runs every lane and reports every
failure.
- `runQualityTests({ failFast })` so the behaviour is reachable from a
test, not just the CLI.
- The skip line now states the consequence and the remedy: lanes were
**NOT RUN**, status **UNKNOWN rather than passing**, and `--all` shows
the full set.
## Both halves pinned
A flag nobody can prove works is the same as no flag:
| test | asserts |
|---|---|
| DEFAULT stops after the first failing lane | `launched === ["one"]`,
`skipped: 2` |
| `failFast:false` runs all three | `launched ===
["one","two","three"]`, `failed === [one, three]` |
The second is the load-bearing one: **lane three ran even though lane
one had already failed**, and both failures are reported rather than
only the first.
**Anti-vacuity control:** reverting the `if (failFast)` plumbing fails
the second test and only it (`1 failed / 5 passed`); restoring passes
`6/6`.
## Scope
Runner and its tests only. No lane contents, no vitest configs, no CI
workflow — CI already invokes lanes individually, so this changes local
behaviour and the shared helper, not what CI runs.
eslint clean; `check-fnxc-future-dates` exit 0.
Suggest #2784 closes on the measured-green half and links here for the
structural half, so the mechanism does not close along with the symptom
that exposed it.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an option to run all quality-test lanes, even when earlier lanes
fail.
* Added `--all` and `--no-fail-fast` command-line options.
* Quality tests now stop on the first failure by default, with clearer
output for skipped lanes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Fixes the red `main`: `planning-browser-e2e.test.ts > places the sole
contextual comment trigger by viewport in embedded and modal Planning`.
**Unclaimed and not a flake.** `check-file-claimed` reported UNCLAIMED,
it is not in the quarantine ledger, it reproduced locally and
deterministically, and it failed identically across three consecutive
Full Suite runs. Quarantine would have been the wrong instrument — that
rule is for flakes, and appeasing a consistent failure buries a real
regression.
## Root cause
`PlanningModeModal.css` sizes dialog Planning as a **full-viewport
sheet**:
```css
.planning-modal:not(.planning-modal--embedded) {
height: 100dvh;
min-height: 100dvh; /* ← beats max-height: 100% */
max-height: 100%;
}
```
That was correct until the modal branch moved **inside
`FloatingWindow`** (`FNXC:ModalTouchGeometry 2026-07-26-14:10`). The
floating host's body is shorter than the viewport — it sits below a
title bar — so the rule now asks the sheet to be *taller than the box
containing it*. `min-height` wins over `max-height`, so the sheet cannot
shrink to its host and overflows.
Measured by walking the ancestor chain at 768×900:
```
BUTTON.btn top=928 h=36 ← 28px past the fold
DIV.planning-actions top=919 h=101
DIV.modal h=900 ← forced to full viewport height
DIV.floating-window__body h=763 sh=900 ← host is 763 tall, content is 900
DIV.floating-window h=765
```
The "Add comment to selection" control needed a scroll to reach —
exactly what the placement case exists to prevent.
## The fix
A scoped override under `.floating-window`, rather than editing the
sheet rule, so Planning rendered **outside** a floating host keeps its
full-viewport sizing:
```css
.floating-window .planning-modal:not(.planning-modal--embedded) {
height: 100%;
min-height: 0;
}
```
## Surface enumeration
Embedded Planning was **never affected** — it is excluded from the sheet
rule, and all four embedded viewports passed throughout. The failure was
modal-only, at every modal viewport (768, 769, 1024, 1280 — it fails
fast at the first).
**No new test.** The existing placement case already asserts this
invariant across **4 viewports × 2 presentations = 8 combinations**,
which is the surface enumeration for this affordance. It was red; it is
now green. Adding a narrower repro-only test would be the anti-pattern
the Fix-the-Invariant rule names.
## A disproven hypothesis, recorded
`min-height: 0` on `.planning-plan-review > .planning-plan-pane` — the
canonical flex-overflow fix, and a pattern used 10+ times in this very
file — **does not fix it**. Measured, not assumed. The overflow is one
level up, at the sheet/host boundary. Noted so the next reader does not
repeat the experiment.
## Verification
```
fix applied Tests 5 passed (5)
fix reverted Tests 1 failed | 4 passed (5) ← the test genuinely holds this fix
fix restored Tests 5 passed (5)
```
Neighbours green: **57 tests across 8 suites** (mobile
footer/bottom-space/pan-containment, terminal keyboard layout,
task-detail tablet width, mission planning modals mobile, mobile
planning input font size, task-detail floating geometry) plus **9**
planning e2e.
Changeset included (`patch`, category `fix`) — this is user-visible
dashboard behaviour.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## What
Pins the CLI board glyph's terminal-lane resolve — **the last flagged
site in the repo-wide resolver audit.**
Two commits: a behaviour-preserving extraction, then the test.
## I was wrong to flag this as unpinnable
In #3236 I recorded this site as not pinnable, reasoning that
*"extracting a pure helper and testing it would look like coverage and
would not be."*
That is true of a helper that **receives** the lane set — such a test
passes with the resolve blinded, which is exactly the `reads.ts` trap
the audit note records. It is **not** true of one that **resolves** it.
Building `resolveReliabilityLanes` in #3237 made the distinction
obvious: the seam has to contain the resolve, and then blinding fails a
test of it.
So the flag was too broad, and correcting it closes the site rather than
leaving a permanent excuse. That is the same failure mode I corrected in
someone else's note earlier today — a caution that hardens into a reason
not to look.
## Measured
```
converted: Tests 5 passed (5)
blinded: Tests 2 failed | 3 passed (5)
```
The two failures are the **renamed complete** and **renamed archive**
lanes. The three survivors are the default-vocabulary control, the
active-lane negative, and the degrade path — all of which should
survive.
```
task-list-board-columns + bin: 82 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```
## Why the sibling file did not cover it
`task-list-board-columns.test.ts` pins `boardColumnsForDisplay`, which
decides **which** lanes print. That function takes no lane set, so it
cannot fail when this resolve is blinded — and its own header says so
honestly. Two tests about the same command, one of which cannot see the
other's bug.
## What breaks without the conversion
On a board whose complete lane is `shipped`, a finished lane renders `●`
— the same glyph as active work. The board says work is in flight when
it shipped. Cosmetic next to the blank-board bug this area already
fixed, but wrong in the direction an operator reads at a glance.
## Also pinned
Two contracts the surrounding comments assert but nothing tested:
- **Cards come from the TASKS, not a resolved IR** — a card must never
depend on resolution succeeding to be *visible*. Asserted with an
unreadable workflow list.
- **A failed resolve degrades to the legacy pair**, with an unresolved
custom lane rendering as active — the documented fail-open direction.
Plus the paired negative: an ACTIVE lane keeps the active glyph under
both vocabularies, so widening the terminal set cannot mark the whole
board finished.
## Audit complete
Every `resolveProjectColumnsForRoles` call site in the repository —
`engine`, `core`, `dashboard`, `cli` — has now been blinded
individually, and every uncovered one is either pinned or has a recorded
reason it cannot be. Nothing is left flagged.
## What
Partial fix for **red main**. Test-only.
`ResearchView.test.tsx` has **4 failing tests on main**; 2 fail with:
```
No "fetchBoardWorkflows" export is defined on the "../../api" mock
```
The `vi.mock("../../api")` factory **replaces the whole module**, so
every import anywhere in the rendered tree must appear in it.
`fetchBoardWorkflows` reached this file *indirectly* — the task modals
ResearchView opens import it — so adding that export to product code
broke four cases that have nothing to do with board workflows.
Stubbed with the flag-OFF payload the server sends when multi-lane
boards are disabled, which is the shape these cases already assume.
## Measured
```
before: Tests 4 failed | 23 passed (27)
after: Tests 2 failed | 25 passed (27)
lint clean; fnxc-future-dates: none added
```
## The remaining 2 are a different cause and are NOT fixed here
They fail with `Number of calls: 0` — the enrich-task and create-task
actions never fire. That is a UI-wiring question, not mock completeness.
**I checked that my stub is not responsible**, rather than assuming:
re-running with `flagEnabled: true` and a populated workflow list
produces the *same* 2 failures, so the payload shape does not gate those
affordances. Left for whoever owns that surface.
## How this was found
While establishing a clean baseline for the resolver audit. That sweep
also reported `lazy-loaded-views-docs.test.ts` red — **it now passes**,
fixed by another worker between my measurement and this PR, which is why
the count here is 3 files rather than the 4 I reported in #3236.
Still red on main, untouched by this PR:
- `src/__tests__/planning-browser-e2e.test.ts` — `expected {
totalButtons: 1, …(7) } to match object { totalButtons: 1, …(6) }`; an
assertion shape gained a field.
- `src/__tests__/register-model-routes-kimi-k3-supplemental.test.ts` —
`Test timed out in 15000ms`. Per the standing rule a timeout with no
corresponding bug in the change is a **quarantine candidate**, not
something to appease with a longer timeout; I am not quarantining it
unilaterally since it is not my subsystem, but flagging it as the shape
that rule describes.
## What
Pins the **Reliability endpoint's three lane reads** — the last
uncovered resolver cluster the repo-wide audit found.
Two commits, deliberately separate:
1. **refactor** — extract the three resolves behind
`resolveReliabilityLanes(store)`. Behaviour-preserving, no test changes.
2. **test** — pin all three through that seam.
## Why a seam was needed
The three resolves lived inline in the `/api/health/reliability` route
closure. Blinding any of them left the **entire dashboard suite green —
21,582 tests** — and the only way to reach them was booting
`createServer` behind a mock-the-world shell the slow-test rule forbids.
**And the obvious test would not have helped.**
`reliability-metrics.test.ts` exercises `countEntriesInto`,
`countBouncesOut` and `inReviewDurationMetrics` with lane sets **passed
in by hand**. That proves the collaborators honour a resolved set; it
says nothing about whether the caller passes one. *A unit test of the
collaborator can never fail when the caller's resolve is blinded* — the
same trap the audit note records for `reads.ts`, where a suite written
for the exact conversion still could not see it.
The seam is the caller. It resolves, so blinding a resolve fails a test
of it.
## Measured — each blind fails exactly its own case
| blinded | fails |
|---|---|
| `REVIEW_ROLES` | "resolves the board's OWN review lane" |
| `["countsTowardWip"]` | "resolves the board's OWN wip lane" |
| `["complete"]` | "resolves the board's OWN complete lane" |
```
converted: Tests 6 passed (6)
each blind: Tests 1 failed (its own case only)
reliability-metrics.test.ts + this file: 28 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```
That isolation is the point: **three resolves in one function invite a
copy-paste that hands the same set to all three**, and every positive
assertion would still pass. There is a paired negative asserting each
renamed lane appears in *its* bucket and nowhere else — without it the
duration metric could silently measure review → review.
Also pinned: the degrade path. An unreadable workflow list must not fail
the endpoint, so the legacy ids still answer.
## What breaks without the conversion
On a board that renames either lane, every underlying query returns `{}`
— so `tasksEnteredInReview` and `tasksBouncedToInProgress` are zero for
every day, and `inReviewFailureRate7d` divides one zero by another and
reports a **healthy** rate. It produces a NUMBER, not an error, and the
number says everything is fine. An operator reading 0% review failures
beside a populated audit list has no reason to suspect the metric is
blind.
## The one observable difference in the refactor, stated not buried
The complete-lane read moves from *after* the counting `Promise.all`
into the same phase as the review/wip pair. These are pure reads of
workflow definitions — no writes, no ordering dependency — so the
resolved values are identical; only the concurrency shape changes (three
parallel reads instead of two-then-one). Flagging it because
"behaviour-preserving" should be a claim someone can check, not an
assertion.
## Audit status
With this, **3 of the 4 flagged sites are closed**. Remaining:
`cli/commands/task.ts:660`, where the glyph decision is inline in
`runTaskList` and the same seam argument applies — but its sibling test
file already documents that driving that function needs the forbidden
shell, and extracting a helper there would produce a test that *looks*
like coverage while leaving the resolve unpinned. Left flagged rather
than faked.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Reliability health metrics now recognize configured review,
work-in-progress, and completion lanes, including renamed workflow
lanes.
- **Bug Fixes**
- Improved fallback behavior when workflow definitions are unavailable,
preserving compatibility with legacy lane configurations.
- Ensured lane resolution remains isolated by role for more accurate
reliability metrics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## What
Pins the **last two uncovered lane reads** in `packages/core`.
Test-only. This closes the per-site core audit.
| site | what it decides |
|---|---|
| `async-mission-store.ts:1179` | is an ARCHIVED card valid terminal
evidence for mission repair? |
| `task-id-integrity.ts:502` | does an archived child still count as a
LIVE lineage child? |
## Measured
```
mission-store: 39 passed clean; 1 failed | 38 passed blinded
lineage: 3 passed clean; 1 failed | 2 passed blinded
lint clean; fnxc-future-dates: none added; census unchanged
```
Both blinds confirmed applied with `git diff --stat` before each run.
## The third adjacent-pair split
`:1179` is the **archived** half of a pair whose **complete** half
(`:1178`, *one line above*) was already covered by a test in the same
file, written for exactly this concern. Terminal evidence is "done OR
supported archived state," so an archived card is equally valid repair
evidence — but on a board whose archive lane is `vaulted` the archived
half could not see it, and reconciliation threw `TASK_NOT_TERMINAL` for
a card that was genuinely filed away. Same refusal the covered case
fixed, reached through the other door.
That is now the third confirmed instance in core (after `team-analytics`
in #3227 and the scheduler pair earlier). **Being adjacent to a covered
resolver is not coverage**, and it is the most reliable place to look.
## What breaks without the lineage read
An archived child is filed away, not live, so it must not hold the
delete gate shut. Renamed, it still counted as live and
`TaskHasLineageChildrenError` blocked the parent's delete **forever** —
the operator archived the child *precisely* to clear the way, and the
gate could not see that they had.
## A fixture detail I got wrong first
My first mission fixture created a live card in a `vaulted` column and
failed with `deleted or archived without a valid retained tombstone and
archive snapshot` — nothing to do with the lane read.
The `archived` verdict requires **all three** of `deletedAt !== null`,
an archive-snapshot row, and `isArchived(column)`. A live card merely
sitting in an archive-trait column is `invalid-deleted`, not `archived`.
The test now archives for real and *then* renames the recorded lane,
which isolates the third condition — the only one under test. Recorded
in the file so the next person does not re-derive it.
## Paired positives
Both files pin the complement: a WORKING child still counts as live.
Recognising the renamed archive lane must not degrade into "no child is
ever live" — that would silently **disable** the lineage gate and let a
parent be deleted out from under real descendants, which is worse than
the bug being fixed.
## Core audit complete
**14 sites blinded individually: 9 already covered, 5 uncovered, all 5
now pinned** (#3233, #3234, this PR).
Every `resolveProjectColumnsForRoles` call site in `packages/engine` and
`packages/core` has now been blinded. Remaining unaudited: `dashboard`
(2 files) and `cli` (1) — I claim nothing about those.
Found by a live browser E2E of the coding workflow in test mode: every
scripted full-task run failed at `steps#0:step-execute` with `Step 4 out
of range (task has 4 steps)`, rebounding through recovery forever.
**Root cause:** `fn_task_update.step` has been **0-based since FN-6607**
(executor.ts FNXC:StepNumbering — the old `step - 1` conversion made
Step 0 impossible to mark). `mock-provider.ts` still sent `index + 1`,
so test mode marked steps 1..N instead of 0..N-1: Step 0 (Preflight)
never completed and step N threw out-of-range. Test mode's full-task
path has been broken since June.
**Also fixes the test that pinned the bug:** `mock-provider.test.ts`
expected `{ step: 1 }` for a fixture whose first unfinished step is
index 0 — the expectation encoded the 1-based off-by-one.
Verified: 12/12 mock-provider tests; the live E2E instance completes the
task after this patch (see follow-up screenshot in the session).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## What
Pins **both** create-time duplicate guards in
`branch-and-pr-entities.ts`. Test-only.
| site | method | excludes |
|---|---|---|
| `:445` | `findRecentTasksByContentFingerprint` | ARCHIVED (unless
`includeArchived`) |
| `:484` | `findRecentTasksBySourceParentTaskId` | COMPLETE and ARCHIVED
|
Blinding either back to its literals left the entire 16-file
lane-detector set green. **No test in `packages/core` reaches either
method.**
## Measured
```
converted: Tests 8 passed (8)
blinded :445 Tests 1 failed | 7 passed (8) <- only the fingerprint case
blinded :484 Tests 2 failed | 6 passed (8) <- only the sibling cases
lint clean; fnxc-future-dates: none added; census unchanged
```
**Each blind fails exactly its own cases.** That matters: it proves the
two resolvers are pinned *independently*, rather than one broad test
appearing to cover both. Blinding `:445` leaves every sibling case green
and vice versa — so neither is riding on the other's coverage.
## They fail in opposite directions
This is why both belong in one file:
- **Fingerprint guard** — a renamed board leaves archived cards in the
candidate set, so filing a new task is **refused as a duplicate** of one
the operator already archived. The create is blocked and the thing
blocking it is invisible.
- **Sibling guard** — a renamed board leaves finished siblings in the
"recent live siblings" set, so completed work keeps counting as active.
One over-includes into a *refusal*, the other over-includes into
*phantom activity*. Neither raises an error.
## Positives pinned too
A LIVE fingerprint match is still a duplicate candidate;
`includeArchived: true` opts the renamed archived lane back in; a
WORKING sibling is still live. Excluding the finished lanes must not
degrade into excluding everything, or the guards stop guarding — the
failure mode a lane-widening change invites.
## A fixture detail that would have made this vacuous
Both queries cut off at `Date.now() - windowMs`, with `windowMs` capped
at 24h. The sibling harness I copied from seeds a **fixed past
timestamp**, which falls outside that window — every case would then
pass on an empty result, including the ones that are supposed to fail
under blinding. Fixtures are seeded at current time instead, and the
reason is recorded in the file so nobody "tidies" it back to a frozen
date.
## Progress
3 of the 5 uncovered core sites are now pinned (`store.ts:1135` in
#3233, these two here). Remaining and unclaimed:
`async-mission-store.ts:1179` and `task-id-integrity.ts:502`.
## What
Follow-up to the #3224 review comment *"reject legacy archived
identifiers for renamed workflows."* Test-only.
That comment had two halves. **The half it got wrong** is already
answered on main: asserting an exact single-column set `["vaulted"]`
*fails*, because `resolveProjectColumnsForRoles` unions
`LEGACY_COLUMN_IDS_BY_ROLE` in as a documented floor — so a row whose
workflow cannot be resolved still classifies. Pinning `["vaulted"]`
would encode the opposite of the design.
**The half it got right was never addressed.** `toContain` also passes
when the set grows a lane nobody intended, and an over-broad archived
set silently classifies *live* rows as archived. So the reviewer's worry
was legitimate even though the proposed fix was not.
The exact set is assertable — it just is not the one the review
proposed:
| board | resolved set |
|---|---|
| default | `["archived"]` |
| renamed | `["archived", "vaulted"]` — legacy floor + the board's own
lane |
`[...].sort()` in the helper makes ordering stable, so these pin the
resolver's whole answer rather than a substring of it.
## Proven to catch what `toContain` missed
Giving the fixture a second `archived`-trait column fails both new
assertions:
```
AssertionError: expected [ 'archived', 'cold-store' ] to deeply equal [ 'archived' ]
AssertionError: expected [ 'archived', 'cold-store', 'vaulted' ] to deeply equal [ 'archived', 'vaulted' ]
Tests 2 failed | 1 passed (3)
```
The previous `toContain` assertions pass unchanged against that same
spurious lane. That is the whole justification for this PR — without the
injection test it would be a stylistic preference.
```
clean: Tests 3 passed (3)
spurious lane: Tests 2 failed | 1 passed (3)
lint clean
```
## Note
I said in the review thread I would tighten this, so this closes that
loop. I also checked before editing whether another worker had already
done it — main already carries the FNXC tag and the legacy-floor
explanation from the same review round, so this PR adds only the part
still missing rather than redoing settled work.
## What
Pins the **open-undo query's finished-lane exclusion** in
`packages/core/src/store.ts`. Test-only.
`findOpenRevertTaskForSource` answers *"is there an OPEN undo task for
this source?"* — the question behind the dashboard's Undo affordance. It
answers by **excluding the finished lanes**, so a prior undo that
already landed does not keep rendering as open.
Blinding that exclusion back to `ne(column,"archived"),
ne(column,"done")` left the entire 16-file lane-detector set green. **No
test in `packages/core` reaches this method at all.** The dashboard-side
twin (`taskRevert.ts`, #3129) is tested; the store-side query behind it
was not.
## Measured
| | default (control) | renamed complete | renamed archived | working
lane |
|---|---|---|---|---|
| converted | pass | pass | pass | pass |
| blinded to `["done","archived"]` | pass | **FAIL** | **FAIL** | pass |
```
converted: Test Files 1 passed (1) / Tests 4 passed (4)
blinded: Test Files 1 failed (1) / Tests 2 failed | 2 passed (4)
lint clean; fnxc-future-dates: none added; census unchanged
```
Blind confirmed applied with `git diff --stat` before the run.
## What breaks without it
On a board whose complete lane is `shipped`, neither literal matches, so
a **done** undo task is never excluded and the query keeps returning it.
The card shows an undo already in flight *forever*, and the real
affordance is unreachable. Nothing errors — the button is just
permanently wrong, which is why it went unnoticed.
## Includes the paired positive
An undo still in a **working** lane IS reported as open. Excluding the
finished lanes must not degrade into excluding everything, or the
affordance breaks in the other direction and no undo is ever reported in
flight. Both new failing cases are renamed-lane cases; both survivors
are cases that should survive.
## Where this came from
Per-site blinding of all 14 remaining `resolveProjectColumnsForRoles`
call sites in `core`, run against a 16-file detector set. **9 covered, 5
uncovered:**
| site | verdict |
|---|---|
| `store.ts:1135` | **uncovered** → pinned here |
| `async-mission-store.ts:1179` (archived) | **uncovered** — its
neighbour `:1178` (complete) is covered |
| `branch-and-pr-entities.ts:445` | **uncovered** |
| `branch-and-pr-entities.ts:484` | **uncovered** |
| `task-id-integrity.ts:502` | **uncovered** |
| `reads.ts` ×3, analytics ×3, `eval-automation`, `task-artifacts-ops`,
`async-mission-store:1178` | covered |
The first run of that probe was **invalid and I nearly published it**:
it reported all 14 sites "COVERED" with *zero failing tests*. zsh does
not word-split unquoted parameter expansions, so `vitest run $DET`
passed 16 paths as one argument and vitest exited 1 with "No test files
found" — which my script read as a failing test. The re-run treats that
string as `INVALID` rather than a result. Third time this session a
wrong reading came from test *selection* rather than from blinding.
## Flagged, not guessed
The four remaining uncovered sites are named above rather than quietly
left; `async-mission-store` shows the same adjacent-pair split as
`team-analytics` in #3227, which is now the third confirmed instance of
that shape.
## What
**Fixes a red main.**
`workflow-reconciliation-production-shape.pg.test.ts` has been failing
with `expected 'todo' to be 'triage'`. Test-only.
Found while establishing a clean baseline for an unrelated coverage
audit — my tree was clean at `origin/main` (`76c73238a0`), so this is
not something I introduced. It is in the non-blocking suite, which is
why it has stayed red.
## It is not a regression — the test was the stale half
The delete path was deliberately fixed to re-home occupants using
`resolveEntryColumnId(resolveDefaultWorkflowIr())` instead of
`BUILTIN_CODING_WORKFLOW_IR`. This assertion was not updated with it.
The two IRs are **not the same board**:
| IR | entry column |
|---|---|
| `BUILTIN_CODING_WORKFLOW_IR` (`builtin:legacy-coding`) | `triage` |
| `resolveDefaultWorkflowIr()` (the catalog default) | `todo` |
Re-homing into `triage` put cards in a column the default board never
declares. It slipped past `moveTask`'s undeclared-target guard **only
because `triage` is a legacy id** and the recovery-rehome path exempts
those — so the guard that exists to stop exactly this could not see it.
So `todo` is the correct behaviour and the literal `"triage"` was what
needed fixing.
## Why it asserts a resolver rather than `"todo"`
Swapping one hardcoded id for another would be the identical trap one
rename later — the same class of defect this whole program exists to
remove. The expectation now derives from **the same two functions the
product path calls**, so it cannot drift out of sync with them again.
I also added the complement: the card must genuinely have **left** the
vanished column, not merely match whatever a resolver returns. Without
it, a resolver that started returning `custom-hold` would pass.
## Proven not appeasement
Reverting the product line to the legacy IR — the original defect —
fails this test:
```
AssertionError: expected 'triage' to be 'todo'
Test Files 1 failed (1) / Tests 1 failed | 6 passed (7)
```
That is the check that matters for a test edit that turns a red green.
It fails on the defect it describes.
## Measured
```
before: Tests 1 failed | 6 passed (7)
after: Tests 7 passed (7)
16-file detector set: 181 passed (16 files) [was 1 failed | 180 passed]
lint clean
```
## Note on the reading
I got this wrong twice before getting it right, and the record is worth
having. My first read was "the test is stale, `triage` was merged away."
My second was "the builtin IR still declares `triage`, so the
*behaviour* is the defect" — which the IR file superficially supports.
Only the third reading, of the FNXC note at the fix site, showed the
file I was reading is the **legacy** IR and not the default one. Two of
those three readings would have produced a confidently wrong PR; the
deciding evidence was the comment the fixing author left at the call
site, which is a good argument for writing them.