c67aafde1ccf673fc9a2ee8f04b13e155646bcff
11422 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c67aafde1c |
revert(fnxc): restore seven author stamps I falsified while chasing a gate bug (#3282)
Closes #3279. Undoes the damage my #3261 did, now that #3277 has landed and made it safe. ## What went wrong #3277 established that last night's "future-dated" stamps were **correct** — the author's local date in a UTC+1 container, written minutes before their commits. The gate compared against the runner's local calendar (PDT) and called them tomorrow. I diagnosed it as author error and repointed seven stamps to turn main green. The values I wrote were **neither the author's local time nor UTC** — invented times chosen to satisfy a broken check. The FNXC record is this project's why-does-this-exist trail, so those stamps misstated when the work happened. ## Restored verbatim | file | mine (wrong) | restored | |---|---|---| | `workflow-column-boundary-capacity.test.ts` | `22:30` | `2026-08-01-00:30` | | `runtimes/in-process-runtime.ts` | `22:20` | `2026-08-01-00:20` | | `scheduler.ts` (`MissionReconciliation`) | `22:00` | `2026-08-01-00:00` | | `workflow-column-boundary-hooks.ts` | `22:20` | `2026-08-01-00:20` | | `workflow-column-boundary.ts` ×2 | `22:20` | `2026-08-01-00:20` | | `workflow-graph-task-runner.ts` | `22:20` | `2026-08-01-00:20` | ## The check that mattered Sequencing was deliberate — #3277 had to land first or this would have re-reddened main. The real question is whether the gate now accepts the **originals**, measured across the rollover boundary at local `2026-07-31 17:23 PDT` / UTC `2026-08-01 00:23`: ``` America/Los_Angeles exit 0 Europe/Paris exit 0 UTC exit 0 Asia/Tokyo exit 0 ``` `123 known future-dated stamp(s), none added`. **No baseline change needed** — #3278's pruning already re-recorded `scheduler.ts`, and these are known stamps rather than new ones. Stamps only: `git diff` shows **zero** non-FNXC lines, 7 insertions / 7 deletions across 6 files. `census --strict` 0, `pnpm test:gate` 0. ## The part worth keeping I argued against exactly this on #3263 — *"it rewrites stamps whose authors are not us"* — and then did it myself six lines later, because I was confident about a cause I had not checked. The commits' timestamps were available the entire time; I read the runner's clock and never asked what timezone the **author** was in. Four of last night's seven PRs were fixing something that was not broken. This is the cleanup for my share of that. |
||
|
|
500f40e65b |
fix: descriptive waiting badges (Queued to revise / Queued behind FN-X) + dependency-free blocked exits replan calmly
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
5acd8e987b |
fix(fnxc): three future-dated scheduler stamps — main red through three closed fixes (#3280)
**`check-fnxc-future-dates` exits 1 on `origin/main`.** ``` packages/engine/src/scheduler.ts: 3 future-dated stamps, baseline allows 2 FNXC:ConcurrencyAdmission 2026-08-06-09:00 (six days out) FNXC:WorkflowLifecycleColumns 2026-08-01-05:00 FNXC:WorkflowScheduling 2026-08-01-01:05 ``` All three repointed to `2026-07-31`, times preserved. Gate now exits 0. ## This red has outlived three owners #3270, #3272 and #3274 were each opened against it and each **closed without merging**. Main has been red on this gate for hours while three fixes came and went. Claimed with `check-file-claimed.mjs` before starting — only #3262 touches `scheduler.ts`, and it is a terminal-role refactor rather than a stamp fix, so this was genuinely unowned. ## Why this keeps recurring Seven incidents in roughly two hours. The mechanism, in one line: **the date check runs only in CI** (`pr-checks.yml:66`, no pre-commit or pre-push hook), so every PR is validated against main's baseline *at its own CI time* and cannot see a concurrent or later change. Two PRs stamping the same file both pass, then compose into a red main. One case (#3273) was a stale branch **reverting** an already-merged fix. Patching instances has not converged — this PR is the eighth attempt at the same class. Two structural options, neither of which I am landing unilaterally since the second changes the gate's contract: - run the date check at **author time** (pre-push); it needs no baseline for "is this date in the future", so it cannot be raced - make the date rule **baseline-free** — a future-dated stamp is always wrong, unlike a lifecycle literal that may be a deliberate fallback `2026-08-06` being six days out also suggests these are not off-by-one timezone slips but stamps written from an intended future date. ## Verification - `check-fnxc-future-dates` — **exit 0** (was exit 1 on main) - `scheduler` suites — **148 pass** - `tsc --noEmit` (engine) — 0 errors - comment-only diff, no behaviour change Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
475bb2d641 |
FN-8637: restrict Quick Add Start to manual-intake workflows
Restrict Quick Add Start eligibility to verified manual intake lanes. - Require the server-derived manualIntake flag instead of hold alone. - Preserve Coding Ideas routing while hiding Start for Coding's merged planning lane. - Cover desktop and mobile eligibility behavior and document the updated rule. - Add a patch changeset for the corrected workflow gating. Files changed: .changeset/fn-8637-quick-add-start-manual-intake.md | 7 ++++ docs/dashboard-guide.md | 2 +- packages/dashboard/app/components/QuickEntryBox.tsx | 21 ++++++------ packages/dashboard/app/components/__tests__/Column.test.tsx | 20 ++++++++--- packages/dashboard/app/components/__tests__/ListView.test.tsx | 40 ++++++++++++++++++---- packages/dashboard/app/components/__tests__/QuickEntryBox.test.tsx | 21 +++++++++--- packages/dashboard/app/utils/__tests__/quickAddStart.test.ts | 38 ++++++++++++++++++-- packages/dashboard/app/utils/quickAddStart.ts | 9 ++++- 8 files changed, 128 insertions(+), 30 deletions(-) Fusion-Task-Id: FN-8637 Fusion-Task-Lineage: 9652d7d8-f954-49e8-9a76-a2421654baae Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
911b7f1c31 |
test(engine): memoize identical full-repo census spawns in lifecycle-column-census
The lifecycle-column-census test file spawned the census CLI ~14 times, each parsing every tracked source file to a TypeScript AST (~2s for ~1960 files). Many spawns were byte-identical, deterministic, read-only real-repo scans: the --json census 4x, the plain report 2x, plus a repeated identical --update-baseline tree-sync across the ratchet cases. Memoize each distinct read-only spawn's output (keyed by argv) and reuse the synced baseline JSON, collapsing duplicate full-AST scans without changing any assertion. File wall-time: 30.5s -> 18.6s (-39%). 57/57 tests still pass. Fusion-Task-Id: FN-slow-test-census |
||
|
|
95410b5de6 |
FN-8638: add Factory Light dashboard theme
Add a daylight industrial theme that persists across dashboard and desktop startup. - Register Factory Light in persisted theme types, selectors, and bootstrap validators. - Define Factory Light tokens and preview swatches for light and dark modes. - Cover theme registration and rendered token contracts, and document the new option. Files changed: .changeset/fn-8638-factory-light-theme.md | 7 ++ docs/dashboard-guide.md | 3 +- docs/settings-reference.md | 2 +- packages/core/src/types/execution-and-ui.ts | 2 + .../app/__tests__/factory-light-theme.test.ts | 106 +++++++++++++++++++++ .../dashboard/app/components/ThemeSelector.css | 14 +++ .../components/__tests__/ThemeDropdown.test.tsx | 2 +- .../components/__tests__/ThemeSelector.test.tsx | 2 +- .../__tests__/CommandCenterControls.test.tsx | 2 +- packages/dashboard/app/components/themeOptions.ts | 1 + packages/dashboard/app/index.html | 2 +- packages/dashboard/app/public/theme-data.css | 86 ++++++++++++++++- packages/desktop/src/renderer/index.html | 1 + 13 files changed, 223 insertions(+), 7 deletions(-) Fusion-Task-Id: FN-8638 Fusion-Task-Lineage: 3b78bc31-0f03-4299-8f5f-1a69ac7c604a Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
5bf9279c5d |
FN-8636: hide unavailable task card cost badges
Hide dash-only cost badges from board task cards when pricing is unavailable. - Suppress unavailable cost labels and their empty layout shells. - Cover priced and unavailable badges with and without Promote across desktop and mobile widths. - Add a patch changeset for the board-card fix. Files changed: .changeset/fn-8636-card-cost-badge-dash.md | 7 ++ packages/dashboard/app/components/TaskCard.tsx | 8 +- .../__tests__/TaskCard.cost-badge.test.tsx | 93 ++++++++++++++++------ .../app/components/__tests__/TaskCard.test.tsx | 14 ++-- 4 files changed, 84 insertions(+), 38 deletions(-) Fusion-Task-Id: FN-8636 Fusion-Task-Lineage: 31fbf82a-7a67-4629-bf82-48faf3c3a9d7 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
91c3854607 |
fix(fnxc): the last future-dated stamp keeping main red (#3269)
**`check-fnxc-future-dates` exits 1 on `origin/main`.** This is the last stamp causing it. ``` packages/core/src/task-store/lifecycle-ops.ts: 1 future-dated FNXC stamp, baseline allows 0 FNXC:Diagnostics 2026-08-01-00:50 (today is 2026-07-31) ``` Corrected to `2026-07-31-00:50`. One character. `check-fnxc-future-dates` now exits 0; `tsc --noEmit` clean. ## Why this was left behind Four PRs converged on this red main — #3262, #3263, #3265, and my own #3266 (closed as superseded). Between them they covered the census rise and the boundary-work stamps. **None touched `lifecycle-ops.ts`**, so the gate stayed red after the others landed. That is the predictable failure of parallel work on one symptom: everyone fixes the part they saw first, and the residue survives because each author checked "is main green now?" against their own branch rather than against main. ## I claimed before working this time ``` node scripts/check-file-claimed.mjs packages/core/src/task-store/lifecycle-ops.ts → UNCLAIMED ``` Then pushed the branch before editing. I did the opposite on #3266 — built it, then discovered #3265 already covered it — which was the sixth duplication of the phase and my third. The tool answers in one command; the discipline is running it *first*. ## Verification - `check-fnxc-future-dates` — **exit 0** (was exit 1 on main) - `census --strict` — exit 0 (already green; #3265's marker landed) - `tsc --noEmit` (core) — 0 errors - one-character diff, no behaviour change Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f95c6d53e |
fix(engine): a Ready card's retained worktree transfers on release instead of blocking it
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8e6b0ad67e |
test(engine): main is red — the census preconditions were MET, repoint two guards that expired on success (#3260)
## main is red, and this is the second half of it Running the full `engine-default` project on `origin/main` — 754 files, 10,538 tests — returns **3 files / 4 tests failing**. One is the parked-seam audit counter, fixed in #3258. The other three are here. All three assert that lifecycle debt **still exists**. It does not: | assertion | expected | actual | | --- | --- | --- | | `finds >10 literal move targets, so the census is not vacuous` | > 10 | **0** | | `both recoveryRehome groups are non-empty` | > 0 each | **0 / 0** | | `still says ROSE when guards genuinely grew` | exit 1 | **exit 0** (mutation was a no-op) | Nothing regressed. The conversion program drove the engine's literal move-target population to zero and the census baseline to zero entries. **Each guard's premise was "the debt still exists", so each expired the moment the work succeeded — and expired by failing, which reads as a regression in the very thing it was guarding.** The third is the sharpest: it manufactured a rise by finding a baseline entry with more than one guard and zeroing it. With `byFile` empty there was nothing to find, so it mutated nothing, the census correctly passed, and the test asserted exit 1 against a correct pass. ## Repointed, not deleted A vacuity guard must not depend on real debt existing. Two halves: - **Vacuity now runs against a synthetic fixture** — a small in-memory source with three `moveTask` literals (one with `recoveryRehome`, one targeting an undeclared column). The collector is exercised forever regardless of how much real debt remains. This is what keeps the rest honest: at a real population of 0 the zero-assertions are trivially true, and **only the fixture proves they would still fire**. - **The real-tree assertions now assert zero**, so a reintroduced literal move target fails them. Same guarantee as before, pointed at the state the tree is actually in. The ROSE case builds its own one-file tree through the `FUSION_CENSUS_FILE_ROOT`/`FILE_LIST` seam from #3230 rather than borrowing a baseline entry that no longer exists — a rise **constructed** instead of borrowed. It also now asserts the failure is *not* the reclassification wording, which is the distinction that file exists to protect. ## Verification | mutation | expected | result | | --- | --- | --- | | reintroduce a literal `moveTask(id, "in-review")` in the engine | fail | exit 1 ✅ | | blind the collector to return `[]` | fail | exit 1 ✅ | | remove the `ROSE` wording from the census script | fail | exit 1 ✅ | | clean tree | pass | exit 0 ✅ | 60 tests green across the three census suites. All eight ratchets exit 0. Test-only; no changeset. **One honest note on my own method.** My first `ROSE` mutation replaced 1 of the 2 occurrences in the script and the test stayed green — which looks exactly like a dead assertion. It was an ineffective mutation, not a dead test; the manual run still printed `ROSE` from the other occurrence. Re-run against both, it failed. A mutation that does not actually change behaviour proves nothing, and it is worth checking that the mutation landed before concluding the test is dead — the same trap as reading a report-only ratchet's exit 0 as a pass. ## Not in scope `defaultColumnIds()` has a pre-existing type error (`Property 'columns' does not exist on WorkflowIrV1` — union narrowing). Untouched by this PR and the test executes fine; flagging rather than fixing, since it is unrelated to the red. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved census analysis reliability when evaluating guard-count increases and reclassification messages. * Improved detection and classification of literal move targets and declared columns. * Enhanced file path reporting for inputs outside the primary source directory. * **Tests** * Added isolated test scenarios using temporary census data and fixtures. * Strengthened validation of recovery, plain, and declared-column classifications. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
e52da740a5 |
fix(core): stale-orphan-dir skip logs at debug, not warn — steady-state per-sweep noise
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
78d411cfe2 |
fix: main is RED on two gates — record the new fallback, repoint six future-dated stamps (#3261)
`9094d1640e` (globalPause gates every graph node entry) reddened **two** lifecycle gates on main. Both are fixed here, in separate commits. ## 1. The census ratchet went 0 → 2 `isTerminalColumnTask` in `scheduler.ts`: ```ts const flags = columnFlagsForTask(task); if (flags) return flags.complete === true || flags.archived === true; return task.column === "done" || task.column === "archived"; // ← counted ``` **The code is correct.** It resolves traits first and falls back only when the workflow is unreadable. The census counts fallback literals on purpose — *"a fallback literal is still a literal and should go when the trait path becomes unconditional"* — and reports them beside the backlog as already-converted. Its own remedy for a legitimate one is a `DELIBERATE-LITERAL` marker at the site. Recorded rather than converted because **there is nothing to convert to**: a task whose workflow cannot be read has no resolved lane, and treating it as non-terminal would count a finished card's retained worktree against live capacity — the opposite of what the surrounding fix does. Marker sits in the declaration's **leading** comments; an inline one attaches to the wrong node and is silently ignored, which cost a miscount once before. Baseline re-recorded in the same commit, since the census tracks deliberate counts and reports a marker addition as `RECLASSIFIED`. ## 2. The stamp gate was red as well Six files stamped `2026-08-01-00:2x` while UTC was `2026-07-31`: ``` workflow-column-boundary.ts 2 workflow-graph-task-runner.ts 1 workflow-column-boundary-hooks.ts 1 in-process-runtime.ts 5 (allows 4) workflow-column-boundary-capacity.test 1 ``` This checkout is UTC-7, so "just after midnight local" is tomorrow in UTC — the case AGENTS.md documents, which passes `pnpm lint` locally *because* the local clock agrees with what was written. Second occurrence today; I fixed the same shape on #3208 for another worker. Repointed to `2026-07-31-22:2x`, preserving relative order. **Zero non-comment lines changed** — 8 lines across 6 files, verified by diffing out FNXC lines. ## Measured | check | before | after | |---|---|---| | `census --strict` | **1** | **0** | | backlog | **2** | **0** (DELIBERATE-LITERAL 148 → 150) | | `check-fnxc-future-dates` | **1** | **0** | | `pnpm test:gate` | 0 | 0 | | `census-reclassification-message` | 2 failed | **1 failed** | That last row is deliberate: the remaining failure is the expired-premise case #3260 fixes, and I have not touched it. The capacity test from `9094d1640e` still passes 9/9. ## Why this landed at all Both gates run in `pr-checks.yml`, so a PR carrying either would have gone red. Worth someone checking how it merged — a stale merge base would explain it, and if so the same hole is open for the next merge. |
||
|
|
9094d1640e |
fix(engine): globalPause gates every graph node entry; maxWorktrees counts planning/review holders
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c8a6af13a0 |
test(census): pin file DISCOVERY, which every existing test was blind to (#3259)
Follow-through on the recommendation I made reviewing #3256: **a gate needs a test for its file discovery, not only its matcher.** ## The gap This suite pinned the matcher and never the scan. Every case either feeds the classifier a source string or drives the CLI against the real tree — so **the file list could return empty and all 53 tests would still pass.** Not hypothetical. `git ls-files` lists tracked files only, so a new file with a plain `task.column === "in-review"` scored 0 until staged (#3254). The identical bug then turned up in the move-target ratchet **behind its own 12 matcher tests** (#3256) — I wrote those 12 specifically to stop that gate regressing, and they could not see it, because they import the matcher and never run a scan. ## Four cases, on a synthetic tree Driven through `FUSION_CENSUS_FILE_ROOT` + `FUSION_CENSUS_FILE_LIST`, so discovery is testable without creating files inside a checkout the operator writes to concurrently: - a guard in a scanned file **reaches the classifier** and is counted - `--strict` fails **for the right reason** (message names the file; not an ENOENT fail-closed) - **every** listed file is counted, not just the first - files are read from the **scan root**, so a listed path and a read path cannot diverge ## The second case earns its wording Its first version asserted only `code === 1` — and **passed while discovery was broken.** With the injected list ignored, paths come from the real repo while reads resolve against the fixture root, every read misses, and the gate fails closed with exit 1. Right code, unrelated cause. A test that cannot tell *"found a guard"* from *"could not read anything"* is not testing the ratchet. Asserting the message is what separates them. I found that only by checking which cases the control actually failed — 3 of 4, not 4 of 4. Had I stopped at "the control fails, ship it", I would have added a test that passes for the wrong reason to a suite whose whole purpose is catching tests that pass for the wrong reason. ## Measured | check | result | |---|---| | suite | **57 passed** (53 + 4) | | anti-vacuity: `injectedList` forced undefined | **all 4 fail** (3/4 before strengthening case 2) | | restored | 57/57 | | `census --strict` / `check-fnxc-future-dates` | 0 / 0 | Tests only — no gate or product change. The same four assertions port directly to the other lifecycle gates once each grows the fixture seam; the move-target ratchet is the obvious next one, and its `.mjs` is currently claimed by #3256. |
||
|
|
f3d7b73741 |
test(engine): main is red — re-record the parked-seam audit at 4-of-6, split by resolution path (#3258)
## main is red `workflow-optional-role-param-caller-audit-live-e2e.pg.test.ts` fails on `origin/main` in `engine-default`: ``` AssertionError: expected 4 to be 2 expect(parkedConverted.length).toBe(2); ``` Found by running the whole live-E2E corpus rather than trusting that it passes — 27 files, 159 tests, 1 red. Not in the merge gate, so it has been sitting there. The alarm fired **downward**, exactly as that file was written to: two more call sites started passing `parkedColumns`, and a counter fails when someone *closes* a gap as well as when someone widens it. ## Why I did not just write 4 Re-recording it at 4 would have laundered an inert conversion through the audit written to catch inert conversions. Walking all six sites before touching the number: | site | `parkedColumns` provenance | verdict | | --- | --- | --- | | `agent-heartbeat.ts:1267` | — | unconverted | | `agent-heartbeat.ts:3796` | — | unconverted | | `self-healing.ts:13184` | `await resolveProjectColumnsForRoles` (:13170) | async-resolved | | `self-healing.ts:13294` | `await resolveProjectColumnsForRoles` (:13293) | async-resolved | | `task-agent-sync.ts:243` | `await resolveLinkSyncColumnRoles` (:225) | async-resolved | | `scheduler.ts:1798` | `resolveTaskParkedColumnsSync` (:1797) | **SYNC — INERT** | `scheduler.ts:1798` resolves through `resolveTaskParkedColumnsSync` → `getTaskWorkflowSelectionImpl`, which is `undefined` for every task under PostgreSQL. The resolver then takes its `!workflowId` branch and returns the **default builtin IR** — not `undefined` falling through to a legacy arm, but a real IR resolving real traits, with full confidence. It answers `hold`/`intake` as `todo`/`triage` on every board, exactly as the literal did. Driven proof: `workflow-scheduler-sync-role-conversion-inert-live-e2e.pg.test.ts`. **The shape count (4) and the live count (3) are different numbers, and only the second is about behaviour.** Both are now asserted, plus the sync-resolved site by name so it cannot quietly become "just one of the four". ## A mistake worth recording, because the test caught it My first draft keyed on `await` appearing inside the call window. The argument is nearly always a variable (`[...driftedParkedColumns]`, `roles.parked`) and the `await` lives in that variable's **assignment**, several lines above. That draft classified all four sites as inert — and it **would have passed** had I written the expected number to match what it measured. It failed only because I asserted 2 live from reading the source first, and the mismatch exposed the detector. Classification is now by provenance: take the root identifier, find where the file assigns it, ask whether *that* is awaited. The same bug recurred in my named-site check and failed the same way. That is the whole hazard of this program in miniature — a source-text audit that measures nothing looks exactly like one that measures everything, and it is the *number you expected* that catches it, not the green. ## Not asserted, deliberately `self-healing.ts:13184` is inert for an unrelated reason: its gate is `hasFreshRun || hasActiveExecution` and never reads `shouldPreserveParkedLink`, so its correctly-resolved set decides nothing today. Its own FNXC note says so. Resolution path is mechanically checkable; "the gate never reads the answer" is not, and asserting it on a string match would produce a number nobody could maintain. Recorded as prose in the header. I also did not touch `scheduler.ts`. It is a real defect, not a deferral, but converting it is not this file's job — it is named in the test so the next person converting it is sent here to move it from the inert list to the live one. ## Verification Mutation-verified — a passing audit proves nothing until it has been seen to fail: | mutation | expected | result | | --- | --- | --- | | convert the scheduler site to the async resolver | fail (3/1 → 4/0) | exit 1 ✅ | | drop `parkedColumns` from a converted site | fail (shape 4 → 3) | exit 1 ✅ | | add a new unconverted caller | fail (calls 6 → 7) | exit 1 ✅ | | clean tree | pass | exit 0 ✅ | Full 27-file live-E2E corpus green (159 tests). Ratchets exit 0. Test-only; no changeset. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Expanded end-to-end audit coverage for workflows with optional role parameters. * Improved validation of caller resolution paths, including asynchronous and synchronous scheduling scenarios. * Updated expectations to reflect all supported conversion paths and strengthened verification of scheduler behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
08a4e418f7 |
fix(dashboard): idle Revising badge explains it is queued for a planning slot
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
37d891e879 |
FN-8633: improve tablet terminal dragging
Give floating tablet terminals a dedicated drag grip while preserving tab-strip panning. - add a touch-sized tablet-only header drag grip and pop-out hit target - preserve floating geometry at the tablet breakpoint and document the gesture - cover grip availability, dragging, and horizontal tab-panning CSS isolation Files changed: .changeset/fn-8633-tablet-terminal-drag.md | 7 ++ docs/dashboard-guide.md | 4 +- packages/dashboard/app/components/TerminalModal.css | 61 +++++++++++++ packages/dashboard/app/components/TerminalModal.tsx | 16 ++++ packages/dashboard/app/components/__tests__/TerminalModal.test.tsx | 99 ++++++++++++++++++++++ 5 files changed, 185 insertions(+), 2 deletions(-) Fusion-Task-Id: FN-8633 Fusion-Task-Lineage: f1c442a8-3302-4f6a-98e9-f1efa4083c12 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
1e50b71255 |
fix(engine): reap leaked fn-verify verification worktrees in the temp-dir sweep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a20ddf6ed6 |
fix(core): refine + duplicate create into the resolved intake lane, not the deleted triage column
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3916e062aa |
fix(dashboard-quality): the lane runner could not be asked to run every lane (#3248)
The structural half of #2784. I re-measured its 123 failures on current main and **all four reported lanes are green** (143 / 2010 / 1961 / 5149 / 2103 passing). This fixes the reason nobody saw them. ## The mechanism `pnpm --filter @fusion/dashboard test` sets `stopScheduling = true` on the first failing lane, so the rest never run — and there was **no flag to ask for a full pass**. The report said: ``` [dashboard-quality] skipped 9 lane(s) after first failure ``` Nine lanes with **unknown** status and nine **passing** lanes produce the same absence of failure text. That is how 123 failures accumulated behind one red lane, and it is why the original issue could only be written by running all twelve lanes by hand. ## What changes, and what deliberately does not Fail-fast stays the **default** — fast feedback on a broken lane is right, and changing it would slow everyone for a rare case. - `--all` (alias `--no-fail-fast`) runs every lane and reports every failure. - `runQualityTests({ failFast })` so the behaviour is reachable from a test, not just the CLI. - The skip line now states the consequence and the remedy: lanes were **NOT RUN**, status **UNKNOWN rather than passing**, and `--all` shows the full set. ## Both halves pinned A flag nobody can prove works is the same as no flag: | test | asserts | |---|---| | DEFAULT stops after the first failing lane | `launched === ["one"]`, `skipped: 2` | | `failFast:false` runs all three | `launched === ["one","two","three"]`, `failed === [one, three]` | The second is the load-bearing one: **lane three ran even though lane one had already failed**, and both failures are reported rather than only the first. **Anti-vacuity control:** reverting the `if (failFast)` plumbing fails the second test and only it (`1 failed / 5 passed`); restoring passes `6/6`. ## Scope Runner and its tests only. No lane contents, no vitest configs, no CI workflow — CI already invokes lanes individually, so this changes local behaviour and the shared helper, not what CI runs. eslint clean; `check-fnxc-future-dates` exit 0. Suggest #2784 closes on the measured-green half and links here for the structural half, so the mechanism does not close along with the symptom that exposed it. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an option to run all quality-test lanes, even when earlier lanes fail. * Added `--all` and `--no-fail-fast` command-line options. * Quality tests now stop on the first failure by default, with clearer output for skipped lanes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
bcaa48390b |
FN-8627: add Sage color theme
Add the Sage palette across persisted dashboard and desktop theme selection paths. - Register Sage in core, dashboard bootstrap, desktop, and selector metadata. - Add dark and light Sage tokens plus independently resolvable swatches. - Cover registration, token, selector, and documentation updates. Files changed: .changeset/fn-8627-sage-theme.md | 7 ++ docs/dashboard-guide.md | 3 +- packages/core/src/types/execution-and-ui.ts | 2 + .../dashboard/app/__tests__/sage-theme.test.ts | 101 +++++++++++++++++++++ .../dashboard/app/components/ThemeSelector.css | 14 +++ .../components/__tests__/ThemeDropdown.test.tsx | 2 +- .../components/__tests__/ThemeSelector.test.tsx | 2 +- .../__tests__/CommandCenterControls.test.tsx | 2 +- packages/dashboard/app/components/themeOptions.ts | 1 + packages/dashboard/app/index.html | 2 +- packages/dashboard/app/public/theme-data.css | 86 +++++++++++++++++- packages/desktop/src/renderer/index.html | 1 + 12 files changed, 217 insertions(+), 6 deletions(-) Fusion-Task-Id: FN-8627 Fusion-Task-Lineage: fd4353b3-1e0c-4c7e-84dd-bcad2815178c Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
46d5019e2d |
FN-8631: remove task card bottom whitespace
Make progress-bearing task cards use their content height without an unused trailing band. - Remove the fixed minimum height from the task-card steps toggle. - Cover trailing-row layout across desktop and mobile task-card variants. - Add a patch changeset for the visual layout fix. Files changed: .changeset/fn-8631-task-card-bottom-space.md | 7 ++ packages/dashboard/app/components/TaskCard.css | 8 +- .../app/components/__tests__/TaskCard.test.tsx | 140 +++++++++++++++++++++ 3 files changed, 153 insertions(+), 2 deletions(-) Fusion-Task-Id: FN-8631 Fusion-Task-Lineage: 408d359f-66ed-4510-8974-3debbf76860f Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
cde02b423d |
FN-8632: align Command Center concurrency controls
Align Command Center capacity controls so their slider tracks remain visually synchronized. - Use a two-column grid for the surviving per-project capacity sliders. - Stretch slider cards and bottom-align range inputs despite optional running-count captions. - Add regression coverage and a patch changeset for the layout correction. Files changed: .changeset/fn-8632-concurrency-layout.md | 7 ++++ .../command-center/CommandCenterControls.css | 22 ++++++---- .../__tests__/CommandCenterControls.test.tsx | 47 +++++++++++++++++++++- 3 files changed, 67 insertions(+), 9 deletions(-) Fusion-Task-Id: FN-8632 Fusion-Task-Lineage: c59f52fc-0e5b-4633-98fa-64b8a60621d0 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
f86d758f9b |
FN-8630: balance Task Detail scrollbar insets
Keep Task Detail content symmetrically inset when its body scrolls. - Reserve stable scrollbar gutters on both inline edges of the scrollable detail body. - Add deterministic coverage for modal, pop-out, and embedded detail inset symmetry. - Publish a patch changeset for the layout correction. Files changed: .changeset/fn-8630-task-detail-right-padding.md | 7 + .../__tests__/task-detail-inset-symmetry.test.ts | 222 +++++++++++++++++++++ .../dashboard/app/components/TaskDetailModal.css | 15 ++ 3 files changed, 244 insertions(+) Fusion-Task-Id: FN-8630 Fusion-Task-Lineage: 9fd0c2bd-3370-4f86-ad12-9f06e5172c5b Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
24ef266e48 |
FN-8628: add Factory Dark dashboard theme
Add a low-light industrial dashboard color theme with first-paint support and release documentation. - Register Factory Dark across persisted theme types, selector metadata, and desktop/dashboard bootstrap validators. - Define dark and light Factory Dark tokens, swatches, and selector styling. - Cover theme registration, tokens, bootstrap behavior, and UI theme-option counts. - Add a minor @runfusion/fusion changeset and document the theme. Files changed: .changeset/fn-8628-factory-dark-theme.md | 7 ++ docs/dashboard-guide.md | 3 +- docs/settings-reference.md | 2 +- packages/core/src/types/execution-and-ui.ts | 2 + .../app/__tests__/factory-dark-theme.test.ts | 106 +++++++++++++++++++++ .../dashboard/app/components/ThemeSelector.css | 14 +++ .../components/__tests__/ThemeDropdown.test.tsx | 2 +- .../components/__tests__/ThemeSelector.test.tsx | 2 +- .../__tests__/CommandCenterControls.test.tsx | 2 +- packages/dashboard/app/components/themeOptions.ts | 1 + packages/dashboard/app/index.html | 2 +- packages/dashboard/app/public/theme-data.css | 86 ++++++++++++++++- packages/desktop/src/renderer/index.html | 1 + 13 files changed, 223 insertions(+), 7 deletions(-) Fusion-Task-Id: FN-8628 Fusion-Task-Lineage: 6f3c7cd9-0130-482d-8aa8-ca47d48b134f Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
efd8454b6c |
fix(planning): the add-comment trigger sat below the fold — the sheet outgrew its floating host (#3242)
Fixes the red `main`: `planning-browser-e2e.test.ts > places the sole
contextual comment trigger by viewport in embedded and modal Planning`.
**Unclaimed and not a flake.** `check-file-claimed` reported UNCLAIMED,
it is not in the quarantine ledger, it reproduced locally and
deterministically, and it failed identically across three consecutive
Full Suite runs. Quarantine would have been the wrong instrument — that
rule is for flakes, and appeasing a consistent failure buries a real
regression.
## Root cause
`PlanningModeModal.css` sizes dialog Planning as a **full-viewport
sheet**:
```css
.planning-modal:not(.planning-modal--embedded) {
height: 100dvh;
min-height: 100dvh; /* ← beats max-height: 100% */
max-height: 100%;
}
```
That was correct until the modal branch moved **inside
`FloatingWindow`** (`FNXC:ModalTouchGeometry 2026-07-26-14:10`). The
floating host's body is shorter than the viewport — it sits below a
title bar — so the rule now asks the sheet to be *taller than the box
containing it*. `min-height` wins over `max-height`, so the sheet cannot
shrink to its host and overflows.
Measured by walking the ancestor chain at 768×900:
```
BUTTON.btn top=928 h=36 ← 28px past the fold
DIV.planning-actions top=919 h=101
DIV.modal h=900 ← forced to full viewport height
DIV.floating-window__body h=763 sh=900 ← host is 763 tall, content is 900
DIV.floating-window h=765
```
The "Add comment to selection" control needed a scroll to reach —
exactly what the placement case exists to prevent.
## The fix
A scoped override under `.floating-window`, rather than editing the
sheet rule, so Planning rendered **outside** a floating host keeps its
full-viewport sizing:
```css
.floating-window .planning-modal:not(.planning-modal--embedded) {
height: 100%;
min-height: 0;
}
```
## Surface enumeration
Embedded Planning was **never affected** — it is excluded from the sheet
rule, and all four embedded viewports passed throughout. The failure was
modal-only, at every modal viewport (768, 769, 1024, 1280 — it fails
fast at the first).
**No new test.** The existing placement case already asserts this
invariant across **4 viewports × 2 presentations = 8 combinations**,
which is the surface enumeration for this affordance. It was red; it is
now green. Adding a narrower repro-only test would be the anti-pattern
the Fix-the-Invariant rule names.
## A disproven hypothesis, recorded
`min-height: 0` on `.planning-plan-review > .planning-plan-pane` — the
canonical flex-overflow fix, and a pattern used 10+ times in this very
file — **does not fix it**. Measured, not assumed. The overflow is one
level up, at the sheet/host boundary. Noted so the next reader does not
repeat the experiment.
## Verification
```
fix applied Tests 5 passed (5)
fix reverted Tests 1 failed | 4 passed (5) ← the test genuinely holds this fix
fix restored Tests 5 passed (5)
```
Neighbours green: **57 tests across 8 suites** (mobile
footer/bottom-space/pan-containment, terminal keyboard layout,
task-detail tablet width, mission planning modals mobile, mobile
planning input font size, task-detail floating geometry) plus **9**
planning e2e.
Changeset included (`patch`, category `fix`) — this is user-visible
dashboard behaviour.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f31a716a2a |
FN-8629: prevent false Grok usage percentages
Prevent omitted Grok billing percentages from being displayed as fully consumed credits. - Require a finite CLI-supplied credit usage percentage before creating a billing window. - Cover omitted, zero, invalid, and non-weekly Grok billing responses. - Add a patch changeset for the corrected usage display. Files changed: .changeset/fn-8629-grok-usage-percent.md | 7 +++ packages/dashboard/src/__tests__/usage.test.ts | 71 ++++++++++++++++++++++++-- packages/dashboard/src/usage.ts | 13 ++--- 3 files changed, 77 insertions(+), 14 deletions(-) Fusion-Task-Id: FN-8629 Fusion-Task-Lineage: b5b7c83b-e34f-43d9-a31a-d1fd769c4eb8 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
56e16d9dea |
test(cli): pin the board glyph's terminal-lane resolve (extract seam + pin) (#3238)
## What Pins the CLI board glyph's terminal-lane resolve — **the last flagged site in the repo-wide resolver audit.** Two commits: a behaviour-preserving extraction, then the test. ## I was wrong to flag this as unpinnable In #3236 I recorded this site as not pinnable, reasoning that *"extracting a pure helper and testing it would look like coverage and would not be."* That is true of a helper that **receives** the lane set — such a test passes with the resolve blinded, which is exactly the `reads.ts` trap the audit note records. It is **not** true of one that **resolves** it. Building `resolveReliabilityLanes` in #3237 made the distinction obvious: the seam has to contain the resolve, and then blinding fails a test of it. So the flag was too broad, and correcting it closes the site rather than leaving a permanent excuse. That is the same failure mode I corrected in someone else's note earlier today — a caution that hardens into a reason not to look. ## Measured ``` converted: Tests 5 passed (5) blinded: Tests 2 failed | 3 passed (5) ``` The two failures are the **renamed complete** and **renamed archive** lanes. The three survivors are the default-vocabulary control, the active-lane negative, and the degrade path — all of which should survive. ``` task-list-board-columns + bin: 82 passed typecheck clean; lint clean; fnxc-future-dates: none added ``` ## Why the sibling file did not cover it `task-list-board-columns.test.ts` pins `boardColumnsForDisplay`, which decides **which** lanes print. That function takes no lane set, so it cannot fail when this resolve is blinded — and its own header says so honestly. Two tests about the same command, one of which cannot see the other's bug. ## What breaks without the conversion On a board whose complete lane is `shipped`, a finished lane renders `●` — the same glyph as active work. The board says work is in flight when it shipped. Cosmetic next to the blank-board bug this area already fixed, but wrong in the direction an operator reads at a glance. ## Also pinned Two contracts the surrounding comments assert but nothing tested: - **Cards come from the TASKS, not a resolved IR** — a card must never depend on resolution succeeding to be *visible*. Asserted with an unreadable workflow list. - **A failed resolve degrades to the legacy pair**, with an unresolved custom lane rendering as active — the documented fail-open direction. Plus the paired negative: an ACTIVE lane keeps the active glyph under both vocabularies, so widening the terminal set cannot mark the whole board finished. ## Audit complete Every `resolveProjectColumnsForRoles` call site in the repository — `engine`, `core`, `dashboard`, `cli` — has now been blinded individually, and every uncovered one is either pinned or has a recorded reason it cannot be. Nothing is left flagged. |
||
|
|
0698ce6f9c |
test(dashboard): restore the missing api mock export in ResearchView tests (#3239)
## What
Partial fix for **red main**. Test-only.
`ResearchView.test.tsx` has **4 failing tests on main**; 2 fail with:
```
No "fetchBoardWorkflows" export is defined on the "../../api" mock
```
The `vi.mock("../../api")` factory **replaces the whole module**, so
every import anywhere in the rendered tree must appear in it.
`fetchBoardWorkflows` reached this file *indirectly* — the task modals
ResearchView opens import it — so adding that export to product code
broke four cases that have nothing to do with board workflows.
Stubbed with the flag-OFF payload the server sends when multi-lane
boards are disabled, which is the shape these cases already assume.
## Measured
```
before: Tests 4 failed | 23 passed (27)
after: Tests 2 failed | 25 passed (27)
lint clean; fnxc-future-dates: none added
```
## The remaining 2 are a different cause and are NOT fixed here
They fail with `Number of calls: 0` — the enrich-task and create-task
actions never fire. That is a UI-wiring question, not mock completeness.
**I checked that my stub is not responsible**, rather than assuming:
re-running with `flagEnabled: true` and a populated workflow list
produces the *same* 2 failures, so the payload shape does not gate those
affordances. Left for whoever owns that surface.
## How this was found
While establishing a clean baseline for the resolver audit. That sweep
also reported `lazy-loaded-views-docs.test.ts` red — **it now passes**,
fixed by another worker between my measurement and this PR, which is why
the count here is 3 files rather than the 4 I reported in #3236.
Still red on main, untouched by this PR:
- `src/__tests__/planning-browser-e2e.test.ts` — `expected {
totalButtons: 1, …(7) } to match object { totalButtons: 1, …(6) }`; an
assertion shape gained a field.
- `src/__tests__/register-model-routes-kimi-k3-supplemental.test.ts` —
`Test timed out in 15000ms`. Per the standing rule a timeout with no
corresponding bug in the change is a **quarantine candidate**, not
something to appease with a longer timeout; I am not quarantining it
unilaterally since it is not my subsystem, but flagging it as the shape
that rule describes.
|
||
|
|
476c5c360c |
test(dashboard): pin the Reliability endpoint's three lane reads (extract seam + pin) (#3237)
## What
Pins the **Reliability endpoint's three lane reads** — the last
uncovered resolver cluster the repo-wide audit found.
Two commits, deliberately separate:
1. **refactor** — extract the three resolves behind
`resolveReliabilityLanes(store)`. Behaviour-preserving, no test changes.
2. **test** — pin all three through that seam.
## Why a seam was needed
The three resolves lived inline in the `/api/health/reliability` route
closure. Blinding any of them left the **entire dashboard suite green —
21,582 tests** — and the only way to reach them was booting
`createServer` behind a mock-the-world shell the slow-test rule forbids.
**And the obvious test would not have helped.**
`reliability-metrics.test.ts` exercises `countEntriesInto`,
`countBouncesOut` and `inReviewDurationMetrics` with lane sets **passed
in by hand**. That proves the collaborators honour a resolved set; it
says nothing about whether the caller passes one. *A unit test of the
collaborator can never fail when the caller's resolve is blinded* — the
same trap the audit note records for `reads.ts`, where a suite written
for the exact conversion still could not see it.
The seam is the caller. It resolves, so blinding a resolve fails a test
of it.
## Measured — each blind fails exactly its own case
| blinded | fails |
|---|---|
| `REVIEW_ROLES` | "resolves the board's OWN review lane" |
| `["countsTowardWip"]` | "resolves the board's OWN wip lane" |
| `["complete"]` | "resolves the board's OWN complete lane" |
```
converted: Tests 6 passed (6)
each blind: Tests 1 failed (its own case only)
reliability-metrics.test.ts + this file: 28 passed
typecheck clean; lint clean; fnxc-future-dates: none added
```
That isolation is the point: **three resolves in one function invite a
copy-paste that hands the same set to all three**, and every positive
assertion would still pass. There is a paired negative asserting each
renamed lane appears in *its* bucket and nowhere else — without it the
duration metric could silently measure review → review.
Also pinned: the degrade path. An unreadable workflow list must not fail
the endpoint, so the legacy ids still answer.
## What breaks without the conversion
On a board that renames either lane, every underlying query returns `{}`
— so `tasksEnteredInReview` and `tasksBouncedToInProgress` are zero for
every day, and `inReviewFailureRate7d` divides one zero by another and
reports a **healthy** rate. It produces a NUMBER, not an error, and the
number says everything is fine. An operator reading 0% review failures
beside a populated audit list has no reason to suspect the metric is
blind.
## The one observable difference in the refactor, stated not buried
The complete-lane read moves from *after* the counting `Promise.all`
into the same phase as the review/wip pair. These are pure reads of
workflow definitions — no writes, no ordering dependency — so the
resolved values are identical; only the concurrency shape changes (three
parallel reads instead of two-then-one). Flagging it because
"behaviour-preserving" should be a claim someone can check, not an
assertion.
## Audit status
With this, **3 of the 4 flagged sites are closed**. Remaining:
`cli/commands/task.ts:660`, where the glyph decision is inline in
`runTaskList` and the same seam argument applies — but its sibling test
file already documents that driving that function needs the forbidden
shell, and extracting a helper there would produce a test that *looks*
like coverage while leaving the resolve unpinned. Left flagged rather
than faked.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Reliability health metrics now recognize configured review,
work-in-progress, and completion lanes, including renamed workflow
lanes.
- **Bug Fixes**
- Improved fallback behavior when workflow definitions are unavailable,
preserving compatibility with legacy lane configurations.
- Ensured lane resolution remains isolated by role for more accurate
reliability metrics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
|
||
|
|
f0a13745b2 |
fix(test): lazy-views doc parser stops at any heading — six phantom views came from the H2 that follows
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fb8d37e30d |
test(core): pin the last two uncovered lane reads (mission archive, lineage gate) (#3235)
## What Pins the **last two uncovered lane reads** in `packages/core`. Test-only. This closes the per-site core audit. | site | what it decides | |---|---| | `async-mission-store.ts:1179` | is an ARCHIVED card valid terminal evidence for mission repair? | | `task-id-integrity.ts:502` | does an archived child still count as a LIVE lineage child? | ## Measured ``` mission-store: 39 passed clean; 1 failed | 38 passed blinded lineage: 3 passed clean; 1 failed | 2 passed blinded lint clean; fnxc-future-dates: none added; census unchanged ``` Both blinds confirmed applied with `git diff --stat` before each run. ## The third adjacent-pair split `:1179` is the **archived** half of a pair whose **complete** half (`:1178`, *one line above*) was already covered by a test in the same file, written for exactly this concern. Terminal evidence is "done OR supported archived state," so an archived card is equally valid repair evidence — but on a board whose archive lane is `vaulted` the archived half could not see it, and reconciliation threw `TASK_NOT_TERMINAL` for a card that was genuinely filed away. Same refusal the covered case fixed, reached through the other door. That is now the third confirmed instance in core (after `team-analytics` in #3227 and the scheduler pair earlier). **Being adjacent to a covered resolver is not coverage**, and it is the most reliable place to look. ## What breaks without the lineage read An archived child is filed away, not live, so it must not hold the delete gate shut. Renamed, it still counted as live and `TaskHasLineageChildrenError` blocked the parent's delete **forever** — the operator archived the child *precisely* to clear the way, and the gate could not see that they had. ## A fixture detail I got wrong first My first mission fixture created a live card in a `vaulted` column and failed with `deleted or archived without a valid retained tombstone and archive snapshot` — nothing to do with the lane read. The `archived` verdict requires **all three** of `deletedAt !== null`, an archive-snapshot row, and `isArchived(column)`. A live card merely sitting in an archive-trait column is `invalid-deleted`, not `archived`. The test now archives for real and *then* renames the recorded lane, which isolates the third condition — the only one under test. Recorded in the file so the next person does not re-derive it. ## Paired positives Both files pin the complement: a WORKING child still counts as live. Recognising the renamed archive lane must not degrade into "no child is ever live" — that would silently **disable** the lineage gate and let a parent be deleted out from under real descendants, which is worse than the bug being fixed. ## Core audit complete **14 sites blinded individually: 9 already covered, 5 uncovered, all 5 now pinned** (#3233, #3234, this PR). Every `resolveProjectColumnsForRoles` call site in `packages/engine` and `packages/core` has now been blinded. Remaining unaudited: `dashboard` (2 files) and `cli` (1) — I claim nothing about those. |
||
|
|
623581837a |
fix(engine): mock provider sends 0-based steps — test mode full-task runs complete again (#3231)
Found by a live browser E2E of the coding workflow in test mode: every
scripted full-task run failed at `steps#0:step-execute` with `Step 4 out
of range (task has 4 steps)`, rebounding through recovery forever.
**Root cause:** `fn_task_update.step` has been **0-based since FN-6607**
(executor.ts FNXC:StepNumbering — the old `step - 1` conversion made
Step 0 impossible to mark). `mock-provider.ts` still sent `index + 1`,
so test mode marked steps 1..N instead of 0..N-1: Step 0 (Preflight)
never completed and step N threw out-of-range. Test mode's full-task
path has been broken since June.
**Also fixes the test that pinned the bug:** `mock-provider.test.ts`
expected `{ step: 1 }` for a fixture whose first unfinished step is
index 0 — the expectation encoded the 1-based off-by-one.
Verified: 12/12 mock-provider tests; the live E2E instance completes the
task after this patch (see follow-up screenshot in the session).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
7f3acf8929 |
test(core): pin both create-time duplicate guards' lane exclusions (#3234)
## What Pins **both** create-time duplicate guards in `branch-and-pr-entities.ts`. Test-only. | site | method | excludes | |---|---|---| | `:445` | `findRecentTasksByContentFingerprint` | ARCHIVED (unless `includeArchived`) | | `:484` | `findRecentTasksBySourceParentTaskId` | COMPLETE and ARCHIVED | Blinding either back to its literals left the entire 16-file lane-detector set green. **No test in `packages/core` reaches either method.** ## Measured ``` converted: Tests 8 passed (8) blinded :445 Tests 1 failed | 7 passed (8) <- only the fingerprint case blinded :484 Tests 2 failed | 6 passed (8) <- only the sibling cases lint clean; fnxc-future-dates: none added; census unchanged ``` **Each blind fails exactly its own cases.** That matters: it proves the two resolvers are pinned *independently*, rather than one broad test appearing to cover both. Blinding `:445` leaves every sibling case green and vice versa — so neither is riding on the other's coverage. ## They fail in opposite directions This is why both belong in one file: - **Fingerprint guard** — a renamed board leaves archived cards in the candidate set, so filing a new task is **refused as a duplicate** of one the operator already archived. The create is blocked and the thing blocking it is invisible. - **Sibling guard** — a renamed board leaves finished siblings in the "recent live siblings" set, so completed work keeps counting as active. One over-includes into a *refusal*, the other over-includes into *phantom activity*. Neither raises an error. ## Positives pinned too A LIVE fingerprint match is still a duplicate candidate; `includeArchived: true` opts the renamed archived lane back in; a WORKING sibling is still live. Excluding the finished lanes must not degrade into excluding everything, or the guards stop guarding — the failure mode a lane-widening change invites. ## A fixture detail that would have made this vacuous Both queries cut off at `Date.now() - windowMs`, with `windowMs` capped at 24h. The sibling harness I copied from seeds a **fixed past timestamp**, which falls outside that window — every case would then pass on an empty result, including the ones that are supposed to fail under blinding. Fixtures are seeded at current time instead, and the reason is recorded in the file so nobody "tidies" it back to a frozen date. ## Progress 3 of the 5 uncovered core sites are now pinned (`store.ts:1135` in #3233, these two here). Remaining and unclaimed: `async-mission-store.ts:1179` and `task-id-integrity.ts:502`. |
||
|
|
3f06d7201b |
test(engine): assert the evaluator's exact archived-lane set, not toContain (#3232)
## What Follow-up to the #3224 review comment *"reject legacy archived identifiers for renamed workflows."* Test-only. That comment had two halves. **The half it got wrong** is already answered on main: asserting an exact single-column set `["vaulted"]` *fails*, because `resolveProjectColumnsForRoles` unions `LEGACY_COLUMN_IDS_BY_ROLE` in as a documented floor — so a row whose workflow cannot be resolved still classifies. Pinning `["vaulted"]` would encode the opposite of the design. **The half it got right was never addressed.** `toContain` also passes when the set grows a lane nobody intended, and an over-broad archived set silently classifies *live* rows as archived. So the reviewer's worry was legitimate even though the proposed fix was not. The exact set is assertable — it just is not the one the review proposed: | board | resolved set | |---|---| | default | `["archived"]` | | renamed | `["archived", "vaulted"]` — legacy floor + the board's own lane | `[...].sort()` in the helper makes ordering stable, so these pin the resolver's whole answer rather than a substring of it. ## Proven to catch what `toContain` missed Giving the fixture a second `archived`-trait column fails both new assertions: ``` AssertionError: expected [ 'archived', 'cold-store' ] to deeply equal [ 'archived' ] AssertionError: expected [ 'archived', 'cold-store', 'vaulted' ] to deeply equal [ 'archived', 'vaulted' ] Tests 2 failed | 1 passed (3) ``` The previous `toContain` assertions pass unchanged against that same spurious lane. That is the whole justification for this PR — without the injection test it would be a stylistic preference. ``` clean: Tests 3 passed (3) spurious lane: Tests 2 failed | 1 passed (3) lint clean ``` ## Note I said in the review thread I would tighten this, so this closes that loop. I also checked before editing whether another worker had already done it — main already carries the FNXC tag and the legacy-floor explanation from the same review round, so this PR adds only the part still missing rather than redoing settled work. |
||
|
|
2780a8ae7b |
test(core): pin the open-undo query's finished-lane exclusion (#3233)
## What Pins the **open-undo query's finished-lane exclusion** in `packages/core/src/store.ts`. Test-only. `findOpenRevertTaskForSource` answers *"is there an OPEN undo task for this source?"* — the question behind the dashboard's Undo affordance. It answers by **excluding the finished lanes**, so a prior undo that already landed does not keep rendering as open. Blinding that exclusion back to `ne(column,"archived"), ne(column,"done")` left the entire 16-file lane-detector set green. **No test in `packages/core` reaches this method at all.** The dashboard-side twin (`taskRevert.ts`, #3129) is tested; the store-side query behind it was not. ## Measured | | default (control) | renamed complete | renamed archived | working lane | |---|---|---|---|---| | converted | pass | pass | pass | pass | | blinded to `["done","archived"]` | pass | **FAIL** | **FAIL** | pass | ``` converted: Test Files 1 passed (1) / Tests 4 passed (4) blinded: Test Files 1 failed (1) / Tests 2 failed | 2 passed (4) lint clean; fnxc-future-dates: none added; census unchanged ``` Blind confirmed applied with `git diff --stat` before the run. ## What breaks without it On a board whose complete lane is `shipped`, neither literal matches, so a **done** undo task is never excluded and the query keeps returning it. The card shows an undo already in flight *forever*, and the real affordance is unreachable. Nothing errors — the button is just permanently wrong, which is why it went unnoticed. ## Includes the paired positive An undo still in a **working** lane IS reported as open. Excluding the finished lanes must not degrade into excluding everything, or the affordance breaks in the other direction and no undo is ever reported in flight. Both new failing cases are renamed-lane cases; both survivors are cases that should survive. ## Where this came from Per-site blinding of all 14 remaining `resolveProjectColumnsForRoles` call sites in `core`, run against a 16-file detector set. **9 covered, 5 uncovered:** | site | verdict | |---|---| | `store.ts:1135` | **uncovered** → pinned here | | `async-mission-store.ts:1179` (archived) | **uncovered** — its neighbour `:1178` (complete) is covered | | `branch-and-pr-entities.ts:445` | **uncovered** | | `branch-and-pr-entities.ts:484` | **uncovered** | | `task-id-integrity.ts:502` | **uncovered** | | `reads.ts` ×3, analytics ×3, `eval-automation`, `task-artifacts-ops`, `async-mission-store:1178` | covered | The first run of that probe was **invalid and I nearly published it**: it reported all 14 sites "COVERED" with *zero failing tests*. zsh does not word-split unquoted parameter expansions, so `vitest run $DET` passed 16 paths as one argument and vitest exited 1 with "No test files found" — which my script read as a failing test. The re-run treats that string as `INVALID` rather than a result. Third time this session a wrong reading came from test *selection* rather than from blinding. ## Flagged, not guessed The four remaining uncovered sites are named above rather than quietly left; `async-mission-store` shows the same adjacent-pair split as `team-analytics` in #3227, which is now the third confirmed instance of that shape. |
||
|
|
dd09e57511 |
test(core): fix red main — assert the delete re-home against the resolver, not "triage" (#3229)
## What **Fixes a red main.** `workflow-reconciliation-production-shape.pg.test.ts` has been failing with `expected 'todo' to be 'triage'`. Test-only. Found while establishing a clean baseline for an unrelated coverage audit — my tree was clean at `origin/main` (`76c73238a0`), so this is not something I introduced. It is in the non-blocking suite, which is why it has stayed red. ## It is not a regression — the test was the stale half The delete path was deliberately fixed to re-home occupants using `resolveEntryColumnId(resolveDefaultWorkflowIr())` instead of `BUILTIN_CODING_WORKFLOW_IR`. This assertion was not updated with it. The two IRs are **not the same board**: | IR | entry column | |---|---| | `BUILTIN_CODING_WORKFLOW_IR` (`builtin:legacy-coding`) | `triage` | | `resolveDefaultWorkflowIr()` (the catalog default) | `todo` | Re-homing into `triage` put cards in a column the default board never declares. It slipped past `moveTask`'s undeclared-target guard **only because `triage` is a legacy id** and the recovery-rehome path exempts those — so the guard that exists to stop exactly this could not see it. So `todo` is the correct behaviour and the literal `"triage"` was what needed fixing. ## Why it asserts a resolver rather than `"todo"` Swapping one hardcoded id for another would be the identical trap one rename later — the same class of defect this whole program exists to remove. The expectation now derives from **the same two functions the product path calls**, so it cannot drift out of sync with them again. I also added the complement: the card must genuinely have **left** the vanished column, not merely match whatever a resolver returns. Without it, a resolver that started returning `custom-hold` would pass. ## Proven not appeasement Reverting the product line to the legacy IR — the original defect — fails this test: ``` AssertionError: expected 'triage' to be 'todo' Test Files 1 failed (1) / Tests 1 failed | 6 passed (7) ``` That is the check that matters for a test edit that turns a red green. It fails on the defect it describes. ## Measured ``` before: Tests 1 failed | 6 passed (7) after: Tests 7 passed (7) 16-file detector set: 181 passed (16 files) [was 1 failed | 180 passed] lint clean ``` ## Note on the reading I got this wrong twice before getting it right, and the record is worth having. My first read was "the test is stale, `triage` was merged away." My second was "the builtin IR still declares `triage`, so the *behaviour* is the defect" — which the IR file superficially supports. Only the third reading, of the FNXC note at the fix site, showed the file I was reading is the **legacy** IR and not the default one. Two of those three readings would have produced a confidently wrong PR; the deciding evidence was the comment the fixing author left at the call site, which is a good argument for writing them. |
||
|
|
d4c25384ae |
test(engine): record what the zero-backlog early return stops testing (#3228)
## A green case that stopped testing what its name says #3226 fixed a red `main` correctly: `fileWithGuards()` now returns `null` at zero backlog, and with nothing to inflate there is no rise to manufacture. Asserting `totals.column === 0` and returning is the honest response. What went unrecorded is the cost. At zero, these two cases: - *"exits 0 and REWRITES the baseline under `--update-baseline`, even when the count rose"* - *"exits 1 and LEAVES the baseline alone on a rise without `--update-baseline`"* no longer exercise the CLI's ordering or exit codes. They assert the backlog is empty and return. **If the write-before-exit ordering regressed — the exact bug those cases were written for — both would still pass.** That matters more than it would elsewhere, because **zero is not a state to wait out.** It is this program's terminal state: the backlog went 126 → 0 and is meant to stay there. So the vacuity is permanent, not transitional. This file already legislates against precisely this, two hundred lines down: > `/* Anti-vacuity: an empty exclusion list would make the assertion below trivially true. */` ## What this PR does Adds a comment on `fileWithGuards()` recording (a) which cases go vacuous at zero and why, (b) that zero is terminal so it will not resolve itself, and (c) the durable fix. **Comment only. No behaviour change — suite stays 53/53.** ## The durable fix, recorded rather than done Point the scan at a synthetic tree so the fixture stops being a function of the real backlog — the same seam `FUSION_CENSUS_BASELINE_PATH` already provides for the baseline, applied to the file list. It needs one CLI correction to work, and that is a genuine bug in my own code regardless of this suite: `triageFindings` and the sync-resolver check read files via `join(REPO_ROOT, f.file)`, where `REPO_ROOT` is derived from the **script's** location. An overridden file list therefore changes which paths are *listed* without changing where they are *read from*, and every read misses with `ENOENT`. I verified that approach works (a three-file fixture yields a stable `1 backlog / 1 deliberate / 1 sync-resolved`) and got two of the six failing cases green with it, then stopped rather than keep guessing in a file being actively revised. Left as a comment so whoever takes it does not re-derive the diagnosis. ## Why this is worth a PR at all This program's recurring failure is instruments that report green while measuring nothing — an inert conversion the census scored as a win, a ratchet wired to nothing, a gate that could not fail. A test asserting `0 === 0` under a name promising ordering coverage is the same shape at the test layer. The suite cannot be fixed in this PR without re-opening work someone else owns, but it can at least stop being silent about it. ## Census before / after ``` before: COLUMN guards (the backlog): 0 after: COLUMN guards (the backlog): 0 ``` ## Verification `test:gate` exit 0 · `lifecycle-column-census.test.ts` **53 passed** · `fnxc-future-dates`, `lifecycle-columns`, `inert-sync-lanes`, `quarantine-ledger`, `inert-flag-seams`, `lane-wiring`, `sql-column-literals` — all exit 0. |
||
|
|
6646c1b95d |
docs: regenerate synced skill tool tables (unblocks all four full-suite shards)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
32edd1421a |
test(core): pin team analytics' in-flight lane read — the other half of the pair (#3227)
## What Pins `aggregateTeamAnalytics`' **in-flight lane read** (`activeLanes`) — the unpinned half of an adjacent resolver pair. Test-only, no product change. `completeLanes` and `activeLanes` are declared **two lines apart**. Every existing case in this file asserts only `totals.tasksCompleted`, so the in-flight query `activeLanes` feeds was never observed. Measured on main: | blinded resolver | result | |---|---| | `completeLanes` → `["done"]` | **FAILS** the file (1 failed / 3 passed) — pinned | | `activeLanes` → `["in-progress","in-review"]` | **entirely GREEN** (4 passed) — unpinned | One resolver held, its neighbour not, in a file named `team-analytics-renamed-lanes`. This is the half-covered-pair shape the program keeps finding; being *next to* a covered resolver is not coverage. ## How I found it Rather than blind core's 17 files one at a time, I made `resolveProjectColumnsForRoles` itself return legacy-only — its own documented degrade path — which blinds **all 116 call sites in one edit**. The full core suite then reported **24 failures across 16 files** out of 4,987 tests, which maps the covered areas in a single run: the analytics renamed-lane pg tests, the archived-lane family, eval-automation, mission-store, and the resolver's own tests. That global probe finds *areas* that are covered, not *resolvers* — so the pairs still needed individual blinding, which is what surfaced this one. `workflow-analytics.ts` has the identical two-resolver shape and **both halves are covered**; the gap is specific to `team-analytics.ts`. ## Two things have to be right, and both are now asserted 1. **The SQL must ASK for the board's real wip lane** — `activeLanes`. 2. **`buildTeamAnalytics` must RECOGNISE the row it gets back.** It classifies via `isWipColumnRole(query.columnFlagsByName?.get(name), name)`, which **without flags falls back to `name === "in-progress"`** and drops a renamed lane it already fetched. So supplying `columnFlagsByName` is part of the caller contract, not test scaffolding: **widening the query alone would still report zero.** A test that only widened the first half would pass while the feature stayed broken. ## Measured ``` converted: Test Files 1 passed (1) / Tests 7 passed (7) blinded activeLanes: Test Files 1 failed (1) / Tests 2 failed | 5 passed (7) lint clean; fnxc-future-dates: none added; census unchanged ``` Blind confirmed applied with `git diff --stat` before each run. ## What breaks without it A per-agent `tasksInProgress: 0` sitting beside a nonzero completed count and real token spend — an agent that looks idle while it is working. Same wrong-but-plausible shape this file's own header describes: nothing errors, and a plausible-looking number is the least likely defect for anyone to file. ## Flagged, not guessed - `packages/core` is not my package; this is additive tests only. I raised the same note on #3225. - The global probe shows core has **substantial** renamed-lane coverage — it is not the uniformly-unpinned surface I implied when I first reported 17 unaudited files. Correcting that here rather than leaving the stronger claim standing. - Still unblinded individually: the resolver pairs in `async-mission-store.ts` (1178/1179) and the archived reads in `task-store/reads.ts` (396/615/793). Their *files* fail under the global blind, so something covers each area — but that is not per-resolver evidence, and I am not claiming it is. |
||
|
|
12c4ab5a6e |
test(engine): pin the evaluator's archived-lane read — the service had no test at all (#3224)
## What Pins the **evaluator's archived-lane read**. Test-only — no product change. `HybridEvaluatorService.evaluateTask` resolves the board's archived lanes and hands them to `collectDeterministicSignals`, which decides which of a task's related rows count as archived when scoring a run. **The service had no test anywhere in the repo.** Four test files import the module; none construct or exercise it. So this conversion was unobservable for the simplest possible reason — *nothing ran the code*. That is a different failure from the ones this audit has been finding (harnesses that run the code but cannot see the difference), and worth distinguishing: no amount of fixture care helps when the entry point is never called. ## Measured | | default (control) | renamed | differential | |---|---|---|---| | converted | pass | pass | pass | | blinded to `["archived"]` | pass | **FAIL** | **FAIL** | ``` converted: Test Files 1 passed (1) / Tests 3 passed (3) blinded: Test Files 1 failed (1) / Tests 2 failed | 1 passed (3) the 4 files importing evaluator.ts, plus this one: 5 files/76 tests, all green lint clean; fnxc-future-dates: none added; census unchanged ``` Per the rule I documented in #3223, the blind was confirmed applied with `git diff --stat` **before** the run rather than trusting the tool's exit code. ## What breaks without it On a board whose archived lane is `vaulted`, the evaluator hands the collector the legacy `{archived}` set. Rows resting in `vaulted` are not recognised as archived, and the deterministic half of every evaluation score is computed from a wrong picture of the task's history. **Nothing errors, the run completes, the number is just wrong** — which is why it survived unnoticed. ## Pinned without faking a provider response The assertion is about what the collector *receives*, which is decided before any model call. `collectDeterministicSignals` is mocked to record its arguments and throw a sentinel; the test asserts the resolved lane set and stops. This is deliberate over the obvious alternative of feeding `runPrompt` a canned AI payload: `deps.runPrompt` is injectable so either approach is offline, but a canned payload has to satisfy `parseAiResponse` and every `EVAL_SCORE_CATEGORIES` entry, and would silently rot into a maintenance burden on a test whose subject is one `Set`. Reversible if someone later wants full end-to-end evaluator coverage — that is a different test, not this one. ## Completes the engine audit With this, every `resolveProjectColumnsForRoles` call site in `packages/engine` has been blinded: | file | resolvers | result | |---|---|---| | `self-healing.ts` | 64 | 21 pinned, 1 recorded inert by construction, remainder mapped | | `executor.ts` | 2 | both already covered | | `scheduler.ts` | 1 | uncovered → pinned (#3219, merged) | | `triage.ts` | 1 | uncovered → pinned (#3221) | | `restart-recovery-coordinator.ts` | 1 | already covered | | `notification-service.ts` | 1 | already covered | | `evaluator.ts` | 1 | uncovered → pinned (this PR) | `project-engine.ts:5154` takes `roles` as a **parameter**, so it is a generic wrapper with no fixed role set to blind — flagged rather than guessed at; its callers are where the question belongs. **`packages/core`'s 17 files remain entirely unaudited** and I am claiming nothing about them. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added regression coverage to verify reliable resolution of archived workflow lanes. * Covered both the default archived-lane name and custom renamed configurations. * Confirmed compatibility with legacy archived-lane naming behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
76c73238a0 |
test(core): pin the engine-downtime shift's wip read (428 tests could not see it) (#3225)
## What Pins the **engine-downtime timing shift's wip read** in `packages/core/src/store.ts`. Test-only — no product change. First audited site in `core`. `reconcileActiveTimingForEngineDowntime` (FN-7011/FN-7975) excludes proven stopped-engine wall-clock from a card's active time. It finds the cards to fix by querying the board's wip lane. **Blinding that read back to `["in-progress"]` left every test that touches the sweep green — 4 in this file plus 424 in the two engine files that exercise it, 428 in total.** ## Why 428 tests were blind to it The existing store double is 10 lines and contains **both** documented anti-patterns, either one sufficient on its own: 1. **`listTasks: vi.fn(async () => tasks)` ignores its `column` argument** — it returns the same rows whichever lane is requested. A fake that ignores its own filter cannot see a filter bug, which is exactly the bug this resolver exists to fix. 2. **No `listWorkflowDefinitions`** — `resolveProjectColumnsForRoles` then returns the legacy ids and nothing else (an intentional degrade in `project-lane-vocabulary.ts` so an unreadable workflow list cannot fail a sweep). The resolved set and the literal set were *equal by construction*. The new double fixes both and changes nothing else. **The existing cases keep the original double on purpose:** they are about heartbeat and threshold arithmetic, not lanes, and rewriting them would put unrelated churn in the same commit. ## Measured | | default (control) | renamed | differential | non-wip card | |---|---|---|---|---| | converted | pass | pass | pass | pass | | blinded to `["in-progress"]` | pass | **FAIL** | **FAIL** | pass | ``` converted: Test Files 1 passed (1) / Tests 8 passed (8) blinded: Test Files 1 failed (1) / Tests 2 failed | 6 passed (8) engine neighbours (project-engine-unpause-active-timing + self-healing): 424 tests, green and unchanged lint clean; fnxc-future-dates: none added; census unchanged ``` Blind confirmed applied with `git diff --stat` before each run, not inferred from the tool's exit code. ## What breaks without it On a board whose wip lane is `building`, the sweep queries `in-progress`, finds **no tasks**, and shifts no anchor. Every card silently absorbs the stopped-engine wall-clock the sweep exists to exclude. The reported active time is simply wrong and nothing fails to signal it — the same silent-wrong-number shape as the evaluator defect in #3224. ## Also covers the complement A held card *outside* the wip lane is **not** shifted. Widening a lane read is the kind of change that can quietly turn a targeted sweep into a board-wide rewrite; a card in `todo` has no stopped-engine time to exclude, and there is now a case saying so. ## Scope note `packages/core` is not my package. This is an additive test file with no product change, so collision risk is low, but I am flagging it rather than assuming: **16 of core's 17 files with resolver call sites remain unaudited** and I claim nothing about them. The audit method and its failure modes are documented in #3223 if core's owner wants to continue it. |
||
|
|
3c7dd9a803 |
test(engine): keep census regressions valid at zero backlog (#3226)
## Summary - keeps lifecycle-census end-to-end fixtures valid after the conversion backlog reaches zero - treats zero backlog as a real protected end state instead of requiring a remaining guard or claim target - preserves nonzero rise/claim assertions when guards remain ## Test plan - `corepack pnpm --filter @fusion/engine exec vitest run src/__tests__/lifecycle-column-census.test.ts --silent=passed-only --reporter=dot --project=engine-default` - `corepack pnpm --filter @fusion/engine typecheck` - verified the same suite at zero backlog on the aggregate runtime (53/53) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Improved lifecycle validation to handle empty backlogs without errors. * Added coverage for zero-item results and changing file sets. * Enhanced baseline and trend checks to report completed backlogs consistently. * Improved resilience when expected result entries are unavailable, ensuring validation completes cleanly with accurate zero counts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
719cf281fd |
test(engine): pin the startup stale-planning sweep's planner-lane read (#3221)
## What Pins the **startup stale-planning sweep's** planner-lane read in `triage.ts`. Test-only — no product change. A card can hold `status: "planning"` while triage specifies it in place. A crash or restart before planning completes leaves that status set, and a startup sweep clears it. If the sweep misses the card, it occupies a planning admission slot **permanently** and new triage work is never admitted. The lane read was converted from `resolvePlannerLanes(store, "")` — called with an **empty task id**, so it could never resolve a task and always answered with the default board — to a project-level `resolveProjectColumnsForRoles(store, ["intake", "hold"])`. **No test could observe that conversion.** All 25 existing triage test files stayed green (375/375) with the resolver neutered — under *both* blinds tried. ## Measured | | default (control) | legacy ids | renamed | differential | |---|---|---|---|---| | converted | pass | pass | pass | pass | | blinded to empty set | pass | pass | **FAIL** | **FAIL** | | blinded to legacy pair | pass | pass | **FAIL** | **FAIL** | ``` converted: Tests 4 passed (4) blinded to empty set: Tests 2 failed | 2 passed (4) blinded to legacy: Tests 2 failed | 2 passed (4) existing 25 files: green under BOTH blinds (only this new file fails) triage suite: 25 files/375 tests -> 26 files/379 tests, all green lint clean; fnxc-future-dates: none added; census unchanged ``` ## A prediction I got wrong, and what it changes The site is **seed-then-union**: ```ts const sweepColumns = [...new Set(["triage", "todo", ...projectPlannerColumns])]; ``` I expected blinding the resolver to its legacy pair `["triage","todo"]` to be a **no-op by construction** — the seed already contains both. It is not. That reasoning holds only for a *default* board; on a renamed board the resolver is the sole contributor of `drafting`, so the legacy blind drops it and the renamed cases fail. The corrected rule, which is narrower than the one I was carrying: **a seed-then-union site hides a defect only while every lane you assert on is already in the seed.** Assert on a lane that is not, and the union stops protecting it. The empty-set blind remains the stricter of the two because it also models a resolver returning nothing at all. Both are recorded in the test file so the next person does not re-derive it. ## A tooling failure worth naming My first triage measurement reported **375/375 green under blinding** — and was a lie. The blinding script had no mapping for `["intake","hold"]` and exited `2` **silently**; the `&&` short-circuited the check while a `;` let vitest run against **unmodified source**. A no-op blind produces a green run that is indistinguishable from real coverage. The script now prints what it substituted and where, fails loudly on an unmapped role or a missing variable, and I self-tested both directions (bogus var → rc 2 with a message; real var → rc 0 with the substitution echoed) before trusting any number above. This is the same standard I have been applying to other people's guards — *a guard that reports success without checking anything is worse than no guard* — and my own instrument failed it. ## Also pinned The **legacy half** of the union. `triage`/`todo` stay in the sweep even on a board whose workflow declares neither, because pre-U11 and Coding (Ideas) rows can rest there. Dropping them in favour of the resolved lanes alone would strand exactly those rows, so there is now a case asserting it. ## Flagged, not guessed - The FNXC stamp at `triage.ts:987` reads `2026-07-31-23:59`, hours ahead of the `date -u` clock. It is one of the 181 grandfathered stamps so the gate is green; I left it rather than widen this PR. |
||
|
|
9ab9822b8d |
test(engine): pin the deleted-blocker WIP-lane read, which no test could see (#3219)
## What Pins the deleted-blocker sweep's **WIP-lane read** in `scheduler.ts`. Test-only — no product change. When a task is soft-deleted, a `task:deleted` listener clears `blockedBy` on every dependent so the work can be scheduled again. It reads two lanes to find them: hold and WIP. The WIP read was already converted to `resolveProjectColumnsForRoles(store, ["countsTowardWip"])`. **Blinding that resolver back to the literal `["in-progress"]` left all 14 existing scheduler test files green — 145/145.** The conversion was load-bearing and unpinned. ## Why nothing could see it Two independent harness properties, **either one sufficient** to hide the defect: 1. **`resolveProjectColumnsForRoles` returns legacy ids and nothing else when the store has no `listWorkflowDefinitions`** — an intentional degrade in `project-lane-vocabulary.ts` so an unreadable workflow list cannot fail a sweep. The shared scheduler harness does not define it, so *the resolved set and the literal set were equal by construction* in every existing test. 2. **`listTasks` was mocked as `vi.fn(async () => tasks)`**, ignoring its `column` filter. A mock that returns every task regardless of the lane asked for cannot detect a wrong lane. This is worth naming because #1 is not a test bug — it is correct production behaviour that happens to erase the difference a test is trying to measure. A harness can satisfy a conversion's *shape* while making its *effect* unobservable. There is a test file named `scheduler-renamed-dependency-and-review-lanes.test.ts` covering dependency satisfaction, file-scope leases, base-branch stacking, PR hydration and mission completion on a renamed board. It does not reach this sweep, and its harness has both properties above. ## What breaks without the conversion On a board whose WIP lane is `building`, the literal read asks for `in-progress`, finds nothing, and the in-flight dependent is never reconciled. It keeps `blockedBy` pointing at a task that no longer exists — **permanently**, because the blocker can never be completed or re-deleted to trigger another sweep. Work stops with nothing to rescue it. ## Measured | | default (control) | renamed | differential | |---|---|---|---| | converted | pass | pass | pass | | blinded to `["in-progress"]` | pass | **FAIL** | **FAIL** | ``` converted: Test Files 1 passed (1) / Tests 3 passed (3) blinded: Test Files 1 failed (1) / Tests 2 failed | 1 passed (3) scheduler suite: 14 files/145 tests -> 15 files/148 tests, all green lint: clean fnxc-future-dates: none added ``` The default-vocabulary control passes in **both** columns by design: it is what keeps a generic break in this path from hiding in the renamed assertion. ## Census Unchanged — **2 / 148 deliberate**. `scheduler.ts` was already at 0. This PR adds coverage, not conversions; the census counts comparisons and would not have moved either way, which is the same blind spot that let #3078 merge green. ## Flagged, not guessed - The `resolveTaskParkedColumns` **hold** read one line above (`scheduler.ts:1402`) is the other half of this pair. My scenario places its dependent in the WIP lane, so it does not exercise the hold read and I am not claiming it is pinned. - The FNXC stamp at `scheduler.ts:1404` reads `2026-08-01-05:00`, which is future-dated. It is one of the 181 grandfathered stamps, so the gate is green and I left it alone rather than widen this PR. |
||
|
|
3146a745bf |
test(engine): cover two reporter resolvers that no test could tell from the literal (#3217)
## What Applied #3214's blinding procedure **outside `self-healing.ts`**, where that measurement has never been run. Two of the five resolvers across the two reporters were uncovered; this covers both. ## The measurement One resolver at a time, blinded back to its legacy ids, against each file's existing suite: | site | blinded to | result | |---|---|---| | `backlog-pressure-reporter.ts:87` hold | `["todo"]` | 2 failed — covered | | **`backlog-pressure-reporter.ts:88` wip** | `["in-progress"]` | **0 failed of 11 — UNCOVERED** | | `backlog-pressure-reporter.ts:89` terminal | `["done","archived"]` | 1 failed — covered | | `stale-task-reporter.ts:59` wip | `["in-progress"]` | 1 failed — covered | | **`stale-task-reporter.ts:60` review** | `["in-review"]` | **0 failed of 7 — UNCOVERED** | Both uncovered resolvers sit in a `Promise.all` **beside one that is covered**, so each sweep reads as converted while half of it was held by nothing. That is rule 1 in the doc — coverage is per-resolver, not per-sweep — and it is why the census cannot answer this: a syntactic scan sees five resolved sites and five is what it counts. `stale-task-reporter.ts` is the sharper case. Its describe block **already declared `signoff` in the fixture IR** and no case ever put a card there, so the review resolver was decorative. ## What they cost on a renamed board - **wip** feeds `inProgressCount`, the *denominator* of `ratio = todoCount / max(inProgressCount, 1)`. Against the literal, busy work in a renamed lane counts as **zero**, the ratio inflates, and the backlog-pressure alert fires on a queue that is draining normally — the operator is paged that the board is jammed while agents work through it. - **review** decides which rows the staleness read *fetches at all*. A review stalled for days in a renamed lane is never queried and never surfaced — precisely the condition this reporter exists to report. ## Following the four rules **Rule 2 — the fixture reaches the guarded branch.** 12 hold cards over 2 wip cards is a ratio of 6, *under* the default threshold of 10, so the correct answer is "no alert"; blinding collapses the denominator to 1, the ratio becomes 12, and it alerts. A fixture whose ratio cleared the threshold either way would exercise the sweep and never touch the line under test. **Rule 3 — assert the path-specific side effect.** `upsertInsight` not called, and `logEntry` called with `column=signoff`. Asserting `alerted === false` alone would also pass if the run bailed for an unrelated reason — missing insight store, cooldown, too few candidates — none of which involve the wip lane. **Rule 4 — the store fake honours `options.column`.** Both harnesses already did; reused rather than replaced. Each new case is paired with a negative so it cannot pass vacuously: the "does not alert" case is backed by a *same renamed board still alerts when in-progress work really is thin* case, so a reporter broken into never firing fails. ## Census **Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0` before and after.** This converts nothing. It closes coverage on conversions the census already counts as done, which is the gap #3214 names: *"the census counts comparisons; it cannot tell a working conversion from one a later merge silently reverted."* ## Verification Blind-verified in both directions — blinding each resolver fails **exactly** the new case and nothing else: ``` backlog-pressure BLIND wip -> 1 failed | 12 passed (13) restored: 13 passed stale-task BLIND review -> 1 failed | 7 passed (8) restored: 8 passed combined 21 passed (2 files) ``` No changeset: test-only, behavior-preserving, no published-package surface. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cfcbba6f81 |
fix(census): 4 RED ratchet tests on main, and the report said nothing at zero (#3218)
Two problems, both caused by the backlog actually shrinking. ## 1. Four failing tests on main **Pre-existing, not introduced here** — running this file on clean `origin/main` gives `49 passed / 4 failed` with identical messages. I checked that before touching anything, because the failures surfaced while I was editing the same file. The ratchet cases build their fixture like this: ```ts Object.entries(baseline.byFile).find(([, c]) => c > 1) // needs a file with MORE THAN ONE guard ``` After the tail reclassification no such entry exists. `find` returns undefined → `byFile[undefined] = NaN` → the baseline is corrupt → every case fails with `expected … to contain 'TIGHTENED'`, a message that points squarely at the CLI when the **fixture** is at fault. That misdirection is why this sat red. The ratchet doesn't care *which* file it tightens, only that an allowance exceeds the measured count. So `inflate` now takes any entry, and synthesises one against a real scanned file when the backlog is empty. `deflate` is the harder half: a RISE needs an allowance **below** the real count, and once every measured count is 0 the only value below is negative. The empty case uses `-1`. That is not a realistic baseline value and the comment says so — it is the sole way to exercise the `measured > allowed` comparison against a tree with nothing left to count, which is the tree this suite now runs on. Same class as the unbounded-slice rot in #3207: **census self-tests coupled to the size of a shrinking backlog.** That is now twice, so it is a pattern rather than an accident. ## 2. The report went silent at the finish line The verdict was two inline branches and neither fired at zero — `CONVERSION QUEUE EMPTY` required `totals.column > 0`. So the one state the entire fleet phase was working toward printed **nothing**, which reads as a broken scan rather than the protected end state. Extracted to a pure `describeBacklogState({ columnGuards, unexaminedGuards })` returning lines, so the caller stays a dumb printer: ``` BACKLOG ZERO: no lifecycle-column guard remains. This is the protected end state, not an empty scan — `--strict` fails on any RISE, so a new guard cannot land silently. Use the role helpers (resolveLifecycleColumns / columnHasRole). ``` Pure **specifically** so the zero state is testable before the tree reaches zero. While it was inline, only the *current* backlog state was observable — and a message nobody can test before they need it is the one that is wrong when they do. ## Evidence | check | result | |---|---| | census test file | **53 passed** (was 49 passed / 4 failed) | | behaviour on today's tree | **unchanged** — identical `CONVERSION QUEUE EMPTY` block | | empty-baseline probe | exits 1, `column-guard count ROSE` | | forced zero verdict | prints `BACKLOG ZERO … not an empty scan` | | `--strict` / `check-fnxc-future-dates` / eslint | 0 / 0 / clean | | `pnpm test:gate` | exit 0 (744 tests) | Four new tests pin all three states, including that the unexamined branch must **not** claim the queue is empty while real work is outstanding. ## Census No guard converted — this is tooling and test repair. Backlog unchanged at 1, which #3215 takes to 0. |
||
|
|
78d87f0a10 |
test(core): pin the search archive-lane WIRING — the predicate was covered, the hand-off was not (#3220)
## The false-green #3160 (mine) proved `liveSearchPredicate` honours a resolved archive set: hand it `Set(["archived","filed"])` and `filed` appears in the bound params. That contract is real and still correct. **Nothing proved `reads.ts` passes one.** It is a unit test of the collaborator, so blinding the resolver at the call site cannot fail it. A conversion, a test that looks like it covers it, and no connection between them. ## The measurement — and the instrument matters | site | vs. the predicate unit test | vs. a test that drives `reads.ts` | |---|---|---| | `reads.ts:396` cold-storage list | 0 failed | **1 failed — covered** | | `reads.ts:615` incremental sync | 0 failed | 0 failed — **UNCOVERED** | | `reads.ts:793` search | 0 failed | 0 failed — **UNCOVERED** | Against `search-excludes-renamed-archive-lane.test.ts` all three read as uncovered — an artefact of asking a file that never executes `reads.ts`. Against `cold-storage-renamed-archive-lane.test.ts`, which drives `listTasksImpl` for real, 396 is covered and the other two genuinely are not. That is rule 2 of #3214 one level up: *the test must reach the site*, and a unit test of the collaborator never does. Had I stopped at the first instrument I would have reported three uncovered resolvers, one of them wrongly. ## What 793 costs on a renamed board `searchTasks` backs the **CREATE-time near-duplicate check**. Without the resolved lanes threaded, search stops excluding the board's archive lane, and creating a task can be refused as a duplicate of one the operator archived long ago — with no way to see why, because the matching card is not on the board. Precisely the symptom #3160 set out to fix; this pins the wiring that delivers it. ## An assertion I got wrong, and the correction I expected an unreadable workflow list to leave `archivedColumns` **undefined** via the call-site `.catch(() => undefined)`. It does not: `resolveProjectColumnsForRoles` catches internally and returns its **legacy-seeded** set, so `Set(["archived"])` is threaded and the `.catch` never fires on that path. Two layers fail soft and the inner one wins. The case now asserts the guarantee that actually holds either way — **never an empty set** (which would exclude nothing and return archived rows in every search), legacy id always excluded. Recorded at the site, because the mechanism is not obvious from the call. ## Flagged, not papered over **`reads.ts:615` is left uncovered on purpose.** It composes Drizzle conditions and runs them against `layer.db` with no injectable seam, so pinning it needs a real database and belongs with the `.pg` suites. A test asserting "the query was built" rather than "the rows were excluded" would satisfy the ratchet and prove nothing. Also flagged from this sweep: `workflow-analytics.ts` and `team-analytics.ts` (4 resolvers) are **unmeasurable in my environment** — their renamed-lane coverage lives in `.pg` suites, and this worktree has no TCP PostgreSQL (`pg_isready` reports a Unix socket; the harness probes TCP, so `pgDescribe` correctly skips). Not claimed either way. ## Census **Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0`.** Converts nothing; closes coverage on a conversion the census already counts as done. ## Verification ``` as written Tests 4 passed (4) BLIND reads.ts:793 Tests 1 failed | 3 passed (4) restored Tests 4 passed (4) ``` Anti-vacuity case included: every other assertion reads a mock's arguments and would pass if the search were never reached, so one case pins that the primary search path actually ran. Typecheck clean. No changeset: test-only, behavior-preserving, no published-package surface. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0bdc9bf4fb |
fix(dashboard): archived tasks stayed in the research picker on a renamed board (#3215)
## The defect The enrich-mode task picker filtered with `task.column !== "archived"`. On a board whose archive lane is renamed, that matched nothing — so filed-away tasks stayed in the picker and an operator could attach research findings to work they had deliberately archived. ## Census before / after | | before | after | |---|---|---| | COLUMN guards (backlog) | 10 | **9** | | `ResearchTaskActionModal.tsx` | 1 | **0 — converted** | Baseline re-recorded in the same commit; `--strict` green. ## This site was declined twice, and I wrote the second wrong estimate #3213 left it counted, correctly, on the note that was here — which was mine. Both prior cost estimates were wrong, so this corrects my own work: 1. **"Needs a data-fetch change"** — reasoned about `columnFlagsByTaskId`, a per-**task** map built from board-resident rows. Right that such a map can't help (archived rows are exactly what a board map omits), but this guard asks a per-**column** question, so it never needed one. 2. **"Needs prop threading, MainContent → ResearchView → here"** — right that the answer is column-keyed, wrong about where it lives. `ListView` builds `columnFlagsById` *inline*, which made it look like the owner. The data is `useBoardWorkflows`, a hook already called from `App`, `Board`, and `HeaderWorkflowSwitcherSlot`. **Measured cost: one file.** The modal already takes `projectId`, and `ResearchView` renders it only when a finding is open (`open` hardcoded beside `if (!finding) return null`) — so the hook cannot fetch for a closed modal, which was the one real objection to calling it here. Union across workflows keyed by column id, first declaration wins — the same convention `ListView` uses, so the two cannot disagree about a shared id. `isArchivedColumnRole` fail-softs to the legacy id when a column has no flags, so an unresolved workflow behaves exactly as the literal did. ## Tests — the invariant, not the repro Per the surface-enumeration rule, four cases: renamed archive lane, legacy id, unresolved workflow (fail-soft), and a second workflow's archive lane through the cross-workflow union. A repro-only test would pass on the legacy board and prove nothing about the case the guard exists for. **Anti-vacuity control:** | | renamed lane | union | legacy id | fail-soft | |---|---|---|---|---| | pre-fix literal | **FAIL** | **FAIL** | pass | pass | | converted | pass | pass | pass | pass | The legacy and fail-soft cases hold in both directions **on purpose** — they pin that this conversion did not change the pre-resolution answer. Flagging that so 4/4 isn't read as four independent proofs. ## Measured | check | result | |---|---| | `census --strict` / `check-fnxc-future-dates` | exit 0 / exit 0 | | `eslint` | clean | | `tsc -p tsconfig.app.json` (the config that actually covers `app/`) | exit 0 | | new tests | 4/4 | | `pnpm test:gate` | exit 0 (744 tests) | ## Note on process My first attempt at the control silently did nothing — the revert script threw a `SyntaxError`, so the "pre-fix" run was the fixed code and reported 4/4. Caught it because the error printed. The table above is from the re-run. |
||
|
|
c66b434b7b |
fix(self-healing): a renamed hold lane re-logged the same overlap blocker on every sweep (#3216)
## The defect `clearStaleBlockedBy` keeps a per-task memo of which overlap blocker it already logged, so a sweep running every few seconds doesn't repeat the same line forever. The memo was retained only while the card sat in a column matching the literal `todo` — so on a renamed board it was dropped on **every** sweep and `still blocked by file scope overlap with <id>` was re-logged each time. ## Census before / after | | before | after | |---|---|---| | COLUMN guards (backlog) | 9 | **8** | | `packages/engine/src/self-healing.ts` | 1 | **0 — converted** | Baseline re-recorded in the same commit; `--strict` green. (Counts follow #3215, which took 10 → 9.) ## The stated blocker was not real The note here declined the conversion because the lane prefetch is keyed on `candidates`, *"which this closure helps build"*. Measured — it does not: ``` 6033| for (const task of blockedTasks) candidates.set(task.id, task); 6034| for (const task of queuedDependencyTasks) candidates.set(task.id, task); 6036| for (const [taskId, lastLoggedBlockerId] of this.preservedQueuedOverlapLogged) { <- only CLEARS memos ``` `candidates` is fully populated two statements earlier, and this loop only clears memo entries. So the prefetch was hoistable; it now sits above the loop. That is a pure move of a read-only computation with no conditional between the two positions. Reaching the lane clause already proves the id is a candidate — `!candidates.has(taskId)` is the first arm of the same `||` chain, so short-circuit means the lane question is only asked for ids the prefetch covered (`referencedIds.add(task.id)` runs for every candidate). `lanesOf` still falls back to the legacy set, so an unresolvable workflow answers exactly as the literal did. This is the second inherited "too expensive" estimate to fail on inspection this session (see #3215). Both were written in good faith and both were checkable in a few minutes. ## One thing typecheck caught that review would not have `memoTask?.column !== "todo"` was **also** the undefined check, and tsc narrowed the later clauses on it. Replacing it without that arm compiled clean to the eye but broke narrowing — `TS18048: 'memoTask' is possibly 'undefined'` on the next line. `|| !memoTask` is now explicit rather than implied. ## Evidence The test drives the sweep **twice**, because a single pass cannot observe a dedup memo at all. | | pre-fix literal | converted | |---|---|---| | `still blocked by file scope overlap` log lines | **2 — FAILS** | **1 — passes** | Failure message against the pre-fix code: `expected [ [ 'FN-DEPENDENT', …(1) ], …(1) ] to have a length of 1 but got 2`. Worth correcting the record: the note called the cost *"a duplicate log line, not a wrong lifecycle decision"*. The lifecycle half is right — but it is a duplicate on **every sweep**, so it is recurring log spam, not a one-off. That is a bigger cost than the note implies, though still not a correctness bug. | check | result | |---|---| | `census --strict` / `check-fnxc-future-dates` | exit 0 / exit 0 | | `eslint` / engine `tsc --noEmit` | clean / exit 0 | | self-healing + overlap suites | 15 / 21 / 6 passed | | `pnpm test:gate` | exit 0 (744 tests) | Reused the existing `RENAMED_BOARD_IR` harness in that file rather than building a new one. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved cleanup of stale workflow blockers, including renamed workflow lanes. - Prevented duplicate overlap warnings during repeated cleanup. - More reliably preserves valid queued overlaps while ignoring missing or inactive tasks. - **Tests** - Added regression coverage for repeated stale-blocker cleanup and duplicate warning prevention. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |