5f6f39e115f3b6933ca20d2fb355ef536e60f7fc
3821 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3f06d7201b |
test(engine): assert the evaluator's exact archived-lane set, not toContain (#3232)
## What Follow-up to the #3224 review comment *"reject legacy archived identifiers for renamed workflows."* Test-only. That comment had two halves. **The half it got wrong** is already answered on main: asserting an exact single-column set `["vaulted"]` *fails*, because `resolveProjectColumnsForRoles` unions `LEGACY_COLUMN_IDS_BY_ROLE` in as a documented floor — so a row whose workflow cannot be resolved still classifies. Pinning `["vaulted"]` would encode the opposite of the design. **The half it got right was never addressed.** `toContain` also passes when the set grows a lane nobody intended, and an over-broad archived set silently classifies *live* rows as archived. So the reviewer's worry was legitimate even though the proposed fix was not. The exact set is assertable — it just is not the one the review proposed: | board | resolved set | |---|---| | default | `["archived"]` | | renamed | `["archived", "vaulted"]` — legacy floor + the board's own lane | `[...].sort()` in the helper makes ordering stable, so these pin the resolver's whole answer rather than a substring of it. ## Proven to catch what `toContain` missed Giving the fixture a second `archived`-trait column fails both new assertions: ``` AssertionError: expected [ 'archived', 'cold-store' ] to deeply equal [ 'archived' ] AssertionError: expected [ 'archived', 'cold-store', 'vaulted' ] to deeply equal [ 'archived', 'vaulted' ] Tests 2 failed | 1 passed (3) ``` The previous `toContain` assertions pass unchanged against that same spurious lane. That is the whole justification for this PR — without the injection test it would be a stylistic preference. ``` clean: Tests 3 passed (3) spurious lane: Tests 2 failed | 1 passed (3) lint clean ``` ## Note I said in the review thread I would tighten this, so this closes that loop. I also checked before editing whether another worker had already done it — main already carries the FNXC tag and the legacy-floor explanation from the same review round, so this PR adds only the part still missing rather than redoing settled work. |
||
|
|
d4c25384ae |
test(engine): record what the zero-backlog early return stops testing (#3228)
## A green case that stopped testing what its name says #3226 fixed a red `main` correctly: `fileWithGuards()` now returns `null` at zero backlog, and with nothing to inflate there is no rise to manufacture. Asserting `totals.column === 0` and returning is the honest response. What went unrecorded is the cost. At zero, these two cases: - *"exits 0 and REWRITES the baseline under `--update-baseline`, even when the count rose"* - *"exits 1 and LEAVES the baseline alone on a rise without `--update-baseline`"* no longer exercise the CLI's ordering or exit codes. They assert the backlog is empty and return. **If the write-before-exit ordering regressed — the exact bug those cases were written for — both would still pass.** That matters more than it would elsewhere, because **zero is not a state to wait out.** It is this program's terminal state: the backlog went 126 → 0 and is meant to stay there. So the vacuity is permanent, not transitional. This file already legislates against precisely this, two hundred lines down: > `/* Anti-vacuity: an empty exclusion list would make the assertion below trivially true. */` ## What this PR does Adds a comment on `fileWithGuards()` recording (a) which cases go vacuous at zero and why, (b) that zero is terminal so it will not resolve itself, and (c) the durable fix. **Comment only. No behaviour change — suite stays 53/53.** ## The durable fix, recorded rather than done Point the scan at a synthetic tree so the fixture stops being a function of the real backlog — the same seam `FUSION_CENSUS_BASELINE_PATH` already provides for the baseline, applied to the file list. It needs one CLI correction to work, and that is a genuine bug in my own code regardless of this suite: `triageFindings` and the sync-resolver check read files via `join(REPO_ROOT, f.file)`, where `REPO_ROOT` is derived from the **script's** location. An overridden file list therefore changes which paths are *listed* without changing where they are *read from*, and every read misses with `ENOENT`. I verified that approach works (a three-file fixture yields a stable `1 backlog / 1 deliberate / 1 sync-resolved`) and got two of the six failing cases green with it, then stopped rather than keep guessing in a file being actively revised. Left as a comment so whoever takes it does not re-derive the diagnosis. ## Why this is worth a PR at all This program's recurring failure is instruments that report green while measuring nothing — an inert conversion the census scored as a win, a ratchet wired to nothing, a gate that could not fail. A test asserting `0 === 0` under a name promising ordering coverage is the same shape at the test layer. The suite cannot be fixed in this PR without re-opening work someone else owns, but it can at least stop being silent about it. ## Census before / after ``` before: COLUMN guards (the backlog): 0 after: COLUMN guards (the backlog): 0 ``` ## Verification `test:gate` exit 0 · `lifecycle-column-census.test.ts` **53 passed** · `fnxc-future-dates`, `lifecycle-columns`, `inert-sync-lanes`, `quarantine-ledger`, `inert-flag-seams`, `lane-wiring`, `sql-column-literals` — all exit 0. |
||
|
|
12c4ab5a6e |
test(engine): pin the evaluator's archived-lane read — the service had no test at all (#3224)
## What Pins the **evaluator's archived-lane read**. Test-only — no product change. `HybridEvaluatorService.evaluateTask` resolves the board's archived lanes and hands them to `collectDeterministicSignals`, which decides which of a task's related rows count as archived when scoring a run. **The service had no test anywhere in the repo.** Four test files import the module; none construct or exercise it. So this conversion was unobservable for the simplest possible reason — *nothing ran the code*. That is a different failure from the ones this audit has been finding (harnesses that run the code but cannot see the difference), and worth distinguishing: no amount of fixture care helps when the entry point is never called. ## Measured | | default (control) | renamed | differential | |---|---|---|---| | converted | pass | pass | pass | | blinded to `["archived"]` | pass | **FAIL** | **FAIL** | ``` converted: Test Files 1 passed (1) / Tests 3 passed (3) blinded: Test Files 1 failed (1) / Tests 2 failed | 1 passed (3) the 4 files importing evaluator.ts, plus this one: 5 files/76 tests, all green lint clean; fnxc-future-dates: none added; census unchanged ``` Per the rule I documented in #3223, the blind was confirmed applied with `git diff --stat` **before** the run rather than trusting the tool's exit code. ## What breaks without it On a board whose archived lane is `vaulted`, the evaluator hands the collector the legacy `{archived}` set. Rows resting in `vaulted` are not recognised as archived, and the deterministic half of every evaluation score is computed from a wrong picture of the task's history. **Nothing errors, the run completes, the number is just wrong** — which is why it survived unnoticed. ## Pinned without faking a provider response The assertion is about what the collector *receives*, which is decided before any model call. `collectDeterministicSignals` is mocked to record its arguments and throw a sentinel; the test asserts the resolved lane set and stops. This is deliberate over the obvious alternative of feeding `runPrompt` a canned AI payload: `deps.runPrompt` is injectable so either approach is offline, but a canned payload has to satisfy `parseAiResponse` and every `EVAL_SCORE_CATEGORIES` entry, and would silently rot into a maintenance burden on a test whose subject is one `Set`. Reversible if someone later wants full end-to-end evaluator coverage — that is a different test, not this one. ## Completes the engine audit With this, every `resolveProjectColumnsForRoles` call site in `packages/engine` has been blinded: | file | resolvers | result | |---|---|---| | `self-healing.ts` | 64 | 21 pinned, 1 recorded inert by construction, remainder mapped | | `executor.ts` | 2 | both already covered | | `scheduler.ts` | 1 | uncovered → pinned (#3219, merged) | | `triage.ts` | 1 | uncovered → pinned (#3221) | | `restart-recovery-coordinator.ts` | 1 | already covered | | `notification-service.ts` | 1 | already covered | | `evaluator.ts` | 1 | uncovered → pinned (this PR) | `project-engine.ts:5154` takes `roles` as a **parameter**, so it is a generic wrapper with no fixed role set to blind — flagged rather than guessed at; its callers are where the question belongs. **`packages/core`'s 17 files remain entirely unaudited** and I am claiming nothing about them. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added regression coverage to verify reliable resolution of archived workflow lanes. * Covered both the default archived-lane name and custom renamed configurations. * Confirmed compatibility with legacy archived-lane naming behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3c7dd9a803 |
test(engine): keep census regressions valid at zero backlog (#3226)
## Summary - keeps lifecycle-census end-to-end fixtures valid after the conversion backlog reaches zero - treats zero backlog as a real protected end state instead of requiring a remaining guard or claim target - preserves nonzero rise/claim assertions when guards remain ## Test plan - `corepack pnpm --filter @fusion/engine exec vitest run src/__tests__/lifecycle-column-census.test.ts --silent=passed-only --reporter=dot --project=engine-default` - `corepack pnpm --filter @fusion/engine typecheck` - verified the same suite at zero backlog on the aggregate runtime (53/53) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Improved lifecycle validation to handle empty backlogs without errors. * Added coverage for zero-item results and changing file sets. * Enhanced baseline and trend checks to report completed backlogs consistently. * Improved resilience when expected result entries are unavailable, ensuring validation completes cleanly with accurate zero counts. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
719cf281fd |
test(engine): pin the startup stale-planning sweep's planner-lane read (#3221)
## What Pins the **startup stale-planning sweep's** planner-lane read in `triage.ts`. Test-only — no product change. A card can hold `status: "planning"` while triage specifies it in place. A crash or restart before planning completes leaves that status set, and a startup sweep clears it. If the sweep misses the card, it occupies a planning admission slot **permanently** and new triage work is never admitted. The lane read was converted from `resolvePlannerLanes(store, "")` — called with an **empty task id**, so it could never resolve a task and always answered with the default board — to a project-level `resolveProjectColumnsForRoles(store, ["intake", "hold"])`. **No test could observe that conversion.** All 25 existing triage test files stayed green (375/375) with the resolver neutered — under *both* blinds tried. ## Measured | | default (control) | legacy ids | renamed | differential | |---|---|---|---|---| | converted | pass | pass | pass | pass | | blinded to empty set | pass | pass | **FAIL** | **FAIL** | | blinded to legacy pair | pass | pass | **FAIL** | **FAIL** | ``` converted: Tests 4 passed (4) blinded to empty set: Tests 2 failed | 2 passed (4) blinded to legacy: Tests 2 failed | 2 passed (4) existing 25 files: green under BOTH blinds (only this new file fails) triage suite: 25 files/375 tests -> 26 files/379 tests, all green lint clean; fnxc-future-dates: none added; census unchanged ``` ## A prediction I got wrong, and what it changes The site is **seed-then-union**: ```ts const sweepColumns = [...new Set(["triage", "todo", ...projectPlannerColumns])]; ``` I expected blinding the resolver to its legacy pair `["triage","todo"]` to be a **no-op by construction** — the seed already contains both. It is not. That reasoning holds only for a *default* board; on a renamed board the resolver is the sole contributor of `drafting`, so the legacy blind drops it and the renamed cases fail. The corrected rule, which is narrower than the one I was carrying: **a seed-then-union site hides a defect only while every lane you assert on is already in the seed.** Assert on a lane that is not, and the union stops protecting it. The empty-set blind remains the stricter of the two because it also models a resolver returning nothing at all. Both are recorded in the test file so the next person does not re-derive it. ## A tooling failure worth naming My first triage measurement reported **375/375 green under blinding** — and was a lie. The blinding script had no mapping for `["intake","hold"]` and exited `2` **silently**; the `&&` short-circuited the check while a `;` let vitest run against **unmodified source**. A no-op blind produces a green run that is indistinguishable from real coverage. The script now prints what it substituted and where, fails loudly on an unmapped role or a missing variable, and I self-tested both directions (bogus var → rc 2 with a message; real var → rc 0 with the substitution echoed) before trusting any number above. This is the same standard I have been applying to other people's guards — *a guard that reports success without checking anything is worse than no guard* — and my own instrument failed it. ## Also pinned The **legacy half** of the union. `triage`/`todo` stay in the sweep even on a board whose workflow declares neither, because pre-U11 and Coding (Ideas) rows can rest there. Dropping them in favour of the resolved lanes alone would strand exactly those rows, so there is now a case asserting it. ## Flagged, not guessed - The FNXC stamp at `triage.ts:987` reads `2026-07-31-23:59`, hours ahead of the `date -u` clock. It is one of the 181 grandfathered stamps so the gate is green; I left it rather than widen this PR. |
||
|
|
9ab9822b8d |
test(engine): pin the deleted-blocker WIP-lane read, which no test could see (#3219)
## What Pins the deleted-blocker sweep's **WIP-lane read** in `scheduler.ts`. Test-only — no product change. When a task is soft-deleted, a `task:deleted` listener clears `blockedBy` on every dependent so the work can be scheduled again. It reads two lanes to find them: hold and WIP. The WIP read was already converted to `resolveProjectColumnsForRoles(store, ["countsTowardWip"])`. **Blinding that resolver back to the literal `["in-progress"]` left all 14 existing scheduler test files green — 145/145.** The conversion was load-bearing and unpinned. ## Why nothing could see it Two independent harness properties, **either one sufficient** to hide the defect: 1. **`resolveProjectColumnsForRoles` returns legacy ids and nothing else when the store has no `listWorkflowDefinitions`** — an intentional degrade in `project-lane-vocabulary.ts` so an unreadable workflow list cannot fail a sweep. The shared scheduler harness does not define it, so *the resolved set and the literal set were equal by construction* in every existing test. 2. **`listTasks` was mocked as `vi.fn(async () => tasks)`**, ignoring its `column` filter. A mock that returns every task regardless of the lane asked for cannot detect a wrong lane. This is worth naming because #1 is not a test bug — it is correct production behaviour that happens to erase the difference a test is trying to measure. A harness can satisfy a conversion's *shape* while making its *effect* unobservable. There is a test file named `scheduler-renamed-dependency-and-review-lanes.test.ts` covering dependency satisfaction, file-scope leases, base-branch stacking, PR hydration and mission completion on a renamed board. It does not reach this sweep, and its harness has both properties above. ## What breaks without the conversion On a board whose WIP lane is `building`, the literal read asks for `in-progress`, finds nothing, and the in-flight dependent is never reconciled. It keeps `blockedBy` pointing at a task that no longer exists — **permanently**, because the blocker can never be completed or re-deleted to trigger another sweep. Work stops with nothing to rescue it. ## Measured | | default (control) | renamed | differential | |---|---|---|---| | converted | pass | pass | pass | | blinded to `["in-progress"]` | pass | **FAIL** | **FAIL** | ``` converted: Test Files 1 passed (1) / Tests 3 passed (3) blinded: Test Files 1 failed (1) / Tests 2 failed | 1 passed (3) scheduler suite: 14 files/145 tests -> 15 files/148 tests, all green lint: clean fnxc-future-dates: none added ``` The default-vocabulary control passes in **both** columns by design: it is what keeps a generic break in this path from hiding in the renamed assertion. ## Census Unchanged — **2 / 148 deliberate**. `scheduler.ts` was already at 0. This PR adds coverage, not conversions; the census counts comparisons and would not have moved either way, which is the same blind spot that let #3078 merge green. ## Flagged, not guessed - The `resolveTaskParkedColumns` **hold** read one line above (`scheduler.ts:1402`) is the other half of this pair. My scenario places its dependent in the WIP lane, so it does not exercise the hold read and I am not claiming it is pinned. - The FNXC stamp at `scheduler.ts:1404` reads `2026-08-01-05:00`, which is future-dated. It is one of the 181 grandfathered stamps, so the gate is green and I left it alone rather than widen this PR. |
||
|
|
3146a745bf |
test(engine): cover two reporter resolvers that no test could tell from the literal (#3217)
## What Applied #3214's blinding procedure **outside `self-healing.ts`**, where that measurement has never been run. Two of the five resolvers across the two reporters were uncovered; this covers both. ## The measurement One resolver at a time, blinded back to its legacy ids, against each file's existing suite: | site | blinded to | result | |---|---|---| | `backlog-pressure-reporter.ts:87` hold | `["todo"]` | 2 failed — covered | | **`backlog-pressure-reporter.ts:88` wip** | `["in-progress"]` | **0 failed of 11 — UNCOVERED** | | `backlog-pressure-reporter.ts:89` terminal | `["done","archived"]` | 1 failed — covered | | `stale-task-reporter.ts:59` wip | `["in-progress"]` | 1 failed — covered | | **`stale-task-reporter.ts:60` review** | `["in-review"]` | **0 failed of 7 — UNCOVERED** | Both uncovered resolvers sit in a `Promise.all` **beside one that is covered**, so each sweep reads as converted while half of it was held by nothing. That is rule 1 in the doc — coverage is per-resolver, not per-sweep — and it is why the census cannot answer this: a syntactic scan sees five resolved sites and five is what it counts. `stale-task-reporter.ts` is the sharper case. Its describe block **already declared `signoff` in the fixture IR** and no case ever put a card there, so the review resolver was decorative. ## What they cost on a renamed board - **wip** feeds `inProgressCount`, the *denominator* of `ratio = todoCount / max(inProgressCount, 1)`. Against the literal, busy work in a renamed lane counts as **zero**, the ratio inflates, and the backlog-pressure alert fires on a queue that is draining normally — the operator is paged that the board is jammed while agents work through it. - **review** decides which rows the staleness read *fetches at all*. A review stalled for days in a renamed lane is never queried and never surfaced — precisely the condition this reporter exists to report. ## Following the four rules **Rule 2 — the fixture reaches the guarded branch.** 12 hold cards over 2 wip cards is a ratio of 6, *under* the default threshold of 10, so the correct answer is "no alert"; blinding collapses the denominator to 1, the ratio becomes 12, and it alerts. A fixture whose ratio cleared the threshold either way would exercise the sweep and never touch the line under test. **Rule 3 — assert the path-specific side effect.** `upsertInsight` not called, and `logEntry` called with `column=signoff`. Asserting `alerted === false` alone would also pass if the run bailed for an unrelated reason — missing insight store, cooldown, too few candidates — none of which involve the wip lane. **Rule 4 — the store fake honours `options.column`.** Both harnesses already did; reused rather than replaced. Each new case is paired with a negative so it cannot pass vacuously: the "does not alert" case is backed by a *same renamed board still alerts when in-progress work really is thin* case, so a reporter broken into never firing fails. ## Census **Unchanged — `CONVERSION QUEUE EMPTY`, `AVAILABLE: 0` before and after.** This converts nothing. It closes coverage on conversions the census already counts as done, which is the gap #3214 names: *"the census counts comparisons; it cannot tell a working conversion from one a later merge silently reverted."* ## Verification Blind-verified in both directions — blinding each resolver fails **exactly** the new case and nothing else: ``` backlog-pressure BLIND wip -> 1 failed | 12 passed (13) restored: 13 passed stale-task BLIND review -> 1 failed | 7 passed (8) restored: 8 passed combined 21 passed (2 files) ``` No changeset: test-only, behavior-preserving, no published-package surface. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cfcbba6f81 |
fix(census): 4 RED ratchet tests on main, and the report said nothing at zero (#3218)
Two problems, both caused by the backlog actually shrinking. ## 1. Four failing tests on main **Pre-existing, not introduced here** — running this file on clean `origin/main` gives `49 passed / 4 failed` with identical messages. I checked that before touching anything, because the failures surfaced while I was editing the same file. The ratchet cases build their fixture like this: ```ts Object.entries(baseline.byFile).find(([, c]) => c > 1) // needs a file with MORE THAN ONE guard ``` After the tail reclassification no such entry exists. `find` returns undefined → `byFile[undefined] = NaN` → the baseline is corrupt → every case fails with `expected … to contain 'TIGHTENED'`, a message that points squarely at the CLI when the **fixture** is at fault. That misdirection is why this sat red. The ratchet doesn't care *which* file it tightens, only that an allowance exceeds the measured count. So `inflate` now takes any entry, and synthesises one against a real scanned file when the backlog is empty. `deflate` is the harder half: a RISE needs an allowance **below** the real count, and once every measured count is 0 the only value below is negative. The empty case uses `-1`. That is not a realistic baseline value and the comment says so — it is the sole way to exercise the `measured > allowed` comparison against a tree with nothing left to count, which is the tree this suite now runs on. Same class as the unbounded-slice rot in #3207: **census self-tests coupled to the size of a shrinking backlog.** That is now twice, so it is a pattern rather than an accident. ## 2. The report went silent at the finish line The verdict was two inline branches and neither fired at zero — `CONVERSION QUEUE EMPTY` required `totals.column > 0`. So the one state the entire fleet phase was working toward printed **nothing**, which reads as a broken scan rather than the protected end state. Extracted to a pure `describeBacklogState({ columnGuards, unexaminedGuards })` returning lines, so the caller stays a dumb printer: ``` BACKLOG ZERO: no lifecycle-column guard remains. This is the protected end state, not an empty scan — `--strict` fails on any RISE, so a new guard cannot land silently. Use the role helpers (resolveLifecycleColumns / columnHasRole). ``` Pure **specifically** so the zero state is testable before the tree reaches zero. While it was inline, only the *current* backlog state was observable — and a message nobody can test before they need it is the one that is wrong when they do. ## Evidence | check | result | |---|---| | census test file | **53 passed** (was 49 passed / 4 failed) | | behaviour on today's tree | **unchanged** — identical `CONVERSION QUEUE EMPTY` block | | empty-baseline probe | exits 1, `column-guard count ROSE` | | forced zero verdict | prints `BACKLOG ZERO … not an empty scan` | | `--strict` / `check-fnxc-future-dates` / eslint | 0 / 0 / clean | | `pnpm test:gate` | exit 0 (744 tests) | Four new tests pin all three states, including that the unexamined branch must **not** claim the queue is empty while real work is outstanding. ## Census No guard converted — this is tooling and test repair. Backlog unchanged at 1, which #3215 takes to 0. |
||
|
|
c66b434b7b |
fix(self-healing): a renamed hold lane re-logged the same overlap blocker on every sweep (#3216)
## The defect `clearStaleBlockedBy` keeps a per-task memo of which overlap blocker it already logged, so a sweep running every few seconds doesn't repeat the same line forever. The memo was retained only while the card sat in a column matching the literal `todo` — so on a renamed board it was dropped on **every** sweep and `still blocked by file scope overlap with <id>` was re-logged each time. ## Census before / after | | before | after | |---|---|---| | COLUMN guards (backlog) | 9 | **8** | | `packages/engine/src/self-healing.ts` | 1 | **0 — converted** | Baseline re-recorded in the same commit; `--strict` green. (Counts follow #3215, which took 10 → 9.) ## The stated blocker was not real The note here declined the conversion because the lane prefetch is keyed on `candidates`, *"which this closure helps build"*. Measured — it does not: ``` 6033| for (const task of blockedTasks) candidates.set(task.id, task); 6034| for (const task of queuedDependencyTasks) candidates.set(task.id, task); 6036| for (const [taskId, lastLoggedBlockerId] of this.preservedQueuedOverlapLogged) { <- only CLEARS memos ``` `candidates` is fully populated two statements earlier, and this loop only clears memo entries. So the prefetch was hoistable; it now sits above the loop. That is a pure move of a read-only computation with no conditional between the two positions. Reaching the lane clause already proves the id is a candidate — `!candidates.has(taskId)` is the first arm of the same `||` chain, so short-circuit means the lane question is only asked for ids the prefetch covered (`referencedIds.add(task.id)` runs for every candidate). `lanesOf` still falls back to the legacy set, so an unresolvable workflow answers exactly as the literal did. This is the second inherited "too expensive" estimate to fail on inspection this session (see #3215). Both were written in good faith and both were checkable in a few minutes. ## One thing typecheck caught that review would not have `memoTask?.column !== "todo"` was **also** the undefined check, and tsc narrowed the later clauses on it. Replacing it without that arm compiled clean to the eye but broke narrowing — `TS18048: 'memoTask' is possibly 'undefined'` on the next line. `|| !memoTask` is now explicit rather than implied. ## Evidence The test drives the sweep **twice**, because a single pass cannot observe a dedup memo at all. | | pre-fix literal | converted | |---|---|---| | `still blocked by file scope overlap` log lines | **2 — FAILS** | **1 — passes** | Failure message against the pre-fix code: `expected [ [ 'FN-DEPENDENT', …(1) ], …(1) ] to have a length of 1 but got 2`. Worth correcting the record: the note called the cost *"a duplicate log line, not a wrong lifecycle decision"*. The lifecycle half is right — but it is a duplicate on **every sweep**, so it is recurring log spam, not a one-off. That is a bigger cost than the note implies, though still not a correctness bug. | check | result | |---|---| | `census --strict` / `check-fnxc-future-dates` | exit 0 / exit 0 | | `eslint` / engine `tsc --noEmit` | clean / exit 0 | | self-healing + overlap suites | 15 / 21 / 6 passed | | `pnpm test:gate` | exit 0 (744 tests) | Reused the existing `RENAMED_BOARD_IR` harness in that file rather than building a new one. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved cleanup of stale workflow blockers, including renamed workflow lanes. - Prevented duplicate overlap warnings during repeated cleanup. - More reliably preserves valid queued overlaps while ignoring missing or inactive tasks. - **Tests** - Added regression coverage for repeated stale-blocker cleanup and duplicate warning prevention. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
aa1655ccd9 |
fleet: reclassify the census tail — 10 → 2 guards, all reasoning already in the code (#3213)
## Census before / after
```
before after
COLUMN guards (backlog) 10 2
DELIBERATE-LITERAL 138 148
```
Baseline re-recorded in the same commit; `--strict` green.
## This converts nothing — the tail was never backlog
All ten remaining guards already carried an explicit in-code decision.
**None carried the `DELIBERATE-LITERAL` marker the census reads**, so
each re-appeared to every fleet pass as if unexamined. That is the whole
defect this fixes.
| site | the reasoning already at the site |
| --- | --- |
| `audit-ops.ts`, `moves.ts` | the degraded fallback arm of an
**already-converted** site; the live arm uses the resolved lane set |
| `scheduler.ts` ×2 | *"LEFT COUNTED"* — an await behind the
`tracked.has` re-entrance guard lets two updates double-start a monitor;
the sibling is the measured-expensive `task:updated` emit path (26 sites
against 7) |
| `notification-service.ts` | this method and its only caller are
**sync**, reached from a listener the store invokes as `(task: Task):
void`; resolving makes the chain async and reorders notification
classification against every other `task:updated` handler |
| `lifecycle-ops.ts` | *"Recorded rather than converted"* — dead code |
| `task-id-integrity.ts` | sync, no store-scoped read; converting alone
would disagree with `getLiveTaskColumn` |
| `triage.ts` | *"LEFT COUNTED until then"* — wants a non-sync-resolved
lane answer |
## Marker placement is load-bearing, and I got it wrong twice
The census reads a node's **leading** comments. A marker in a nearby
block comment attaches to the wrong node and is **silently ignored** —
it reads as reviewed while the count still lists the site.
- `task-id-integrity.ts` — my first marker went into the block comment
above the `const`; the literal is in the `return`. Count stayed at 1
until I moved it.
- `ResearchTaskActionModal.tsx` — marker added, **measured that it did
not register**, reverted.
Every edit was verified by re-running the census, not assumed. That is
the only reason the count actually moved.
## Two sites deliberately left counted
- **`ResearchTaskActionModal.tsx`** — the literal sits mid-expression
inside a `.then()` chain, so no marker can attach. The census's own
guidance is to hoist it into a named helper; the site's note asks for
that to be someone's deliberate change rather than a drive-by, so it
stays counted and honest.
- **`self-healing.ts`** — the memo closure I converted and reverted in
#3049. Its note: a renamed board costs a duplicate log line, not a wrong
lifecycle decision.
## Correction I owe on the measurement itself
For many turns I reported "zero unclaimed guards". That came from a bug
in **my own** query — `byFile` is an array of `[file, count]` pairs and
I had switched to `Object.entries()`, which yields `[index, pair]`, so
`n > 0` was always false and the filter returned zero regardless of
state. It agreed with reality while open PRs held every file, which is
why it went unnoticed; it was still wrong, and a constant zero against a
falling backlog should have prompted me to check it sooner.
## Verification (measured)
- engine `self-healing` + `scheduler` suites — **1003 passed / 56
files**
- core `task-id` / `moves` suites — green
- `tsc --noEmit` clean in core, engine and dashboard; `eslint` clean
- `pnpm test:gate` — green
- `lifecycle-column-census --strict`, `check-lane-wiring`,
`check-sql-column-literals`, `check-fnxc-future-dates` — green
No changeset: `@fusion/core`, `@fusion/engine` and `@fusion/dashboard`
are private, and no runtime behaviour changes.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Clarified internal annotations for archived, in-progress, and
in-review workflow states.
* Documented fallback behavior and timing safeguards across lifecycle,
scheduling, notification, and triage flows.
* **Chores**
* Updated internal lifecycle tracking baselines to reflect current
annotations and state coverage.
* **Bug Fixes**
* No user-visible behavior changes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
|
||
|
|
f94ff68391 |
docs(engine): record that agentParkedColumns is inert at this call site (passed-but-unread) (#3212)
Closes the last entry I could not pin on the #3115 coverage map — by establishing **why** it cannot be pinned, rather than leaving it open or forcing a green test around it. ## `agentParkedColumns` is inert at this call site It influences exactly one output — `shouldPreserveParkedLink` — and `recoverAgentsRunningOnInactiveTasks` **never reads it**. The gate is: ```ts if (proof.hasFreshRun || proof.hasActiveExecution) continue; ``` Neither depends on the lane. So passing a resolved set changes nothing today, and blinding it back to the legacy ids leaves every test green **because there is no behaviour to observe**. That is not a coverage gap — there is nothing there to cover. ## Kept, not deleted - Removing it makes this call site read as **unwired** to the lane-wiring ratchet, inviting the next worker to "fix" it by re-adding exactly this. - If the gate ever adopts `shouldPreserveParkedLink` — the parked-specific semantics the sibling sweep uses (`recoverDriftedAgentTaskLinks`, wired in #3208 an hour ago) — the resolved set is already correct here. ## The shape worth naming **Passed-but-unread** is the mirror of the **resolved-gate, literal-branch** defect #3208 fixed. Both read as converted while deciding nothing — and only one of them is a bug. A ratchet that counts call sites cannot tell them apart: #3208's site looked *unwired* and was a live defect; this one looks *wired* and is dead code. That is why the distinction belongs in a comment at the site rather than in a baseline number. ## Map status **21 of 26 pinned**, 1 established as unpinnable-by-construction, 4 remaining with obstacles recorded: - `reclaimHoldColumns` / `reclaimReviewColumns` — both audit paths emit the same `branch:auto-reclaim` type, differing only by a `trigger` string the branch-level scan also produces. - `completedHoldColumns`, `wsDoneColumns`, `doneMetaColumns` and others have since gone green from other workers' PRs. ## Verification `pnpm test:gate` 13 + 161 + 499 + 71 · lint · census `--strict` · fnxc-dates (TZ=UTC) — green. Comment-only change. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added an explanatory note clarifying parked-column handling and its connection to future parked-link preservation behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
c027a72d23 |
fix(engine): a live agent lost its task link on a renamed hold lane (found while testing, not converting) (#3208)
**A defect, not a coverage gap** — found while trying to pin `agentParkedColumns` from the #3115 map. `recoverDriftedAgentTaskLinks` enters its preservation branch on a **resolved** question (`isPreWipColumn`) and then decided it on a **literal** one: `evaluateParkedAgentTaskLink` was called without `parkedColumns`, so parked-ness fell back to `todo`/`triage`. The sibling sweep passes the resolved set; this call site did not. On a renamed board: the card is pre-wip, `isParkedTaskColumn` says no, `shouldPreserveParkedLink` is false, and **an agent with a fresh heartbeat run has its task link cleared — while it is working.** `task-agent-sync.ts` predicted this in writing when the parameter was introduced: > *"turning a stale-link bug into a dropped-link bug, since the card would be treated as unparked and its live agent link cleared"* That is what an unpassed optional lane parameter costs — the same missed-pair shape as #2956, #2963 and #3186. ## The test needed two fixture corrections, both caught by failing - the renamed IR had **no hold column**, so no card could be pre-wip at all; - the **per-task selection readers** were missing, so `isPreWipColumn` resolved the default IR and the branch was never entered. Either alone made the case pass while exercising nothing. Third time today a fixture passed for a reason unrelated to the resolver — a pattern, not an anecdote. ## Measured 15 pass; removing `parkedColumns` from the call fails exactly this case. ## Note on how it was found I had discarded a probe at this sweep earlier for failing to discriminate. Coming back with the obstacle understood — the fixture must reach the branch the resolver gates — turned a coverage miss into a defect find. The five discards this session were not wasted; three of them named the obstacle that made a later attempt work. ## Verification `self-healing-agent-link-drift` **15 passed** · `pnpm test:gate` 161 + 13 + 499 + 71 · lint · lane-wiring — green. |
||
|
|
8661b739ff |
fix(scheduler): a board with TWO complete columns left dependents waiting forever (#3210)
## The defect On a board declaring more than one complete-trait column — a merged lane and a shipped lane, say — a card landing in the **second** one was never recognised as finished, so nothing unblocked its dependents. Silent: no error, the dependent just waits. Two problems, the same shape: 1. **`TaskMoveLanes` carried one id per role.** That is right for *"where should this card go"* and wrong for *"is this column one of the finished lanes"*, which is a **membership** question. The payload could not express such a board at all. 2. **`mergeParkedColumns` rebuilt `terminal` as `new Set([complete, archived])`** — discarding `base.terminal`, which the sync IR path had already resolved correctly, and narrowing a membership set back to first-match-per-role. Point 2 contradicted the note sitting directly above it in the same file: > `terminal` is a MEMBERSHIP set, and it is not the same question as `complete`/`archived`. […] A workflow may declare more than one complete-trait column […] and `to === parked.complete` sees only the first and silently skips the rest. The reasoning was already written down. The overlay added later didn't honour it. ## Fix `TaskMoveLanes.terminal?: readonly string[]`, filled from `columnsWithFlag(ir, "complete"|"archived")` rather than the first-match `resolveLifecycleColumns`, and the merge is now a **union** of base, payload, and the single lanes. Optional, so all 12 emitters and every listener keep compiling — a listener that ignores it is exactly as correct as before. Union is the direction `scheduler.ts` already argues for at line ~422: a superset costs one extra query; a subset **silently withholds work from a finished card**. `complete` deliberately stays first-match — a set would be the wrong shape for a move *target*. Both questions now coexist rather than one replacing the other. ## How it was found, and what it corrects Supplying `task:moved` lanes fixed every *other* renamed-board case in `scheduler-renamed-hold-events` — measured **10 passed / 1 failed** — and left exactly this one broken. That same measurement is why I narrowed my earlier claim on #3082: the other behaviours were never broken in production, because all 12 emitters already carry lanes. This is the residue that was genuinely broken. ## Tests — both with anti-vacuity controls | control | result | |---|---| | revert `toTaskMoveLanes` | **2 of 4** core tests fail (the terminal pair) | | revert the scheduler union | the new engine test fails, **and only it** (1 failed / 11 passed) | | both restored | 4 passed, 12 passed | The 2 core tests that pass either way are shape invariants asserted on purpose (`complete` must stay first-match; a column-less IR returns `undefined` rather than an invented lane) — flagging that so the control isn't read as 4-of-4. The pre-existing scheduler case emits **without** lanes, which no production emitter does, so it exercises the sync fallback. The new one emits `toTaskMoveLanes(ir)` — the shape that actually ships. ## Measured | check | result | |---|---| | `@fusion/core` / `@fusion/engine` tsc | exit 0 / exit 0 | | eslint | clean | | `census --strict`, `check:fnxc-future-dates`, `check:changesets` | exit 0 | | every `TaskMoveLanes` consumer | 24 passed | | `pnpm test:gate` | **exit 0 — 744 tests, up 12** | ## Census No guard converted; this is a payload-shape fix. Backlog unchanged at 11, all deferred. |
||
|
|
215f09d88f |
fix(census): the bare command could not say the conversion queue is EMPTY — and a test fix for main (#3207)
## Why this exists
The fleet instruction is *"claim the largest unclaimed census file
cluster (`node scripts/lifecycle-column-census.mjs`)"*. That command
cannot answer it. The availability verdict lived **only** behind
`--claims`, which shells to `gh`:
```
line 342: if (claims && !json) {
```
So a worker following the instruction literally sees per-file counts,
reads a nonzero backlog as a work queue, and picks a file whose guard is
already documented as deferred. Counts alone cannot separate *work left*
from *debt left*.
**Measured cost:** the queue reached **zero unexamined guards** while
dispatch continued. I re-audited the last three candidates —
`merge-queue-ops-2`, `lifecycle-ops`, `notification-service` — and all
three were already documented. Only one was reclassifiable, and by
**deletion** rather than conversion (#3205).
## What the bare command prints now
```
COLUMN guards (the backlog): 11
CONVERSION QUEUE EMPTY: all 11 remaining column guard(s) carry a documented deferral note.
There is no unexamined guard to claim. A nonzero backlog above is DEBT, not a work queue.
Re-read the note at a site before converting it; run --claims to also check open-PR ownership.
```
Or, when work does exist: `N unexamined guard(s) remain (no deferral
note) — run --triage to list them by file.`
**Local signals only**, so it is honest offline. It reports what it can
prove — no *unexamined* guard remains — and explicitly does **not**
claim the files are unclaimed, because only `--claims` sees open PRs. No
count, no exit code, `--strict`/`--json` untouched.
## Three commits, deliberately separated
1. **`refactor`** — move `FLAG_MARKERS` + the 40-line window into the
lib as `hasDeferralNote()`, verbatim. It was a private const plus an
inline `.slice()` in the CLI, so the rule deciding where the fleet is
sent had **no test in either direction**. Proven identical on the real
tree: `11 documented / 0 unexamined` before and after.
2. **`feat`** — the verdict + 6 tests.
3. **`fix`** — an unrelated pre-existing failure (below).
## The test fix — this one is turning main red
`attributes a remaining file to the open PR that touches it` asserted
over `out.slice(out.indexOf("UNCLAIMED:"))`, which runs to **end of
output** and so also covers the `SYNC-RESOLVED` section printed
afterward. That section legitimately lists `scheduler.ts`.
Latent until `topRemainingFile()` returned `scheduler.ts` — which
happened as the backlog shrank, **a state every conversion moves
toward**. Confirmed pre-existing: clean `origin/main` runs `42 passed /
1 failed` with the identical message.
## Evidence
| check | result |
|---|---|
| `hasDeferralNote` tests | both directions, boundary exact at 40 above
/ not below, 5 real phrasings |
| verdict control (by hand) | one tracked undocumented guard → **11 →
12**, verdict flips to `1 unexamined`; removed → restored |
| test-fix anti-vacuity | claim split broken → **FAILS**; restored →
passes |
| census file | **49 passed** (was 42 passed / 1 failed) |
| `census --strict` / `check:fnxc-future-dates` | exit 0 / exit 0 |
| `pnpm test:gate` | **exit 0** (732 tests) |
The verdict control was **invalid on the first attempt** — my probe file
was untracked and `git ls-files` never scanned it, so the verdict did
not flip and nothing was proven. Recording that because a control that
silently proves nothing is the exact failure this PR is about.
## Census before / after
No guard converted here; this is tooling. Backlog unchanged at 11, all
deferred.
|
||
|
|
23403e1426 |
test(engine): pin the agent sweep's terminal skip (21st resolver, after a corrected fixture) (#3206)
`agentLinkTerminalColumns` was uncovered on the #3115 map. The existing case uses `todo` and `in-progress`, so the terminal skip is never the deciding branch. ## The corrected fixture is the lesson An earlier attempt of mine put the card in a renamed **wip** lane and stayed green when blinded — **correctly**. Such a card is caught by the wip∪review set first, so the terminal resolver never decides anything. The card has to rest in a renamed **complete** lane for this guard to be the one that matters. That is the same class as my two discards on the branch-conflict sweeps: **the fixture has to reach the branch the resolver gates.** A test can exercise the sweep, pass, and still never touch the line under test. ## What the literal costs A finished task's agent is not skipped, so the sweep **unlinks an agent from a task that completed normally** — churn on a row that needed no repair, and a lost link if that agent was about to be reused. ## Observable Asserts `syncExecutionTaskLink` — the action the guard prevents — rather than a return value, per the rule from #3202. ## Measured 421 pass; blinding `agentLinkTerminalColumns` fails exactly this case. **21 of 26 pinned** across 20 merged PRs. ## Still open, with the obstacle recorded `reclaimHoldColumns` / `reclaimReviewColumns` resist: both audit paths in that sweep emit the same `branch:auto-reclaim` type, differing only by a `trigger` string the branch-level scan also produces, so no observable I found isolates the bucket resolvers from the branch scan. `agentParkedColumns` needs a fixture where parked-ness changes the outcome — mine forced the proof true via a fresh run. ## Verification `self-healing.test.ts` **421 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
d6079970e8 |
fix(self-healing): 18 recovery rebounds hardcoded todo and THREW on a renamed board (#3150, first slice) (#3152)
First slice of #3150. `self-healing.ts` held **26** `moveTask` calls with a legacy literal target; this converts the **18 `todo` rebounds**. ## Why this is worse than a guard, and documented already `task-store/moves.ts` records it from a previous incident: > `moveTaskInternal` **REJECTS** a target the workflow does not declare (`TransitionRejectionError: unknown-column`) … completion handoff did not silently no-op — it **THREW**. Every one of these 18 is a **recovery**. On a renamed board they threw instead of rebounding, so the strand each sweep exists to clear survived *and* the sweep reported failure. The reliability layer meant to be the backstop was the layer that broke. ## Why the census never saw it It counts **comparisons** against legacy ids. A move target is an **argument**. That is the third blind spot of the same instrument, and all three have now produced real defects found by hand: | blind spot | found this session | |---|---| | definitions | `GITHUB_TRACKING_EDITABLE_COLUMNS` — tracking unreachable on renamed boards (#3149) | | collections | swept: 30 sites, 29 already correct, 1 defect (the above) | | **targets** | **this** — 26 in one file, 31 tree-wide | ## Why 18 sites at once is safe `resolveReboundTargetForTask` **degrades to `"todo"`** when no workflow resolves, and `self-healing.ts` already used it at line 745. On every board we ship, the resolved answer *is* `todo` — so default behaviour is unchanged **by construction**, not by inspection. The control case pins exactly that, and it is the reason this can land as one change rather than eighteen. ## Scope, and what I deliberately did not touch Converted: the 18 `todo` rebounds. **Not** converted: the `done`, `archived` and `in-review` targets. They need different helpers and genuine reasoning about which lane a completion or an archive belongs in — converting them by analogy is exactly the half-conversion this program keeps paying for. Sites with no resolver in scope are unchanged. The audit behind the split is in the commit: of 26 sites, 5 had resolved lanes in scope, 4 had an IR, 17 had nothing — and `lanesOfReclaim` returns **Sets**, which is the wrong arity for a target (a move takes exactly one column, per the `moves.ts` note). ## Verification | | result | |---|---| | engine `tsc` | **0 errors** | | **all 43 self-healing suites** | **843 passed** | | census `--strict` | exit 0, **unchanged** — invisible to it | | `check-inert-sync-lanes` | exit 0 | | differential | restoring the literal → **1 failed \| 1 passed**, renamed case only | The new test drives a **public entry point** (`reconcileInReviewUnmetDependencies`, the FN-6793 contract) rather than calling the helper directly, so it covers the producer path too. One harness note worth keeping: the first version of the test failed **upstream** of the target, because the sweep selects rows via `resolveProjectColumnsForRoles` — a *project-level* resolver reading `listWorkflowDefinitions`, not the task's own selection. Without that mocked, the renamed card was never considered and the failure looked like the fix not working. That distinction (project-level vocabulary vs per-task IR) will bite the next slices too. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Tasks now move to workflow-specific rebound, completion, and archive columns instead of fixed default destinations. * Retrying and recovering tasks works correctly on boards with renamed lifecycle columns. * Added safe fallback behavior for workflows without custom lifecycle settings. * **Tests** * Added coverage to prevent legacy hardcoded task destinations and verify renamed-column recovery scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
97b945f980 |
docs(engine): the scheduler flag's reason went stale, and two deferrals read as unexamined (#3142)
My own flag on these two literals went stale, in exactly the way I have spent this session cataloguing in other people's notes. ## What the note said, and why it is now wrong It said these two stay because converting them would be **inert** — the sync resolver answers with the default board. True when written. #3128 then converted the rest of this listener by deferring each resolve into a `void (async () => ...)` block, which reaches the **async** resolver and is genuinely correct. So async resolution *is* available here now, and my stated reason no longer explains why these two are different. ## The real reason, which #3128 itself states Three branches down, in its own note: > The `planningTaskIds.delete` stays SYNCHRONOUS — it is the edge-trigger bookkeeping, and deferring it would let a second update re-enter this branch. Both remaining literals are that case: | literal | why it cannot move behind an await | |---|---| | `failedTaskIds.add` | edge-trigger bookkeeping raced against `moveTask` clearing the failure metadata — its own comment says so. Deferring the add can miss that window. | | PR-monitoring guard | it gates `getTrackedPrs()` / `startMonitoring()`, where `tracked.has(task.id)` **is** the re-entrance guard. Move the lane answer behind an await and two updates for the same task can both pass that check before either starts — **double-starting a monitor**. | ## Why the distinction is worth a PR "Blocked on a resolver" invites the next person to wait for the sync reader. What these actually need is somewhere to put the answer that is **not behind an await** — the emitter-carried `lanes` #3109 added to `task:moved`, whose extension to `task:updated` is measured as expensive rather than impossible (#3123: 26 emit sites against 7, on the hottest write path). Those are different tickets with different owners. Leaving the wrong one written down is how a blocker outlives its cause — the failure I have now found in five separate notes this session, including two of my own. ## Measured - Comment-only. - `src/__tests__/scheduler*` — **14 files / 144 tests pass**. - `tsc --noEmit -p packages/engine` clean; `check-inert-sync-lanes` and census `--strict` clean. - `check-fnxc-future-dates` is red from `main`'s own #3128 stamps — **#3139** fixes that; this branch inherits it and does not add to it. ## Census No movement. Both literals stay counted, now with the correct reason attached. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aef88a2976 |
test(engine): pin the PR-conflict sweep's worktree-owner index (20th resolver, after two discarded attempts) (#3202)
`prConflictWipColumns` builds the worktree-owner index behind `ownedByOtherInProgressTask` — the guard that stops this sweep **deleting a worktree another live task is executing in**. Keyed on the id, that index is empty on a renamed board, so every worktree reads as unowned. ## Two discarded attempts, and why they matter more than the fix **1. Asserted `result.outcome !== "reclaimed"`.** It failed *with the fix in place* — `reclaimed` is reachable through a second path this guard does not gate. **An outcome assertion cannot isolate a guard in a sweep with several routes to the same outcome.** That also explains my earlier discard on `reclaimSelfOwnedBranchConflicts`, which has the same shape. **2. Asserted `removeWorktree` was not called — but overrode the task's branch while leaving its id.** The reclaim path also requires `branchOwnerTaskId === taskIdUpper`, so the branch was never reachable and the case passed **blinded**: vacuous for a reason that had nothing to do with lanes. The shipped version asserts `removeWorktree`, which runs **only** on the guarded branch and is the irreversible part, and keeps the default id/branch pair so that branch is genuinely reachable. ## Measured 16 pass; blinding `prConflictWipColumns` fails exactly this case. **20 of 26 pinned** across 19 merged PRs. ## Generalisation For sweeps with multiple paths to one outcome, the discriminating observable is a **path-specific side effect** — `removeWorktree`, a `task:reconcile-*` audit type, a specific `reason` string — not the return value. Every case I landed today that stuck used one; both discards asserted a return value. ## Verification `self-healing-pr-conflict` **16 passed** · `pnpm test:gate` 13 + 161 + 487 + 71 · lint — green. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved protection for active worktrees during pull request conflict recovery, including tasks in renamed workflow lanes. * **Tests** * Added regression coverage to verify that worktrees owned by other tasks are not removed incorrectly. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
9234ca2402 |
fix(triage): the startup sweep resolved its columns from a SENTINEL task id — a characterization test already pinned it (#3201)
Fourth and last convertible site in `triage.ts`. This one needed a
different fix from the other three, and the codebase already said so.
## The defect
```ts
const sweepLanes = resolvePlannerLanes(this.store, "");
const sweepColumns = [...new Set(["triage", "todo", sweepLanes.intake, sweepLanes.hold])];
```
There is no task `""`. No selection can be read for it, no board
resolved — the lanes come back as the **default** board's and the union
collapses to the legacy pair `{triage, todo}`. On a renamed board the
sweep queries columns the card is not in, so its stale `planning` status
survives and it **holds a planning admission slot indefinitely**.
## A characterization test already pinned this, and called the fix
correctly
`workflow-sweep-sentinel-task-id-live-e2e.pg.test.ts` documents it as a
third inert-conversion mechanism — *"inert by construction rather than
by environment"* — and its header says:
> making `resolvePlannerLanes` async would **NOT** repair this site,
because the defect is the argument, not the resolver
That is right, and it is why this fix differs from #3191 / #3193 /
#3195, which all used the async twin. Here the sweep has **no task to
resolve against** and wants every column playing these roles **anywhere
in the project** — so the correct resolver is
`resolveProjectColumnsForRoles(store, ["intake", "hold"])`, the same
helper `self-healing.ts` already uses for the same purpose.
The legacy pair stays in the union deliberately: the note at the site
explains that `triage` and `todo` must both be swept for pre-U11 and
Coding (Ideas) rows, and extra columns are free because the sweep only
**reads** and filters on `status === "planning"` first.
## The test is inverted, not deleted
It asserted `"planning"` survives — the bug. It now asserts the status
is cleared. Keeping the case with its original reasoning intact
preserves the file's record of what the defect *was*.
## Measured
| | result |
|---|---|
| broad suite (triage / planning / self-healing) | **76 files, 1302
tests passed** |
| differential | restoring the sentinel call → **1 failed \| 2 passed**
|
| `census --strict`, `check-fnxc-future-dates` | exit 0 |
**Inert count unchanged at 4 for `triage.ts`** — these lanes fed an
*array literal*, not a comparison, so the ratchet never counted them.
Third fix this session in that blind-spot class, stated so the number is
not read as the whole picture.
## What remains in this file
Two sites: the `task:moved` wake handler and the evacuation handler —
both **synchronous arrow callbacks** whose answers are consumed in-tick.
Genuinely blocked on the emitter-side work in #3082, with corroborating
evidence attached there.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
09edce2366 |
test(engine): cover #3112's executor lane conversion — three renamed-board cases main lacks (#3118)
**I flagged these four as unconvertible in #3104. #3109 landed and dissolved both of my reasons, so the flag comes off.** Leaving a "blocked" note standing behind a blocker that no longer exists is the exact decay this program keeps paying for — I have now found three other people's deferrals in that state this session, and I am not adding a fourth of my own. ## Both blockers, and why they are gone | my stated blocker | why it is gone | |---|---| | **A.** `trackTaskDisposal` writes `pendingTaskDisposals` in *this* tick, and the wip branch reads that map to serialise a fast bounce (FN-5256). Deferring branch selection to a microtask reopens that race. | Reading `lanes` off the payload needs **no await**. The prologue stays synchronous and the race stays closed. | | **B.** It is an if / else-if **chain**, so the guards are entangled and convert together or not at all. | They convert together here. | #3109 made the **emitter** carry the resolved lanes, which is the one route that removes the dilemma instead of trading one horn for the other. `lanes` is optional and fail-soft to `undefined` — *"unknown, never legacy"* — so each guard keeps its literal as the fallback, following the `mergeParkedColumns` convention #3109 established in `scheduler.ts`. An emit path that cannot resolve is no worse than before. ## What it fixes On a renamed board: execution never started on a move into the board's own wip lane, terminal session release never ran on a move into its archive lane, and neither `from` guard fired — so in-flight work was not aborted when a card left implementation. Nothing errored; the engine simply stopped reacting. ## Census | | before | after | |---|---|---| | `executor.ts` | 4 | **0** | | repo backlog | 45 | **41** | ## Measured - 3 new cases added to the FN-7717 suite; file **13/13 pass**. - **MUTATION**: restoring the `archived` literal fails the renamed case. - **The paired negative is the load-bearing one.** `done`/`in-review` deliberately keep their merge leases across the transition (FN-6736 / Phase C–D). The renamed **complete** lane must therefore *not* release — a conversion that released on every terminal-ish lane would satisfy the positive case and quietly break the guarantee that file already exists to protect. - A **fail-soft** case pins that an emit carrying no `lanes` behaves exactly as before. - `src/__tests__/executor*` — **84 files / 853 tests pass**. - `tsc --noEmit -p packages/engine` clean; census `--strict`, `check-lane-wiring`, `check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean. ## Note on #3104 That PR (merged) added the flag and the sharpened reasoning. This one removes it. The reasoning there was correct at the time and is what made it possible to check quickly whether #3109 actually addressed it — a flag that states its blocker precisely is cheap to retire, which is the argument for writing them that way. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
339d4451af |
test(engine): drop four redundant sync-reader stubs — they fed the broken reader the right answer (#3198)
Completes the audit filed as #3197. ## Why a stub here is not neutral `resolveTaskWorkflowIrSync` answers with the **default board for every task** in production — its selection reader returns `undefined` unconditionally under PostgreSQL. A test that stubs it with a working IR proves its call site's *logic* while being structurally unable to notice that the real path resolves nothing. The suite stays green even if the site goes inert, which is the failure this whole phase has been chasing. ## The audit, complete Deleted each stub and checked whether the suite still discriminates: | file | without the stub | verdict | |---|---|---| | `planner-lane-resolution` | 7 passed | redundant → **removed** | | `triage-undeclared-column-rescue` | 7 passed | redundant → **removed** | | `recover-approved-intake-post-u11` | 6 passed | redundant → **removed** | | `workflow-scheduler-parked-columns-live-e2e.pg` | 2 passed | redundant → **removed** | | `planner-lanes-async-resolution` | 1 failed | **legitimate** — the stub is its subject | | `scheduler-renamed-hold-events` | **3 failed** | **masking** — see #3082 | | `triage.test.ts`, `triage-release-renamed-hold` | — | resolved in #3191 / #3193 / #3195 | **Only the redundant four are touched.** `planner-lanes-async-resolution` stubs the reader *deliberately*, to contrast the two resolvers given the same store and task — removing it would delete the point of the file. That is the case that makes this a hand audit rather than a ratchet: a hit is not presumptively a defect. `scheduler-renamed-hold-events` is left alone because its three failures **are the finding, not the fix**. They correspond to the 13 inert guards `check-inert-sync-lanes` counts in `scheduler.ts` — two independent instruments agreeing that those handlers are green in tests and dead in production on renamed boards. They live in synchronous `task:*` listeners, so they need the emitter-side work in #3082, not a stub edit. ## Verification - 4 files / **22 tests pass** without the stubs - engine `tsc` 0 errors - `census --strict`, `check-inert-sync-lanes`, `check-fnxc-future-dates`: exit 0 The PG e2e was the one file I had marked unaudited when filing #3197; it ran here and is included rather than left as an open question. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f5f296c11 |
test(engine): pin the workspace land-lease owner check (19th resolver) (#3199)
Nineteenth resolver from the coverage map on #3115. The terminal-owner reclaim directly above it proves the behaviour with `column: "done"` — **the id** — so blinding `leaseOwnerCompleteColumns` left all 20 tests green. ## What the literal costs That set is what `isWorkspaceOwnerLive` consults. Keyed on the id, an owner resting in a renamed completion lane reads as **live**, so its land lease is never reclaimed. The workspace repo stays leased by a task that has finished, and **every later land against that repo waits behind a phantom**. ## A note on how this resolver came to exist `isWorkspaceOwnerLive` is one of the sites I flagged earlier today as **unconvertible** — synchronous, no store handle, converting it would mean a signature change I had excluded from that PR's scope. Someone threaded the resolved set through its callers instead. That is the better answer than either converting in place or leaving it, and this test pins it — so the threading cannot be undone silently. ## Measured 21 pass; blinding `leaseOwnerCompleteColumns` fails exactly this case. **19 of 26 pinned** across 18 merged PRs. ## Verification `self-healing-workspace` **21 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
73bff5f88c |
test(engine): pin the orphan-only sweep's project query — the harness could not see a filter bug (#3196)
Eighteenth resolver from the coverage map on #3115, and **why** it was uncovered is the interesting part. ## A fake that ignores its own filter cannot see a filter bug Every case in this file stubs `listTasks` to return the same task **whatever column is asked for**: ```ts (store.listTasks as ...).mockResolvedValue([failedReviewTask()]); ``` So the project query is never exercised. Blinding `orphanReviewColumns` changes which column is *requested*, the fake answers identically, and nothing fails. Eight passing tests, and the selection logic among them was untested. That is the same blindness the production sweep had — querying a column that does not exist and finding nothing — reproduced in the harness that was supposed to catch it. ## The case `listTasks` honours the column, so a card resting in a renamed review lane is found **only if the query asked for that lane**. Keyed on the id, the sweep asked for `in-review`, got nothing, and a failed orphan-only card **stayed failed forever**. ## Measured 9 pass; blinding `orphanReviewColumns` fails exactly this case. **18 of 26 pinned** across 17 merged PRs. ## Generalisation worth checking elsewhere Any sweep whose test stubs `listTasks` with a flat `mockResolvedValue` has this hole. The fix is a store fake that filters on `options.column` — the shape `self-healing-query-filter-blindness.test.ts` already uses. I would look there first for the remaining map entries. ## Verification `self-healing-orphan-only-scope` **9 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
a6d67844b8 |
fix(triage): the "unconvertible" site was convertible — the blocker was two test harnesses (#3191)
#3141 measured this site as unconvertible, and I twice reported the cause as a production constraint. It was not. This is the instrumented answer to the probe I recommended there and then ran myself. ## The isolation | configuration | result | |---|---| | flag only, no conversion | **8 passed** → the orphan arm is *not* the cause | | flag + conversion | **5 failed** → the conversion is | | same, with a realistic mock store | **8 passed** → the mock was the cause | `triage-stuck-requeue-preserve-draft.test.ts` defined neither `getTaskWorkflowSelection` nor its async twin — exactly like `triage.test.ts` did before #3189. Both made `resolveWorkflowIrForTaskWithProvenance` **throw** and take its catch branch: the *"could not ask"* shape, which a production store never presents. So the 5 failures I deferred as a possible semantics change were the same harness gap in a second file — confirmed, not argued. ## What changes **`selectionAbsent`** marks the determinate case: the store *answered* "no selection", so the workflow is the default and its IR is in hand. Added as a **separate field, not a third `source` value** — `source === "default"` is compared in **31 places** in `self-healing.ts` meaning "be conservative", and a new enum value would silently stop matching every one of them while still compiling and still passing on a default board. **`recoverApprovedTask`** now accepts a legacy `triage` row *explicitly* (its workflow does not declare that column) instead of depending on `resolvePlannerLanes` **failing** and falling back to legacy ids. Correctness resting on a resolver's failure mode is what this removes. ## Measured | | result | |---|---| | broad suite (triage / self-healing / recovery / planning) | **77 files, 1302 tests passed** | | the three directly affected suites, post-rebase | **245 passed** | | the flag is load-bearing | conversion **without** it: **18 failed \| 221 passed** | | `census --strict`, `check-fnxc-future-dates` | exit 0 | **The inert-sync-lane count is unchanged at 7 for `triage.ts`.** This site was never among the counted guards, so this is **not** a ratchet reduction — stating that rather than letting a conversion imply one. It removes a real inert dependency the ratchet cannot see, which is the blind-spot class this phase has been mapping. ## Why this took four attempts I described this blocker at four levels: merged intake/hold, orphan-arm scoping, identity verification (filed as **#3187**, closed as wrong), and finally the harness. **The two I instrumented held; the two I reasoned to did not.** The fix here is the probe I wrote down for someone else — which is where it should have started. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c8268a6454 |
test(engine): pin the temp-merge sweep's terminal grace (17th resolver) (#3192)
Seventeenth resolver from the coverage map on #3115. The two cases around this one use `done` and `archived` — **the ids** — so blinding `mergeTempTerminalColumns` left all 21 tests green. ## What the literal costs The terminal check selects the **shorter grace**: a finished task's temp merge worktree is reaped after `DONE_TASK_TEMP_WORKTREE_GRACE_MS` instead of the full stale window. Keyed on the ids, a card in a renamed completion lane never qualified, so its worktree lingered for the long window — **disk held by work that already finished**. ## The second cost, which is why this asserts on the audit reason Without the resolver the sweep eventually acts, but records `reason: "stale"` instead of `"done-task-stale"`. So its own trail **misattributes why it acted**. A sweep that does roughly the right thing under the wrong label is the kind of defect nobody notices until they are reading audit events during an incident — and then the record actively misleads. Asserting only on the file being gone would have passed either way. ## Measured 22 pass; blinding `mergeTempTerminalColumns` fails exactly this case. **17 of 26 pinned** across 16 merged PRs. Also re-measured this turn: `wsDoneColumns` and `doneMetaColumns` have gone green independently, so the map keeps drifting as the fleet adds coverage — re-run before picking the next entry. ## Verification `self-healing-tempdir-sweep` **22 passed** · `pnpm test:gate` full pass · lint — green. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed stale-task handling for tasks in terminal columns of custom workflows. * These tasks now correctly follow the done-task grace period and record the appropriate audit reason. * **Tests** * Added regression coverage to verify the corrected behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
f755f44734 |
test(engine): pin the completion fan-out's review dependent bucket (16th resolver) (#3190)
Sixteenth resolver from the coverage map on #3115. `completedReviewColumns` reads the **dependents** resting in review when a blocker completes. No case in this file put a dependent in a renamed review lane, so blinding it left all 13 tests green. ## What the literal costs A dependent sitting in review is never read, so its `blockedBy` is never cleared when the blocker finishes. **It stays blocked by work that is already done** — the most visible form of this class, because the board simply stops moving. ## Measured 14 pass; blinding `completedReviewColumns` fails exactly this case. ## Note for anyone continuing the map `completedHoldColumns` in this same sweep measured as **already covered**, so only the review bucket was owed. Three buckets, three resolvers, covered independently — the same per-resolver granularity that found the missing halves in #3138 and #3186, where my own earlier tests pinned one resolver of a pair and I had recorded the sweep as done. **16 of 26 pinned** across 15 merged PRs. ## Verification `self-healing-completion-fanout` **14 passed** · `pnpm test:gate` full pass · lint — green. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed task completion reconciliation for workflows with renamed lanes. * Dependent tasks in review lanes are now correctly unblocked when their blocker moves to a custom completion lane. * **Tests** * Added regression coverage for custom workflow lane configurations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
da10131d3f |
test(triage): the mock store could not be ASKED for a selection — 231 cases exercised a shape production cannot produce (#3189)
`createMockStore` in `triage.test.ts` defined **neither** `getTaskWorkflowSelection` nor its async twin. So `resolveWorkflowIrForTaskWithProvenance` **threw** calling them and took its catch branch, reporting `source: "default"` in the sense of *"the lookup failed"*. Production stores always expose both readers — every case in this file was exercising a store shape that cannot exist. Returning `undefined` models the real answer: the store **can** be asked and says there is no selection row, which is what a pre-U11 card actually presents. ## Why it mattered `triage.ts`'s post-U11 intake recovery gates on that provenance. A *failed* lookup correctly refuses to claim a workflow lacks `triage`, so the orphan arm stayed off and the recovery depended on `resolvePlannerLanes` **failing** and falling back to legacy ids — correctness resting on a resolver's failure mode. In #3141 I measured the async conversion of that site as failing 13 cases and **twice reported it as a production constraint**. It was this harness. That is the concrete cost of a mock that cannot answer a question production always can. ## Behaviour-preserving on its own **380 passed across 26 triage/recovery suites.** ## What this deliberately does NOT do It does not convert the site. I prototyped the full unblock — a `selectionAbsent` flag on the determinate `!workflowId` branch, its single consumer, and the async conversion — and it works: the previously-failing suite goes **237 passed**. But with a realistic store the orphan arm starts firing for no-selection rows, which changes recovery flow in **5 `triage-stuck-requeue-preserve-draft` cases** that currently assert the refusing behaviour. Whether accepting a legacy `triage` row there is correct is a lifecycle-semantics decision about migration, not a harness fix. So it is reverted and reported rather than bundled. Findings and the measured branch table are on #3141. ## One correction carried from this work I filed #3187 claiming provenance verifies resolution via `ir.id === workflowId`, which cannot pass for builtins. **That was wrong** — the live code uses a symbol marker, and the text I quoted was historical prose describing what was removed. Closed with the measurement: ``` store with NO selection readers -> source: default (catch: could not ask) store answering builtin selection -> source: selection ✓ ``` That is the same class of error this PR fixes — reasoning from what something says rather than what it does. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Improved workflow-resolution test coverage by supporting stores with no selected workflow. * Added synchronous and asynchronous test readers for workflow selection. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
893b6421be |
test(engine): pin the contamination sweep's WIP bucket (15th resolver) (#3188)
Fifteenth resolver from the coverage map on #3115. Every case in this file seeds the candidate in `in-review`, so only the review bucket was exercised — blinding `contaminationWipColumns` left the file green. ## Why the WIP bucket matters A card sent back for a fix **re-enters execution while its branch still carries the foreign commits**, so contamination is discovered there as often as in review. Keyed on the id, that bucket read nothing on a renamed board and the card kept a branch built on someone else's work — which is what this sweep exists to re-anchor. ## Two facts the fixture had to learn, both from failing first - **This is an ACTION site and deliberately skips a card whose own board cannot be read**, rather than guessing from the project union. A fake with only `listWorkflowDefinitions` resolves the default IR, the card is reported unclassifiable, and the case fails for a reason unrelated to the resolver under test. The per-task selection readers are required. - **The WIP bucket's predicate is not the review bucket's.** It additionally requires `paused === true` with `pausedReason` of `branch-cross-contamination` or `branch-conflict-unrecoverable`. A card merely resting in the wip lane is not a candidate — the FN-5704 manual-review contract this sweep mirrors. Neither is guessable from the resolver. Both came from the test failing twice, and I would have shipped something that exercised nothing had the first version passed. ## Measured 3 pass; blinding `contaminationWipColumns` fails exactly this case. **15 of 26 pinned** across 14 merged PRs. ## Verification `self-healing-foreign-only-contamination` **3 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
6bc90ccbe2 |
fix(core): allow-list the legacy workflow IR — it found a fourth bug my grep missed (#3185)
## The name is the defect `BUILTIN_CODING_WORKFLOW_IR` reads like the default and **is** the legacy workflow (`builtin:legacy-coding`). Post-U11 they differ by exactly one column — `triage` — the one a caller most often wants absent. **Four bugs have come from reaching for it by name:** 1. two move-path resolvers disagreed on the no-selection default → *"workflow move policy preflight is stale"* on every flag-on move (recorded in `resolveDefaultWorkflowIr`'s own header) 2. the TUI board rendered a `triage` lane the default board lacks — #3178 3. `deleteWorkflow` re-homed occupants into `triage` — #3183 4. **`board-workflows.ts`** described a *custom* workflow whose definition failed to load using legacy columns — the #3178 symptom through the dashboard route. **Fixed here.** It type-checks, it is the obvious identifier, and on the five shared columns it behaves correctly. The mistake only shows on the column that differs. ## I said the sweep was complete last round. It wasn't. My grep excluded paths and truncated at `head -10`; it missed two sites. **The allow-list found both on its first run.** That is the lesson the sibling sync-resolver ratchet already records — *"FOUND BY THIS RATCHET, not by the grep that seeded the list"* — and I had just quoted that file while repeating the mistake. ## One site is allow-listed rather than fixed, and I tried the fix first `workflow-graph-executor.run()`'s default `ir` is unreachable in production (both callers pass it explicitly). But `workflow-graph-executor-parity.test.ts`, in the **engine-core gate suite**, drives the method *without* the argument to assert the historical seam sequence. Switching it to the catalog default rewrites what "parity" means: **measured, 6 gate tests fail** with `expected 'failure' to be 'success'`. Reverted, and recorded at the call site *and* in the allow-list entry so nobody repeats the experiment. That is what an allow-list is for: a legitimate narrow use next to a plausible-looking wrong one. ## Guard construction Follows the repo's existing call-site allow-lists (sync resolver, engine blocking-shellout, detached-spawn script guard). - **Comments stripped before scanning** — `activity-analytics.ts` and `TaskContextMenu.tsx` name this constant in notes *about past bugs* while correctly avoiding it. Counting prose would train readers to allow-list mentions. - **Anti-vacuity**: the scan still sees the catalog's own uses, so a renamed constant or broken walker cannot make the guard pass by finding nothing. - **Stale-entry**: the list cannot rot into files that no longer touch it — the decay every ledger in this repo has hit. ## Measured - Guard **3/3**; `tsc --noEmit` clean in core, engine, dashboard. - census `--strict`, `check-fnxc-future-dates` clean. ## Census **No movement — that is the point.** This class has no column literal to count, which is why the census never saw any of the four bugs. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
39a2e0481a |
test(engine): pin the merged-review sweep's HOLD bucket (14th resolver — the half my own test missed) (#3186)
Fourteenth resolver from the coverage map on #3115, and it is the other half of a sweep **I converted and tested myself**. The file already pinned `mergedReviewColumns`. Blinding `mergedHoldColumns` back to `["todo"]` left all 71 tests green — no case put a merge-confirmed card in a renamed hold lane. ## The lane is not hypothetical A merge-confirmed card gets **rebounded to hold** by other recovery paths — a failed post-merge step, a requeue. So *merged but sitting in hold* is exactly the state this sweep's second bucket exists to finalize. Keyed on the id, that bucket read nothing on a renamed board and the card stayed unfinished **while its commit was already on the base branch**. ## The lesson, repeated This is #3138's finding again: a test that pins one resolver of a pair reads as covering the sweep. I wrote the earlier case, recorded the sweep as done, and it was half-done. **Only blinding each resolver separately finds this.** A single passing revert proves one guard — which is why the map is keyed by resolver, not by sweep. ## Measured 72 pass; blinding `mergedHoldColumns` fails exactly this case. **14 of 26 pinned** across 13 merged PRs. ## Verification `self-healing-query-filter-blindness` **72 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
9c00699e61 |
test(engine): pin the stalled-card watchdog's terminal skip on a renamed board (13th resolver) (#3182)
Thirteenth resolver from the verified coverage map on #3115. `sweepTerminalColumns` was uncovered: the case directly above it asserts the terminal skip using `done` and `archived` — **the ids** — so blinding the resolver left the file green. ## What the literal costs The skip matched nothing on a renamed board, so **finished cards were scanned as live**, and a card parked in a renamed completion lane could be reported stalled. A watchdog that cries about completed work is worse than a quiet one: it trains operators to ignore the alert. That is the exact failure this sweep's own dedup logic was built to avoid, reintroduced through the lane vocabulary. ## The case The renamed twin of the existing terminal-skip test — same assertion, same shape, different vocabulary. That is the whole point: the original passes either way, so it cannot see the conversion. **Measured:** 10 pass; blinding `sweepTerminalColumns` fails exactly this case. ## Map status **13 of 26 pinned** across 12 merged PRs, plus `starvedWaitingColumns` now covered by another worker independently. Re-measure before picking the next one — the map drifts green as the fleet adds coverage, and I have already caught it stale once today. ## Verification `self-healing-stalled-card-watchdog` **10 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
5eec7dc73b |
test(engine): pin the completed-blocked park release on a renamed board (12th resolver) (#3180)
Twelfth resolver from the verified coverage map on #3115. `completedBlockedHoldColumns` was uncovered: every case in this file seeds the park in `todo`, where the literal is correct, so blinding the resolver left all 21 tests green. ## What the literal costs A completed-blocked park rests in the board's **hold** lane, which is only called `todo` on the built-in workflow. Keyed on the id, the sweep selects nothing on a renamed board — so **finished work stays parked behind a blocker that has already cleared**, stranded exactly as FN-7926 describes. Silently: a sweep that selects no rows reports success. ## Two fixture facts, found by the test failing first - **The completion-blocker gate resolves the *blocker's* own workflow**, so the per-task selection readers are required too. `listWorkflowDefinitions` alone leaves the renamed complete lane unrecognised and the park is rejected for the wrong reason — a green-for-the-wrong-reason test, which is the exact thing this effort removes. - **The blocker must rest in the renamed complete lane**, not the legacy one, or the case proves nothing about the board it claims to test. I only learned both because the first version failed. Had it passed, I would have shipped a test that exercised none of this. ## Measured 21 pass; blinding `completedBlockedHoldColumns` fails exactly this case. ## Map status **12 of 26 pinned.** Also re-measured six entries this turn: `starvedWaitingColumns` is **now covered by another worker's test** (#3128-era, peer-progress vocabulary), so the map is drifting green underneath me as the fleet adds coverage too — worth re-running before anyone picks the next entry. ## Verification `execute-requeue-loop-guard` **21 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
f8fb9b1473 |
test(engine): pin the unmet-dependency rebound on a renamed board (11th resolver from the coverage map) (#3176)
Eleventh resolver from the verified coverage map on #3115. `unmetDepReviewColumns` was uncovered: the existing FN-6778/FN-6779 case uses `in-review`, where the literal is correct, so blinding the resolver left the file green. ## What the literal costs The sweep selects **no card**. A review card whose dependency is still unmet is never rebounded — it sits in review, **eligible for merge, ahead of the work it depends on**. That is precisely the ordering violation this sweep exists to prevent, and it fails silently: no error, no audit event, nothing to notice. ## Measured 3 pass; blinding `unmetDepReviewColumns` fails exactly the new case. ## Map status **11 of 26 resolvers pinned** across 10 merged PRs. The remaining 15 need real harness work — I threw away two probes earlier today that passed while proving nothing (`reconcileInReviewBranchRebind` never entered its loop; `recoverAgentsRunningOnInactiveTasks` stayed green under both blindings), and recorded them on #3164 rather than shipping green decoration. ## Verification `in-review-unmet-dependency-reconcile` **3 passed** · `pnpm test:gate` full pass · lint — green. |
||
|
|
10a0c5848f |
fix(executor): planner-evacuation lanes come from the emitter — executor leaves the inert list (16 → 12) (#3137)
`executor.ts` was the last file besides `triage.ts` and `scheduler.ts`
on `check-inert-sync-lanes`, holding **4 guards that read as converted
and behave as literals**. Neither cause turned out to be "needs an async
resolver".
## 1. Two of the four were in code with no caller
`isPlannerColumnFor` is a **private method with zero production
callers**. `tsc` reports it unused; the only things reaching it were two
tests casting through `executor as unknown as { … }`, which is exactly
what let it look alive. Its doc comment described the
planning-evacuation branch — but that branch calls
`isBackwardMoveOutOfPlanning` and never called this.
Deleted, along with the two tests whose subject it was. Converting
guards in unreachable code would have "fixed" behaviour that cannot run
and left two more sites to maintain; a test whose subject has no caller
pins nothing.
## 2. The other two no longer need to resolve anything
`isBackwardMoveOutOfPlanning` resolved its own lanes via
`resolvePlannerLanes`, whose selection reader returns `undefined`
unconditionally under PostgreSQL — so it answered with the **default
board for every task**, and both its guards were inert.
Its comment justified the sync resolver by the synchronous `task:moved`
emitter. That was true and **is no longer binding**: the emitter now
resolves lanes once, asynchronously (`moves.ts` →
`resolveWorkflowIrForTask`), and hands them on the payload — which #3112
already reads in this same listener. Reading a parameter is as
synchronous as reading `from`, so nothing reorders and no listener
resolves.
`lanes` is **required, not optional**. An optional parameter that the
one production caller happens to pass is the seam-with-no-supplier shape
this program keeps finding; required means a future caller fails
typecheck instead of silently getting a default board. When the emitter
itself could not resolve, the legacy ids answer — exactly what
`resolvePlannerLanes` degraded to anyway.
## Measured
| | before | after |
|---|---|---|
| `check-inert-sync-lanes` | **16** guards, 3 files | **12** guards, 2
files |
| `executor.ts` on that list | 4 | **0 — off the list** |
| census | 18 | 18 (`--strict`: every file matches baseline exactly) |
**The census is deliberately unchanged.** This targets the inert
population, which the census cannot see by construction: those guards
already read as converted. That gap is the argument in #3082 — 12 guards
still behave as literals while the census shows them as done.
## The producer half, which I nearly shipped without
The predicate's own suite covers it thoroughly — and every case calls it
**directly**. Mutation testing exposed that this proves nothing about
the listener: replacing the listener's `lanes` argument with `undefined`
left `planning-evacuation` at **20/20 green**. That is the fifth failure
shape in this program's learnings verbatim — a converted consumer with
an unconverted producer passing every instrument.
So there is now a case driving the **real listener** on a board whose
planner lanes share no id with the legacy pair (`queued` holds,
`drafting` intakes), withdrawing a card to a non-lifecycle column — the
reported symptom (`todo -> Ideas`) in that board's vocabulary.
## Verification
- engine `tsc` — **0 errors**
- `executor-planner-lanes-resolved` — **12 passed**
- `executor-archive-releases-active-session` — **14 passed**; listener
passing `undefined` → **1 failed | 13 passed**
- `planning-evacuation` + `triage-planning-wake` + archive suite — **47
passed**
- `check-inert-flag-seams`, `check-fnxc-future-dates`, census `--strict`
— exit 0
- `eslint` on changed files — 0 errors
The predicate tests are also **stronger than before**, not merely
adapted: they now build lanes with `toTaskMoveLanes`, the same function
`moves.ts` uses for the payload. Previously they reached the predicate
through the store-backed sync reader, so renamed-lane assertions passed
in the harness while the real path could never see a renamed lane.
## Not done here
The inert baseline still reads 29 against a tree of 12 and the gate
advises re-recording. I left it: a stale allowance is a real hazard, but
re-recording is a one-line change that conflicts with every lane, and it
should land once rather than in each of our branches.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4365a3b10b |
test(engine): pin the paused-scope-decay lane filter (plus two probes I threw away) (#3164)
Next entry from the verified coverage map on #3115. `scopeDecayWipColumns` was uncovered — the existing case uses `in-progress`, where the literal is correct, so blinding the resolver left all 420 tests green. ## What the literal costs A paused holder resting in a renamed wip lane is **never selected**. The loop does not run, nothing is recorded, and its file scope decays with nothing to rebound it — so its followers stay blocked behind a card that is not coming back. ## The observable The audit event. Reaching a no-action record proves the holder was **selected by the lane filter**, which is the only thing this resolver controls. Asserting on the rebound itself would have needed triple-proof to succeed, dragging in state the resolver has nothing to do with. **Measured:** 420 pass; blinding `scopeDecayWipColumns` fails exactly this case. ## Two attempts thrown away first Worth recording, because the remaining map entries are not uniform with the ones already closed: - **`reconcileInReviewBranchRebind`** — a git-free probe (workspace task, rejected before any git runs) returned `{outcomes: [], repaired: 0}`. The loop never ran. I confirmed the `merge` trait does map to `mergeOrchestration`, so the filter should have matched; something else short-circuits and I could not establish what. - **`recoverAgentsRunningOnInactiveTasks`** — the test passed, then **both** resolvers stayed green when blinded. `agentLinkTerminalColumns` never fires because the live card is caught by the wip∪review set first; `agentParkedColumns` only feeds `evaluateParkedAgentTaskLink`, whose result my fixture already forced true via a fresh run. Both would have been green, plausible, and worthless. They were reverted rather than adjusted until they passed — which is the failure this whole effort exists to remove, and the one I committed myself in #3078. ## Map status Closed: 9 resolvers across 7 PRs. **~18 remain.** The easy ones are done; what is left needs real harness work — an `execAsync` git fixture, and understanding how `evaluateParkedAgentTaskLink` weighs run-freshness against lane. Budget for that rather than expecting the pattern that closed the first nine. ## Verification `self-healing.test.ts` **420 passed** · `pnpm test:gate` 161 + 13 + 487 + 71 · lint — green. |
||
|
|
44a67df65d |
test(engine): pin both dependency-lease resolvers on a renamed board (one case covered only half the conversion) (#3138)
Next two entries from the verified coverage map on #3115. `reconcileDependencyBlockingLeases` had **both** of its resolvers uncovered — blinding either `leaseWipColumns` or `leaseHoldColumns` back to its legacy id left all 825 self-healing tests green, because every fixture in that block uses `in-progress` / `todo`, where the literals happen to be correct. ## What the literals cost The holder scan matches no card **and** the dependency scan matches no card. A stale file-scope lease blocking a real dependency is never rebounded, so the dependent stays `overlapBlockedBy` behind a holder that is not coming back. That is the deadlock this sweep exists to break — silently not broken, no error, no log. ## Two cases, because one did not cover both — measured, not assumed My first case (holder in a renamed wip lane, dependency marked `overlapBlockedBy`) pinned `leaseWipColumns`. I then blinded `leaseHoldColumns` against it and **it stayed green**. The reason is in the control flow: the `overlapBlockedBy === holder.id` branch short-circuits and `break`s **before** the hold membership is consulted. So that fixture can never reach the guard `leaseHoldColumns` feeds. The second case drops the marker, leaving an unmarked dependency resting in a renamed hold lane, which falls through to the overlapping-hold-dependency branch. | blinded | result | |---|---| | `leaseWipColumns` | **1 failed** | | `leaseHoldColumns` | **1 failed** | Before the second case, that table read `1 failed` / `still green`. Checking each resolver separately is the only reason I noticed — a single "the suite fails when reverted" would have looked like proof and covered half the conversion. ## Remaining 23 uncovered resolvers on the map. Next by risk: `reclaimStaleActiveBranches` (deletes branches) and `reconcileInReviewBranchRebind` (rebinds branches of live cards), both needing a git-shelling harness. ## Verification `self-healing.test.ts` **415 passed** · `pnpm test:gate` 13 + 161 + 487 + 71 · lint — green. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved recovery of stalled workflow tasks when dependency-blocking leases become stale. * Added support for workflows using customized task status lanes, including marked and unmarked overlap blockers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
7cdd3f3e87 |
test(engine): pin the last two worktree-metadata resolvers — completes the sweep with 3 of 3 uncovered (#3148)
Completes `reconcileTaskWorktreeMetadata` — the sweep with the most uncovered resolvers on the #3115 map (3 of 3). #3132 pinned the wip half; these are **terminal** and **review**. Both were uncovered: blinding either back to its legacy ids left all 825 self-healing tests green, because no fixture in that block used a renamed lane. | resolver | what the literal cost | |---|---| | terminal | a finished card is not this sweep's business. Keyed on the ids the skip never fired, so finished cards were reconciled on every pass | | review | the other half of the **FN-5256** liveness guard. Keyed on the id it went silent — and this sweep nulls `worktree`/`branch`/`sessionFile` on a live row | ## Blinded separately, not as a pair Each resolver was blinded on its own and measured on its own. **#3138 is exactly why**: there, one case pinned `leaseWipColumns` and left `leaseHoldColumns` green, because the control flow short-circuited before the second guard was ever reached. A single revert that fails proves *one* resolver, not the conversion. That is the finer-grained version of the lesson from #3078, where a whole suite passing proved nothing at all. | blinded | result | |---|---| | `worktreeReconcileTerminalColumns` | **1 failed** | | `worktreeReconcileReviewColumns` | **1 failed** | ## Map progress Closed: `archiveStaleDoneTasks` ×2, `reconcileOrphanedPendingStepResults`, `recoverDriftedAgentTaskLinks`, `reconcileDependencyBlockingLeases` ×2, `reclaimStaleActiveBranches`, `reconcileTaskWorktreeMetadata` ×3. **20 uncovered resolvers remain** of the original 26. Next: `reconcileInReviewBranchRebind` (rebinds branches of live cards), then `recoverAgentsRunningOnInactiveTasks` ×2. ## CI note This will show red on Lint until **#3145** merges — main carries FNXC stamps dated 2026-08-01 through 08-06 while UTC now is 07-31, so every open PR inherits it. #3145 fixes it; this PR touches none of those files. ## Verification `self-healing.test.ts` **416 passed** · `pnpm test:gate` 161 + 487 + 13 + 71 · lint — green locally. |
||
|
|
1136474a63 |
test(engine): pin the archive skip in branch reclaim — the first uncovered resolver whose failure deletes a branch (#3144)
Next entry from the verified coverage map on #3115 — and the first one whose failure mode is **irreversible**. ## The gap `reclaimArchivedColumns` was uncovered: blinding it back to the id `archived` leaves all 825 self-healing tests green, because no fixture in this suite puts a card in a renamed archive lane. ## Why it matters more than the other 22 That guard **skips** archived cards — their branches belong to archive cleanup, not to branch reclaim. Keyed on the id, a card filed in a renamed archive lane fails the skip, and this sweep reaches: ``` git branch -D "fusion/<id>" ``` Every other uncovered resolver I have pinned so far causes a wrong lifecycle decision — a card not requeued, a lease not released, a diagnostic not surfaced. All of those are recoverable from the task row. **A deleted branch is not.** ## Measured 415 pass. Blinding `reclaimArchivedColumns` fails exactly this case, and the assertion that fails is the one checking `git branch -D` was never called — so the failure *is* the branch being deleted, not a proxy for it. ## Progress on the map Closed so far: `archiveStaleDoneTasks` ×2 (#3115), `reconcileOrphanedPendingStepResults` (#3090), `recoverDriftedAgentTaskLinks` (#3102), `reconcileTaskWorktreeMetadata` wip (#3132), `reconcileDependencyBlockingLeases` ×2 (#3138), and this one. **22 uncovered resolvers remain.** Every sweep probed so far has been uncovered, and one (`reconcileDependencyBlockingLeases`) was only half-covered by its own first test — the branch short-circuited before the second resolver was ever consulted. That is why I now blind each resolver separately rather than trusting a single revert. Next: `reconcileInReviewBranchRebind` (rebinds branches of live cards) and the two remaining `reconcileTaskWorktreeMetadata` resolvers (terminal, review). ## Verification `self-healing.test.ts` **415 passed** · `pnpm test:gate` 13 + 161 + 487 + 71 · lint — green. |
||
|
|
920d68e10f |
fix(dashboard): expose column roles to browser bundle (#3151)
## Summary - export the browser-safe `@fusion/core/column-roles` subpath - keep Vite/Vitest aliases ahead of broad `@fusion/core` aliases - restore production dashboard builds after task undo classification adopted shared column-role helpers ## Test plan - `node scripts/check-no-node-only-core-imports-in-dashboard.mjs` - `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest run app/utils/__tests__/taskRevert.test.ts --pool=threads --maxWorkers=1` - `pnpm --filter @fusion/core typecheck` - `pnpm --filter @fusion/dashboard typecheck` - `CI=true pnpm check:changesets` - `pnpm build` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed dashboard build compatibility for browser-based environments. * Improved reliability when importing column role functionality across supported application components. * **Refactor** * Made column role utilities available through a dedicated browser-safe entry point. * **Chores** * Updated development and test configurations to consistently resolve the new entry point. * Documented the browser-safe module classification and recorded the release patch. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
ada62a7c4a |
census: --claims shows which remaining files an open PR already holds (two duplicate claims today) (#3124)
The census says **where** the work is but not **who has it**, and duplicate claims are now the dominant coordination cost of this phase. This adds an opt-in `--claims` report mapping each remaining file to the open PRs already touching it. ## The problem is measured, not suspected - **`self-healing.ts` took three overlapping conversions** from different lanes while one branch was open (#3049, #3075, #3078). Each forced a full rebuild of #3094, and every conflict was the same shape: *same guard, two spellings, different variable names*. That PR's body asks, in as many words, for one lane to own the file. - **`executor.ts` took two independent conversions today** — #3112 and #3118 — same four literals, same payload-lanes fix, two branches. Two workers each read the census, saw the top cluster, and started. Neither could see the other; I only caught it because both appeared in one `gh pr list`. The census is what sends everyone to the same file, so the claim signal belongs here rather than in a side channel nobody reads. `--triage` (#3097) already measured the underlying fact — 53 of 88 guards sat inside an open PR — one step short of being actionable. ## Measured on current main (29 guards) ``` CLAIMED by an open PR: 6 files holding 15 guards 6 packages/engine/src/self-healing.ts ← #3121 #3116 4 packages/engine/src/executor.ts ← #3118 #3112 2 packages/engine/src/auto-merge-finalization.ts ← #3107 1 packages/core/src/task-store/task-artifacts-ops.ts ← #3120 #3119 #3091 … UNCLAIMED: 12 files holding 14 guards — start here 2 packages/dashboard/app/utils/taskRevert.ts 2 packages/engine/src/scheduler.ts … ``` It independently reproduces **both** collisions I found by hand today, which is the strongest evidence I can offer that it works: `executor.ts ← #3118 #3112` and `self-healing.ts ← #3121 #3116`. It also answers the standing fleet instruction empirically. "Claim the largest unclaimed cluster" currently resolves to **12 files holding 14 guards, none larger than 2** — and one of those two (`scheduler.ts`) is in the SYNC-RESOLVED list, where conversion is inert. That is a materially different picture from the headline `29`. ## Design decisions **Report-only and fail-soft**, on the same terms as `--triage`: opt-in, printed beside the totals, changes no count and no exit code. It shells to `gh`, so it is unavailable offline, in CI without a token, and in sandboxes — all of which print a notice and continue. A gate must not depend on network state; this is a work-selection aid, not a gate. **The fail-soft path is loud on purpose**, and it is the case I care most about. A claim report that silently degrades to "nothing is claimed" is *worse than no report*, because it actively sends the reader into work another lane holds — the exact failure the flag exists to prevent. So when `gh` cannot answer it prints `POSSIBLY CLAIMED` and suppresses the start-here list entirely rather than rendering it empty. **Heuristic, and says so.** A PR touching a file is not proof it converts *that file's* guards — it may edit an unrelated function. It over-reports rather than misses, which is the safe direction: a false claim costs one comment asking, a missed one costs a rebuilt branch. **One bulk `gh pr list` call**, not a request per PR — the per-PR shape was too slow to become habitual, and a report nobody runs is not a fix. ## Verification - `lifecycle-column-census.test.ts` — **42 passed** (was 40) - Differential: disabling the flag gives **2 failed | 40 passed**. Both new tests fail on the defect they were written for. - `--strict` and `check-fnxc-future-dates` — exit 0 - Tests stub `gh` on PATH, so no network call and no dependency on the live PR list. The fixture reads the census's **own current top file** rather than a hardcoded path, so it cannot rot as the backlog shrinks (same self-maintaining discipline as #3106). ## What this does not do It does not reserve anything — there is no lock, and two workers who both run it can still collide if they start simultaneously. It reports what is already visible in the PR list, which is enough to catch the every-case-so-far pattern of *starting work on a file someone has held for hours*. A real reservation would need shared mutable state, and I would not add that without an owner asking for it. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d09a856941 |
test(engine): pin the FN-5256 liveness guard on a renamed board (the sweep that clears a live task's worktree) (#3132)
Top item from the verified coverage map on #3115. `reconcileTaskWorktreeMetadata` had **three** uncovered resolvers — the most of any sweep in the file — and it is the one that nulls `worktree`/`branch`/`sessionFile` on a live row. ## Why this sweep first Its own header names FN-5256: the incident where clearing worktree metadata yanked a checkout out from under a running shell. The guard that prevents it is `scopeOverrideMergeActiveSafe`, and that guard is exactly what the wip/review resolvers feed. The existing guard test uses `column: "in-progress"` — **the literal**. So blinding `worktreeReconcileWipColumns` back to `["in-progress"]` leaves all 825 self-healing tests green. The guard is converted; nothing in the suite could tell. On a renamed board the pre-conversion form matched nothing, `scopeOverrideMergeActiveSafe` became true for a card an executor was actively running, and the sweep cleared its metadata. ## The case The renamed twin of the existing FN-5256 test: a `scopeOverride` task live in a **renamed wip lane** keeps its metadata. Same shape, same assertions, different vocabulary — which is the whole point, since the original passes either way. **Measured:** 414 pass; blinding `worktreeReconcileWipColumns` to the legacy id fails **exactly this test**. ## Remaining from the map 25 uncovered resolvers left. Next by risk: `reclaimStaleActiveBranches` (deletes branches) and `reconcileInReviewBranchRebind` (rebinds branches of live cards) — both need a git-shelling harness, so they are slower to pin than this one was. Then the two `reconcileDependencyBlockingLeases` resolvers. I will keep working down that list. The map is on #3115 with verified names; anyone can pick an entry and check it the same way — blind one resolver, run `vitest run src/__tests__/self-healing`, and if it stays green that conversion has nothing behind it. ## Verification `self-healing.test.ts` **414 passed** · `pnpm test:gate` 13 + 161 + 487 + 71 · lint — green. |
||
|
|
ce84aa48d0 |
test(self-healing): cover the renamed-board starved-refinement wake that main's conversion lacked (#3116)
**Rebased onto current `main`, and it shrank to one test.** Was "self-healing consolidated (45 → 39)". ## What happened **Every code change in this PR landed independently from other workers** while it was open, and in each case theirs is equal or better. I took theirs and dropped mine: | My change | Landed on `main` as | |---|---| | pre-execution worktree seizure | `preExecLiveColumns` — same "dangerous direction" reasoning | | FN-5256 liveness cluster | `worktreeReconcileWipColumns` / `worktreeReconcileReviewColumns` | | agent-link membership | `agentLinkLiveColumns` / `agentLinkTerminalColumns` | | starved-refinement peer progress | `starvedWaitingColumns` — a project union covering both duplicated sites | Resolving the rebase by taking `main` left two orphaned declarations (`activeOrQueuedColumns`, `holdPeerIds`) that nothing referenced. `tsc` doesn't flag unused locals here, so I checked references by hand and removed them rather than ship dead code that reads as converted. ## What's worth landing **Their starved-refinement conversion has no renamed-board test — the suite had zero.** This adds one. A candidate resting in a renamed **intake** lane, with its peers in a renamed **hold** lane, must still escalate. The two are deliberately distinct columns so a wrong role set resolves no peers and escalates nothing; a fixture where they coincide would pass either way. The fake needed `listWorkflowDefinitions` — `starvedWaitingColumns` is a **project union**, so per-task selection readers alone leave it resolving nothing and the test would pass for the wrong reason. That mismatch is how I found the gap: my original test failed against their implementation. ## Verification - Green against **their** code - **Revert-proof against theirs:** restoring the literal fails it — 0 escalations against 1 expected - 8 tests in the suite green ## Note for the fleet This is the second PR of mine to shrink to a test on rebase (#3096 was the first). Both times the duplicated work was real and mine was the later arrival. The pattern is worth acting on at the coordination level, not by me working faster. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6483f9ce2b |
fix(scheduler): resolve task:updated / task:deleted lanes asynchronously (scheduler inert 5 → 0) (#3128)
The last inert guards in `scheduler.ts`. Independent of my other branches. ## Inert-guard ratchet | Scope | Before | After | |---|---:|---:| | `scheduler.ts` | 5 | **0** | | total | 12 | **7** (triage.ts 8 → other worker; executor.ts 4 → #3112) | ## The live bug These read `resolveTaskParkedColumnsSync`, which answers with the **default** workflow in production. On a renamed board the scheduler **never woke** on unpause or planning-finish, and a **deleted blocker never unblocked its dependents** — the card sat behind a task that no longer existed. ## The criterion, restated because I got it wrong before **What blocks a guard is whether its answer is consumed synchronously — not whether the enclosing listener is declared sync.** I assumed the latter earlier in this program and reverted for it. All three fail that test: two only gate `schedule()`, which is itself `async`, fire-and-forget and re-entrance-guarded; the third already sits below an `await getSettings()`. The edge-trigger bookkeeping (`planningTaskIds.delete`) **stays synchronous** on purpose — deferring *that* would let a second update re-enter the branch. ## The union is load-bearing, not defensive Post-U11 the default lineage has no `triage` column, so a **resolved** answer returns `intake: "todo"` where the inert path fell back to `"triage"`. Converting without unioning the legacy ids silently **narrowed** the wake set and stopped waking cards in a legacy-named lane — caught by *"schedules when planning clears in triage"*. **A resolved conversion must be a superset of what it replaces, or it is a behaviour change wearing a vocabulary change's clothes.** That's the reusable lesson here. ## Tests - Drained with the repo's existing **`flushAsyncHandlers`** helper — written for exactly this fire-and-forget shape — rather than loosening any assertion. - **The characterization test flipped, as designed.** `workflow-scheduler-parked-columns-live-e2e.pg.test.ts` asserted *"a dependent in a RENAMED hold column is NEVER unblocked"*, with its author noting: *"expected to flip to null the moment the resolver is fixed — and that flip is the whole point of writing it down."* It flipped. Inverted to a REGRESSION case so the assertion holds the fix rather than the defect; it now matches its own CONTROL arm, which still guards against a vacuous pass. ## Verification - 21 scheduler suites — **361 green**, including the live PostgreSQL e2e - **`pnpm test:gate` green**; eslint and `tsc` clean - Changeset added; `check:changesets` passes 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ad5172afd5 |
fix(engine): main is red on check:inert-sync-lanes — #3114's triage conversion is inert, revert the arm (#3126)
## `main` is red on `check:inert-sync-lanes` right now ``` inert-sync-lane: NEW inert conversions — a lane guard now reads a sync resolver that always answers with the DEFAULT board. packages/engine/src/triage.ts: 7 -> 8 ``` Verified on a clean `origin/main` checkout, not on my branch. #3114 converted this guard's third arm to `disposeLanes.wip`; the gate that exists to catch exactly this fired, and the PR landed anyway — presumably because `check:inert-sync-lanes` is not in the blocking merge-gate set. ## The change did not change behaviour `disposeLanes` comes from `resolvePlannerLanes`, which resolves through `resolveTaskWorkflowIrSync` — inert under PostgreSQL for two independent reasons (#3103). So `disposeLanes.wip` evaluates to `in-progress`: **the same value as the literal it replaced.** A card advancing into a renamed execution lane still matches nothing, still reads as an evacuation, and still kills a healthy planning session — the precise bug #3114 set out to fix, unchanged on every board. So the arm goes back to the literal. The gate's own failure text rules out the alternative: > Do NOT re-record the baseline to clear this — that is the same false green one layer up. ## #3114's analysis is kept — only the code reverts Its behavioural description is **correct** and is the clearest statement of this bug anywhere in the file. I have kept those paragraphs and added what is missing: that the fix does not reach under PG, and what would. Whoever supplies a lane answer that is not sync-resolved should make this line read `disposeLanes.wip` and delete the note. The specification is sitting right there for them. ## It also reconciles two contradictory notes, one of them mine My #3108 flag said converting the third arm this way adds an inert comparison and removes a census entry that is telling the truth. #3114 then converted it and added a note saying it fixes the bug. **Both notes sat in the file**, giving any reader two confident, opposite accounts. They are now one account with the evidence attached. ## Read this file's census count carefully #3114 took it to **0** while the inert count went to **8**. The census's own `--triage` output warns about exactly this shape: > for a sync-resolved file, a count of 0 is the WORST case, not the best — the file reads as fully converted Reverting restores it to 1, which is the honest signal. ## Census | | before | after | |---|---|---| | `triage.ts` | 0 | **1** | | repo backlog | 26 | **27** | **The number going up is the point.** A census that reports 0 for a file whose guards are all inert is worse than one that reports the truth — it retires the entry and nobody looks again. ## Measured - `check-inert-sync-lane-conversions`: **exits 1 on `main`, 0 here** (8 → 7). - `src/__tests__/triage*` — **25 files / 374 tests pass**. - `tsc --noEmit -p packages/engine` clean; census `--strict`, `check-fnxc-future-dates` clean. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4b61170a51 |
fix(executor): read task:moved lanes from the payload (executor.ts 4 → 0) (#3112)
**Stacked on #3109** — merge that first; this is its first consumer. ## Census | Metric | Before | After | |---|---:|---:| | COLUMN guards (backlog) | 47 | **43** | | `executor.ts` | 4 | **0** | `executor.ts` is off the census top-files list. ## Why these four could not be converted in place This listener is synchronous and its branches **start execution**, dispose worktrees and release sessions. An await ahead of them defers the `execute()` dispatch itself. The sync IR resolver isn't an option either — it answers with the default workflow under PostgreSQL, so a guard written through it is inert. Reading the lanes the emitter already resolved costs nothing and leaves the prologue synchronous. This listener is the reason #3109 has the shape it does. ## The archive branch is the one with teeth `to === "archived"` matched nothing on a board with a renamed terminal lane, so **archiving never released the task's active-session registry entry** — and that entry is what blocks a **successor** task from acquiring the same path. Not cosmetic: the next task wanting that path fails to register. ## Verification - **Revert-proof:** the new case drives a `shipped` terminal lane (matching no legacy id) and asserts the release. Reverting the branch to the literal leaves the entry held — `expected [Array(1)] to have a length of 0`. - 43 executor suites — **483 green** - **`pnpm test:gate` green**; eslint clean ## Note on shape Lanes are read as **single ids, not sets**, because each branch here is a lane-identity test on one column — exactly what the literals were. Widening to membership would change behaviour, not just vocabulary. Fail-soft to the legacy ids when the emit path could not resolve, matching every other consumer of this payload. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
218086bea2 |
fleet(engine): self-healing 6 → 1 — the board-stall counter, the last guard that needed a sync answer (#3121)
The last fan-out guard, and the one I explicitly said needed a synchronous answer. #3109 made that answer available without an await, so the flag comes off. ## Why this one was last The other two guards in this listener gated work the listener **already `void`s**, so they moved onto the async resolver in #3094. This one increments in-memory state **in the handler's own tick**, so it genuinely needed a synchronous answer. The sync IR path was never that answer: `resolveTaskWorkflowIrSync` cannot resolve a **custom** workflow at all — two independent blockers, #3103 — which is why I wrote that conversion, measured it, and withdrew it. #3109's emitter-carried `lanes` removes the dilemma rather than trading one horn for the other: reading them needs **no await**, so the increment stays in the same tick *and* the guard becomes correct. ## What it fixes On a renamed board this counter read **zero**. The board-stall watchdog was blind to a board whose cards were moving out of implementation the whole time — the signal it exists to raise was never raised. ## Census | | before | after | |---|---|---| | `self-healing.ts` | 6 | **1** | | repo backlog | 29 | **24** | The remaining 1 is the log-dedup closure — a pre-existing flag whose degraded answer costs a duplicate log line, not a lifecycle decision. ## Measured - 3 new cases; `self-healing-completion-fanout.test.ts` **13/13 pass**. - **MUTATION**: restoring the literal pair fails the renamed case. - **The paired negative is the load-bearing one.** The guard means *"left implementation for somewhere that is not implementation"*, so a move **between two non-wip lanes** must not count. Without that case, a conversion that counted every move would pass the positive and inflate the watchdog's denominator — breaking it in the opposite direction, which is harder to notice than a zero. - A **fail-soft** case pins that an emit carrying no `lanes` still counts on the legacy ids. - **Asserted through the counter itself**, not a downstream alert. The increment *is* what this guard decides; routing the assertion through the watchdog would let an unrelated threshold change mask a regression here. - `src/__tests__/self-healing*` + `task-agent*` — **42 files / 848 tests pass**. - `tsc --noEmit -p packages/engine` clean; census `--strict`, `check-lane-wiring`, `check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean. ## On the withdrawal this reverses #3094 withdrew a sync-IR conversion of this listener and recorded why, precisely. That record is what made this cheap: I could tell in one read that #3109 addressed the *specific* obstacle rather than a general "async is hard". A flag that names its blocker exactly is a flag that can be retired the day the blocker goes. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
56b5cdfeed |
test(notifications): cover the wedge-episode renamed-lane clear that main's conversion lacked (#3096)
**Rebased onto current `main`, and it shrank to a test.** Was "close the wedge-episode race, then resolve its lanes." ## What happened Another worker landed **both halves of this PR independently** while it was open. Rebasing showed their versions are better, so I took theirs and dropped mine: - **The serialisation** — theirs is `enqueueWedgeHandling(taskId, run)`, a general callback; mine was wedge-specific. - **The conversion** — theirs is **project-union membership** over the four roles; mine was first-match-per-role via `resolveLifecycleColumns`. Membership is correct: more than one lane can fill a role on a renamed board, and first-match silently ignores the rest. My rebased branch initially compiled to a **duplicate `wedgeHandlingChains` field and duplicate method** — caught by `tsc`, removed. Nothing of my implementation survives, and it shouldn't. ## What's left is worth landing Their conversion has **no renamed-board test**. This adds one. A card recovering into a renamed hold lane must **clear** its episode. Asserted through the *second* notification, because a stale active episode also **refuses the next genuine wedge its claim** — so the visible symptom is a real wedge going unannounced, not merely a stale alert. The fixture needed `listWorkflowDefinitions`: `resolveProjectColumnsForRoles` unions across the project's workflows, so the per-task selection readers alone leave it resolving nothing and the test would pass for the wrong reason. That's how I found the mismatch — my original test failed against their implementation. ## Verification - Green as written against **their** implementation - **Revert-proof against theirs:** restoring the four literals fails it — 1 delivered, 2 expected - 7 notification suites — **80 green**; `tsc` clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed wedge notifications so they can trigger again after a task recovers into a renamed workflow’s hold lane. * **Tests** * Added regression coverage confirming that recovered tasks correctly clear their wedge state and support subsequent notifications. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6050d6eb83 |
chore(engine): mark the auto-merge-finalization reviewed literals DELIBERATE (census 47→45) (#3107)
Fleet phase. `packages/engine/src/auto-merge-finalization.ts` was the last census file with no branch, worktree, or open PR against it. Claim published by pushing the branch before starting. ## Census before / after | | total | this file | |---|---|---| | before | **47** | 2 | | after | **45** | 0 | `--strict` exits 0, baseline re-recorded. **Reclassification, not conversion** — both lines are unchanged. ## Both sites were already reasoned, in a note that calls them non-defects - **Line 30** is the resolver's **degraded fallback arm**, inside `catch`. The live arm two lines up calls `columnHasFlag(ir, columnId, "complete")`. The literal is reached only when IR resolution throws, where the legacy id is the only answer left — removing it would make a failed resolve return nothing. - **Line 99** picks an **error string**. The note above it works through threading `isCompleteColumn` in and concludes the signature widening costs more than the sharper diagnostic buys. I did not revisit either judgement. The gap was mechanical: prose the census cannot read, so both stayed in `byFile` as apparent debt for the next pass to re-derive. ## This is the fourth, and it closes the set With #3056, #3060, and #3063, **every census file that was unclaimed during this phase has now been examined, and not one needed a conversion.** Each site was a three-state fallback arm, or a site a prior pass had already reviewed and kept. The corollary is the finding I would most want carried forward: the remaining count is not a work queue. A worker told to "claim the largest cluster" reads the number, finds most of it already reasoned, and reaches for whatever moves it — which is how three PRs converted guards to a synchronous resolver that is inert under PostgreSQL. One exception worth preserving: **`taskRevert.ts` should stay counted.** I claimed, inspected, and released it without marking. Converting it would classify a *neighbour* row using the modal task's flags — wrong on data, not merely stale on vocabulary — and its note correctly calls the entry **accurate debt** blocked on a per-neighbour flag map. Fallback arms and dead paths → mark. Placeholders awaiting a capability → leave counted. ## Verification - `census --strict` exit 0; `tsc --noEmit` (engine) **0 errors** - No dedicated test file for this module (`vitest` reports none), so no suite to run — comment-only diff, no behaviour change Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
89b21e2906 |
fleet: triage's planning-evacuation check uses the resolved wip lane (census 45 → 44) (#3114)
## Census
| | column guards |
|---|---|
| before | **45** |
| after | **44** |
## What changed
```ts
if (task.column === disposeLanes.hold || task.column === disposeLanes.intake
|| task.column === "in-progress") return;
```
Two role questions and one id question on the same line.
`resolvePlannerLanes` is **already called immediately above**, and its
result carries `wip` — so this needs no new resolution and no new await.
The literal just stops being the odd one out among its neighbours.
## What it cost on a renamed board
This handler aborts a planning session when a card leaves the planner
lanes. `in-progress` is excluded because *a card advancing into
execution is not an evacuation* — that's stated in the note directly
above it.
Against the literal, that exclusion **never matched** on a board whose
execution lane is renamed. So a legitimate advance into execution read
as an evacuation and **killed a healthy planning session** — precisely
the case the comment says must not abort.
`wip` is optional by design (PR #2628: a missing role stays `undefined`
so callers refuse rather than invent a column). Undefined here means the
board declares no execution lane, so there's no advance-into-execution
to exclude and the comparison is correctly false.
## Not addressed, and pre-existing
This line resolves through `resolvePlannerLanes` — the **sync** twin,
which returns the default workflow's lanes under PostgreSQL. That
affects all three lanes on the line equally and predates this change:
the handler is `(task: Task) => {}` with no await available, so fixing
it needs the same emitter-side change as #3082.
Making the third lane consistent with the other two doesn't deepen that,
and it leaves **one** shape to fix there rather than two.
## Measured
| check | result |
|---|---|
| triage / evacuation / planner-lane suites | **453 tests green** |
| four gates + strict census | green |
| engine `tsc` | clean |
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6eeeb43d4b |
test(engine): pin #3047's archive-sweep conversion — measured uncovered (825 tests passed against the reverted fix) (#3115)
Not a conversion — the fleet's conversions are landing faster than their coverage, and this is the audit that shows which ones actually have any. ## Method For each of today's fleet commits to `self-healing.ts`: revert that single commit, re-run the file's suites, see whether anything fails. If nothing fails, the conversion has no regression protection and the "N tests passed" cited on its PR was measuring something else. | commit | reverted | verdict | |---|---|---| | #3075 pause-abort recovery | suites **fail** | covered | | **#3047 archiveStaleDoneTasks** | **825 tests all pass** | **uncovered** | | #3078 (mine) | 204 tests all passed | was uncovered — closed by #3090, #3102 | Every fixture in the `archiveStaleDoneTasks` describe block uses the id `done`, where the literal is correct, so none of them could see the conversion at all. ## What the literal cost The sweep's dependent scan skips tasks in terminal lanes. On a renamed board **nothing matched `done`/`archived`, so every task read as active** — which means every archive candidate looked like it had active dependents, and the sweep archived **nothing**. The board quietly stops auto-archiving: no error, no log line, no failing test. ## The case covers both halves of #3047 - a stale card in a **renamed complete lane** is archived — the `complete` role - a card whose dependent is still live in the **renamed wip lane** is **not** archived, and a dependent already in the **renamed archive lane** does not count as live — the `terminal` role That second assertion is the one that matters: it stops the fix from degenerating into "archive everything", which is the failure mode a one-sided test would miss. ## Measured **413 pass** on current main; reverting #3047 fails **exactly this test**. ## Remaining audit I have now audited 3 of ~13 fleet conversions to this file this way. The method is cheap (one revert, one 17s suite run) and I will keep working through the rest unless someone else picks it up. #3049 could not be auto-reverted — later commits overlap its hunks — so it needs a manual read rather than a mechanical revert. ## Verification `self-healing.test.ts` **413 passed** · `pnpm test:gate` 13 + 161 + 487 + 71 · lint — green. |