60bfebdc98b1d88db02a7df307a75cbec87ea794
11201 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
60bfebdc98 |
fix(reliability): the duration query hid its lane ids inside a SQL template (#2875)
The Reliability panel's **third and last** blind input — and my own loose end. #2861 fixed the two counts beside it, so the panel went from uniformly wrong to **partially** wrong: entries and bounces populated, duration reporting `no-in-review-entries` forever. Partial blindness is harder to notice than total, which is why finishing it matters more than one site suggests. ```sql metadata->>'to' = 'in-review' OR (metadata->>'from' = 'in-review' AND metadata->>'to' = 'done') ``` ## The class, not just the site **This shape is invisible to every check we have.** The lifecycle census scans `===`/`!==` comparisons; the unwired-lane-parameter guard scans declarations. Neither sees a lane id inside a `sql` template, so this class is **not in the backlog total at all** — the number is a floor for this reason as well as the usual one. `scripts/check-sql-column-literals.mjs` (#2841, in flight) is the detector for exactly this: it freezes the surface at 30 sites rather than converting any, so this one was unowned. That PR and this one are complementary — it stops the surface growing, this shrinks it by one. ## The fix Lanes resolve **once per call** via `resolveProjectColumnsForRoles` and arrive as parameterised equality fragments, one branch per id — no interpolated list, no string building. Resolution lives in `getInReviewDurationEventsImpl` because that is where the store is; `async-audit.ts` takes a bare `db` handle and cannot resolve anything. Best-effort, defaulting to the legacy pair, so a caller that cannot resolve keeps exactly today's query. **The union is correct rather than a widening hack**, for the same reason as #2861: these are *move records*, and a past move recorded the column name as it was at the time. A board renamed last month has rows under both ids, so the honest query covers both — which is precisely what `resolveProjectColumnsForRoles` returns. ## Tested against real PostgreSQL, deliberately This is a **SQL predicate** change. A mocked store would assert the arguments and prove nothing about the query that actually runs — which is the entire risk when the literal lives inside `sql`. The new case inserts real `activity_log` rows on a renamed board and reads them back through the real store method. The legacy-lane case in the same file stays green, which is the compatibility half. **Revert proof, measured:** restore the hardcoded fragments and the new case fails with ``` expected [] to deeply equal [ 'renamed-entered', 'renamed-done' ] ``` ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit` (`@fusion/core`) — clean - `activity-log-parity.pg.test.ts` — 5 passed against real PostgreSQL With this, all three Reliability inputs read the board's own lanes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Reliability duration metrics now work correctly with renamed workflow lanes. * Completion tracking recognizes configured completion lanes instead of relying on fixed defaults. * Improved handling of transitions between multiple review lanes and review-to-work-in-progress movements. * Legacy lane behavior remains supported when configured lane information is unavailable. * **Tests** * Added coverage for renamed lanes, historical lane IDs, and transition edge cases. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
de677b231a |
fix(tests): settings-mobile queries container; SettingsModal renders through a portal (#2890)
## What this clears All **17 failures** in `settings-mobile.test.tsx` — the second-largest block in the dashboard `app:backfill 3/4` shard after the TaskDetailModal file (#2885), and the **same root cause**. ## Probed, not assumed Every failure read `expected null to be truthy`, which looks like the modal never rendered: ``` PROBE container.settings-layout=false document.settings-layout=true document.modal=true bodyLen=36637 ``` `SettingsModal` mounts through `createPortal`, so its subtree hangs off `document.body`, not the container `render()` returns. The markup is there; the container-rooted lookup cannot see it. ## Nine of these assertions could never have failed They are **absence** checks: ```ts expect(container.querySelector(".settings-scope-banner")).toBeNull(); expect(container.querySelector("#settings-mobile-section")).toBeNull(); ``` `container` is empty for this component no matter what, so these passed on an **empty root** rather than on absence — they would have kept passing if the element appeared. Converting them makes them mean what they say. All nine still pass, so they were correct, just unproven. ## Two rounds of my own errors, both caught by measuring 1. A blanket `container` → `document` replace also rewrote `renderResult.container.querySelector` into `renderResult.document...` — **not a thing**. That broke the two *"embedded Settings surface"* star tests. Because the embedded surface is genuinely not portalled, this first read as *"embedded needs container"*. It doesn't; the JS was simply invalid. 2. Fixed by restoring those seven, then converting them to **bare `document`** once I confirmed each test unmounts its surface before rendering the next (`modalRender.unmount()` precedes the embedded render), so a document-rooted query cannot match a stale instance. **Verified no regressions rather than assuming** — diffed the failing-test list before and after: 13 fixed / 0 new, then the remaining 4 fixed. Every failure in the final state was already failing at the start. ## Evidence | | result | |---|---| | the file | **41/41** (was 17 failed) | | shard `3/4` | **55 → 38** with this change alone | | mutation: rename `.settings-layout` in `SettingsModal` | **1 failed** | `pnpm lint` clean. Test-only; `SettingsModal.tsx` restored clean. ## Together with #2885 #2885 clears the TaskDetailModal file (30). Both are the same portal/query-root defect in different files, and both were sitting behind `app:app` in the runner's fail-fast — which is why they went unnoticed. Shard `3/4` should land at **~8** with both applied. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d1ea33ee79 |
docs(lanes): three measured claims of mine had gone stale — date them or delete them (#2904)
Comment-only. No source change, no behaviour change. #2903 corrected a note that named a caller which had since been converted. This applies the same check to my own notes, and all three measured claims I wrote are now wrong: | claim | where | actual | |---|---|---| | "`self-healing.ts` alone issues **49** such reads" | `project-lane-vocabulary.ts` + its test | **37** | | "**51** such destinations exist in production" | `workflow-lifecycle-traits.ts` | ~34 | | "**22** deliberately pass `recoveryRehome: true`" | same | ~18 | All were accurate when measured. The shq fleet has been converting `self-healing.ts` since, and this program has been converting `moveTask` destinations all day. The repo-wide read-shaped total is now 37 *in total*, so "49 in one file" could not have remained true regardless. ## Two different repairs, because the claims differ in one way that matters **The self-healing figure has a reproduction.** `node scripts/lifecycle-column-census.mjs --json` reports `queryByFile` and `queryRoles`. So the number is kept, marked explicitly as a dated measurement, and the reader is pointed at the command rather than asked to trust the figure. **The `moveTask` counts have none.** Nothing regenerates them — the census cannot see call arguments, which is the very point the note is making. So they are **deleted** rather than refreshed, with the grep that approximates them inlined and labelled approximate. Refreshing an un-reproducible number just resets the clock on the same failure. The shape of the finding is what the note is for; the count was decoration that decays. ## Two process notes worth recording **Notes that assert facts about other files are a decay class with no detector.** The census counts literals; the unwired-lane guard counts declarations; neither reads prose. Two of these corrections in a row (#2903 and this) came from *reading a note and checking its claim*, which is not something the toolchain will ever do for us. The durable form is: cite a command, or state the shape without the number. **I nearly published a wrong replacement figure.** My first probe against the census AST returned `0` because I wired `summarize()` incorrectly — the third time this session a probe has been wrong before the product was. That is why the corrected note cites `--json` output rather than another hand count. ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit` (`@fusion/core`) — clean - `project-lane-vocabulary.test.ts` — 9 passed - census `--strict` — exit 0, counts unchanged 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8e505dc4e1 |
docs(recovery): the reason this parameter is optional stopped being true (#2903)
Comment-only. No source change, no behaviour change.
The note on `isInReviewMissingWorktreeSessionStartFailure` said:
> Optional rather than required because the other caller
(`extension.ts`) still asks BOTH questions with the literal.
**It doesn't.** All three production callers pass the resolved answer:
```
packages/cli/src/extension.ts:1927 retryReviewColumns.has(task.column)
packages/cli/src/commands/task.ts:1390 retryReviewColumns.has(task.column)
packages/dashboard/src/routes/register-task-workflow-routes.ts:2885 retryReviewColumns.has(task.column)
```
Left standing, that sentence tells the next reader an unconverted caller
exists — and "we keep the fallback because someone still needs it" is
exactly the justification that keeps an inert-conversion shape alive. It
is the specific failure this program has spent the day removing, in the
form of a comment rather than code.
## The parameter stays optional, for a reason that does not rot
I checked whether to make it **required** — the unwired-lane-parameter
guard's own failure message suggests exactly that ("make the parameter
required so the compiler finds the call sites") — and decided against
it, for measured reasons:
- **25 test call sites** use the optional form, several of them
*precisely* to pin the degraded mode (`cli-active-count-lanes.test.ts`
exercises the no-argument path on both a legacy and a renamed lane).
Requiring the parameter deletes that coverage.
- The enforcement it would buy already exists: `isReviewColumn` is in
the guard's vocabulary, so if any of those three callers stops passing
it, the build fails.
So the note now gives the durable reason instead of the expired one.
## How this surfaced
Not from the census — the count here is unchanged, and an *omitted
argument* is invisible to it anyway. It came from reading an audit note
that named a specific caller and checking the claim. Notes that assert
facts about other files decay silently; this one had.
## Verification
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/engine`) — clean
- `restart-recovery-coordinator.test.ts` — 12 passed
- `cli-active-count-lanes.test.ts` — 10 passed
- unwired-lane guard — 9/9, no new entries
- SQL-literal gate — green
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
10f9df1600 |
fix(overseer): the whole oversight loop was inert on a renamed board (#2898)
`resolveWatchedStage` keyed on the literals `in-progress`/`in-review`, so on a board that renames either it returned `null` for **every** card. That is three literals with an outsized blast radius. `observeTask` returns early on a null stage, so: - no `OverseerStageObservation` is recorded, - no `overseer:intervention` entry is emitted, - and `PlannerRecoveryController`, which consumes those observations, has nothing to steer, retry or targeted-fix. **The entire oversight loop was inert and silent about it** — the same shape as the self-healing sweeps whose queries returned empty arrays. ## I deferred this myself, on a cost argument that was wrong The audit note I wrote for this site said resolving inside `observeTask` "buys a workflow read per card per poll". Then I read the caller: the poll **already awaits `resolveEffectiveSettings` per task**. It is a per-task async loop regardless, so with an IR cache keyed by workflow the addition is *(distinct workflows)* resolutions, not *(cards)*. Pricing the fix before checking the caller cost a deferral. Worth recording, because "this needs a cost judgement" is the most comfortable place in this program to leave something. ## The review test is the three-trait union, deliberately `isReviewColumnRole` checks only `mergeBlocker || humanReview`. A board whose review lane carries `merge` (**mergeOrchestration**) — the built-in default's own shape — would classify as *not in review* and be skipped. Reaching for the obvious helper would have reintroduced the bug this change removes, through the helper meant to fix it. There is a case asserting exactly that. ## Wiring Both call sites, because either alone leaves a hole: | site | why it matters | |---|---| | the poll (`project-engine.ts`) | per-poll IR cache — a workflow edit is picked up next tick rather than served stale | | the manual nudge | otherwise a renamed board answers `no-active-stage` to an operator pressing the button | `columnFlags` is in the `unwired-lane-parameter` vocabulary, so the wiring cannot silently rot — the guard reports it if a future change drops the argument. Fail-soft throughout: an unresolvable workflow yields `undefined` and the callee falls back to the legacy ids, which is exactly today's behaviour. A v1 IR declares no columns, so it takes the same path. ## Revert proof (measured) Drop the `columnFlags` branch and **exactly the three renamed-lane cases fail**: ``` expected null to be "executor" expected null to be "merger" (mergeOrchestration lane) expected null to be "merger" (humanReview lane) ``` The legacy-id and neither-role cases stay green — the gate must still gate, and watching every column would be its own defect. ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit` (`@fusion/engine`) — clean - `planner-overseer.test.ts` + `planner-recovery-controller-human-control.test.ts` — 64 passed - unwired-lane guard — 9/9, no new entries Carries the one-line SQL-baseline re-record (`team-analytics.ts: 6 → 3`) that #2864 left behind, same as my other open branches — main is red on it, and identical changes to that line merge without conflict. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
defe48d30f |
fix(core): per-workflow metrics read zero on a renamed board (#2866)
Second of the 14 lane-bound SQL sites from #2839, after #2864. Independent of it — different file, different caller argument. ## The defect `aggregateWorkflowAnalytics` filtered in SQL on `t."column" = 'done'` and `IN ('in-progress','in-review')`. On a renamed board those match nothing, so `tasksCompleted`, `tasksInProgress` and `tasksInReview` come back **zero for every workflow** while the board is busy. Nothing errors. Same shape and same fix as #2864: resolve per **project** via `resolveProjectColumnsForRoles`, bind an `IN` list, and thread the store from the single Command Center caller so the parameter has a supplier immediately rather than becoming an inert seam. ## What the test caught that I had not **The renamed case still failed with the query fixed.** The bucketing at lines 296–297 already uses `isWipColumnRole` / `isReviewColumnRole` — correctly converted — but those read `query.columnFlagsByName`, which production supplies and my fixture did not. So: - the **SQL** decides *which rows come back*; - the **trait map** decides *which bucket each row lands in*. Both halves have to be right. Fixing only the query would have shipped a "conversion" that still reported zero on a renamed board, and the file would have scored as converted twice over. That is exactly the partial-conversion shape this program keeps re-finding — caught here only because the test asserts `tasksInReview` alongside `tasksCompleted`, since those two paths take **different** resolved sets (complete vs wip+human-review). Asserting the completed count alone would have left the second conversion unproven. ## Measured Reverted, only the renamed case flips: ``` ✓ default vocabulary: completed and in-review work are counted × renamed vocabulary: completed and in-review work are counted ✓ renamed vocabulary: a card in the HOLD lane counts as neither ✓ without a lane store, the legacy ids still answer Tests 1 failed | 3 passed (4) ``` The hold-lane negative is there so resolving real lanes cannot degrade into "every column counts" — trading an undercount for an overcount is harder to notice than the original bug. ## Scope The sync SQLite arm in the same file keeps its literals: it throws in backend mode and has no production caller, the same dead-arm conclusion reached for `cleanupStaleMergeQueueRowsImpl` on #2839. ## Verification `pnpm test:gate` green · both Command Center analytics suites 8/8 · `tsc` core 0, dashboard 0 · lint 0 · changeset included. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c143327d4b |
fix(core): the archived-document guards failed in OPPOSITE directions on a renamed lane (#2886)
Two of the four convertible sites my own learnings doc **miscounted as sentinels** — the #2877 review corrected "8 of 9 must not be converted" to "5 of 9", and these are two of the three that correction freed. They read `task.column` straight off a row `select`, so they are board lanes by exactly the test that document gives, and a renamed archived column is simply not seen. What makes the pair worth fixing together is that they fail in **opposite directions**: | guard | on a renamed archived lane | consequence | |---|---|---| | `upsertTaskDocument` | fails to **reject** | an archived card's documents stay **writable** — the read-only contract silently does not hold | | `publishArchivedTaskDocumentAddition` | fails to **accept** | a legitimate archived-document publication is refused as `parent-not-archived` | The second is the sharper one: valid operator work refused, and refused with a message that reads as a data-integrity error rather than a lifecycle mismatch. ## Shape Both take an `AsyncDataLayer` and can resolve nothing themselves; their store-level impls hold the store, so the lane set arrives as a parameter resolved once per call — the shape #2875 used for the SQL predicate. **One shared `resolveArchivedLanes` for both paths**, deliberately: if the write guard and the publication guard could disagree about whether a card is archived, a card ends up both read-only *and* un-publishable. ## The revert proof caught my own fixture first My first version set `deletedAt` alongside the renamed column, and **the revert proof passed with the fix removed**. Both guards are `column-is-archived || deletedAt != null`, so a soft-deleted fixture short-circuits the exact comparison under test — the assertion was holding for an unrelated reason. Dropping `deletedAt` isolates it, and is also the *real* shape: a live row in a workflow-declared archived lane is what a renamed board produces, and what `getLiveTaskColumn` was written to catch. Revert proof, measured honestly the second time: restore `task.column === "archived"` and the renamed-lane case fails — the upsert resolves instead of rejecting. ## Real PostgreSQL, deliberately These are row predicates inside a transaction. A mocked store would assert the arguments and prove nothing about the comparison that runs — the same reasoning as #2875. Three cases: the renamed lane rejects, the **legacy** `archived` id still rejects (most boards never rename anything), and a live card is still allowed through (a guard that rejects everything is its own bug). ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit` (`@fusion/core`) — clean - new `archived-document-lanes.pg.test.ts` + existing `artifacts-documents-evals.pg.test.ts` — 12 passed against real PostgreSQL Note: the SQL-literal baseline is untouched here — #2881 owns re-recording it after #2864's conversion left main's gate red. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9366bc8382 |
fix(workflow): the review handoff killed the walk on a renamed review lane (#2900)
The sharpest lane defect left in the backlog, and the one I have been
deferring since the first sweep.
```ts
if (seam === "review-handoff") {
const result = await primitives.transitionTask(primitiveCtx, context.task, {
column: "in-review", // ← post-U12 this is a rejected destination on a renamed board
```
Post-U12 `moveTask` **rejects** a destination the workflow does not
declare. So on any board with a renamed review lane, the handoff threw
`TransitionRejectionError` and **killed the workflow walk mid-run**. Not
a silent wrong answer for once — a hard failure in the middle of a task,
which is why it outranked everything else once it became reachable.
**Why it was deferred:** every fix threads a resolver out of
`executor.ts`, and #2820 was editing that file. It merged at 22:08, so
this was finally free of the conflict.
## The role travels, not the column
Seam handlers in `workflow-node-handlers.ts` are pure functions over an
IR node and a task — no store, no task id to resolve from — so a handler
can only ever name a literal. The runtime primitive in `executor.ts`
**does** hold the store, so the seam now asks for `columnRole: "review"`
and the primitive resolves it against the task's **own** selection.
One authority, deliberately. Answering one question with two reads is
what took #2843 five review rounds, and I would rather not relearn it
here.
Compatibility is preserved in both directions:
- `column` still wins when both are supplied — an explicit destination
is an explicit destination;
- an unresolvable role falls back to the legacy `in-review` rather than
failing the transition, which is exactly the behaviour every caller had
before.
## The test asserts the literal is *gone*, not merely accompanied
`column` takes precedence over `columnRole` downstream, so a diff that
added the role while leaving the literal would look converted and be
completely inert. That is the exact shape this program keeps finding — a
documented fallback in front of a literal that still decides everything
— so the assertion is:
```ts
expect(input.columnRole).toBe("review");
expect(input.column).toBeUndefined(); // ← the half that matters
```
**Revert proof, measured:** restore `column: "in-review"` in the seam
and it fails with `expected undefined to be 'review'`.
## Verification
- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/engine`) — clean
- new `review-handoff-lane.test.ts` plus the two neighbouring seam
suites — 41 passed
Carries the one-line SQL-baseline re-record (`team-analytics.ts: 6 → 3`)
that #2864 left behind, same as my other open branches — main is red on
it, and identical changes to that line merge without conflict.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a453912ddf |
self-healing: merged-but-unfinished tasks never finalized on a renamed board (fifteenth sweep) (#2897)
`recoverMergedReviewTasks` finalizes a task whose merge is **confirmed** but which never reached the complete lane. Two literal reads meant that on a renamed board it was never found, so a card whose commit is already on the base branch sat in review or hold indefinitely — merged work the board still shows as unfinished. ## The two redundant guards convert, they don't get deleted Both `t.column === …` checks were redundant while the query pinned the column. Under a resolved read they become the per-card verdict. Deleting them would have silently widened the sweep — the same trap called out in #2891. ## Carries the two shapes review established earlier in this series - **Narrow when the card can answer, broad when it cannot** (#2891). `resolveWorkflowIrForTask` *substitutes* the built-in IR rather than failing, so a card with an unreadable selection would otherwise be rejected by the very verdict that the project-scoped query had just admitted it under. It falls back to the project sets instead. - **Deduped across the buckets** (#2879), so a column carrying both a review role and the hold role cannot finalize one card twice. Both were review findings on earlier PRs in this series, applied here up front rather than waiting to be caught again. ## Revert results Each applied alone and the file re-run: | conversion | reverted → | | --- | --- | | the resolved reads | fails — the card is never listed | | the per-card review verdict | fails — the renamed review lane does not match | Observable is `resolveSelfHealingMergeTarget`, a private method called once per candidate, so the assertion sits downstream of both halves without a git fixture. A non-vacuous companion (merge-confirmed card in the wip lane → untouched) rules out a read that returns everything. ## Verification `pnpm test:gate` 161 + 487 + 13 + 71, plus `self-healing.test.ts` 412; `tsc` engine clean; `pnpm lint`, `check:changesets`, census `--strict` clean, each run explicitly. |
||
|
|
cfb713bda1 |
notification: record the measured reason four wedge-progress ids stay literal, and un-red main's gate (#2882)
Two small things, neither of which changes behaviour. ## 1. A conversion I attempted, measured, and reverted `hasProgressed` in the wedge-episode path names four column ids outright. I converted them to a resolved lane set. It **broke an existing gate test** — `task-wedge-notification.test.ts` → *"sends one actionable push and mailbox message per active terminal episode"*: 1 message delivered, 2 expected. The note already in that file was right, and stronger than it read. The hazard is **not** specific to the resolve/claim ordering — it is **any `await` added before the resolve**. Column resolution needs one. `task:updated` listeners fire synchronously, so a re-wedge arriving close behind a recovery reaches `claim` while the first episode is still open, and the operator's second alert is dropped. Product change reverted; only the comment lands, now carrying the measurement and naming the failing test as the acceptance check for whoever owns the wedge-episode contract. **Left counted, not exempted** — the census should keep pointing here. Worth stating: the pre-existing note was a warning written speculatively. Attempting the conversion is what turned it into evidence, and the evidence says the blocker is real but sits somewhere else (per-task serialisation) than the note implied. ## 2. `main`'s gate is red, and not from this branch `pnpm test:gate` fails on a clean `origin/main` tree at `check-sql-column-literals`: ``` packages/core/src/team-analytics.ts: 3 site(s) now, baseline still allows 6 — re-record it ``` A reduction landed without re-recording the baseline in the same commit, which that check explicitly asks for. Reproduced on `origin/main` with my changes stashed, so it is not mine — but it blocks **every** open PR until recorded. Ratchets **31 → 28** sites across 14 files, downward only. ## The vacuous assertion this round (sixth) The first version of the reverted test passed **with the fix reverted**. `hasProgressed` is a three-clause OR, and the middle clause — *status is a string and is not `failed`* — is true for a recovered task on any board, so `status: "in-progress"` in the fixture satisfied it regardless of column. Same shape as the other five: something the code does anyway. Found by running the revert, not by reading it. ## Verification `pnpm test:gate` 161 + 487 + 13 + 71 (green only with the baseline commit); `tsc` engine clean; notification suite 77 passed; `pnpm lint` and census `--strict` clean. |
||
|
|
890e1f87e7 |
fix(core): issue panels reported nothing fixed on a renamed board (#2871)
Fourth and last of the lane-bound analytics sites from #2839, after #2864, #2866 and #2870. ## The defect `aggregateGithubIssueAnalytics` and its GitLab twin filtered their resolved-issue query on `"column" = 'done'`. On a renamed board that matches nothing, so `fixed` is **zero**, the resolved-issue list is empty, and `net` reports every filed issue as still outstanding — while the team closes issues all week. Nothing errors. Same fix as the previous three: resolve per **project** via `resolveProjectColumnsForRoles`, bind an `IN` list, thread the store from each Command Center caller so the parameters have suppliers immediately. ## Both providers in one change, deliberately These two files are **copies** — same query, only the provider literal differs — and a copy is exactly what gets half-fixed. Converting one and not the other type-checks, passes that provider's test, and leaves the second silently broken with no signal anywhere. The suite runs every case against both, so the pair cannot drift. ## Measured Reverted, exactly the two renamed cases fail — **one per provider** — while both default-vocabulary controls, both WIP-lane negatives, and both omitted-store legacy cases stay green: ``` ✓ github: default vocabulary counts a resolved issue × github: renamed vocabulary counts a resolved issue ✓ github: renamed vocabulary does NOT count an issue still in the WIP lane ✓ github: without a lane store, the legacy id still answers ✓ gitlab: default vocabulary counts a resolved issue × gitlab: renamed vocabulary counts a resolved issue ✓ gitlab: renamed vocabulary does NOT count an issue still in the WIP lane ✓ gitlab: without a lane store, the legacy id still answers Tests 2 failed | 6 passed (8) ``` That the failures are symmetric is itself the check on the copy-paste risk. ## Scope The sync SQLite arms keep their literals: they throw in backend mode and have no production caller, the same dead-arm conclusion as `cleanupStaleMergeQueueRowsImpl` on #2839. ## Verification `pnpm test:gate` green · Command Center + GitLab issue analytics suites 10/10 · `tsc` core 0, dashboard 0 · lint 0 · changeset included. --- **This closes the lane-bound half of #2839.** All 14 sites the hand-review identified as genuinely vocabulary-bound are now converted across four PRs. What remains there is the 11 `!= 'archived'` exclusions, which are probably correct as literals — archiving writes `task.column = 'archived'` unconditionally as a state rather than a lane — plus one dead SQLite arm. Those need per-site judgment, not conversion. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
216632bd3a |
fix(core): task-duration stats were computed from an empty set on a renamed board (#2870)
Third of the 14 lane-bound SQL sites from #2839, after #2864 and #2866. Independent of both. ## The defect `aggregateProductivityAnalytics` filtered its duration query on `"column" = 'done'`. On a renamed board that matches nothing, so the entire task-duration distribution — median, p90, average, total — is computed from an **empty row set** and reports zeros while the project ships work. Nothing errors. Same shape and fix as the previous two: resolve per **project** via `resolveProjectColumnsForRoles`, bind an `IN` list, thread the store from the single Command Center caller so the parameter has a supplier immediately rather than becoming an inert seam. ## Measured Reverted, only the renamed case flips: ``` ✓ default vocabulary: a finished task contributes to the duration stats × renamed vocabulary: a task in the RENAMED complete lane contributes ✓ renamed vocabulary: a task still in the WIP lane does NOT contribute ✓ without a lane store, the legacy id still answers Tests 1 failed | 3 passed (4) ``` ## The negative asserts the median, not just the count This fix's failure mode is **worse than the bug it fixes**. Resolving too many lanes would pull unfinished work into the distribution and produce a plausible-but-wrong median — a number nobody questions — where the bug produces an obvious zero. So the WIP-lane case asserts `medianMs` is null as well as `completedTasks` being 0. ## A fixture error worth naming My first version asserted `taskDuration.count`. `TaskDurationSummary` exposes `completedTasks`. Every case failed with `expected undefined to be 1` — **including the controls** — which reads exactly like a broken product until you notice the control is failing too. A control that fails is a fixture bug, not a finding; that asymmetry is the fastest way to tell them apart. ## Scope The sync SQLite arm keeps its literal: it throws in backend mode and has no production caller, the same dead-arm conclusion as `cleanupStaleMergeQueueRowsImpl` on #2839. ## Verification `pnpm test:gate` green · both Command Center analytics suites 8/8 · `tsc` core 0, dashboard 0 · lint 0 · changeset included. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ab15e5f9f7 |
docs(lanes): audit the three files the census points at with no reason attached (#2873)
**No source change.** Every literal stays counted and none gets an exemption marker. What changes is that the census now points at these three with the analysis attached, instead of making each worker who reaches them re-derive it. Peers have already done this well for `notification-service.ts` — converted it, *measured* a real delivery regression, reverted, and left it counted with the reason. These three had nothing at all, and one of them is the highest-impact unowned site I found. ## `planner-overseer.ts` (3 guards) — REAL, and larger than three literals suggest On a renamed board `resolveWatchedStage` returns `null` for every card. `observeTask` returns early on a null stage, so **no observation is recorded**, no `overseer:intervention` entry is emitted, and `PlannerRecoveryController` — which consumes those observations — has nothing to steer, retry, or targeted-fix. **The entire oversight loop is inert and silent about it**, exactly like the self-healing sweeps whose queries returned empty arrays. Not mechanical, which is why it is flagged rather than converted. `resolveWatchedStage` is a pure sync function over a `Partial<OverseerTaskRef>` with no store and no task id, so the lane answer has to arrive as a parameter. Its only production caller, `observeTask`, *is* async and the monitor *does* hold a store — but it runs **once per task per poll**, so resolving inside it buys a workflow read per card on a timer. The shape that works is the one the board-load enrichment landed on in #2845: resolve at the **poll**, once, with an IR cache keyed by workflow, and pass the flags down. That makes it a change to `project-engine.ts`'s poll as much as to this file — a cost judgement about a periodic engine loop, not a rename. `columnFlags` is in the unwired-lane-parameter vocabulary, so whoever adds the parameter cannot leave it unwired. ## `async-mission-store-queries.ts` (1 of 3) — REAL `getTerminalTaskEvidence` tests only `column === "done"` for its `done` verdict, so a completed card on a renamed board falls through every branch to `{ kind: "nonterminal" }`. The caller is mission **terminal evidence repair**, so a finished feature reads as unfinished — a wrong *verdict*, not an error. The `archived` test beside it has the same defect, masked for soft-deleted rows by its `deletedAt` companion, which is why only the `done` half bites in practice. Takes a bare `QueryHandle`: no store, no task object, no workflow. The fix is a resolved terminal-lane set threaded in by the caller — the same shape `getLiveTaskColumn` needs, and it should land *with* it so the two cannot disagree about what "finished" means. ## `audit-ops.ts` (2) — one sentinel, one real, and they look identical ```ts if (state === "archived") // ← getLiveTaskColumn's MANUFACTURED value: do NOT convert if (pgRow.column === "archived") // ← a real board lane: convertible ``` The first compares against a string `getLiveTaskColumn` *fabricates* for an archived-or-soft-deleted parent, so converting it to `isArchivedColumnRole` would keep passing on the built-in board and start **failing** on a renamed one — a soft-deleted task's log would become writable. The second reads the task row, so a renamed archived column keeps accepting log writes; its `deletedAt` companion masks that in practice. Two lines that look the same and need opposite treatment is precisely the reason these notes are worth more than the count. ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit` (`@fusion/core`, `@fusion/engine`) — clean - `planner-overseer.test.ts` — 50 passed - census `--strict` — exit 0, **counts unchanged** (that is the point) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b85e5f90e1 |
fix(create): two task-CREATE destinations named a lane the board does not have (#2843)
Both files sat at **census-zero** and both wrote real cards into columns
no workflow declares. The census scores `===` comparisons, so a lane
literal passed as a **call argument** is invisible to it — one of the
four census-blind classes. These are the only two explicit-`column`
creates in production:
```
packages/dashboard/src/routes/register-gitlab.ts:108 column: "triage"
packages/cli/src/extension.ts:5243 column: "todo"
```
## The two defects
**`register-gitlab.ts` — `column: "triage"`, a column U11 DELETED.**
This one is broken on *every* board, not only renamed ones: the default
lineage is now `todo | in-progress | in-review | done | archived`.
`createTask` already resolves the intake column of the workflow it
selects (`resolvedEntryColumn`), and an explicit `column` **overrides**
that resolution — which is why the stale literal survived U11. Nothing
rejects the write and nothing logs it: the route answers `201` with a
task id and the imported card is simply not on the board. Same shape as
the `task-update.ts` triage defect fixed earlier in this program.
Fix: omit `column` and let `createTask` resolve intake.
**`extension.ts` `fn_delegate_task` — `column: "todo"`.** On a workflow
whose ready lane is named anything else, the delegated card goes to an
undeclared column: written, reported to the caller as delegated, never
visible to the agent it was delegated to.
Fix: resolve the selected workflow's `hold` lane. **Deliberately not**
"omit the column like the GitLab route" — the tool's own contract is
*"the task goes to the ready-to-work lane and the target agent picks it
up on its next heartbeat"*, so inheriting intake resolution would park a
delegated card in a manual-intake lane waiting for a human. That would
be a behaviour change; `hold` is the role that names the lane the
literal meant.
## New helper: `resolveWorkflowColumnForRole(store, role, workflowId?)`
The **write**-shaped counterpart to `resolveProjectColumnsForRoles`. The
read helper unions in the legacy ids because an extra id in a query set
is inert; here the same trick is a silent wrong write (post-U12 an
undeclared column is a `TransitionRejectionError` on move, a phantom
lane on create), so it returns one column from one workflow, or
`undefined`.
**A contract I got wrong twice, now pinned by a test.** `undefined`
means *"this workflow declares no such column"* and nothing else.
`resolveWorkflowIrById` never throws and never returns nothing — an
unregistered builtin id, a missing definition row and a failing read all
resolve to the default coding IR (branded via `markFellBack`). So an
unreadable workflow yields the **built-in** hold lane, not `undefined`,
and both call sites' `?? "todo"` fallbacks are narrower than they look.
Two of my first test cases asserted the opposite and failed; the
behaviour is the resolver's, and the write it produces is identical to
the caller's own legacy fallback either way.
## Revert proofs (measured, not asserted)
| revert | failure |
|---|---|
| `column: holdColumn` -> `column: "todo"` | `extension.test.ts`:
`expected 'todo' to be 'queued'` |
| omitted column -> `column: "triage"` | `routes-gitlab.test.ts`:
`expected 'triage' to be undefined` |
The two neighbouring `fn_delegate_task` cases stay green under the first
revert, because the built-in board and the test's `linearWorkflowIr`
both call the lane `todo` — which is exactly why this literal survived
every previous pass.
The GitLab case asserts **absence** of the key rather than a resolved
id: the store there is a fake whose `createTask` echoes its input, so
asserting a resolved value would be testing the fake. Absence is the
property that hands the decision to the real `createTask`.
## Census
| file | before | after |
|---|---|---|
| `packages/dashboard/src/routes/register-gitlab.ts` | 1 | 0 |
| `packages/cli/src/extension.ts` | 1 | 0 |
Baseline tightened. It also picks up
`packages/core/src/task-store/moves.ts` 2 -> 0, which was **already true
on main** — not from this diff.
## Noted, deliberately not changed
- `validateAssignableAgentId`'s synthetic probe a few lines above still
uses `{ id: "<new>", column: "todo" }`. It feeds `isImplementationTask`,
whose `IMPLEMENTATION_TASK_COLUMNS` set an earlier worker documented as
deliberately-not-converted (converting it makes the routing policy async
— an agent-admission behaviour change). On a renamed board the probe is
now *stricter* than the real destination, which is the safe direction
and matches the pre-existing behaviour.
- The third site from this bucket, `workflow-node-handlers.ts:455`
(`transitionTask({ column: "in-review" })` on the `review-handoff`
seam), is a **hard** failure rather than a silent one — `transitionTask`
routes through `moveTask`, which post-U12 throws
`TransitionRejectionError` for an undeclared destination, so the
workflow walk dies at the handoff on any renamed review lane. It is
engine (`batch-engine`) and fixing it properly touches `executor.ts`,
which #2820 is also editing. Left for that batch rather than opened as a
conflicting edit.
## Verification
- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit -p tsconfig.json` for `@fusion/core`,
`@runfusion/fusion`, `@fusion/dashboard` — clean
- `node scripts/lifecycle-column-census.mjs --strict` — exit 0
- targeted: `project-lane-vocabulary.test.ts` 14/14,
`routes-gitlab.test.ts` 8/8, `extension.test.ts -t fn_delegate_task` 9/9
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* GitLab-imported cards now appear in the workflow’s configured intake
lane.
* Delegated tasks now move to the workflow’s configured hold lane,
including workflows with custom lane names or separate intake and hold
lanes.
* Delegation reports the task’s final lane and provides an error when it
cannot be moved successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7763063cbd |
docs(task-revert): the recorded conversion "blocker" is a correctness boundary, not a cost (#2868)
`findOpenUndoTaskForSource` carries a note saying its conversion is blocked on **hoisting the flags derivation** in `TaskDetailModal` — framed as a hook-ordering cost in a 5000-line component, i.e. something somebody could pay. I went to pay it, and the framing is wrong in a way that would have produced a **worse defect than the literal**. ## Why hoisting would not unblock it This function scans the `tasks` list for **other** tasks pointing back at the source, so the column it classifies belongs to a **neighbour**. `detailColumnFlags` describes the **modal's own** task — its own FNXC note says exactly that, and it is guarded by `detailFlagsAreForThisTask` *precisely because* using it for anything else is wrong. So supplying it here would answer *"is this neighbour finished?"* with the modal task's traits. On a project where two workflows reuse a column id, an open undo task gets classified by a workflow it does not belong to, and the "Undo task" affordance vanishes or persists wrongly. **Wrong flags are worse than a stale vocabulary.** The literal is at least uniformly legacy; this would be wrong *on data*. It is the flags-for-the-wrong-row shape — the same question I have been applying all program: *do these flags describe the column of the row this guard is about?* ## What a correct conversion needs Per-**neighbour** flags: the caller would have to resolve each candidate's own workflow, which the modal does not have and should not fetch mid-render. Until a per-task lane map exists at that call site, **the literal is the right answer**. ## Why this is worth a PR rather than silence The previous note is an invitation. Somebody following it would land a plausible-looking conversion, drop the census count by one, and introduce a data-dependent bug that only appears on multi-workflow projects — the exact "conversion that scores as a win" pattern documented in `docs/solutions/workflow-learnings/`. The census entry stays and is now labelled as **accurate debt** rather than a missed conversion. No behaviour change; comment only. Dashboard app `tsc` clean, `pnpm lint` clean, `app/utils` suites 59 files / 689 tests green. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ea477f3ada |
fix(core): team analytics reported an idle project on a renamed board (#2864)
First of the 14 genuinely lane-bound SQL sites from #2839. I had been deferring these as "the owner is active in those files" — then checked, and **no open PR touches them**. The batch-core commit had already landed, so there was nothing in flight to collide with. The deferral was an assumption I could have tested two turns earlier. ## The defect `aggregateTeamAnalytics` filtered in SQL on `"column" = 'done'` and `IN ('in-progress','in-review')`. On a board whose lanes are renamed those match nothing, so per-agent completed counts, the project total, and the in-flight breakdown all came back **zero**. Nothing errors. A dashboard reading *"0 tasks completed"* for a team that shipped all week looks like an idle project, not a bug — which is why this survives review and never gets filed. **Why the sweep missed it:** the lifecycle census parses TypeScript comparisons; these ids live inside SQL strings, which are string data. The batch-core conversion fixed this file's TS guards **today** and left the queries untouched. The file scored as converted. ## Resolved per project, not per task `resolveProjectColumnsForRoles` gives the union of a role's columns across the project's workflows, which is the right set here because analytics aggregates a whole project — so a bound `IN` list is sufficient. The merge-queue cleanup needed the *superset-then-decide-in-JS* shape instead, because its lanes are genuinely per task and SQL cannot know a task's workflow. Same program, two correct answers; worth not copying the wrong one. ## Measured Reverted, exactly one case flips: ``` ✓ default vocabulary: a completed task is counted × renamed vocabulary: a task in the RENAMED complete lane is counted ✓ renamed vocabulary: a task in the WIP lane is NOT counted as completed ✓ without a lane store, the legacy ids still answer Tests 1 failed | 3 passed (4) ``` The three controls are deliberate: the default vocabulary (a generally broken aggregator cannot hide behind the renamed case), a WIP task that must **not** count as completed on the renamed board (resolving real lanes must not degrade into "every column counts" — an undercount turned overcount is harder to notice), and an omitted lane store that must keep the legacy answer byte-identical. ## Two mistakes the first attempt made, both caught by running it - **`= ANY(${array})` does not work.** Drizzle expands a JS array in a template into a comma tuple, so PostgreSQL rejected `(($1,$2,$3))` with *"op ANY/ALL (array) requires array on right side"*. An `IN` list of individual bound parameters is the working shape. Each id stays a parameter — these come from operator-authored workflow definitions and are never interpolated as SQL text. - The fixture's `ON CONFLICT (id)` had no matching constraint on `project.agents`; the sibling suite seeds with explicit `created_at`/`updated_at`. ## Scope One production caller (`register-command-center-routes.ts`), threaded in this same change so the parameter has a supplier from the start rather than becoming another inert seam of exactly the class this program keeps finding. The sync SQLite arm in the same file keeps its literals: that path throws in backend mode and has no production caller, the same dead-arm conclusion reached for `cleanupStaleMergeQueueRowsImpl` on #2839. ## Verification `pnpm test:gate` green · both Command Center analytics suites 8/8 · `tsc` core 0, dashboard 0 · lint 0 · changeset included. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cb731420c |
test(engine): re-point the stale-spec ratchet at the improved guard (#2863)
## Main is red; this fixes it
`executor-stale-spec-active-lanes.test.ts` — 2 failures on `main`:
```
× resolves the task's lifecycle columns before deciding the skip
→ expected source to contain 'const activeLifecycle = resolveLifecy…'
× adds the wip, review and complete lanes to the active set
```
**Nothing regressed — the guard got better and the ratchet didn't
follow.**
## What changed in the product
The stale-spec skip used to build its active set from
`resolveLifecycleColumns(...)`, taking `?.wip`, `?.review` and
`?.complete`. That returns the **first** column carrying each trait, so
a board with two wip lanes — or a review lane plus a second
merge-blocking one — had only one of each recognised as active. A card
in the other read as **inactive**, and its prompt file was treated as
reclaimable.
It now resolves the IR once and unions `columnsWithFlag` over five
flags, which returns **every** column carrying each:
```ts
const activeIr = await resolveWorkflowIrForTask(this.store, task.id);
const activeColumns = new Set<string>(["in-progress", "in-review", "done"]);
if (activeIr) {
for (const flag of ["countsTowardWip", "mergeOrchestration", "mergeBlocker", "humanReview", "complete"] as const) {
for (const lane of columnsWithFlag(activeIr, flag)) activeColumns.add(lane);
```
That is a real fix to an arity bug, and the legacy trio stays unioned in
for the documented reason: under-reporting active is the destructive
direction.
## Why re-point rather than loosen
This is a **source ratchet** in the `engine-no-blocking-shellout` style,
and its entire value is that it fails on a revert. A substring loose
enough to match both the old and new shapes would keep the file green
through exactly the regression it exists to catch.
So the assertions now name the new shape precisely, including the **flag
list**, so dropping one of the five is caught here too.
**Mutation-verified:** replacing the IR resolution with `undefined`
fails the ratchet. It still does its job.
## Scope
The file's own header notes this is *not* a behavioural proof — the
guard sits inside `execute()` behind worktree and session setup a unit
test has no business standing up, and it asks whoever next touches that
scaffolding to add the end-to-end case. Re-pointing is maintenance; I
have not taken on that harness here, and the note still stands.
## Verification
- `executor-stale-spec-active-lanes.test.ts` — **4/4**, fails on revert
- `pnpm lint` — clean
Test-only; no changeset. Found by running the full `engine-default`
project after #2855 merged — the remaining `main` red is my own audit
case, fixed by the pending #2857.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
9e1deffa72 |
fix(reliability): the review-failure headline read 0% because both inputs were zero (#2861)
The Reliability panel's headline metric is computed from two queries
that name lanes:
```ts
scopedStore.getTaskMovedCountsByDay({ …, toColumn: "in-review" }),
scopedStore.getTaskMovedCountsByDay({ …, fromColumn: "in-review", toColumn: "in-progress" }),
…
const headline = inReviewFailureRate7d(enteredByDay, bouncedByDay, nowMs);
```
On a board that renamed either lane, **both return `{}`**. Every per-day
count is zero, and the headline then divides one zero by another and
reports a healthy rate.
**This is the worst shape a lifecycle defect takes.** It produces a
*number*, not an error, and the number is reassuring. An operator
reading a 0% review-failure rate beside a populated audit event list has
no reason to suspect the metric is blind — the same failure mode as the
analytics `tasksInProgress: 0` sitting next to correct cost totals.
## Why a union is correct here, not a compromise
`getTaskMovedCountsByDay` takes **one** column per side, so the lanes
are resolved to sets and the query is issued per `(from, to)` pair and
summed.
The important part is *why* summing over a union is the right answer
rather than a widening hack. These read **move history**, and a past
move recorded the column name as it was at the time — the same reasoning
that keeps `tasksEnteredInReviewPerDay` in this module matching recorded
values verbatim, marked DELIBERATE-LITERAL. A board renamed last month
therefore has old rows under the old id and new rows under the new one,
so the correct query covers **both**. That is exactly the set
`resolveProjectColumnsForRoles` returns, with the legacy id always
unioned in.
Asking for either name alone is the bug — not a choice between them.
No double-counting: a move event has exactly one `(from, to)` pair, so
the queries partition the events rather than overlapping.
**The common path does not get more expensive.** On the built-in board
this issues the same two queries as before, and there is a test
asserting the call count so a future change cannot quietly turn one
query into N.
## Revert proof (measured)
Collapse the sets back to the single literals:
```
AssertionError: expected { '2026-07-01': 1 } to deeply equal { '2026-07-01': 2, '2026-07-02': 3 }
AssertionError: expected { '2026-07-01': 1 } to deeply equal { '2026-07-01': 2, '2026-07-03': 5 }
```
The helpers live in `reliability-metrics.ts` rather than inline in
`server.ts` specifically so they are testable without booting an express
app.
## Verification
- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm smoke:boot` — PASS (`fn --help`, real `serve` `/api/health` 200,
clean shutdown)
- `pnpm lint` — clean
- `tsc --noEmit` (`@fusion/dashboard`) — clean
- `reliability-metrics.test.ts` — 15 passed
- census `--strict` — exit 0 (unchanged: query filters are not
comparisons)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b9785ec10f |
fix(tests): stop the census guard reddening main on every legitimate conversion (#2856)
## This is the cause of four main reds today, not a fifth instance of them I have now fixed the census baseline on `main` three times (#2814, plus a withdrawn branch, plus watching #2811 and #2844 do the same). Rather than do it a fourth time, here is why it keeps happening. `census-baseline-corruption-guard` asserted: ```ts expect(result).toContain("every file matches its baseline exactly"); ``` That demands the **committed baseline be byte-in-step with the tree at all times**. **It is not, by design.** A conversion PR that removes guards leaves the tree holding *fewer* than the baseline allows, and the CLI treats that as the good case — it tightens the pin and exits 0. Measured directly: ``` simulated drop → EXIT ON DROP: 0 "The baseline file has been rewritten downward. COMMIT IT … in CI this write is discarded with the runner, which is why the gate is green and not silent." ``` So the ratchet was already happy while this test went red. Every conversion that did not *also* re-record the baseline turned `main` red for a condition that was never a defect. That is the mechanism behind **#2783's markers, #2837's query split and two more** — plus three collisions between workers racing to re-record the same file (#2811/#2814, #2844, and a branch of mine I deleted rather than open as a duplicate). ## The fix matches the guard's own stated intent Its comment says: *"a guard that always fails is no guard"* — its job is to prove the **corruption** diagnosis does not false-positive on a healthy file. **A tightened baseline is healthy.** So it now asserts what that needs: - the run **succeeds** — `execFileSync` throws on a non-zero exit, so a **rise still fails before any assertion runs**; a rise is real debt and must stay loud - the corruption diagnosis is **absent** - the outcome is one of the two healthy shapes the CLI can report ## Measured discrimination — all four cases | scenario | before | after | |---|---|---| | **drop** (legitimate conversion) | ❌ 1 failed — *the false red* | ✅ 3 passed | | **rise** (real debt) | ❌ fails | ❌ **still fails** | | **corrupt JSON** | ❌ fails | ❌ still fails (case unchanged) | | **unreadable file** | ❌ fails | ❌ still fails (case unchanged) | Only *"somebody converted guards and has not re-recorded the pin yet"* stops being a red. ## Scope Engine **11022 passed / 0 failed** · gate **732 green** · lint clean. Test-only. **Does not change** the CLI, the ratchet, or what `--strict` reports. Re-recording the baseline on a drop is still the right thing to do — it just stops being an emergency that reddens main and blocks everyone else while three people race to fix it. The baseline file is restored byte-clean after every simulation above (`git diff` verified). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Tests** - Expanded baseline validation coverage to detect unreadable or invalid baseline data. - Added support for both exact baseline matches and successfully tightened baselines. - Improved health checks for baseline verification outcomes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9d103fb171 |
test(engine): record the partial planning-continuation conversion the audit caught (#2857)
## The audit fired, as designed `workflow-planning-continuation-terminal-gap-live-e2e.pg.test.ts` went red on `main`: ``` × AUDIT — one of the two classifier call sites is converted, and the inner predicate is not → expected 2 to be 1 ``` That is the alarm working. #2799 built this case to fail **in both directions** — a new unconverted caller *or* a conversion of an existing one — precisely so a change here cannot land silently. Someone converted the inner half: ```ts if (!isPlanningContinuationTaskDispatchable(task, terminal)) { ``` ## What actually changed, and what did not | site | before | now | |---|---|---| | `drainDuePlanningContinuations` | passes `{ terminalColumns }` | unchanged ✅ | | `resolvePlanningContinuationCandidate` → inner predicate | **not threaded** | **threaded** ✅ | | `selectActionablePlanningContinuations` | passes nothing | **still passes nothing** ❌ | **The behavioural case is untouched and still passes.** A card in a renamed COMPLETE lane still comes back `actionable` and still re-enters plan-review, because `selectActionablePlanningContinuations` is the site that decides that. That combination is the useful shape of a **partial conversion**: the code reads more converted than before, and the operator-visible defect is exactly where it was. Without the behavioural case sitting next to the audit, the arity change would have looked like the fix. The fixer's own comment agrees, and names the narrow case the inner half reaches — a board declaring `done` as a *non-terminal* column id, where the outer check passes and the inner one skipped the continuation as "paused". Real, and not the case the characterization covers. ## The change Arity assertion updated `1` → `2`, with the reasoning recorded at the assertion and in the header. **Updated deliberately, not loosened** — arity is still the property asserted, so the next change to either site lands here again. Nothing else moved. ## Verification - suite — **4/4** - full live-PG E2E surface — **171/171** - `pnpm lint` — clean Test-only; no changeset. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7784cb1fe8 |
self-healing: six recovery sweeps that never ran on a renamed board — and the guards widening their queries activates (#2838)
**Six self-healing sweeps did not run at all on a renamed board. Each is a recovery path — the thing that unsticks a card when something has already gone wrong.** #2800 measured this class and could not fix it: a read happens *before* any task is in hand, so there is nothing to resolve a per-task lane from. `resolveProjectColumnsForRoles` (landed separately) is the seam that was missing. ## What was silently dead | sweep | what stayed broken on a renamed board | | --- | --- | | `reconcileDoneTaskIntegrity` | a landed card kept **no commit sha**, forever | | `recoverAlreadyMergedReviewTasks` | a card whose merge **succeeded** stayed parked with `status: "failed"` | | `recoverStuckMergeDeadlocks` | **doubly blind** — no candidates *and* no dependents | | `recoverInterruptedMergingTasks` | a task interrupted mid-merge sat in `merging` indefinitely | | `recoverMergeableReviewTasks` | a card ready to merge was never re-enqueued | | `recoverReviewTasksWithFailedPreMergeSteps` | a card parked on a failed review step was never revived | The census scored the `task.column === "..."` re-assertion *inside* each loop, never the query above it. Converting those comparisons would have dropped six counts and changed nothing — the loop bodies were already unreachable. ## The conversion shape — five parts, three of which review taught me Documented in `self-healing-sweeps-are-blind-on-a-renamed-board.md`, because the second sweep **drifted from the first**: I wrote it from the pre-review version and reproduced a flaw already fixed one commit earlier. 1. **Read** — project union, query each column, dedupe by id. Legacy ids unioned so a board mid-rename is not skipped. 2. **Verdict** — per card against **its own** workflow. Widening the read and widening the verdict are different decisions: *a missed row is invisible, a wrong row is a write.* Using the project union as a per-card test claims a card because some **other** board calls its column that role. 3. **Provenance** — the resolver **substitutes** the built-in IR rather than failing, so `length > 0` reads as "this card answered" when nobody did. It does not change the verdict (measured: identical) — it makes the unrepaired card **reportable**. 4. **The log strings** — widening a query invalidates every message naming the old literal. One logged `"stale merging task(s) in in-review"` after its read covered several lanes. 5. **The guards the query ACTIVATES.** ## Part 5 is the one that bites A guard downstream of a literal query is **unreachable** on a renamed board — and unreachable is indistinguishable from correct. That is why these sit unwired indefinitely. `recoverReviewTasksWithFailedPreMergeSteps` filters on `blocker !== "task has failed pre-merge workflow steps"` — an **exact string match**. Unwired, the blocker returns `"task is in 'checking', must be in 'in-review'"`, so widening the query alone would have made the sweep **find every card and reject every card**. Measured: **6 sweeps hold both a literal query and an unwired lane guard**; 30 hold a literal query with no such guard. All six are named in the doc. **One of the six was my own already-converted sweep.** I widened `recoverAlreadyMergedReviewTasks` two commits before noticing its `getTaskHardMergeBlocker` was unwired — so for two commits it found renamed-board cards and declined them. The scan must run **before** widening; I did it after, and only caught it because the next sweep forced the question. `getTaskHardMergeBlocker` was the blind spot for four of the six: a wrapper, no lane parameter at all, every caller behind a literal query. ## Corrections to my own work, kept visible - The project union used as a **per-card verdict** — the flat-set mistake `project-lane-vocabulary.ts` warns about in its own header, which I quoted while writing it. - A **provenance fix that was a no-op**: measured identical verdicts in every state, revert passed its own new test, so it was thrown away rather than shipped with a comment claiming otherwise. - The second sweep **reproducing the first's pre-fix shape**. - Three assertions that were **vacuous until the revert exposed them** — including one where the write needed a real git repo, so `commitSha` could not distinguish accepted from rejected. ## Verification - `pnpm test:gate` — 161 + 487 + 13 + 71 - `self-healing.test.ts` 412, query-blindness suite 12 - `tsc` on core and engine; `pnpm lint`; `check:changesets`; census `--strict` — all clean, each run explicitly - Every conversion revert-measured, **each direction independently** where a sweep has two (read and guard) ## Scope **42 queries remain**, 5 of the 6 activation-risk sweeps among them. Each is per-sweep work — its own filter semantics, its own downstream guards, its own log strings — so they land one at a time with the pattern proven, never swept. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Self-healing workflows now work correctly on boards with renamed lifecycle columns. - Improved recovery for completed, in-review, interrupted, stalled, and failed-merge tasks. - Prevented tasks from being incorrectly classified using another workflow’s columns. - Added warnings when a task’s workflow lanes cannot be resolved. - **Documentation** - Expanded guidance on renamed-board recovery behavior and related diagnostic limitations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
68fd781eb8 |
test(engine): stop engine-default exiting 1 with zero failing tests (#2855)
## The problem
On `main`, the whole `engine-default` project exits **1 while reporting
736 files passed and 0 failed**:
```
Test Files 736 passed | 1 skipped (737)
Tests 9944 passed | 14 skipped | 1 todo (9959)
Errors 9 errors
```
Nine unhandled rejections — `TypeError: this.store.listTasks is not a
function`, all from `self-healing-db-corruption.test.ts`, raised *after*
its tests had passed.
That is worse than a plain failure: the lane is red, **nothing names a
test**, and a genuine regression landing later arrives in an already-red
lane where it reads as more of the same noise.
## Why the existing isolation didn't hold
The suite stubs every other maintenance sweep so only the corruption
path runs. But `runMaintenance` invokes the surfacing family like this:
```ts
{ name: "surface-in-review-stalled", fn: () => this.surfaceInReviewStalled(maintenanceSurfacing()) },
```
`maintenanceSurfacing()` lazily calls `openSurfacingCycle()`, which does
`await this.store.listTasks({ slim: false })`. **Spying on
`surfaceInReviewStalled` replaces the method, but the call site still
evaluates its argument** — so the shared cycle opened regardless,
against a mock store that deliberately implements only what corruption
surfacing needs.
## Two fixes that look right and are not
Both tried, both reverted — recorded because each is the obvious next
move:
| attempt | result |
|---|---|
| add `listTasks` to the mock | **3 tests fail** — the sweeps then run
far enough to record audit events the assertions don't expect |
| stub the 27 maintenance methods missing from the suite's lists | **5
tests fail, and all 9 errors remain** — `openSurfacingCycle` isn't among
`runMaintenance`'s own calls, so the drift was a real but unrelated gap
|
The second result is what located the fault: stubbing everything
`runMaintenance` calls does not silence the rejections, so they don't
originate there. The stack confirmed it —
`SelfHealingManager.openSurfacingCycle` at `self-healing.ts:8015`. I
should have read that stack before guessing twice.
## The fix
Stub the cycle itself. `null` is exactly what `openSurfacingCycle`
returns when the engine is paused, so every consumer already handles it,
and the corruption-surfacing path this suite exists to test is
untouched.
## Verification
- `self-healing-db-corruption.test.ts` — **6/6, 0 errors, exit 0**
- full `engine-default` project — **736 files passed, 9944 tests, 0
errors, exit 0** (was exit 1)
- `pnpm lint` — clean
Test-only change; no production file is touched, so no changeset.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
92d82b7a17 |
fix(guard): the unwired-lane check reported 0 because its question was too weak (#2852)
A guard nobody has proven can fail is a number, not a check. This one
was returning a clean `[]` while **18** real unwired lane declarations
sat on `main` — including one it was specifically built to catch.
## The escape that found it
`diffSnapshots` in the glasses plugin:
```ts
opts: { notifyOnColumns: ReadonlySet<ColumnId>; completeColumnsByTaskId?: ReadonlyMap<...> }
...
const completeColumns = opts.completeColumnsByTaskId?.get(task.id);
const isComplete = completeColumns ? completeColumns.has(task.column) : task.column === "done";
```
**No file anywhere builds that map.** The conversion was decorative —
the literal decided every real poll, so on a renamed board the wearer is
notified of every column transition *except the card finishing*, the one
they care about. The name was already in the guard's vocabulary list,
the declaration was exported and optional; it satisfied every condition
the guard checks, and the guard said nothing.
## Three independent blind spots
| # | blind spot | why it mattered |
|---|---|---|
| 1 | `SCANNED_PACKAGES` omitted `plugins/` | plugins hold lane logic
like anything else — this one resolves workflow IRs and decides what
"finished" means |
| 2 | inline options-object types were not walked | only bare parameters
and *named* interfaces were. Whether a lane answer arrives as a
parameter, an interface property, or an inline field is a style choice —
**a check evadable by a style choice is decorative** |
| 3 | the mention rule was `source.includes(parameter)` **anywhere** |
satisfied by coincidence for any ordinarily-named parameter |
Fixing 1 or 2 alone would still have missed it: **measured on `main`,
the guard found 0 unwired across 1753 files, and 0 again across 2114
once `plugins` was added**, because the shape was invisible too.
### On (3), I proved it on myself
Renaming the unwired parameter from `completeColumnsByTaskId` to
`completeColumns` — a better name, chosen for good reasons — **silenced
the guard instantly**, because 15 unrelated production files declare a
local called `completeColumns`. The check had not been satisfied; it had
been switched off by a rename. That is exactly the failure the code
comment two lines up condemns, committed one edit later.
The fix is the cause, not a name blocklist: a file that never references
`diffSnapshots` cannot be the thing that wires `diffSnapshots`'s
options. Still deliberately loose — a co-occurrence test, not call-graph
analysis — which keeps the low false-positive rate that makes the guard
bearable while removing a false **negative** that scaled with how
ordinary a parameter's name was.
## The 17 this uncovered
Tightening (3) surfaced 17 further unwired declarations across core,
engine and dashboard. **Spot-checked, not assumed**:
`buildUnblockWeightMap` in `task-priority.ts` declares `terminalColumns`
and `reviewColumns`, and the only files that pass either are its own
tests — the production caller silently uses the built-in `{done,
archived}` default. That is the inert-conversion shape this module
exists to name.
They span three other batches, so they are recorded as a **ratcheted
baseline** in the shape `scripts/lifecycle-column-census.mjs` already
uses here — keyed on `file + parameter` so an unrelated edit above them
cannot manufacture a failure. A new one fails immediately; these can
only leave the list. Listing them beats pretending for another week that
they do not exist.
## The glasses fix
`completeColumnsByTaskId` -> a flat `completeColumns` set, matching its
sibling `notifyOnColumns` in the same options object, resolved **once
per poll** by `notifier.ts` via `resolveProjectColumnsForRoles`.
Project-scoped and not per task because this runs on a polling timer
over the whole board — a per-card workflow read would scale with the
board on every tick. Best-effort: a failed resolve leaves the diff on
its documented legacy default rather than dropping a poll.
Still gated by `alsoNotifyOnDone`, which the production caller passes as
`false`, so it remains unobservable at runtime. Wired anyway: the day
someone enables the flag the resolution must already be right — and now
the guard will say so if the wiring is removed.
## Revert proofs (measured, one per fix)
| revert | failure |
|---|---|
| unwire `notifier.ts` | baseline gains `plugins/…/diff.ts
completeColumns` |
| drop the inline-options walk | "covers an INLINE options-object type"
fails `expected [] to deeply equal [ 'completeColumns' ]`, and the repo
scan loses the glasses entry |
| drop the owner scoping | the repo scan loses **all 17** pre-existing
entries |
| restore `task.column === "done"` | both new `diff.test.ts` cases fail
|
The new diff cases assert **both** directions — a card in the resolved
lane fires, and a card in the legacy `done` does *not* once the caller
resolved other lanes. The second is what proves the resolved set
replaces the default rather than being unioned with it.
## Verification
- `pnpm test:gate` — 161 / 487 / 13 / 71 passed
- `pnpm lint` — clean
- `tsc --noEmit` — clean
- guard suite — 9/9; full glasses plugin — 188 passed across 19 files
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
eed8ca55fc |
batch-dashboard-src: the planner metrics tool froze active runtime on a renamed execution lane (186 → 185) (#2842)
`packages/cli` and the plugin packages are at **zero** lifecycle guards, so this picks up the nearest unowned work: the `packages/dashboard/src/` remainder. ## The defect `activeRuntimeMs` adds the wall-clock since `executionStartedAt` only while the card is accruing work — the **WIP role** — but it was keyed on the literal `in-progress`. On a board whose execution lane is renamed, that live tail was dropped, so `fn_task_planner_get_task_metrics` reported active time frozen at whatever the last completed segment left in `cumulativeActiveMs`. The number stayed plausible, which is why nothing surfaced it. ## The part worth reading: the wiring had no watcher, from either direction I wired the producer (`chat.ts` resolves the task's own lanes via `wipColumnsForTask`) in the same commit, then checked whether that wiring was actually covered. It was not: - **Deleting the `wipColumns:` argument left the entire 3830-test dashboard suite green.** The formatter's own tests inject the set by hand, so they prove the *guard* and are structurally blind to whether production fills it. - **`check-inert-flag-seams.mjs` does not see it either.** It tracks trailing optional **parameters**; this is a property inside an options bag. That is a real gap in the checker — every seam expressed as an options-bag property is currently unguarded. Reported here rather than fixed, because #2822 and #2830 both already modify that script and a third change would guarantee a three-way conflict. So `createTaskPlannerMetricsTool` is exported and a second test drives it, letting it do its **own** resolution against a renamed board. Deleting the argument now fails 1 of 2. ## Census | | before | after | |---|---|---| | COLUMN guards | 186 | **185** | | `packages/dashboard/src/task-planner-chat-metrics.ts` | 1 | **0** | Baseline re-recorded; `--strict` exits 0. ## Two findings I did NOT act on, deliberately **1. `github-tracking-state.ts` keeps 2 counted guards and should.** They are the documented degraded-mode arms of a fully-resolved classifier (`completeLanes === undefined ? columnId === "done" : ...`). Marking them `DELIBERATE-LITERAL` would drop the count by **reclassification rather than conversion** — the exact move the census's own strict-check warns about. Related: the census reports `0 are trait-fallback branches (already converted)`, yet these are precisely that shape, so the trait-fallback classifier appears not to recognise a ternary whose fallback arm is the literal. Worth a look by whoever owns the census. **2. Three pre-existing failures in `packages/dashboard/src/__tests__`, unrelated to this change** — measured identically on `origin/main` before and after: - `planning-browser-e2e.test.ts:353` - `register-model-routes-kimi-k3-supplemental.test.ts:60` - `routes-tasks-near-duplicate.test.ts:274` Flagging rather than touching them; per the standing rule they are quarantine candidates, not appeasement candidates. ## Verification Dashboard `tsc` clean, `pnpm lint` clean, census `--strict` 0, `check-inert-flag-seams` 21/21 supplied, changeset lint clean. Targeted suites: `task-planner-chat-metrics.test.ts` 8/8, `task-planner-metrics-tool-wip-lanes.test.ts` 2/2. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
89aaf341d0 |
the unwired-seam audit: 9 defects the census cannot see, incl. a reviewed card that cannot merge (#2820)
**Nine operator-visible defects in a class the census cannot see, plus
the audit method that found them.**
The census scans for lifecycle-column **comparisons**. This PR is about
guards that have no literal to find: a helper takes an optional
*resolved* lane set, its own test passes it, the census entry is gone —
and the callers pass nothing. **A resolved seam nobody wired is
indistinguishable from no seam at all.**
## What was broken
| defect | operator sees |
| --- | --- |
| `getTaskMergeBlocker` unwired in `mergeTaskImpl` | `Cannot merge FN-1:
task is in 'checking', must be in 'in-review'` — **a reviewed card
cannot merge** |
| …and in the completion move | `Cannot move FN-1 to done: …` — **and
cannot complete** |
| `isParkedTaskColumn` unwired ×2 (`agent-heartbeat`) | a durable agent
keeps claiming a parked card; **Health Check renders it RUNNING** |
| `resolveLinkSyncColumnRoles` first-per-role | link hygiene skips a
**second hold lane** entirely |
| `executor` active-task predicate first-per-role | a card in a **second
wip lane reads as INACTIVE**; its prompt file becomes reclaimable |
| `isPlanningContinuationTaskDispatchable` partially threaded | a board
declaring `done` as *non-terminal* stalls its cards — **stalled by a
lane name** |
| `default-workflow-hooks:72`, `executor:2404` | resolved gate admits
the move, unresolved blocker refuses it |
## The recurring shape, which is sharper than "a caller forgot an
argument"
Four sites resolve the lane and then re-ask with the literal, **a few
lines apart in the same function**:
- `task-artifacts-ops` resolves `completeColumn`, then asks the blocker
with the literal.
- `default-workflow-hooks:72` gates on `lifecycleColumns?.review`, then
the literal.
- `executor:2404` compares `resolveResumeLanes(…).review`, then the
literal.
- `resolvePlanningContinuationCandidate` applies the caller's terminal
set, then delegates without it.
**Grep for the helper, not the literal.** The literal is one function
away, correctly annotated as a fallback — which is exactly why the
census is blind to all of it.
## The arity trap, named and measured (six occurrences, one caught by
review here)
`resolveLifecycleColumns` answers *"which column is **the** hold
lane?"*. A `.includes()`/`.has()` test asks *"is this **any** hold
lane?"*. Nothing distinguishes them — same types, no literal.
**A default-vs-renamed differential cannot catch it**, because the
default board declares one column per role and therefore cannot express
the failing shape. It needs a *structurally* different fixture. That is
a sharper rule than "test both vocabularies", and it would have caught
all six.
Scanned: 12 candidate sites. **4 fixed · 3 blocked (2 on the inert sync
IR reader; `triage:833` also query-shaped) · 1 needs a hook-contract
change · 3 not defects (a returned tuple; an ordering-sensitive
precedence list) · 1 false positive of my own scan.**
A sweep over all twelve would have broken the ordering-sensitive pair,
delivered nothing at the sync-blocked ones, and "fixed" a site that was
already correct.
## Two traps in fixing this class — I hit both here
1. **The legacy id is a FALLBACK, not a member.** Pre-seeding
`"in-review"` admits a board that *declares* `in-review` as its WIP
column — a card mid-implementation merges prematurely. A real resolved
answer must **replace** the default. (Caught by review; it is the same
unscoped-legacy-acceptance the glasses plugin's review caught earlier,
which I had read and reintroduced.)
2. **Two guards, one assertion.** `toContain("must be in")` passed with
`mergeTaskImpl` reverted, because the *completion* guard caught the card
instead. The assertion now names the site (`Cannot merge` vs `Cannot
move … to done`) so the two fail independently.
## Corrections I made to my own work, recorded rather than quietly fixed
- My first PG test was **vacuous three ways**:
`saveWorkflowDefinition?.()`/`setTaskWorkflowSelection?.()` do not exist
(the `?.` swallowed both, so the task kept the builtin workflow),
`updateTask({column})` does not move a card, and a two-node IR made
every setup move illegal. Premise is now **asserted**, not assumed.
- My doc claimed the audit was complete. It enumerated **helpers**, not
every **caller** — `getTaskMergeBlocker` alone has 13 call sites.
Corrected in place, with the still-unwired ones listed by file and line
and a note to distrust any "audit complete" claim including mine.
- A severity correction to another worker's E2E:
`selectActionablePlanningContinuations` has **no production caller**, so
its stated consequence is latent, not live.
## Verification
- `pnpm test:gate` — 161 + 487 + 13 + 71
- `tsc` on core and engine; `pnpm lint`; `check:changesets`; census
`--strict` — all clean, each run explicitly
- Every fix revert-measured; each has a non-vacuous companion. The
two-hold-lane and repurposed-`in-review` cases exist because the default
board cannot express those shapes.
## Deliberately not done, with reasons in
`resolved-seams-nobody-wired.md`
`isTaskReadyForMerge` (dead in production — wiring it would be the
anti-pattern itself); `getTaskHardMergeBlocker` (3 of 4 callers are
query-gated sweeps); `getInReviewStallReason` (needs a **batch
prefetch**, not a per-task resolve — its callers decorate every task on
every list read; the in-review stall badge is wrong on renamed boards
until then); `default-workflow-hooks` planning/live-work sets (needs
`DefaultWorkflowMoveContext` to carry the IR — a shared contract
change).
|
||
|
|
aca25494a5 |
test(dashboard): App.test.tsx is fully green — an incomplete mock was crashing NewTaskModal (#2846)
Closes #2829. `App.test.tsx` has been red since **2026-07-25**; #2833 took it 10 failures → 2, and this takes it to **0**. ## The cause, and it was not what I guessed `vi.mock("../../hooks/useViewportMode")` did not export `isShortViewport`, which `NewTaskModal` imports. An incomplete module mock **does not fail at the mock**. It throws inside whichever component imports the missing export, and the nearest `ErrorBoundary` swallows that into *"This section encountered an error"* — so the test failed on a **missing heading** with a DOM that looks perfectly healthy. Structurally the same presentation as the workflow-skeleton bug in #2833: the real failure sits two layers below what the assertion reports. Both remaining tests had this one cause. ## Found by probing, not theorising I had already been wrong twice on this file (`renders nothing`, then i18n), and my standing note said the modal was "plausibly FN-8620's FloatingWindow rework, **unverified**". That would have been a third wrong guess. Instead I dumped what was actually in the DOM at the assertion point: ``` headings: ["Fusion"] | dialogs: 0 newtask-ish: [..., "error-boundary error-boundary--modal"] ``` The `error-boundary--modal` class was the entire answer; reading its text gave the exact missing export. Same technique that cracked #2833 in one shot — the difference between the two halves of this investigation is that I stopped guessing. **App.test.tsx: 141 passed (141).** ## Also patched, and one deliberately not - **`Header.mobile-project-favorites.test.tsx`** mocks the same module with an inline factory and lacks the export. It is **latent, not broken** — its component does not import `isShortViewport` yet, and the day it does, the failure would be this same swallowed crash. One line. - **`FloatingWindow.touch-geometry.test.tsx`** greps as missing but is **not** — it uses `{ ...(await vi.importActual(...)) }`, so its surface is complete by construction. My grep-based classification was a false positive. That spread is the pattern that prevents this class outright, and it is the better fix **where it is available**. It is not available here: App.test.tsx's mock exists precisely so tests need no `window.matchMedia` in jsdom, and spreading the real module would reintroduce that dependency for every unstubbed export. So the scoped one-line stub is the right fix for this file, and the spread is the recommendation for new mocks. A guard asserting "a module mock exports everything the real module does" would close this class repo-wide. I did not build it — 34 files mock this module and only one was genuinely wrong, so the population does not yet justify another instrument, by the same measure-first rule that killed two other candidate gates this week. ## Verification App.test.tsx 141/141 · the two touched siblings 9/9 · `tsc` 0 · lint 0. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0e647f320 |
test(dashboard): name the skeleton in board-mobile-view-switch instead of the shared class (#2848)
Closes the last item I left open on #2834 — the two cases that were passing against a skeleton. ## The weakness They asserted `.list-view` to mean *"we are in list view"*. The workflow **skeleton** carries that same class, so they matched it and passed **without the list ever drawing**. I found this while adding the `list-view-body` marker in #2834, tried to fix it, hit a dead end, and reverted rather than leave two green tests red. ## Fixed by making the assertion honest, not by forcing the real body This harness renders `<ListView>` with a deliberately small `../../api` mock and no workflows, so the skeleton **is** what appears — and that is fine, because the subject of these cases is the **board's** structure after switching, not the list's contents. Naming the skeleton says what the test actually observes. ## Why not the real body — measured, and it explains the earlier dead end Supplying the shared `DEFAULT_BOARD_WORKFLOWS` fixture makes ListView render its full body, which **throws** under this file's api mock. The harness has no `ErrorBoundary`, so React unmounts the entire root: the DOM goes completely empty and every later query fails with a misleading `Unable to find switch-to-board`. That is precisely the failure I could not explain on #2834. A probe placed after the await settled it: ``` PROBE after-await testids: [] | boundary: none ``` Empty DOM, no boundary — an uncaught render error taking the root down, presenting as a missing button three lines later. Upgrading this to the real body means expanding an api mock inside a file about board CSS structure, which is a bigger change than the assertion is worth. The reasoning is recorded at the site so the next person does not rediscover it the way I did. ## The assertion is load-bearing, not just renamed Mutating the skeleton's own `data-testid` in `ListView.tsx`: ``` TestingLibraryElementError: Unable to find an element by: [data-testid="list-workflows-skeleton"] (x2) ``` Restoring it: `Tests 3 passed`. ## Verification 3/3 · `tsc` 0 · lint 0. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bebbdf9083 |
fix(tests): main red — archive restore returns to the ARCHIVED lane now (#2832) (#2847)
## Red on main ``` store-archive-reads > TaskStore archived read parity (PostgreSQL) > rebuilds a missing live row before consuming its archive snapshot AssertionError: expected 'done' to be 'todo' ``` ## The product change is a real fix **#2832** found that `preArchiveColumn` has no database column — it exists on the `Task` type and in the archive snapshot and nowhere else — so the old code fell through to a literal and decided the destination the same way for **every unarchive that has ever run**. On a custom board that meant a card archived mid-implementation came back marked *finished*. That PR flipped its own characterization cases. This one, in a different file, was missed. The fixture creates the card in `done`, so `done` is the answer now. ## I measured a second sample before encoding a rule The behaviour is **narrower than "restores to the lane it came from"**. With the fixture changed to `in-progress`, restore returns **`todo`** — not `in-progress`: | archived from | restored to | |---|---| | `done` | `done` | | `in-progress` | **`todo`** | A terminal lane is preserved; a WIP lane is re-queued. That is plausible product behaviour — a card cannot resume mid-execution after a restore — but it is **not what #2832's summary describes**, so I have flagged it there rather than encoding it here. If re-queueing WIP is deliberate it deserves its own case; if it is not, this snapshot-rebuild path still carries the defect #2832 fixed elsewhere. ## An honest limitation, recorded in the test This assertion is **weaker than it looks and cannot be strengthened here**: `done` is also the complete lane, which is exactly what the pre-#2832 *"no usable history"* branch returned. A card archived from `done` therefore reads identically under both the fixed and the broken implementation. My first attempt "strengthened" it by moving the fixture to `in-progress` — that is what surfaced the second behaviour above, and shipping it would have encoded a rule inferred from two samples. Reverted; the limitation is documented instead. ## Scope Only the line-205 case is touched. Line 218 asserts `todo` for a card genuinely archived from `todo` and still passes — the two are not the same claim. Core **4813 passed / 0 failed** · gate **732 green** · lint clean. Test-only. ## Pre-flight results for the current queue Merged-with-main, engine + core on each: | PR | result | |---|---| | #2822, #2819, #2823, #2818 | only the 2 inherited `workflow-ir-resolver` failures (fixed in #2836, now merged) | | #2805 | clean | | #2828 | inherited only, once this and #2836 land | | #2830 | inherited only | | #2820, #2803, #2808 | **conflict** with main — census baseline; told the owners to regenerate rather than hand-merge | 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
51934931e1 |
fix(board): the awaitingPlanning badge only ever worked on a lane named "todo" (#2845)
Converts the one site in `register-task-workflow-routes.ts` that a previous pass **deliberately deferred**, and does it in the shape that note asked for. ## What the deferral said ``` FNXC:WorkflowResolvedColumns 2026-07-29-00:00 (U12 — R8, DELIBERATELY NOT CONVERTED): … I converted it and then REVERTED: resolving each task's hold column needs a per-task workflow read, and this is the board-load path whose own comment above exists because unbounded reads here "turn a board load into thousands of reads". Converting it properly needs the hold column resolved per WORKFLOW from data the board payload already carries, not per task from the store. ``` That was the right call and the right diagnosis. `resolveProjectColumnsForRoles(store, ["hold"])` is exactly the project-scoped shape it names: **one** `listWorkflowDefinitions()` read per board load, flat in task count. The expensive part — a PROMPT.md read per row — is untouched and still bounded by `AWAITING_PLANNING_ENRICH_LIMIT`. The test asserts the flatness directly (`listWorkflowDefinitions` called exactly once), so a future per-task regression fails here rather than being discovered as board latency. ## What was broken The filter named `todo`, so on a board whose waiting lane is called anything else **no row was enriched at all** — no error, no log line, just a silent fall back to the client's `steps.length === 0` heuristic. That heuristic is precisely what this enrichment was added to correct, so the card most likely to be mislabelled — real spec, zero parsed steps, already a scheduler dispatch candidate — sat on "Queued to plan" indefinitely. Over-inclusion is the safe direction and is chosen deliberately: a card in some other workflow's hold lane gets annotated as waiting, which is what a waiting card in a waiting lane should show. ## Revert proof (measured) Restore `task.column === "todo"`: ``` FAIL register-task-workflow-routes.awaiting-planning.test.ts > enriches a card in a RENAMED hold lane, not only one literally named todo expected undefined to be false ``` The other 8 cases in the file stay green — their harness store declares no `listWorkflowDefinitions`, so they run the degraded legacy-`todo` path. That compatibility is half the contract, which is why the new case brings its own store rather than widening the shared harness. ## Census | file | before | after | |---|---|---| | `packages/dashboard/src/routes/register-task-workflow-routes.ts` | 3 | 2 | The 2 remaining in that file are documented trait-fallback branches, not unconverted debt. The baseline also picks up `packages/core/src/task-store/moves.ts` 2 -> 0, already true on main and not from this diff. ## Verification - `pnpm test:gate` — 161 / 487 / 13 / 71 passed - `pnpm lint` — clean - `tsc --noEmit -p tsconfig.json` (`@fusion/dashboard`) — clean - `node scripts/lifecycle-column-census.mjs --strict` — exit 0 - targeted: `register-task-workflow-routes.awaiting-planning.test.ts` 9/9 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cea9637dfc |
feat(census): split the query class into read-shaped and write — most of the remainder must not be converted (#2837)
Splits the census's query class into **read-shaped** (convertible) and
**write** (must not be converted), reported beside the existing total.
On current `main`:
```
QUERY filters (column: "<legacy>"): 63
of those: 48 read-shaped (convertible), 5 writes (do NOT convert), 10 other
```
## Why the single number misleads
`column:` sits in an options-shaped object for both a source query and a
write, so the existing definition-vs-query rule cannot separate them.
The result reads as "dead reads to convert" — and after #2818 landed,
**48 of the 63 are `self-healing.ts` and the rest are largely not
convertible at all.**
Converting a write in this class is **harmful, not merely pointless**.
`async-persistence.ts` soft-deletes with `.set({ column: "archived",
deletedAt, … })`, and `getLiveTaskColumn` returns `"archived"` as a
**sentinel** for any soft-deleted row — the write and the sentinel have
to agree. A sweep that "finished the query class" by converting all 63
would break live-column resolution for every deleted task.
That is the same shape #2808 flagged for `recoveryRehome` moves. **Two
of the census-invisible classes now have members that must not be
fixed**, and in both cases the count alone cannot tell you which.
## Reported, not ratcheted
`properties.query` and `queryByFile` are byte-identical, so the pinned
baseline does not move and no open PR's Lint changes. The split is one
extra line of output.
**Changing what a ratchet enforces is the owner's call; improving what
it says is not.** Same line I drew when making marker-only failures
legible without loosening them.
## Honest limit
Read-shaped is a better filter than the raw count and **still not a
verdict**. `auto-merge-finalization.ts:242` is classified read-shaped
and must NOT be converted — its own comment records that it is
`getTaskHardMergeBlocker`'s review-eligible sentinel, deliberately not
re-keyed. Nothing mechanical would catch that; only the comment beside
it does. The split narrows a haystack to a readable list; it does not
decide the list.
## Verification
4 new cases — a `listTasks` filter counts read, a `.set()` tombstone
counts write, IR node definitions stay excluded from the class entirely
(the pre-existing rule must keep working), and the pinned total is
unchanged by the split. Revert proof: dropping the write branch fails
the tombstone case.
Census's own suites **87 passed**, gate **161 / 13 / 487 / 71**, lint
clean, `--strict` exits 0.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f53c9dbd39 |
fix(core): merge re-enqueue threw on every board with a renamed review column (#2819)
The single most consequential finding from the u12 seam-gate work, picked up because batch-core (#2783) merged without it and it is now unowned. ## The defect `enqueueMergeQueueInTransaction` gates on the task's column being a review column, and takes the board's review columns as an optional trailing argument. - `moves.ts:487` and `moves.ts:1153` — the automatic handoff-to-review path — resolve and pass them. - The public `enqueueMergeQueue` wrapper (`async-merge-coordination.ts:246`), reached through `store.enqueueMergeQueue`, **did not**, so it fell back to `new Set(["in-review"])`. This is not the quiet legacy-id degradation most of these seams produce. The reject branch records `mergeQueue:enqueue-rejected` and **throws** `MergeQueueInvalidColumnError`. Its production callers are `merger.ts:7251` and `self-healing.ts:10329` — so on any board whose review lane is renamed, the merge and recovery re-enqueue paths failed outright while the handoff path kept working. ## Measured, not asserted With the fix reverted, the renamed case fails with the exact predicted error and the controls stay green: ``` × renamed vocabulary: a task in the RENAMED review lane enqueues for merge MergeQueueInvalidColumnError: Task KB-001 is in column 'checking', not 'in-review'; cannot enqueue ✓ default vocabulary: a task in the review lane enqueues for merge ✓ renamed vocabulary: a task in the WIP lane is still REJECTED ✓ default vocabulary: a task in the WIP lane is still REJECTED Tests 1 failed | 3 passed (4) ``` With it: `Tests 4 passed (4)`. ## About the suite **Differential.** The fixture is the builtin coding workflow with only its column ids renamed, so the sole difference between the two runs is vocabulary — a hand-built graph would test the fixture's own transition table as much as the code. It asserts the rename actually landed (`checking` present, `in-review` absent), so a surviving literal cannot pass by luck, and it walks the graph rather than jumping, because moves are transition-validated. **Both negatives included.** A WIP-lane task must still be REJECTED under each vocabulary. Supplying the real columns must not degrade into "every column is a review column", which would let work merge straight out of the WIP lane — the failure mode a careless version of this fix would introduce. ## Why nothing caught it Partial supply. Two of three call sites passed the argument, so a check asking "does SOME caller supply this?" reported the seam as satisfied, and the lifecycle-column census counted the conversion as done. Closing that one-supplier floor in `scripts/check-inert-flag-seams.mjs` (on #2772) is what surfaced it. ## Verification - `pnpm test:gate` green - new suite 4/4; neighbouring merge-queue suites (`taskstore-lifecycle`, `store-in-review-stall`, `runtime-lifecycle-async`) 29/29 - `tsc -p packages/core` 0, lint 0 - changeset included (`@runfusion/fusion` patch) Note: #2772 still carries a TEMPORARY per-call-site exemption for this seam. Once this lands, that exemption's staleness check will fail and I will remove it there — it cannot outlive the fix. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4c5080c1cf |
fix(tests): main red — the IR fallback is a BRANDED COPY, so identity can't hold (#2836)
## Red on main
```
workflow-ir-resolver > resolveWorkflowIrForTask > falls back to the built-in default when the definition is missing
workflow-ir-resolver > resolveWorkflowIrById > falls back to the canonical IR for an unknown built-in id
AssertionError: expected { version: 'v2', …(6), …(1) } to be { version: 'v2', …(6) }
Received: serializes to the same string
Compared values have no visual difference.
```
That message is the signature of an **identity-only** break, and that is
exactly what it is.
## Why identity can never hold again
Both asserted `toBe` against the exported builtin constant. **#2815**
added `markFellBack`:
```ts
function markFellBack(ir: WorkflowIr): WorkflowIr {
const copy = { ...ir } as WorkflowIr;
Object.defineProperty(copy, FELL_BACK_TO_DEFAULT, { value: true, enumerable: false, configurable: true });
return copy;
}
```
It brands the fallback so a caller can tell a **resolved** workflow from
a **guessed** one — the provenance contract #2618 introduced. Copying
*is* the mechanism, so these two paths cannot return the shared object.
Worth noting: the sibling `toBe` assertions in the same file **still
pass**. The no-selection and throwing-selection paths return the
constant unbranded, so only the two `markFellBack` paths changed — which
is why this presents as two failures rather than five, and why it is a
genuine contract change rather than a blanket refactor.
## Not just loosening to `toEqual`
Swapping `toBe` → `toEqual` alone would delete a real assertion. The
**brand is asserted too**, via the public provenance API rather than by
reaching for the private symbol:
```ts
expect(ir).toEqual(BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR);
const provenance = await resolveWorkflowIrForTaskWithProvenance(store, "t1");
expect(provenance.source).toBe("default");
```
This is **stronger** than the identity check it replaces:
| mutation | result |
|---|---|
| fallback no longer branded | **fails** — `expected 'selection' to be
'default'` |
| fallback returns a different IR entirely | **fails** structurally |
The first is the case the old `toBe` could not articulate: an unbranded
fallback still equals the constant structurally, so a caller asking
*"was this actually resolved?"* would get **yes for a guess** —
precisely the lie #2815 exists to prevent.
Core **4792 passed / 0 failed** (was 2 failed) · gate **732 green** ·
lint clean. Test-only; the resolver is restored clean after the
mutations.
## How it was found
Pre-flighting **#2822, #2819, #2823 and #2818** merged-with-main. All
four reported the *same two* failures — the signature of an inherited
red rather than four independent regressions. Confirmed directly on
`origin/main`. Each of those four is otherwise green (engine 11003
passed on all of them); I have noted that on the PRs.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
2c24966d0c |
fleet: the app-side remainder 18 → 0 — Archive/Revert and diff stats were silently absent on a renamed board (#2731)
Three **genuinely free** clusters in one layer and one idiom —
`Column.tsx` (7), `ListView.tsx` (6), `useTaskDiffStats.ts` (5). I built
the claimed-file set from every open PR's diff before starting, having
duplicated a claimed cluster last round.
## Census
| file | before | after |
|---|---:|---:|
| `Column.tsx` | 7 | **0** |
| `ListView.tsx` | 6 | **0** |
| `useTaskDiffStats.ts` | 5 | **0** |
**16 converted; 2 reclassified with a reason** — the two are accounted
for separately below so the numbers stay honest.
## Three silent failures, not three style nits
- **`ListView` Archive and Revert** were gated on `task.column ===
"done"` / `=== "archived"`, so on a board with renamed terminal lanes
**they did not render at all**. No error, no log — the operator simply
cannot archive or revert from the list.
- **`useTaskDiffStats`** compared a bare `column: string` to
`done`/`in-progress`/`in-review`, so on a renamed board it **fetched
nothing** and the row showed no changes.
- **`ListView` progress display** had the same shape for the WIP lane.
## The `?? {}` is the whole subtlety
Every `Column.tsx` site was `workflowMode ? <trait> : column ===
"<id>"`. One adapter now feeds the shared helpers:
```ts
const columnRoleFlags = workflowMode ? (columnFlags ?? {}) : undefined;
```
`workflowMode` means **traits are the only authority**, so a
workflow-mode column with no resolved flags must answer `false` — which
`Boolean(columnFlags?.archived)` did. Passing `undefined` to a role
helper instead selects its **legacy id fallback**, so a flagless
workflow-mode column would start matching on its id. An empty object
keeps the helper on its trait branch. Legacy mode passes `undefined`
deliberately: there the id fallback *is* the answer, and routing it
through the helpers is the point.
## Two things I deliberately did not do
**`isTodoLikeColumn` keeps its own trait arm.** Adopting
`isPreImplementationColumnRole` would widen its fallback from `todo`
alone to `{todo, triage}`, handing a legacy `triage` column a bulk
replan affordance it does not have today — a behaviour change hiding
inside a de-duplication. Only its *fallback* is routed through a helper.
**The `mode === "done"` pair is reclassified, not converted.** It is the
hook's own `"done" | "active"` discriminant, assigned three lines from
`shouldFetchDoneTask` — not a column id, with no trait to resolve. The
census counts it because the receiver is compared to the string `done`,
which is a classifier limit. Marked deliberate and **recorded in
`deliberateByFile`**, so that file's `byFile` drop is 5 while its
conversion count is 3.
One genuine simplification fell out: `workflowMode ? isReviewColumn :
column === "in-review"`, where `isReviewColumn` is *itself* that same
ternary. Both arms already agreed with it — collapsing is
behaviour-identical.
## Revert proof
Restoring the id comparisons on the ListView row menu fails the new
renamed-lane case with `Unable to find an accessible element with the
role "menuitem" and name "Archive"`.
Driven through the **real `fetchBoardWorkflows` seam** with a renamed
vocabulary — payload → `listColumns` → `columnFlagsById` → row menu —
rather than by injecting flags, so the assertion covers the path the
component actually uses. The DEFAULT-vocabulary path passes either way,
which is exactly why the renamed case has to exist.
## Verification
`pnpm test:gate` **GREEN** (158 + 10 + 487 + 71) · **375 passed** across
Column / ListView / useTaskDiffStats / role-invariance / columnRoles ·
dashboard `tsc -p tsconfig.app.json` clean · `pnpm lint` clean · census
`--strict` exits 0.
`TaskCard.tsx` is touched only to pass the new optional `columnFlags`
through; its own census count is unchanged at 3. The 2 `TaskCard` reds
in that suite are the known pre-existing CSS-var geometry assertions.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved workflow lane handling when columns are renamed or assigned
roles through workflow settings.
* Archive and Revert actions now remain available for completed and
archived tasks in renamed lanes.
* Corrected task progress and diff-stat behavior across active, review,
completed, and archived lanes.
* Updated bulk actions, sorting controls, and auto-merge controls to
respond consistently to workflow roles.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7bf3df9477 |
chore(core): delete the dead merge-blocker guard, and catch exemptions for deleted seams (#2830)
Closes the last of the five findings I reported to batch-core (#2783), which merged without them. This one I held back twice on purpose; the reason is now resolved. ## The dead code `evaluateMergeBlockerGuard` had **exactly one reference in the repo: its own declaration.** Never called, never re-exported from `index.ts`/`index.gate.ts`, never registered as a trait hook. Its `lifecycleColumns` conversion was applied to dead code, and the census counted it as progress. ## Why I would not delete it earlier, and what changed I twice declined this, because no production `"guard"` trait hook is registered anywhere in the codebase — the only `registerTraitHookImpl(..., "guard", ...)` is in a test. If a registration had been dropped, that would be a real product bug and this function would be its evidence, so deleting it would have destroyed the breadcrumb. **It is not missing.** Merge blocking is enforced inline in `task-store/moves.ts` (~645 and ~821) via `getTaskMergeBlocker`, gated on the **resolved trait flags** (`toFacts.flags.complete` + `fromFacts.flags.mergeBlocker`) rather than on column ids — a better implementation than the one being deleted. The logic moved; this function, and a file-header line crediting a never-existent `evaluateDefaultWorkflowGuards` reader, were left behind. The header now records where the guard actually lives, so the next person does not repeat the investigation I just did. Also removed: `GuardVerdict` (its return type, used nowhere else) and the orphaned `getTaskMergeBlocker` import — the latter caught by lint, not by me. ## The gap this exposed The allow-list staleness check only fired for a seam the scan still **finds** — it asks *"is this site supplied now?"*. Delete the declaration and the name is never iterated, so its entry sits in the list forever, exempting nothing and misreporting what is tolerated. I found this by deleting the function and watching the gate stay **silent** about its leftover entry. **Measured:** an `ALLOWED` entry naming a non-existent seam now fails with `no such seam declared any more; remove its ALLOWED entry`; removing the probe returns exit 0. The failure header is reworded, since "the sites are supplied now" no longer covers both reasons. This is the fourth blind spot closed in this check, and like the others it was found by exercising the gate rather than reading it. ## Verification `pnpm test:gate` green · `tsc -p packages/core` 0 · lint 0 · gate now reports **20** seams (was 21 — the drop is this deletion). No changeset: deleting unreachable internal code with no exported surface is behaviour-preserving. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3603da731f |
fix(core): duplicate markers were never cleared on a board with renamed terminal columns (#2823)
Third unowned finding picked up after batch-core (#2783) merged without addressing its reports. Same class as #2819, opposite failure direction. ## The defect `clearNearDuplicateReferencesTo` runs on every complete/archive/delete of a canonical task and clears the `nearDuplicateOf` markers pointing at it. It first asks `isNearDuplicateCanonicalInactive` whether the canonical really is finished — a safety check, so a live canonical's markers are not cleared out from under an operator. That call omitted the canonical's resolved column flags, so it fell back to the legacy `done`/`archived` ids. On a renamed board a just-completed canonical (`shipped`, `filed`) read as **still active**, the guard early-returned, and the markers were **never cleared**. The flagged duplicates stayed parked behind a user decision that could never arrive — the exact stranding the predicate's own FNXC note says it was written to prevent. ## I had the direction backwards, and it matters My first report of this seam described it as *markers cleared against a live canonical*. That is wrong. The legacy fallback errs toward "still active", so the failure is the opposite: markers that never clear at all. Same seam, same missing argument, entirely different symptom to look for — which is why the direction is worth pinning in a test rather than reasoning about. ## Measured With the fix reverted, exactly one case flips: ``` ✓ default vocabulary: completing the canonical clears the duplicate's marker × renamed vocabulary: completing the canonical clears the duplicate's marker ✓ renamed vocabulary: a canonical still in the WIP lane does NOT clear the marker ✓ default vocabulary: a canonical still in the WIP lane does NOT clear the marker ✓ a soft-deleted canonical clears the marker under a renamed board Tests 1 failed | 4 passed (5) ``` With it: `Tests 5 passed (5)`. ## Both negatives included A canonical still in the WIP lane must **not** clear its duplicates' markers, under each vocabulary. Resolving the real flags must not degrade into "every column is terminal", which would clear markers out from under an operator who has not made the duplicate decision yet. The soft-deleted path is covered too, since that branch never consults column flags at all. ## A fixture trap worth keeping `sourceMetadata` is `jsonb` and must be seeded as an **object**. Seeding a stringified value reads back fine through `getTask` — it parses either shape — while the production query's `source_metadata->>'nearDuplicateOf'` matches nothing. The fixture looks correctly seeded and the code under test can never find the row. My first version had this, and the self-check in the seed helper is what caught it. ## Scope Five of this predicate's six production call sites already resolved flags. This was the sixth, and the one that runs on every archive/complete transition. Not touched here: the same function's SQL predicate excludes duplicates by literal `ne(column, "archived")` / `ne(column, "done")`. That is a per-row question across many rows in one statement, not a one-line supply, so it is flagged rather than guessed at. The five **engine** call sites of this predicate are separately reported on #2785 and surfaced only via #2822. ## Verification `pnpm test:gate` green, new suite 5/5, `tsc -p packages/core` 0, lint 0, changeset included. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a6af3188df |
fix(core): startup recovery deadlocked on its own per-task lock (#2809)
## The bug `recoverStaleTransitionPendingImpl` runs its whole per-task body inside `store.withTaskLock(id, …)`. On the PostgreSQL arm it then read the task with `store.getTask(id)` — and `getTaskImpl` opens with `store.withTaskLock(id, …)` too. **The per-task lock is non-reentrant.** This codebase states that invariant in prose in two other files: > "nesting inside `withTaskLock` would deadlock since the lock is non-reentrant" — `branch-and-pr-entities.ts:561` > "because the per-task lock is non-reentrant" — `workflow-ops.ts:464` So the sweep waited forever on a lock its own frame was holding. **PostgreSQL-only — which is every production install.** The SQLite arm on the very next line reads through `readTaskFromDb`, a lock-free row read. The backend-mode port swapped only the PostgreSQL arm to `getTask`. The fix restores a lock-free read (`readTaskRow`) on that arm; nothing else changes. ## Why it survived until now The branch is entered **only** when a stale marker names a plugin hook the trait registry still knows (`hasSurvivingPluginHook`). Three nearby cases all miss it: | marker | path | |---|---| | none | the row is never scanned | | only `default-workflow:postCommit` | `hasSurvivingPluginHook` false — marker just cleared | | names an **uninstalled** plugin hook | reconciled away as degraded; nothing survives to re-run | | names a **registered** plugin hook | **reaches the in-lock read → deadlock** | Those first three are what the existing tests cover. The fourth is precisely the state a crash mid-hook leaves behind. All four are asserted in the new suite so the path cannot be re-narrowed and called covered. ## Impact This sweep runs at **startup**. A task left with such a marker deadlocks startup recovery — and because it deadlocks *while holding the task's lock*, that task is also left permanently unlockable. ## How it was found, including a correction By **bisection**, not by reading. An earlier attempt of mine to drive this recovery reported that "the sweep never returns". That was wrong in a way worth recording: the sweep returns fine in three of the four cases, and generalising the one hang to the whole function is what hid the actual trigger across several sessions. Narrowing case by case — empty store, plain task, default-only marker, unknown-hook marker, registered-hook marker — put the fault on one line. ## Verification - **Mutation-verified against the real defect.** With the fix reverted, the regression case fails by name — `recoverStaleTransitionPendingImpl did not settle within 8000ms — deadlock` — while the other three stay green. That is the actual pre-fix behaviour, not a simulation of it. - Every case is **timeboxed** on purpose: a deadlock otherwise surfaces as a suite-level timeout naming no case, which is useless for locating the fault. The deadline is not a flake knob — the fixed code settles in ~150 ms and the broken code never settles, so there is no value in between to tune. - A **vacuity guard** (no markers → scans nothing) so a change that stopped listing marked rows can't leave the other cases green. - `pnpm test:gate` — **exit 0** - full live-PG E2E surface — **152/152** - `pnpm lint` — clean Changeset included (`patch`, category `fix`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
109204c590 |
fix: the query class — three sweeps that never ran on a renamed board (#2818)
Three sweeps that **never ran at all** on a renamed board, plus the shared answer the rest of the class needs. Consolidated from three handoff branches so the helper appears once. #2811 merged, so this is my only open PR. `#2800` measured this class and shipped evidence deliberately without conversions: `listTasks({ column: "<literal>" })` filters in the store, so on a renamed board the read returns an **empty array** and the sweep it feeds does nothing. The census scores the comparison *inside* the loop, never the query above it. ## What was broken | file | census count | what actually happened on a renamed board | |---|---|---| | `backlog-pressure-reporter.ts` | **0** | both reads empty, ratio computed as 0/0 — **the alert never fired**, on a board that may be under exactly the pressure it reports | | `stale-task-reporter.ts` | **0** | both reads empty — **no stale-task signal ever raised**, where work is most likely sitting unnoticed | | `restart-recovery-coordinator.ts` | flagged | sweep never ran — **an engine restart left interrupted tasks stuck with no requeue** | Two of the three have a census count of **zero**. They contain no lifecycle comparison at all, so they have never appeared in the backlog, in a per-file list, or in any "N → 0" claim — and were completely inert. **A file at zero is not evidence of anything.** ## The shared answer, and what it is not Every existing resolver answers a **per-task** question. A query has no task in hand, so it needs the project-level one: every column any workflow declares for a role, unioned with the legacy ids so a board mid-rename still finds rows under the old ones. The set is never empty, so a caller cannot accidentally query nothing. The header states what it is **not**: answering a per-card question from the union would mark a card as review because some *other* workflow calls its column review — the flat-set mistake this program has made four times. ## The finding that generalises: the query is rarely the whole defect `stale-task-reporter` **still reported zero after the query was fixed** — `getTaskAgeStalenessSignal` defaults to the legacy pair, so a card the query now returned was refused inside the signal. Converting only the query would have looked like a fix and changed nothing. That is a caveat on #2800's approach, offered as refinement rather than correction: **asserting the query ARGUMENT is right when pinning a known defect** (the outcome is 0 either way) **and insufficient when proving a fix**, because the outcome is the only thing that distinguishes a real conversion from a deeper one. All three conversions here assert outcomes. `restart-recovery` had three layers — query, a redundant re-assertion (deleted; a test pins the `paused` guard it did contribute), and a move destination that was **already** resolved but whose warning comment was stale. A stale warning is its own hazard: it told the next reader a defect existed where none did. ## Verification - helper **8 passed** · three reporter/coordinator suites **29 passed** - `pnpm test:gate` **161 / 13 / 487 / 71** · lint clean · `--strict` exits 0 · four `tsc` targets clean - each conversion revert-proven independently; the failing case is named in each test header ## Two mistakes worth recording **The helper's own test caught a bug in it.** My first draft wrapped the definition loop in one `try`, and `parseWorkflowIr` **validates** rather than parses — one malformed row would have returned legacy-only lanes for *every* workflow, indistinguishable from the bug it exists to fix. Now isolated per definition. **I clobbered the core barrel** by taking `index.ts` wholesale from a handoff branch, dropping two exports `main` had added since; three packages stopped compiling. Taking a file from another branch takes its whole contents, including what is now stale — for a barrel that is nearly always wrong. Re-applied as a single edit on top of `main`. ## Not included `self-healing.ts`'s 49 — actively owned and mid-conversion; an outside refactor there produces conflicting halves of one sweep. `project-engine.ts` (7) and `executor.ts` (2) need their own read of what each sweep does with the rows, which these three are the argument for. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1824c04584 |
fix(core): restore an archived card to the lane it came from (#2832)
## What Fixes the defect #2824 measured. That PR's characterization cases — merged and asserting the wrong-but-real behaviour — are flipped here to the correct lanes, which is what they were written to do. ## The bug ```ts // archive-lifecycle-2.ts const preArchiveColumn = task.preArchiveColumn ?? "todo"; ``` **`preArchiveColumn` has no database column.** It exists on the `Task` type and in the archive snapshot, and nowhere else — so the in-place restore cannot carry it, `store.getTask(id)` reads a live row that never had it, and the literal decided the destination for **every unarchive that has ever run**. | board | what happened | |---|---| | **default** | `todo` is declared, so the resolver returned it. Restores landed in the queue and **looked right**. | | **custom** | `todo` is declared nowhere, so the resolver took its "no usable history" branch and returned the **complete** lane. A card archived mid-implementation came back marked **finished**. | That coincidence is why this survived **three** separate fixes to `resolveUnarchiveTargetColumnImpl` — a `?? "done"` that invented a column, an `isColumn` legacy-enum gate, and the same gate one function over. Every one was correcting how the resolver interprets a value that never arrived. ## The fix is two halves, and either alone does nothing 1. **Capture** — `taskToArchiveEntryImpl` records `task.column` into the snapshot. That is the last place the original is still in hand, since the entry's own `column` is set to `"archived"` on the line above. 2. **Read** — `unarchiveTaskImpl` reads the **snapshot it already loaded**, not the restored row. I shipped half of this first and watched the destination stay wrong, which is how I found that the field has no row to live on. Mutation matrix: | state | result | |---|---| | both halves | **5/5 pass** | | capture only (read reverted) | **3 fail** | | read only (capture reverted) | **3 fail** | I also tried carrying it through `restoreTaskFromArchive`'s row update — that fails to typecheck, which is the proof that no such column exists and the snapshot is the only source. ## Behaviour changes, deliberately - **Custom boards** — a card returns to the lane it was archived from instead of appearing finished. - **Default board** — a card archived from `done` restored to `todo` under the literal and now restores to `done`. Returning finished work to the queue was the fallback showing through, not a rule anyone chose; the resolver's own branches say a card archived from a declared column goes back to it. ## One expectation of mine was wrong, and the resolver was right I expected a card archived from the review lane to return to **hold**. It returns to the review lane, and that is correct: `.review` is derived from the `mergeOrchestration` flag, **not** from `human-review`. The fixture's review column declares `human-review` + `merge-blocker` only, so it is not a `.review` lane to the resolver — just a declared column with usable history. The case now asserts that with the reasoning attached, so the next reader does not "fix" it back to hold. ## Verification - unarchive suite — **5/5**, mutation matrix above - `pnpm test:gate` — **exit 0** - full live-PG E2E surface — **164/164** - `pnpm lint` — clean Changeset included (`patch`, category `fix`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9f61180dfc |
test(dashboard): restore 8 of 10 red App tests — the fixture went stale at the U12 flag cutover (#2833)
Fixes most of #2829. `App.test.tsx` has been red on `main` since 2026-07-25 — deterministically, on every commit since, identically with and without any batch branch applied. ## Root cause `ListView` early-returns a workflow skeleton when the board has no workflows: ```ts if (boardWorkflows === null || boardWorkflows.workflows.length === 0) { return renderListWorkflowSkeleton(boardWorkflows !== null); } ``` That skeleton carries the **same `list-view` class** as the real body. So every test that waits on `.list-view` and then reaches for a control inside it passed its `waitFor` and failed on the control — presenting a DOM that looks like a perfectly healthy list. That is why this read as "the board renders nothing" and then as an i18n problem; both were wrong. The probe that settled it: ``` cluster: false | listview: true | buttons: [] ``` The fixture returned `workflows: []` with `flagEnabled: false`. That **used to be correct** — the flag-off path rendered legacy columns and needed no workflow. **U12 deleted that path**, so an empty `workflows` array now means "this board has no lanes". The fixture went stale, not the product. The mock now mirrors the builtin coding lanes with the trait flags the board actually reads. ## Measured | | before | after | |---|---|---| | App.test.tsx | 10 failed / 131 passed | **2 failed / 139 passed** (141) | `tsc` 0 · lint 0 · rest of the `dashboard-app-quality-app` project unaffected. ## Two remain — different root cause, deliberately not chased here - **`opens the NewTaskModal from the list view new-task button`** now gets *past* the button (that was the skeleton bug) and fails on `role="heading"` name `"New Task"`. The heading exists as `<h3>{t("newTaskModal.title", "New Task")}</h3>`, so the modal is not rendering at all — plausibly FN-8620's FloatingWindow rework, **unverified**. - **`keeps the 'subtask' background-session route unchanged`** — uncharacterised. Two of my earlier root-cause guesses on this file were wrong, so I am not offering a third. #2829 stays open for these two with the evidence attached. ## Not quarantined `test-quarantine.json` is for flakes. This failed deterministically, and quarantining would have started a 14-day deletion clock on 141 tests covering deep-link handling, view switching, board branch filters, and the FN-5817 mobile auto-merge shell — while hiding a fixture that had silently stopped exercising the list view at all. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
654d28c104 |
test(engine): live-PG coverage for the timing role — and it is correct (#2828)
## What One new live-PostgreSQL E2E suite, 2 tests. **No production file is touched** — evidence, per the E2E worker's remit. `packages/engine/src/__tests__/workflow-timing-trait-live-e2e.pg.test.ts` ## This one is a clean bill of health Coverage for the last lifecycle role that had none — and the seam turns out to be **correct**. Recorded deliberately: a working conversion with no test is one refactor away from a silent regression, and this program should be able to say which seams are right, not only which are broken. `cumulativeActiveMs` accrues while a card sits in the WIP lane, and both halves of the segment boundary resolve that lane by role: ```ts const isWip = (column) => inRole(column, ctx.lifecycleColumnSets?.wip, ctx.lifecycleColumns?.wip, "in-progress"); ``` Note the literal at the end. That is the **same optional-parameter shape** this series found broken at four other seams — `shouldHoldActiveFileScopeLease` (#2795), `evaluateParkedAgentTaskLink` (#2798), `resolvePlanningContinuationCandidate` (#2799), `hasAutoHealableVerificationBufferFailure` (#2802) — where a caller failed to supply the resolved answer and the literal silently took over. Here the move path **does** supply it. Measured on a live store: ``` default: after exit cumulativeActiveMs=84 renamed: after exit cumulativeActiveMs=75 <- accrues on `building`, not just `in-progress` ``` ## What it guards If the move path ever stops populating the resolved columns, the literal takes over and a renamed board's cards accrue **exactly zero** active time — silently, because zero is a legitimate value for a card that has not run. The operator sees no error, just wrong numbers on every custom board. **Mutation-verified against precisely that regression:** blinding the resolved answer (`inRole(column, undefined, undefined, "in-progress")`) fails the renamed case and leaves the control green. This is a guard with demonstrated power, not a test that happens to pass. ## Both boundaries are asserted `cumulativeActiveMs` is `0` while the card is still in WIP and only accrues when it leaves; `firstExecutionAt` is stamped on entry. A test that checked only the final number would pass against an implementation that accrued at the wrong boundary, so both are pinned. ## Verification - new suite — **2/2 passed**, mutation-verified - full live-PG E2E surface — **161/161** - `pnpm lint` — clean Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is reachable, so the merge gate is unaffected. Throwaway per-file database; never port 4040. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d56c635c77 |
test(dashboard): cover the three lane resolvers nobody was testing (#2826)
## Why this exists #2821's review found a bug that lived entirely in a lane **builder** while every test drove the **guard** that consumed it. That is a structural blind spot, not a one-off: injecting a resolved value into a synchronous guard makes the guard testable and the resolver invisible. So I audited every lane helper I added this session. Three had **no direct coverage at all** — `archivedColumnsForTask`, `wipColumnsForTask`, `preWipColumnsForTask`. Their callers were tested; the functions were not. ## They shared the defect that review named Each read `resolved.length > 0 ? resolved : legacyId`, which conflates two different boards: - a **v1 upgrade** — `synthesizeDefaultColumns` emits `traits: []` on every column, so the legacy id is the only vocabulary that exists, and falling back is correct; - a **v2 board that expresses traits** and declares no lane of that role — where the legacy id names a column the board may still *have* and deliberately did not give the role. Falling back there widens the guard onto a role the board explicitly withheld. `declaresAnyLifecycleTrait` separates them, matching the shape #2821's review established for `resolveNodeOverrideLanes`. ## The fixture trap, which is the part worth reading **My first fixture could not see the bug.** It traited the role under test — and where the role *is* traited, the two shapes agree: both return the traited lane. Mutating a helper back to the old shape left all 15 cases green. The shapes diverge only when the resolved set is **empty while traits are expressed**. Each helper now has that case explicitly, with a fixture that traits something *other* than the role under test. **Mutation-verified per helper:** all three reverted independently now fail. Before the extra case, none did. This is the second time this session a fixture built with the production path normalised away the very thing under test. Worth stating as a rule: a renamed-lane fixture proves the resolver reads traits; only a *traits-expressed-but-role-absent* fixture proves what it does when the answer is legitimately nothing. ## Verification - `task-lifecycle-lanes.test.ts` → 18 passed (was 15, none covering these three) - consumer suites (`github-issue-comment`, `planning-board-tools`, `register-git-github.review-lanes`) → 64 passed together - `pnpm test:gate` → 161 + 487 + 13 + 71 - `--strict` → 0; `tsc --noEmit` and `pnpm lint` → 0 errors Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e5c9ea3870 |
fix(core): resolve the task's own terminal node in the node-override guard (#2812)
## What
Fixes the defect **#2793 measured** but deliberately did not fix. That
PR's two characterization tests are flipped here to assert the correct
behaviour — which is what they were written to do.
## The bug
`updateTask({ nodeId })` passes through `validateNodeOverrideChange`
**twice**, and both calls were wrong in different ways:
| call | how it answered "is this terminal?" |
|---|---|
| `branch-and-pr-entities.ts:568` | `resolveTaskWorkflowIrSync` — the
**default** workflow for every task under PostgreSQL |
| `task-update.ts:53` | no options at all → `defaultIsTerminalNodeId`,
the bare literal `nodeId === "end"` |
On a board whose terminal node is not named `end`, the FN-7641 guard
inverted in both directions:
- an override to a **non-terminal** node that happens to be named `end`
was **rejected** with a merge-proof error about finalizing a card the
operator was not finalizing;
- an override to the board's **real** terminal node was **written
verbatim**, no error, card unadvanced — the silent no-op FN-7641 exists
to prevent.
## The fix
A new `isTaskTerminalNodeIdAsync` resolves the task's own workflow, with
the **identical** literal fail-soft for an unresolvable one. Both call
sites use it — pre-resolved, because `validateNodeOverrideChange`'s
callback is synchronous and it asks the question at most once.
**Nothing forced the sync call at either site**: both frames are already
`async` and already awaiting. That is the same finding as #2809's
review, one file over.
The sync helper is **deleted, not kept as a fallback**. Keeping both
would re-create the half-conversion this program keeps finding — one
caller resolved, one not, and no way to tell from a call site which it
got. `branch-and-pr-entities.ts` also leaves the sync-resolver call-site
allow-list (ratchet green, 3/3), the **second** of the six allow-listed
sites to close.
## Why both halves were needed — and how that is proven
#2793's mutation matrix showed the rejected-`end` case is
**over-determined**: both guards independently called it terminal, so
correcting either one alone changed nothing an operator could see. That
is why fixing only the allow-listed sync site would have looked like
progress and delivered none.
Re-measured here, on the fixed tree:
| state | result |
|---|---|
| both guards fixed | **3/3 pass** |
| inner guard reverted to no-options | **1 fails** |
| outer guard reverted to the literal | **1 fails** |
## Tests
#2793's two cases now assert the fixed behaviour and keep their
reasoning:
- the non-terminal `end` override is **written**, and the card stays in
review — a routing change, not a finalize;
- the real terminal `finish` override is **refused** without merge
proof, **and** the field is not written on the way to refusing.
The fixture-integrity case (`finish` is the end node, `end` is not) is
unchanged — it is what stops both assertions passing for the wrong
reason.
## Verification
- terminal-node suite — **3/3**, both-halves matrix above
- sync-resolver call-site allow-list ratchet — **3/3**
- `pnpm test:gate` — **exit 0**
- full live-PG E2E surface — **151/151**
- `pnpm lint` — clean
Changeset included (`patch`, category `fix`).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7c30b54255 |
test(engine): live-PG evidence that restore lands every custom-board card in DONE (#2824)
## What One new live-PostgreSQL E2E suite, 5 tests. **No production file is touched** — evidence, per the E2E worker's remit. `packages/engine/src/__tests__/workflow-unarchive-target-live-e2e.pg.test.ts` ## The finding I went looking for coverage of the one lifecycle role that had none — `archived` — in the function with the worst track record in this program. `resolveUnarchiveTargetColumnImpl`'s own comments record **three** separate defects fixed on those few lines: a `?? "done"` that invented an undeclared column, an `isColumn` legacy-enum gate that rejected every renamed id, and the same gate one function over that dropped a renamed board's stored history on read. All three were reasoned from source. **None had a live-store test.** With one, the path is still broken — and the cause is upstream of everything those fixes touched: ```ts // archive-lifecycle-2.ts:441 const preArchiveColumn = task.preArchiveColumn ?? "todo"; ``` **`preArchiveColumn` is never written.** Across `packages/core`, every occurrence *reads* it or copies it through — into the archive entry, back out of it, through serialization. Nothing ever sets it from `task.column` when a card is archived. Measured on both boards: ``` default: after archive column=archived preArchiveColumn=undefined -> resolver target=todo renamed: after archive column=archived preArchiveColumn=undefined -> resolver target=shipped ``` So the fallback fires for **every restore that has ever happened**, and the two boards diverge on what `"todo"` means to each: | board | what happens | |---|---| | **default** | `todo` is declared, so the resolver returns it. Every restore lands in the queue — right for a card archived mid-implementation, wrong for one archived from `done`, and it **looks** right. | | **renamed** | `todo` is declared nowhere, so the resolver takes its "no usable history" branch and returns the **complete** lane. Archive a card mid-implementation, restore it, and it comes back marked **finished**. | That coincidence on the default board is why this survived three rounds of fixes to the resolver: they were correcting how it interprets a value that never arrives. ## Evidence discipline - **Observed state** — the persisted `column` after a real `archiveTask` + `unarchiveTask` against a real store and real stored workflows. - **Characterization**: the cases assert today's wrong-but-real behaviour so the defect is executable rather than argued, and they are written to flip when the write is added. - **The default-board control** is what shows the coincidence, not just the failure. - **Mutation-verified**: changing the hardcoded fallback from `"todo"` to the renamed hold id fails **all five** cases — every assertion binds to that literal, which is the claim. ## The fix, proposed rather than included One line: persist the card's column when archiving, so `preArchiveColumn` carries real history and the resolver — already correct — can use it. I did not include it because **it also changes default-board behaviour**: a card archived from `done` currently restores to `todo` (the fallback), and with real history it would restore to `done`. That is the better outcome, but it is a visible placement change for every existing install, and that decision belongs to whoever owns archive semantics rather than to an evidence PR. The five cases above are the acceptance criteria for it — the three renamed expectations become the lane the card was actually in. ## Verification - new suite — **5/5 passed**, mutation-verified - full live-PG E2E surface — **159/159** - `pnpm lint` — clean Lane: `.pg.test.ts`, skipped via `pgDescribe` when no PostgreSQL is reachable, so the merge gate is unaffected. Throwaway per-file database; never port 4040. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9a5155c71a |
fix(tests): main red — a census source guard that a COMMENT could invert (#2825)
## Red on main
```
lifecycle-column-census > the baseline can always be re-recorded
> writes the baseline BEFORE the rise check can exit
AssertionError: expected 19345 to be greater than 26374
```
Read literally, that says the CLI now runs its rise check *before* the
`--update-baseline` write — which would break the one command whose
entire job is re-recording, and would be a genuine bug worth stopping
for.
**It does not.** The order in code is correct and unchanged:
| | line |
|---|---|
| `if (updateBaseline) { … writeBaseline() … process.exit(0)` |
`scripts/lifecycle-column-census.mjs:487` |
| `"column-guard count ROSE"` + `process.exit(1)` | `:510` |
## What actually moved was a comment
Line **359** explains this exact failure mode and quotes the marker
verbatim:
> …`"column-guard count ROSE"`, which is the opposite of what happened
and sends the reader looking for…
So `cli.indexOf("column-guard count ROSE")` found the **prose**, 7000
characters before the branch it was meant to locate.
A guard that a comment can invert is not measuring control flow. And the
honest-looking fix — reword the comment — silently re-arms the same trap
for whoever explains this next.
`cliSource()` now strips comments before indexing. The same defence is
already used by `archived-column-gate-parity.test.ts`, for the same
reason: notes documenting *why* a literal is dangerous have to mention
the literal.
## Kept, not deleted
The end-to-end block below these does cover the contract — it drives the
real CLI and asserts exit code, baseline content and printed output, and
its own comment names the ordering bug. It would have been defensible to
delete the two source-text cases as redundant.
I kept them because two guards at different levels is the point: **e2e
proves the behaviour, these locate the branch that provides it.** They
only needed to stop being defeated by prose.
## Evidence
The real ordering bug — make a rise exit before the update branch writes
— fires **all three**:
| guard | failure |
|---|---|
| source order | `expected 9454 to be greater than 9505` |
| slice / uniqueness | `expected 10433 to be -1` |
| end-to-end | `expected 1 to be +0` (exit code) |
**My first mutation attempt was invalid** and I nearly reported it as
evidence: it moved the block by line range, mangled the file, and both
markers disappeared — the resulting `-1`s look like a firing guard but
prove nothing. A mutation that corrupts its target is not evidence that
a guard works.
Engine **10991 passed / 0 failed** · gate **732 green** · lint clean.
Test-only; the CLI is restored clean.
## Note on duplicated effort
#2811 and my #2814 both re-recorded the census baseline for #2783's
rise, concurrently. No harm done — but this file is now a fleet-wide
contention point, and the per-file baseline shape exists precisely to
avoid that. Worth one owner for census/ratchet fixes rather than whoever
notices first.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ed83fd6ec3 |
batch-core: node-override guards let a running task be re-routed on a renamed board (45 → 43) (#2821)
## The defect Two guards in `node-override-guard.ts` answered a **role** question with a **column name**: - **`task.column === "in-progress"`** refuses changing a task's node override mid-flight. On a renamed board it never matched, so an operator could re-route a **running** task — precisely what the guard exists to prevent, and the failure is silent because the guard simply returns `allowed: true`. - **`task.column !== "done"`** gates overriding *to* the terminal node. On a renamed board it never matched either, so the override was refused for exactly the tasks that had legitimately reached the end node. Both fail in the direction that looks like normal behaviour rather than an error. ## Why the lanes are injected rather than resolved in place `validateNodeOverrideChange` is **synchronous by design**, and its existing `isTerminalNodeId` option already establishes the pattern: callers with cheap IR access inject, callers without keep a documented literal fallback. **Both production callers now supply the lanes** — `branch-and-pr-entities.ts:594` (which already injected `isTerminalNodeId`) and `task-update.ts:53`. That was the deciding factor: an optional parameter that only tests fill is the inert-injection shape this program keeps finding, where a guard reads as converted, its test passes because the test injects the value, and production keeps the literal. I checked both call sites had a store in scope *before* adding the option. `resolveNodeOverrideLanes` lives beside the guard rather than in the callers, so the two cannot drift about what "executing" and "completed" mean. ## Fallbacks A workflow expressing **no trait on any column** is a v1 upgrade — `synthesizeDefaultColumns` emits `traits: []` everywhere — not a board without these roles, so it keeps the legacy ids. Same for an unresolvable workflow. Both preserve exactly the behaviour the literals already had. ## Verification - **Mutation-verified per guard:** restoring `task.column === "in-progress"` fails a case; restoring `task.column !== "done"` fails a different one. - The suite also pins the paired negative — resolving lanes must not turn the guard into a blanket refusal for a task outside every WIP lane. - `node-override-guard.test.ts` → 27 passed - `pnpm test:gate` → 161 + 487 + 13 + 71 - `--strict` → exit 0; `tsc --noEmit` and `pnpm lint` → 0 errors Census: batch-core scope **45 → 43**; repo total **255**. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Node overrides now correctly recognize workflow-defined in-progress and completed lanes, including renamed columns. - Override validation falls back safely for legacy or unresolved workflows. - Prevented validation from using stale task-column information during updates. - **Tests** - Added coverage for workflow lane resolution, legacy fallbacks, renamed lanes, and override eligibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8c79da364c |
fix(engine): FN-8356's duplicate-marker cleanup was inert on a renamed board (5 call sites) (#2827)
Fourth unowned finding picked up after both batch PRs (#2783 core, #2785 engine) merged without addressing their reports. Completes the `isNearDuplicateCanonicalInactive` seam alongside #2823, which fixed the sixth (core) site. ## The defect All five engine call sites — `self-healing.ts` ×2, `triage.ts` ×3 — called the predicate without the canonical's resolved column flags, so it fell back to the legacy `done`/`archived` ids. A canonical resting in a renamed complete column (`shipped`) read as **still active**, so every *"the canonical is inactive, so clear the marker"* branch failed to fire. The user-visible result is precisely the stranding FN-8356 was written to remove: a card keeps its **"Needs your decision"** duplicate badge pointing at work that shipped days ago, and no decision can resolve it — the detail banner deliberately offers none for an inactive canonical. ## Measured With the fix reverted, exactly one case flips: ``` ✓ clears the FN-8353-shaped hidden decision for every inactive canonical state × renamed vocabulary: clears the decision for a canonical resting in a RENAMED complete column ✓ renamed vocabulary: leaves the decision alone while the canonical is still in the WIP lane ✓ leaves active canonical decisions, user pauses, unrelated reasons, non-marker sources untouched Tests 1 failed | 3 passed (4) ``` With it: `Tests 4 passed (4)`. ## Wiring is proven separately from behaviour A behaviour test on one call site says nothing about the other four — that is the failure this whole lane keeps re-finding, so I did not rely on it. With #2822's barrel-import fix applied locally, the seam gate's staleness check **fails both engine allow-list entries as "now supplied"** — it can no longer find an omitting call site in either file. That is the proof for all five. Consequence worth flagging: **when the second of #2822 / this PR lands, the two engine entries in #2822 must be deleted.** CI fails until they are, by design — the exemption cannot outlive its fix. ## Both negatives included A canonical still in the WIP lane must **not** have its decision cleared, under each vocabulary. Resolving real flags must not degrade into "every column is terminal", which would dismiss a duplicate decision the operator has not made yet. ## Design note The helper is module-private in each file rather than shared. `findColumn` is already duplicated exactly this way in `hold-release.ts`, `merge-trait.ts`, and `workflow-capacity.ts` — following the established shape beat adding a cross-module abstraction for a bug fix. ## Verification `pnpm test:gate` green · 27 triage suites / 387 tests green · engine `tsc` 0 · lint 0 · changeset included. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bcc8bf17c1 |
fix(core): brand the unregistered built-in workflow fallback (#2815)
> **Corrected after opening.** The first version of this PR claimed the
id cross-check reported `"default"` for *every* authored workflow and
stranded cards in triage recovery forever. That was wrong, and the
correction is below in full. The reachable defect is the branding hole;
removing the id check is correctness, not a live bug fix.
## The bug (reachable)
`resolveWorkflowIrById` has four ways to substitute the default coding
IR. Three brand the result via `markFellBack`. The fourth did not:
```ts
if (isBuiltinWorkflowId(workflowId)) {
const builtin = getBuiltinWorkflow(workflowId);
const ir = builtin?.ir ?? defaultCodingWorkflowIr(); // <- unmarked substitution
```
An id that *looks* built-in but is not registered — a workflow removed
between releases, a typo'd selection — lands there and silently gets the
default coding IR. `resolveWorkflowIrForTaskWithProvenance` then reports
**`source: "selection"`** for it, handing a caller the default board's
graph under the selected workflow's name. That is precisely the lying
signal the API exists to prevent.
Found by review on this PR (thanks — see the thread), not by me.
## The correction to my own claim
I also deleted an id cross-check that ran after the marker check and
reported `"default"` when the resolved IR's `id` differed from the
requested workflow id. I justified that by saying it misfired for every
authored workflow, because `createWorkflowDefinition` stores an IR
verbatim while minting `WF-NNN` separately:
```
store workflow id = WF-001 stored ir.id = custom:prov
PROVENANCE source = default resolved ir.id = custom:prov <- the CORRECT IR, called a guess
```
**That measurement is real but it came from this suite's own fixture.**
Neither `WorkflowIrV1` nor `WorkflowIrV2` declares an `id` field. An
editor-authored workflow carries none, so `resolvedId` is `undefined`
and the check passed it as `"selection"` — the misfire never reached
`triage.ts`'s post-U11 intake recovery, the one production consumer. My
"declined forever, not deferred" claim does not hold.
Removing the check is still right: it is **unreliable** (it interrogates
a property the IR types do not declare, and when one is present it is
the author's id, unrelated to the store-minted row id) and **redundant**
(all four substitutions are now branded, and the marker is checked
first). But it is a cleanup, not a fix — and saying otherwise is how a
narrow change gets backported as a critical one.
## Tests
| case | asserts |
|---|---|
| unregistered **built-in** id | **`source: "default"`** — the reachable
defect |
| missing definition | still `"default"` |
| no selection | still `"default"` |
| IR carrying its own id | `"selection"`, with the id mismatch asserted
explicitly so it can't pass for the wrong reason |
| the resolved IR is the task's own board | contains the renamed hold
column, not a default-board id |
**Mutation-verified:** removing only the branding fails the
unregistered-builtin case and nothing else — which shows it covers that
specific hole rather than overlapping the others. Restoring the id check
fails only the carries-its-own-id case, leaving both fallback cases
green, which shows the deletion did not widen trust.
## Blast radius
`triage.ts:1211` is the only production consumer of the provenance
result across core, engine, dashboard and CLI. Every other `.source ===`
hit is an unrelated field on an unrelated type; the two index hits are
re-exports.
## Verification
- new suite — **5/5**, mutation-verified in both directions
- `pnpm test:gate` — **exit 0**
- `pnpm lint` — clean
Changeset included (`patch`, category `fix`), rewritten to describe the
branding fix rather than the overstated one.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d2f47acedd |
fix(tests): the core half of #2783's bookkeeping — archived-gate inventory (#2817)
## The 6th red from #2783 #2814 cleared the 5 **engine** reds #2783 left on `main`. This is the sixth, in **core** — I found it after #2814 was already open, and it merged before I could fold this in. ``` archived-column-gate-parity > all three encodings of the archived gate stay in lockstep with the audited inventory AssertionError: TypeScript encoding changed. ``` ## Same cause, same shape as #2786 #2783 converted three more sites off raw `column === "archived"` comparisons and did not update the inventory in the same commit — which the guard's own failure text explicitly asks for. | file | before → after | |---|---| | `async-mission-store.ts` | 2 → 0 | | `task-store/symbol-locks.ts` | 1 → 0 | | `task-store/archive-lifecycle-2.ts` | 2 → 1 | **Verified each is a real conversion, not a dropped gate.** All three now seed a legacy lane set and extend it from the workflow: ```ts const lanes = new Set<string>(["done", "archived"]); … for (const id of columnsWithFlag(ir, "archived")) lanes.add(id); ``` The literal still in `archive-lifecycle-2.ts:46` is `column: "archived"` as a **move destination**, not a gate comparison — the same distinction the planner-lane move targets get. ## Verified NOT a split-brain That is the thing this file exists to catch — TypeScript moving to the resolved role while the SQL sides keep comparing the raw string. The **Drizzle and raw-sql inventories are unchanged and both pass**. Worth stating explicitly because those assertions run *after* the TypeScript one: a plain red tells you nothing about them, so they had to be re-run green to know. ## One thing I nearly got wrong My first edit was a whole-file string replace and it threw on an assertion count. That turned out to be load-bearing: **these paths appear in more than one inventory in this file** (`AUDITED_TS_SITES` and the raw-sql inventory both list `async-mission-store.ts`). An unscoped replace would have silently edited the raw-sql inventory too — making the parity guard agree with itself and defeating the exact cross-encoding check it exists for. The edit is now scoped to `AUDITED_TS_SITES` by line range. ## Evidence - Guard still bites: appending a real `task.column === "archived"` to an audited file → **fails**. - Core **4773 passed / 0 failed** · engine **10988 passed / 0 failed** · gate **732 green** · lint clean. Test-only; `agent-store.ts` restored clean after the mutation. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ba40942a10 |
batch-dashboard-app: 75 → 2 across packages/dashboard/app — the last two are deliberate, not missed (#2772)
**Batch branch is live: `batch-dashboard-app`.** Push conversions here as commits rather than opening per-file PRs — that is the CI-run bottleneck this model removes. **One-line ownership note for you to arbitrate:** you have addressed me as U11, U12 and U7 at different points, so the `u12 worker -> batch-dashboard-app` mapping is ambiguous from my side. I claimed it because `dashboard/app` is where I have done the most work this session (TaskContextMenu, Column, TaskCard, TaskDetailModal, columnRoles, taskActivity) and I know which of its guards are load-bearing fallbacks. **If another worker is the intended owner, say so and I will hand the branch over rather than both of us pushing to it** — two workers on one shared branch is exactly what silently discarded a reviewed fix in #2645 today. ## The work order (measured at branch point, tests excluded) **75 guards across 32 files.** Largest: `TaskContextMenu.tsx` 9 · `Column.tsx` 7 · `ListView.tsx` 6 · `TaskDetailModal.tsx` 4 · then a long tail of 3s, 2s and 1s. Full per-file list is in the committed work order so feeders can claim without re-measuring. ## Two rules this surface keeps tripping on **1. A literal after `??`, or in the `else` of a `flags ?` ternary, is a DEGRADED-MODE answer — not an unconverted guard.** Two real states reach it: the **pre-load window** (board renders before the workflows fetch resolves) and a card stranded on an id its workflow no longer declares. In both, `columnFlagsById` has no entry at all. Deleting the fallback does not remove a decision — it substitutes "no role" silently, and affordances vanish during first paint. Those sites reach 0 by **marking**, not deleting. Expect `TaskContextMenu.tsx` and the `utils` files to be **mostly marks**. A "9 → 0" that deleted 9 fallbacks is a regression wearing a green census. **2. A marker excuses ONLY the construct it is attached to** — the statement or function holding the literal, not a sibling declaration. This has cost three passes, two of them mine; my first attempt on `reliability-metrics.ts` scored **1 of 6**. **Verify by the count moving, not by the comment existing.** With the ratchet gate-blocking, a mis-marked batch either wedges the gate or locks the miss into a re-recorded baseline. ## Status Opening commit is the work order only — **0 of 75 converted so far.** I am near the end of my context, so I am establishing the branch and the shared list rather than starting conversions I cannot finish cleanly. Feeders can begin immediately; I will keep the branch rebased. My other PR **#2762** (`live-agent-count.ts` 6 → 0) is green and unconflicted — per your rule it should land rather than fold into a batch, and it is `packages/core` so it belongs to batch-core anyway. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Task UI now resolves workflow “column roles” per task to drive diffs/merge details, routing/steering, progress/runtime visibility, and review badges. * Right-dock/overflow views and dev-server now use per-task column traits for “executing” behavior and dependency-based “Up Next” eligibility. * **Bug Fixes** * Fixed bulk action selection/delete/archive eligibility and prevented cross-workflow role leakage. * Made in-review/stale-paused-review, stuck, and effective executor/validator model logic role-aware. * **Tests** * Added regression coverage for degraded-flag behavior and ensured resolved-flag props aren’t ignored. * Added a static check to fail builds on inert optional flag seams. * **Documentation** * Updated batch work-order and mega-batch branch guidance. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ## Late addition: the seam gate was masking a real offender `scripts/check-inert-flag-seams.mjs` matched call sites by NAME, so two same-named functions in different modules were conflated. I had documented that as a known false-positive source and moved on — reports mentioning `sortTasksForDisplayColumn` are noise, read past them. That annotation was the damage. Core's `sortTasksForDisplayColumn` genuinely never receives its `columnFlags` argument outside its own tests. The dashboard's separate function of the same name (`app/components/taskSorting.ts`), called with up to five arguments from `Lane`/`Board`/`ListView`, was raising the arg-count max and clearing core's seam. The offender was behind a row everyone had been told to skip. The gate now records the module each callee is imported from and matches it against the seam's declaring module. **Measured, by reverting the change:** the scan prints `17 seams, all supplied` and emits **no row** for the function. With the change, it is reported. Both directions watched. Reported on #2783 rather than fixed from outside — core owns it, and "wire the flags" vs "drop the parameter and let the literal stay counted" is their judgment call. TEMPORARY allow-list entry carries it meanwhile; the existing staleness check fails the moment the site becomes supplied, so the entry cannot outlive the fix. Two known limits remain, both inherent to name matching and both documented in the script: the one-supplier floor, and the `__tests__` exclusion (hence the two permanent `ALLOWED` entries). ## And the one-supplier floor, closed the same way I wrote in the section above that the floor "hasn't cost anything yet." That is verbatim the reasoning that kept the imported-shadow bug alive, so I closed it instead of leaving the note. `best < arity` asked only whether SOME caller supplied the argument. One correct call site cleared the seam while every sibling took the legacy fallback — the `isTaskStuck` defect class, where two of three sites omitted the flags and the gate stayed green because the third was right. Review caught that one. A partially-supplied seam is the harder of the two: wholly-unsupplied is uniformly wrong, this works on the board you tested and degrades on the column you did not. **Measured:** dropping the flags argument at `Column.tsx`'s supplied call site produces `supplied by 5/6 call sites; omitted at packages/dashboard/app/components/Column.tsx:1 (of 2)`; restoring returns `all supplied at every call site`. Red and green both watched. Two real omissions found, both on `isNearDuplicateCanonicalInactive`: - **`TaskDetailModal.tsx`** — deliberate, and it **corrects a note I left at that site**. The old note said hoisting the flags state was "the actual fix." It is not, for this call: the flags in scope describe the *modal's* task, and the canonical is a **different task** on a column this component never resolves. Passing them would type-check, read as a conversion, and answer about the wrong task — exactly what `column-role-degraded-flags.test.ts` exists to catch. Supplying it correctly needs a fetch, which is a data change and out of scope. - **`core/task-store/branch-group-ops.ts`** — genuinely wireable (the impl is async and already holds `store` and `canonicalId`). Reported on #2783, not edited from outside. Exemptions for this class are keyed by **call site** (`<file>::<function>`), not by function name. A name-level entry would waive every site of a partially-supplied seam, which is backwards — its other sites are correct and are the reason the omission is worth reporting. Both entries carry the same staleness check as the name-level list and cannot outlive their fix. Remaining known limit, now the only one: the `__tests__` exclusion, which makes a test-only export read as having no callers. That is what the two permanent `ALLOWED` entries are. ## The `__tests__` exclusion, and two allow-list entries built on false reasons Named as the "last remaining limit" above, so it got closed too. The scan now reads test files for call sites — but counts them **separately**, and a test never clears a seam. That direction is the dangerous one: counting test callers as suppliers would have re-hidden core's `sortTasksForDisplayColumn`, whose only suppliers are its own tests. Measured by lifting its exemption: still reported. Both permanent allow-list entries claimed the scanner couldn't see their callers. **Both reasons were false**, and reading tests is what proved it: - **`evaluateMergeBlockerGuard`** — zero callers in tests either. Its only reference in the repo is its own declaration; never registered as a trait hook; the `evaluateDefaultWorkflowGuards` reader its file header credits does not exist. The `lifecycleColumns` conversion went onto dead code, and its note describes a crossing the guard cannot make. Reported on #2783, including the two things I am explicitly *not* concluding (no `"guard"` hook is registered in production; whether that is residue or a dropped registration needs core's intent). - **`isRecoverableMissingWorktreeReviewFailure`** — 5 test call sites. It wraps `...WithProgress`/`...NoProgress`, the live pair called from `self-healing.ts`, both supplying `reviewColumns`. Entry kept, true reason recorded. ### A wrong turn, recorded because it is the failure mode this PR is about I first classified no-production-caller seams as *informational* when they weren't re-exported from a package index, reasoning that a public export might be called externally. That silently downgraded `sortTasksForDisplayColumn` — a confirmed real offender — from failing to a footnote. Publication status has nothing to do with whether there is production behaviour to be wrong. Reverted to the simple rule: no production caller means inert, and it fails. It is worth stating plainly because it is the exact shape of everything else in this PR: a change that made the gate read *cleaner* while making it catch *less*, and it type-checked, passed every test, and would have reviewed fine. ### Where that leaves the check Every blind spot named in this PR has now been closed, and **each one produced a real defect within minutes of closing it** — imported shadows, the one-supplier floor, the `__tests__` exclusion. Four verified findings went to core, one to engine. I would not read the remaining ~240 guards' green gates as evidence that they are clean; I would read them as untested. ## Two guards for one question, one of them worse Having hardened the script, I checked its older twin rather than assuming it was fine. `resolved-flags-seams-have-suppliers.test.ts` carried its own copy of the trailing-flags-parameter check — written before the script existed — with **all three** holes the script has since closed. **Measured on one reintroduced defect** (dropping the flags argument at `Column.tsx`'s supplied `isNearDuplicateCanonicalInactive` call): | | result | |---|---| | `scripts/check-inert-flag-seams.mjs` | `supplied by 5/6 call sites; omitted at .../Column.tsx:1 (of 2)` | | this test's arity half | **3 passed** | Deleted the arity half. Redundancy between a strong and a weak check isn't redundancy — it's a green result available to whoever runs the weak one, and there was no signal at the call site telling you which you were looking at. The **props-shape half stays**: it has no twin in the script, and I confirmed it still fires by reintroducing the original `PrPanel` defect (outer component stops destructuring `taskColumnFlags`) — it reports `PrPanel declares taskColumnFlags but never takes it`. Dashboard app suite: **113 files / 3921 tests** (was 3922 — the deleted case is the difference). ## The gate started catching defects as they landed Syncing with main brought in three fresh conversions from other workers. The hardened check flagged all three immediately — the first time these guards have fired on someone else's landed code rather than on my own. - **`TaskCard`** — `getRunningOptionalGateBadge(task)` omitted flags while *both* `ListView` sites supplied. Fixed, and `taskColumnFlags` added to the `useMemo` deps: no `exhaustive-deps` rule here, so a memo that reads flags without listing them keeps the first-paint `undefined` answer and reproduces the bug through staleness instead of omission. - **`TaskTokenStatsPanel`** — `getTotalAgentActiveMs` omitted while `TaskCard` supplied, so the same runtime number came from the real column on a card and from legacy ids in the detail modal. Now takes `columnFlags`, supplied from `detailColumnFlags` — correct here because the panel renders the modal's **own** task, unlike the near-duplicate canonical above. - **`ListView` ×2** — passed `columnFlagsById.get(task.column)`, the cross-workflow **union**. A task whose own workflow doesn't declare that column gets a *neighbour workflow's* traits. The landed comment justified it as "this list already owns `columnFlagsById`" — exactly the reasoning `column-role-degraded-flags.test.ts` exists to reject. It failed on merge and is how I found this. Also: the `getTotalAgentActiveMs` exemption I was carrying **self-retired**. Main wired the seam, the staleness check failed the entry, and I removed it. That mechanism has now paid for itself once. ### Pre-existing, NOT from this PR: `App.test.tsx` is red on main `app/components/__tests__/App.test.tsx` fails **10 of 141** identically with my changes, with my changes stashed, and with main's own `App.tsx` restored. Not mine, and not in the merge gate. **Bisected on clean `main` checkouts, so this is measured rather than inferred:** | commit | date | result | |---|---|---| | `main~400` (`41d60f0355`) | 2026-07-25 | **140 passed** (140 tests) | | `main~275` (`74d6513fae`) | 2026-07-27 | 3 failed / 141 | | `main~210` (`d2ce1ba8b5`) | 2026-07-29 | 10 failed / 141 | | `main` (`6fc98fd6c7`) | 2026-07-30 | 10 failed / 141 | So it is **not one regression** — it degraded in two stages across 2026-07-25 → 07-29, and the test file itself changed in that window (140 → 141 tests). Three commits touched it there: `73b2a32e2b`, `f26cbedf4f`, `f157bf7460`. That window overlaps the workflow-owned lifecycle migration, which is suggestive but not something I confirmed. The failures are render-level, not assertion-level — `Unable to find an element with the text: + New Task`, `Unable to find role="dialog"`, `Unable to find ... Back nav task`. The board appears to render nothing. That reads like a real regression or a harness mismatch after the lifecycle migration, not a flake, so I have deliberately **not** quarantined it — quarantine is for flakes, and using it here would hide the signal. Flagging for whoever owns `App.tsx`. My suites: `app/__tests__` **113 files / 3921 tests** green, `tsc` 0, lint 0, census `--strict` 0, seam gate 0. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2ccd78abbc |
fix: main is red on the lifecycle ratchet — re-record the census baseline (#2811)
**`main` is RED on the lifecycle ratchet right now.** `node scripts/lifecycle-column-census.mjs --strict` exits **1** on pristine `origin/main`, which is the `Lint` job's *Lifecycle-column ratchet* step — so **every open PR fails Lint** until this lands, regardless of its own contents. Verified on a detached checkout of `origin/main`, not on a branch of mine. ## Cause Eight `DELIBERATE-LITERAL` markers were added across seven files without re-recording the baseline: ``` packages/core/src/task-move-disposer.ts (in-progress, todo) packages/core/src/task-store/archive-lifecycle-2.ts (archived) packages/dashboard/src/github-tracking-comments.ts (done) packages/dashboard/src/gitlab-tracking-comments.ts (in-progress) packages/dashboard/src/server.ts (archived) packages/dashboard/src/task-planner-chat-context.ts (done) packages/dashboard/src/test/mockCoreEngine.ts (in-review) ``` Adding a marker RECLASSIFIES a site (column-guard → deliberate), so the tracked deliberate totals move and `--strict` fails until the baseline records the new shape. It is the same mechanism that turned #2775 red earlier today — a marker landing without its baseline — which is worth noting because it has now happened twice from different PRs. ## The fix Baseline re-recorded, nothing else. Zero source changes; the diff is one derived file. - `--strict` exits **0** - `pnpm test:gate` — **161 / 13 / 487 / 71** - `pnpm lint` clean ## Worth a follow-up by whoever owns the ratchet The failure is structural rather than careless: a PR that adds a marker is *doing the right thing*, and the baseline requirement is only discovered when CI goes red — after merge, for everyone else. Two options, neither of which I am taking unilaterally on a red-main fix: 1. have `--strict` treat a marker-only reclassification as an accepted rise (it is not new debt — the count of unconverted guards goes **down**); 2. or fail the PR that adds the marker, by comparing against the base ref rather than the recorded baseline — the machinery for that already exists in this script. I would take (1): a marker is the documented way to close a site, and requiring a second mechanical step to record it is a trap that catches good behaviour. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved lifecycle census error messages to distinguish genuine increases in column-guard debt from reclassified deliberate literals. * Added clearer remediation guidance for reclassified results, including when to update the baseline. * Updated lifecycle census baseline mappings to reflect current classifications. * **Tests** * Added coverage for unchanged baselines, genuine guard-count increases, and marker-only reclassification scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |