b752f9014dc7afed31f8b478ee01746b3ffe4865
12505 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b752f9014d |
fix(core): a hold-only column vanished from the SDLC funnel (+ un-red the census on main) (#2674)
Two things, both small, one urgent. ## 1. Hold-only columns disappeared from the funnel `hold` was absent from `TRAIT_TO_STAGE`, so a column whose only pre-implementation trait is `hold` — a renamed board's wait-for-capacity lane — resolved to `OTHER` and vanished from the SDLC funnel. Measured before the fix: ``` stageForTraits(["hold"]) === "other" ``` The default lineage hid it: its Planning column also carries `intake` and `reset-on-entry`, so it always matched something. Only a board that names its wait lane separately was affected — **exactly the custom shape this trait mapping exists to support**. Revert check: removing the entry gives `expected 'other' to be 'todo'`. ## What I deliberately did NOT fix, and why The merged default Planning column carries `["intake","hold","reset-on-entry"]`, and `stageForTraits` prefers the earliest stage in flow order — so `intake` wins and it still resolves to `triage`. The `todo` stage therefore stays empty on every default board since U11, and the funnel shows a **phantom 100% drop between Triage and Todo**. That is a real defect. It is also not a reversible call: changing which stage Planning reports would retroactively alter how historical analytics read. Flagged on #2669 for a product decision. Adding `hold` does not touch it — `intake` still outranks — and a second test **pins the current behaviour** so the larger question gets answered deliberately rather than drifted into by a future edit to this map. ## 2. `check:lifecycle-columns` is RED on pristine `origin/main` — again ``` census exit on pristine main = 1 packages/engine/src/executor.ts: allows 87, tree has 85 ``` `executor.ts` is a file this PR does not touch, so a merge lowered the count without re-recording and the blocking PR check is failing for **every open PR**. The re-record is mechanical and is included here to unblock it — called out explicitly because it is unrelated to the funnel fix and should not ride along unexplained. This is the second time the baseline has gone stale on main this way. The rule works (`--strict` caught it immediately); what is missing is that it caught it *after* the merge. Worth considering whether the census should run on the merge queue rather than only on PR head — otherwise a PR that is green when opened can still land a stale baseline. ## Verification `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm check:lifecycle-columns` exits 0 after the re-record. `tsc -p packages/core/tsconfig.json` clean. `pnpm lint` clean. `sdlc-funnel-default-columns.test.ts` 8/8. No changeset: the funnel entry is a correctness fix with no user-facing API change, and the baseline re-record is internal. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved SDLC funnel classification for hold-only columns, placing them in the Todo stage instead of Other. * Preserved correct Planning column behavior when hold-related traits are combined. * **Tests** * Added coverage for hold-related funnel stage mapping and trait ordering. * Updated lifecycle column census baselines to reflect current results. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
fe7e68bc13 |
fix(core): wedge notifications could never be resolved on PostgreSQL (42P18) (#2669)
Not a U12 change — found while **attributing** the pre-existing live-PG
failures during U12's closing verification, and it turned out to be a
product bug rather than a stale test.
## The defect
`resolveWedgeNotification` builds its UPDATE with:
```ts
jsonb_build_object('status', 'resolved', 'transitionedAt', ${transitionedAt})
```
`jsonb_build_object` is variadic `"any"`, so there is no signature for
PostgreSQL to resolve the bind parameter against. It rejects the
statement at **parse time**:
```
42P18: could not determine data type of parameter $1
```
Parse-time is the important part: this failed on **every call**, not on
unusual data. Wedge notifications could not be resolved at all in
PostgreSQL mode.
Casting the parameter to `::text` fixes it.
## Evidence
- `store-wedge-resolution.pg.test.ts` goes **0/7 → 7/7**. That suite has
been red on `main`.
- **Causally verified, not assumed:** removing the cast reproduces
`42P18` exactly. The fix is the cast, not something incidental to the
edit.
- Checked the rest of `packages/core` for the same shape — this is the
only `jsonb_build_object` call site, so there is no second instance
hiding.
## Why it survived
The failure is in a live-PG suite that was already red, so it read as
part of the ambient noise. I only found it because the closing
verification required me to attribute each failing suite to a cause
rather than count them — and "these 4 fail on main too" is an
attribution of *whose*, not of *what*.
Worth flagging for whoever owns the remaining three
(`agent-logs-and-monitor`, `central-archive-secrets`,
`workflow-settings-project-identity`): the same reasoning applies. A
suite failing on main is not evidence that the code is fine.
## Verification
`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0. `tsc -p packages/core/tsconfig.json`
clean. `pnpm lint` clean.
Independent of #2655; either order merges.
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
a099813e94 |
fix(test): decouple the audit-emitter assertion from log formatting (last of the 4 red PG suites) (#2675)
Last of the four long-red live-PG suites. **This one is a stale test — the only one of the four that is.** ## The cause `withSeverityMarker` (`logger.ts:31`) deliberately wraps every message in a machine-readable severity marker so the TUI log pane can colour by level. The emitted string carries a `fnlvl=warn` marker and a `[core-async-secrets-store]` subsystem tag ahead of the real text. The assertion pinned the raw message with `toHaveBeenCalledWith`, so it broke when that convention landed. It was coupled to log **formatting**, not to the behaviour it exists to check. ## The fix Rewritten to assert what it actually cares about: exactly one warning, whose message **contains** the subsystem-tagged text, carrying the underlying cause. Both halves of the behaviour stay pinned — the `resolves.toMatchObject` above proves the secret is still created when the audit emitter fails, and this proves the failure is surfaced rather than swallowed. **Mutation-verified rather than assumed green:** deleting the `severityAuditLog.warn` call in `async-secrets-store.ts` fails with `expected "warn" to be called 1 times, but got 0 times`. A `stringContaining` assertion that passes because it matches nothing would be worse than the brittle one it replaces. ## The four, complete | suite | verdict | |---|---| | `store-wedge-resolution` | **product bug** — `42P18`, total runtime failure of wedge resolution in PG (#2669) | | `workflow-settings-project-identity` | **stale docs** — resolver contradicted its own documented order (#2671) | | `agent-logs-and-monitor` | **real defect** — funnel mis-bucketing from the U11 merge; half fixed in #2674, half needs a product call | | `central-archive-secrets` | **stale test** — this PR | **Three of four were real problems**, sitting behind "pre-existing, fails on main too". That phrase answers *whose* problem it is, not *what* is wrong. ## Census, again `check:lifecycle-columns` is **still** exiting 1 on `origin/main` — `executor.ts: allows 87, tree has 85` — the same staleness flagged on #2674. Re-recorded here too, because the blocking check stays red for every open PR until some PR carries it, and I do not know which of #2674 / this one lands first. ## Verification `pnpm test:gate` green (10 / 158 / 487 / 71). Suite **14/15 → 15/15**. `pnpm check:lifecycle-columns` exits 0 after the re-record. `pnpm lint` clean. No changeset: test-only plus an internal baseline. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e711fbab15 |
The ratchet's baseline could not be re-recorded once a file rose — the one state that blocks a correct conversion (#2668)
Unowned (no open PR touches the census CLI — only its baseline JSON) and **live**, since #2654 gates CI on `--strict`. ## The problem `--update-baseline` sat **behind** the rise exit, so the only supported way to re-record was unavailable in exactly the situation that needs it. That matters because **a conversion legitimately adds a literal.** The correct shape for a caller that may have no traits is `flags ? flags.x : columnId === "legacy"`, and each one raises a file's count by one. Measured on current main: `columnRoles.ts` went **0 → 1** from precisely that shape (added by #2647, documented at the site, correct code). So a worker doing the right thing meets a red gate whose only escape is hand-editing the JSON. That is how a ratchet becomes something people route around rather than run — and then it guards nothing. This is the same failure mode as a guard that cannot fire, arrived at from the other side. ## The change `--update-baseline` is an explicit operator action, so it re-records **unconditionally** and prints what it accepted under `ACCEPTED RISES`. Swallowing a rise silently is the real danger; refusing to let anyone re-record is the same danger one step later, wearing a red check nobody trusts. **The rise check is unchanged** and still exits 1 without the flag. **One writer now.** The old second `writeFileSync` behind the rise exit is deleted rather than left unreachable — two writers for one artifact is how they drift. The `!deliberateTracked && updateBaseline` special case went with it, since the unconditional block covers the legacy-shape migration too. ## Exercised end to end On a real rise injected into `live-agent-count.ts`: ``` rise + plain --strict exit 1 (the ratchet still bites) rise + --strict --update-baseline exit 0 "ACCEPTED RISES live-agent-count.ts: 6 -> 7" ``` Four cases assert the CLI's own source, because exit codes are the contract and the pure summarizer cannot express them: the write precedes the rise check, the branches exit 0 and 1 respectively, accepted rises are **named**, and there is exactly **one** writer. ## A note on the revert proof, because it caught me twice My first attempt to move the block back was a **no-op**: the marker I sliced on (`if (regressions.length > 0) {`) also appears *inside* the update block, so the "revert" reassembled the file unchanged and the suite stayed green. **A revert proof that does not go red can mean the guard is vacuous *or* that the revert did not land** — and the second is easy to miss when you are expecting the first. The real revert fails **2 of 27**, and the assertions now verify marker *uniqueness* before slicing on it. ## Verification - 27/27 census suites; `--strict` exits 0; `pnpm test:gate` **71/71**; `pnpm lint` clean - census on this tree: 748 column guards, **4 triage** (all in `moves.ts`'s flag-OFF block, deletion-scheduled with #2655) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
13bf7e001d |
Closing-bar verification pass on origin/main — one tree, one report (+ the E2E red it found) (#2660)
**Closing-bar item 4, run on one clean tree at `origin/main` (`be63e72f1`).** Nobody was assigned this and my own work is merged, so I took it. > **This PR is now REPORT-ONLY — net zero file changes.** I found the planning-lane E2E red, fixed it, then discovered **#2658 (gsxdsm) makes byte-for-byte the same change** to the same helper and was opened first. I reverted mine rather than leave two identical edits to one function to conflict. **The E2E result below depends on #2658 landing** — on `origin/main` without it, that suite is 2 failed / 5 passed. > > The duplication is worth one note for the fleet: two workers independently hit the same control-card failure and independently traced it to FN-7648's unplanned-seed gate plus a fixture that never wrote a spec. Independent confirmation of the diagnosis, but also ~an hour spent twice — the census-style work order exists to stop exactly that, and E2E fixture defects are not on it. ## Report — all four, one tree | Check | Result | |---|---| | `pnpm test:gate` | **PASS** (132 + 10 + 487 + 71 tests) | | `pnpm verify:fast` | **PASS** — 13 steps green in 89.9s, boot smoke `GET /api/health 200`, clean shutdown | | E2E families | **13 files / 109 tests PASS** — *after* the fix below; **2 failed** before it | | census | total **787**, triage **10** | ## Two corrections to the bar itself **1. It is not "all-8 E2E" any more — there are 13 families.** The suite grew while the bar was being written: ``` agent-count · agent-link · lease-rebound · lifecycle · merge-family · merge-rebound merge-safeguards · merged-board · planner-lane · planner-lane-resolution planning-lane · rebound-family · stranded-column ``` A verification pass scoped to 8 would have skipped 5 families — including the one that was red. Worth fixing the number in the bar so the final pass globs rather than counts. **2. `DELIBERATE-LITERAL (reviewed)` reads 3, and I chased it — RESOLVED, no gap.** I flagged the drop from an earlier "7" as a possible fleet-safety hole. It is not one. Reconciled against `--json byFile`: | File | markers | counted `deliberate` | counted `column` | |---|---|---|---| | `hold-release.ts` | 2 | **2** | **0** | | `live-agent-count.ts` | 1 | **1** | 6 | | `replan-target.ts` | 2 | 0 | 4 | `deliberate: 3` = hold-release 2 + live-agent-count 1, which is exactly the set of marker-covered **comparisons**. `replan-target.ts`'s two markers sit above `return "triage"` **return-value** literals, not comparisons — the census correctly does not count those as guards at all, so they are neither `deliberate` nor `column`. The earlier "7" was simply a different tree state before conversions landed; I was quoting a stale number. Worth noting the marker matcher is already hardened for the subtle case: `hasDeliberateMarker` walks every **ancestor** rather than the enclosing statement, because the real markers sit above the enclosing *function* while the comparison is a `return` inside it — a statement-only lookup "silently reclassified three reviewed literals as backlog". That is the guard-cannot-fire pattern, already caught and fixed by whoever wrote the AST version. **Consequence for the fleet: the census's categories are trustworthy as-is.** No pre-launch action needed on this. ## The red it found `workflow-planning-lane-live-e2e.pg.test.ts` — **2 failed / 5 passed**, including its own **control** case: ``` releases an ordinary held card on a default board (the control) → AssertionError: expected [] to include 'FN-OK' ``` `seedHeldTask` never wrote a `PROMPT.md`, so task creation's bootstrap seed stood, and FN-7648's `isUnplannedForExecution` correctly refused to release an unspecified card. **The sweep was right; the fixture was asking it to release a card that had never been specified.** **This is the second instance of the identical defect** — same cause and same fix as `workflow-lifecycle-live-e2e`'s `seedTask` in #2634. This suite was written after that fix and did not inherit it. The graph-entry contract doc already states the rule: *"Scheduler/release test fixtures must model a card that cleared the gate ... A held unreviewed card is the gate working."* Both failures had one cause — the mid-sweep approval-park case was downstream of the control never releasing. **5 → 7 passed**, and cards that are *supposed* to be held still are, held by their own status/marker, which is what those cases assert. Given it has now happened twice, a shared `seedPlannedTask` helper in the E2E fixture module would prevent a third. I did not add one here: it touches suites owned by U7 and U11 mid-consolidation, and this PR should stay the verification pass plus its one finding. ## Bar status after this - **gate / verify:fast / E2E** — green on one tree, with this commit. - **triage → 0** — still **10**, all in U12's `moves.ts` (4) and `register-task-workflow-routes.ts` (1) per file scan; flag resolution in flight. - **ratchet tightened (item 2)** — not done, U12's. - Once triage hits 0 and the ratchet lands, re-running this exact pass is a ~4-minute job and I can produce the final report. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
632d10a9b4 |
fix(engine): a completed-blocked guard was inert on renamed boards — plus one owner for the terminal pair (#2568)
Two commits: a behaviour-preserving extraction, then the behaviour change. ## ⚠️ Stack note worth acting on **#2550 and #2554 both report MERGED, but their content is not on `main`.** They merged into their *base branches*, and the bottom of that stack (**#2544**) is still open. Nothing in this chain has reached `main` yet. Nothing is lost — everything is in `origin/feature/workflow-e2e-merge-rebound`, which is why this PR targets it. But "merged" reads as "landed" and here it doesn't. **Merging #2544 flows the whole chain down.** ## The bug `parkCompletedBlockedTask` opens with *"is this card already finished?"* and answered it with: ```ts if (task.column === "done" || task.column === "archived") return false; ``` On a renamed board neither matches, so **the guard was inert** — and the very next branch (`if (task.column !== "todo")`) would then have **moved a completed card back out of its own terminal column**. A guard that never fires does not fail a test. This one was found by tracing the last ledger site, not by anything going red. ## Why a shared owner, not a local fix `merger-ai`'s `isAlreadyFinalizedColumn` held the **only** copy of the per-role terminal-pair rule — a P1 learned the hard way (PR #2471 review): a per-**set** fallback collapses to one element for a workflow declaring `complete` but no `archived`, silently dropping the archived half of every already-finished check. Executor's guard was the raw literal pair, so **whoever converted it next would have re-made exactly that mistake** — the lesson lived in a comment in another file. Hence `resolveTerminalColumns(ir)` in core: one owner, one place for the rule. ## Evidence, and its limits **Commit 1 (extraction) is proven behaviour-preserving**: `workflow-already-finalized-live-e2e` is unchanged and green through the delegation, and the per-set mutation **still fails** through the shared helper. **Commit 2 (the fix) is unproven at the call site, and I'm labelling it rather than implying otherwise.** `parkCompletedBlockedTask` is private and reached only from inside executor dispatch — I could not drive it end to end. So the shared helper gets its **own** tests, in both partial-role directions, precisely because its other consumer can't vouch for it. The call site is a one-line delegation to a tested function. Weaker evidence than the rest of this unit's work. Saying so, because quietly counting it as proven is the exact failure this unit exists to catch. ## Census 417 → 416. That ratchet (#2557) is a **ceiling**, so it stays green without coordination; lower the pin when convenient. ## Verification - E2E suites 10/10; helper unit tests 5/5 - core + engine `tsc --noEmit` clean - `pnpm test:gate` green (414 + 10 + 71) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --- ## Note on the conflict status (2026-07-31) GitHub reports this PR `CONFLICTING / DIRTY`. **It is not.** Three independent checks: - `git rebase origin/main` on the pushed head reports *"up to date"* and leaves the SHA unchanged — the branch is already on top of main. - `git merge-tree` against the merge base produces **zero** conflict markers. - `origin/main` is unchanged at the commit this was rebased onto. The remote SHA matches the local head, so the push landed. The `mergeable` field is a **stale computation** — it goes stale after a force-push and doesn't always recompute. This branch has now been rebased and force-pushed four times against that cached value. Worth guarding at the source: the auto-retry treats `mergeable` as ground truth, so a stale value generates conflict notices indefinitely. Confirming with a trial rebase or `git merge-tree` before dispatching distinguishes "actually conflicting" from "GitHub hasn't recomputed" — one command, and it ends the loop. Verification on the current head: merge gate green (487 + 158 + 10), engine + core tsc clean, lint clean, 20 tests in the affected suite, zero unresolved threads. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8d84cee11e |
fix(core): the workflow-settings identity resolver contradicted its own docs (2 long-red tests) (#2671)
Second of the four long-red live-PG suites, after #2669. This one is **stale documentation making a stale test look like a code bug** — behaviour is unchanged. ## What was wrong `getWorkflowSettingsProjectIdImpl` documented a three-step resolution order: ``` (a) store.asyncLayer?.projectId — central-registry id (PG) (b) store.db.getProjectIdentity()?.id — legacy SQLite identity (c) store.rootDir — last-resort key ``` The code does (a), then returns `rootDir`. **Step (b) was removed** by `FNXC:SqliteDualPathCleanup 2026-07-26-14:15` — but the doc block kept describing it, and a comment three lines above the return still said *"Only the true legacy (non-backend) path consults the SQLite identity"*, which has been false for every caller since. ## Which side was wrong — settled by construction, not judgement In my triage on #2669 I said I would not guess between "the test is stale" and "the code lost a needed branch", because the two have opposite consequences and the stale comments made the intent unreadable from outside. That was the right call then; it is now answerable: `dbImpl` **throws unconditionally and ignores its store argument** (`task-id-integrity.ts:58`): ```ts export function dbImpl(_store: TaskStore): Database { throw new Error("TaskStore.db: SQLite Database is not available in backend mode …"); } ``` There is no mode in which `store.db` yields a usable SQLite handle. Step (b) is unreachable **by construction**, not merely unused — so the code is right and the documentation was wrong. ## Why the tests passed review originally They build a store double whose `getProjectIdentity()` **returns** a value: ```ts db: { getProjectIdentity() { return { id: "legacy_identity_id" }; } } ``` Production cannot produce that shape. The double made an unreachable branch look testable, which is how the assertion survived the cleanup that deleted the branch. Rewritten to the shipped contract. A neighbouring case that already asserted `rootDir` *when the stub throws* was passing all along — the two forms of the same store disagreed inside one file. ## Verification Suite **7/9 → 9/9**. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm check:lifecycle-columns` exits 0. `tsc -p packages/core/tsconfig.json` clean. `pnpm lint` clean. No changeset: no behaviour change, and no user-visible effect. ## Remaining from the four - ✅ `store-wedge-resolution` — real product bug, fixed in #2669 - ✅ `workflow-settings-project-identity` — this PR - ⬜ `agent-logs-and-monitor` — `expected +0 to be 2` on an aggregation - ⬜ `central-archive-secrets` — an assertion on `warn` arguments Two of four were real problems hiding behind "pre-existing". The other two are still unruled-out. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a6138abeff |
U12: DELIBERATE-LITERAL counts key on file AND column — closing the P1 left on merged #2661 (#2666)
Closes the P1 that was still open when #2661 merged. ## The hole A per-file integer is offset **within a single file**: remove one reviewed `todo` exemption, add an `in-review` one beside it, and the number never moves. The fresh guard is invisible to the column counts too, because deliberate findings are excluded from them — so `--strict` goes green with a new lifecycle-column guard hiding inside an existing marker. Now keyed on **file AND column id**. **Proven with the exact scenario:** swapping a marked `triage` for `done` inside `TaskCard.tsx` leaves the per-file total unchanged and now fails with ``` packages/dashboard/app/components/TaskCard.tsx (DELIBERATE-LITERAL: done): 0 -> 1 ``` ## The pattern worth naming This is the **third** time this instrument has been defeated by an aggregate: | version | defeated by | |---|---| | repo-wide `totals.deliberate` | an addition in file A offset by a removal in file B | | per-file integer | an addition offset by a removal **in the same file** | | per-file per-column | — | Each step narrows what can offset silently, and I walked into the next one twice by fixing the *reported case* rather than the *shape*. Writing it down because the same reflex will produce a fourth if someone adds another aggregate here. **The residual is deliberate, not an oversight:** a same-file **same-column** swap still offsets. Two `todo` exemptions in one file are interchangeable by definition, so there is nothing a reviewer could act on. That is recorded at the site so the next person doesn't rediscover it as a bug. ## Migration, again The key **shape** changed (`file` → `file\0columnId`), which is the same hazard as a missing field: comparing new keys against old reports every existing marker as a fresh rise and pushes people to convert already-reviewed literals. I hit it on the first run here — `TaskCard (DELIBERATE-LITERAL: triage): 0 -> 2` — exactly as I did one shape earlier in #2661. Detected by the delimiter rather than a version field, since old keys have none, and re-seeded on the next `--update-baseline`. 15 file+column entries recorded. ## Verification `pnpm lint` clean. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm check:lifecycle-columns` exits 0. Independent of #2655; either order merges. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1f149d21de |
fix(engine): TAKING spec-staleness.ts + mission-feature-sync.ts — planner lanes (2 triage guards → 0) (#2616)
**Claiming `packages/engine/src/spec-staleness.ts` and `packages/engine/src/mission-feature-sync.ts`.** Deliberately *not* `self-healing.ts` (contended) or `task-creation.ts` (#2589 in flight). ## Guard counts | scope | before | after | |---|---|---| | `spec-staleness.ts` | 1 | **0** | | `mission-feature-sync.ts` | 1 | **0** | | repo-wide `column === / !== "triage"` in `packages/*/src` (excl. tests) | 26 | **24** | **21** once #2612 (comments-ops, 3 guards) also lands. ## What was silently broken **mission-feature-sync** — *"has this task returned to a planner lane?"* decided whether a mission feature drops from `in-progress` back to `triaged`. Keyed on the legacy pair, a card sent back for re-planning on a renamed board left its feature stuck at `in-progress` **forever**: the mission board showed work in flight that nobody was doing, and nothing said so. **spec-staleness** — the preserved-progress skip refuses to fire for an **intake** card, since a card being specified has no progress to protect. Keyed on `triage`, a renamed-board intake card looked like started work and its stale spec was skipped instead of re-planned. ## The union is a deliberate call, and the existing suite forced it My first cut *replaced* the legacy pair with the resolved lanes. That broke a real case: **post-U11 the default lineage has no `triage`**, so a legacy row still resting there stopped counting as a planner lane. `usage-limit-detector` already made this call for the same situation and wrote down why — **over-inclusion is the safe direction**. Marking a feature `triaged` for a card in a legacy planner column is recoverable; a mission board permanently showing phantom work is the bug. So the legacy pair stands and resolved lanes are *added* to it. Worth noting the existing test is what caught this, not review — which is the argument for converting against a real suite rather than in isolation. ## A parameter, and why that needs the ratchet `spec-staleness`'s predicate is **pure** (a task, no store), so the role arrives as a parameter and both callers resolve it. That optionality is exactly the caller-omission hazard this program has already shipped twice, so the function is also registered in core's `role-parameter-caller-audit` (#2588). **This PR's tests prove the parameter is honoured; the audit proves it is passed. Neither alone is enough** — that split is the whole lesson of #2586. ## Mutation-verified | mutation | result | |---|---| | mission-feature-sync → legacy pair only | the two renamed cases fail | | spec-staleness → restore the `triage` literal | its renamed case fails | Negatives included in both: demoting a **WIP** card's feature would report running work as un-started, and never-skipping would discard every card with real progress. ## Verification - new suites 4/4 and 3/3; `mission-feature-sync` + `spec-staleness` 27/27 - engine `tsc --noEmit` clean; `pnpm test:gate` green (482 + 132 + 10) - `executor-prompt` reports 3 failures **both with and without** this change — pre-existing, baselined by stashing rather than assumed 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dca20496f4 |
consolidate/u7: plugins to zero + 8 executor rebound guards + resume lanes (supersedes #2607, #2635, #2640) (#2644)
Consolidation branch for U7, per the new one-branch working mode. **Supersedes #2607, #2635, #2640** — the three of my PRs that were stuck on review threads. My other seven (#2602, #2605, #2606, #2611, #2621, #2628, #2633) are green with **zero unresolved threads** and are deliberately left alone for the merge sweep. ## What is in here, file by file | file | change | guards before → after | |---|---|---| | `plugins/…/glasses/src/agent-actions.ts` | gates, destinations and degraded-resolution refusal all resolve from the task's own workflow | 2 → 0 | | `plugins/…/glasses/src/quick-capture.ts` | accepted capture columns come from the board; default no longer names the deleted column | 1 → 0 | | `plugins/…/glasses/src/settings.ts` | quick-capture default was `triage`, the column #2515 removed | (assignment, uncounted) | | `plugins/…/dependency-graph/src/GraphTaskNode.tsx` | redundant column condition deleted | 1 → 0 | | `packages/engine/src/executor.ts` | 8 rebound guards compare the resolved column; 4 resume-eligibility literals share one resolver | 151 → 143 (+4 off-bar) | | `packages/engine/src/__tests__/` | 4 new suites, 26 cases | — | `plugins/` reaches **zero** column guards with this branch. ## The three threads it closes **#2607 — five findings, all mine, all the same rule.** I kept *qualifying* a legacy-id fallback instead of removing it: | attempt | rule | hole review found | |---|---|---| | 1 | fall back to `todo` when the role is missing | moved cards to phantom columns | | 2 | …only if the workflow **declares** `todo` | aliased **review** lane named `todo` | | 3 | …and only if no other role is assigned to it | **traitless** parking column named `todo` | The qualifications were the mistake. Once `resolveLanes` returns a lane set the workflow *has* a column vocabulary, so "no column carries the hold trait" is a complete answer — refuse. `destination()` is two lines now, with no aliasing surface left to qualify. Plus a sixth, which is a genuinely different state: **degraded resolution is indistinguishable from the default board.** `resolveWorkflowIrForTask` is total by design — a missing definition silently returns the *default* coding IR — so a card on a custom board whose definition could not be read resolved to `todo`/`in-progress`. `undefined` lanes cannot express that (it means "no workflow at all", where the legacy ids *are* the answer). The actions now refuse with 409. #2618 would replace this check with resolver provenance; it is not merged, so this does not depend on it. **#2635 — "seven rebound sites remain untested."** Fair; my "same shape" note was an assertion, not coverage. Seven of the eight need a live graph run to reach, so the *shape* is pinned instead: a static check that no guard in front of a rebound move compares against a column literal, with a vacuity case (the same detection run against the original shape) and a match-count floor (≥8), because a guard reporting success on zero matches is worse than no guard. **#2640 — duplicate workflow resolution.** Framed as I/O; it is also a correctness bug. Eligibility and re-entry are two halves of one decision and resolved the workflow separately, so a workflow edit landing between them has the halves reading *different boards*. Now one caller-owned memo per decision — caller-owned because a process-lifetime cache would have to guess when a mid-flight workflow edit invalidates it. ## Behavioural findings, not tidying - **The last-resort recovery for completed-but-stranded work did not exist off the default lineage.** `promotedFromPlannerColumn` was false on a renamed board, so finished work resting in planning was never promoted; the code fell through to a review handoff that role adjacency rejects, and the card stayed stuck with its work complete. - **Rebound guards could not see the column their own move targeted.** U5b converted the move target; the eight `column !== "todo"` checks in front of it were left literal, so on a renamed board the engine moved a card into the column it was already in — and `moveTaskInternal` runs reset-on-entry on every real move, so at the `preserveProgress: false` site it reset step progress a second time. - **The FN-1404 `task:move` audit row was lying**, recording `to: "todo"` while the move target was resolved. A run-audit trail that disagrees with the move it describes is worse than none. Not a comparison, so no census counts it. - **A task interrupted by an engine pause never resumed on a renamed board** (off-bar, `in-review`/`in-progress` literals): four comparisons decided one question and had to agree; two of them disagreed on a renamed board, so re-entry silently never fired. ## Revert proofs, isolated per site | reverted | result | |---|---| | `destination()` back to attempt 3 | 3 of 38 fail | | degraded-resolution refusals removed | 2 of 42 fail | | capture set back to the legacy five | 2 of 3 fail (renamed-board suite) | | forward exclusions → literals | 1 of 14 fails | | missing-wip refusal removed | 2 of 14 fail | | `promotedFromPlannerColumn` → literals | 3 of 7 fail | | promotion target → `"in-progress"` | 3 of 7 fail | | one rebound guard → `!== "todo"` | 1 of 3 fails (static shape) | | resume lanes → legacy trio | 1 of 5 fails | Every conversion is paired with a negative — a forward move, a not-a-planner-lane card, a default-lineage card, an unresolvable workflow — so neither "always fire" nor "never fire" can pass for "resolve the role". ## Commit discipline Twelve commits, each one thing: the code move (`resolvePlannerLanes` out of `triage.ts`) is separate from every behavior change, and each review fix is its own commit with its own revert proof. ## Verification - `pnpm test:gate` **71/71** - 162/162 across the glasses plugin's 19 files; 26/26 across the four new engine suites - engine + glasses typecheck clean; `pnpm lint` clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Engine recovery and retries now work correctly with renamed or customized workflow columns. * Tasks in manual-intake columns are no longer automatically planned. * Agent actions and quick capture now respect each board’s declared columns and lifecycle stages. * Awaiting-approval tasks are recognized regardless of their current column. * Command Center SDLC funnel stages now accurately reflect customized workflows. * **Documentation** * Added guidance for safely changing workflow-column logic and interpreting lifecycle-column checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
76e92f33c4 |
fix(core): review badges were silent on renamed boards — the same P1 as #2470, one role over (#2586)
Independent of my other open PRs. ## The defect PR #2470's review caught `getStalePausedTodoSignal` gaining a `holdColumn` parameter in B1 while **both** hydration sites in `reads.ts` omitted it — a correct guard comparing against the literal, so the badge was silent for a paused card in a renamed hold column. **That P1 was fixed for `holdColumn` and not for its sibling.** `getStalePausedReviewSignal` and `getInReviewStalledSignal` both take `reviewColumn`, and **all six call sites in the same file** left it defaulted to `"in-review"`. So on a renamed board (`checking`) both review badges were silent — the identical defect, in the identical file, one role over, *after* the pattern had already been found, written down, and fixed next door. ## The transferable part: this class is invisible to the census My column-literal census (#2557) cannot see this. The literal lives in a **parameter default**, and the offending call site **contains no literal at all** — it's defined by what it *omits*. The audit that finds it is different in kind: *"for every role-parameterised signal, does each caller pass the role?"* — run across the **callers**, not the definitions. Result on `reads.ts`: ``` PASSES holdColumn x2 <- fixed by #2470 OMITS reviewColumn x6 <- never fixed ``` ## Why threading differs per path Not one helper call, because the three list paths differ: - `listTasksImpl` / `searchTasksImpl` map **asynchronously** → resolve inline through a per-pass IR cache - `listTasksModifiedSinceImpl` maps **synchronously** → pre-resolve into a Map beside the existing `holdColumnByTaskId`, which exists for exactly the same reason One IR per workflow per pass in all three. ## Evidence Proven against a real store through the **real hydration paths** (`listTasks` and `listTasksModifiedSince`), mirroring the sibling renamed-hold suite because the defect lives in hydration rather than in the pure signal. **Mutation-verified:** reverting the threading fails the two renamed cases and leaves the negative and the builtin regression floor green. The fixture asserts itself — an unpaused or unaged card produces no signal for reasons unrelated to the column, which would let the suite pass while testing nothing. ## Verification - new suite 4/4 - full core PG: **1048 passed / 3 failed** — the same three that reproduce with this change stashed - core `tsc --noEmit` clean; `pnpm test:gate` green (414 + 10 + 71) 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f6010ef558 |
fix(test): planning-lane E2E is RED on main (2/7, incl. its own control) — same unplanned-spec fixture defect #2634 fixed next door (#2658)
Found while establishing the pre-closing E2E baseline for the final verification pass (closing bar, item 4). **Two of this family's seven cases are failing on `origin/main` right now**, and one of them is its own control. ``` releases an ordinary held card on a default board (the control) expected [] to include 'FN-OK' holds a card parked for approval MID-SWEEP, after the snapshot read expected false to be true ``` ## Cause — the same defect #2634 repaired in the file next door `seedHeldTask` creates the task and never writes a `PROMPT.md`, so the card carries only the bootstrap seed. FN-7648's `isUnplannedForExecution` reads that file for any card resting in an intake- or hold-trait column and refuses to move an unplanned card into a processing column, so the sweep released nothing. **Being held was the gate working.** The fixture was exercising the gate rather than the sweep — which is exactly why the *control* failed, and a failing control means the rest of the family's assertions cannot be trusted either. `workflow-lifecycle-live-e2e` had the identical problem and #2634 fixed it the same way. This file landed alongside it (#2611) and did not get the same treatment. Worth stating twice because it is a general rule for this directory: **a release/scheduler fixture that does not model a card which cleared specification is testing the gate, not the sweep.** ## The check that matters more than the fix #2611's stated value is "3/7 red without the guard". Making red tests green is the easiest thing in the world to do wrongly, so I verified the family still discriminates *after* the seed — disabling `isTaskBlockedOnApproval` in `hold-release.ts` still kills exactly three, and the same three: | killed by mutation | |---| | does NOT release a card blocked on manual plan approval on a **default** board | | does NOT release a card blocked on manual plan approval on a **renamed** board | | holds a card parked for approval **MID-SWEEP**, after the snapshot was read | Two cases turned green, zero discriminating power lost. Without that mutation this change would be indistinguishable from weakening the tests until they passed, which the standing rule forbids. Note the third killed case is also one of the two that were failing: it was red for the fixture reason **and** genuinely proves the guard. ## Why it is worth a PR of its own The closing bar's final verification pass (gate, `verify:fast`, all E2E, census) has to run on a green tree. Two red E2E cases on main would otherwise show up in that report as a new failure and cost a diagnosis at exactly the wrong moment. Pre-closing baseline for the record — **13 E2E families, 109 tests, these 2 the only failures**: ``` green agent-count 13 · agent-link 5 · lease-rebound 6 · lifecycle 24 · merge-family 7 merge-rebound 4 · merge-safeguards 10 · merged-board 5 · planner-lane 5 planner-lane-resolution 3 · rebound-family 15 · stranded-column 5 RED planning-lane 7 (2 failing) ``` ## Verification 7/7 green, mutation 3/7 as designed, engine typecheck clean (0 lines), `pnpm lint` exit 0, `pnpm test:gate` exit 0. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cef1b08af3 |
U12: the census baseline follows the count down — and goes in the merge gate (#2661)
Coordinator item 2. The census had the right mechanism and no teeth. ## The gap `--strict` already fails on a rise **and** on an unrecorded drop — that logic was correct. But nothing blocking ran it, so the baseline drifted to **854 while the tree held 787**. That is **67 guards of regression that would have merged silently**: a high-water mark wearing a ratchet's name. This is the same shape as the ceilings I tightened in #2647, one level up. Worth saying plainly: I fixed the vitest ratchet's slack by hand and did not check whether the *authoritative* instrument had the same problem. It did, and by a much larger margin. ## Three changes 1. **`--strict` runs in `test:gate`.** The baseline cannot go stale again without a red gate. 2. **Baseline re-recorded: 854 → 785** across 14 files (`triage` 38 → 9). 3. The single RISE is resolved honestly rather than absorbed. ## The +3 investigation One file rose: `register-task-workflow-routes.ts` **22 → 23**. #2621 replaced one `task.column === "todo"` with `task.column === "triage" || task.column === "todo"` — a net **+1** that also reintroduced a `triage` literal, while the PR title reported *"count 0 → 0"*. Not an accusation. There was no gate for the author to check against, and a hand-counted claim in a PR title is exactly the thing that goes wrong without one. Change 1 is the fix. **The literal is justified and stays**, marked `DELIBERATE-LITERAL` rather than converted. It is the **v1-IR arm**: a v1 workflow yields no role assignments, so `resolveLifecycleColumns` returns nothing and the legacy pre-implementation ids are the only pre-WIP signal available. The `else` branch directly below already resolves intake/hold for every v2 workflow. Converting this arm would not finish anything — it would delete the only answer v1 boards have and admit `in-progress`/`in-review` cards into a rebound that clears worktree, branch and retry counters, which is the regression #2621 was fixing. ## Both directions proven | direction | probe | result | |---|---|---| | rise | add `t.column === 'in-review'` | `live-agent-count.ts: 6 -> 7`, exit 1 | | drop | convert one guard | `self-healing.ts: allows 111, tree has 110`, exit 1 | **The drop probe took three attempts to test honestly, and the first two "passed" while proving nothing:** 1. I renamed a receiver (`task.column` → `Probe`) — the classifier is **fail-closed**, so an unknown receiver is still counted and the number never moved. 2. I targeted a site in `hold-release.ts` that carries a `DELIBERATE-LITERAL` marker — not counted as a column guard at all, so removing it changed nothing. Only removing a counted comparison outright moved the number. Both false negatives came from me assuming the probe worked because the command exited the way I expected. ## On auto-rewrite vs fail-and-instruct You offered either. The script already does **fail-and-instruct**, with `--update-baseline` as the explicit re-record, and I kept it that way rather than making the test rewrite the baseline during a run. Reason: a silent downward rewrite means a conversion PR's own diff never shows the number moving, so "census before/after in the PR body" becomes unverifiable — the reviewer would have to re-derive it. Failing with the new number in the message puts it in the diff where a human sees it, and it costs one command. ## Verification `pnpm lint` clean. `pnpm test:gate` green with the census in it — `every file matches its baseline exactly` (10 / 132 / 487 / 71). Note for the fleet launch: with `--strict` gating, **every** conversion PR must now re-record the baseline in the same PR. That is the intended cost, and it makes the fleet's "baseline must shrink by exactly the converted count" rule mechanically enforced instead of a review instruction. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2fb0df9da8 |
docs(replan-target): the flagged follow-up is done — the note said otherwise (#2665)
Comment only. The block above `resolveReplanTargetColumn` still reads: > **STILL A REAL FOLLOW-UP** … the `return "triage"` fallbacks on the no-match and throw paths name a column the default lineage no longer declares … flagged rather than fixed That directly contradicts the code six lines below it. **#2598 landed the fix:** the no-match path now returns `roles?.hold ?? roles?.intake`, and the throw path returns `undefined`. I wrote that note. A stale *"not fixed yet"* sitting above a fixed implementation is worse than no note — the next reader either distrusts the code or re-does work that is already done. This is the closing-bar item 3 I was assigned, and I nearly re-did it myself: I had the change written and reverted before checking whether main had overtaken me. ## One thing worth recording about #2598's version Its catch-path answer is **stronger than the one I had drafted**. I was going to return `"todo"` — the better guess, since post-U11 the default lineage declares `todo` and not `triage`. #2598 returns `undefined` instead, which forces callers to handle "this workflow could not be resolved" explicitly rather than papering over it with a plausible column id that the move path may then reject. That is the same lesson as the sync-reader audit in #2653: **a defective lookup that returns a valid-looking answer is worse than one that admits it does not know.** Recorded in the comment so the reasoning survives. ## Verification Engine typecheck clean · **51/51** across both replan-target suites · comment-only, no executable change. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9dbc98f1b3 |
Audit: every sync workflow-IR read answers for the DEFAULT workflow (not a PG-only problem) (#2653)
Docs only. This came out of a #2593 review thread that reported the problem as PostgreSQL-specific. **It is unconditional**, and it has consequences well outside the guard I was fixing — including one that looks like a live production break for custom workflows. ## The chain, each link checkable 1. `TaskStore.getTaskWorkflowSelection(taskId)` delegates straight to `getTaskWorkflowSelectionImpl` — **no mode branch** (`store.ts:2545`). 2. `getTaskWorkflowSelectionImpl` **returns `undefined` unconditionally** (`workflow-definitions.ts:505-512`). Its own comment: *"sync selection reader is incomplete-PG; use getTaskWorkflowSelectionAsync."* A PG-cutover stub that never got finished. 3. So `resolveTaskWorkflowIrSyncImpl` always takes its `if (!workflowId)` branch and returns `resolveDefaultWorkflowIr()`. Its `isBuiltinWorkflowId` and `SELECT ir FROM workflows` branches are **unreachable in production**. `resolveTaskWorkflowIrSync` is typed `WorkflowIr`, non-optional — so callers cannot detect the substitution. There is no `undefined` to check and the IR that arrives looks valid. **Why tests don't catch it:** test stores stub `getTaskWorkflowSelection` with a real selection, so the reader works under test and substitutes only in production. Any test written against a stubbed store proves the caller's logic and never the reader's behavior. ## Consequences, severity descending 1. **Custom fields appear to be rejected on custom workflows.** `resolveTaskCustomFieldDefsSyncImpl` returns `ir.fields` — the DEFAULT workflow's. `task-update.ts:128-136` validates against them, and its own comment states the outcome: *"a write against a workflow with no fields (the default) is rejected with a typed CustomFieldRejectionError."* 2. **Per-workflow capacity pools collapse** — `resolveEffectiveWorkflowIdSyncImpl` reads the same selection, so every task resolves to `resolveCapacityPoolId(undefined)`. 3. **Plugin transition hooks re-run against the wrong IR** (`lifecycle-ops.ts:1052`, crash recovery). 4. **Terminal-node detection degrades** to `nodeId === "end"` (`branch-and-pr-entities.ts:578`). 5. **A U7 guard was inert** — fixed in #2593. Its fail-closed arm was `workflowIr ? … : true`, dead code against a non-optional return. **#1 and #2 are REASONED FROM SOURCE, NOT OBSERVED.** I did not execute those paths, and I am labelling them that way in the doc rather than reporting them as confirmed. No test in `packages/core` covers `CustomFieldRejectionError` or `resolveTaskCustomFieldDefsSync` — consistent with the gap, but absence of a test is not proof of a break. **Reproduce before fixing.** I would rather hand you a labelled hypothesis than a confident claim I did not verify. ## Why this matters for the fleet, specifically The census work replaces column literals with trait lookups. A conversion that resolves its traits through a **sync** reader produces a guard that reads the DEFAULT workflow's traits for every task — plausible, wrong, and invisible. **It converts a visible literal into a hidden bug**, and the ratchet counts it as progress. Suggested addition to the fleet brief: conversions must resolve through `resolveWorkflowIrForTaskWithProvenance` and branch on `source`; `resolveTaskWorkflowIrSync` is never acceptable in a converted guard. ## Not fixed here Each consequence needs its sync call path made async — a real slice per site, not an end-of-turn edit. #2593 fixed only the one that was mine. Census unchanged (781 / triage 5); this PR adds and converts no guards. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added an architecture-pattern finding documenting a workflow-reading limitation that can cause synchronous reads to use the default workflow. * Described resulting effects on custom workflow updates, crash recovery, capacity-pool handling, and terminal-node detection. * Documented testing gaps and guidance to avoid synchronous task workflow reads. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
174eb22534 |
cleanup: delete the dead sync capacity-pool helper rather than document it (#2656)
Follow-up to the #2653 audit, and the one item there that is better deleted than described. ## Why delete rather than annotate `resolveEffectiveWorkflowIdSync` **has no callers.** Verified across every `.ts`/`.tsx` in `packages` (excluding `dist`): only its own impl, the `store.ts` import and public method, and one comment naming it. Not exported from the core index, not referenced by any test. It is also **wrong**. It reads `getTaskWorkflowSelection` — the sync selection reader that has returned `undefined` unconditionally since the PG cutover — so it always resolved `resolveCapacityPoolId(undefined)`: the default pool for every task, regardless of workflow. The binding capacity path reads the selection asynchronously inside its transaction and does not use this. That combination is the argument. A dead function is clutter; a dead function that returns a **plausible wrong answer** is a trap. The next person to need "which capacity pool is this task in?" would find a public method with exactly the right name, call it, and get default-pool behavior with no signal that anything degraded. #2653 documents it, but documentation loses to autocomplete. ## Provenance of the claim greptile's P2 on #2653 corrected my first draft, which called this a live capacity collapse — it isn't, precisely because nothing calls it. I verified the no-callers claim myself before accepting, and this PR is the logical end of that correction: if it is unreachable, it should not exist. ## Removed - the impl in `task-store-helpers.ts` - the `resolveEffectiveWorkflowIdSync` public method on `TaskStore` - the import specifier in `store.ts` - the now-unused `resolveCapacityPoolId` import (its only use was the deleted function) - updated the `workflow-definitions.ts` comment that named it ## Verification core / engine / dashboard typechecks clean · eslint clean on both touched files · `pnpm --filter @fusion/core build` exit 0 · **`pnpm test:gate` green (487 + 71)**. The engine and dashboard typechecks are the ones that matter here: removing a public method from `TaskStore` would surface immediately in any consumer that called it, and neither reports anything. ## Census Unchanged (776 / triage 5) — no guards added or converted. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b6b2fdcdc6 |
test(U7): rescue the orphan-triage regression test — main has the fix but not its test (#2663)
Main already carries **every other artifact** from #2593 — the provenance fix, the `DELIBERATE-LITERAL` markers in `TaskCard`/`TaskDetailModal`/`register-routes`, the audit doc. The one thing missing is the test. That is the same artifact class that vanished when #2645's branch was force-pushed, so I rebased #2593 onto current main, found every commit conflicting because the work had landed by other routes, and rescued the one piece that had not. **#2593 can now be closed** — it carries nothing else main lacks. **#2654 needs rebasing onto main** rather than stacking on it. ## What makes this test worth rescuing It took three attempts to write honestly, and the reason is pinned in the test body: on a bare mock, `resolvePlannerLanes` reads `resolveTaskWorkflowIrSync`, which the mock does not define, so it returns `LEGACY_PLANNER_LANES` (`intake: "triage"`) and a `triage` card matches the **first** arm — the orphan arm is never reached. Every earlier fixture I wrote passed through that short-circuit and proved nothing. All three cases stub that reader with the merged default (`intake: "todo"`), which is what production resolves, leaving the orphan arm as the only thing deciding. They differ **only** in the workflow readers. | case | role | |---|---| | **C** — workflow declares `triage` as a review lane | **the discriminator.** Pre-fix, the sync reader ignores the selection, returns the default IR declaring no `triage`, so the arm fires and a card is finalized out of a custom workflow's code-review column | | **B** — workflow resolves, declares no `triage` | positive control; without it "returns false" is unfalsifiable | | **A** — workflow unresolvable | **behavior pin, NOT a regression test** — passes in both worlds | I had A labelled "REGRESSION" until the mutation said otherwise. It is relabelled with the null result documented, because a future edit making it flip would mean the arm's scope changed. ## Verification, stated precisely **231/231** against main's implementation. The mutation that proved C discriminates was run on the branch where the pre-fix code still compiled. **It cannot be re-run against main**: the `WorkflowIr` type import was removed along with the fix, so a naive revert no longer transforms. I am stating that rather than implying I re-verified it here — the discrimination was demonstrated, just not on this base. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
efbbc45eb0 |
U12: the LAST triage guard — Plan was offered on executing cards named triage (#2664)
The final `column === "triage"` in production source, and it was a live
defect rather than dead vocabulary.
## The defect
`isPreExecutionHoldColumn` ORed the legacy id with the traits
**unconditionally**:
```ts
return column === "triage" || flags?.intake === true || flags?.hold === true;
```
That is not a fallback. A resolved column merely *named* `triage`
answered true even when its own traits said work was underway — so the
context menu offered **Plan**, which re-plans, on a card that is already
executing.
Now flags-first, with the id as the documented no-metadata answer.
## Why the file's earlier conversion missed it
Every existing case in `TaskContextMenu.test.tsx` passes a column with
**no flags**, or with `hold`/`intake` set. All of them agree under both
forms, so the suite could not distinguish them. Nothing exercised a
column whose **name and traits disagree**, which is the only shape that
separates an OR from a fallback.
Three new cases cover it. Revert check: restoring the OR form fails the
first one — Plan reappears on a mid-flight card.
## The asymmetry is preserved, and now tested
The degraded set stays `{triage}` **alone**, deliberately not the
`{todo, triage}` used by `isPreImplementationColumnRole`. That helper
drives the preserve-progress prompt, where a flagless `todo` *should*
prompt because losing steps is unrecoverable. This drives Plan, where a
flagless `todo` must **not** offer to re-plan a card that may already be
planned. The file documented that difference; nothing asserted it. Now a
test does.
## On reaching zero honestly
The surviving literal is marked `DELIBERATE-LITERAL`. It is the degraded
answer, not an unconverted guard — there is no trait to read when
`flags` is `undefined`, which happens during first paint and for a card
in a column its workflow no longer declares. Deleting it would silently
withdraw Plan from exactly the stranded cards that most need
re-planning.
So **`triage → 0` means "no unconverted guards remain", not "the string
is gone"**, and I would rather say that than move a number by deleting a
fallback.
| branch | triage |
|---|---:|
| `origin/main` | 5 |
| this PR | **4** |
| #2655 (flag resolution, removes 4 in `moves.ts`) | 1 → **0** combined
|
I found it with the census's own AST classifier rather than grep — my
grep of the same tree returned only comment prose and would have had me
report the bar as met while a real defect sat in
`TaskContextMenu.tsx:179`.
## Verification
`pnpm lint` clean. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0 with the baseline re-recorded in this
PR (column 769 → 768, deliberate 12 → 13). `tsc -p tsconfig.app.json`
clean. `TaskContextMenu.test.tsx` 18/18.
Depends on nothing; stacks cleanly with #2655 and #2661.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
3bf9bf5f74 |
collapse the plan-admission-throttle payload to one gate (+ AGENTS.md) (#2562)
The cross-project semaphore is deleted, so `task:plan-admission-throttled` was describing a gate that no longer exists. Nothing wires `options.semaphore` any more, which left three things dead-but-visible: - `semaphoreAvailable` was permanently `Infinity`, so `Math.min(projectRoom, …)` was a no-op keeping a deleted limiter in the arithmetic - `blockedBy` was a **discriminator** between `"running-agent cap"` and `"global semaphore"`; only the first can occur - four `semaphore*` metadata fields were always `undefined`, and two more terms in the dedupe signature were constant ## `blockedBy` is kept, not dropped Even though it is now a constant. The event exists (FN-8600) to answer *“why did this card sit queued to plan?”* after the fact — a named reason answers that even when there is one gate, whereas a payload with **no** reason field reads as “unknown”. It costs nothing and preserves the shape if a second gate is ever added. The dedupe signature drops the two semaphore terms and keeps the eligible task IDs — that term is what stops a **new** card’s stall being swallowed when the counts land on an unchanged tuple, which is the property the event depends on. ## AGENTS.md It documented the removed field names verbatim, so it is updated in the same commit. Leaving docs describing a payload the code cannot emit is exactly the readable-but-wrong artifact this program keeps deleting. ## Verification `pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green · triage suites **234/234**. --- **Correction I owe on `concurrency.ts`, measured rather than estimated.** I earlier told the coordinator ~75% of its 886 lines could go with the cross-project cap. That was line-range arithmetic and it was wrong. With the cap now fully removed, `concurrency.ts` is **still 886 lines**, because `AgentSemaphore` has four consumers unrelated to it — `verification-concurrency` (maxConcurrentVerifications), `research-orchestrator` (research runs), `experiment-executor` (maxConcurrentExperiments), `step-session-executor` (parallel steps) — plus `ProjectAdmissionCoordinator`, which is FN-8453 oldest-first **ordering**, not a limiter. The real remaining win there is the pre-held-slot bookkeeping and the idle-semaphore leak recovery, which existed to service the global instance; I will measure that as its own slice rather than quote a fraction. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Updated plan admission throttling to consistently use the project’s running-agent capacity. * Improved throttle audit events by reporting stable capacity details and removing obsolete semaphore information. * Preserved accurate deduplication for repeated throttling events, including changes in stalled tasks. * **Documentation** * Updated run-audit guidance to match the revised throttling event format. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8393bba7dc |
U7: the replan rebound targets a column the workflow declares (R7) — re-landed on main (#2598)
> Based on `main`, no dependencies. Re-landed after closing the stacked chain (#2517, and #2551 below) that never reached main. ## What main already has, and what it lacks Main independently converted the planner-lane **parameters** in this file — and **better than I had**: it splits `plannerColumn` from `roles.mergedPlanningColumn`, because a merged lane joins the FN-8596 arrival-order rescue but *not* the "planner column is never advanced" shortcut. That work is main's and untouched here. What main still lacks is the **R7 fix**: `resolveReplanTargetColumn` returns `"triage"` **by fiat** for any workflow declaring neither legacy id — `builtin:marketing` (ideation/backlog/drafting/…) and every fully renamed set. A Plan Review REVISE therefore moves the card into a column its workflow **does not declare**, for `reconcileUndeclaredTaskColumns` to clean up after. A move the engine makes on purpose, not drift. | Workflow | Target | Changed? | |---|---|---| | `builtin:coding` / stepwise | `todo` | no | | Coding (Ideas) | `todo` | no | | `builtin:marketing` | `backlog` (its own hold) | **yes** — was `triage`, undeclared | | declares no planning lane | `undefined` → park | **yes** — was `triage` by fiat | ## The ordering the existing suite taught me Legacy ids stay preferred **first**, and the trait resolution prefers **hold over intake**. That is not arbitrary: Coding (Ideas) declares `ideas` as its intake, and `ideas` is **manual capture with no AI** (plan R10) — a rejected plan sent there stops being replanned at all. The old code got Ideas right **by accident**: it never recognised `ideas` as intake and fell through to `todo`. An "intake first" trait rule would have shipped that regression dressed as a cleanup, and three existing Ideas tests were the only thing between me and doing it. ## An inverted comment, corrected The function's own U11 note read: *"the second lookup asks for `todo`, which U11 deletes… the first lookup still matches `triage` (which U11 keeps)"*. **That is backwards.** #2515 keeps `todo` and deletes `triage`, so the consequence is the opposite of what was written — the `todo` branch is what saves builtin coding. Fixed rather than left, because a comment that inverts a merge's direction sends the next reader to the wrong branch. ## Fail-closed callers `undefined` means "nowhere to replan" (plan U5: *skipped with a log rather than moved arbitrarily*). All four call sites park **visibly** rather than log a move they did not make. The scheduler's rebound still writes `needs-replan` — deliberately, since that is what blocks dispatch and the branch has already decided the card must not be released — with only the *log* made conditional. ## The superseded test is deleted, not skipped A skipped test is a guard that cannot fire. Its replacement asserts the new contract **and** the R7 invariant directly — *"a column this workflow declares"*, not just an id — plus a new case for a workflow with no planning lane at all. ## Verification | Check | Result | |---|---| | replan-target | 41/41 | | with scheduler-trait-dispatch + pre-release-plan-review | 55/55 | | `tsc --noEmit` (engine) | clean | | `pnpm lint` | clean | | `pnpm test:gate` | green (482 + 10 + 71) | | `pnpm check:changesets` | clean | `triage.test.ts` still shows main's **8 pre-existing #2515 failures** — unchanged by this, fixed by **#2576**. ## Closing #2551 Its parameter work is superseded by main's better version; this PR carries the only part main lacked. Same story as #2517: a stacked PR outlived the surface it was converting. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f8c053c3fa |
fix(core): TAKING comments-ops.ts — re-triage on renamed planner lanes (3 triage guards → 0) (#2612)
**Claiming `packages/core/src/task-store/comments-ops.ts`** from the shared backlog so nobody collides. ## Guard count | scope | before | after | |---|---|---| | `comments-ops.ts` | **3** | **0** | | repo-wide `column === / !== "triage"` in `packages/*/src` (excl. tests) | **26** | **23** | ## Why this file, and why it matters more than its size `addComment`'s post-comment **re-triage** decides, from the card's column, whether a user comment should invalidate an approved spec or send already-planned work back for re-specification. It asked with three legacy literals: ```ts task.column === "todo" || task.column === "triage" task.column === "triage" && status === "awaiting-approval" hasRealPrompt && (todo || (triage && status !== "awaiting-approval")) ``` On a renamed board none match, so a user comment on planned work does **nothing**: no approval invalidation, no re-specification, no error. **The operator types a correction and the agent never sees it.** This is the surface a human actually touches, which makes it the worst place in the program for a silent guard. Now resolved per task via `resolveLifecycleColumns`, fail-soft to the legacy pair — this phase is documented best-effort (*"failures are logged but never fail the comment add"*), so an unresolvable workflow must behave exactly as before rather than skip re-triage. ## Red-green, not green-only The suite was written **first** and failed **3 of 6** against the literals — precisely the three renamed cases — while the two negatives and the default-vocabulary floor passed throughout. Both negatives earn their place: re-triaging a **WIP** card would discard an in-flight session, and the **author gate** (agent comments must not re-triage) has to survive the conversion. ## Fixture guards its own preconditions `PROMPT.md` is written where the guard reads it rather than relying on task creation's side effects. `hasRealPrompt` gates two of the three branches, so a bootstrap stub would make those cases pass for the wrong reason — the trap that has produced two vacuous tests in this program already. ## Verification - new suite 6/6; `store-comments` 14/14 - full core PG: **1050 passed / 3 failed** — the same three that reproduce with this change stashed (`central-archive-secrets`, `workflow-settings-project-identity`) - core `tsc --noEmit` clean; `pnpm test:gate` green (482 + 132 + 10) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved comment-driven re-triage for workflows with renamed planning columns. - Comments on planned or awaiting-approval tasks now correctly move eligible tasks to “Needs re-plan.” - Prevented re-triage for tasks actively in progress. - Preserved existing re-triage behavior for standard workflow columns. - Non-user comments no longer incorrectly trigger re-triage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cf6133da8b |
consolidate/e2e — E2E evidence: already-finalized terminal roles (real merge entry, no git) + ledger corrections (#2648)
Consolidation branch for the E2E-evidence worker. Two commits, both engine test/comment only — **no production code, census unchanged**. ## Census (the authoritative instrument) `node scripts/lifecycle-column-census.mjs` on this branch: **triage 10, total 784** — identical to its base. `lifecycle-column-census-ast.test.ts` and `lifecycle-column-census.test.ts` pass (15). This PR neither shrinks nor grows the backlog; it is evidence. ## What it contains, file by file | file | change | |---|---| | `packages/engine/src/__tests__/workflow-already-finalized-live-e2e.pg.test.ts` | **new** — 3 cases, live PG store + real `runAiMerge` | | `packages/engine/src/__tests__/workflow-lifecycle-live-e2e.pg.test.ts` | comment only — retires two unproven-ledger entries | ## The evidence: `isAlreadyFinalizedColumn` never needed the real-git lane My unproven-sites ledger listed it as requiring a git harness because it is module-private inside `runAiMerge`. Reading the function instead of costing the lane: `runAiMerge` reaches it after only `store.getTask`, a pure workspace assert, and a pure branch resolve — **before** the merge blocker, settings, and any branch sync. `projectRootDir` is never touched on that path, and the short-circuit returns a `noOp` rather than throwing. Reachable through the real public entry point with no repository at all. ### Why two cases and not one The guard resolves terminal columns **per role**: ```ts terminal = [lifecycle.complete ?? "done", lifecycle.archived ?? "archived"] ``` #2471's P1 caught the first cut replacing the whole legacy **pair** as soon as *any* terminal role resolved — a workflow declaring `complete` but no `archived` collapsed to one element, silently lost the archived short-circuit, and an archived card then threw *"must be in 'in-review'"* for a card whose real state was "already done, nothing to do". A per-set rule passes for whichever role **is** declared and fails the other, so a single case cannot tell the two rules apart. The shared fixture declares `complete` (renamed `shipped`) and **no** `archived`, so it is exactly that partially-declared shape — resolved half and fallback half live on one board. Mutation-verified, each killing only its own case: | mutation | kills | |---|---| | per-**set** replacement (the #2471 defect) | the legacy-`archived` fallback case | | legacy pair only (conversion reverted) | the renamed-`shipped` case | Plus a differential: a renamed **review** card must not report already-finalized. Without it both cases above would pass for a guard that finalizes everything — turning every merge into a silent no-op, the worst failure this function has. Evidence strength is stated in the file header rather than overclaimed: this reads a returned **decision**, not a persisted row, so it proves the renamed board resolves and short-circuits — not that a card moves. ## Ledger corrections (comment only) Two entries retired, both wrong the same way — each stated a **lane cost** as if it were an impossibility: - `columnIsIntakeOrHold` — "consumers are dashboard-side" is true and irrelevant; its one consumer is an exported pure function. Proven on merged and renamed boards by work already merged in #2631. - `register-task-workflow-routes.ts` — "standing up the route shell is mock-the-world" was false; `createApiRoutes` + `test-request.js` is this repo's established convention with ten existing suites for that file. Covered by its owner in #2614. Counting this PR's own subject, that is **seven** wrong lane-cost inferences in that ledger. The rule it keeps violating is unchanged and now recorded in the file: read what the FUNCTION touches before costing a lane for it. ## Verification `pnpm test:gate` exit 0 (695), `pnpm lint` exit 0, engine typecheck clean, 28 tests green across the AST ratchet and the three merged-board/planner-lane/already-finalized families. ## Not in scope here The `performWorkflowRerunBounce` E2E. The harness exists (`new TaskExecutor(store, "/tmp/test", {})`), but `executor.ts:4305` gates the rebound on the legacy `in-progress`/`in-review` pair while resolving its target by role — so on a renamed board the bounce never fires and the resolved target is unreachable. An E2E asserting today's behaviour would cement that. It is executor.ts's owner's fix; evidence should follow it. Detail in #2632's thread. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
20878e9d5f |
census: count column: "<legacy>" query filters as a separate, separately-pinned instrument (backlog unchanged at 784) (#2650)
Pre-launch input for the 779-guard fleet. **The backlog number does not
move: 784 before, 784 after.** This adds a second number beside it.
## The problem it measures
A guard is not the only way a legacy column id decides behaviour:
```ts
const todo = await this.store.listTasks({ column: "todo", slim: true });
```
That is a **source query** — it selects the rows a sweep considers *at
all*. On a renamed or merged board it returns nothing, so a sweep whose
per-task predicate was correctly converted still does nothing, while
looking converted. `self-healing.ts:2849` names the pairing in prose,
and #2560 had to repair exactly that combination after a converted
predicate was left with a literal query.
The census walks comparison `BinaryExpression`s. A `PropertyAssignment`
is not one, so this class was invisible to the instrument **and to its
ratchet** — it could grow silently.
Measured: **83 query filters, 43 IR node definitions.**
I proved one live consequence earlier on #2648:
`recoverStuckMergeDeadlocks` cannot see a renamed board at all — the
renamed rows exist and none appear in its three-literal union
(`renamedInsideUnion=0`, on a live PG store).
## Why this matters *before* the fleet is briefed
The fleet rule is *"the baseline ratchet must shrink by exactly the
converted count."* In `self-healing.ts` — the largest batch at 111 —
both classes sit in the same functions, so today a worker either:
- converts only the comparisons → arithmetic is clean, and sweeps whose
source query still filters a dead literal stay blind; or
- converts the query too → the count does **not** move by the converted
amount, and a more-correct PR looks like a miscount.
The second punishes the better worker. With a second pinned number,
converting a query becomes visible work instead of an apparent error.
## Counted separately, deliberately
`totals.column` is a published shape — the baseline, the reporter, and
other workers' in-flight PRs read it, and the completion bar is defined
against it. Growing it would move a number the program is actively
driving to zero.
So the new counts live in `summary.properties` / `queryByFile`, under
their own baseline keys, with their own both-directions ratchet (same
rule as #2633's, including the stale-allowance half). `totals` keeps its
**exact** shape — two existing tests assert it with `toEqual`, and
breaking a contract others depend on mid-flight to add a number is not
worth it.
## Definitions are not queries
Workflow IR graph nodes carry `column:` to declare where a node lives —
`{ id: "review", kind: "...", column: "in-review" }`. That is the
lineage describing itself: not a lookup, not convertible, and ~43 of the
raw matches. They are told apart **structurally** (an `id`/`kind`
sibling in the same object literal), not by filename, so a definition
written anywhere classifies the same way.
## Baseline seeding, stated plainly
`--update-baseline` could not pin a **new** category: the regression
check runs before the write, and with no prior key every file reads as a
rise. I seeded the three new keys once, directly, leaving every guard
field byte-identical. The diff is purely additive — no removals.
## Finding, not caused by this change
**`--strict` is already red on clean main**:
`register-task-workflow-routes.ts` is **23** against a baseline of
**22**. Verified by stashing this branch and re-running on an unmodified
tree. Until that is reconciled the guard ratchet is passing nothing —
worth fixing before the fleet starts relying on it as the work order.
## Verification
- census suites **44 green**, 6 new cases: counted; kept out of the
backlog; definition-not-query; both instruments independent (a bug
routing comparisons into the query bucket would otherwise look clean on
both); `DELIBERATE-LITERAL` honoured; non-legacy id ignored
- `node scripts/lifecycle-column-census.mjs` → backlog still 784
- `pnpm lint` exit 0, `pnpm test:gate` exit 0 (695)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
2771408bba |
ci: enforce the lifecycle-column ratchet — it has never actually run (#2654)
**The ratchet was advisory.** `scripts/lifecycle-column-census.mjs` existed only as `pnpm census:lifecycle-columns` — without `--strict` — and **no workflow invoked it**. Nothing has ever compared the tree to the baseline. Every "the baseline ratchet holds them" assumption in this program rested on a check that does not run. That explains both classes of hole: **1. Three PRs lowered counts without re-recording,** leaving allowances the deleted guards could return through while every check stayed green. I've tightened them across #2593 and earlier PRs, but nothing stops the next one. **2. #2621 GREW the count while its own title claimed "count 0 → 0".** It added `column === "triage"` and `column === "todo"` at `register-task-workflow-routes.ts:2681`, taking that file to **23 against an allowance of 22**. It landed unchallenged. This is the failure mode the ratchet exists to prevent, and it happened *inside this program*, in a PR that asserted the opposite. ## The change Adds `check:lifecycle-columns` (the census with `--strict`) to the `pr-checks.yml` lint job, next to `check:changesets` and `check:routes-modular` — the established pattern. **~1.8s over ~1950 files**, so this is not a slow-test addition. ## Proven to fail, in both directions A guard that reports success without checking anything is worse than no guard, so: | injected defect | result | |---|---| | `const __probe = (c: string) => c === "triage"` added to `moves.ts` | `count ROSE — moves.ts: 39 -> 40`, exit 1 | | run against main's current baseline | exit 1 on `mission-feature-sync.ts: allows 5, tree has 0` | Both reverted; exit 0 restored. Note the second row: **this check is RED on main right now**, which is the point. ## Merge order **Stacked on #2593**, which carries the `DELIBERATE-LITERAL` marker for the #2621 site (a v1 IR declares no roles, so no trait can answer that question) plus the baseline re-record. Standalone on main this PR is red — correctly. **Merge #2593 first**, then this. I stacked rather than duplicating those two edits because I already caused one conflict today by appending related content from two branches, and #2651 merged a correction ahead of the section it corrected. Same-content edits in two PRs is the same mistake. ## Census Unchanged by this PR: **776 total, triage 5, reviewed 16** — it adds no guards and converts none. It only makes the numbers enforceable. ## For the fleet This should land before the 776-guard fleet launches. The brief says "the baseline ratchet must shrink by exactly the converted count" — until now nothing verified that claim, so a batch worker could report a shrink that did not happen, or grow the count while converting, and CI would agree. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
642a4fa264 |
consolidate/u12 — U12 consolidation: 4 live defects, the AST ratchet fail-closed, and the moves.ts flag scoped (#2647)
One branch, one PR, per the consolidation directive. Contents file-by-file below. **Supersedes #2625** (its overlapping conversions landed via U11's #2624/#2626/#2636; only the parts nobody else did are folded here). **#2630 and #2639 stay open** — both green with zero threads, per rule 3. ## Four live defects, each measured **1. Every planning card renders an actions menu.** `TaskContextMenu.tsx` still had `shouldShowActionsMenu: task.column !== "triage"` on main *after* the rest of that file was converted. Since #2515 removed the id, the condition is TRUE for every card, so the suppression stopped applying anywhere — including on cards whose menu is empty, the orphaned click target the Surface Enumeration rule exists to catch. Found **twice independently**: by reading the guard, and again by the invariance test below, which failed on main with `shouldShowActionsMenu` true on one lineage and false on another. That is the argument for an invariance property over per-site conversion — the file had already been converted "2 → 1" and the survivor was the live one. **2. Worktree upcoming-work list empty on renamed boards.** `groupByWorktree` filtered `t.column === "todo"`. On the default board the id and the role coincide so every existing test passed; renamed, it matched nothing and a whole panel read as idle. **3. Hold-lane FIFO ordering lost on renamed boards.** `sortTasksForDisplayColumn` gated priority-then-FIFO on `column === "todo"`, degrading to the generic id-ordered sort elsewhere. Cards simply appear in the wrong order, silently. **4. The AST ratchet still failed open** — fourth time in that file, third found by review. `receiverName` understood only one-level property access and bare identifiers, so `task["column"]`, `metadataColumn(entry, "to")`, ternaries, `(task!.column)` and backtick literals were dropped. **Measured on main: `in-progress` 196 → 197, `in-review` 211 → 213** — three real guards nobody counted, including `metadataColumn(entry, "to") === "in-review"` in `reliability-metrics.ts`. Now walks wrappers, resolves calls to the callee name, and emits a `<SyntaxKind>` **sentinel** for anything unnameable: counted *and* trips the classification guard, so a human judges it instead of it vanishing. ## Per-file guard counts | file | before | after | |---|---:|---:| | `app/components/TaskContextMenu.tsx` | 1 | **0** | | `app/utils/worktreeGrouping.ts` | 1 | **0** | | `app/components/taskSorting.ts` | 1 | **0** | The other dashboard files I had converted reached 0 via U11's PRs; where our work overlapped I took theirs during the rebase, including two places where theirs was **stronger** than mine — they deleted Column's unreachable quick-create arm outright (with fixtures migrated) where I had converted it, and they verified the same `isPreExecutionHoldColumn` degraded-set asymmetry I did, independently. ## Flip precondition: the moves.ts flag is scoped, not flipped `move-target-declared-census.test.ts` answers precondition 2 with measurement. 41 engine `moveTask` calls have literal targets — `todo` 27, `in-progress` 7, `done` 6, `archived` 1 — and **all four are declared by the default lineage**, so the default board is not the exposure. `triage` appears only in a comment noting `replan-target.ts` used to hardcode it. My own grep had said `todo=29`; the AST says 27, because grep counts comments. The exposure is **custom** lineages: 20 of the 41 carry no `recoveryRehome` and would reject with unknown-column post-flip; 21 are exempt via the #1411 carve-out, which makes that carve-out load-bearing. I did not flip the flag. It is six seams, not the `789`/`837` pair every summary including mine described, and seam 2 turns on *new refusals* rather than swapping equivalent implementations — a green suite says nothing about that. #2639 pins the blast radius. ## Tests - `column-role-id-invariance.test.tsx` — hold traits fixed, vary only the column id across MERGED / LEGACY / RENAMED; every decision must agree. Drives the real consumers, so a component keeping an inline comparison fails it. Includes a unanimous-and-**false** case so it can't be satisfied by a predicate hardwired to true. **This is the test that caught defect 1 on main.** - `worktreeGrouping.test.ts` — includes two cards both in a column named `staging`, one hold and one not, asserting opposite answers. That assertion is impossible under a board-wide column-id set, which is why hold resolution is keyed per task via `getEffectiveTaskWorkflowId` (#2625 review). - `taskSorting.test.ts` — discriminates on the **tiebreak**, not priority: both branches sort by priority, so my first version passed for the wrong reason. Equal-priority cards whose `createdAt` order disagrees with their id order. - `no-hardcoded-lifecycle-columns.test.ts` — 16 detector cases: 11 shapes counted, 4 legitimate ignored, one asserting the sentinel path. Revert checks, all run: menu suppression → diff names the field; worktree → `expected [] to include 'FN-50'`; sort → `FN-2, FN-9` instead of `FN-9, FN-2`; ratchet → the 3 recovered guards disappear. ## One site that should never be converted `MissionControlPanel.tsx:46` — `{ id: "triage", match: (c) => c === "triage" || c === "signal" || c === "backlog" }` is a deliberate name-similarity heuristic for the SDLC funnel; it matches synonyms and folds unknown columns into an "other" bucket so custom columns still contribute. Converting it changes what the funnel displays. Like the `live-agent-count` fallbacks, it belongs in a documented floor — **the ratchet's target is that floor, not zero.** `DocumentsView.tsx:73` is convertible but the file has no column flags at all, so a real fix means plumbing board-workflow metadata into a view that doesn't fetch it — its own unit of work. ## Verification `pnpm lint` clean. `pnpm test:gate` green (10 / 482 / 71). `tsc -p packages/dashboard/tsconfig.app.json` and `packages/core/tsconfig.json` clean. Core ratchet + seam suites 24/24. Dashboard target suites 37/38 — the one failure is the pre-existing `"Back to In Progress"` label casing, confirmed identical on the base. --- ## Added after the initial push **5. `TaskCard` lost inline editing on renamed boards; `TaskDetailModal` kept it.** Still live on main: the modal resolved field editability from traits in U10/R8, the card used a hardcoded `{triage, todo}` set with **no trait path at all** — even though `taskColumnFlags` was already in scope. On a renamed board the title was editable in the modal and the pencil was missing from the card. Body moved unchanged into `isFieldEditableColumnRole` so the two surfaces cannot drift again. The veto traits are the substance: a column can legally carry `hold` **and** a WIP or review trait, and a plain `intake || hold` check would let an operator rewrite a description while a session executes against it. Coverage gap **measured, not assumed**: mutating `canEdit` back to the hardcoded set left `TaskCard*` at the same failure count as the unmutated run — nothing caught it. The four render cases assert the real `aria-label`; that mutation now fails with `Unable to find an accessible element ... name 'Edit task'`. **6. The ratchet's target is a documented FLOOR, not zero** — and this changes the completion bar. Zero is not reachable, and chasing it means breaking working code. Two categories are permanent, now protected as positive assertions so a future sweep cannot "finish the job" by deleting them: - `MissionControlPanel.tsx`'s `FUNNEL_STAGES` is a deliberate **name-similarity** heuristic — it matches `signal`, `backlog`, `to-do`, `ready`, `shipped` and folds unrecognised columns into an "other" bucket so a custom board still contributes counts. It is not asking whether a column has the intake trait; it buckets arbitrary column *names* for display. Asserted on the **synonym list**, because the synonyms are what prove it is name matching — if they disappear the site has changed character and the exemption stops applying. - `live-agent-count.ts`'s no-flags arm is reachable (a remote store is deliberately given an empty flag map; a card in an undeclared column has no flags at all) and deleting the literal makes such a card match **no** arm, so the queued total silently under-reports a stranded card. A count with an undocumented floor invites someone to drive it to zero. **Not done, and why:** `DocumentsView.tsx:73` is convertible but that file has no column flags anywhere, so a real fix means plumbing board-workflow metadata into a view that does not fetch it — its own unit of work, not something to smuggle into a conversion. **Re-verified after these commits:** `pnpm lint` clean, `pnpm test:gate` green (10 / 482 / 71), `tsc` clean on core and `tsconfig.app.json`, core ratchet suite 26/26, `columnRoles` 10/10, `TaskCard.test.tsx` 384/386 (the 2 are pre-existing CSS assertions). `TaskDetail*` is 130 failed / 551 passed **both with and without** this change — verified by stashing, so pre-existing and unrelated. --- ## Flag resolution: preconditions 1 and 2 are now DISCHARGED. Precondition 3 is blocked, and by evidence. **Precondition 1 — the side-effect equivalence proof — done.** `moves-flag-equivalence.test.ts` runs the same journey under both flag states against live PG and diffs the persisted row. **Result: identical** — whole-row equality across 128 fields plus an equal timing shape, over `todo → in-progress → in-review → todo → in-progress`. That test was **wrong twice** before it meant anything, and both times it was passing: 1. **It proved nothing.** `experimentalFeatures` is **global-only**, and `moves.ts` reads `getSettingsFast()`, which filters global-only keys out of the project layer. My `updateSettings` write was silently discarded, `useWorkflow` was false in *both* runs, and the "proof" compared the legacy path against itself. Found by stamping the flag-ON branch and observing the test still passed. Now written via `updateGlobalSettings`, and the helper **asserts the flag took effect** before the journey runs. 2. **The journey was forward-only**, so it never reached the reopen hook's field resets (`status`, `error`, `blockedBy`, pause clearing) — a mutation there passed. Extended with a backward move and a re-entry. Mutation-verified after both fixes: stamping seam 3, and diverging the reopen hook, each fail the comparison. **Precondition 2 — done, and its answer is a blocker.** The census says the default board is safe: all 41 literal engine move targets are declared by the default lineage. But **20 of those 41 carry no `recoveryRehome`**, so on a custom lineage that does not declare `todo` / `in-progress` / `done`, seam 2 would start rejecting them with unknown-column. That is a user-facing break on custom boards, not a theoretical one, and it is not fixed by the equivalence proof — seam 2 adds *new refusals* rather than swapping implementations. **So the flip is one step away, and the step is not mine to take alone:** those 20 call sites need to resolve their target from the task's workflow (or justify `recoveryRehome`), and they live across engine lanes in `moves.ts` caller territory — U2b/MAIN. Flipping before that trades a dormant flag for broken custom boards. What remains for precondition 3 once those land: flip both readers **atomically** (`moves.ts` + `workflow-task-create-ops.ts`, since the latter computes the preflight the former consumes), delete the flag-OFF branch with its guards, and drop the settings key. --- ## CORRECTION: seam 2 is not a blocker. My earlier claim was wrong. I stated in #2639 and above that "with the flag off there is **no** target-column validation on the move path", so flipping would introduce new refusals. **That is not what happens.** Reproduced against live PG: the identical custom-lineage move rejects with the flag **OFF** as well — ``` Error: Invalid transition: 'backlog' -> 'todo'. Valid targets: building ``` Transition validation is already in force on the flag-OFF path. So for the shape in question — an engine move to a column the task's own workflow does not declare — **the move already fails today**, and seam 2 introduces no new break for it. The 20 census sites lacking `recoveryRehome` are broken on a custom lineage *now*, not broken by the flip. I found this because the discriminator I added to prove "the flag is the cause" failed. Had I written the test to my assumption it would have passed and the false claim would have shipped — the same way the equivalence test passed while proving nothing until I tried to make it fail. **Revised precondition status:** | precondition | status | |---|---| | 1 — side-effect equivalence | **discharged** — identical rows, mutation-verified both directions | | 2 — seam-2 exposure census | **discharged, and it is not a blocker** — the rejection predates the flag | | 3 — flip both readers atomically, delete the flag-OFF branch, drop the settings key | **the remaining work** | So the flip is no longer gated on fixing 20 engine call sites. What it is still gated on is precondition 3 being done atomically across `moves.ts` and `workflow-task-create-ops.ts` (the latter computes the preflight the former consumes), which is `moves.ts` caller territory. Three cases now cover seam 2: the flag-ON rejection, the flag-OFF rejection (asserting the error *message*, so a change in which guard rejects stays visible rather than reading as agreement), and the #1411 `recoveryRehome` carve-out succeeding — pinning why that carve-out is load-bearing and must not be tidied away. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f5cc416ae4 |
U7 item 3: the replan no-match fallback named a column no lineage declares (#2659)
**Item 3 from the closing bar.** Behaviour change, own commit. ## The defect `resolveReplanTargetColumn` fell back to the literal `"triage"` when a workflow declared neither legacy planner id. That names a column the workflow doesn't declare — and since #2515 the **default lineage doesn't declare it either**, so the fallback pointed at a column that exists nowhere. The replan move then either failed outright or put the card somewhere no sweep owns. Resolved through `resolveReboundTarget` (KTD-10: hold → intake → first declared) — the same helper every other rebound path uses, so replan lanes and rebound lanes stay consistent instead of drifting. ## The catch path keeps its literal, deliberately It's reached only when resolution **throws** — not when it silently falls back to the default IR, which returns a real workflow and takes the `todo` branch above. With no IR there's nothing to resolve, and swapping one arbitrary literal for another changes behaviour without evidence about the workflow. Documented at the site so the asymmetry reads as a decision, not an oversight. ## Test Written first and observed **red**. It asserts the target is a column the workflow actually declares: ```ts expect(workflowHasColumn(ir, target)).toBe(true); ``` rather than pinning a specific id — so it can't pass by naming a *different* wrong column, which is how the previous version of this test stayed green while the fallback was broken. **Mutation-verified:** restoring the literal fails it. ## Verification 40 replan-target tests green, engine tsc clean, lint clean, merge gate green (487 + 132 + 10). No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
be63e72f10 |
U11 [E2E evidence]: live-PG proof for the stranded-column rescue and the planner-lane asymmetry (8 tests, test-only) (#2629)
**Completion bar #3 for my phases.** Test-only, no production changes, no guard-count movement — the two live-PG E2E suites I held during the freeze. ## Why these exist Every U11 slice I shipped closed with the same caveat: *all evidence is unit-level*. Three claims in particular were argued from reading code, and each is the kind a mock would happily confirm: 1. #2515 left `triage` a legal id but removed it from the default lineage. 2. #2603 — `createTask` resolves the workflow's intake column, and an explicit `column` **overrides** it. Nine write sites were removed on that reasoning. 3. #2591 — a card stranded on a legacy planner id is admitted by planning discovery, which is what lets it heal with no data migration. Both suites drive a **real PostgreSQL TaskStore** (per-file throwaway database) and the **real shipped workflows**, not fixture IRs. Claim 3 goes through the real `discoverReadyPlanningTasks` — the method the poll calls. Every assertion is on **observed persisted state** (fresh `getTask` after clearing the task cache), the rule inherited from `workflow-lifecycle-live-e2e.pg.test.ts`, because "a function was called" is exactly what has passed falsely on this program before. ## Two things the E2E found that unit tests did not **The shared fixture's "merged" shape was not #2515's.** Omitting `separateIntake` leaves the hold column with *no* intake trait, so the resolver reports `undefined` — "I have no intake to name" — whereas the shipped merged lineage carries intake **and** hold on one column and reports `[]` — "intake exists and *is* the hold column". Callers treat those differently: `undefined` keeps their legacy default, `[]` positively asserts no dedicated planner lane. Assuming the plain shape was the merged shape is how a test appears to cover #2515 while covering something else. Added an opt-in `mergedIntake` to model the real thing; the third shape is now asserted explicitly. **`insertWorkflowDefinitionSync` throws in backend mode** — it's the SQLite path. The suites use `createWorkflowDefinition` + `writeTaskWorkflowSelection` like the other live E2Es, including binding to the id the *store* allocated rather than the one passed in, which the lifecycle suite documents as a way a renamed-workflow fixture silently resolves to the default IR. ## Fixture changes are opt-in Both new options follow the existing `mergeOrchestration` precedent: seven suites build on this builder and a shared fixture must not silently change an existing suite's subject. ## Naming `workflow-planner-lane-**resolution**-live-e2e` deliberately, to stay distinguishable from #2611's `workflow-planning-lane-live-e2e`. Different subjects — that one drives the real hold-release sweep, this one drives the resolvers the lane guards consume. Near-identical names would invite someone to delete one as a duplicate. ## Verification - 8 new tests green against a real PG store - **Mutation-verified:** disabling the #2591 rescue in `discoverReadyPlanningTasks` fails claim 3, and only claim 3 - Merge gate green (482 + 132 + 10), engine tsc clean, lint clean **Pre-existing failures, not from this PR:** the full live-E2E sweep is 82 tests / 2 failed, both in `workflow-lifecycle-live-e2e.pg.test.ts`. Verified by swapping main's `_workflow-vocabulary-fixture.ts` in and re-running: 2 failed either way, identical. They are main's, and they appeared since my earlier clean run of that suite — worth a look against bar #2. ## What this does not cover Neither suite runs a planning **session** — that lane is the AI, substituted here as `testMode` does in production. So this proves a card is *admitted* and re-homable, not that a full plan-and-release round trip happens. The release half is covered by the existing lifecycle E2E. No changeset: `@fusion/engine` is private and this is test-only. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
5481c27729 |
docs(solutions): finding 6 — read the implementation before claiming its output is wrong (#2649)
Completes `proving-a-code-path-actually-runs.md` (merged as #2642) with the rule its own author broke three times while writing it. **Docs only.** ## Why this belongs in that document rather than a new one Findings 1-5 are about proving **your own** claim: does this path run, can this test fail, is this negative result observable. Finding 6 is the mirror image — the claims we make against **other people's** work — and it is the same underlying error pointed outward. Splitting them would let a reader take the first five as "be rigorous about my code" and miss that the identical discipline applies when reviewing someone else's. ## The three cases, all mine, all in one day | What I claimed | What was actually true | |---|---| | The census undercounts triage guards, 13 vs 10 | `summarize()` counts `byColumnId` only for `kind === "column"`. My patched counter summed `role`, `status` and `deliberate` too. The three "missing" ones were exactly the ones it classifies correctly — and I reported this against the instrument the program had just adopted as authoritative. | | `resolvePlannerLanesForTask` silently disables two recovery paths for legacy cards — escalated across four messages | The file's own header had already reasoned it through and documented why that answer is correct. And `TaskStore` implements `getTaskWorkflowSelectionAsync`, which the resolver prefers — so real projects never take the path my `{ getTask }`-only probe forced. | | `executor.ts` is clean of triage guards | A receiver-specific grep missed three under `from` and `originColumn`. Same error one step earlier: trusting a reconstruction of the thing instead of the thing. | Every one was: reconstruct behaviour from outside → compare to actual output → find a difference → report a defect, **without reading the implementation.** ## The rules it adds - Read the implementation and its header comment before reporting anything as wrong. On this codebase the reasoning is usually already written down, and the FNXC note frequently answers the exact objection — twice today it answered mine verbatim. - **A fixture is not a measurement of production.** When a probe and the real system disagree, suspect the probe: ask what it had to stub, and whether production ever supplies that shape. - Retract precisely and immediately. A false defect report against shared infrastructure costs more than the bug would have — it sends people to verify something already correct, and spends the credibility needed for the next report that is real. Also updates the count in the intro (five → six) and adds an `applies_when` entry so the doc surfaces for "about to report a tool as defective", which is when it is needed and not when someone is already debugging. `pnpm lint` clean. No changeset — internal documentation. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
177c2309d9 |
consolidate/u11: a hold column is a planner lane only if it precedes wip (real defect in merged code) + funnel aliases (triage 11 -> 10) (#2645)
**Consolidation branch for u11/u7.** Supersedes #2624. Contents changed substantially while it sat unmerged — this body reflects what is actually in it now. ## Measured with the authoritative census, not grep `node scripts/lifecycle-column-census.mjs` — **triage guards 11 → 10.** ## 1. A real defect in merged code: a hold column is only a planner lane if it precedes implementation Found by greptile on #2616, verified by me, fixed here at the source because that PR cannot land. `resolveLifecycleColumns` returns `hold` as the **first** hold-trait column in declared order, with no positional constraint relative to wip (`workflow-lifecycle-traits.ts`: `hold: first(LIFECYCLE_ROLE_FLAGS.hold)`). A workflow using a hold trait for a **mid-pipeline wait** — a pause after implementation starts — therefore had that column returned as its planner lane, and `reconcileMissionFeatureState` demoted the feature to `triaged`. The mission board reported started work as not-yet-started: silent, and wrong in the direction that makes a roadmap lie. This is my defect, introduced in #2610. **Why it survived:** every lineage anyone has tested puts the hold *in front* of wip, so the default and Ideas boards are unaffected and no existing test could see it. **The fix is positional, with a deliberate asymmetry.** A hold column counts only when it appears before wip in declared order. When wip cannot be located the hold is left **out** rather than guessed — including it wrongly demotes live work on the roadmap, while excluding it wrongly costs only a `triaged` transition the next reconcile re-applies. Mutation-verified: dropping the positional test fails the mid-pipeline case and nothing else. ## 2. MissionControlPanel funnel aliases Assessed and **deliberately not trait-converted**. These are heuristic *name aliases* for a canonical SDLC stage — the matcher already accepts `signal`/`backlog`/`ready`/`shipped` because it buckets arbitrary boards, with an `other` fallback. Post-#2515 a default board's planning cards sit in `todo` and count at the Todo stage, leaving Planning at zero: the funnel reporting where cards *are*, not a guard that stopped firing. Hoisted to a named set so it stops reading as unconverted. This is the **DISPLAY-ALIAS** class the census still lacks — receiver *is* a column id, purpose is presentation rather than a lifecycle decision. `DocumentsView`'s status dot is the other one. Without that bucket a ratchet will keep demanding conversions that make the product worse. ## What I dropped, because main's version was better The original #2624 carried a `TaskContextMenu` conversion. #2626 landed `isPureIntakeColumn` — intake **without** hold — while mine treated any intake-flagged column as intake. That's wrong for a **merged Planning column**: it carries both traits, cards there wait for capacity and have real actions, so I would have suppressed the menu where it belongs — a new regression in place of the one I was fixing. Theirs is correct. Mine is gone, along with its now-invalid test and a helper nothing else used. ## Verification - merge gate green (482 + 132 + 10), engine tsc clean, dashboard tsc clean, lint clean - 7 planner-lane tests green, mutation-verified No changeset: `@fusion/engine` and `@fusion/dashboard` are private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
3e8f604848 |
test(engine): census the UNCONVERTED lifecycle surface — 417 legacy column literals, ratcheted (#2557)
Test-only, no production change. Independent of my other open PRs.
## The number nobody was counting
This program has two censuses, and **both count converted things**: the
unproven-sites ledger (callers of the lifecycle-role resolvers) and
`raw-workflow-columns-flag-census` (reads of the `workflowColumns`
flag).
Neither counts what is still keyed to a legacy column id — **which is
where every defect this program has found actually lived**:
| defect | the literal |
|---|---|
| pool-id sentinel (capacity gate never bound) | `?? "builtin:coding"`
vs the counter's sentinel |
| agent-link leak (slot consumed forever) | terminal column matched
against a fixed id set |
| stale-paused badge silent on renamed boards | `task.column !== "todo"`
|
| merge chokepoint threw on a finished card | the `done`/`archived` pair
|
| recovered card stranded harder | `?? "todo"` |
Every one was found **by hand, one at a time, by whoever happened to
look.**
## Measured
**438 lifecycle decisions keyed to a legacy column name** (417
comparisons + 31 `??` column fallbacks, minus 8 agent-id false positives
and 2 lines carrying both shapes), across 85+ production files — 94 in
`self-healing.ts`, 70 in `executor.ts`, 26 in the dashboard
task-workflow routes.
That is the real size of the remaining surface. It dwarfs the 15-site
resolver census I've spent this unit closing, which is worth knowing
before anyone calls the vocabulary work finished.
## A hit is not a bug
Many are correct — documented legacy fallbacks, the legacy-adoption
path, code genuinely about the built-in workflow. The census claims only
that each site decides by **name** rather than by **role**, and
therefore needs a human judgment. Reporting 417 as a bug count would be
exactly the overclaiming this program keeps correcting.
## A ceiling, not an equality — deliberate
The sibling flag census fails in both directions. That number moves only
when two units touch it. **This** one moves whenever any of a dozen
concurrent conversion slices lands, and an exact-equality assertion
would go red on work heading the *right* way.
A test that's red for good reasons gets suppressed, and a suppressed
ratchet is worse than none — the failure mode AGENTS.md's quarantine
rule exists to prevent. So the count may fall freely and may never rise;
when it falls, the failure message says to lower the pin.
## Verified in both directions
- green at 417
- adding **one** literal to `replan-target.ts` → `census ROSE to 418
(ceiling 417)`
- the regex is unit-tested to count a **decision**, not a mention: a
column id in a fixture, a log line, or a `moveTask` argument is not
counted — inflating the number into noise is how a census stops being
acted on
- unreadable sources **fail closed** rather than silently shrinking the
count
## Follow-up (a8c150b12): the census was blind to three of the five
defects it cites
I ran the census against its own header. It lists five motivating
defects; the comparison-only regex counted **two**. The pool-id
sentinel, the rebound strand and the terminal fallback are all `??`
**defaults** — invisible to a `.column === "x"` pattern.
A census that cannot see three of the five bugs it names as its reason
to exist is worse than none: it reports a number that *feels* like
coverage. That is precisely the overclaim this unit keeps catching in
other people's work — caught here in mine, and only because the header
wrote the examples down somewhere they could be tested against.
It now counts two shapes — deciding **by** a name (`===`/`!==`) and
**defaulting** to one (`??`) — and pins the five motivating examples as
a test case, so the pattern cannot narrow back without failing.
**Measured: 417 comparisons + 31 fallbacks, of which 2 lines carry both
shapes → 446 lines.** Ceiling raised 417 → 446 to cover the missing
shape, not to excuse new debt.
`?? "builtin:coding"` stays deliberately uncounted: it defaults a
*workflow* id rather than a column and is legitimately correct at most
sites. It already has a stronger guard —
`scripts/check-capacity-pool-id.mjs` bans it only where the value
reaches a capacity counter, which is the only place it's wrong.
Verified both directions: green at 446; adding one fallback of the
newly-counted shape → `census ROSE to 447 (ceiling 446)`.
## Follow-up 2 (98f4264fd): 8 false positives removed — 446 → 438
Then I checked the census against real source instead of trusting the
pattern. Its top-scoring fallback file was `triage.ts` with 8 hits — and
**every one is `agentId: task.assignedAgentId ?? "triage"`**, an *agent*
id, not a column. `"triage"` is both a column id and the synthetic agent
id triage stamps on its audit rows.
Eight of ~34 fallbacks is a quarter of that shape: enough to make the
number **wrong** rather than merely imprecise. A census with known false
positives is one people learn to discount — the same end state as not
having one, which is exactly what its own header warns about.
Excluded, and the exclusion is **pinned as a test case** so it can't
creep back: the three agent-id spellings must match the raw shape *and*
be filtered, while a genuine column fallback that also mentions triage
(`first("intake") ?? "triage"`) must still count.
**Residual imprecision is stated rather than tuned away.** A couple of
counted lines are display defaults (a column rendered in CLI output).
They stay: the census claims each site *needs a human judgment*, and a
display default passes that judgment in seconds. Chasing them costs more
than the precision buys and makes the pattern too clever to trust.
Agent-ids were excluded because they're a quarter of the shape — not
because any false positive is intolerable.
Ceiling 446 → **438**. Verified both directions: green at 438; one new
fallback → `census ROSE to 439`.
## Verification
- census 3/3; engine `tsc --noEmit` clean; `pnpm test:gate` green (414 +
10 + 71)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
2a4013b723 |
consolidate/u9 — review+merge lane: E2E evidence, re-greens, and the conversion blocker (#2646)
**U9's consolidation branch.** Supersedes nothing — #2637 and #2643 are green with zero threads and left for your sweep per rule 3. ## Contents | File | Change | Before → After | |---|---|---| | `__tests__/executor-step-numbering-zero-based.test.ts` | isolate the review-handoff `moveTask` call so the assertion is attributable | **1 failed / 3 passed → 4 passed** | | `__tests__/ce-workflow-step-executor.test.ts` | re-green against the block-first merge boundary | **3 failed / 48 passed → 51 passed** | | `__tests__/goal-anchoring-audit.test.ts` | swallow path reports at debug, not `console.warn` | **1 failed / 6 passed → 7 passed** | Triage-guard counts: **no change**. My lane has no remaining column receivers — the rest belong to the capacity/U7/U8/U11/U12 workers, or are deliberate compat retentions I verified individually (`spec-staleness.ts` carries its own "U11 proof" block; `live-agent-count.ts`'s literal fallback is reachable by flag-less callers). Commits kept small and separated: signature fix, then attribution fix, then the boundary re-green, then the debug-channel fix. **Census reconciliation:** `node scripts/lifecycle-column-census.mjs` reports **11** triage guards on main, and **none are in the review/merge lane** — they are the `moves.ts` flag-OFF branch plus the dashboard cluster. Nothing in this branch moves that number, and I am not chasing the 779 non-triage guards per your instruction. ## 1. The review-handoff assertion (and a lesson) The handoff gained a third argument (workflow move provenance), so a two-arg `toHaveBeenCalledWith` failed on the extra options object while the card moved correctly. My first fix used `expect.anything()` — and I *documented in the comment* that six mutations couldn't make it fail, then shipped it anyway. Greptile (P2) correctly called that out: this flow records two `moveTask` calls, so the assertion is satisfied by the boundary move even if the handoff regresses. **Documenting a weakness is not removing it.** Now the test selects the handoff call by its own marker (`workflowMoveMetadata.reason === "workflow-review-handoff"`), asserts exactly one such call, and asserts its target column: | Mutation | Before | After | |---|---|---| | change the seam's `reason` | green | **NEW=1**, this test only | | retarget the seam to `"done"` | green | **NEW=1**, this test only | ## 2. The merge boundary changed shape `ensureWorkflowMergeBoundaryTask` (`executor.ts:7808`) now **refuses** a foreach step-execute region with incomplete pre-merge node proof — logging `"Workflow merge boundary blocked: <reason>"` and returning **without moving**. The move-then-check sequence this file pinned is gone: `"Workflow merge boundary moved task to in-review before requesting merge"` no longer exists anywhere in production. Three fixes, one per failure: 1. **negative case** pinned the retired move-first log. Now pins the *stronger* property the new order gives: an unproven card is **not moved into review at all**. The old assertion could only say "it was moved, then blocked". Log text asserted by stable prefix — the reason clause enumerates missing instance ids, which is legitimately volatile. 2. **"moves direct-to-merge tasks into in-review"** got zero calls: its fixture recorded no node results, so the gate blocked it. Added one `steps#0:step-execute` pre-merge result. 3. **"completes graph-native checklist projection"** also got zero calls. Its existing `plan` result proves *some* pre-merge node ran but not the per-instance work; the gate additionally requires an instance per foreach step-execute. Added the two matching its two steps. (2) and (3) are the same class as the lifecycle E2E `seedTask` fix in #2634: a fixture that never modelled completed work, asking the engine to advance it, and reading the correct refusal as a failure. Proof shape matched to the evaluator (`source: "node"`, `phase: "pre-merge"`, terminal = `passed`/`skipped`) rather than guessed. Verified the gate is what these fixtures exercise: disabling the boundary proof check fails the negative case (`NEW=1`, that test only). `pnpm test:gate` green, `pnpm lint` clean. ## Where U9 actually stands The conversion (S06/S07/S08) is **not** done, and is now precisely characterised rather than "blocked on U8": `workflow-graph-executor.ts:310` short-circuits every `MERGE_REGION_KINDS` entry to the legacy merge seam, so `merge-gate`, `merge-attempt`, `manual-merge-hold`, `retry-backoff`, `recovery-router` and both `branch-group-*` handlers **never execute**. `createMergeGateHandler` does read `task.autoMerge` and emit auto-on/auto-off — and is never called. The builtin IR's `outcome:auto-*` edges are unreachable. **U9's conversion, concretely: stop short-circuiting `MERGE_REGION_KINDS` and let those nodes run.** S06/S07/S08 all hang off that one change. Safeguard 2 has no node-level representation today, so enabling the region without carrying the `autoMerge` contract into it would let an `autoMerge:false` card merge on PR-readiness alone. Full write-up in `docs/plans/workflow-owned-merge-stack/u9-safeguard-baseline.md` (#2634). ## 3. A recurring class worth a shared helper `goal-anchoring-audit`'s swallow path now reports via `log.debug` (a deliberate demotion of log noise), and `debug` is FUSION_DEBUG-gated so vitest emits nothing — the test asserted a channel that was both wrong *and* disabled. I kept both halves of the contract (swallowed **and** reported) by enabling the flag for that case, rather than deleting the awkward assertion. **This is the third instance this session** — `worktree-pool`, `self-healing`'s auto-archive line, and now this. If a fourth appears it deserves a shared test helper rather than three bespoke fixes. ## Two failing files I could NOT responsibly take — flagged, not touched **`executor-prompt.test.ts` (3 failures) — I ESCALATED THIS AND I WAS WRONG. Retracting.** I flagged these as a possible real pause-contract violation: an agent session spawning while an operator has globally paused the engine. I then finished the diagnosis, and the evidence goes the other way. Recording the retraction with the same detail as the alarm, because a false alarm aimed at another unit costs them a chase. **The discriminator I asked for, resolved.** Six tests in that file assert `expect(mockedCreateFnAgent).not.toHaveBeenCalled()` during global pause; 3 fail. Splitting them by what they drive: | Assertions | Drives | Result | |---|---|---| | `does not resume unpaused in-progress task while global pause is active` (+2 siblings) | no executor method — `task:updated` / resume paths | **pass** | | `parks todo tasks in in-progress when fn_task_done…` (+2 siblings) | `executor.execute(...)` **directly** | **fail** | So the guard holds on every event-driven path and is absent only from the direct `execute()` entry. **And `execute()` is not the guard site — the scheduler is.** `scheduler.ts:1491` is an explicit hard stop (*"Global pause (hard stop): halt all scheduling activity"*), with a second gate at `:1055`, and the scheduler never calls `.execute(` at all — dispatch routes through the runtime. In production a global pause halts scheduling before anything reaches the executor. **Conclusion: the pause contract is intact in production.** The 3 failing tests call `execute()` directly, bypassing the upstream gate, and assert a defence-in-depth check *inside* `execute()` that is not there. They are testing a path production does not take during a pause. What that leaves is a real but much smaller question, and a design one rather than a defect: should `execute()` carry its own pause check as defence-in-depth, given non-scheduler callers exist (self-healing, manual retry)? If yes, add the guard and all six assertions pass. If no, the 3 direct-`execute` assertions are asserting a guarantee the architecture places elsewhere and should be retired. **I have not changed either the code or the tests** — but nobody needs to hunt a pause-contract regression, because there isn't one. **`executor-fast-mode-workflows.test.ts` (1 failure) — mechanism not isolated.** `visitedNodeIds` is `['review']` where the test expects `['start','review']`. Three probes failed to explain it: giving the review node an explicit `column: "in-progress"` changed nothing (so it is not column-based entry resolution), and swapping `seam: "review"` for a plain prompt config did not isolate it either. Two structurally identical sibling tests in the same file still pass with `['start', ...]`, so something in graph traversal distinguishes them that I did not find. That is U8/graph-executor territory; I am not asserting a `visitedNodeIds` shape I cannot explain. ## Still open and green - **#2637** — `task-delete-notice` 21 failed → 34 passed. - **#2643** — shellout allowlist re-pin. **Merge early:** it re-drifts whenever `executor.ts`/`self-healing.ts` line counts shift, with no git conflict to warn you. It already drifted once while open (`executor.ts:17106 → 17198`) and I re-pinned it. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6ed284f36a |
drop the dead semaphore parameter from dropPreHeldExecutorSlot (#2574)
Small follow-on to the cross-project cap removal. `dropPreHeldExecutorSlot(taskId, semaphore?)` released a cross-project semaphore slot. That semaphore is deleted, and **all 16 production call sites passed `this.options.semaphore`**, which nothing wires any more — so the release was a no-op on an always-undefined value: an optional parameter that reads as if it does something. ## What is *not* deleted Pre-held slots are **dual-purpose**: a cross-project semaphore slot **and** the FN-8453 per-project coordinator reservation. Only the first is gone. The reservation is the half that matters — every rejection path funnels through this helper so an early scheduler/triage return cannot permanently consume a project slot — and it stays. That is why this is a parameter change, not a helper deletion. Sites that still hold a semaphore reference release it **explicitly** next to their drop, so behaviour is unchanged for any caller that supplies one. Nothing wires one in production today, but silently leaking a slot for a caller that does is not a trade a cleanup is allowed to make. ## One real leak fixed — found by a failing test, not by reading `ProjectAdmissionCoordinator.admitOldest`’s release lambda took the pre-held branch and **returned**, relying on the deleted parameter to hand the host slot back. With the parameter gone, that branch unwound the registration and the reservation while **leaking the host slot** the attempt had acquired. The release is now unconditional across both branches. Worth noting how it surfaced: the test that caught it (`drops a declined candidate’s pre-held executor slot`) asserted `semaphore.activeCount`, which I had initially assumed was just coupling to the deleted half. It was not — it was pinning a real invariant. ## Tests Five cases in `concurrency.test.ts` pinned `sem.activeCount` through a drop. Each is re-pointed at the surviving contract — registration and reservation unwound, nothing left for a later pass to “take” — with the semaphore assertions moved to the sites that now own the release. ## Verification `pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green · `concurrency.test.ts` **56/56**. The 8 `triage.test.ts` failures are **pre-existing** — reproduced identically with this branch’s `triage.ts` replaced by main’s. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3aa942ee5f |
capacity: spawned agents count against the project agent count (#2579)
Two configurable numbers per project. `maxSpawnedAgentsPerParent` (5) and `maxSpawnedAgentsGlobal` (20) were a **third and fourth** limiter with private budgets invisible to both. ## This closes a hole, not just knobs A spawned child **is** an agent and gets **its own git worktree** (branched from the parent’s — the tool’s own description says so), but children were counted by **neither** capacity gate. A fan-out could put up to 20 extra worktrees on disk while the scheduler believed the project was at its configured limit. The operator’s two numbers were simply wrong about what was running. ## The old caps also measured the wrong thing `totalSpawnedCount` decrements on child cleanup, but the per-parent **set** is cleared only when the **parent task** ends. So `maxSpawnedAgentsPerParent` throttled *cumulative* spawns across a task’s life rather than *concurrent* ones — a long-running task could exhaust its budget with five children that had all long since finished, and the operator had no way to see why. ## Fix `fn_spawn_agent` gates on the same project agent count every other lane uses (`computeTopLevelConcurrencyClaimedFromStore`) plus live children. One number, one answer, no private budget that can disagree with the board. The refusal names **Max Concurrent Tasks** — a control the operator actually has. The old messages pointed at settings that no longer exist, which is worse than no message: it sends someone hunting for a knob that is not there. ## Verification **Revert-proof, measured:** restoring the private budgets turns **3 of the 4** new cases red — a project at 1/1 could still spawn, which is precisely the hole. `executor.ts` restored byte-identical. `pnpm lint` clean · core + engine `tsc` clean · `pnpm test:gate` green (414 + 10 + 71) · new suite 4/4 · `settings-default-descriptions` 4/4. There was no spawn-capacity test before this; the file is new. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Spawned agents now count toward the project’s **Max Concurrent Tasks** capacity. * Agent spawning is blocked when capacity is reached, including concurrent spawn attempts. * **Bug Fixes** * Prevented over-allocation during simultaneous agent spawns. * Restored available capacity when agent creation fails. * **Changes** * Removed separate per-parent and global spawned-agent limits. * Updated settings to reflect the revised capacity controls. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
152fedbd32 |
record the detector audit: gridlock and stuck-task are keeps, with evidence (#2581)
Answering the review question *“does gridlock detection still have a job?”* — with evidence rather than assumption, and recording it so the question is not re-opened by someone reading the name. **No behaviour change.** Comments only. ## Gridlock detector — KEEP `GridlockEvent.reasons` is typed `"dependency" | "overlap"`. It detects **dependency deadlock** and **file-scope overlap deadlock** via the scheduler’s `pathsOverlap` / `filterPathsByIgnoreList`. That has nothing to do with limiters arbitrating against each other — two tasks can still block on a dependency cycle or a shared file scope no matter how many agents the operator allows. The hypothesis that gridlock ≈ competing limiters deadlocking was reasonable from the name, and wrong. ## Stuck-task detector — KEEP Detects a stuck **agent** — a live session repeating the same tool call, or emitting no activity signal — via tool fingerprints and inactivity windows. Orthogonal to how many agents may run: a single agent on an unlimited board can still wedge. ## Evidence Measured for both: **zero** references to `maxConcurrent` / `maxWorktrees` / `semaphore` / `capacity` / `slot`. Both are live and wired — gridlock via `project-engine.ts → notifier.notifyGridlock`, stuck-task via `in-process-runtime.ts`. The note lives in each file because the natural reading of “gridlock” is “limiters deadlocking”, and deleting a live detector on that reading would remove real coverage silently. Each note states the question a future cleanup should actually ask — *is dependency/overlap deadlock still possible?* — rather than *is capacity simpler now?* `pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green · detector suites **108/108**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
15b21dead1 |
fix(dashboard): reconcile task state through live API (#2595)
## Summary - add a project-scoped live API route for updating individual task checklist steps - add an atomic live API route for resolving stale durable wedge episodes - prevent operator repair tooling from opening a second embedded store that can diverge from the running dashboard backend ## Why Legacy graph-native workflow runs can retain successful `workflowStepResults` while their narrative checklist remains at 0/N. The existing `fn task update` fallback may open a separate embedded store, producing split-brain writes that do not accumulate in the live dashboard backend. There was also no API surface for the existing atomic wedge-episode resolver. ## Verification - `pnpm exec vitest run src/routes/__tests__/register-task-workflow-routes.step-update.test.ts` — 5/5 passing - `pnpm build` in `packages/dashboard` — passing - full managed runtime workspace build — passing - deployed to the managed local runtime and used to reconcile six legacy review-deadlock tasks - live board audit: zero `in-review-stall-deadlock` paused reasons - exact local and Tailscale dashboard roots: HTTP 200 with 16,926-byte bodies <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added live API endpoints to update individual task checklist steps with validation (step index and allowed status values). * Added an endpoint to reconcile/resolve stale task “wedge” episodes, resolving only the matching active episode and returning conflicts on mismatches. * **Tests** * Expanded route tests for step updates and wedge resolution, including consistent 404 behavior for soft-deleted and missing tasks, plus conflict and invalid-input cases. * Expanded PostgreSQL coverage for wedge resolution persistence and concurrent episode replacement scenarios. * **Bug Fixes** * Improved task-lookup error handling so soft-deleted tasks are consistently treated as “not found” (HTTP 404). <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
0c07584d51 |
U11 fallout: disprove the coding-ideas column collapse, and correct a U11 note that recorded the merge backwards (#2651)
Two findings, no behavior change. Both are about **recorded reasoning that was wrong** — the kind that sends the next person the wrong way. ## 1. The coding-ideas column collapse does not work (IR change reverted) I implemented it — deleted `ideas`, moved its `intake`/`autoTriage: false` onto Planning, repointed the `start` anchor, updated the IR suites to the merged shape (they went green, 44/44). Then the wider suites failed and showed why it cannot work. **The manual gate IS the column boundary.** `replan-target.ts` names the discriminator in its own comment: *"The real discriminator is which lane the triage service SCANS, which depends on the intake column's `autoTriage` config."* So `ideas` is unscanned, `todo` is scanned, and "promote" means moving the card from one into the other. Merge them and one column must be both: | if… | consequence | |---|---| | `autoTriage: false` wins | never scanned → nothing is ever planned → the capacity hold releases an **unplanned** card into `in-progress`, violating FN-7648 | | scanning wins | `autoTriage: false` is meaningless → the manual gate is gone → the preset duplicates the default Coding workflow | **8 tests fail, and they are not fixtures** — they encode the promotion flow itself, e.g. `store-create-intake-column.test.ts` › *"promotes an Ideas-parked task to todo without planning it (still bootstrap-stub PROMPT.md)"*. Rewriting them would have meant inventing what "promote" means with no destination column, which is how a broken flow gets blessed by a green suite. **What it would actually take:** a promoted flag the triage scan reads, so one column can hold both "not yet promoted" and "being planned". That is a new lifecycle signal, not a column merge — the same shape as the deferred `needs-replan` follow-up. Happy to scope it. **I also corrected my own earlier checklist** in this doc, which said to delete the now-dead `isUnplannedStartCreate` arm. Wrong: `autoTriage` is a general trait field (`builtin-traits.ts`), so any custom workflow can declare a manual intake with `intake !== hold`. The arm is dead only for this preset. ## 2. `replan-target.ts` recorded the U11 merge backwards The note claimed U11 deletes `todo` and keeps `triage`. It is the reverse — Shape B kept the id `todo` and deleted `triage`, precisely so the ~120 `column === "todo"` guards kept their meaning and no data migration shipped. The default lineage now declares `todo, in-progress, in-review, done, archived`. The lookups are correct today, but **for the opposite reason to the one recorded**: the default lineage falls *through* the `triage` lookup and lands on `todo`, its merged planning column. `triage` still matches the workflows that genuinely declare it (Lead generation, PR review). Also flagged without changing (it would be a behavior change): the `return "triage"` fallbacks on the no-match and throw paths name a column the default lineage no longer declares, so a workflow with neither `triage` nor `todo` gets a nonexistent target. ## Census **Unchanged: 781 total, triage 5.** This PR adds no guards and converts none — `workflowHasColumn(ir, "triage")` is a call argument, not a comparison, so it is outside what the census counts either way. ## Verification 41/41 engine replan-target suites (including the existing `replan-target-merged-planning-column` suite that covers the corrected behavior) · engine typecheck clean · the reverted IR restores the tree to main's content for those three files, verified by `git checkout --`. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
22a66c3a51 |
fix(test): re-pin the blocking-shellout allowlist after source lines moved (5 lines, 0 new sanctions) (#2643)
**A failing ratchet, fixed without widening it.** 5-line diff, no
production change. `pnpm test:gate` green.
`engine-no-blocking-shellout.test.ts` was red on main, reporting **5
unaudited synchronous shellouts** (1 in `executor.ts`, 4 in
`self-healing.ts`).
## No new violation — the allowlist went stale
The allowlist is keyed by `(file, LINE, primitive, signature)`.
`self-healing.ts` shrank during the U4 extractions, so the recorded
lines drifted. Every flagged signature was **already sanctioned**:
self-healing's three sat at 4445/4451/4488 and are now at
4127/4133/4170. The file carries an FNXC note for precisely this case —
*"Re-pin all audited shellouts after current main moved source lines
without changing the sanctioned short-git-plumbing calls."*
**Re-pinned by signature, not by hand:** for each of the 33 entries,
find the line whose trimmed text equals the recorded signature and take
the occurrence closest to the old line — which stays stable when a
signature repeats, e.g. `merger.ts`'s six identical `git reset --merge`
calls. Result: **5 lines moved, 0 signatures unfound**, so nothing
became sanctioned that was not sanctioned before.
## A wrong turn worth recording
I first read the guard's *"only after proving timeout and maxBuffer
bounds"* as applying to every sanctioned site. I checked all 5, found
none had `timeout` or `maxBuffer`, and was about to add bounds to
`executor.ts` and `self-healing.ts` — an unnecessary production change
in two files other workers are actively editing.
Re-reading the guard's own comment corrected it: the bounds criterion
belongs to `BOUNDED_GIT_DIFF` (data-dependent diff output, where the
buffer can grow with the repo), not to `SHORT_GIT_PLUMBING`. All 5 are
`rev-parse` / `merge-base --is-ancestor` / `rev-list --count` / `branch
--list` — fixed-size output, already the sanctioned category.
## Verified the ratchet still bites
A re-pin could silently widen a guard, so I checked rather than assumed:
injecting an unaudited `execSync("git log --all")` into
`integration-branch.ts` **fails** the ratchet and it names the offending
signature; reverting returns it to green.
## Why this one was worth taking
A red ratchet is the worst failure mode for a guard — it stops being a
signal and starts being noise someone silences. This one had already
caught a real class of defect (unbounded sync shellouts on the shared
event loop), and it was red for a purely mechanical reason.
`workflow-lifecycle-live-e2e` (#2634, merged) and
`executor-review-verdicts` (#2641) clear two more of main's 13 failing
engine-default files; this is a third.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8d3b8262c0 |
U2b: the second move-path divergence — legal targets differ, not just message shape (blocks the useWorkflow flip) (#2638)
Tests only. Advances U2b's equivalence proof **without touching `moves.ts`**, whose edit order is still being agreed between U12 and MAIN. ## Why this one decides the sequencing The equivalence suite already records one divergence: rejection **type and message** differ. That is a shape difference and easy to absorb. This second one is a difference in **which moves are legal**, and it is workflow-dependent. U11 removed `triage` from the default coding lineage, so rows left there sit in a column their own workflow no longer declares. #2515 added an escape hatch to `resolveAllowedColumns` so such a card has a legal move — its workflow's rebound target — instead of `Valid targets: none`. **That hatch lives inside the `useWorkflow` block, so it only runs on the hooks path.** Mutation-verified in #2597: stubbing it back to `[]` left an operator-move test green, because the inline path answers from the legacy `VALID_TRANSITIONS` map instead, whose `triage` row happens to permit the move for unrelated reasons. So **flipping `useWorkflow` changes move validation for every stranded card**, not just side-effect routing. ## What that means for "flip the flag, then delete the flag-OFF branch" That plan is not the mechanical cleanup it looks like, and KTD-6's Phase A escalation correction already ruled on this exact shape once: > Deleting the inline branch would have swapped every project onto an untravelled code path and called it a cleanup. > > The convergence is its own unit with an equivalence proof (**U2b**), and it **blocks Phase B**. Nothing downstream may assume the trait-hook path runs until it lands. Flipping the flag and deleting the other branch *is* that deletion, reached from the other side. The tests won't catch a divergence because both paths have tests and only one of them runs — which is precisely why U2b was scoped as a proof rather than a refactor. **Recommended order:** U2b's equivalence proof completes → U12 flips `useWorkflow` → the flag-OFF branch and its 5 guards go away wholesale. That still gets the "strictly less work" outcome, just after the proof instead of instead of it. ## A note on how this is asserted Deliberately a **positive assertion about the inline path**, not a comparison of two target lists. The two paths do not reach the same rejection — inline enumerates the legacy table and reports it in the message; hooks throws the typed unknown-column rejection first. That asymmetry **is** the divergence. Comparing two lists would hide it behind two empty arrays and read as equivalence, which is the failure mode this suite exists to prevent. 11 tests green; lint clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
14c73ab727 |
U11 [tool-availability + skill-resolver + cli/task]: name the 3 literals that are NOT columns (48 -> 45) (#2619)
**Taking: `engine/tool-availability.ts`, `engine/skill-resolver.ts`, `cli/commands/task.ts`** — the three census hits that are not board columns. ## Census | file | before | after | |---|---:|---:| | `packages/engine/src/tool-availability.ts` | 1 | **0** | | `packages/engine/src/skill-resolver.ts` | 1 | **0** | | `packages/cli/src/commands/task.ts` | 1 | **0** | | **repo total (comment-stripped)** | **48** | **45** | ## These are not lifecycle guards — converting them would have been wrong - **`tool-availability`** — `surface: "triage" | "executor"` is an **agent lane**. The lane that writes specs keeps its name whatever the board calls its planning column. Resolving it from a workflow IR would make an agent's prompt depend on board configuration. - **`skill-resolver`** — `sessionPurpose === "triage"` is an **agent role**. Same argument: a role doesn't move when a board renames a column. - **`cli task list`** — the glyph chain distinguished **active** columns from the rest and nothing else; all four active ids mapped to the same `●`. Each is now named (`AgentResearchSurface`, `ROLE_FALLBACK_SESSION_PURPOSES`, `ACTIVE_COLUMN_GLYPH_IDS`) so the next person working the census sees at a glance that they're out of scope, rather than re-deriving it as I had to. ## A real divergence my own equivalence test caught I first wrote the glyph as the tempting inverse: ```ts const dot = col === "done" || col === "archived" ? "○" : "●"; ``` That is equivalent across all six lifecycle ids and **not** equivalent for anything else — the original chain fell through to `"○"` for an unrecognised id, while the inverse renders it as **active**. The loop only walks the six `COLUMNS` today, so nothing would have caught it in practice; a renamed workflow reaching this code later would have silently changed how its columns render. Shipped as an explicit ACTIVE set that mirrors the fallthrough exactly. The test asserts equivalence over the six ids **and** over unknown ids, which is where the difference lives. That's the point of testing a "pure rename" at its edges rather than only where it's currently exercised. ## Verification - 71 tests green across skill-resolver / heartbeat-skills / tool-availability / the new equivalence suite - merge gate green (482 + 132 + 10), engine + CLI tsc clean, lint clean ## Note on the remaining count Of the 45 left, `replan-target.ts` (2), `board-workflows.ts` (2) and `archive-planning.ts` (1) show up in a **raw** grep but are **0** real — every hit is inside a comment. A raw grep reports 53; comment-stripped is 45. Real remaining work concentrates in `self-healing.ts` (11), `register-task-workflow-routes.ts` (6), and the parked `moves.ts` / `default-workflow-hooks.ts` (9). No changeset: `@fusion/engine` is private; the CLI change is display-identical. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --- ## ⚠️ Read before merging — these are ROLE renames, not column conversions The coordinator's hand classification says several literals in this PR "must be left exactly as they are" because they compare an **agent role**, not a task column, and resolving them to a column trait would be a bug. **I agree, and this PR does not do that.** What it does: replaces a bare `=== "triage"` with a **named role predicate** — `isPlanningAgentLane`, `AgentResearchSurface`, `ROLE_FALLBACK_SESSION_PURPOSES`. Behaviour is **byte-identical** for every input. No IR is consulted, no trait is resolved, no column is involved. The reason to keep it rather than revert: the danger isn't the literal, it's that nothing at the call site tells the next person `"triage"` here means a *lane*. A list of exceptions maintained elsewhere only helps someone who finds the list; a call named `isPlanningAgentLane` helps whoever is reading the line. It also shrinks what the #2630 ratchet's ignore list has to carry. Reversible: if the preference is to leave the literals untouched, say so and I'll strip these hunks — but then the ratchet's ignore list must carry **all twelve** role sites or it can never reach zero, because those six are correct code. ## Classification finding Bucketing by the **receiver** of the comparison (not the literal) mechanically separates guards from roles, and it found **six role sites currently listed as "real column guards"**: | site | receiver | what it actually is | |---|---|---| | `usage-limit-detector.ts:144, 207` | `agentType` | agent type — the column test one line above is *already* trait-driven | | `skill-resolver.ts:432` | `sessionPurpose` | session purpose | | `tool-availability.ts:32` | `surface` | agent surface (`"triage" \| "executor"`) | | `effective-model-resolution.ts:148` | `entry.agent` | agent-log lane | | `useTasks.ts:162` | `entry.agent` | agent-log lane | So the real bar is roughly **39**, not 45. The rule that found all twelve without judgement calls: `column`/`toColumn`/`taskColumn`/`c` are guards; `role`/`agent`/`agentType`/`surface`/`sessionPurpose` are not. Worth teaching #2630's ratchet directly. |
||
|
|
f91b8a4178 |
TAKING mission-feature-sync.ts: roadmap reconciliation resolves lifecycle roles (unowned drift site) (#2602)
> **Taking `packages/engine/src/mission-feature-sync.ts`** from the shared backlog — announced in the title per the collision protocol. Based on `main`, no dependencies. It is in **no unit's file list**: absent from the plan's per-file census *and* from the drift review's ownership split (self-healing, dashboard, triage/replan-target, core, executor). It is a planning-lane reader. ## What was broken `reconcileMissionFeatureState` maps a task's lifecycle **position** onto its mission feature's roadmap status, and read five column literals: `done`, `archived`, `in-progress`, `in-review`, `triage`/`todo`. On a renamed workflow **every branch answers "no"**, so the function collapses to a permanent `noop`. **What an operator sees:** a mission roadmap frozen at whatever status it last held, while the tasks underneath it run to completion. Nothing errors, nothing retries. Worse than a wrong status, because a stale roadmap reads as a stable one. ## Guard counts (per the reporting requirement) | Metric | Before | After | |---|---:|---:| | `column === / !== "triage"` in this file | **1** | **1** | | role comparisons converted | — | **5** | **The metric does not move here, and I am not claiming it does.** The five role comparisons are converted; the one literal that remains is the deliberate scoped migration acceptance this change *adds*. That is the third time on this program the real fix has been invisible to the convergence count — the count finds the site, it does not define done. Worth knowing while the shared backlog is being tracked by that number: repo-wide it currently reads **29** triage comparisons (including 4 in `plugins/`, which are also unowned). ## Fallback direction matters **Unresolvable workflow falls back to the legacy ids, not to `noop`.** A mission whose workflow cannot be read should keep tracking on the default vocabulary rather than go silent — going silent *is* the failure being fixed, so the fallback must not reproduce it. **The planner-lane branch also accepts an orphaned legacy id.** A pre-existing test asserted a card in `triage` returns its feature to `triaged`; that stopped holding for the default lineage after #2515 — the migration-window population again. Accepting `triage`/`todo` additively keeps those rows tracked, **scoped to ids the workflow does not declare**, for the reason greptile gave on #2593: a custom workflow may legitimately name its **review** lane `triage`, and mapping a card there to `triaged` would walk the roadmap backwards while the task is awaiting merge. ## A test of mine that proved nothing until fixed The scoping case first used a `triaged` feature. The planner-lane branch only fires for an **in-progress** feature, so the fixture fell through to the review branch and **passed under both implementations**. It discriminates only once the feature status lets the wrong branch win — verified by reverting the scoping and watching exactly that case fail. ## Revert proofs | Reverted | Result | |---|---| | all five literals restored | **5 of 15 fail** — every renamed case; every default case passes | | legacy acceptance unscoped | **1 of 15 fails** — the custom-`triage`-as-review case | ## Verification | Check | Result | |---|---| | new suite | 15/15 | | pre-existing mission-feature-sync + mission-autopilot + scheduler-trait-dispatch | 94/94, **no expectation edits** | | `tsc --noEmit` (engine) | clean | | `pnpm lint` | clean | | `pnpm test:gate` | green (482 + 10 + 71) | | `pnpm check:changesets` | clean | ## Next from the shared backlog Taking `plugins/fusion-plugin-even-realities-glasses/src/agent-actions.ts` (3) and `plugins/fusion-plugin-dependency-graph/src/GraphTaskNode.tsx` (1) next — 4 sites in `plugins/`, which no unit owns and which the #2587 ratchet now scans. Shout if anyone is already there. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
347107d8e4 |
Phase B — TaskContextMenu.tsx: intake by ROLE (2 → 1), and two conversions I dropped rather than force (#2626)
**Claimed:** `packages/dashboard/app/components/TaskContextMenu.tsx`
| file | before | after |
|---|---:|---:|
| `TaskContextMenu.tsx` | **2** | **1** |
## The real bug
`shouldShowActionsMenu: task.column !== "triage"` meant *"a bare card in
a pure intake lane has no actions worth showing yet."*
Post-U11 the literal does not go dead — it **inverts**. A default
Planning card is `todo`, so the condition is true and the menu shows
unconditionally. That is right for the hold half (cards waiting for
capacity do have actions), but the guard has stopped distinguishing
anything — and it would show a full action menu on a bare Coding (Ideas)
capture, which is the case it existed to suppress.
Resolved to `intake AND NOT hold` — a *pure* intake lane — which
reproduces all four shapes rather than picking a winner:
| workflow | column traits | menu |
|---|---|---|
| legacy `triage` | intake only | suppressed *(as before)* |
| legacy `todo` | hold only | shown *(as before)* |
| merged Planning | intake + hold | shown *(matches the Todo half, where
cards wait)* |
| Ideas `ideas` | intake only | suppressed *(a bare captured idea)* |
Its degraded arm now defers to `isIntakeColumnRole`, so the legacy
intake id lives in `utils/columnRoles.ts` only.
## The remaining site is audited, not overlooked
I routed `isPreExecutionHoldColumn` through
`isPreImplementationColumnRole` — same question, one definition — **and
then reverted it.** Its degraded-mode answer is wider: its legacy set is
`{todo, triage}`, this predicate's was `{triage}` alone.
They differ **for a reason.** That helper drives the preserve-progress
prompt, where a flagless `todo` *should* prompt because losing steps is
unrecoverable. This one drives the Plan affordance, where a flagless
`todo` must **not** offer to re-plan a card that may already be planned.
Consolidating added `plan` to flagless `todo` cards — caught by
*"exposes Plan only for pre-execution hold columns"*. Identical trait
path, non-interchangeable fallbacks. Kept separate with the difference
recorded rather than made to look shared.
## Two conversions I dropped rather than force
**1. A `ListView.tsx` 5 → 0 conversion.** Main changed underneath it:
the U12 worker centralized the same fallbacks into
`utils/columnRoles.ts`. Their approach is on main and other files
already call it, so I took theirs and dropped mine rather than fight for
my version through a rebase conflict.
**2. A `strandedColumnFlags.ts` seam** that resolved an undeclared
column's role from the workflow's **rebound target**, so the degraded
arms could be *deleted* rather than documented.
I built it, tested it, wired it into ListView — and then their
`columnRoles.ts` identified a state my seam cannot serve: the **pre-load
window**, where the board renders before the workflows fetch resolves
and there are no columns at all, hence no rebound target to borrow from.
Their analysis is more complete than mine, the fallback is genuinely
undeletable, and shipping an unused module is worse than shipping
nothing.
Worth recording because I twice reported these arms as permanently
unconvertible, then thought I had a way to convert them, and was wrong
for a reason worth knowing: **there are two degraded states, not one.**
The stranded-card half is resolvable; the pre-load half is not.
## Verification
10 of 11 green in this suite. The one failure — `"Back to in-progress"`
vs `"Back to In Progress"` — is **pre-existing**, verified by stashing
this change and re-running against clean `main`.
Dashboard app typecheck and lint clean.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1d0f21b428 |
U12 R12: lifecycle-column literal ratchet — and the raw count's floor is not zero (#2630)
The anti-regression ratchet U12 R12 calls for. Counts lifecycle-column **literal comparisons** in production source and fails when the count rises. ## Two jobs **1. Ratchet.** A converted guard cannot silently come back as a literal. Ceilings only go down. **Mutation-verified both ways:** adding one `t.column === "triage"` fails with *"rose to 49 (ceiling 48)"*; lowering the ceiling to 47 fails with *"rose to 48 (ceiling 47)"*. The number is exact, not approximately right. **2. Honest denominator — the finding.** `triage` is overloaded in this codebase: a column id, an **agent role**, a **session purpose**, a **prompt-template family**, and a **CLI glyph key**. A raw grep counts them together, which makes "reach zero" unreachable by construction — converting `role === "triage"` in `agent-prompts.ts` would break the planning agent's prompt-template resolution, and the failure would look nothing like a column bug. | | count | |---|---:| | raw `triage` matches | 72 | | **not a column at all** | **10** | | genuine column comparisons | **48** | The 10: `agent-prompts.ts` ×3 (`role`), `usage-limit-detector.ts` ×2 (`agentType`), `skill-resolver.ts` (`sessionPurpose`), `tool-availability.ts` (`surface`), `cli/commands/task.ts` (a glyph key), plus two in comments. Ceilings recorded for all four ids — **`in-progress` (133) and `in-review` (200) were untracked entirely.** ## A measurement error of mine that writing this caught A grep over `packages/<pkg>/src` **misses `packages/dashboard/app`**, where the board components live. That undercounted `triage` as 43 in an earlier audit of mine when it was 62. The source roots are now listed explicitly in code so the number cannot drift with someone's glob. ## The classifier is under test, not trusted A ratchet that matched nothing would pass forever while measuring nothing — the failure mode this program keeps finding. Three self-tests prevent it: - it asserts **positively** that `agent-prompts` / `skill-resolver` / `tool-availability` are excluded, so a classifier change that swallowed them would fail rather than quietly shrink the number; - it asserts the classifier still **matches** real column comparisons; - it pins `live-agent-count.ts`'s two sites as **permanent** no-flags fallbacks, with the reason, so a future edit that deletes them has to argue with it rather than silently drop stranded cards from the footer's queued total. The classifier keys on the **left-hand side naming a column** — deliberately syntactic, so a reader can audit it against the source without running anything, and conservative: an unrecognised shape counts **as** a column comparison, erring toward demanding conversion rather than excusing it. 7 tests green; lint clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bcb9782d3f |
fix(test): re-green task-delete-notice after the SQLite-arm deletion (21 → 0) (#2637)
**Unowned work, picked up.** `task-delete-notice.test.ts` was **21 failed / 13 passed** on main → now **34 passed**. Test-only. `pnpm test:gate` green, `pnpm lint` clean. ## Cause `deleteTaskImpl` and `deleteTaskIfImpl` are now **thin delegators**. The SQLite arms were deleted in the PG cutover (`FNXC:SqliteDualPathCleanup 2026-07-26`) and both forward unconditionally to `store.deleteTaskBackend` / `store.deleteTaskIf` — `deleteTaskImpl` is literally *"throw if self-delete; return `store.deleteTaskBackend()`"*, with **no `backendMode` branch left**. The suite drove them against a `makeSqliteStore` fake providing neither method, so every case threw `store.deleteTaskBackend is not a function` **before reaching any notice logic**. All 21 failures were measuring a crash, not a decision. ## Fixes - the PG fake gains `deleteTaskBackend` / `deleteTaskIf`, wired to the **real backend impls** rather than stubbed — so the delegating paths still prove the delegation preserves the notice decision instead of asserting against a mock; - plus `withTaskLock` (`deleteTaskIf` wraps the conditional delete in the per-task lock), running the body inline so the predicate and short-circuit paths execute for real; - every remaining `makeSqliteStore` call site retargeted, and the now-dead factory **deleted** so it cannot rot back in. ## Corrected a claim the file was making The header's Surface Enumeration said the three paths prove *"the behavior cannot depend on backend mode"*. **There is one backend now.** The enumeration is still worth driving — a caller reaching the public entry point must get the same notice as one reaching the backend directly, and these paths prove exactly that — but it is a different claim, and the file now states it, with the two paths renamed from `(SQLite)` to what they actually are. ## A bug in my own patch, caught by re-running My first edit inserted the method assignments **after** the `return`, so they were unreachable and the symptom didn't change. I only found it because the failure count stayed identical and I checked the file instead of assuming the edit had landed. Worth noting because "the patch applied" and "the patch took effect" are different facts, and this session has now produced three variants of that same mistake. ## Deliberately not fixed here `task-delete-caller-attribution` (13 failed) and `task-delete-nonblocking-cleanup` (2 failed) share the root cause, but their `makeDeleteStore` fake carries `backendMode: false` and lacks the PG surface the real backend impl needs (`asyncLayer` / `transactionImmediate` / `rowToTask` …). Wiring them means either building that surface out or re-pointing the suites at `deleteTaskBackendImpl` directly — a judgement about what those suites are *for*, and worth making deliberately rather than folding into this fix. Whoever owns the PG cutover cleanup will know which; the diagnosis above is the whole of it. Verified no collateral: `task-merge` and `legacy-adoption` unaffected (166 passed across the three files). 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9edc746f96 |
E2E evidence: the MERGED board (third completion criterion) — plus a RETRACTION of my #2613 escalation (#2632)
This is the merged-board half of the evidence assignment. **It is red, deliberately, and the red is the finding.** Do not merge it to make the red go away — the assertions are correct and `main` is broken. ## Escalation first: #2613 broke the default board, and the gate did not notice `6a33d8f8c` — *"Phase B — TAKING task-creation.ts: intake classification by trait (4 sites → 0)"* (#2613) — regressed four E2E cases, including **the default-vocabulary full lifecycle**, which is scenario 1 of the whole E2E assignment. Attribution is a clean single-file revert, not a guess: ``` HEAD (main): 4 failed | 39 passed HEAD with ONLY 6a33d8f8c's task-creation.ts reverted: 28 passed (both files fully green) ``` Failing: 1. `scenario 1 — DEFAULT vocabulary … persists the card in the expected column at every stage` 2. `scenario 2 — RENAMED vocabulary … writes the same column-transition audit trail as the default` 3. `releases a card out of the merged lane on capacity — the release is not a self-move` 4. `does not re-release a card that already left the merged lane` **`pnpm test:gate` is green on this branch — exit 0, 695 tests.** #2613 merged through a green gate, and its own tests pass. This is the eighth time this program a test has passed without exercising its subject, and the first one an E2E family caught rather than review. ### Mechanism `isIntakeColumn` in `task-creation.ts` decides whether a new card gets a **bootstrap** prompt (freeform, "triage will plan this later") or a **specified** prompt (planned, executable). #2613 rewrote it as: ```ts const isIntakeColumn = (intakeFacts.intake !== undefined && task.column === intakeFacts.intake) || … ``` where `intakeFacts.intake` falls back to the **default workflow's** intake when the create supplies no `workflowId`. Post-U11 the default workflow's intake **is `todo`**. So any card created directly in `todo` is now classified as intake and gets a bootstrap prompt — unplanned. Unplanned cards do not advance through the graph (no `NodeEntered` audit rows → failures 1 and 2) and hold-release will not release them (FN-7648: no unplanned card enters a processing column → failures 3 and 4). Before U11 this was safe: `triage` was intake and `todo` was a distinct lane, so creating in `todo` meant "planned work". The merge deleted that distinction. ### Why this is the exact trap you warned about You said you did not want *"a conversion that swaps the literal for a trait lookup WITHOUT checking what the guard was for."* The old `task.column === "triage"` guard meant **"is this card unplanned?"** On a merged board, intake-vs-hold **cannot answer that question at all** — one column is both. The distinguishing fact is not the column; it is whether the caller supplied a spec. Resolving the role faithfully still gets the wrong answer, because the question was never really about the column. Not fixing it from here: `task-creation.ts` is #2613's owner's file, and the fix is a design call about which fact replaces the column test. ## What the evidence itself adds Three families extended to the U11 shape — one column carrying **both** intake and hold. That breaks a class of guard renamed boards structurally cannot reveal: | shape | consequence | |---|---| | `intake && !hold` | **unsatisfiable** — silent | | hold → intake release | **self-move**, re-fires every poll — loops | | `intake && column !== "triage"` | inverts to **always-true** — silent | Two are silent and one loops, so every case sweeps **twice** and asserts no re-release; a single pass cannot tell a no-op from a self-move. ### A fixture that could not fail My first merged row used `MERGED_VOCAB`, which is *faithful* to U11 — it reuses the legacy ids, because that is what the default lineage has. That fidelity **destroyed its discriminating power**: its hold column *is* `todo`, so a guard falling back to the `todo` literal returns the same answer as one resolving the role. The "hold but not intake" mutation left all 23 green. Added `MERGED_RENAMED_VOCAB` — merged *structure*, renamed *vocabulary* — the only combination where the collapse is observable **and** the literal is wrong. Same mutation now fails exactly 1 of 23. Both vocabularies stay: one asks *"does the collapse break the release path"*, the other *"is the role actually resolved"*. One rebound mutation was **genuinely unobservable** rather than undetected — `hold` is also the first column in that fixture, so the fallback chain lands there regardless. Pointing rebound at `complete` instead fails 9 of 15. Recorded rather than papered over. ## For the CAPACITY worker before `self-healing.ts` is marked done `self-healing.ts:2952` and `:9134` query `listTasks({ column: "triage" })`. Converting the 10 guards leaves those sweeps **blind** — they never see a renamed card, so the guard is correct and unreachable. Query and guard convert together or not at all. There are **137** such `column: "<legacy id>"` sites repo-wide, 52 in that one file, and the 45→0 grep counts none of them because they are object properties, not comparisons. **The bar can reach zero with sweeps still unable to fire.** 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7ab6506c0f |
docs(solutions): proving a code path actually runs — the five ways U8 shipped code that never executed (#2642)
Durable write-up of U8's verification findings. **Docs only — no code change, no CI risk beyond lint.** These currently exist only in PR bodies, which nobody greps. `docs/solutions/` is where this project keeps exactly this kind of thing, and every one of the five will recur: the handler-pair shape and the resolved-vs-guessed fork both have more call sites than U8 touched. ## The five 1. **Two prompt-node handlers exist; only one runs.** `createDefaultNodeHandlers` prefers the primitives handler whenever `deps.primitives` is set, and `executeWorkflowGraph` always sets it — so every seam entry in `createAuthoritativeWorkflowSeams` is unreachable for prompt nodes. A lifecycle announcement sat there through two PRs. It type-checked and its unit tests passed, because a seam-level test calls the seam object directly and therefore always can. 2. **A negative instrumentation result is worthless without a control.** No output from an instrumented seam is only evidence once you have shown writes from that module are visible under the harness. One `process.stderr.write` at module load separates "never ran" from "output swallowed" — opposite conclusions. 3. **Source-string ratchets prove syntax, not behavior.** Three were torn down in review. The sharpest guarded a never-executed-code bug with a source search, reproducing the bug one level up; measured, the behavioural version fails an inverted dispatch and the textual one passes it. Includes the sub-rules paid for the hard way: use the AST not regex (a brace in a string truncated an extraction to 13 lines and every count read a *passing* zero), guard the guard, anchor by index rather than a character window. 4. **A green test on first try, on a path with no prior coverage, is a warning.** Two conversions were reverted in one day because their tests passed with the change reverted. Negative assertions succeed trivially when the method returns early — `recoverCompletedTask` has seven guards before the converted line, and the fixture has to satisfy all of them. 5. **A named workflow selection is not a resolved one.** Provenance cannot be inferred from the returned value, because a fallback IR and a valid id-less IR are structurally identical — the resolver that knows has to report it. This is the fork every remaining lifecycle-column conversion hits. ## Why this rather than another conversion Everything left in my area is now owned and further along than I could take it: `executor.ts` → #2628 (which solved the `recoverCompletedTask` fixture I could not), `self-healing.ts` → #2560 (independently hit all three traps I catalogued), the dashboard cluster → #2625/#2626/#2636. Duplicating that would be motion, not progress. Turning findings that cost real cycles into something greppable is the useful thing I can still add. `pnpm lint` clean. No changeset — internal documentation. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a best-practices guide for verifying that workflow code paths actually execute. * Covers reliable behavioral assertions, instrumentation controls, regression-proof tests, source validation, and detection of fallback behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7397cea2dc |
U12: pin the move-path flag blast radius (6 seams, not 1) before flipping it (#2639)
Per the sequencing agreed in-thread — **U12 resolves the flag first, then the flag-OFF branch is deleted wholesale** — this is the step before the flip, not the flip. ## The flag is six switches, not one `moves.ts:363` reads the raw compatibility flag nothing in production source writes, and gates the hottest lifecycle path in the system. Every summary so far has under-scoped it, mine included: I described it as the `789`/`837` pair. Measured, it is six decision points: | seam | what flipping turns on | |---|---| | 392 | resolves the task's workflow IR — `undefined` when off, so **every IR-dependent guard below is inert** | | 489 | typed **rejections**: unknown-column and adjacency validation | | 789 | column side effects route through the trait hooks instead of the inline legacy block (timing, reset-on-entry, abort-on-exit, `merge.onEnter`) | | 1092 | writes the transition-pending marker that **capacity counting reads** | | 1330 | runs **plugin hooks** on column change | | 1395 | records `workflowId` on the emitted move payload | ## The risk is seam 2, and it is not an equivalence question With the flag off there is **no target-column validation on the move path at all**. Flipping introduces new refusals for moves that succeed today, on the path every engine lane uses. That is not "do the two implementations agree" — it is new behaviour, and a green suite is not evidence about it. `recoveryRehome` already carves out legacy targets (#1411); nothing proves the other callers are covered. ## What this PR asserts - **The seam count.** Mutation-checked, not assumed: replacing one gate with `if (true)` fails with `expected 5 to be 6`. (My first mutation attempt silently didn't apply — `str.replace` with no assert — so the anchor is verified now.) - **All six read ONE flag.** If a seam were rewritten to consult settings directly, a flip would move five behaviours and leave one behind, and nothing else in the suite would notice because both states are individually valid. - **The flag-OFF branch is still inline**, so the delete-with-the-branch step has a test naming the plan if someone converts its guards instead. - **The atomically-coupled second reader is named.** `workflow-task-create-ops.ts` computes the `movePolicyPreflight` that `moves.ts` consumes, so un-gating either alone either evaluates workflow move policies whose result is ignored, or validates against a preflight never computed. Comments cannot inflate the count — it is an AST walk. Parse failure fails loudly via `parseDiagnostics` rather than a try/catch, since `createSourceFile` is error-tolerant and a partial tree would undercount and read as "seams were removed". ## Preconditions for the flip, recorded in the file header 1. An equivalence proof for seam 3 across timing, reset-on-entry, abort-on-exit and `merge.onEnter`. **Neither implementation is the observed baseline** — they have never both run in production. 2. A census of the moves seam 2 would newly reject. 3. Both raw-flag readers flipped atomically. ## Why not just flip it here Because I cannot honestly claim the equivalence proof from a source read, and the flip is on every task move. Landing the blast radius as a test first means the flip PR has something to be proven against, and it means the next person cannot under-scope it the way this has been under-scoped four times. ## Verification `pnpm lint` clean, `pnpm test:gate` green (10 / 482 / 71), `tsc -p packages/core/tsconfig.json` clean, suite 5/5. No changeset: test-only, no behaviour change. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
50ebf3c543 |
TAKING cli/project.ts (fn project reported 0 running agents) + two test fixes — dashboard conversions WITHDRAWN in favour of #2626 and #2636 (#2631)
Three app-cluster conversions plus the evidence that they behave on a renamed AND a merged board. ## Per-file guard counts | file | before | after | note | |---|---|---|---| | `packages/cli/src/commands/project.ts` | 0 | 0 | not a comparison site — see below | | `packages/dashboard/app/components/TaskContextMenu.tsx` | 2 | 2 | **count does not move — deliberate, see below** | | `packages/dashboard/app/components/Column.tsx` | 2 | 2 | **count does not move — deliberate, see below** | **Read this before scoring the PR against the bar.** You said a claim that does not move your number is not done, so I am telling you up front that *this PR does not move it*, and why. Both dashboard conversions are **fallback-preserving**: ```ts const isIntakeColumn = columnFlags ? columnFlags.intake === true : column === "triage"; ``` The literal survives as the no-flags branch, so the grep still counts it. That is the shape the sibling code already uses (`isPreExecutionHoldColumn`, same file, converted earlier in the program), and dropping the fallback would make an unresolved-column render *lose* the affordance a second way. What changes is the **behaviour when flags exist** — which is what the mutation results below measure. If you want these to zero out the count, the fallback has to go, and that is a separate decision about whether an unresolved column should fail open or closed. Say the word and I will do it as a follow-up; I did not make that call unilaterally because it is not reversible from a rendering standpoint. `cli/project.ts` was never a comparison site at all — it fed **raw rows** to `isRunningAgentTaskShape`, so the helper's own internal legacy fallback kicked in and `fn project` reported **0 running agents** on any renamed board. Fixed by resolving the IR per task before counting. Nothing to subtract. ## Two of the three had a test that looked like coverage and was not - **`Column.tsx`** — the quick-create gate is `workflowMode || isIntakeColumn`. Every pre-existing intake case in `Column.test.tsx` *also* passes `workflowMode`, so the `||` short-circuited and **none of them ever reached the trait lookup**. Added cases that omit `workflowMode`, the only path where the conversion changes the answer. - **`TaskContextMenu.tsx`** — the intake suppression was asserted only for the legacy `triage` id, the one board shape where a broken conversion still returns the right answer. Mutation-verified rather than asserted: | mutation | result | |---|---| | `isIntakeColumn` → `column === "triage"` | **2 of 88 fail** (exactly the renamed and merged cases) | | menu suppression → `task.column !== "triage"` | **1 of 12 fail** | ## A pre-existing red I fixed on the way past `uses VALID_TRANSITIONS and in-review back-to-progress labels` was **already failing on origin/main**. #2521 correctly moved the "Back to X" label onto the host's `columnLabel` function; this file's stub is `(column) => column`, so the hardcoded `"Back to In Progress"` expectation was left over from the pre-#2521 hardcode and nothing had updated it. Matching the raw id would have made it pass while proving nothing, so instead that one case gets a display-like label function — the assertion now fails both if the "Back to" prefix regresses **and** if the label stops routing through `columnLabel`. Strengthened, not relaxed. Counts against completion criterion #2. ## Verification - `Column.test.tsx` + `TaskContextMenu.test.tsx`: **100 passed** - `tsc -p tsconfig.app.json` (the root config does not cover `app/`) and the CLI typecheck: clean - `pnpm test:gate`: green 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8e211d1870 |
TAKING scripts/: parse instead of grep — an AST classifier for the lifecycle-column bar, cross-checked by a second implementation (#2633)
The program's completion bar is "`column === "triage"` reaches zero".
This measures what that bar actually covers, and checks the measurement
in so it cannot drift.
## The number, measured by the checked-in tool
```
lifecycle-column-census: scanned 1956 source files
COLUMN guards (the backlog): 1031
ROLE comparisons (not guards): 10
DELIBERATE-LITERAL (reviewed): 4
by column id:
313 done
217 in-review
201 in-progress
177 archived
83 todo
40 triage
top files:
151 packages/engine/src/executor.ts
136 packages/engine/src/self-healing.ts
50 packages/dashboard/app/components/TaskCard.tsx
44 packages/core/src/task-store/moves.ts
34 packages/dashboard/app/components/TaskDetailModal.tsx
```
**`triage` is under 4% of the class.** Every one of those 1031 sites is
the same defect: a lifecycle decision made by column NAME, which stops
matching the moment a board renames a column. The bar can be met in full
while 991 identical guards remain — and two files hold a quarter of
them.
## The tracked count is wrong in three directions at once
Each of these cost real work this week, which is why this is a PR and
not a comment.
1. **Vocabulary.** It measures one of six legacy ids.
2. **Receiver.** It is anchored on locals named
`column`/`toColumn`/`fromColumn`, so it never saw the three real guards
in `executor.ts` written against `from` and `originColumn`. One of those
meant completed-but-stranded work was never recovered on a renamed
board, with nothing else owning that state (converted in #2628).
3. **Collision.** `role === "triage"`, `agentType === "triage"`,
`entry.agent === "triage"` compare an **AGENT ROLE**. The planner *lane*
is named `triage` and keeps that name — U11 removed the *column*. Ten
such sites were counted as backlog, and the "obvious" fix (renaming the
role) silently empties the planner's prompt template and mis-binds its
model markers.
A count that is too high and too low simultaneously sends work to the
wrong files while hiding the files that need it. So the census reports
**three separate numbers** and never nets them.
## Proven to fail on the original defect
Not asserted — exercised:
```
$ # reintroduce `task.column === "triage" || task.column === "todo"` into live-agent-count.ts
$ node scripts/lifecycle-column-census.mjs --strict; echo "exit=$?"
packages/core/src/live-agent-count.ts: 10 -> 12
exit=1
$ # restore the file
$ node scripts/lifecycle-column-census.mjs --strict >/dev/null; echo "exit=$?"
exit=0
```
The CLI also exits 1 when its own file list comes back empty — a guard
that reports success without checking anything is worse than no guard.
## 12 regression cases, split by what they defend
Must catch: all six ids; a guard on a local named `from`/`originColumn`
(verbatim the executor.ts shape); single quotes; negation; several
comparisons on one line.
Must **not** catch: role comparisons; comment prose (two tracked
"guards" in `replan-target.ts` were prose about a filter that lives in
another file); a trailing `// … === "triage"` on a code line; sites
carrying a `DELIBERATE-LITERAL` marker.
Plus: **one marker cannot launder a distant guard in the same file** —
that is how allowlists rot.
## Report-only, deliberately
`--strict` compares per-file counts against
`scripts/lib/lifecycle-column-census-baseline.json` and fails when any
file's count **rises**. It is **not** wired into the merge gate: a
thousand-site backlog cannot be a blocking check the day it is first
measured, and a guard nobody can pass is a guard everyone disables.
Owners tightening their own area re-record the baseline in the PR that
lowers it. This is the ratchet shape the `DELIBERATE-LITERAL` markers
scattered through the program already anticipate.
## Stated limitation
Classification is by receiver **name**, so a future field named `agent`
that holds a column would be misclassified as a role comparison.
Recorded at the site, and it is precisely why the two classes are
reported separately instead of netted into one figure.
## Verification
- 12/12 new cases
(`packages/engine/src/__tests__/lifecycle-column-census.test.ts`)
- `pnpm test:gate` **71/71**; `pnpm lint` clean
- `pnpm census:lifecycle-columns`, `--json`, and `--strict` all
exercised end to end
- documented in `docs/testing.md`; no production code touched
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6d10683dbd |
docs(solutions): store fakes that lie — six fixture defects that each looked like a production bug (#2534)
Six consecutive slices of U7 produced **six test-fixture defects, and every one first presented as a bug in the code under test.** Not one was real. Each cost 15–60 minutes debugging the wrong file. **Two would have shipped a false green** — a test passing while asserting nothing — if the failure had happened to look plausible rather than implausible. This is not a story about carelessness. Every one of these fakes was modelled on an existing fixture in this repo, and the repo's fixtures are inconsistent about exactly the things that matter. ## The catalogue | # | Defect | How it presented | Real cause | |---|---|---|---| | 1 | `moveTaskIf` ignores its predicate | Test passed; in-txn guard untested and indistinguishable from absent | Fake never invoked the callback | | 2 | `updateTaskAtomic: vi.fn()` never invokes its callback | *Every* finalize bailed before the branch under test | Success is derived from whether the callback ran | | 3 | Harness default parameter swallows the input | "Task vanished" case became a duplicate of the control | `harness(undefined)` triggers the default | | 4 | `logEntry: vi.fn()` returns `undefined` | Sweep appeared to match only one column | `.catch` on a non-promise throws, aborting the loop after item one | | 5 | Harness lets `poll()` reach the real `specifyTask` | **exit 1 with every test green** | Real agent path threw *asynchronously*, after assertions passed | | 6 | `updateTask: vi.fn()` returns `undefined` | Branch "did not run" | Same as #4 | **4 and 6 are the same shape, found a week apart, because nothing prevented the second.** That is the argument for writing this down. ## The three rules 1. **Every store method a fake exposes returns what the real one returns** — overwhelmingly a promise. Production writes `await store.m(...).catch(h)` as a fail-soft idiom; `.catch` on `undefined` throws a `TypeError` that unwinds into a broad *"never let housekeeping break the poll"* handler and vanishes. Symptom is never "your fake is wrong" — it is *"the loop only processed the first item"*. 2. **A fake handed a predicate or callback must invoke it.** Ignoring it makes the guarded and unguarded implementations *indistinguishable*, so a test named for the guard cannot detect the guard's removal. Includes the `onLockedRead` hook, without which an in-transaction recheck stays untestable even once the predicate is invoked. 3. **Stub the agent-dispatch boundary.** `poll()` ends in "start an agent", which in a unit test throws *after* the test resolved — `17 passed`, exit code 1, which on CI reads as infrastructure noise. > Never accept a non-zero exit on a green run. It is the only signal that something escaped your assertions entirely. ## Also covered - **How to spot a fixture defect fast** — the tell is *failing for the wrong reason*. Three concrete checks before you open the production file. - **Why differential tests earn their keep** even when they feel redundant: the default-vocabulary half doubles as a fixture self-check, because it asserts behavior that is by definition already shipping. On this program, "both halves failed" was the signal that found three of the six. - **The connection to guards that cannot fire** — six of those on this program too, including a ratchet I wrote that matched only a double-quoted literal (#2527). Same discipline either way: *prove the check fails on the thing it claims to catch before trusting that it passes.* Including the warning that one ratchet injection silently failed to apply, leaving a green run that would have "proven" the ratchet worked. ## The concrete next step, stated plainly A shared `createTaskStoreFake({ tasks, workflowIr })` with promise-resolving, callback-invoking defaults would remove this whole class in one small PR. **It is not built here** because it is cross-unit and needs adopters — building it inside U7 and hoping others find it is how conventions die. The doc says: if you are about to hand-roll a seventh store fake, build the helper instead and link it. Docs-only; no changeset (AGENTS.md excludes internal docs). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added guidance on six store-fake defect patterns that can resemble production bugs during testing. * Documented best practices for creating reliable store fakes, including promise handling, callback invocation, and async dispatch isolation. * Added diagnostic techniques for distinguishing fixture issues from genuine application defects. * Included guidance for validating production guards and links to related documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26c82ebc18 |
ratchet the planner-liveness gate so a fourth door fails CI (FN-6756) (#2540)
Test-only follow-up to the merged P0 (#2531). No production change, no changeset (internal). ## Why This bug reached users **three times**, each as the same mistake in a new place: | | What happened | |---|---| | FN-8600 | the reclaim sweep removed a worktree a live **planner** was using — fixed by registering planning paths and teaching *that* sweep `isPathActive` | | FN-6756 | the leaked-slot reaper never got the same signal; its last line of defense computed liveness from four TaskExecutor-owned maps, so a triage planner matched none of them | | (same PR) | fixing that was not enough — `recoverPausedAbortFailures` **discarded** the refusal and still logged `"Auto-recovered…"`, audited and counted it. The whole bug again, while reporting success | The shared cause is not any one sweep: **“liveness” was re-derived per call site**, so closing one door left the next open and nothing failed. Every one of those fixes was found by review, not by CI. This makes the next one a CI failure. ## Four properties, each written to fail on the exact defect that got through 1. **Every `clearPhantomExecutorBinding?.(` call site consumes its return** — a bare expression statement (including `void`/`await`-prefixed) is the signature of the pause-abort defect. 2. **The destructive path delegates to `hasLiveSessionSurface`** rather than inlining the session-map disjunction — a second copy can drift from the one callers gate on, which is precisely how each sweep got “fixed” without fixing the next. 3. **The probe is wired** in `in-process-runtime`. `self-healing.ts` already records `releaseExecutorWorktreeOwnership` as a declared-but-never-wired option that silently no-opped; an unwired *probe* is worse, since `?.() === true` is `false` when unwired and every gate would quietly stop deferring with nothing failing. 4. **The probe counts registered session paths**, not just executor maps — a triage planner appears in no executor-owned map, so that term is the only thing that sees it. Grep-level, comment-stripped, production source only; no engine boot and no fixtures (FN-5048). Fails closed on an empty/moved source file so a rename cannot make it silently check nothing. ## Proven, one injection at a time **The first draft of property 1 was worthless** — its filter chain was convoluted enough to discard every candidate, so the injected bare call passed. Caught by actually running the injection instead of trusting the green, and rewritten as a single “is this a bare expression statement” rule. | Injection | Result | |---|---| | discard the return value | fails, naming the call site | | re-derive liveness inline | fails on the delegation assertion | | unwire the probe | fails, naming `in-process-runtime` | | drop the registry term | fails, naming `activeSessionRegistry` | Clean tree passes 4/4; all three sources restored byte-identical (`git status` shows only the new file). **Verified:** `pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (414 + 10 + 71). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added safeguards to ensure liveness checks remain consistently enforced. * Verified phantom executor cleanup uses shared session-liveness detection. * Added coverage for registered session paths to prevent false inactive states. * Added fail-closed checks when required runtime source is unavailable. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |