Commit Graph

12521 Commits

Author SHA1 Message Date
gsxdsm
dfbab18fd5 fix(scheduler): a renamed wip column held NO file-scope lease — two agents could edit the same files (#2693)
## The defect

`activeScopes` is the file-scope lease registry the dispatch path reads
(`scheduler.ts:2167`) to decide whether a candidate overlaps work
already in flight. Two column-id literals kept it empty on any board
whose columns are renamed:

1. the lease loop gated on `task.column !== "in-progress"`;
2. `shouldHoldActiveFileScopeLease` keyed **both** its branches on
`in-progress` / `in-review`, so it returned `false` for *every* card on
a renamed board.

Forty lines above that loop, the same sweep resolves `countsTowardWip`
from the workflow IR for capacity arithmetic. **The scheduler was
simultaneously right about capacity and wrong about leases.**

Consequence: a second task sharing a file scope **dispatched instead of
queueing** — two agents editing the same files, which is precisely what
`groupOverlappingFiles` exists to prevent.

## Measured, differential

Same workflow *shape* under two vocabularies with identical traits; only
the column ids differ, so any difference is attributable to a surviving
literal. No renamed id collides with a legacy one, so a surviving `===
"in-progress"` cannot pass by luck.

| | default vocabulary (control) | renamed vocabulary |
|---|---|---|
| fix reverted | queued on lease ✓ | **dispatched into the wip column**
✗ |
| fix applied | queued on lease ✓ | queued on lease ✓ |

`2 of 3 fail` reverted → `3 of 3 pass` applied. The control passes on
**both** sides, so a change that breaks overlap protection generally
cannot hide behind this test.

I checked the test wasn't vacuous before trusting it: instrumented the
run to print the actual `moveTask` calls, and confirmed the renamed case
really produced `[["FN-CAND","building"]]` — a genuine dispatch — rather
than the candidate simply never being considered. Both failure modes
look identical in the assertion.

## Why optional booleans, not a flags object

`shouldHoldActiveFileScopeLease` is **exported** and shared with the
self-healing / repair paths (`self-healing.ts:4488`, `:5406`) — its own
comment says those "must use this same predicate so stale
`overlapBlockedBy` cleanup does not preserve blockers the scheduler
would ignore". So the role questions became optional parameters that
**default to today's literals**: a caller that resolved the traits
passes the answer, a caller that has not gets exactly current behaviour.
No existing call site changes meaning, and no dependency on #2690.

## Verification

| Check | Result |
|---|---|
| scheduler / capacity / hold-release / overlap / self-healing | **59
test files green** |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint`, engine `tsc --noEmit` | clean |

`self-healing-advanced-triage`, `-agent-link-drift`,
`-starved-refinement` are **7 failed / 19 passed both before and after**
— verified pre-existing on clean `origin/main` by reverting only
`scheduler.ts` and re-running. Flagged, not fixed, and not in scope
here.

## Census

**722 → 721**, `scheduler.ts` 28 → 27. Baseline re-recorded in the same
commit.

To be precise about what that −1 is: the *loop* literal is gone, while
the two literals **inside** the predicate remain by design as the
documented defaults. So this is not "scheduler is now trait-aware" — it
is one site, plus the seam that lets callers be.

## Merge-order note

**#2690 also records `scheduler.ts` 28 → 27**, converting a *different*
site (`isWipColumnTask`'s hand-rolled flags-first copy, `:1690`). The
two are independent and do not double-count: if both land,
`scheduler.ts` is **26**, and whichever merges second will conflict on
`scripts/lib/lifecycle-column-census-baseline.json` and must re-record
to 26 rather than resolve to 27. Flagging so the merger does not take
one side blindly.

## Still broken, flagged for an owner

The **in-review** half. `activeScopes` is also populated for review-lane
cards via `t.column === "in-review"` (`scheduler.ts:1751`, `:1757`), and
this PR leaves those literals in place: the sweep's flags map holds only
`countsTowardWip`, so no review-role answer is available to pass in.
Fixing it needs the flags-object change in #2690, after which the same
optional parameter added here carries it. Until then a renamed review
column still holds no lease.
2026-07-30 03:19:02 -07:00
gsxdsm
e9e63d8e0f consolidate/capacity: --strict was red on main (my #2621), 14 stale baselines, routines seeding a deleted column, worktrees-off audit (#2652)
Capacity unit consolidation. Three coherent themes, small commits
inside.

## Census before/after (`node scripts/lifecycle-column-census.mjs`)

| | before | after |
|---|---:|---:|
| triage column guards (the bar) | 10 | **10** |
| `--strict` on main | ❌ **RED** | ✅ green |
| baseline staleness | 14 files stale | **0** |

This branch does **not** move the triage bar — its remaining 10 are
moves.ts (dies with the flag), the dashboard cluster, and one deliberate
site. It fixes the instrument that measures the bar, plus a live defect
the comparison count cannot see.

---

## 1. `--strict` was RED on clean `origin/main`, and it was my fault

```
packages/dashboard/src/routes/register-task-workflow-routes.ts: 22 -> 23
```

My merged #2621 added a v1-IR pre-WIP fallback answering a greptile P1
and shipped no marker or baseline update, so the program's measuring
instrument has been failing on main since it landed.

Fixed **at the site** with a `DELIBERATE-LITERAL` marker, not by bumping
the baseline. That branch runs only when the IR declares no columns and
no nodes, so there is no role to resolve — `resolveLifecycleColumns`
returns nothing and the legacy pre-implementation ids are the only
pre-WIP signal that exists there. It is *unconvertible*, not unfinished;
the sibling `else` two lines down is the trait path for every IR that
can answer. A rise that is genuinely correct belongs where a reader will
see it.

## 2. The baseline was stale for 14 files — a hole, not cosmetics

A stale allowance lets converted guards return while the check stays
green. Measured gaps:

```
self-healing.ts          allows 126, tree has 111
executor.ts              allows 112, tree has 104
moves.ts                 allows  44, tree has  39
default-workflow-hooks   allows  25, tree has   7
mission-feature-sync     allows   5, tree has   0
MissionControlPanel      allows   4, tree has   0        (+8 more)
```

**Only two of the fourteen are mine.** The other twelve are
already-merged conversions by other workers where nobody re-recorded.
Re-recorded all fourteen here rather than waiting for twelve PRs,
because until it happens the ratchet is not holding the 779 it exists to
hold. Flagging it plainly: those drops are other people's work being
locked in, not mine being claimed.

## 3. Routines created tasks into the column U11 deleted

The routine editor's "Target Column" defaulted to `triage`. That value
is submitted as the create step's `taskColumn`, and an **explicit**
column bypasses the workflow entry-column resolution added for
column-less creates (#2589) — so every routine saved with the untouched
default seeded its tasks into a column the board does not declare.

Defaulting to `todo` would be the same mistake one column over: a custom
workflow declaring no `todo` is seeded into an undeclared column just as
surely, because an explicit column overrides entry resolution whatever
its value. So the default sends **nothing** and each workflow's own
intake resolution decides.

The `triage` **option** is removed too, not merely un-defaulted — fixing
the initializer alone left the operator able to pick the deleted column
one click away, and it was the option labelled "Planning", the name the
merged `todo` column now displays. Removing it retires that label
inversion as well.

Found by scanning **membership** forms rather than comparisons: the
comparison census cannot see a `?? "triage"` default, so no count showed
this and nobody was looking. Revert-proof — restoring the default fails
with *"the default must not name a column at all"*.

## 4. "Worktrees off is INERT" had one unaudited reader

The constraint was that `maxWorktrees` become genuinely inert, "not set
very high and not skipped by convention". `resolveWorktreeCapacityLimit`
returns `null` for that, and its unit tests can only prove the
**resolver** is right — they cannot see a second reader, which is the
only way the constraint breaks.

Audited every `maxWorktrees` read that bounds anything. **Exactly two:**
`scheduler.ts` (the admission gate, via the resolver, single call site,
optional gate snapshot) and `self-healing.ts`'s `enforceWorktreeCap` —
`(settings.maxWorktrees ?? 4) * 2`, a **raw** read.

The second is **not a bug** and is left alone: it bounds worktree
*directories on disk* and only removes *idle* ones. Worktrees still
exist in OFF mode, so that bound must keep applying or idle directories
accumulate unbounded. Recorded consequence: in OFF mode the number still
governs disk retention while gating no admission — an edge you scoped
out. The note says explicitly **not** to unify the two readers: routing
hygiene through the resolver returns `null` in OFF mode and silently
removes the disk bound, which is a leak dressed as a simplification.

New ratchet requires every file bounding on `maxWorktrees` to be named
with a reason, and rejects a **stale** allowlist entry. Proven by
injecting `active >= (settings.maxWorktrees ?? 4)` into
`hybrid-executor.ts`.

---

## Deliberately NOT included

- **My own census script.** #2633 landed the canonical one, and it is
better than mine — an AST classifier *plus* an independent text
classifier with `--compare`, and a baseline that fails on unrecorded
**drops** as well as rises. Mine only caught rises. I deleted mine
rather than ship a second measuring instrument; three copies of "strip
comments" is the drift shape this program keeps paying for, so the
worktree ratchet now imports #2633's `stripComments`.
- **My TaskContextMenu fix.** Superseded, and by a better answer: main's
`isPureIntakeColumn` (intake *without* hold) keeps the merged Planning
column shown and suppresses only a bare Ideas capture, which resolves
the exact hold-lane objection coderabbit raised against my version. I
briefly clobbered that merged work by checking my old file out
wholesale, caught it in the diff, and reverted.

## Verification

`pnpm lint` clean · core + dashboard `tsc` clean · census suite 23/23 ·
worktree ratchet 8/8 · RoutineEditor 49/49 ·
`routes-task-retry-planning-column` 16/16 · `lifecycle-column-census
--strict` exits 0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Added after review (all four greptile threads were real, and two of
them mattered)

**The routine fix was half a fix.** `routine-runner.ts:515` *and*
`cron-runner.ts:982` both did `column: (step.taskColumn as Column) ||
"triage"` **after** the step is read, so every routine — including ones
saved through the fixed editor — still created tasks into the deleted
column. Both now omit it.

**The advanced steps editor MANUFACTURED the defect.**
`ScheduleStepsEditor.tsx` had three `triage` defaults: the new-step
template (`:64`), the per-step initializer (`:95`), and the select still
offering it (`:344`). So the path I had *not* fixed produced the bug by
default, on fresh data. Template names no column; initializer coerces a
persisted `triage`; `triage` removed from the options; empty submits
`undefined`.

**Four pre-existing tests pinned the defect** and are rewritten to the
corrected invariant rather than appeased:

| test | asserted |
|---|---|
| `cron-runner`: "defaults column to triage when taskColumn is not set"
| `column: "triage"` |
| `ScheduleStepsEditor`: "adds a create-task step..." | `taskColumn`
toBe `"triage"` |
| `ScheduleStepsEditor`: "allows saving create-task step..." | the
legacy column is **resubmitted** |
| plus the explicit-column case added beside each, so the fix cannot
swallow a deliberate choice |

**The allowlist hole was the worst finding.** `AUDITED_BOUNDS` was keyed
by FILE, so every bounding expression in an allowlisted file was exempt
— a second raw bound in `scheduler.ts` stayed green, the one case that
ratchet exists for. Per-expression now, and making it so **immediately
surfaced a real second bound the file-level version was hiding**
(`maxWorktreesGate.used >= maxWorktreesGate.limit`, safe by construction
since the snapshot is `undefined` in OFF mode). Proven by injection.

## Found while re-reading my own deletion, not reported

A **rendered tooltip** still named a deleted cap. The "Queued to plan"
badge read *"planning starts when a concurrency slot frees up
(maxConcurrent / globalMaxConcurrent)"*. The cross-project cap is gone —
capacity is two numbers per project — so it told operators their
planning waited on a limiter they can no longer find a setting for.
Names the surviving dimension only now.

## Coding (Ideas): enforcing #2651 rather than repeating it

I took the unowned coding-ideas IR merge, concluded it must not be done,
then found **#2651 had already implemented, reverted and documented
exactly that** — with better grounding than my own argument. It added no
test, so nothing stops the next person reaching the same dead end.

So this ships their reasoning as a ratchet, not a second opinion: triage
discovery keys on the column's `autoTriage`, so a merged column is
either never scanned (cards sit on a bootstrap stub until the **capacity
hold** releases them, sending **unplanned** work into in-progress —
worse than stalling) or scanning wins and the manual gate is gone. Their
scope caveat is kept: `autoTriage` is a general trait field, so only
*this preset's* collapse is dead, not manual intake as a concept. The
registry does not reject the merged shape, which is why prose was not
enough.

## Verification (re-run)

`pnpm lint` clean · core + engine + dashboard-app `tsc` clean ·
`lifecycle-column-census --strict` exits 0 ("every file matches its
baseline exactly") · routine-runner 24/24 · cron-runner 156/156 ·
ScheduleStepsEditor 41/41 · RoutineEditor 49/49 · worktree +
coding-ideas 12/12. TaskCard has 2 failures **pre-existing on main** —
confirmed identical with my changes stashed.

---

## Bears directly on the closing bar: this PR already removes the
67-guard ratchet slack

Measured on current `origin/main` with the census itself:

```
tree total: 787   baseline total: 854   SLACK: 67

FILES ABOVE BASELINE (1):
   +1  packages/dashboard/src/routes/register-task-workflow-routes.ts  (22 -> 23)

FILES BELOW BASELINE: 13, totalling 68 unrecorded conversions
   -18  core/default-workflow-hooks.ts (25->7)   -15  engine/self-healing.ts (126->111)
    -8  engine/executor.ts (112->104)             -5  core/task-store/moves.ts (44->39)
    -5  engine/mission-feature-sync.ts (5->0)     -4  core/live-agent-count.ts (10->6)
```

**The slack is not regression — it is 13 files of merged conversions
nobody re-recorded**, against exactly **one** rise. This PR re-records
the baseline **854 → 782 across 140 files**, which closes it.

**And the "+3 that slipped in" is +1, and it is mine.**
`register-task-workflow-routes.ts 22 → 23` is the v1-IR pre-WIP fallback
my #2621 added; it is justified (that branch runs only when the IR
declares no columns or nodes, so there is no role to resolve) but it
shipped with no marker and no baseline update — which is why `--strict`
has been **red on main since it merged**. Fixed here at the site with a
`DELIBERATE-LITERAL` marker rather than by bumping the baseline, because
a rise that is genuinely correct belongs where a reader will see it.

Sequencing note for the auto-lowering change: if this lands first, that
work is purely the mechanism (auto-lower, or fail with tighten
instructions) rather than a cleanup, and the two re-records will not
collide in the same file.

Also worth carrying into that mechanism, from building the same guard
here: **`--update` must refuse to RAISE.** An earlier version of mine
wrote current counts verbatim, so a developer who added a literal and
ran the documented update command locked the regression in as the new
ceiling — the mirror of the high-water problem. Lowering can be
unattended; raising should be a hand edit with the reason recorded.

## Third piece of residue from my own deletion

`updateGlobalConcurrency` in the dashboard API client PUT to
`/api/global-concurrency`, a route removed when the machine-wide cap
went. Zero callers; the only reference was the `legacy.ts` barrel
re-export. Deleted both. `fetchGlobalConcurrency` **survives on
purpose** — the GET route remains and serves live utilization telemetry
to the footer and Command Center; nothing gates on it.

That is the third: after the second raw `maxWorktrees` reader and the
"Queued to plan" tooltip. A deletion is not finished when the
enforcement goes — the client, the label and the tooltip outlive it.

---

## Re-greened the dashboard API tests: 117 failures on main, ONE root
cause

These would have polluted the closing verification pass, and nobody
owned them.

`api()` builds headers via `new Headers(...)` and returns
`Object.fromEntries(headers.entries())` — and `Headers.entries()`
**lowercases every key**, so the object reaching `fetch` is
`content-type`, not `Content-Type`. `ab87d0d80` then added
`x-fusion-client: dashboard-ui` for run-audit attribution. Both changes
are correct; neither is visible at a call site, so **114 assertions
across 7 files** kept asserting the old shape and went red together.

Fixed by naming the shape **once** in `app/test/apiRequestHeaders.ts`
rather than patching 114 literals — restating a shared fact 114 times is
what made a two-line client change look like 117 failures. Deliberately
not a loose `objectContaining`: these tests are the only thing pinning
that the attribution header is sent *at all*.

**117 → 4.** The remaining 4 are unrelated pre-existing CSS failures
(`task-detail-modal-tablet-width` ×3, `space-token-defined` ×1) —
confirmed identical on clean main with my changes stashed.

### A gap this surfaced, recorded not papered over

Three routes failed in the *opposite* direction — they send the old
shape because they call `fetch()` **directly**, bypassing `api()`, so
they never get the attribution header. `client.ts` claims the opposite:

> "Applied once here rather than per-call so no future mutation route
has to remember it."

That does not hold for a route that bypasses the helper it is applied
in. **Measured in `app/api/`: 8 files make direct `fetch()` calls and 7
include mutations (POST/DELETE)** — among them `ai-sessions.ts`'s
DELETE, which is the same class as the four-delete incident the header
was added for. So the attribution fix has a hole in exactly its
motivating case.

Not fixed here: routing those onto `api()` is a behaviour change across
the API layer and belongs to its owner, not to a test re-green. Those
assertions use a separate `API_JSON_HEADERS_NO_ATTRIBUTION` constant so
the gap stays **visible** — if a route is later moved onto `api()`, its
test fails and points at the note explaining why.

---

## This branch takes the triage bar 10 → 5, and makes `--strict` green

`node scripts/lifecycle-column-census.mjs` on this branch reports
**triage 5**, against **10** on `origin/main`. The five removed are the
ScheduleStepsEditor template/initializer/option and the RoutineEditor
default/option — the automation paths that were creating tasks into the
deleted column.

**`--strict` was also RED on clean main, twice over, and both causes
were the same mistake:** a thorough written rationale the tool cannot
read, because the marker was not where the census looks. The census
reads a comparison node's **leading comments**; a `DELIBERATE-LITERAL`
in the JSDoc above the enclosing function or declaration does not reach
the comparison inside it.

| site | why it is legitimate | why the tool could not see it |
|---|---|---|
| `columnRoles.ts:80` `isHoldColumnRole` | degrades to `columnId ===
"todo"` only when a column has **no resolved traits** — identical in
kind to `LEGACY_PRE_IMPLEMENTATION_COLUMN_IDS` directly above, which
escapes counting only because a Set is a membership form | rationale
written, **no marker token** |
| `MissionControlPanel.tsx` ×3 | the SDLC funnel **alias table** — maps
`to-do`/`ready`/`review`/`shipped` onto one display stage with an
explicit `other` bucket, and nothing branches on it | marker in the
JSDoc; the comparisons are arrow bodies **inside the array literal**,
which it does not reach |

The second only surfaced because converting the `triage` stage to a Set
removed its count and exposed the siblings — red gate, justification
sitting three lines above, unreachable.

Both are markers, no behaviour change. Neither is a conversion
candidate: resolving the funnel table to traits would **drop the
non-column aliases it exists to accept**.

**For the auto-lowering work:** the marker-placement rule is now the
recurring trap — three instances, three different authors, including me.
A marker that does not register is indistinguishable from no marker, and
the failure mode is a red gate with a written explanation nobody can act
on. If the census accepted a marker anywhere in the enclosing
declaration's comments, none of the three would have happened.

Baseline re-recorded per the tool's own instruction ("Re-record the
baseline in the SAME PR that lowered the count").


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Task “Actions” menus no longer appear on bare cards in the Planning
column.
- Routines, scheduled tasks, and create-task steps now respect each
board’s configured workflow intake column instead of using a retired
default.
- Legacy tasks saved with the retired intake column are migrated to
automatic workflow resolution.
- Target-column selection now offers only “Automatic (workflow intake)”
and “Planning,” removing the obsolete option.
- Capacity/planning messaging and related UI tooltip text were
clarified; concurrency cap updates are managed per project.

- **Tests**
- Added/updated coverage for workflow intake resolution, create-task
target column behavior (including legacy coercion), capacity safeguards,
and API request consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 03:10:35 -07:00
gsxdsm
aa02db5782 fleet: scheduler.ts 28 → 27 + make the column-role predicates reachable from the other 80% of the backlog (#2690)
## Census

**722 → 721**; `packages/engine/src/scheduler.ts` **28 → 27**. Exactly
the one site converted. Baseline re-recorded in the same commit —
`--strict` flagged the stale allowance itself, and `in-progress` went
138 → 137.

## The unblocker (commit 1)

The role helpers live in `packages/dashboard/app/utils/columnRoles.ts`,
a dashboard-**app** module. Measured against the census:

| Location | Guards | Share | Helpers importable? |
|---|---:|---:|---|
| `packages/engine/**` | 316 | 43% | no |
| `packages/dashboard/app/**` | 150 | 20% | **yes** |
| `packages/core/**` | 148 | 20% | no |
| `packages/dashboard/src/**` | 78 | 10% | no |
| `packages/cli/**` | 24 | 3% | no |

**Only 20% of the backlog can call them at all.** #2685 widens the
helper *set* correctly; that is coverage, not location.
`packages/core/src/column-roles.ts` is the same flags-first /
legacy-id-fallback predicate placed where the other 80% can reach it —
core already exports `resolveColumnFlags`, so no new resolution
machinery comes with it.

**Semantics are mirrored from #2685, not invented**, so the two sets
cannot answer the same question differently: `complete` EXCLUDES
`archived`; `wip` keys on `countsTowardWip` (the same flag capacity
arithmetic uses); `review` accepts `mergeBlocker` OR `humanReview`. One
addition — `isTerminalColumnRole` for the `!== "done" && !== "archived"`
union, the most repeated shape in the backlog.

10 tests cover both modes of all 8 predicates, including the **degraded
no-flags fallback** — the half with no coverage when these lived only in
the dashboard app — plus the two cases that prove the predicate does
something rather than nothing: a renamed column carrying the right trait
answers yes, and a legacy id carrying the WRONG trait answers no.

## A trap every fleet worker converting engine code will hit

A new core export must be added to **both** `index.ts` and
`index.gate.ts`.

The `engine-core` gate project resolves `@fusion/core` to a bundle built
from `index.gate.ts` (`scripts/build-engine-core-gate-bundle.mjs`). An
export present only in `index.ts` is `undefined` at runtime under the
gate: 13 `scheduler-workflow-cutover` tests failed with `isWipColumnRole
is not a function`, in a file that does not mock `@fusion/core` at all.
The symptom points at the consumer, the cause is the barrel.

I nearly mis-attributed this. Baseline first:
`scheduler-workflow-cutover` is **42 passed on clean main**, so the 13
were mine — not pre-existing. That measurement is the only reason I
looked at the barrel instead of "fixing" the tests.

## The conversion (commit 2)

`scheduler.ts:1690`'s `isWipColumnTask` was a hand-rolled copy of
`isWipColumnRole` — it stored only `countsTowardWip` as a boolean and
re-implemented flags-first-then-legacy-id inline. It now stores the
resolved flags object and lets the shared predicate decide.

Behaviour is identical in all four states: column present with the flag
true or false (flags win), column absent from a resolved IR, and IR
resolution failed (both defer to the legacy id).

| Check | Result |
|---|---|
| `scheduler-workflow-cutover` | **42 passed** before and after |
| 21 scheduler/capacity/hold-release files | **372 passed** |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint` | clean |

## Flagged and skipped, not guessed

**`scheduler.ts:1736` — a latent legacy-vocabulary defect, not a
conversion.** `if (task.column !== "in-progress") continue;` gates the
file-scope-lease loop on the literal, ~40 lines below capacity
arithmetic that is trait-aware. On a renamed WIP column the loop
silently does nothing while capacity counts the same cards correctly.
Converting it *changes behaviour* on renamed boards (from wrong to
right), which the fleet rules put out of scope — so it is flagged here
for whoever owns that fix. It is the same class as U10's six
legacy-vocabulary defects.

**`hold-release.ts:343`** — already marked `DELIBERATE-LITERAL`. It is
the legacy half of FN-5719's dual-accept pair; converting it would make
both halves compute the same answer, deleting the compatibility signal
*and* its divergence detector while looking like a cleanup. Untouched.

**`task-merge.ts:254`** — the documented fallback for callers that have
not proven lane identity; trait-aware callers pass
`skipColumnIdentityCheck`. Untouched.

**The other 14 `scheduler.ts` sites** have no flags in scope (e.g.
`isLegacyDependencySatisfied(dep: Task | undefined)`,
`shouldHoldActiveFileScopeLease(...)` — task-only pure functions).
Threading an IR in changes signatures and call graphs: behaviour change,
out of scope. This is why the cluster is 28 → 27 and not 28 → 0, and the
reachability measurement behind it is #2687.

No changeset: `@fusion/*` are private and no `@runfusion/fusion`
behaviour changes.
2026-07-30 03:01:36 -07:00
gsxdsm
bb3bdab999 The ratchet follows the count down — a drop tightens instead of reddening the gate (coordinator item 2) (#2679)
Taken after asking twice for reassignment with no reply, and after the
same failure bit a **third** time. No open PR touches the census CLI, so
this is unowned in practice — **U12, say so if you have started and I
will close this in favour of yours.**

## What changed

A **drop** now tightens the baseline instead of failing. Failing hard
was defensible in isolation — a stale allowance is a hole, since those
guards can return up to the old count while the check stays green. What
it missed:

**The drop is almost never the failing author's to fix.** Eleven files
dropped during one merge wave, none of those PRs re-recorded, and none
of their authors did anything wrong. Measured three times since CI began
gating this: `columnRoles.ts` 0 → 1, then `executor.ts` twice.

A permanently-red gate is a bigger hole than a stale allowance, because
it gets ignored and then nothing is guarded at all. **The rise check —
the ratchet's actual purpose — is untouched and still fails hard.**

## The residual, named rather than glossed

In CI the write is discarded with the runner, so the committed baseline
stays stale until someone commits a tightened one. The exposure is
bounded (regrowth only up to the old count), printed on every run, and
strictly smaller than the exposure from a check people route around.
`--strict --exact` restores hard failure for the pinned end state.

**One writer:** the write is now a named `writeBaseline()` shared by the
tighten path and `--update-baseline`, rather than a second
`writeFileSync`. Two writers for one artifact is how they drift — a
lesson this file already learned once.

## Exercised end to end

| scenario | result |
|---|---|
| drop, `--strict` | exit **0**, `TIGHTENED`, allowance rewritten 9 → 6
|
| drop, `--strict --exact` | exit **1**, baseline untouched |
| rise, `--strict` | exit **1** |
| clean | exit **0** |

Pinned through the real CLI with an isolated baseline. Revert proof:
restoring the hard failure fails **1 of 32**.

## Two of my own mistakes, recorded

**A vacuous assertion, in the case that guards against vacuity.** I
first wrote `expect(allowedAfter).toBeLessThan(4 + allowedAfter)` — true
for every number. Replaced with a comparison against the inflated value
the fixture started from. This file documents that trap repeatedly and I
still walked into it, which is the argument for the mechanical revert
check over careful reading.

**The env override is `FUSION_CENSUS_BASELINE_PATH`**, not the
`FUSION_CENSUS_BASELINE` I used in the first draft — so the first
version of these cases silently ran against the **real** baseline and
passed for the wrong reason. A test whose fixture never took effect is
the same failure as a test whose fixture can't fail.

## Verification

32/32 census suites, `pnpm test:gate` **71/71**, `--strict` exits 0,
`pnpm lint` clean, `docs/testing.md` updated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Update — the base-ref ratchet (review round 2, commit `4895845579`)

The first version of this PR shipped a **named residual**: the
tightening write dies with the CI runner, so the committed allowance
stays high and a later PR can regrow guards up to it while `--strict`
prints green. I called the exposure bounded and moved on. Greptile
flagged it P1 and was right — naming a hole is not closing one.

`--strict` now stops trusting the committed number for files the branch
touched. It measures each **changed** file at the base commit
(`FUSION_CENSUS_BASE_REF`, else the PR base branch, else `origin/main`)
and fails if the file carries more guards than the base ref has. **The
enforced ceiling is what main has today**, so a stale, missing, or
long-unrecorded baseline no longer opens a window.

| decision | why |
|---|---|
| changed files only, `<ref>...HEAD` | untouched files have main's
counts by construction; censusing all ~400 at the base ref is ~400 `git
show` calls to re-derive numbers that cannot have moved. Three-dot also
stops charging this branch for guards that landed on main after the
fork. |
| a new file's base allowance is **0** | "absent at the base ref" as
unbounded would make a new file the cheapest place to hide a fresh guard
|
| fails **open** on an unresolvable ref, printing `SKIPPED` | a shallow
clone cannot produce an honest comparison; a degraded run must not read
as a clean one. The baseline comparison still applies. |
| merged into the existing `regressions` list | one failure per file,
and `--update-baseline` keeps working as the deliberate escape hatch. No
new exit path. |

**Revert proof, measured both ways.** With the base-ref block removed,
the regrowth fixture — base commit 2 guards, HEAD 5, baseline allowing 9
— exits **0** with `TIGHTENED`, which is precisely the reported
scenario. With it: exit **1**, `column-guard count ROSE`, `above its
count on the base ref`, baseline left at 9. **3 of the 4** end-to-end
cases go red on revert. The fourth passes without the fix by design — it
is the genuine-conversion case the auto-tighten exists to keep green,
and a case that reddens either way proves nothing.

The end-to-end suite builds a throwaway two-commit `git init` repo under
the temp dir, because this exploit is a property of the **plumbing**,
not of the comparison: resolving a ref, working out the changed set,
reading base source through `git show`. The comparator itself is pure
with the reader injected (`findRegrowthAgainstBase`), with its own cases
in `lifecycle-column-census-ast.test.ts` — including the one that would
silently pass everything, looking up the wrong key in
`summarize().byFile`.

**Rebased onto `origin/main` @ bc782d8d92** (the branch was forked
before the recent merge wave; its baseline read 746 against a tree of
722).

Verification on the rebased branch: census **722** / `--strict` exit 0 ·
**70/70** across both census suites · `pnpm test:gate` **71/71** · `pnpm
lint` clean.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:56:08 -07:00
gsxdsm
ae23be79f7 fleet: scheduler.ts 28 → triaged (NOT converted) + repo-wide reachability measurement — the work order sorts on a number that doesn't predict convertibility (#2687)
## Claim

`packages/engine/src/scheduler.ts` — the largest **unclaimed** cluster
(28). Triaged, **not converted**, for the reason below. Census
unchanged: **722 → 722**. No baseline movement is claimed, because
nothing was converted.

## Why not converted

Three of us independently hit the same wall on our first file — #2683
and #2684 (`self-healing.ts`), #2685 (helper coverage). This measures
the whole backlog **once** so the remaining workers don't each pay that
cost.

Two constraints gate conversion. Neither is visible in the per-file
counts the work order sorts on.

### 1. Location — the helpers aren't importable from 80% of the backlog

`isIntakeColumnRole` / `isPreImplementationColumnRole` /
`isHoldColumnRole` live in
`packages/dashboard/app/utils/columnRoles.ts`, a dashboard-**app**
module.

| Location | Guards | Share | Importable? |
|---|---:|---:|---|
| `packages/engine/**` | 316 | 43% | no |
| `packages/dashboard/app/**` | 150 | 20% | **yes** |
| `packages/core/**` | 148 | 20% | no |
| `packages/dashboard/src/**` | 78 | 10% | no |
| `packages/cli/**` | 24 | 3% | no |
| plugins | 6 | 1% | no |

**150 of 722 (20%)** can call them at all. Widening the helper *set*
(#2685, correctly) does not move this number — it's the module's
location, not its coverage. Core already exports `resolveColumnFlags`,
so a core-side predicate module would be *the same* abstraction made
reachable, not a new one. It is a prerequisite for the other 80% and
**not sufficient** — see below.

### 2. Flag scope — the binding constraint

A role predicate needs resolved trait flags. Most guards run in
functions handed a bare task row with no IR to resolve from. Threading
one in changes a signature and its call graph: a **behavior change, out
of scope**.

File-level proxy over the 572 non-dashboard guards: **339** in files
that reference an IR/flags resolver, **233** in files with none.

**That proxy overstates convertibility, and the overstatement is the
finding.** Reachability varies *within* one file, so a file-level
verdict is unusable. In my claimed cluster:

| Site | Context | Convertible? |
|---|---|---|
| `scheduler.ts:1690` | `resolveWorkflowIrById(...)` +
`resolveColumnFlags(c)` in the same block | **yes** |
| `scheduler.ts:231` | `isLegacyDependencySatisfied(dep: Task \|
undefined)` | no — task only |
| `scheduler.ts:341` | `shouldHoldActiveFileScopeLease(...)` | no — task
only |

So "convert the file" is not a unit of work that exists in this backlog,
and the rule *"the baseline must shrink by exactly your converted
count"* cannot be satisfied per-file until the count is per-site.

## A guard that must be skipped, not guessed

`packages/core/src/task-merge.ts:254`:

```ts
if (!options.skipColumnIdentityCheck && task.column !== "in-review") {
```

The parameter is `Pick<Task, "column" | "paused" | ...>` — no IR,
deliberately. The in-source FNXC comment records that callers who *have*
resolved the `merge-blocker` trait pass `skipColumnIdentityCheck` rather
than spoofing `{ ...task, column: "in-review" }`. The trait-aware path
already exists *beside* this literal.

Converting it wouldn't remove a legacy id — it would delete the fallback
the option was introduced to make explicit. **Flagged and skipped.**

## Suggested census upgrade (not done here)

Emit per-site whether trait flags are resolvable in the enclosing scope.
That turns the work order from "largest file" into "largest
**convertible** cluster" and makes baseline shrinkage predictable. I did
not touch `scripts/lifecycle-column-census.mjs` — it is the shared
instrument and changing it unannounced would invalidate everyone's
in-flight before/after numbers.

## Method correction worth propagating to every fleet worker

Claim-collision scans must compare a branch to its **merge base**, not
to `origin/main`. `git diff origin/main origin/<branch> -- <file>`
reports a difference when the branch is merely *stale* (the file didn't
exist at its base) — it flagged dashboard and CLI PRs as touching engine
E2E files. I nearly skipped a free cluster on that false signal.

`feature/code-organization-wave17` is excluded from collision checks:
**1556 files, 150 commits behind main, already `DIRTY`**. It must rebase
wholesale regardless, and counting it as a claim marks *every* cluster
in the backlog as taken.

Docs-only — no source, no test, no census change.
2026-07-30 02:53:18 -07:00
gsxdsm
da77e61118 A rate-limited provider kept getting hammered — the executor and merger lane checks were still literals (#2672)
`taskUsesProvider` resolves a task's **active lane** to decide which
providers it is running on. The **planner** half was converted to traits
— its own note in that function describes this exact failure and says it
was fixed — and the **executor** and **merger** halves were left as
`task.column === "in-progress"` / `=== "in-review"`.

So on a renamed board an actively-executing card resolved **no
providers**, and a provider rate limit never paused it: the engine kept
sending work to the limited provider with that card.

## Measured

Renamed board (`building` = wip, `checking` = review), limit triggered
by a peer:

| lane | before | after |
|---|---|---|
| executor | `["FN-TRIGGER"]` | `["FN-TRIGGER","FN-PEER"]` |
| merger | `["FN-TRIGGER"]` | `["FN-TRIGGER","FN-PEER"]` |

The trigger was paused either way — but only through the
*always-include-the-trigger* fallback, and **that is what made the hole
quiet**: one task always gets paused, so the behaviour looks like it
works.

Resolved from the **same per-workflow IR cache** the planner lane
already uses, so this adds no reads. Both halves fail soft to the legacy
literal, so an unresolvable workflow behaves exactly as before.

## How it was found — the part worth keeping

A scan for *"legacy literal within a few lines of a role-resolved call"*
flagged this file. That is the same heuristic that produced #2670, and
it is now 2 for 3.

The first thing I suspected here was the `done`/`archived` terminal
filter. I wrote that fix, and **its revert stayed green.** That is not a
dead end — chasing *why* it would not go red showed the lane check
already excludes finished cards, so the terminal literal there is
genuinely redundant, and the real defect was one line over in the lane
check itself. **A revert that stays green is information: either the
guard is vacuous, or you are looking at the wrong line.**

I dropped the unprovable change and kept the provable one. The
`done`/`archived` filter is deliberately unchanged, with that reason
recorded in the test.

## Revert proofs, isolated

- wip lane back to the literal → **1 of 55 fails** (the executing peer)
- review lane back to the literal → **1 of 55 fails** (the merger peer)

Three paired cases pass under both, so "pause everything with that
provider" cannot pass for "resolve the lane": a card parked in the
renamed planning lane is **not** paused on an executor limit, a wip card
is **not** paused on a merger limit, and the legacy vocabulary still
pauses peers when no workflow resolves.

## Verification

- 55/55 `usage-limit-detector`; `pnpm test:gate` **71/71**; engine
typecheck clean; `pnpm lint` clean
- census unchanged at 748 — the fail-soft literals remain by design,
which is why the count is not the measure of this fix

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:50:30 -07:00
gsxdsm
89d084e5bd fix(scripts): --compare was accusing the parser of a blind spot it does not have (#2682)
## The problem

`census --compare` fails on `origin/main` — I verified it at
`bc782d8d92` and at every commit on the branch where I found it. Its
failure message reads:

> The parser has a blind spot; its count cannot be the bar until this is
closed.

That matters because the parser's count **is** the bar the program just
used to declare the closing bar met (triage 0, backlog 722). A red
cross-check asserting the instrument is untrustworthy had to be settled
in one direction or the other.

## It is settled: there is no blind spot

**Measured — 13 divergent sites, all 13 seen by the parser:**

| parser's classification | count |
|---|---|
| `deliberate` | 4 |
| `role` | 5 |
| `status` | 4 |
| **missed entirely** | **0** |

## The bug is in the check

It compared per-bucket totals and failed when the regex's `column` total
exceeded the parser's. That conflates the two things it most needs to
separate:

- the parser **missed** a site → a real hole, the failure worth having;
- the parser **classified it better** → `role`/`status`/`deliberate`
instead of `column`.

The second is the parser's entire reason for existing. So the old form
fired *more* the better the parser got, while accusing it of the one
defect it did not have. The regex is knowingly weaker at telling an
agent role from a column guard — that asymmetry is why the parser was
adopted, and the check was penalising it.

The intent was never wrong; the comment above the check already said the
contract was "a site the REGEX found and the parser missed". Only the
implementation disagreed with it.

## After

```
text classifier:  {"column":728,"role":0,"status":182,"deliberate":14}
AST classifier:   {"column":722,"role":5,"status":186,"deliberate":17}
parser sees every site the regex does (+131 sites the regex cannot see).
14 the regex calls a column guard, the parser classifies as {"role":5,"status":4,"deliberate":4,"definition":1}.
```

Fails only on a genuinely missed site now, printing the first ten.
Reclassifications are reported rather than failed.

## Scope

Report-only. `--compare` is not in the merge gate — the gate runs
`--strict`, which is why this stayed red and unwatched. Verified:
`--compare` exit 0, `--strict` exit 0, lint clean. No census numbers
change.
2026-07-30 02:50:18 -07:00
gsxdsm
d75de0fb80 fleet: complete the column-role helper set — 680 of 722 guards had no helper to convert to (#2685)
**Fleet blocker, measured before claiming a file — this unblocks 94% of
the work order.**

## The gap

The work order says *"conversion pattern: the existing role helpers
ONLY; no new abstractions"*. Measured against the census, those helpers
cover **42 of 722** backlog guards:

| role | guards | helper? |
|---|---:|---|
| `intake` / `hold` → `todo` | 42 | ✅ `isIntakeColumnRole`,
`isPreImplementationColumnRole`, `isHoldColumnRole` |
| `in-review` | 200 | ❌ |
| `done` | 195 | ❌ |
| `archived` | 147 | ❌ |
| `in-progress` | 138 | ❌ |

**680 guards — 94% — had no helper to convert to.** Every fleet worker
hits this on their first file. I hit it claiming `TaskCard.tsx`, whose
42 guards are `done` 13, `archived` 12, `in-progress` 9, `in-review` 7,
`todo` 1.

## Why this is not "a new abstraction"

It is the **same** abstraction — flags-first, legacy id only as the
documented no-metadata fallback — applied to the roles it did not yet
cover. The alternative is inlining a flags-plus-fallback expression at
680 sites, which recreates exactly the copy-paste drift these helpers
exist to remove: **three inline copies in `ListView` are what started
this file.**

Widening `ColumnRoleFlags` threads nothing new through any call site.
Callers already pass these flags — `TaskContextMenuColumnFlags` declares
all of them — the interface had only *declared* the two the earlier
helpers needed, so the type was dropping the rest on the floor.

## A correction I made mid-change

My first draft of that comment claimed the flags were "already carried
on `ColumnRoleFlags`". `tsc` disproved it immediately — `complete`,
`archived` and `countsTowardWip` did not exist on the type. Corrected
rather than quietly patched, because that claim *was* the justification
for calling this a completion rather than an addition.

## Each distinction is asserted, not just documented

- **`isCompleteColumnRole` does not count `archived`** — an archived
card is finished but not *completed*; surfaces counting throughput would
double-count it.
- **`isWipColumnRole` keys on `countsTowardWip`**, the same flag
capacity arithmetic uses, so a board cannot have a column that counts
toward WIP for capacity but not for this predicate.
- **`isReviewColumnRole` accepts either `mergeBlocker` or
`humanReview`** — separable traits, but every converted caller asks "is
this card in review", for which both qualify. A caller needing one and
not the other should read the flag directly rather than widen this.

Both directions are asserted in every case, so a helper returning
`false` unconditionally cannot pass.

## Verification

`columnRoles.test.ts` **10 → 14**. `pnpm test:gate` green (10 / 158 /
487 / 71). `pnpm check:lifecycle-columns` exits 0. `tsc -p
tsconfig.app.json` clean. `pnpm lint` clean.

No census movement — this adds capability, converts nothing. My
`TaskCard.tsx` conversion (42 → 0) follows on top of it.

No changeset: internal helpers, no user-facing change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:47:26 -07:00
gsxdsm
3e80dcb8ef fix(test): a Vite prefix-match alias silently unresolved a core subpath (greens full-suite shard 1) (#2686)
## What

`full-suite.yml` shard 1 on main fails with **zero test failures** — it
dies on a resolution error:

```
Failed to resolve import "@fusion/core/task-delete-attribution" from "packages/dashboard/app/api/client.ts"
```

**Root cause.** Vite string aliases match by **PREFIX**. So `find:
"@fusion/core"` → `core/src/index.ts` rewrites
`@fusion/core/task-delete-attribution` into
`core/src/index.ts/task-delete-attribution`, which cannot resolve. The
narrower subpath alias has to come *first*.

The module exists and *is* correctly declared in
`packages/core/package.json` exports — this is purely a test-config
trap, and `packages/dashboard/vitest.config.ts` already documents it in
a comment. Six configs alias `@fusion/dashboard` (whose
`app/api/client.ts` imports that browser-safe leaf) while lacking the
narrower alias, so they inherited the trap. This carries the same
one-line pattern to all six.

## Measured

`dependency-graph` — the project actually red on main:

| | Test files | Tests collected |
|---|---|---|
| before | 3 failed \| 17 passed | 147 |
| after | **20 passed** | **180** |

**33 tests were never collected** — neither passing nor reported as
failing. That is the part worth flagging: an unresolved import removes
tests from the run silently, and the shard's own summary printed no
`Tests N failed` line at all, which is why this red looked like
infrastructure noise rather than a real defect.

No regressions: `reports` 110, `cli-printing-press` 41,
`compound-engineering` 317, **gate 726** — all green. `pnpm lint` clean.

`@fusion/desktop` is `1 failed | 264 passed` **both before and after**;
verified pre-existing on clean `origin/main` by reverting just that one
config and re-running. Cause is `@fusion-plugin-examples/roadmap` entry
resolution, unrelated — **flagged, not fixed.**

## Deliberately not changed

Engine's *second* `@fusion/core` alias (the `.gate-bundle/core.mjs`
entry) is untouched: that lane bundles core on purpose, and pointing it
at source would defeat the isolation the gate bundle exists to provide.

## Full-suite triage this came out of (for whoever owns the rest)

Reading the four red shards of the last completed run on main
(`30523568756`):

| Shard | Real cause | Owner |
|---|---|---|
| 1/4 | **this PR** — resolution error, 0 test failures | — |
| 2/4 | 23 failed: `store-wedge-resolution.pg`,
`central-archive-secrets`, `task-delete-caller-attribution`,
`task-delete-nonblocking-cleanup` | #2669 / #2675 cover the first two |
| 3/4 | **watchdog SIGKILL** mid-`@fusion/engine [1/2]` — no test
failures, no summary | unowned |
| 4/4 | 17 failed, all in `@runfusion/fusion` CLI (`project.test.ts` 8,
`task.test.ts` 5, `extension.test.ts` 2, +2) | unowned |

Two of the four shard reds contain **no failing test at all**, so
"main's full-suite failure count" cannot be read off the shard
conclusions — it has to be read off `Tests N failed` summary lines, and
shards 1 and 3 emit none.
2026-07-30 02:47:14 -07:00
gsxdsm
bb30d37e59 fleet: self-healing.ts 110 → scoped (NOT converted) — sync workflow reads make this cluster unsafe to batch (#2683)
Claiming the largest unclaimed cluster per the work order, then
**handing it back sized rather than half-converted.** Docs only; census
unchanged (722 / triage 0).

## The cluster

`packages/engine/src/self-healing.ts` — **110 guards**, largest single
file in the order.

```
by column:   in-review 48 · in-progress 20 · done 17 · todo 13 · archived 12
by receiver: column 100 · to 7 · from 3
```

## Why the mechanical conversion is unsafe here

**The engine has no synchronous way to learn a task's workflow.**
`resolveTaskWorkflowIrSync` returns the DEFAULT IR for every task in
production — `getTaskWorkflowSelection` returns `undefined`
unconditionally (a PG-cutover stub), so the reader always takes its
`!workflowId` branch. It is typed non-optional, so **no caller can
detect the substitution.**

A conversion routed through it: compiles, reads better than the literal,
**counts as census progress**, and is wrong for every custom workflow,
silently. That is strictly worse than leaving the literal — the literal
is at least honest about being one. It is the "guard that cannot fire"
pattern wearing better clothes, and the ratchet would score it as a win.

The correct form uses `resolveTaskLifecycleColumns(store, taskId)`
(async, store-aware), which needs resolved lanes **in scope per
method**. Sampled sites (926, 932, 984) do sit in `async` methods so it
is reachable — but that is a per-sweep restructuring, not a per-line
substitution, and these sweeps iterate task lists, so a naive per-task
resolve turns one sweep into N store reads.

**In-tree precedent:** `triage.ts` `discoverReadyPlanningTasks` solved
this exact problem — store-free `couldBeCandidate` prefilter, bounded
(8) concurrent resolve over the survivors, decision stays synchronous
over a resolved map. Any batch here should follow that shape per sweep.

## Recommended split, by SWEEP not by column

110 sites cannot honour *"census before/after, baseline shrinks by
exactly the converted count"* while also restructuring six-plus sweeps
in one PR.

1. **the review/merge sweeps** (`in-review` 48) — largest, and the one
where a wrong lane silently changes **merge eligibility**. First and
alone.
2. **WIP/rebound sweeps** (`in-progress` 20, `todo` 13).
3. **terminal sweeps** (`done` 17, `archived` 12) — read
`complete`/`archived`; most mechanical of the three.
4. **the 10 `from`/`to` sites** — these are MOVE-transition arms, not
task-column reads. Different question (*"is this transition into a
review lane?"*), so they must not ride along with the `task.column`
work.

## Why I am not doing item 1 myself

I am near the end of a long session — this is the same context in which
I produced a confidently-wrong structural finding earlier today
(retracted in #2667, where I trusted a hand-rolled brace counter over a
comment in the file). A 48-site restructuring of the merge-eligibility
sweeps is exactly the work that should not be done by a worker in that
state, and the fleet rules' *flag-and-skip* discipline is the right call
over guessing.

**What a fresh worker gets from this PR:** the site census, the
async-scope survey, the hazard with its root cause, the in-tree pattern
to copy, and a four-way split with the risky piece isolated. That is the
expensive part of the job already done.

## Fleet rule this cluster proves, worth adding to the brief

**Never resolve a workflow synchronously in a converted guard.** Use
`resolveWorkflowIrForTaskWithProvenance` (branch on `source`) or
`resolveTaskLifecycleColumns`; if neither is reachable at the site, flag
and skip.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:44:27 -07:00
gsxdsm
dc50425e98 docs: correct 104 future-dated FNXC timestamps across 61 files (#2680)
## What

The FNXC convention exists so a reader can place a note against the
change that motivated it. A stamp dated *after* the edit landed defeats
exactly that.

This is program-wide drift, not one author's slip — I contributed to it
in my own commits this week, which is how I noticed it.

## Measured, on this tree

**104 stamps across 61 files** dated later than the day they were
written, from one day ahead to **2026-10-19 (81 days)**:

| count | date | count | date | count | date |
|---|---|---|---|---|---|
| 50 | 2026-07-31 | 6 | 2026-08-05 | 3 | 2026-08-13 |
| 17 | 2026-08-01 | 1 | 2026-08-07 | 1 | 2026-08-19 |
| 7 | 2026-08-02 | 1 | 2026-08-12 | 2 | 2026-08-26 |
| 11 | 2026-08-03 | | | 3 | 2026-10-19 |

An earlier number I circulated was ~70. That came from a narrower
pathspec and was wrong; **104** is the measurement.

## How

Each stamp is rewritten to the date of the commit that introduced **that
line**, via per-line `git blame` — deliberately *not* stamped uniformly
with today's date. A uniform stamp swaps a wrong date for a different
wrong date and flattens the ordering that makes these comments
navigable; blame preserves it. Times of day are untouched, and a blame
date in the future is clamped rather than trusted.

## Why the verification is listed

A docs sweep across 61 files is precisely where a stray edit hides, so
the safety claims are mechanical rather than asserted:

- every changed line begins with a comment marker — **no code touched**;
- **no test asserts an FNXC date later than today**, so no `toContain`
assertion on embedded source text can be silently invalidated (several
such assertions do exist);
- CSS files, which carry several of those assertions, are outside the
pathspec.

## Verified

lint clean · merge gate green (487 + 158 + 10 + 71) · `census --strict`
exit 0 · tsc clean for core, engine, and dashboard
(`tsconfig.app.json`).

**No behavior change.** Comment text only.

## Not done here

A guard preventing recurrence. A check that rejects an FNXC stamp dated
after the commit would stop this returning, but it needs a decision
about where it runs (lint rule vs. gate) and it is a behavior change to
CI — it does not belong riding inside the sweep it would police.
2026-07-30 02:38:58 -07:00
gsxdsm
543f4a556c Tell an already-converted fallback literal from an unconverted guard — 19 of 19 dashboard scan hits were the former (#2677)
The batch phase is about to hand per-file guard lists to cheap workers,
and the census currently cannot distinguish **"not yet converted"** from
**"converted, with a documented degradation."**

## The measurement that makes this a class, not a preference

A proximity scan for *"legacy literal near a role-resolved call"* — the
heuristic that produced #2670 and #2672 from the engine — returned **19
hits across the dashboard and zero defects.** Every one was:

```ts
if (flags) return flags.hold === true || flags.countsTowardWip === true;
return column === "todo" || column === "in-progress";      // reachable only without traits
```

That literal is **correct**: it answers for callers with no resolved
column metadata, which is the case `resolveLifecycleColumns` returns
`undefined`-for-the-whole-struct to preserve. A worker told to "convert"
it would delete the only answer available when traits are absent.

## And the difference is structural, so the parser can see it

In **both** engine defects the literal sat in a **separate statement
beside resolved data**, not in a fallback branch. Proximity cannot tell
those apart; an AST can.

`traitFallback` flags the ternary form and the **early-return** form
(which is how most are actually written), and deliberately does **not**
flag a fallback whose test is itself a column-*name* check — otherwise
any `if/else` over column names would launder itself.

## Reported beside the backlog, not subtracted from it

```
COLUMN guards (the backlog):   746
  of the column guards, 9 are trait-fallback branches (already converted)
```

A fallback literal is still a literal and should go when the trait path
becomes unconditional. This only says which **kind** of work it is.

**Advisory, and structurally so:** `traitFallback` never changes `kind`,
and the count lives *outside* `totals`. My first attempt put it in
`totals` and broke two existing suites that correctly deep-equal that
shape — an advisory number does not belong in the structure that defines
the bar.

## Revert proof

Forcing `traitFallback: false` fails **3 of 33** (both fallback forms,
plus the kind-unchanged case). The two *negative* cases pass under the
revert — which is the point: they assert what must **not** be flagged,
and a classifier that flags nothing satisfies them trivially. Worth
stating, because a revert proof that only counts failures would look
stronger than it is.

## Baseline

Re-recorded: `executor.ts` 87 → 85 was **main's own drift** from #2568
landing, so `--strict` was red on main again. #2668 made the re-record
possible; the auto-tighten (coordinator item 2) is still open, and this
is the third time in this program that a legitimate merge has left the
gate red for everyone else.

## Verification

62/62 across both census suites, `pnpm test:gate` **71/71**, `--strict`
exits 0, `pnpm lint` clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added reporting for legacy column comparisons found in trait-fallback
branches.
* Census results now include a separate count for these fallback-related
column guards.
* Human-readable reports display the new metric alongside the existing
backlog totals.

* **Tests**
* Added coverage for fallback detection across ternary, early-return,
and conditional patterns.
* Added safeguards to prevent false positives in resolved-data and
column-name checks.

* **Maintenance**
  * Updated baseline census metrics to reflect revised classifications.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:33:19 -07:00
gsxdsm
30e287a29e fix(test): decouple a second logger assertion from log formatting (#2681)
Second instance of the defect fixed in #2675 — found by applying the
same attribution pass to the **core** suite's long-red files rather than
counting them.

## The defect

`withSeverityMarker` (`logger.ts:31`) deliberately wraps every message
in a machine-readable severity marker so the TUI log pane can colour by
level: an `fnlvl=<level>` marker plus a `[core-merge-policy]` subsystem
tag ahead of the real text.

This case pinned the raw string with `toHaveBeenCalledWith`, so it broke
when that convention landed. It was coupled to log **formatting**, not
to the behaviour it exists to check.

## Two instances is a pattern

`toHaveBeenCalledWith` on a logger is brittle **by construction** in
this codebase, because decorating the message is the logger's entire
job. Any assertion pinning an exact logged string will break the next
time the format changes — and both instances found so far were long-red,
i.e. nobody noticed they had stopped testing anything.

Worth a lint rule or a shared helper if a third appears. I have not
added one for two instances.

## What is preserved

Warn-once semantics and the requirement that the warning names both the
legacy value and its replacement are unchanged and still fully asserted.
Only the exact-prefix coupling is removed.

**Mutation-verified:** deleting the `severityAuditLog.warn` call in
`merge-policy.ts` fails with `expected "warn" to be called 1 times, but
got 0 times`. A contains-check that passed because it matched nothing
would be worse than the brittle assertion it replaces.

## Still red in the core suite, not addressed here

Two neighbours in the same cluster, both needing an owner's context
rather than a guess:

- `settings-parity.test.ts` — `expected [ 'testMode', 'voiceInput',
…(15) ] to deeply equal [ …(14) ]`. A settings key now appears in
**both** global and project scope without being listed as intentional.
That is either a real scoping mistake or a stale allowlist, and the
difference matters.
- `workflow-ir-settings.test.ts` — `expected 10 to strictly equal 3` on
the moved-key catalog.

Both are genuine signals, not noise. Flagging rather than guessing, same
as the funnel decision on #2674.

## Verification

`settings-defaults.test.ts` **39/40 → 40/40**. `pnpm test:gate` green
(10 / 158 / 487 / 71). `pnpm check:lifecycle-columns` exits 0. `pnpm
lint` clean.

Test-only; no changeset.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:30:20 -07:00
gsxdsm
bc782d8d92 U12: resolve the move-path compatibility flag — trait hooks unconditional, legacy branch deleted (#2655)
U12's headline goal. The raw `experimentalFeatures.workflowColumns` flag
gated **every task move**; its six seams are now unconditional and the
flag, its last two readers, and the 124-line inline legacy branch are
deleted.

## Deleted, not converted

The flag-OFF branch goes with the gate. Converting a branch we intended
to delete would have left a second definition of every column side
effect alive to drift — the defect this program has spent its length
removing.

Both readers flip in **one commit** because they are not separable: the
preflight in `workflow-task-create-ops.ts` computes the
`movePolicyPreflight` that `moves.ts` consumes and validates. Un-gating
either alone either evaluates workflow move policies — with their
plugin-gate side effects — whose result is ignored, or validates against
a preflight that was never computed.

## Evidence, not assertion

**Equivalence (precondition 1).** `moves-flag-equivalence.test.ts`
(commit 1) ran the same journey under both flag states against live
PostgreSQL and diffed the persisted row: **identical across 128 fields**
plus an equal timing shape, over `todo → in-progress → in-review → todo
→ in-progress`. Mutation-verified both ways — stamping the flag-ON
branch, and diverging the reopen hook, each fail it.

**And there was stronger evidence already on main that isn't mine.**
U2b's `move-path-equivalence.pg.test.ts` ran *every* scenario once per
path and has been green across ~10 of them: `preserveStatus`,
`preservePause`, timing accounting, `preserveProgress`,
`preserveWorktree`, engine-source rehome, `in-progress → todo`. Two
independently built harnesses agreeing is the best evidence this
question has had.

**The flag was read by nothing in production.** `experimentalFeatures`
is global-only and no module writes it, so this path had never run for
any project without a stale persisted value. That is also why a green
suite was never evidence on its own — both paths were individually valid
and only one was live.

## Two claims of mine this PR corrects

**1. Seam 2 does not introduce new rejections.** I said in #2639 and in
the census that with the flag off there is *no* target validation, so
flipping would add refusals. Reproduced the opposite: a move to an
undeclared column already rejects on the legacy path with `Invalid
transition: … Valid targets: …`. I found it because the discriminator I
wrote to prove "the flag is the cause" failed.

**2. My first equivalence test proved nothing.** It used
`updateSettings`; `experimentalFeatures` is **global-only**, so
`getSettingsFast()` filtered the write out and `useWorkflow` was false
in *both* runs. Caught by stamping the flag-ON branch and watching the
test stay green. It now writes via `updateGlobalSettings` and **asserts
the flag took effect** before the journey. U2b's harness carries the
same warning independently — `MUST be updateGlobalSettings, NOT
updateSettings`.

## The user-visible change

Move rejections now report **workflow-resolved** targets instead of the
hardcoded legacy adjacency table. Concretely: `Valid targets:
in-progress, triage, archived` becomes `Valid targets: archived,
in-progress`. That is the fix, not a regression — the legacy table still
advertised `triage`, a column the default lineage stopped declaring at
#2515, so an operator following the old message was told to move
somewhere impossible.

Likewise a move *into* `triage` is now refused rather than stranding the
card in a column with no trait flags, invisible to every trait-driven
sweep until reconciliation re-homes it.
`live-move-path-undeclared-target.test.ts` characterised exactly that
defect and carried `it.todo("should REFUSE a move into a column the
task's workflow does not declare (U2b)")` — **this fulfils it.**

## Test migration

| file | change |
|---|---|
| `move-path-equivalence.pg.test.ts` | deleted — every scenario ran once
per path; purpose fully discharged |
| `workflow-capacity-invariant.pg.test.ts` |
`setPath("inline"\|"hooks")` → `assertMovePathLive()`; the probe is
**kept** so capacity cannot pass because moves were broken for an
unrelated reason |
| `store-movement.pg.test.ts` | asserts the refusal **and** that the
legitimate backward move still works, so it reads as a narrowing |
| `raw-workflow-columns-flag-census.test.ts` | deleted per its own
instructions — it was built to fail in both directions and fired exactly
as designed: `expected [] to deeply equal [3 readers]` |
| `moves-workflow-flag-seams.test.ts` | deleted — it pinned the six
seams this removes |

## Verification

Full core suite: **33 failed / 10 files — byte-identical to main's
baseline**, with **zero** files failing exclusively on this branch. I
measured the baseline by checking out `origin/main` and running the same
command, because the first comparison I made was by count alone and
would have blamed the flip for 8 files that were already red.

`pnpm lint` clean. `pnpm test:gate` green (10 / 132 / 482 / 71). `tsc -p
packages/core/tsconfig.json` clean. Core builds.

## Left in place deliberately

The `workflowColumns` settings key stays schema-tolerated and is already
in `HIDDEN_EXPERIMENTAL_FEATURE_KEYS`, so an upgraded project carrying a
stale value renders nothing and loads cleanly. Removing it from the
schema would risk rejecting those projects for no benefit now that
nothing reads it.

---

## Rebased onto current main — and `triage` reaches ZERO

| metric | before | after |
|---|---:|---:|
| `moves.ts` column guards | 39 | **15** |
| repo column total | 745 | **741** |
| **`triage` column guards** | 1 | **ABSENT (0)** |

`triage` is now absent from `byColumnId` entirely: no unconverted
`triage` guard remains anywhere in production source. Combined with
#2664 (the last one, in `TaskContextMenu`) this closes bar item 1.

The census behaved exactly as designed on the rebase: the flip *deletes*
guards, so `--strict` reported `moves.ts: allows 39, tree has 15` rather
than leaving a stale allowance, and the re-record lands in this PR's
diff.

## Verification on the rebased tree

- Full core suite: **33 failed / 10 files — identical to main's
baseline**, zero files failing exclusively on this branch (measured by
checking out `origin/main` and diffing the failing-file sets, not by
comparing counts).
- `pnpm test:gate` green (10 / 158 / 487 / 71).
- `pnpm check:lifecycle-columns` exits 0.
- `pnpm lint` clean, `@fusion/core` builds.

## Two review fixes carried in this PR

**P1 — optionless engine moves lost their bypass.**
`resolveWorkflowBypassGuardsImpl` did `void moveSource;` — it discarded
the resolved parameter and re-read `options?.moveSource`, so
`moveTask(id, target)` resolved to `"engine"` at the call site and
computed `bypassGuards === false`. Latent while the flag gated
validation; with the gate gone, an internal executor/merger/recovery
move made without an options object would be judged as a user move.

**P2 — the absence signal.** Emitting `workflowId` unconditionally would
have stamped `builtin:coding` onto every task with no explicit
selection, reporting a fallback as authoritative. Now emits the
selection directly, so absent still means "not resolved here".


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
  - Task moves are now validated against the task’s declared workflow.
- Invalid destinations are rejected with a clear error, and tasks remain
in their original column.
  - Valid backward moves continue to work as expected.
- Move behavior and lifecycle updates are now handled consistently
across workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:19:32 -07:00
gsxdsm
87a4dbdc60 test(u9): a release-leg E2E fixture that diagnoses itself (same defect found 3x independently) (#2678)
## What

The planned-spec release-leg fixture defect has now been diagnosed
**three times independently** — #2634 (`workflow-lifecycle`), #2643
(`workflow-merged-board`), and again in `workflow-planning-lane`. Each
time it cost real time, because it presents as a *scheduler* bug rather
than a fixture bug.

The mechanism: `createTaskWithReservedId` leaves a bootstrap seed (`#
<id>\n\n<description>`), `isUnplannedForExecution` (`hold-release.ts`)
refuses to move an unplanned card out of any intake- or hold-trait
column, and the sweep reports `held: [{ reason:
"move-rejected-or-no-slot" }]` while releasing nothing. That is **the
gate working.**

This extracts the write into `seedPlannedSpec`
(`_planned-spec-fixture.ts`) which **self-checks against the real
predicate the gate uses** (`isUnplannedSeedPrompt`, imported — not
restated) and throws naming the fixture as the cause. Three call sites
converted; their ~12-line comments collapse to a pointer, so the
diagnosis and both dead-end hypotheses live once, at the seam that
causes them.

## Measured (real PostgreSQL, not estimated)

| Check | Result |
|---|---|
| 3 converted families + new ratchet | **41 pass / 41** |
| Guard neutered (`if (false)`) | **4 of 5** ratchet tests fail |
| Original defect reproduced faithfully | **3 of 7** planning-lane tests
fail |
| `pnpm lint`, engine `tsc --noEmit` | clean |

**The ratchet fails on the original defect.** It drives the two real
seed shapes through the production builders (`buildBootstrapPrompt`,
`buildRefinementSeedPrompt`) rather than local imitations, so it cannot
keep passing if the seed shape drifts — which is exactly the drift the
fixture absorbs. The one test that survives the neutered guard is the
happy path, which should not move.

**The fixture is load-bearing, not decorative.** Proven by reproducing
the defect end-to-end: production `buildBootstrapPrompt` on the created
row with the guard bypassed → 3 planning-lane tests fail, including its
own control case.

A false mutation is worth recording, because it nearly produced a wrong
"not load-bearing" verdict: a *hand-written* seed passed 7/7.
`isUnplannedSeedPrompt` is **byte-equality** against a prompt built from
the task's own title/description, so only a byte-exact seed reproduces
it. A *missing* prompt does not either — `isUnplannedForExecution`
catches the read error and returns `false`.

## Deliberately not converted

`workflow-rebound-family`'s `PROMPT.md` write. It is a
content-preservation artifact asserted byte-identical across a re-home,
not a release-gate fixture; the helper would overwrite the very bytes
under assertion.

## Reversible decisions taken (per standing authority)

- **`opts.content` seam.** Exists so the ratchet can drive a known seed
and prove the throw fires. Without it the guard could only ever be
observed passing — the "guard that reports success without checking
anything" failure mode. Documented as test-only.
- **`title`/`description` optional.** The check is shape-based: both
recognised seed forms are `<heading>\n\n<description>` with no section
headings, so the written spec cannot match either for *any* description.
Omitting them cannot mask a positive, and a test asserts that directly.
- **`merged-board`'s spec text lengthened** to match the other two (both
were already non-seed, so behaviour is unchanged; verified by the
41-pass run).

## Method correction worth propagating

My collision scan was wrong and I nearly acted on it. `git diff
origin/main origin/<branch> -- <file>` reports a difference when a
branch is merely **stale** (the file did not exist at its base), so it
flagged dashboard and CLI PRs as touching engine E2E files. Diffing
against each branch's **merge base** is correct. Re-run under the fixed
method: all five files here are uncontested, and
`feature/code-organization-wave17` genuinely does touch
`packages/engine/src/triage.ts` (so that one stays hands-off).

## Lane

`.pg.test.ts` under engine-default, `pgDescribe`-skipped without
PostgreSQL — the merge gate is unaffected. Throwaway per-file database,
never port 4040, no temp-root walk.
2026-07-30 02:19:17 -07:00
gsxdsm
07c29757a3 The archived half of the terminal pair, end to end (#2568's fix landed first and is better) (#2670)
**Live on `main`.** `parkCompletedBlockedTask` opens with *"is this card
already finished?"* and answered it with:

```ts
if (task.column === "done" || task.column === "archived") return false;
```

On a renamed board neither matches, so the guard was **inert** — and
inert here is not a missed rescue, it is active damage: the very next
block rebounds the card to its planning lane. **A completed card sitting
in a renamed complete or archived column was moved backwards out of
it.**

## Why this survived two reviews

#2644 (mine) converted the **rebound** half to resolve its target by
role. This **terminal** half stayed a literal. A role-resolved rebound
behind a name-matched guard means the renamed board takes the rebound
and never the guard — the same half-conversion shape as the evacuation
branch, except the two halves were owned by different changes, so
neither review saw both.

Worth keeping as a review heuristic for the remaining conversions:
**when a converted site sits next to an unconverted one, the conversion
can make the neighbour worse — and a diff showing only one of them looks
complete.**

## Relationship to #2568

The fix exists there, stranded four deep in a stack whose bottom (#2544)
has not merged, so nothing in that chain has reached `main`. This
re-lands **only** the guard, directly against `main`. I have noted it on
#2568 so its author can drop that hunk rather than resolve it twice; I
took no other part of that PR (its extraction and the terminal-pair
ownership change are still theirs).

## Deliberate choices

- **Fail-soft to `["done", "archived"]`** — an unresolvable workflow
behaves exactly as the literal pair did.
- **An unclassifiable column is NOT terminal.** Being unable to prove a
card is finished must not be the same as proving it is; that is what
keeps a stranded card moving.

## Revert proof

Restoring the literal pair fails **2 of 5** — the renamed complete and
archived cases. Three paired cases pass under the revert, so neither
"never park" nor "always park" can pass for resolving the lanes: a
genuinely mid-pipeline card is still parked, the legacy pair still
answers when no workflow resolves, and an unclassified column still gets
the park.

## Verification

- 5/5 new; 22/22 with `executor-rebound-already-there` and
`executor-task-done-blocked`
- `pnpm test:gate` **71/71**; engine typecheck clean; `pnpm lint` clean
- census: `executor.ts` **87 → 85** column guards; baseline re-recorded
in the same commit

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Reliability**
* Improved handling of completed/blocked tasks when workflow column
naming changes, ensuring tasks aren’t parked or moved incorrectly in
archived or non-terminal lanes.
* **Tests**
* Added Vitest coverage to verify terminal-lane decisions, including
cases where workflow resolution occurs mid-operation and lane placement
changes during the wait.
* **Maintenance**
* Updated lifecycle column census baseline numbers to match current
workflow usage.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:16:30 -07:00
gsxdsm
b752f9014d fix(core): a hold-only column vanished from the SDLC funnel (+ un-red the census on main) (#2674)
Two things, both small, one urgent.

## 1. Hold-only columns disappeared from the funnel

`hold` was absent from `TRAIT_TO_STAGE`, so a column whose only
pre-implementation trait is `hold` — a renamed board's wait-for-capacity
lane — resolved to `OTHER` and vanished from the SDLC funnel.

Measured before the fix:

```
stageForTraits(["hold"]) === "other"
```

The default lineage hid it: its Planning column also carries `intake`
and `reset-on-entry`, so it always matched something. Only a board that
names its wait lane separately was affected — **exactly the custom shape
this trait mapping exists to support**.

Revert check: removing the entry gives `expected 'other' to be 'todo'`.

## What I deliberately did NOT fix, and why

The merged default Planning column carries
`["intake","hold","reset-on-entry"]`, and `stageForTraits` prefers the
earliest stage in flow order — so `intake` wins and it still resolves to
`triage`. The `todo` stage therefore stays empty on every default board
since U11, and the funnel shows a **phantom 100% drop between Triage and
Todo**.

That is a real defect. It is also not a reversible call: changing which
stage Planning reports would retroactively alter how historical
analytics read. Flagged on #2669 for a product decision.

Adding `hold` does not touch it — `intake` still outranks — and a second
test **pins the current behaviour** so the larger question gets answered
deliberately rather than drifted into by a future edit to this map.

## 2. `check:lifecycle-columns` is RED on pristine `origin/main` — again

```
census exit on pristine main = 1
  packages/engine/src/executor.ts: allows 87, tree has 85
```

`executor.ts` is a file this PR does not touch, so a merge lowered the
count without re-recording and the blocking PR check is failing for
**every open PR**. The re-record is mechanical and is included here to
unblock it — called out explicitly because it is unrelated to the funnel
fix and should not ride along unexplained.

This is the second time the baseline has gone stale on main this way.
The rule works (`--strict` caught it immediately); what is missing is
that it caught it *after* the merge. Worth considering whether the
census should run on the merge queue rather than only on PR head —
otherwise a PR that is green when opened can still land a stale
baseline.

## Verification

`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0 after the re-record. `tsc -p
packages/core/tsconfig.json` clean. `pnpm lint` clean.
`sdlc-funnel-default-columns.test.ts` 8/8.

No changeset: the funnel entry is a correctness fix with no user-facing
API change, and the baseline re-record is internal.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved SDLC funnel classification for hold-only columns, placing
them in the Todo stage instead of Other.
* Preserved correct Planning column behavior when hold-related traits
are combined.

* **Tests**
* Added coverage for hold-related funnel stage mapping and trait
ordering.
* Updated lifecycle column census baselines to reflect current results.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:08:26 -07:00
gsxdsm
fe7e68bc13 fix(core): wedge notifications could never be resolved on PostgreSQL (42P18) (#2669)
Not a U12 change — found while **attributing** the pre-existing live-PG
failures during U12's closing verification, and it turned out to be a
product bug rather than a stale test.

## The defect

`resolveWedgeNotification` builds its UPDATE with:

```ts
jsonb_build_object('status', 'resolved', 'transitionedAt', ${transitionedAt})
```

`jsonb_build_object` is variadic `"any"`, so there is no signature for
PostgreSQL to resolve the bind parameter against. It rejects the
statement at **parse time**:

```
42P18: could not determine data type of parameter $1
```

Parse-time is the important part: this failed on **every call**, not on
unusual data. Wedge notifications could not be resolved at all in
PostgreSQL mode.

Casting the parameter to `::text` fixes it.

## Evidence

- `store-wedge-resolution.pg.test.ts` goes **0/7 → 7/7**. That suite has
been red on `main`.
- **Causally verified, not assumed:** removing the cast reproduces
`42P18` exactly. The fix is the cast, not something incidental to the
edit.
- Checked the rest of `packages/core` for the same shape — this is the
only `jsonb_build_object` call site, so there is no second instance
hiding.

## Why it survived

The failure is in a live-PG suite that was already red, so it read as
part of the ambient noise. I only found it because the closing
verification required me to attribute each failing suite to a cause
rather than count them — and "these 4 fail on main too" is an
attribution of *whose*, not of *what*.

Worth flagging for whoever owns the remaining three
(`agent-logs-and-monitor`, `central-archive-secrets`,
`workflow-settings-project-identity`): the same reasoning applies. A
suite failing on main is not evidence that the code is fine.

## Verification

`pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0. `tsc -p packages/core/tsconfig.json`
clean. `pnpm lint` clean.

Independent of #2655; either order merges.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:05:27 -07:00
gsxdsm
a099813e94 fix(test): decouple the audit-emitter assertion from log formatting (last of the 4 red PG suites) (#2675)
Last of the four long-red live-PG suites. **This one is a stale test —
the only one of the four that is.**

## The cause

`withSeverityMarker` (`logger.ts:31`) deliberately wraps every message
in a machine-readable severity marker so the TUI log pane can colour by
level. The emitted string carries a `fnlvl=warn` marker and a
`[core-async-secrets-store]` subsystem tag ahead of the real text.

The assertion pinned the raw message with `toHaveBeenCalledWith`, so it
broke when that convention landed. It was coupled to log **formatting**,
not to the behaviour it exists to check.

## The fix

Rewritten to assert what it actually cares about: exactly one warning,
whose message **contains** the subsystem-tagged text, carrying the
underlying cause.

Both halves of the behaviour stay pinned — the `resolves.toMatchObject`
above proves the secret is still created when the audit emitter fails,
and this proves the failure is surfaced rather than swallowed.

**Mutation-verified rather than assumed green:** deleting the
`severityAuditLog.warn` call in `async-secrets-store.ts` fails with
`expected "warn" to be called 1 times, but got 0 times`. A
`stringContaining` assertion that passes because it matches nothing
would be worse than the brittle one it replaces.

## The four, complete

| suite | verdict |
|---|---|
| `store-wedge-resolution` | **product bug** — `42P18`, total runtime
failure of wedge resolution in PG (#2669) |
| `workflow-settings-project-identity` | **stale docs** — resolver
contradicted its own documented order (#2671) |
| `agent-logs-and-monitor` | **real defect** — funnel mis-bucketing from
the U11 merge; half fixed in #2674, half needs a product call |
| `central-archive-secrets` | **stale test** — this PR |

**Three of four were real problems**, sitting behind "pre-existing,
fails on main too". That phrase answers *whose* problem it is, not
*what* is wrong.

## Census, again

`check:lifecycle-columns` is **still** exiting 1 on `origin/main` —
`executor.ts: allows 87, tree has 85` — the same staleness flagged on
#2674. Re-recorded here too, because the blocking check stays red for
every open PR until some PR carries it, and I do not know which of #2674
/ this one lands first.

## Verification

`pnpm test:gate` green (10 / 158 / 487 / 71). Suite **14/15 → 15/15**.
`pnpm check:lifecycle-columns` exits 0 after the re-record. `pnpm lint`
clean.

No changeset: test-only plus an internal baseline.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:59:53 -07:00
gsxdsm
e711fbab15 The ratchet's baseline could not be re-recorded once a file rose — the one state that blocks a correct conversion (#2668)
Unowned (no open PR touches the census CLI — only its baseline JSON) and
**live**, since #2654 gates CI on `--strict`.

## The problem

`--update-baseline` sat **behind** the rise exit, so the only supported
way to re-record was unavailable in exactly the situation that needs it.

That matters because **a conversion legitimately adds a literal.** The
correct shape for a caller that may have no traits is `flags ? flags.x :
columnId === "legacy"`, and each one raises a file's count by one.
Measured on current main: `columnRoles.ts` went **0 → 1** from precisely
that shape (added by #2647, documented at the site, correct code).

So a worker doing the right thing meets a red gate whose only escape is
hand-editing the JSON. That is how a ratchet becomes something people
route around rather than run — and then it guards nothing. This is the
same failure mode as a guard that cannot fire, arrived at from the other
side.

## The change

`--update-baseline` is an explicit operator action, so it re-records
**unconditionally** and prints what it accepted under `ACCEPTED RISES`.
Swallowing a rise silently is the real danger; refusing to let anyone
re-record is the same danger one step later, wearing a red check nobody
trusts. **The rise check is unchanged** and still exits 1 without the
flag.

**One writer now.** The old second `writeFileSync` behind the rise exit
is deleted rather than left unreachable — two writers for one artifact
is how they drift. The `!deliberateTracked && updateBaseline` special
case went with it, since the unconditional block covers the legacy-shape
migration too.

## Exercised end to end

On a real rise injected into `live-agent-count.ts`:

```
rise + plain --strict              exit 1   (the ratchet still bites)
rise + --strict --update-baseline  exit 0   "ACCEPTED RISES  live-agent-count.ts: 6 -> 7"
```

Four cases assert the CLI's own source, because exit codes are the
contract and the pure summarizer cannot express them: the write precedes
the rise check, the branches exit 0 and 1 respectively, accepted rises
are **named**, and there is exactly **one** writer.

## A note on the revert proof, because it caught me twice

My first attempt to move the block back was a **no-op**: the marker I
sliced on (`if (regressions.length > 0) {`) also appears *inside* the
update block, so the "revert" reassembled the file unchanged and the
suite stayed green. **A revert proof that does not go red can mean the
guard is vacuous *or* that the revert did not land** — and the second is
easy to miss when you are expecting the first. The real revert fails **2
of 27**, and the assertions now verify marker *uniqueness* before
slicing on it.

## Verification

- 27/27 census suites; `--strict` exits 0; `pnpm test:gate` **71/71**;
`pnpm lint` clean
- census on this tree: 748 column guards, **4 triage** (all in
`moves.ts`'s flag-OFF block, deletion-scheduled with #2655)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:46:36 -07:00
gsxdsm
13bf7e001d Closing-bar verification pass on origin/main — one tree, one report (+ the E2E red it found) (#2660)
**Closing-bar item 4, run on one clean tree at `origin/main`
(`be63e72f1`).** Nobody was assigned this and my own work is merged, so
I took it.

> **This PR is now REPORT-ONLY — net zero file changes.** I found the
planning-lane E2E red, fixed it, then discovered **#2658 (gsxdsm) makes
byte-for-byte the same change** to the same helper and was opened first.
I reverted mine rather than leave two identical edits to one function to
conflict. **The E2E result below depends on #2658 landing** — on
`origin/main` without it, that suite is 2 failed / 5 passed.
>
> The duplication is worth one note for the fleet: two workers
independently hit the same control-card failure and independently traced
it to FN-7648's unplanned-seed gate plus a fixture that never wrote a
spec. Independent confirmation of the diagnosis, but also ~an hour spent
twice — the census-style work order exists to stop exactly that, and E2E
fixture defects are not on it.

## Report — all four, one tree

| Check | Result |
|---|---|
| `pnpm test:gate` | **PASS** (132 + 10 + 487 + 71 tests) |
| `pnpm verify:fast` | **PASS** — 13 steps green in 89.9s, boot smoke
`GET /api/health 200`, clean shutdown |
| E2E families | **13 files / 109 tests PASS** — *after* the fix below;
**2 failed** before it |
| census | total **787**, triage **10** |

## Two corrections to the bar itself

**1. It is not "all-8 E2E" any more — there are 13 families.** The suite
grew while the bar was being written:

```
agent-count · agent-link · lease-rebound · lifecycle · merge-family · merge-rebound
merge-safeguards · merged-board · planner-lane · planner-lane-resolution
planning-lane · rebound-family · stranded-column
```

A verification pass scoped to 8 would have skipped 5 families —
including the one that was red. Worth fixing the number in the bar so
the final pass globs rather than counts.

**2. `DELIBERATE-LITERAL (reviewed)` reads 3, and I chased it —
RESOLVED, no gap.** I flagged the drop from an earlier "7" as a possible
fleet-safety hole. It is not one. Reconciled against `--json byFile`:

| File | markers | counted `deliberate` | counted `column` |
|---|---|---|---|
| `hold-release.ts` | 2 | **2** | **0** |
| `live-agent-count.ts` | 1 | **1** | 6 |
| `replan-target.ts` | 2 | 0 | 4 |

`deliberate: 3` = hold-release 2 + live-agent-count 1, which is exactly
the set of marker-covered **comparisons**. `replan-target.ts`'s two
markers sit above `return "triage"` **return-value** literals, not
comparisons — the census correctly does not count those as guards at
all, so they are neither `deliberate` nor `column`. The earlier "7" was
simply a different tree state before conversions landed; I was quoting a
stale number.

Worth noting the marker matcher is already hardened for the subtle case:
`hasDeliberateMarker` walks every **ancestor** rather than the enclosing
statement, because the real markers sit above the enclosing *function*
while the comparison is a `return` inside it — a statement-only lookup
"silently reclassified three reviewed literals as backlog". That is the
guard-cannot-fire pattern, already caught and fixed by whoever wrote the
AST version.

**Consequence for the fleet: the census's categories are trustworthy
as-is.** No pre-launch action needed on this.

## The red it found

`workflow-planning-lane-live-e2e.pg.test.ts` — **2 failed / 5 passed**,
including its own **control** case:

```
releases an ordinary held card on a default board (the control)
  → AssertionError: expected [] to include 'FN-OK'
```

`seedHeldTask` never wrote a `PROMPT.md`, so task creation's bootstrap
seed stood, and FN-7648's `isUnplannedForExecution` correctly refused to
release an unspecified card. **The sweep was right; the fixture was
asking it to release a card that had never been specified.**

**This is the second instance of the identical defect** — same cause and
same fix as `workflow-lifecycle-live-e2e`'s `seedTask` in #2634. This
suite was written after that fix and did not inherit it. The graph-entry
contract doc already states the rule: *"Scheduler/release test fixtures
must model a card that cleared the gate ... A held unreviewed card is
the gate working."*

Both failures had one cause — the mid-sweep approval-park case was
downstream of the control never releasing. **5 → 7 passed**, and cards
that are *supposed* to be held still are, held by their own
status/marker, which is what those cases assert.

Given it has now happened twice, a shared `seedPlannedTask` helper in
the E2E fixture module would prevent a third. I did not add one here: it
touches suites owned by U7 and U11 mid-consolidation, and this PR should
stay the verification pass plus its one finding.

## Bar status after this

- **gate / verify:fast / E2E** — green on one tree, with this commit.
- **triage → 0** — still **10**, all in U12's `moves.ts` (4) and
`register-task-workflow-routes.ts` (1) per file scan; flag resolution in
flight.
- **ratchet tightened (item 2)** — not done, U12's.
- Once triage hits 0 and the ratchet lands, re-running this exact pass
is a ~4-minute job and I can produce the final report.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:38:29 -07:00
gsxdsm
632d10a9b4 fix(engine): a completed-blocked guard was inert on renamed boards — plus one owner for the terminal pair (#2568)
Two commits: a behaviour-preserving extraction, then the behaviour
change.

## ⚠️ Stack note worth acting on

**#2550 and #2554 both report MERGED, but their content is not on
`main`.** They merged into their *base branches*, and the bottom of that
stack (**#2544**) is still open. Nothing in this chain has reached
`main` yet.

Nothing is lost — everything is in
`origin/feature/workflow-e2e-merge-rebound`, which is why this PR
targets it. But "merged" reads as "landed" and here it doesn't.
**Merging #2544 flows the whole chain down.**

## The bug

`parkCompletedBlockedTask` opens with *"is this card already finished?"*
and answered it with:

```ts
if (task.column === "done" || task.column === "archived") return false;
```

On a renamed board neither matches, so **the guard was inert** — and the
very next branch (`if (task.column !== "todo")`) would then have **moved
a completed card back out of its own terminal column**.

A guard that never fires does not fail a test. This one was found by
tracing the last ledger site, not by anything going red.

## Why a shared owner, not a local fix

`merger-ai`'s `isAlreadyFinalizedColumn` held the **only** copy of the
per-role terminal-pair rule — a P1 learned the hard way (PR #2471
review): a per-**set** fallback collapses to one element for a workflow
declaring `complete` but no `archived`, silently dropping the archived
half of every already-finished check.

Executor's guard was the raw literal pair, so **whoever converted it
next would have re-made exactly that mistake** — the lesson lived in a
comment in another file. Hence `resolveTerminalColumns(ir)` in core: one
owner, one place for the rule.

## Evidence, and its limits

**Commit 1 (extraction) is proven behaviour-preserving**:
`workflow-already-finalized-live-e2e` is unchanged and green through the
delegation, and the per-set mutation **still fails** through the shared
helper.

**Commit 2 (the fix) is unproven at the call site, and I'm labelling it
rather than implying otherwise.** `parkCompletedBlockedTask` is private
and reached only from inside executor dispatch — I could not drive it
end to end. So the shared helper gets its **own** tests, in both
partial-role directions, precisely because its other consumer can't
vouch for it. The call site is a one-line delegation to a tested
function.

Weaker evidence than the rest of this unit's work. Saying so, because
quietly counting it as proven is the exact failure this unit exists to
catch.

## Census

417 → 416. That ratchet (#2557) is a **ceiling**, so it stays green
without coordination; lower the pin when convenient.

## Verification

- E2E suites 10/10; helper unit tests 5/5
- core + engine `tsc --noEmit` clean
- `pnpm test:gate` green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Note on the conflict status (2026-07-31)

GitHub reports this PR `CONFLICTING / DIRTY`. **It is not.** Three
independent checks:

- `git rebase origin/main` on the pushed head reports *"up to date"* and
leaves the SHA unchanged — the branch is already on top of main.
- `git merge-tree` against the merge base produces **zero** conflict
markers.
- `origin/main` is unchanged at the commit this was rebased onto.

The remote SHA matches the local head, so the push landed. The
`mergeable` field is a **stale computation** — it goes stale after a
force-push and doesn't always recompute.

This branch has now been rebased and force-pushed four times against
that cached value. Worth guarding at the source: the auto-retry treats
`mergeable` as ground truth, so a stale value generates conflict notices
indefinitely. Confirming with a trial rebase or `git merge-tree` before
dispatching distinguishes "actually conflicting" from "GitHub hasn't
recomputed" — one command, and it ends the loop.

Verification on the current head: merge gate green (487 + 158 + 10),
engine + core tsc clean, lint clean, 20 tests in the affected suite,
zero unresolved threads.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:33:03 -07:00
gsxdsm
8d84cee11e fix(core): the workflow-settings identity resolver contradicted its own docs (2 long-red tests) (#2671)
Second of the four long-red live-PG suites, after #2669. This one is
**stale documentation making a stale test look like a code bug** —
behaviour is unchanged.

## What was wrong

`getWorkflowSettingsProjectIdImpl` documented a three-step resolution
order:

```
(a) store.asyncLayer?.projectId       — central-registry id (PG)
(b) store.db.getProjectIdentity()?.id — legacy SQLite identity
(c) store.rootDir                     — last-resort key
```

The code does (a), then returns `rootDir`. **Step (b) was removed** by
`FNXC:SqliteDualPathCleanup 2026-07-26-14:15` — but the doc block kept
describing it, and a comment three lines above the return still said
*"Only the true legacy (non-backend) path consults the SQLite
identity"*, which has been false for every caller since.

## Which side was wrong — settled by construction, not judgement

In my triage on #2669 I said I would not guess between "the test is
stale" and "the code lost a needed branch", because the two have
opposite consequences and the stale comments made the intent unreadable
from outside. That was the right call then; it is now answerable:

`dbImpl` **throws unconditionally and ignores its store argument**
(`task-id-integrity.ts:58`):

```ts
export function dbImpl(_store: TaskStore): Database {
  throw new Error("TaskStore.db: SQLite Database is not available in backend mode …");
}
```

There is no mode in which `store.db` yields a usable SQLite handle. Step
(b) is unreachable **by construction**, not merely unused — so the code
is right and the documentation was wrong.

## Why the tests passed review originally

They build a store double whose `getProjectIdentity()` **returns** a
value:

```ts
db: { getProjectIdentity() { return { id: "legacy_identity_id" }; } }
```

Production cannot produce that shape. The double made an unreachable
branch look testable, which is how the assertion survived the cleanup
that deleted the branch.

Rewritten to the shipped contract. A neighbouring case that already
asserted `rootDir` *when the stub throws* was passing all along — the
two forms of the same store disagreed inside one file.

## Verification

Suite **7/9 → 9/9**. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0. `tsc -p packages/core/tsconfig.json`
clean. `pnpm lint` clean.

No changeset: no behaviour change, and no user-visible effect.

## Remaining from the four

- ✅ `store-wedge-resolution` — real product bug, fixed in #2669
- ✅ `workflow-settings-project-identity` — this PR
- ⬜ `agent-logs-and-monitor` — `expected +0 to be 2` on an aggregation
- ⬜ `central-archive-secrets` — an assertion on `warn` arguments

Two of four were real problems hiding behind "pre-existing". The other
two are still unruled-out.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:32:43 -07:00
gsxdsm
a6138abeff U12: DELIBERATE-LITERAL counts key on file AND column — closing the P1 left on merged #2661 (#2666)
Closes the P1 that was still open when #2661 merged.

## The hole

A per-file integer is offset **within a single file**: remove one
reviewed `todo` exemption, add an `in-review` one beside it, and the
number never moves. The fresh guard is invisible to the column counts
too, because deliberate findings are excluded from them — so `--strict`
goes green with a new lifecycle-column guard hiding inside an existing
marker.

Now keyed on **file AND column id**.

**Proven with the exact scenario:** swapping a marked `triage` for
`done` inside `TaskCard.tsx` leaves the per-file total unchanged and now
fails with

```
packages/dashboard/app/components/TaskCard.tsx (DELIBERATE-LITERAL: done): 0 -> 1
```

## The pattern worth naming

This is the **third** time this instrument has been defeated by an
aggregate:

| version | defeated by |
|---|---|
| repo-wide `totals.deliberate` | an addition in file A offset by a
removal in file B |
| per-file integer | an addition offset by a removal **in the same
file** |
| per-file per-column | — |

Each step narrows what can offset silently, and I walked into the next
one twice by fixing the *reported case* rather than the *shape*. Writing
it down because the same reflex will produce a fourth if someone adds
another aggregate here.

**The residual is deliberate, not an oversight:** a same-file
**same-column** swap still offsets. Two `todo` exemptions in one file
are interchangeable by definition, so there is nothing a reviewer could
act on. That is recorded at the site so the next person doesn't
rediscover it as a bug.

## Migration, again

The key **shape** changed (`file` → `file\0columnId`), which is the same
hazard as a missing field: comparing new keys against old reports every
existing marker as a fresh rise and pushes people to convert
already-reviewed literals. I hit it on the first run here — `TaskCard
(DELIBERATE-LITERAL: triage): 0 -> 2` — exactly as I did one shape
earlier in #2661.

Detected by the delimiter rather than a version field, since old keys
have none, and re-seeded on the next `--update-baseline`. 15 file+column
entries recorded.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0.

Independent of #2655; either order merges.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:02:25 -07:00
gsxdsm
1f149d21de fix(engine): TAKING spec-staleness.ts + mission-feature-sync.ts — planner lanes (2 triage guards → 0) (#2616)
**Claiming `packages/engine/src/spec-staleness.ts` and
`packages/engine/src/mission-feature-sync.ts`.** Deliberately *not*
`self-healing.ts` (contended) or `task-creation.ts` (#2589 in flight).

## Guard counts

| scope | before | after |
|---|---|---|
| `spec-staleness.ts` | 1 | **0** |
| `mission-feature-sync.ts` | 1 | **0** |
| repo-wide `column === / !== "triage"` in `packages/*/src` (excl.
tests) | 26 | **24** |

**21** once #2612 (comments-ops, 3 guards) also lands.

## What was silently broken

**mission-feature-sync** — *"has this task returned to a planner lane?"*
decided whether a mission feature drops from `in-progress` back to
`triaged`. Keyed on the legacy pair, a card sent back for re-planning on
a renamed board left its feature stuck at `in-progress` **forever**: the
mission board showed work in flight that nobody was doing, and nothing
said so.

**spec-staleness** — the preserved-progress skip refuses to fire for an
**intake** card, since a card being specified has no progress to
protect. Keyed on `triage`, a renamed-board intake card looked like
started work and its stale spec was skipped instead of re-planned.

## The union is a deliberate call, and the existing suite forced it

My first cut *replaced* the legacy pair with the resolved lanes. That
broke a real case: **post-U11 the default lineage has no `triage`**, so
a legacy row still resting there stopped counting as a planner lane.

`usage-limit-detector` already made this call for the same situation and
wrote down why — **over-inclusion is the safe direction**. Marking a
feature `triaged` for a card in a legacy planner column is recoverable;
a mission board permanently showing phantom work is the bug. So the
legacy pair stands and resolved lanes are *added* to it.

Worth noting the existing test is what caught this, not review — which
is the argument for converting against a real suite rather than in
isolation.

## A parameter, and why that needs the ratchet

`spec-staleness`'s predicate is **pure** (a task, no store), so the role
arrives as a parameter and both callers resolve it. That optionality is
exactly the caller-omission hazard this program has already shipped
twice, so the function is also registered in core's
`role-parameter-caller-audit` (#2588).

**This PR's tests prove the parameter is honoured; the audit proves it
is passed. Neither alone is enough** — that split is the whole lesson of
#2586.

## Mutation-verified

| mutation | result |
|---|---|
| mission-feature-sync → legacy pair only | the two renamed cases fail |
| spec-staleness → restore the `triage` literal | its renamed case fails
|

Negatives included in both: demoting a **WIP** card's feature would
report running work as un-started, and never-skipping would discard
every card with real progress.

## Verification

- new suites 4/4 and 3/3; `mission-feature-sync` + `spec-staleness`
27/27
- engine `tsc --noEmit` clean; `pnpm test:gate` green (482 + 132 + 10)
- `executor-prompt` reports 3 failures **both with and without** this
change — pre-existing, baselined by stashing rather than assumed

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:54:01 -07:00
gsxdsm
dca20496f4 consolidate/u7: plugins to zero + 8 executor rebound guards + resume lanes (supersedes #2607, #2635, #2640) (#2644)
Consolidation branch for U7, per the new one-branch working mode.
**Supersedes #2607, #2635, #2640** — the three of my PRs that were stuck
on review threads. My other seven (#2602, #2605, #2606, #2611, #2621,
#2628, #2633) are green with **zero unresolved threads** and are
deliberately left alone for the merge sweep.

## What is in here, file by file

| file | change | guards before → after |
|---|---|---|
| `plugins/…/glasses/src/agent-actions.ts` | gates, destinations and
degraded-resolution refusal all resolve from the task's own workflow | 2
→ 0 |
| `plugins/…/glasses/src/quick-capture.ts` | accepted capture columns
come from the board; default no longer names the deleted column | 1 → 0
|
| `plugins/…/glasses/src/settings.ts` | quick-capture default was
`triage`, the column #2515 removed | (assignment, uncounted) |
| `plugins/…/dependency-graph/src/GraphTaskNode.tsx` | redundant column
condition deleted | 1 → 0 |
| `packages/engine/src/executor.ts` | 8 rebound guards compare the
resolved column; 4 resume-eligibility literals share one resolver | 151
→ 143 (+4 off-bar) |
| `packages/engine/src/__tests__/` | 4 new suites, 26 cases | — |

`plugins/` reaches **zero** column guards with this branch.

## The three threads it closes

**#2607 — five findings, all mine, all the same rule.** I kept
*qualifying* a legacy-id fallback instead of removing it:

| attempt | rule | hole review found |
|---|---|---|
| 1 | fall back to `todo` when the role is missing | moved cards to
phantom columns |
| 2 | …only if the workflow **declares** `todo` | aliased **review**
lane named `todo` |
| 3 | …and only if no other role is assigned to it | **traitless**
parking column named `todo` |

The qualifications were the mistake. Once `resolveLanes` returns a lane
set the workflow *has* a column vocabulary, so "no column carries the
hold trait" is a complete answer — refuse. `destination()` is two lines
now, with no aliasing surface left to qualify.

Plus a sixth, which is a genuinely different state: **degraded
resolution is indistinguishable from the default board.**
`resolveWorkflowIrForTask` is total by design — a missing definition
silently returns the *default* coding IR — so a card on a custom board
whose definition could not be read resolved to `todo`/`in-progress`.
`undefined` lanes cannot express that (it means "no workflow at all",
where the legacy ids *are* the answer). The actions now refuse with 409.
#2618 would replace this check with resolver provenance; it is not
merged, so this does not depend on it.

**#2635 — "seven rebound sites remain untested."** Fair; my "same shape"
note was an assertion, not coverage. Seven of the eight need a live
graph run to reach, so the *shape* is pinned instead: a static check
that no guard in front of a rebound move compares against a column
literal, with a vacuity case (the same detection run against the
original shape) and a match-count floor (≥8), because a guard reporting
success on zero matches is worse than no guard.

**#2640 — duplicate workflow resolution.** Framed as I/O; it is also a
correctness bug. Eligibility and re-entry are two halves of one decision
and resolved the workflow separately, so a workflow edit landing between
them has the halves reading *different boards*. Now one caller-owned
memo per decision — caller-owned because a process-lifetime cache would
have to guess when a mid-flight workflow edit invalidates it.

## Behavioural findings, not tidying

- **The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** `promotedFromPlannerColumn` was false
on a renamed board, so finished work resting in planning was never
promoted; the code fell through to a review handoff that role adjacency
rejects, and the card stayed stuck with its work complete.
- **Rebound guards could not see the column their own move targeted.**
U5b converted the move target; the eight `column !== "todo"` checks in
front of it were left literal, so on a renamed board the engine moved a
card into the column it was already in — and `moveTaskInternal` runs
reset-on-entry on every real move, so at the `preserveProgress: false`
site it reset step progress a second time.
- **The FN-1404 `task:move` audit row was lying**, recording `to:
"todo"` while the move target was resolved. A run-audit trail that
disagrees with the move it describes is worse than none. Not a
comparison, so no census counts it.
- **A task interrupted by an engine pause never resumed on a renamed
board** (off-bar, `in-review`/`in-progress` literals): four comparisons
decided one question and had to agree; two of them disagreed on a
renamed board, so re-entry silently never fired.

## Revert proofs, isolated per site

| reverted | result |
|---|---|
| `destination()` back to attempt 3 | 3 of 38 fail |
| degraded-resolution refusals removed | 2 of 42 fail |
| capture set back to the legacy five | 2 of 3 fail (renamed-board
suite) |
| forward exclusions → literals | 1 of 14 fails |
| missing-wip refusal removed | 2 of 14 fail |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| promotion target → `"in-progress"` | 3 of 7 fail |
| one rebound guard → `!== "todo"` | 1 of 3 fails (static shape) |
| resume lanes → legacy trio | 1 of 5 fails |

Every conversion is paired with a negative — a forward move, a
not-a-planner-lane card, a default-lineage card, an unresolvable
workflow — so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Commit discipline

Twelve commits, each one thing: the code move (`resolvePlannerLanes` out
of `triage.ts`) is separate from every behavior change, and each review
fix is its own commit with its own revert proof.

## Verification

- `pnpm test:gate` **71/71**
- 162/162 across the glasses plugin's 19 files; 26/26 across the four
new engine suites
- engine + glasses typecheck clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Engine recovery and retries now work correctly with renamed or
customized workflow columns.
  * Tasks in manual-intake columns are no longer automatically planned.
* Agent actions and quick capture now respect each board’s declared
columns and lifecycle stages.
* Awaiting-approval tasks are recognized regardless of their current
column.
* Command Center SDLC funnel stages now accurately reflect customized
workflows.

* **Documentation**
* Added guidance for safely changing workflow-column logic and
interpreting lifecycle-column checks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:52:55 -07:00
gsxdsm
76e92f33c4 fix(core): review badges were silent on renamed boards — the same P1 as #2470, one role over (#2586)
Independent of my other open PRs.

## The defect

PR #2470's review caught `getStalePausedTodoSignal` gaining a
`holdColumn` parameter in B1 while **both** hydration sites in
`reads.ts` omitted it — a correct guard comparing against the literal,
so the badge was silent for a paused card in a renamed hold column.

**That P1 was fixed for `holdColumn` and not for its sibling.**
`getStalePausedReviewSignal` and `getInReviewStalledSignal` both take
`reviewColumn`, and **all six call sites in the same file** left it
defaulted to `"in-review"`.

So on a renamed board (`checking`) both review badges were silent — the
identical defect, in the identical file, one role over, *after* the
pattern had already been found, written down, and fixed next door.

## The transferable part: this class is invisible to the census

My column-literal census (#2557) cannot see this. The literal lives in a
**parameter default**, and the offending call site **contains no literal
at all** — it's defined by what it *omits*.

The audit that finds it is different in kind: *"for every
role-parameterised signal, does each caller pass the role?"* — run
across the **callers**, not the definitions. Result on `reads.ts`:

```
PASSES holdColumn    x2      <- fixed by #2470
OMITS  reviewColumn  x6      <- never fixed
```

## Why threading differs per path

Not one helper call, because the three list paths differ:

- `listTasksImpl` / `searchTasksImpl` map **asynchronously** → resolve
inline through a per-pass IR cache
- `listTasksModifiedSinceImpl` maps **synchronously** → pre-resolve into
a Map beside the existing `holdColumnByTaskId`, which exists for exactly
the same reason

One IR per workflow per pass in all three.

## Evidence

Proven against a real store through the **real hydration paths**
(`listTasks` and `listTasksModifiedSince`), mirroring the sibling
renamed-hold suite because the defect lives in hydration rather than in
the pure signal.

**Mutation-verified:** reverting the threading fails the two renamed
cases and leaves the negative and the builtin regression floor green.

The fixture asserts itself — an unpaused or unaged card produces no
signal for reasons unrelated to the column, which would let the suite
pass while testing nothing.

## Verification

- new suite 4/4
- full core PG: **1048 passed / 3 failed** — the same three that
reproduce with this change stashed
- core `tsc --noEmit` clean; `pnpm test:gate` green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:48:08 -07:00
gsxdsm
f6010ef558 fix(test): planning-lane E2E is RED on main (2/7, incl. its own control) — same unplanned-spec fixture defect #2634 fixed next door (#2658)
Found while establishing the pre-closing E2E baseline for the final
verification pass (closing bar, item 4). **Two of this family's seven
cases are failing on `origin/main` right now**, and one of them is its
own control.

```
releases an ordinary held card on a default board (the control)       expected [] to include 'FN-OK'
holds a card parked for approval MID-SWEEP, after the snapshot read   expected false to be true
```

## Cause — the same defect #2634 repaired in the file next door

`seedHeldTask` creates the task and never writes a `PROMPT.md`, so the
card carries only the bootstrap seed. FN-7648's
`isUnplannedForExecution` reads that file for any card resting in an
intake- or hold-trait column and refuses to move an unplanned card into
a processing column, so the sweep released nothing.

**Being held was the gate working.** The fixture was exercising the gate
rather than the sweep — which is exactly why the *control* failed, and a
failing control means the rest of the family's assertions cannot be
trusted either.

`workflow-lifecycle-live-e2e` had the identical problem and #2634 fixed
it the same way. This file landed alongside it (#2611) and did not get
the same treatment. Worth stating twice because it is a general rule for
this directory: **a release/scheduler fixture that does not model a card
which cleared specification is testing the gate, not the sweep.**

## The check that matters more than the fix

#2611's stated value is "3/7 red without the guard". Making red tests
green is the easiest thing in the world to do wrongly, so I verified the
family still discriminates *after* the seed — disabling
`isTaskBlockedOnApproval` in `hold-release.ts` still kills exactly
three, and the same three:

| killed by mutation |
|---|
| does NOT release a card blocked on manual plan approval on a
**default** board |
| does NOT release a card blocked on manual plan approval on a
**renamed** board |
| holds a card parked for approval **MID-SWEEP**, after the snapshot was
read |

Two cases turned green, zero discriminating power lost. Without that
mutation this change would be indistinguishable from weakening the tests
until they passed, which the standing rule forbids.

Note the third killed case is also one of the two that were failing: it
was red for the fixture reason **and** genuinely proves the guard.

## Why it is worth a PR of its own

The closing bar's final verification pass (gate, `verify:fast`, all E2E,
census) has to run on a green tree. Two red E2E cases on main would
otherwise show up in that report as a new failure and cost a diagnosis
at exactly the wrong moment.

Pre-closing baseline for the record — **13 E2E families, 109 tests,
these 2 the only failures**:

```
green  agent-count 13 · agent-link 5 · lease-rebound 6 · lifecycle 24 · merge-family 7
       merge-rebound 4 · merge-safeguards 10 · merged-board 5 · planner-lane 5
       planner-lane-resolution 3 · rebound-family 15 · stranded-column 5
RED    planning-lane 7  (2 failing)
```

## Verification

7/7 green, mutation 3/7 as designed, engine typecheck clean (0 lines),
`pnpm lint` exit 0, `pnpm test:gate` exit 0.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:47:43 -07:00
gsxdsm
cef1b08af3 U12: the census baseline follows the count down — and goes in the merge gate (#2661)
Coordinator item 2. The census had the right mechanism and no teeth.

## The gap

`--strict` already fails on a rise **and** on an unrecorded drop — that
logic was correct. But nothing blocking ran it, so the baseline drifted
to **854 while the tree held 787**. That is **67 guards of regression
that would have merged silently**: a high-water mark wearing a ratchet's
name.

This is the same shape as the ceilings I tightened in #2647, one level
up. Worth saying plainly: I fixed the vitest ratchet's slack by hand and
did not check whether the *authoritative* instrument had the same
problem. It did, and by a much larger margin.

## Three changes

1. **`--strict` runs in `test:gate`.** The baseline cannot go stale
again without a red gate.
2. **Baseline re-recorded: 854 → 785** across 14 files (`triage` 38 →
9).
3. The single RISE is resolved honestly rather than absorbed.

## The +3 investigation

One file rose: `register-task-workflow-routes.ts` **22 → 23**. #2621
replaced one `task.column === "todo"` with `task.column === "triage" ||
task.column === "todo"` — a net **+1** that also reintroduced a `triage`
literal, while the PR title reported *"count 0 → 0"*.

Not an accusation. There was no gate for the author to check against,
and a hand-counted claim in a PR title is exactly the thing that goes
wrong without one. Change 1 is the fix.

**The literal is justified and stays**, marked `DELIBERATE-LITERAL`
rather than converted. It is the **v1-IR arm**: a v1 workflow yields no
role assignments, so `resolveLifecycleColumns` returns nothing and the
legacy pre-implementation ids are the only pre-WIP signal available. The
`else` branch directly below already resolves intake/hold for every v2
workflow. Converting this arm would not finish anything — it would
delete the only answer v1 boards have and admit
`in-progress`/`in-review` cards into a rebound that clears worktree,
branch and retry counters, which is the regression #2621 was fixing.

## Both directions proven

| direction | probe | result |
|---|---|---|
| rise | add `t.column === 'in-review'` | `live-agent-count.ts: 6 -> 7`,
exit 1 |
| drop | convert one guard | `self-healing.ts: allows 111, tree has
110`, exit 1 |

**The drop probe took three attempts to test honestly, and the first two
"passed" while proving nothing:**

1. I renamed a receiver (`task.column` → `Probe`) — the classifier is
**fail-closed**, so an unknown receiver is still counted and the number
never moved.
2. I targeted a site in `hold-release.ts` that carries a
`DELIBERATE-LITERAL` marker — not counted as a column guard at all, so
removing it changed nothing.

Only removing a counted comparison outright moved the number. Both false
negatives came from me assuming the probe worked because the command
exited the way I expected.

## On auto-rewrite vs fail-and-instruct

You offered either. The script already does **fail-and-instruct**, with
`--update-baseline` as the explicit re-record, and I kept it that way
rather than making the test rewrite the baseline during a run.

Reason: a silent downward rewrite means a conversion PR's own diff never
shows the number moving, so "census before/after in the PR body" becomes
unverifiable — the reviewer would have to re-derive it. Failing with the
new number in the message puts it in the diff where a human sees it, and
it costs one command.

## Verification

`pnpm lint` clean. `pnpm test:gate` green with the census in it — `every
file matches its baseline exactly` (10 / 132 / 487 / 71).

Note for the fleet launch: with `--strict` gating, **every** conversion
PR must now re-record the baseline in the same PR. That is the intended
cost, and it makes the fleet's "baseline must shrink by exactly the
converted count" rule mechanically enforced instead of a review
instruction.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:47:31 -07:00
gsxdsm
2fb0df9da8 docs(replan-target): the flagged follow-up is done — the note said otherwise (#2665)
Comment only. The block above `resolveReplanTargetColumn` still reads:

> **STILL A REAL FOLLOW-UP** … the `return "triage"` fallbacks on the
no-match and throw paths name a column the default lineage no longer
declares … flagged rather than fixed

That directly contradicts the code six lines below it. **#2598 landed
the fix:** the no-match path now returns `roles?.hold ?? roles?.intake`,
and the throw path returns `undefined`.

I wrote that note. A stale *"not fixed yet"* sitting above a fixed
implementation is worse than no note — the next reader either distrusts
the code or re-does work that is already done. This is the closing-bar
item 3 I was assigned, and I nearly re-did it myself: I had the change
written and reverted before checking whether main had overtaken me.

## One thing worth recording about #2598's version

Its catch-path answer is **stronger than the one I had drafted**. I was
going to return `"todo"` — the better guess, since post-U11 the default
lineage declares `todo` and not `triage`. #2598 returns `undefined`
instead, which forces callers to handle "this workflow could not be
resolved" explicitly rather than papering over it with a plausible
column id that the move path may then reject.

That is the same lesson as the sync-reader audit in #2653: **a defective
lookup that returns a valid-looking answer is worse than one that admits
it does not know.** Recorded in the comment so the reasoning survives.

## Verification

Engine typecheck clean · **51/51** across both replan-target suites ·
comment-only, no executable change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:47:19 -07:00
gsxdsm
9dbc98f1b3 Audit: every sync workflow-IR read answers for the DEFAULT workflow (not a PG-only problem) (#2653)
Docs only. This came out of a #2593 review thread that reported the
problem as PostgreSQL-specific. **It is unconditional**, and it has
consequences well outside the guard I was fixing — including one that
looks like a live production break for custom workflows.

## The chain, each link checkable

1. `TaskStore.getTaskWorkflowSelection(taskId)` delegates straight to
`getTaskWorkflowSelectionImpl` — **no mode branch** (`store.ts:2545`).
2. `getTaskWorkflowSelectionImpl` **returns `undefined`
unconditionally** (`workflow-definitions.ts:505-512`). Its own comment:
*"sync selection reader is incomplete-PG; use
getTaskWorkflowSelectionAsync."* A PG-cutover stub that never got
finished.
3. So `resolveTaskWorkflowIrSyncImpl` always takes its `if
(!workflowId)` branch and returns `resolveDefaultWorkflowIr()`. Its
`isBuiltinWorkflowId` and `SELECT ir FROM workflows` branches are
**unreachable in production**.

`resolveTaskWorkflowIrSync` is typed `WorkflowIr`, non-optional — so
callers cannot detect the substitution. There is no `undefined` to check
and the IR that arrives looks valid.

**Why tests don't catch it:** test stores stub
`getTaskWorkflowSelection` with a real selection, so the reader works
under test and substitutes only in production. Any test written against
a stubbed store proves the caller's logic and never the reader's
behavior.

## Consequences, severity descending

1. **Custom fields appear to be rejected on custom workflows.**
`resolveTaskCustomFieldDefsSyncImpl` returns `ir.fields` — the DEFAULT
workflow's. `task-update.ts:128-136` validates against them, and its own
comment states the outcome: *"a write against a workflow with no fields
(the default) is rejected with a typed CustomFieldRejectionError."*
2. **Per-workflow capacity pools collapse** —
`resolveEffectiveWorkflowIdSyncImpl` reads the same selection, so every
task resolves to `resolveCapacityPoolId(undefined)`.
3. **Plugin transition hooks re-run against the wrong IR**
(`lifecycle-ops.ts:1052`, crash recovery).
4. **Terminal-node detection degrades** to `nodeId === "end"`
(`branch-and-pr-entities.ts:578`).
5. **A U7 guard was inert** — fixed in #2593. Its fail-closed arm was
`workflowIr ? … : true`, dead code against a non-optional return.

**#1 and #2 are REASONED FROM SOURCE, NOT OBSERVED.** I did not execute
those paths, and I am labelling them that way in the doc rather than
reporting them as confirmed. No test in `packages/core` covers
`CustomFieldRejectionError` or `resolveTaskCustomFieldDefsSync` —
consistent with the gap, but absence of a test is not proof of a break.
**Reproduce before fixing.** I would rather hand you a labelled
hypothesis than a confident claim I did not verify.

## Why this matters for the fleet, specifically

The census work replaces column literals with trait lookups. A
conversion that resolves its traits through a **sync** reader produces a
guard that reads the DEFAULT workflow's traits for every task —
plausible, wrong, and invisible. **It converts a visible literal into a
hidden bug**, and the ratchet counts it as progress.

Suggested addition to the fleet brief: conversions must resolve through
`resolveWorkflowIrForTaskWithProvenance` and branch on `source`;
`resolveTaskWorkflowIrSync` is never acceptable in a converted guard.

## Not fixed here

Each consequence needs its sync call path made async — a real slice per
site, not an end-of-turn edit. #2593 fixed only the one that was mine.
Census unchanged (781 / triage 5); this PR adds and converts no guards.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added an architecture-pattern finding documenting a workflow-reading
limitation that can cause synchronous reads to use the default workflow.
* Described resulting effects on custom workflow updates, crash
recovery, capacity-pool handling, and terminal-node detection.
* Documented testing gaps and guidance to avoid synchronous task
workflow reads.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:37:59 -07:00
gsxdsm
174eb22534 cleanup: delete the dead sync capacity-pool helper rather than document it (#2656)
Follow-up to the #2653 audit, and the one item there that is better
deleted than described.

## Why delete rather than annotate

`resolveEffectiveWorkflowIdSync` **has no callers.** Verified across
every `.ts`/`.tsx` in `packages` (excluding `dist`): only its own impl,
the `store.ts` import and public method, and one comment naming it. Not
exported from the core index, not referenced by any test.

It is also **wrong**. It reads `getTaskWorkflowSelection` — the sync
selection reader that has returned `undefined` unconditionally since the
PG cutover — so it always resolved `resolveCapacityPoolId(undefined)`:
the default pool for every task, regardless of workflow. The binding
capacity path reads the selection asynchronously inside its transaction
and does not use this.

That combination is the argument. A dead function is clutter; a dead
function that returns a **plausible wrong answer** is a trap. The next
person to need "which capacity pool is this task in?" would find a
public method with exactly the right name, call it, and get default-pool
behavior with no signal that anything degraded. #2653 documents it, but
documentation loses to autocomplete.

## Provenance of the claim

greptile's P2 on #2653 corrected my first draft, which called this a
live capacity collapse — it isn't, precisely because nothing calls it. I
verified the no-callers claim myself before accepting, and this PR is
the logical end of that correction: if it is unreachable, it should not
exist.

## Removed

- the impl in `task-store-helpers.ts`
- the `resolveEffectiveWorkflowIdSync` public method on `TaskStore`
- the import specifier in `store.ts`
- the now-unused `resolveCapacityPoolId` import (its only use was the
deleted function)
- updated the `workflow-definitions.ts` comment that named it

## Verification

core / engine / dashboard typechecks clean · eslint clean on both
touched files · `pnpm --filter @fusion/core build` exit 0 · **`pnpm
test:gate` green (487 + 71)**.

The engine and dashboard typechecks are the ones that matter here:
removing a public method from `TaskStore` would surface immediately in
any consumer that called it, and neither reports anything.

## Census

Unchanged (776 / triage 5) — no guards added or converted.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:37:46 -07:00
gsxdsm
b6b2fdcdc6 test(U7): rescue the orphan-triage regression test — main has the fix but not its test (#2663)
Main already carries **every other artifact** from #2593 — the
provenance fix, the `DELIBERATE-LITERAL` markers in
`TaskCard`/`TaskDetailModal`/`register-routes`, the audit doc. The one
thing missing is the test.

That is the same artifact class that vanished when #2645's branch was
force-pushed, so I rebased #2593 onto current main, found every commit
conflicting because the work had landed by other routes, and rescued the
one piece that had not.

**#2593 can now be closed** — it carries nothing else main lacks.
**#2654 needs rebasing onto main** rather than stacking on it.

## What makes this test worth rescuing

It took three attempts to write honestly, and the reason is pinned in
the test body: on a bare mock, `resolvePlannerLanes` reads
`resolveTaskWorkflowIrSync`, which the mock does not define, so it
returns `LEGACY_PLANNER_LANES` (`intake: "triage"`) and a `triage` card
matches the **first** arm — the orphan arm is never reached. Every
earlier fixture I wrote passed through that short-circuit and proved
nothing.

All three cases stub that reader with the merged default (`intake:
"todo"`), which is what production resolves, leaving the orphan arm as
the only thing deciding. They differ **only** in the workflow readers.

| case | role |
|---|---|
| **C** — workflow declares `triage` as a review lane | **the
discriminator.** Pre-fix, the sync reader ignores the selection, returns
the default IR declaring no `triage`, so the arm fires and a card is
finalized out of a custom workflow's code-review column |
| **B** — workflow resolves, declares no `triage` | positive control;
without it "returns false" is unfalsifiable |
| **A** — workflow unresolvable | **behavior pin, NOT a regression
test** — passes in both worlds |

I had A labelled "REGRESSION" until the mutation said otherwise. It is
relabelled with the null result documented, because a future edit making
it flip would mean the arm's scope changed.

## Verification, stated precisely

**231/231** against main's implementation.

The mutation that proved C discriminates was run on the branch where the
pre-fix code still compiled. **It cannot be re-run against main**: the
`WorkflowIr` type import was removed along with the fix, so a naive
revert no longer transforms. I am stating that rather than implying I
re-verified it here — the discrimination was demonstrated, just not on
this base.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:37:20 -07:00
gsxdsm
efbbc45eb0 U12: the LAST triage guard — Plan was offered on executing cards named triage (#2664)
The final `column === "triage"` in production source, and it was a live
defect rather than dead vocabulary.

## The defect

`isPreExecutionHoldColumn` ORed the legacy id with the traits
**unconditionally**:

```ts
return column === "triage" || flags?.intake === true || flags?.hold === true;
```

That is not a fallback. A resolved column merely *named* `triage`
answered true even when its own traits said work was underway — so the
context menu offered **Plan**, which re-plans, on a card that is already
executing.

Now flags-first, with the id as the documented no-metadata answer.

## Why the file's earlier conversion missed it

Every existing case in `TaskContextMenu.test.tsx` passes a column with
**no flags**, or with `hold`/`intake` set. All of them agree under both
forms, so the suite could not distinguish them. Nothing exercised a
column whose **name and traits disagree**, which is the only shape that
separates an OR from a fallback.

Three new cases cover it. Revert check: restoring the OR form fails the
first one — Plan reappears on a mid-flight card.

## The asymmetry is preserved, and now tested

The degraded set stays `{triage}` **alone**, deliberately not the
`{todo, triage}` used by `isPreImplementationColumnRole`. That helper
drives the preserve-progress prompt, where a flagless `todo` *should*
prompt because losing steps is unrecoverable. This drives Plan, where a
flagless `todo` must **not** offer to re-plan a card that may already be
planned. The file documented that difference; nothing asserted it. Now a
test does.

## On reaching zero honestly

The surviving literal is marked `DELIBERATE-LITERAL`. It is the degraded
answer, not an unconverted guard — there is no trait to read when
`flags` is `undefined`, which happens during first paint and for a card
in a column its workflow no longer declares. Deleting it would silently
withdraw Plan from exactly the stranded cards that most need
re-planning.

So **`triage → 0` means "no unconverted guards remain", not "the string
is gone"**, and I would rather say that than move a number by deleting a
fallback.

| branch | triage |
|---|---:|
| `origin/main` | 5 |
| this PR | **4** |
| #2655 (flag resolution, removes 4 in `moves.ts`) | 1 → **0** combined
|

I found it with the census's own AST classifier rather than grep — my
grep of the same tree returned only comment prose and would have had me
report the bar as met while a real defect sat in
`TaskContextMenu.tsx:179`.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0 with the baseline re-recorded in this
PR (column 769 → 768, deliberate 12 → 13). `tsc -p tsconfig.app.json`
clean. `TaskContextMenu.test.tsx` 18/18.

Depends on nothing; stacks cleanly with #2655 and #2661.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:37:08 -07:00
gsxdsm
3bf9bf5f74 collapse the plan-admission-throttle payload to one gate (+ AGENTS.md) (#2562)
The cross-project semaphore is deleted, so
`task:plan-admission-throttled` was describing a gate that no longer
exists. Nothing wires `options.semaphore` any more, which left three
things dead-but-visible:

- `semaphoreAvailable` was permanently `Infinity`, so
`Math.min(projectRoom, …)` was a no-op keeping a deleted limiter in the
arithmetic
- `blockedBy` was a **discriminator** between `"running-agent cap"` and
`"global semaphore"`; only the first can occur
- four `semaphore*` metadata fields were always `undefined`, and two
more terms in the dedupe signature were constant

## `blockedBy` is kept, not dropped

Even though it is now a constant. The event exists (FN-8600) to answer
*“why did this card sit queued to plan?”* after the fact — a named
reason answers that even when there is one gate, whereas a payload with
**no** reason field reads as “unknown”. It costs nothing and preserves
the shape if a second gate is ever added.

The dedupe signature drops the two semaphore terms and keeps the
eligible task IDs — that term is what stops a **new** card’s stall being
swallowed when the counts land on an unchanged tuple, which is the
property the event depends on.

## AGENTS.md

It documented the removed field names verbatim, so it is updated in the
same commit. Leaving docs describing a payload the code cannot emit is
exactly the readable-but-wrong artifact this program keeps deleting.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green · triage
suites **234/234**.

---

**Correction I owe on `concurrency.ts`, measured rather than
estimated.** I earlier told the coordinator ~75% of its 886 lines could
go with the cross-project cap. That was line-range arithmetic and it was
wrong. With the cap now fully removed, `concurrency.ts` is **still 886
lines**, because `AgentSemaphore` has four consumers unrelated to it —
`verification-concurrency` (maxConcurrentVerifications),
`research-orchestrator` (research runs), `experiment-executor`
(maxConcurrentExperiments), `step-session-executor` (parallel steps) —
plus `ProjectAdmissionCoordinator`, which is FN-8453 oldest-first
**ordering**, not a limiter. The real remaining win there is the
pre-held-slot bookkeeping and the idle-semaphore leak recovery, which
existed to service the global instance; I will measure that as its own
slice rather than quote a fraction.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Updated plan admission throttling to consistently use the project’s
running-agent capacity.
* Improved throttle audit events by reporting stable capacity details
and removing obsolete semaphore information.
* Preserved accurate deduplication for repeated throttling events,
including changes in stalled tasks.

* **Documentation**
* Updated run-audit guidance to match the revised throttling event
format.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:31:27 -07:00
gsxdsm
8393bba7dc U7: the replan rebound targets a column the workflow declares (R7) — re-landed on main (#2598)
> Based on `main`, no dependencies. Re-landed after closing the stacked
chain (#2517, and #2551 below) that never reached main.

## What main already has, and what it lacks

Main independently converted the planner-lane **parameters** in this
file — and **better than I had**: it splits `plannerColumn` from
`roles.mergedPlanningColumn`, because a merged lane joins the FN-8596
arrival-order rescue but *not* the "planner column is never advanced"
shortcut. That work is main's and untouched here.

What main still lacks is the **R7 fix**: `resolveReplanTargetColumn`
returns `"triage"` **by fiat** for any workflow declaring neither legacy
id — `builtin:marketing` (ideation/backlog/drafting/…) and every fully
renamed set. A Plan Review REVISE therefore moves the card into a column
its workflow **does not declare**, for `reconcileUndeclaredTaskColumns`
to clean up after. A move the engine makes on purpose, not drift.

| Workflow | Target | Changed? |
|---|---|---|
| `builtin:coding` / stepwise | `todo` | no |
| Coding (Ideas) | `todo` | no |
| `builtin:marketing` | `backlog` (its own hold) | **yes** — was
`triage`, undeclared |
| declares no planning lane | `undefined` → park | **yes** — was
`triage` by fiat |

## The ordering the existing suite taught me

Legacy ids stay preferred **first**, and the trait resolution prefers
**hold over intake**. That is not arbitrary:

Coding (Ideas) declares `ideas` as its intake, and `ideas` is **manual
capture with no AI** (plan R10) — a rejected plan sent there stops being
replanned at all. The old code got Ideas right **by accident**: it never
recognised `ideas` as intake and fell through to `todo`. An "intake
first" trait rule would have shipped that regression dressed as a
cleanup, and three existing Ideas tests were the only thing between me
and doing it.

## An inverted comment, corrected

The function's own U11 note read: *"the second lookup asks for `todo`,
which U11 deletes… the first lookup still matches `triage` (which U11
keeps)"*.

**That is backwards.** #2515 keeps `todo` and deletes `triage`, so the
consequence is the opposite of what was written — the `todo` branch is
what saves builtin coding. Fixed rather than left, because a comment
that inverts a merge's direction sends the next reader to the wrong
branch.

## Fail-closed callers

`undefined` means "nowhere to replan" (plan U5: *skipped with a log
rather than moved arbitrarily*). All four call sites park **visibly**
rather than log a move they did not make. The scheduler's rebound still
writes `needs-replan` — deliberately, since that is what blocks dispatch
and the branch has already decided the card must not be released — with
only the *log* made conditional.

## The superseded test is deleted, not skipped

A skipped test is a guard that cannot fire. Its replacement asserts the
new contract **and** the R7 invariant directly — *"a column this
workflow declares"*, not just an id — plus a new case for a workflow
with no planning lane at all.

## Verification

| Check | Result |
|---|---|
| replan-target | 41/41 |
| with scheduler-trait-dispatch + pre-release-plan-review | 55/55 |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (482 + 10 + 71) |
| `pnpm check:changesets` | clean |

`triage.test.ts` still shows main's **8 pre-existing #2515 failures** —
unchanged by this, fixed by **#2576**.

## Closing #2551

Its parameter work is superseded by main's better version; this PR
carries the only part main lacked. Same story as #2517: a stacked PR
outlived the surface it was converting.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:31:03 -07:00
gsxdsm
f8c053c3fa fix(core): TAKING comments-ops.ts — re-triage on renamed planner lanes (3 triage guards → 0) (#2612)
**Claiming `packages/core/src/task-store/comments-ops.ts`** from the
shared backlog so nobody collides.

## Guard count

| scope | before | after |
|---|---|---|
| `comments-ops.ts` | **3** | **0** |
| repo-wide `column === / !== "triage"` in `packages/*/src` (excl.
tests) | **26** | **23** |

## Why this file, and why it matters more than its size

`addComment`'s post-comment **re-triage** decides, from the card's
column, whether a user comment should invalidate an approved spec or
send already-planned work back for re-specification. It asked with three
legacy literals:

```ts
task.column === "todo" || task.column === "triage"
task.column === "triage" && status === "awaiting-approval"
hasRealPrompt && (todo || (triage && status !== "awaiting-approval"))
```

On a renamed board none match, so a user comment on planned work does
**nothing**: no approval invalidation, no re-specification, no error.

**The operator types a correction and the agent never sees it.** This is
the surface a human actually touches, which makes it the worst place in
the program for a silent guard.

Now resolved per task via `resolveLifecycleColumns`, fail-soft to the
legacy pair — this phase is documented best-effort (*"failures are
logged but never fail the comment add"*), so an unresolvable workflow
must behave exactly as before rather than skip re-triage.

## Red-green, not green-only

The suite was written **first** and failed **3 of 6** against the
literals — precisely the three renamed cases — while the two negatives
and the default-vocabulary floor passed throughout.

Both negatives earn their place: re-triaging a **WIP** card would
discard an in-flight session, and the **author gate** (agent comments
must not re-triage) has to survive the conversion.

## Fixture guards its own preconditions

`PROMPT.md` is written where the guard reads it rather than relying on
task creation's side effects. `hasRealPrompt` gates two of the three
branches, so a bootstrap stub would make those cases pass for the wrong
reason — the trap that has produced two vacuous tests in this program
already.

## Verification

- new suite 6/6; `store-comments` 14/14
- full core PG: **1050 passed / 3 failed** — the same three that
reproduce with this change stashed (`central-archive-secrets`,
`workflow-settings-project-identity`)
- core `tsc --noEmit` clean; `pnpm test:gate` green (482 + 132 + 10)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Improved comment-driven re-triage for workflows with renamed planning
columns.
- Comments on planned or awaiting-approval tasks now correctly move
eligible tasks to “Needs re-plan.”
  - Prevented re-triage for tasks actively in progress.
  - Preserved existing re-triage behavior for standard workflow columns.
  - Non-user comments no longer incorrectly trigger re-triage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:27:43 -07:00
gsxdsm
cf6133da8b consolidate/e2e — E2E evidence: already-finalized terminal roles (real merge entry, no git) + ledger corrections (#2648)
Consolidation branch for the E2E-evidence worker. Two commits, both
engine test/comment only — **no production code, census unchanged**.

## Census (the authoritative instrument)

`node scripts/lifecycle-column-census.mjs` on this branch: **triage 10,
total 784** — identical to its base.
`lifecycle-column-census-ast.test.ts` and
`lifecycle-column-census.test.ts` pass (15). This PR neither shrinks nor
grows the backlog; it is evidence.

## What it contains, file by file

| file | change |
|---|---|
|
`packages/engine/src/__tests__/workflow-already-finalized-live-e2e.pg.test.ts`
| **new** — 3 cases, live PG store + real `runAiMerge` |
| `packages/engine/src/__tests__/workflow-lifecycle-live-e2e.pg.test.ts`
| comment only — retires two unproven-ledger entries |

## The evidence: `isAlreadyFinalizedColumn` never needed the real-git
lane

My unproven-sites ledger listed it as requiring a git harness because it
is module-private inside `runAiMerge`. Reading the function instead of
costing the lane: `runAiMerge` reaches it after only `store.getTask`, a
pure workspace assert, and a pure branch resolve — **before** the merge
blocker, settings, and any branch sync. `projectRootDir` is never
touched on that path, and the short-circuit returns a `noOp` rather than
throwing. Reachable through the real public entry point with no
repository at all.

### Why two cases and not one

The guard resolves terminal columns **per role**:

```ts
terminal = [lifecycle.complete ?? "done", lifecycle.archived ?? "archived"]
```

#2471's P1 caught the first cut replacing the whole legacy **pair** as
soon as *any* terminal role resolved — a workflow declaring `complete`
but no `archived` collapsed to one element, silently lost the archived
short-circuit, and an archived card then threw *"must be in
'in-review'"* for a card whose real state was "already done, nothing to
do".

A per-set rule passes for whichever role **is** declared and fails the
other, so a single case cannot tell the two rules apart. The shared
fixture declares `complete` (renamed `shipped`) and **no** `archived`,
so it is exactly that partially-declared shape — resolved half and
fallback half live on one board.

Mutation-verified, each killing only its own case:

| mutation | kills |
|---|---|
| per-**set** replacement (the #2471 defect) | the legacy-`archived`
fallback case |
| legacy pair only (conversion reverted) | the renamed-`shipped` case |

Plus a differential: a renamed **review** card must not report
already-finalized. Without it both cases above would pass for a guard
that finalizes everything — turning every merge into a silent no-op, the
worst failure this function has.

Evidence strength is stated in the file header rather than overclaimed:
this reads a returned **decision**, not a persisted row, so it proves
the renamed board resolves and short-circuits — not that a card moves.

## Ledger corrections (comment only)

Two entries retired, both wrong the same way — each stated a **lane
cost** as if it were an impossibility:

- `columnIsIntakeOrHold` — "consumers are dashboard-side" is true and
irrelevant; its one consumer is an exported pure function. Proven on
merged and renamed boards by work already merged in #2631.
- `register-task-workflow-routes.ts` — "standing up the route shell is
mock-the-world" was false; `createApiRoutes` + `test-request.js` is this
repo's established convention with ten existing suites for that file.
Covered by its owner in #2614.

Counting this PR's own subject, that is **seven** wrong lane-cost
inferences in that ledger. The rule it keeps violating is unchanged and
now recorded in the file: read what the FUNCTION touches before costing
a lane for it.

## Verification

`pnpm test:gate` exit 0 (695), `pnpm lint` exit 0, engine typecheck
clean, 28 tests green across the AST ratchet and the three
merged-board/planner-lane/already-finalized families.

## Not in scope here

The `performWorkflowRerunBounce` E2E. The harness exists (`new
TaskExecutor(store, "/tmp/test", {})`), but `executor.ts:4305` gates the
rebound on the legacy `in-progress`/`in-review` pair while resolving its
target by role — so on a renamed board the bounce never fires and the
resolved target is unreachable. An E2E asserting today's behaviour would
cement that. It is executor.ts's owner's fix; evidence should follow it.
Detail in #2632's thread.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:27:30 -07:00
gsxdsm
20878e9d5f census: count column: "<legacy>" query filters as a separate, separately-pinned instrument (backlog unchanged at 784) (#2650)
Pre-launch input for the 779-guard fleet. **The backlog number does not
move: 784 before, 784 after.** This adds a second number beside it.

## The problem it measures

A guard is not the only way a legacy column id decides behaviour:

```ts
const todo = await this.store.listTasks({ column: "todo", slim: true });
```

That is a **source query** — it selects the rows a sweep considers *at
all*. On a renamed or merged board it returns nothing, so a sweep whose
per-task predicate was correctly converted still does nothing, while
looking converted. `self-healing.ts:2849` names the pairing in prose,
and #2560 had to repair exactly that combination after a converted
predicate was left with a literal query.

The census walks comparison `BinaryExpression`s. A `PropertyAssignment`
is not one, so this class was invisible to the instrument **and to its
ratchet** — it could grow silently.

Measured: **83 query filters, 43 IR node definitions.**

I proved one live consequence earlier on #2648:
`recoverStuckMergeDeadlocks` cannot see a renamed board at all — the
renamed rows exist and none appear in its three-literal union
(`renamedInsideUnion=0`, on a live PG store).

## Why this matters *before* the fleet is briefed

The fleet rule is *"the baseline ratchet must shrink by exactly the
converted count."* In `self-healing.ts` — the largest batch at 111 —
both classes sit in the same functions, so today a worker either:

- converts only the comparisons → arithmetic is clean, and sweeps whose
source query still filters a dead literal stay blind; or
- converts the query too → the count does **not** move by the converted
amount, and a more-correct PR looks like a miscount.

The second punishes the better worker. With a second pinned number,
converting a query becomes visible work instead of an apparent error.

## Counted separately, deliberately

`totals.column` is a published shape — the baseline, the reporter, and
other workers' in-flight PRs read it, and the completion bar is defined
against it. Growing it would move a number the program is actively
driving to zero.

So the new counts live in `summary.properties` / `queryByFile`, under
their own baseline keys, with their own both-directions ratchet (same
rule as #2633's, including the stale-allowance half). `totals` keeps its
**exact** shape — two existing tests assert it with `toEqual`, and
breaking a contract others depend on mid-flight to add a number is not
worth it.

## Definitions are not queries

Workflow IR graph nodes carry `column:` to declare where a node lives —
`{ id: "review", kind: "...", column: "in-review" }`. That is the
lineage describing itself: not a lookup, not convertible, and ~43 of the
raw matches. They are told apart **structurally** (an `id`/`kind`
sibling in the same object literal), not by filename, so a definition
written anywhere classifies the same way.

## Baseline seeding, stated plainly

`--update-baseline` could not pin a **new** category: the regression
check runs before the write, and with no prior key every file reads as a
rise. I seeded the three new keys once, directly, leaving every guard
field byte-identical. The diff is purely additive — no removals.

## Finding, not caused by this change

**`--strict` is already red on clean main**:
`register-task-workflow-routes.ts` is **23** against a baseline of
**22**. Verified by stashing this branch and re-running on an unmodified
tree. Until that is reconciled the guard ratchet is passing nothing —
worth fixing before the fleet starts relying on it as the work order.

## Verification

- census suites **44 green**, 6 new cases: counted; kept out of the
backlog; definition-not-query; both instruments independent (a bug
routing comparisons into the query bucket would otherwise look clean on
both); `DELIBERATE-LITERAL` honoured; non-legacy id ignored
- `node scripts/lifecycle-column-census.mjs` → backlog still 784
- `pnpm lint` exit 0, `pnpm test:gate` exit 0 (695)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:27:18 -07:00
gsxdsm
2771408bba ci: enforce the lifecycle-column ratchet — it has never actually run (#2654)
**The ratchet was advisory.** `scripts/lifecycle-column-census.mjs`
existed only as `pnpm census:lifecycle-columns` — without `--strict` —
and **no workflow invoked it**. Nothing has ever compared the tree to
the baseline. Every "the baseline ratchet holds them" assumption in this
program rested on a check that does not run.

That explains both classes of hole:

**1. Three PRs lowered counts without re-recording,** leaving allowances
the deleted guards could return through while every check stayed green.
I've tightened them across #2593 and earlier PRs, but nothing stops the
next one.

**2. #2621 GREW the count while its own title claimed "count 0 → 0".**
It added `column === "triage"` and `column === "todo"` at
`register-task-workflow-routes.ts:2681`, taking that file to **23
against an allowance of 22**. It landed unchallenged. This is the
failure mode the ratchet exists to prevent, and it happened *inside this
program*, in a PR that asserted the opposite.

## The change

Adds `check:lifecycle-columns` (the census with `--strict`) to the
`pr-checks.yml` lint job, next to `check:changesets` and
`check:routes-modular` — the established pattern. **~1.8s over ~1950
files**, so this is not a slow-test addition.

## Proven to fail, in both directions

A guard that reports success without checking anything is worse than no
guard, so:

| injected defect | result |
|---|---|
| `const __probe = (c: string) => c === "triage"` added to `moves.ts` |
`count ROSE — moves.ts: 39 -> 40`, exit 1 |
| run against main's current baseline | exit 1 on
`mission-feature-sync.ts: allows 5, tree has 0` |

Both reverted; exit 0 restored. Note the second row: **this check is RED
on main right now**, which is the point.

## Merge order

**Stacked on #2593**, which carries the `DELIBERATE-LITERAL` marker for
the #2621 site (a v1 IR declares no roles, so no trait can answer that
question) plus the baseline re-record. Standalone on main this PR is red
— correctly. **Merge #2593 first**, then this.

I stacked rather than duplicating those two edits because I already
caused one conflict today by appending related content from two
branches, and #2651 merged a correction ahead of the section it
corrected. Same-content edits in two PRs is the same mistake.

## Census

Unchanged by this PR: **776 total, triage 5, reviewed 16** — it adds no
guards and converts none. It only makes the numbers enforceable.

## For the fleet

This should land before the 776-guard fleet launches. The brief says
"the baseline ratchet must shrink by exactly the converted count" —
until now nothing verified that claim, so a batch worker could report a
shrink that did not happen, or grow the count while converting, and CI
would agree.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:20:31 -07:00
gsxdsm
642a4fa264 consolidate/u12 — U12 consolidation: 4 live defects, the AST ratchet fail-closed, and the moves.ts flag scoped (#2647)
One branch, one PR, per the consolidation directive. Contents
file-by-file below.

**Supersedes #2625** (its overlapping conversions landed via U11's
#2624/#2626/#2636; only the parts nobody else did are folded here).
**#2630 and #2639 stay open** — both green with zero threads, per rule
3.

## Four live defects, each measured

**1. Every planning card renders an actions menu.**
`TaskContextMenu.tsx` still had `shouldShowActionsMenu: task.column !==
"triage"` on main *after* the rest of that file was converted. Since
#2515 removed the id, the condition is TRUE for every card, so the
suppression stopped applying anywhere — including on cards whose menu is
empty, the orphaned click target the Surface Enumeration rule exists to
catch.

Found **twice independently**: by reading the guard, and again by the
invariance test below, which failed on main with `shouldShowActionsMenu`
true on one lineage and false on another. That is the argument for an
invariance property over per-site conversion — the file had already been
converted "2 → 1" and the survivor was the live one.

**2. Worktree upcoming-work list empty on renamed boards.**
`groupByWorktree` filtered `t.column === "todo"`. On the default board
the id and the role coincide so every existing test passed; renamed, it
matched nothing and a whole panel read as idle.

**3. Hold-lane FIFO ordering lost on renamed boards.**
`sortTasksForDisplayColumn` gated priority-then-FIFO on `column ===
"todo"`, degrading to the generic id-ordered sort elsewhere. Cards
simply appear in the wrong order, silently.

**4. The AST ratchet still failed open** — fourth time in that file,
third found by review. `receiverName` understood only one-level property
access and bare identifiers, so `task["column"]`, `metadataColumn(entry,
"to")`, ternaries, `(task!.column)` and backtick literals were dropped.
**Measured on main: `in-progress` 196 → 197, `in-review` 211 → 213** —
three real guards nobody counted, including `metadataColumn(entry, "to")
=== "in-review"` in `reliability-metrics.ts`. Now walks wrappers,
resolves calls to the callee name, and emits a `<SyntaxKind>`
**sentinel** for anything unnameable: counted *and* trips the
classification guard, so a human judges it instead of it vanishing.

## Per-file guard counts

| file | before | after |
|---|---:|---:|
| `app/components/TaskContextMenu.tsx` | 1 | **0** |
| `app/utils/worktreeGrouping.ts` | 1 | **0** |
| `app/components/taskSorting.ts` | 1 | **0** |

The other dashboard files I had converted reached 0 via U11's PRs; where
our work overlapped I took theirs during the rebase, including two
places where theirs was **stronger** than mine — they deleted Column's
unreachable quick-create arm outright (with fixtures migrated) where I
had converted it, and they verified the same `isPreExecutionHoldColumn`
degraded-set asymmetry I did, independently.

## Flip precondition: the moves.ts flag is scoped, not flipped

`move-target-declared-census.test.ts` answers precondition 2 with
measurement. 41 engine `moveTask` calls have literal targets — `todo`
27, `in-progress` 7, `done` 6, `archived` 1 — and **all four are
declared by the default lineage**, so the default board is not the
exposure. `triage` appears only in a comment noting `replan-target.ts`
used to hardcode it. My own grep had said `todo=29`; the AST says 27,
because grep counts comments.

The exposure is **custom** lineages: 20 of the 41 carry no
`recoveryRehome` and would reject with unknown-column post-flip; 21 are
exempt via the #1411 carve-out, which makes that carve-out load-bearing.

I did not flip the flag. It is six seams, not the `789`/`837` pair every
summary including mine described, and seam 2 turns on *new refusals*
rather than swapping equivalent implementations — a green suite says
nothing about that. #2639 pins the blast radius.

## Tests

- `column-role-id-invariance.test.tsx` — hold traits fixed, vary only
the column id across MERGED / LEGACY / RENAMED; every decision must
agree. Drives the real consumers, so a component keeping an inline
comparison fails it. Includes a unanimous-and-**false** case so it can't
be satisfied by a predicate hardwired to true. **This is the test that
caught defect 1 on main.**
- `worktreeGrouping.test.ts` — includes two cards both in a column named
`staging`, one hold and one not, asserting opposite answers. That
assertion is impossible under a board-wide column-id set, which is why
hold resolution is keyed per task via `getEffectiveTaskWorkflowId`
(#2625 review).
- `taskSorting.test.ts` — discriminates on the **tiebreak**, not
priority: both branches sort by priority, so my first version passed for
the wrong reason. Equal-priority cards whose `createdAt` order disagrees
with their id order.
- `no-hardcoded-lifecycle-columns.test.ts` — 16 detector cases: 11
shapes counted, 4 legitimate ignored, one asserting the sentinel path.

Revert checks, all run: menu suppression → diff names the field;
worktree → `expected [] to include 'FN-50'`; sort → `FN-2, FN-9` instead
of `FN-9, FN-2`; ratchet → the 3 recovered guards disappear.

## One site that should never be converted

`MissionControlPanel.tsx:46` — `{ id: "triage", match: (c) => c ===
"triage" || c === "signal" || c === "backlog" }` is a deliberate
name-similarity heuristic for the SDLC funnel; it matches synonyms and
folds unknown columns into an "other" bucket so custom columns still
contribute. Converting it changes what the funnel displays. Like the
`live-agent-count` fallbacks, it belongs in a documented floor — **the
ratchet's target is that floor, not zero.**

`DocumentsView.tsx:73` is convertible but the file has no column flags
at all, so a real fix means plumbing board-workflow metadata into a view
that doesn't fetch it — its own unit of work.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 482 / 71). `tsc -p
packages/dashboard/tsconfig.app.json` and `packages/core/tsconfig.json`
clean. Core ratchet + seam suites 24/24. Dashboard target suites 37/38 —
the one failure is the pre-existing `"Back to In Progress"` label
casing, confirmed identical on the base.

---

## Added after the initial push

**5. `TaskCard` lost inline editing on renamed boards; `TaskDetailModal`
kept it.** Still live on main: the modal resolved field editability from
traits in U10/R8, the card used a hardcoded `{triage, todo}` set with
**no trait path at all** — even though `taskColumnFlags` was already in
scope. On a renamed board the title was editable in the modal and the
pencil was missing from the card. Body moved unchanged into
`isFieldEditableColumnRole` so the two surfaces cannot drift again.

The veto traits are the substance: a column can legally carry `hold`
**and** a WIP or review trait, and a plain `intake || hold` check would
let an operator rewrite a description while a session executes against
it.

Coverage gap **measured, not assumed**: mutating `canEdit` back to the
hardcoded set left `TaskCard*` at the same failure count as the
unmutated run — nothing caught it. The four render cases assert the real
`aria-label`; that mutation now fails with `Unable to find an accessible
element ... name 'Edit task'`.

**6. The ratchet's target is a documented FLOOR, not zero** — and this
changes the completion bar.

Zero is not reachable, and chasing it means breaking working code. Two
categories are permanent, now protected as positive assertions so a
future sweep cannot "finish the job" by deleting them:

- `MissionControlPanel.tsx`'s `FUNNEL_STAGES` is a deliberate
**name-similarity** heuristic — it matches `signal`, `backlog`, `to-do`,
`ready`, `shipped` and folds unrecognised columns into an "other" bucket
so a custom board still contributes counts. It is not asking whether a
column has the intake trait; it buckets arbitrary column *names* for
display. Asserted on the **synonym list**, because the synonyms are what
prove it is name matching — if they disappear the site has changed
character and the exemption stops applying.
- `live-agent-count.ts`'s no-flags arm is reachable (a remote store is
deliberately given an empty flag map; a card in an undeclared column has
no flags at all) and deleting the literal makes such a card match **no**
arm, so the queued total silently under-reports a stranded card.

A count with an undocumented floor invites someone to drive it to zero.

**Not done, and why:** `DocumentsView.tsx:73` is convertible but that
file has no column flags anywhere, so a real fix means plumbing
board-workflow metadata into a view that does not fetch it — its own
unit of work, not something to smuggle into a conversion.

**Re-verified after these commits:** `pnpm lint` clean, `pnpm test:gate`
green (10 / 482 / 71), `tsc` clean on core and `tsconfig.app.json`, core
ratchet suite 26/26, `columnRoles` 10/10, `TaskCard.test.tsx` 384/386
(the 2 are pre-existing CSS assertions). `TaskDetail*` is 130 failed /
551 passed **both with and without** this change — verified by stashing,
so pre-existing and unrelated.

---

## Flag resolution: preconditions 1 and 2 are now DISCHARGED.
Precondition 3 is blocked, and by evidence.

**Precondition 1 — the side-effect equivalence proof — done.**
`moves-flag-equivalence.test.ts` runs the same journey under both flag
states against live PG and diffs the persisted row. **Result:
identical** — whole-row equality across 128 fields plus an equal timing
shape, over `todo → in-progress → in-review → todo → in-progress`.

That test was **wrong twice** before it meant anything, and both times
it was passing:

1. **It proved nothing.** `experimentalFeatures` is **global-only**, and
`moves.ts` reads `getSettingsFast()`, which filters global-only keys out
of the project layer. My `updateSettings` write was silently discarded,
`useWorkflow` was false in *both* runs, and the "proof" compared the
legacy path against itself. Found by stamping the flag-ON branch and
observing the test still passed. Now written via `updateGlobalSettings`,
and the helper **asserts the flag took effect** before the journey runs.
2. **The journey was forward-only**, so it never reached the reopen
hook's field resets (`status`, `error`, `blockedBy`, pause clearing) — a
mutation there passed. Extended with a backward move and a re-entry.

Mutation-verified after both fixes: stamping seam 3, and diverging the
reopen hook, each fail the comparison.

**Precondition 2 — done, and its answer is a blocker.** The census says
the default board is safe: all 41 literal engine move targets are
declared by the default lineage. But **20 of those 41 carry no
`recoveryRehome`**, so on a custom lineage that does not declare `todo`
/ `in-progress` / `done`, seam 2 would start rejecting them with
unknown-column. That is a user-facing break on custom boards, not a
theoretical one, and it is not fixed by the equivalence proof — seam 2
adds *new refusals* rather than swapping implementations.

**So the flip is one step away, and the step is not mine to take
alone:** those 20 call sites need to resolve their target from the
task's workflow (or justify `recoveryRehome`), and they live across
engine lanes in `moves.ts` caller territory — U2b/MAIN. Flipping before
that trades a dormant flag for broken custom boards.

What remains for precondition 3 once those land: flip both readers
**atomically** (`moves.ts` + `workflow-task-create-ops.ts`, since the
latter computes the preflight the former consumes), delete the flag-OFF
branch with its guards, and drop the settings key.

---

## CORRECTION: seam 2 is not a blocker. My earlier claim was wrong.

I stated in #2639 and above that "with the flag off there is **no**
target-column validation on the move path", so flipping would introduce
new refusals. **That is not what happens.** Reproduced against live PG:
the identical custom-lineage move rejects with the flag **OFF** as well
—

```
Error: Invalid transition: 'backlog' -> 'todo'. Valid targets: building
```

Transition validation is already in force on the flag-OFF path. So for
the shape in question — an engine move to a column the task's own
workflow does not declare — **the move already fails today**, and seam 2
introduces no new break for it. The 20 census sites lacking
`recoveryRehome` are broken on a custom lineage *now*, not broken by the
flip.

I found this because the discriminator I added to prove "the flag is the
cause" failed. Had I written the test to my assumption it would have
passed and the false claim would have shipped — the same way the
equivalence test passed while proving nothing until I tried to make it
fail.

**Revised precondition status:**

| precondition | status |
|---|---|
| 1 — side-effect equivalence | **discharged** — identical rows,
mutation-verified both directions |
| 2 — seam-2 exposure census | **discharged, and it is not a blocker** —
the rejection predates the flag |
| 3 — flip both readers atomically, delete the flag-OFF branch, drop the
settings key | **the remaining work** |

So the flip is no longer gated on fixing 20 engine call sites. What it
is still gated on is precondition 3 being done atomically across
`moves.ts` and `workflow-task-create-ops.ts` (the latter computes the
preflight the former consumes), which is `moves.ts` caller territory.

Three cases now cover seam 2: the flag-ON rejection, the flag-OFF
rejection (asserting the error *message*, so a change in which guard
rejects stays visible rather than reading as agreement), and the #1411
`recoveryRehome` carve-out succeeding — pinning why that carve-out is
load-bearing and must not be tidied away.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:13:58 -07:00
gsxdsm
f5cc416ae4 U7 item 3: the replan no-match fallback named a column no lineage declares (#2659)
**Item 3 from the closing bar.** Behaviour change, own commit.

## The defect

`resolveReplanTargetColumn` fell back to the literal `"triage"` when a
workflow declared neither legacy planner id. That names a column the
workflow doesn't declare — and since #2515 the **default lineage doesn't
declare it either**, so the fallback pointed at a column that exists
nowhere. The replan move then either failed outright or put the card
somewhere no sweep owns.

Resolved through `resolveReboundTarget` (KTD-10: hold → intake → first
declared) — the same helper every other rebound path uses, so replan
lanes and rebound lanes stay consistent instead of drifting.

## The catch path keeps its literal, deliberately

It's reached only when resolution **throws** — not when it silently
falls back to the default IR, which returns a real workflow and takes
the `todo` branch above. With no IR there's nothing to resolve, and
swapping one arbitrary literal for another changes behaviour without
evidence about the workflow. Documented at the site so the asymmetry
reads as a decision, not an oversight.

## Test

Written first and observed **red**. It asserts the target is a column
the workflow actually declares:

```ts
expect(workflowHasColumn(ir, target)).toBe(true);
```

rather than pinning a specific id — so it can't pass by naming a
*different* wrong column, which is how the previous version of this test
stayed green while the fallback was broken.

**Mutation-verified:** restoring the literal fails it.

## Verification

40 replan-target tests green, engine tsc clean, lint clean, merge gate
green (487 + 132 + 10).

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 00:13:39 -07:00
gsxdsm
be63e72f10 U11 [E2E evidence]: live-PG proof for the stranded-column rescue and the planner-lane asymmetry (8 tests, test-only) (#2629)
**Completion bar #3 for my phases.** Test-only, no production changes,
no guard-count movement — the two live-PG E2E suites I held during the
freeze.

## Why these exist

Every U11 slice I shipped closed with the same caveat: *all evidence is
unit-level*. Three claims in particular were argued from reading code,
and each is the kind a mock would happily confirm:

1. #2515 left `triage` a legal id but removed it from the default
lineage.
2. #2603 — `createTask` resolves the workflow's intake column, and an
explicit `column` **overrides** it. Nine write sites were removed on
that reasoning.
3. #2591 — a card stranded on a legacy planner id is admitted by
planning discovery, which is what lets it heal with no data migration.

Both suites drive a **real PostgreSQL TaskStore** (per-file throwaway
database) and the **real shipped workflows**, not fixture IRs. Claim 3
goes through the real `discoverReadyPlanningTasks` — the method the poll
calls. Every assertion is on **observed persisted state** (fresh
`getTask` after clearing the task cache), the rule inherited from
`workflow-lifecycle-live-e2e.pg.test.ts`, because "a function was
called" is exactly what has passed falsely on this program before.

## Two things the E2E found that unit tests did not

**The shared fixture's "merged" shape was not #2515's.** Omitting
`separateIntake` leaves the hold column with *no* intake trait, so the
resolver reports `undefined` — "I have no intake to name" — whereas the
shipped merged lineage carries intake **and** hold on one column and
reports `[]` — "intake exists and *is* the hold column". Callers treat
those differently: `undefined` keeps their legacy default, `[]`
positively asserts no dedicated planner lane. Assuming the plain shape
was the merged shape is how a test appears to cover #2515 while covering
something else. Added an opt-in `mergedIntake` to model the real thing;
the third shape is now asserted explicitly.

**`insertWorkflowDefinitionSync` throws in backend mode** — it's the
SQLite path. The suites use `createWorkflowDefinition` +
`writeTaskWorkflowSelection` like the other live E2Es, including binding
to the id the *store* allocated rather than the one passed in, which the
lifecycle suite documents as a way a renamed-workflow fixture silently
resolves to the default IR.

## Fixture changes are opt-in

Both new options follow the existing `mergeOrchestration` precedent:
seven suites build on this builder and a shared fixture must not
silently change an existing suite's subject.

## Naming

`workflow-planner-lane-**resolution**-live-e2e` deliberately, to stay
distinguishable from #2611's `workflow-planning-lane-live-e2e`.
Different subjects — that one drives the real hold-release sweep, this
one drives the resolvers the lane guards consume. Near-identical names
would invite someone to delete one as a duplicate.

## Verification

- 8 new tests green against a real PG store
- **Mutation-verified:** disabling the #2591 rescue in
`discoverReadyPlanningTasks` fails claim 3, and only claim 3
- Merge gate green (482 + 132 + 10), engine tsc clean, lint clean

**Pre-existing failures, not from this PR:** the full live-E2E sweep is
82 tests / 2 failed, both in `workflow-lifecycle-live-e2e.pg.test.ts`.
Verified by swapping main's `_workflow-vocabulary-fixture.ts` in and
re-running: 2 failed either way, identical. They are main's, and they
appeared since my earlier clean run of that suite — worth a look against
bar #2.

## What this does not cover

Neither suite runs a planning **session** — that lane is the AI,
substituted here as `testMode` does in production. So this proves a card
is *admitted* and re-homable, not that a full plan-and-release round
trip happens. The release half is covered by the existing lifecycle E2E.

No changeset: `@fusion/engine` is private and this is test-only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 00:07:20 -07:00
gsxdsm
5481c27729 docs(solutions): finding 6 — read the implementation before claiming its output is wrong (#2649)
Completes `proving-a-code-path-actually-runs.md` (merged as #2642) with
the rule its own author broke three times while writing it. **Docs
only.**

## Why this belongs in that document rather than a new one

Findings 1-5 are about proving **your own** claim: does this path run,
can this test fail, is this negative result observable. Finding 6 is the
mirror image — the claims we make against **other people's** work — and
it is the same underlying error pointed outward. Splitting them would
let a reader take the first five as "be rigorous about my code" and miss
that the identical discipline applies when reviewing someone else's.

## The three cases, all mine, all in one day

| What I claimed | What was actually true |
|---|---|
| The census undercounts triage guards, 13 vs 10 | `summarize()` counts
`byColumnId` only for `kind === "column"`. My patched counter summed
`role`, `status` and `deliberate` too. The three "missing" ones were
exactly the ones it classifies correctly — and I reported this against
the instrument the program had just adopted as authoritative. |
| `resolvePlannerLanesForTask` silently disables two recovery paths for
legacy cards — escalated across four messages | The file's own header
had already reasoned it through and documented why that answer is
correct. And `TaskStore` implements `getTaskWorkflowSelectionAsync`,
which the resolver prefers — so real projects never take the path my `{
getTask }`-only probe forced. |
| `executor.ts` is clean of triage guards | A receiver-specific grep
missed three under `from` and `originColumn`. Same error one step
earlier: trusting a reconstruction of the thing instead of the thing. |

Every one was: reconstruct behaviour from outside → compare to actual
output → find a difference → report a defect, **without reading the
implementation.**

## The rules it adds

- Read the implementation and its header comment before reporting
anything as wrong. On this codebase the reasoning is usually already
written down, and the FNXC note frequently answers the exact objection —
twice today it answered mine verbatim.
- **A fixture is not a measurement of production.** When a probe and the
real system disagree, suspect the probe: ask what it had to stub, and
whether production ever supplies that shape.
- Retract precisely and immediately. A false defect report against
shared infrastructure costs more than the bug would have — it sends
people to verify something already correct, and spends the credibility
needed for the next report that is real.

Also updates the count in the intro (five → six) and adds an
`applies_when` entry so the doc surfaces for "about to report a tool as
defective", which is when it is needed and not when someone is already
debugging.

`pnpm lint` clean. No changeset — internal documentation.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:07:06 -07:00
gsxdsm
177c2309d9 consolidate/u11: a hold column is a planner lane only if it precedes wip (real defect in merged code) + funnel aliases (triage 11 -> 10) (#2645)
**Consolidation branch for u11/u7.** Supersedes #2624. Contents changed
substantially while it sat unmerged — this body reflects what is
actually in it now.

## Measured with the authoritative census, not grep

`node scripts/lifecycle-column-census.mjs` — **triage guards 11 → 10.**

## 1. A real defect in merged code: a hold column is only a planner lane
if it precedes implementation

Found by greptile on #2616, verified by me, fixed here at the source
because that PR cannot land.

`resolveLifecycleColumns` returns `hold` as the **first** hold-trait
column in declared order, with no positional constraint relative to wip
(`workflow-lifecycle-traits.ts`: `hold:
first(LIFECYCLE_ROLE_FLAGS.hold)`). A workflow using a hold trait for a
**mid-pipeline wait** — a pause after implementation starts — therefore
had that column returned as its planner lane, and
`reconcileMissionFeatureState` demoted the feature to `triaged`. The
mission board reported started work as not-yet-started: silent, and
wrong in the direction that makes a roadmap lie.

This is my defect, introduced in #2610.

**Why it survived:** every lineage anyone has tested puts the hold *in
front* of wip, so the default and Ideas boards are unaffected and no
existing test could see it.

**The fix is positional, with a deliberate asymmetry.** A hold column
counts only when it appears before wip in declared order. When wip
cannot be located the hold is left **out** rather than guessed —
including it wrongly demotes live work on the roadmap, while excluding
it wrongly costs only a `triaged` transition the next reconcile
re-applies.

Mutation-verified: dropping the positional test fails the mid-pipeline
case and nothing else.

## 2. MissionControlPanel funnel aliases

Assessed and **deliberately not trait-converted**. These are heuristic
*name aliases* for a canonical SDLC stage — the matcher already accepts
`signal`/`backlog`/`ready`/`shipped` because it buckets arbitrary
boards, with an `other` fallback. Post-#2515 a default board's planning
cards sit in `todo` and count at the Todo stage, leaving Planning at
zero: the funnel reporting where cards *are*, not a guard that stopped
firing. Hoisted to a named set so it stops reading as unconverted.

This is the **DISPLAY-ALIAS** class the census still lacks — receiver
*is* a column id, purpose is presentation rather than a lifecycle
decision. `DocumentsView`'s status dot is the other one. Without that
bucket a ratchet will keep demanding conversions that make the product
worse.

## What I dropped, because main's version was better

The original #2624 carried a `TaskContextMenu` conversion. #2626 landed
`isPureIntakeColumn` — intake **without** hold — while mine treated any
intake-flagged column as intake. That's wrong for a **merged Planning
column**: it carries both traits, cards there wait for capacity and have
real actions, so I would have suppressed the menu where it belongs — a
new regression in place of the one I was fixing. Theirs is correct. Mine
is gone, along with its now-invalid test and a helper nothing else used.

## Verification

- merge gate green (482 + 132 + 10), engine tsc clean, dashboard tsc
clean, lint clean
- 7 planner-lane tests green, mutation-verified

No changeset: `@fusion/engine` and `@fusion/dashboard` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-30 00:00:46 -07:00
gsxdsm
3e8f604848 test(engine): census the UNCONVERTED lifecycle surface — 417 legacy column literals, ratcheted (#2557)
Test-only, no production change. Independent of my other open PRs.

## The number nobody was counting

This program has two censuses, and **both count converted things**: the
unproven-sites ledger (callers of the lifecycle-role resolvers) and
`raw-workflow-columns-flag-census` (reads of the `workflowColumns`
flag).

Neither counts what is still keyed to a legacy column id — **which is
where every defect this program has found actually lived**:

| defect | the literal |
|---|---|
| pool-id sentinel (capacity gate never bound) | `?? "builtin:coding"`
vs the counter's sentinel |
| agent-link leak (slot consumed forever) | terminal column matched
against a fixed id set |
| stale-paused badge silent on renamed boards | `task.column !== "todo"`
|
| merge chokepoint threw on a finished card | the `done`/`archived` pair
|
| recovered card stranded harder | `?? "todo"` |

Every one was found **by hand, one at a time, by whoever happened to
look.**

## Measured

**438 lifecycle decisions keyed to a legacy column name** (417
comparisons + 31 `??` column fallbacks, minus 8 agent-id false positives
and 2 lines carrying both shapes), across 85+ production files — 94 in
`self-healing.ts`, 70 in `executor.ts`, 26 in the dashboard
task-workflow routes.

That is the real size of the remaining surface. It dwarfs the 15-site
resolver census I've spent this unit closing, which is worth knowing
before anyone calls the vocabulary work finished.

## A hit is not a bug

Many are correct — documented legacy fallbacks, the legacy-adoption
path, code genuinely about the built-in workflow. The census claims only
that each site decides by **name** rather than by **role**, and
therefore needs a human judgment. Reporting 417 as a bug count would be
exactly the overclaiming this program keeps correcting.

## A ceiling, not an equality — deliberate

The sibling flag census fails in both directions. That number moves only
when two units touch it. **This** one moves whenever any of a dozen
concurrent conversion slices lands, and an exact-equality assertion
would go red on work heading the *right* way.

A test that's red for good reasons gets suppressed, and a suppressed
ratchet is worse than none — the failure mode AGENTS.md's quarantine
rule exists to prevent. So the count may fall freely and may never rise;
when it falls, the failure message says to lower the pin.

## Verified in both directions

- green at 417
- adding **one** literal to `replan-target.ts` → `census ROSE to 418
(ceiling 417)`
- the regex is unit-tested to count a **decision**, not a mention: a
column id in a fixture, a log line, or a `moveTask` argument is not
counted — inflating the number into noise is how a census stops being
acted on
- unreadable sources **fail closed** rather than silently shrinking the
count


## Follow-up (a8c150b12): the census was blind to three of the five
defects it cites

I ran the census against its own header. It lists five motivating
defects; the comparison-only regex counted **two**. The pool-id
sentinel, the rebound strand and the terminal fallback are all `??`
**defaults** — invisible to a `.column === "x"` pattern.

A census that cannot see three of the five bugs it names as its reason
to exist is worse than none: it reports a number that *feels* like
coverage. That is precisely the overclaim this unit keeps catching in
other people's work — caught here in mine, and only because the header
wrote the examples down somewhere they could be tested against.

It now counts two shapes — deciding **by** a name (`===`/`!==`) and
**defaulting** to one (`??`) — and pins the five motivating examples as
a test case, so the pattern cannot narrow back without failing.

**Measured: 417 comparisons + 31 fallbacks, of which 2 lines carry both
shapes → 446 lines.** Ceiling raised 417 → 446 to cover the missing
shape, not to excuse new debt.

`?? "builtin:coding"` stays deliberately uncounted: it defaults a
*workflow* id rather than a column and is legitimately correct at most
sites. It already has a stronger guard —
`scripts/check-capacity-pool-id.mjs` bans it only where the value
reaches a capacity counter, which is the only place it's wrong.

Verified both directions: green at 446; adding one fallback of the
newly-counted shape → `census ROSE to 447 (ceiling 446)`.


## Follow-up 2 (98f4264fd): 8 false positives removed — 446 → 438

Then I checked the census against real source instead of trusting the
pattern. Its top-scoring fallback file was `triage.ts` with 8 hits — and
**every one is `agentId: task.assignedAgentId ?? "triage"`**, an *agent*
id, not a column. `"triage"` is both a column id and the synthetic agent
id triage stamps on its audit rows.

Eight of ~34 fallbacks is a quarter of that shape: enough to make the
number **wrong** rather than merely imprecise. A census with known false
positives is one people learn to discount — the same end state as not
having one, which is exactly what its own header warns about.

Excluded, and the exclusion is **pinned as a test case** so it can't
creep back: the three agent-id spellings must match the raw shape *and*
be filtered, while a genuine column fallback that also mentions triage
(`first("intake") ?? "triage"`) must still count.

**Residual imprecision is stated rather than tuned away.** A couple of
counted lines are display defaults (a column rendered in CLI output).
They stay: the census claims each site *needs a human judgment*, and a
display default passes that judgment in seconds. Chasing them costs more
than the precision buys and makes the pattern too clever to trust.
Agent-ids were excluded because they're a quarter of the shape — not
because any false positive is intolerable.

Ceiling 446 → **438**. Verified both directions: green at 438; one new
fallback → `census ROSE to 439`.

## Verification

- census 3/3; engine `tsc --noEmit` clean; `pnpm test:gate` green (414 +
10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:54:49 -07:00
gsxdsm
2a4013b723 consolidate/u9 — review+merge lane: E2E evidence, re-greens, and the conversion blocker (#2646)
**U9's consolidation branch.** Supersedes nothing — #2637 and #2643 are
green with zero threads and left for your sweep per rule 3.

## Contents

| File | Change | Before → After |
|---|---|---|
| `__tests__/executor-step-numbering-zero-based.test.ts` | isolate the
review-handoff `moveTask` call so the assertion is attributable | **1
failed / 3 passed → 4 passed** |
| `__tests__/ce-workflow-step-executor.test.ts` | re-green against the
block-first merge boundary | **3 failed / 48 passed → 51 passed** |
| `__tests__/goal-anchoring-audit.test.ts` | swallow path reports at
debug, not `console.warn` | **1 failed / 6 passed → 7 passed** |

Triage-guard counts: **no change**. My lane has no remaining column
receivers — the rest belong to the capacity/U7/U8/U11/U12 workers, or
are deliberate compat retentions I verified individually
(`spec-staleness.ts` carries its own "U11 proof" block;
`live-agent-count.ts`'s literal fallback is reachable by flag-less
callers).

Commits kept small and separated: signature fix, then attribution fix,
then the boundary re-green, then the debug-channel fix.

**Census reconciliation:** `node scripts/lifecycle-column-census.mjs`
reports **11** triage guards on main, and **none are in the review/merge
lane** — they are the `moves.ts` flag-OFF branch plus the dashboard
cluster. Nothing in this branch moves that number, and I am not chasing
the 779 non-triage guards per your instruction.

## 1. The review-handoff assertion (and a lesson)

The handoff gained a third argument (workflow move provenance), so a
two-arg `toHaveBeenCalledWith` failed on the extra options object while
the card moved correctly.

My first fix used `expect.anything()` — and I *documented in the
comment* that six mutations couldn't make it fail, then shipped it
anyway. Greptile (P2) correctly called that out: this flow records two
`moveTask` calls, so the assertion is satisfied by the boundary move
even if the handoff regresses. **Documenting a weakness is not removing
it.**

Now the test selects the handoff call by its own marker
(`workflowMoveMetadata.reason === "workflow-review-handoff"`), asserts
exactly one such call, and asserts its target column:

| Mutation | Before | After |
|---|---|---|
| change the seam's `reason` | green | **NEW=1**, this test only |
| retarget the seam to `"done"` | green | **NEW=1**, this test only |

## 2. The merge boundary changed shape

`ensureWorkflowMergeBoundaryTask` (`executor.ts:7808`) now **refuses** a
foreach step-execute region with incomplete pre-merge node proof —
logging `"Workflow merge boundary blocked: <reason>"` and returning
**without moving**. The move-then-check sequence this file pinned is
gone: `"Workflow merge boundary moved task to in-review before
requesting merge"` no longer exists anywhere in production.

Three fixes, one per failure:

1. **negative case** pinned the retired move-first log. Now pins the
*stronger* property the new order gives: an unproven card is **not moved
into review at all**. The old assertion could only say "it was moved,
then blocked". Log text asserted by stable prefix — the reason clause
enumerates missing instance ids, which is legitimately volatile.
2. **"moves direct-to-merge tasks into in-review"** got zero calls: its
fixture recorded no node results, so the gate blocked it. Added one
`steps#0:step-execute` pre-merge result.
3. **"completes graph-native checklist projection"** also got zero
calls. Its existing `plan` result proves *some* pre-merge node ran but
not the per-instance work; the gate additionally requires an instance
per foreach step-execute. Added the two matching its two steps.

(2) and (3) are the same class as the lifecycle E2E `seedTask` fix in
#2634: a fixture that never modelled completed work, asking the engine
to advance it, and reading the correct refusal as a failure. Proof shape
matched to the evaluator (`source: "node"`, `phase: "pre-merge"`,
terminal = `passed`/`skipped`) rather than guessed.

Verified the gate is what these fixtures exercise: disabling the
boundary proof check fails the negative case (`NEW=1`, that test only).

`pnpm test:gate` green, `pnpm lint` clean.

## Where U9 actually stands

The conversion (S06/S07/S08) is **not** done, and is now precisely
characterised rather than "blocked on U8":

`workflow-graph-executor.ts:310` short-circuits every
`MERGE_REGION_KINDS` entry to the legacy merge seam, so `merge-gate`,
`merge-attempt`, `manual-merge-hold`, `retry-backoff`, `recovery-router`
and both `branch-group-*` handlers **never execute**.
`createMergeGateHandler` does read `task.autoMerge` and emit
auto-on/auto-off — and is never called. The builtin IR's
`outcome:auto-*` edges are unreachable.

**U9's conversion, concretely: stop short-circuiting
`MERGE_REGION_KINDS` and let those nodes run.** S06/S07/S08 all hang off
that one change. Safeguard 2 has no node-level representation today, so
enabling the region without carrying the `autoMerge` contract into it
would let an `autoMerge:false` card merge on PR-readiness alone. Full
write-up in
`docs/plans/workflow-owned-merge-stack/u9-safeguard-baseline.md`
(#2634).

## 3. A recurring class worth a shared helper

`goal-anchoring-audit`'s swallow path now reports via `log.debug` (a
deliberate demotion of log noise), and `debug` is FUSION_DEBUG-gated so
vitest emits nothing — the test asserted a channel that was both wrong
*and* disabled. I kept both halves of the contract (swallowed **and**
reported) by enabling the flag for that case, rather than deleting the
awkward assertion.

**This is the third instance this session** — `worktree-pool`,
`self-healing`'s auto-archive line, and now this. If a fourth appears it
deserves a shared test helper rather than three bespoke fixes.

## Two failing files I could NOT responsibly take — flagged, not touched

**`executor-prompt.test.ts` (3 failures) — I ESCALATED THIS AND I WAS
WRONG. Retracting.**

I flagged these as a possible real pause-contract violation: an agent
session spawning while an operator has globally paused the engine. I
then finished the diagnosis, and the evidence goes the other way.
Recording the retraction with the same detail as the alarm, because a
false alarm aimed at another unit costs them a chase.

**The discriminator I asked for, resolved.** Six tests in that file
assert `expect(mockedCreateFnAgent).not.toHaveBeenCalled()` during
global pause; 3 fail. Splitting them by what they drive:

| Assertions | Drives | Result |
|---|---|---|
| `does not resume unpaused in-progress task while global pause is
active` (+2 siblings) | no executor method — `task:updated` / resume
paths | **pass** |
| `parks todo tasks in in-progress when fn_task_done…` (+2 siblings) |
`executor.execute(...)` **directly** | **fail** |

So the guard holds on every event-driven path and is absent only from
the direct `execute()` entry.

**And `execute()` is not the guard site — the scheduler is.**
`scheduler.ts:1491` is an explicit hard stop (*"Global pause (hard
stop): halt all scheduling activity"*), with a second gate at `:1055`,
and the scheduler never calls `.execute(` at all — dispatch routes
through the runtime. In production a global pause halts scheduling
before anything reaches the executor.

**Conclusion: the pause contract is intact in production.** The 3
failing tests call `execute()` directly, bypassing the upstream gate,
and assert a defence-in-depth check *inside* `execute()` that is not
there. They are testing a path production does not take during a pause.

What that leaves is a real but much smaller question, and a design one
rather than a defect: should `execute()` carry its own pause check as
defence-in-depth, given non-scheduler callers exist (self-healing,
manual retry)? If yes, add the guard and all six assertions pass. If no,
the 3 direct-`execute` assertions are asserting a guarantee the
architecture places elsewhere and should be retired. **I have not
changed either the code or the tests** — but nobody needs to hunt a
pause-contract regression, because there isn't one.

**`executor-fast-mode-workflows.test.ts` (1 failure) — mechanism not
isolated.** `visitedNodeIds` is `['review']` where the test expects
`['start','review']`. Three probes failed to explain it: giving the
review node an explicit `column: "in-progress"` changed nothing (so it
is not column-based entry resolution), and swapping `seam: "review"` for
a plain prompt config did not isolate it either. Two structurally
identical sibling tests in the same file still pass with `['start',
...]`, so something in graph traversal distinguishes them that I did not
find. That is U8/graph-executor territory; I am not asserting a
`visitedNodeIds` shape I cannot explain.

## Still open and green

- **#2637** — `task-delete-notice` 21 failed → 34 passed.
- **#2643** — shellout allowlist re-pin. **Merge early:** it re-drifts
whenever `executor.ts`/`self-healing.ts` line counts shift, with no git
conflict to warn you. It already drifted once while open
(`executor.ts:17106 → 17198`) and I re-pinned it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:50:59 -07:00
gsxdsm
6ed284f36a drop the dead semaphore parameter from dropPreHeldExecutorSlot (#2574)
Small follow-on to the cross-project cap removal.

`dropPreHeldExecutorSlot(taskId, semaphore?)` released a cross-project
semaphore slot. That semaphore is deleted, and **all 16 production call
sites passed `this.options.semaphore`**, which nothing wires any more —
so the release was a no-op on an always-undefined value: an optional
parameter that reads as if it does something.

## What is *not* deleted

Pre-held slots are **dual-purpose**: a cross-project semaphore slot
**and** the FN-8453 per-project coordinator reservation. Only the first
is gone. The reservation is the half that matters — every rejection path
funnels through this helper so an early scheduler/triage return cannot
permanently consume a project slot — and it stays. That is why this is a
parameter change, not a helper deletion.

Sites that still hold a semaphore reference release it **explicitly**
next to their drop, so behaviour is unchanged for any caller that
supplies one. Nothing wires one in production today, but silently
leaking a slot for a caller that does is not a trade a cleanup is
allowed to make.

## One real leak fixed — found by a failing test, not by reading

`ProjectAdmissionCoordinator.admitOldest`’s release lambda took the
pre-held branch and **returned**, relying on the deleted parameter to
hand the host slot back. With the parameter gone, that branch unwound
the registration and the reservation while **leaking the host slot** the
attempt had acquired. The release is now unconditional across both
branches.

Worth noting how it surfaced: the test that caught it (`drops a declined
candidate’s pre-held executor slot`) asserted `semaphore.activeCount`,
which I had initially assumed was just coupling to the deleted half. It
was not — it was pinning a real invariant.

## Tests

Five cases in `concurrency.test.ts` pinned `sem.activeCount` through a
drop. Each is re-pointed at the surviving contract — registration and
reservation unwound, nothing left for a later pass to “take” — with the
semaphore assertions moved to the sites that now own the release.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green ·
`concurrency.test.ts` **56/56**. The 8 `triage.test.ts` failures are
**pre-existing** — reproduced identically with this branch’s `triage.ts`
replaced by main’s.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:52 -07:00
gsxdsm
3aa942ee5f capacity: spawned agents count against the project agent count (#2579)
Two configurable numbers per project. `maxSpawnedAgentsPerParent` (5)
and `maxSpawnedAgentsGlobal` (20) were a **third and fourth** limiter
with private budgets invisible to both.

## This closes a hole, not just knobs

A spawned child **is** an agent and gets **its own git worktree**
(branched from the parent’s — the tool’s own description says so), but
children were counted by **neither** capacity gate. A fan-out could put
up to 20 extra worktrees on disk while the scheduler believed the
project was at its configured limit. The operator’s two numbers were
simply wrong about what was running.

## The old caps also measured the wrong thing

`totalSpawnedCount` decrements on child cleanup, but the per-parent
**set** is cleared only when the **parent task** ends. So
`maxSpawnedAgentsPerParent` throttled *cumulative* spawns across a
task’s life rather than *concurrent* ones — a long-running task could
exhaust its budget with five children that had all long since finished,
and the operator had no way to see why.

## Fix

`fn_spawn_agent` gates on the same project agent count every other lane
uses (`computeTopLevelConcurrencyClaimedFromStore`) plus live children.
One number, one answer, no private budget that can disagree with the
board.

The refusal names **Max Concurrent Tasks** — a control the operator
actually has. The old messages pointed at settings that no longer exist,
which is worse than no message: it sends someone hunting for a knob that
is not there.

## Verification

**Revert-proof, measured:** restoring the private budgets turns **3 of
the 4** new cases red — a project at 1/1 could still spawn, which is
precisely the hole. `executor.ts` restored byte-identical.

`pnpm lint` clean · core + engine `tsc` clean · `pnpm test:gate` green
(414 + 10 + 71) · new suite 4/4 · `settings-default-descriptions` 4/4.

There was no spawn-capacity test before this; the file is new.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Spawned agents now count toward the project’s **Max Concurrent Tasks**
capacity.
* Agent spawning is blocked when capacity is reached, including
concurrent spawn attempts.
* **Bug Fixes**
  * Prevented over-allocation during simultaneous agent spawns.
  * Restored available capacity when agent creation fails.
* **Changes**
  * Removed separate per-parent and global spawned-agent limits.
  * Updated settings to reflect the revised capacity controls.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:41 -07:00
gsxdsm
152fedbd32 record the detector audit: gridlock and stuck-task are keeps, with evidence (#2581)
Answering the review question *“does gridlock detection still have a
job?”* — with evidence rather than assumption, and recording it so the
question is not re-opened by someone reading the name.

**No behaviour change.** Comments only.

## Gridlock detector — KEEP

`GridlockEvent.reasons` is typed `"dependency" | "overlap"`. It detects
**dependency deadlock** and **file-scope overlap deadlock** via the
scheduler’s `pathsOverlap` / `filterPathsByIgnoreList`. That has nothing
to do with limiters arbitrating against each other — two tasks can still
block on a dependency cycle or a shared file scope no matter how many
agents the operator allows.

The hypothesis that gridlock ≈ competing limiters deadlocking was
reasonable from the name, and wrong.

## Stuck-task detector — KEEP

Detects a stuck **agent** — a live session repeating the same tool call,
or emitting no activity signal — via tool fingerprints and inactivity
windows. Orthogonal to how many agents may run: a single agent on an
unlimited board can still wedge.

## Evidence

Measured for both: **zero** references to `maxConcurrent` /
`maxWorktrees` / `semaphore` / `capacity` / `slot`. Both are live and
wired — gridlock via `project-engine.ts → notifier.notifyGridlock`,
stuck-task via `in-process-runtime.ts`.

The note lives in each file because the natural reading of “gridlock” is
“limiters deadlocking”, and deleting a live detector on that reading
would remove real coverage silently. Each note states the question a
future cleanup should actually ask — *is dependency/overlap deadlock
still possible?* — rather than *is capacity simpler now?*

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green ·
detector suites **108/108**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:30 -07:00