Commit Graph

4136 Commits

Author SHA1 Message Date
gsxdsm
dca20496f4 consolidate/u7: plugins to zero + 8 executor rebound guards + resume lanes (supersedes #2607, #2635, #2640) (#2644)
Consolidation branch for U7, per the new one-branch working mode.
**Supersedes #2607, #2635, #2640** — the three of my PRs that were stuck
on review threads. My other seven (#2602, #2605, #2606, #2611, #2621,
#2628, #2633) are green with **zero unresolved threads** and are
deliberately left alone for the merge sweep.

## What is in here, file by file

| file | change | guards before → after |
|---|---|---|
| `plugins/…/glasses/src/agent-actions.ts` | gates, destinations and
degraded-resolution refusal all resolve from the task's own workflow | 2
→ 0 |
| `plugins/…/glasses/src/quick-capture.ts` | accepted capture columns
come from the board; default no longer names the deleted column | 1 → 0
|
| `plugins/…/glasses/src/settings.ts` | quick-capture default was
`triage`, the column #2515 removed | (assignment, uncounted) |
| `plugins/…/dependency-graph/src/GraphTaskNode.tsx` | redundant column
condition deleted | 1 → 0 |
| `packages/engine/src/executor.ts` | 8 rebound guards compare the
resolved column; 4 resume-eligibility literals share one resolver | 151
→ 143 (+4 off-bar) |
| `packages/engine/src/__tests__/` | 4 new suites, 26 cases | — |

`plugins/` reaches **zero** column guards with this branch.

## The three threads it closes

**#2607 — five findings, all mine, all the same rule.** I kept
*qualifying* a legacy-id fallback instead of removing it:

| attempt | rule | hole review found |
|---|---|---|
| 1 | fall back to `todo` when the role is missing | moved cards to
phantom columns |
| 2 | …only if the workflow **declares** `todo` | aliased **review**
lane named `todo` |
| 3 | …and only if no other role is assigned to it | **traitless**
parking column named `todo` |

The qualifications were the mistake. Once `resolveLanes` returns a lane
set the workflow *has* a column vocabulary, so "no column carries the
hold trait" is a complete answer — refuse. `destination()` is two lines
now, with no aliasing surface left to qualify.

Plus a sixth, which is a genuinely different state: **degraded
resolution is indistinguishable from the default board.**
`resolveWorkflowIrForTask` is total by design — a missing definition
silently returns the *default* coding IR — so a card on a custom board
whose definition could not be read resolved to `todo`/`in-progress`.
`undefined` lanes cannot express that (it means "no workflow at all",
where the legacy ids *are* the answer). The actions now refuse with 409.
#2618 would replace this check with resolver provenance; it is not
merged, so this does not depend on it.

**#2635 — "seven rebound sites remain untested."** Fair; my "same shape"
note was an assertion, not coverage. Seven of the eight need a live
graph run to reach, so the *shape* is pinned instead: a static check
that no guard in front of a rebound move compares against a column
literal, with a vacuity case (the same detection run against the
original shape) and a match-count floor (≥8), because a guard reporting
success on zero matches is worse than no guard.

**#2640 — duplicate workflow resolution.** Framed as I/O; it is also a
correctness bug. Eligibility and re-entry are two halves of one decision
and resolved the workflow separately, so a workflow edit landing between
them has the halves reading *different boards*. Now one caller-owned
memo per decision — caller-owned because a process-lifetime cache would
have to guess when a mid-flight workflow edit invalidates it.

## Behavioural findings, not tidying

- **The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** `promotedFromPlannerColumn` was false
on a renamed board, so finished work resting in planning was never
promoted; the code fell through to a review handoff that role adjacency
rejects, and the card stayed stuck with its work complete.
- **Rebound guards could not see the column their own move targeted.**
U5b converted the move target; the eight `column !== "todo"` checks in
front of it were left literal, so on a renamed board the engine moved a
card into the column it was already in — and `moveTaskInternal` runs
reset-on-entry on every real move, so at the `preserveProgress: false`
site it reset step progress a second time.
- **The FN-1404 `task:move` audit row was lying**, recording `to:
"todo"` while the move target was resolved. A run-audit trail that
disagrees with the move it describes is worse than none. Not a
comparison, so no census counts it.
- **A task interrupted by an engine pause never resumed on a renamed
board** (off-bar, `in-review`/`in-progress` literals): four comparisons
decided one question and had to agree; two of them disagreed on a
renamed board, so re-entry silently never fired.

## Revert proofs, isolated per site

| reverted | result |
|---|---|
| `destination()` back to attempt 3 | 3 of 38 fail |
| degraded-resolution refusals removed | 2 of 42 fail |
| capture set back to the legacy five | 2 of 3 fail (renamed-board
suite) |
| forward exclusions → literals | 1 of 14 fails |
| missing-wip refusal removed | 2 of 14 fail |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| promotion target → `"in-progress"` | 3 of 7 fail |
| one rebound guard → `!== "todo"` | 1 of 3 fails (static shape) |
| resume lanes → legacy trio | 1 of 5 fails |

Every conversion is paired with a negative — a forward move, a
not-a-planner-lane card, a default-lineage card, an unresolvable
workflow — so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Commit discipline

Twelve commits, each one thing: the code move (`resolvePlannerLanes` out
of `triage.ts`) is separate from every behavior change, and each review
fix is its own commit with its own revert proof.

## Verification

- `pnpm test:gate` **71/71**
- 162/162 across the glasses plugin's 19 files; 26/26 across the four
new engine suites
- engine + glasses typecheck clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Engine recovery and retries now work correctly with renamed or
customized workflow columns.
  * Tasks in manual-intake columns are no longer automatically planned.
* Agent actions and quick capture now respect each board’s declared
columns and lifecycle stages.
* Awaiting-approval tasks are recognized regardless of their current
column.
* Command Center SDLC funnel stages now accurately reflect customized
workflows.

* **Documentation**
* Added guidance for safely changing workflow-column logic and
interpreting lifecycle-column checks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:52:55 -07:00
gsxdsm
efbbc45eb0 U12: the LAST triage guard — Plan was offered on executing cards named triage (#2664)
The final `column === "triage"` in production source, and it was a live
defect rather than dead vocabulary.

## The defect

`isPreExecutionHoldColumn` ORed the legacy id with the traits
**unconditionally**:

```ts
return column === "triage" || flags?.intake === true || flags?.hold === true;
```

That is not a fallback. A resolved column merely *named* `triage`
answered true even when its own traits said work was underway — so the
context menu offered **Plan**, which re-plans, on a card that is already
executing.

Now flags-first, with the id as the documented no-metadata answer.

## Why the file's earlier conversion missed it

Every existing case in `TaskContextMenu.test.tsx` passes a column with
**no flags**, or with `hold`/`intake` set. All of them agree under both
forms, so the suite could not distinguish them. Nothing exercised a
column whose **name and traits disagree**, which is the only shape that
separates an OR from a fallback.

Three new cases cover it. Revert check: restoring the OR form fails the
first one — Plan reappears on a mid-flight card.

## The asymmetry is preserved, and now tested

The degraded set stays `{triage}` **alone**, deliberately not the
`{todo, triage}` used by `isPreImplementationColumnRole`. That helper
drives the preserve-progress prompt, where a flagless `todo` *should*
prompt because losing steps is unrecoverable. This drives Plan, where a
flagless `todo` must **not** offer to re-plan a card that may already be
planned. The file documented that difference; nothing asserted it. Now a
test does.

## On reaching zero honestly

The surviving literal is marked `DELIBERATE-LITERAL`. It is the degraded
answer, not an unconverted guard — there is no trait to read when
`flags` is `undefined`, which happens during first paint and for a card
in a column its workflow no longer declares. Deleting it would silently
withdraw Plan from exactly the stranded cards that most need
re-planning.

So **`triage → 0` means "no unconverted guards remain", not "the string
is gone"**, and I would rather say that than move a number by deleting a
fallback.

| branch | triage |
|---|---:|
| `origin/main` | 5 |
| this PR | **4** |
| #2655 (flag resolution, removes 4 in `moves.ts`) | 1 → **0** combined
|

I found it with the census's own AST classifier rather than grep — my
grep of the same tree returned only comment prose and would have had me
report the bar as met while a real defect sat in
`TaskContextMenu.tsx:179`.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 158 / 487 / 71). `pnpm
check:lifecycle-columns` exits 0 with the baseline re-recorded in this
PR (column 769 → 768, deliberate 12 → 13). `tsc -p tsconfig.app.json`
clean. `TaskContextMenu.test.tsx` 18/18.

Depends on nothing; stacks cleanly with #2655 and #2661.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:37:08 -07:00
gsxdsm
8393bba7dc U7: the replan rebound targets a column the workflow declares (R7) — re-landed on main (#2598)
> Based on `main`, no dependencies. Re-landed after closing the stacked
chain (#2517, and #2551 below) that never reached main.

## What main already has, and what it lacks

Main independently converted the planner-lane **parameters** in this
file — and **better than I had**: it splits `plannerColumn` from
`roles.mergedPlanningColumn`, because a merged lane joins the FN-8596
arrival-order rescue but *not* the "planner column is never advanced"
shortcut. That work is main's and untouched here.

What main still lacks is the **R7 fix**: `resolveReplanTargetColumn`
returns `"triage"` **by fiat** for any workflow declaring neither legacy
id — `builtin:marketing` (ideation/backlog/drafting/…) and every fully
renamed set. A Plan Review REVISE therefore moves the card into a column
its workflow **does not declare**, for `reconcileUndeclaredTaskColumns`
to clean up after. A move the engine makes on purpose, not drift.

| Workflow | Target | Changed? |
|---|---|---|
| `builtin:coding` / stepwise | `todo` | no |
| Coding (Ideas) | `todo` | no |
| `builtin:marketing` | `backlog` (its own hold) | **yes** — was
`triage`, undeclared |
| declares no planning lane | `undefined` → park | **yes** — was
`triage` by fiat |

## The ordering the existing suite taught me

Legacy ids stay preferred **first**, and the trait resolution prefers
**hold over intake**. That is not arbitrary:

Coding (Ideas) declares `ideas` as its intake, and `ideas` is **manual
capture with no AI** (plan R10) — a rejected plan sent there stops being
replanned at all. The old code got Ideas right **by accident**: it never
recognised `ideas` as intake and fell through to `todo`. An "intake
first" trait rule would have shipped that regression dressed as a
cleanup, and three existing Ideas tests were the only thing between me
and doing it.

## An inverted comment, corrected

The function's own U11 note read: *"the second lookup asks for `todo`,
which U11 deletes… the first lookup still matches `triage` (which U11
keeps)"*.

**That is backwards.** #2515 keeps `todo` and deletes `triage`, so the
consequence is the opposite of what was written — the `todo` branch is
what saves builtin coding. Fixed rather than left, because a comment
that inverts a merge's direction sends the next reader to the wrong
branch.

## Fail-closed callers

`undefined` means "nowhere to replan" (plan U5: *skipped with a log
rather than moved arbitrarily*). All four call sites park **visibly**
rather than log a move they did not make. The scheduler's rebound still
writes `needs-replan` — deliberately, since that is what blocks dispatch
and the branch has already decided the card must not be released — with
only the *log* made conditional.

## The superseded test is deleted, not skipped

A skipped test is a guard that cannot fire. Its replacement asserts the
new contract **and** the R7 invariant directly — *"a column this
workflow declares"*, not just an id — plus a new case for a workflow
with no planning lane at all.

## Verification

| Check | Result |
|---|---|
| replan-target | 41/41 |
| with scheduler-trait-dispatch + pre-release-plan-review | 55/55 |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (482 + 10 + 71) |
| `pnpm check:changesets` | clean |

`triage.test.ts` still shows main's **8 pre-existing #2515 failures** —
unchanged by this, fixed by **#2576**.

## Closing #2551

Its parameter work is superseded by main's better version; this PR
carries the only part main lacked. Same story as #2517: a stacked PR
outlived the surface it was converting.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:31:03 -07:00
gsxdsm
2771408bba ci: enforce the lifecycle-column ratchet — it has never actually run (#2654)
**The ratchet was advisory.** `scripts/lifecycle-column-census.mjs`
existed only as `pnpm census:lifecycle-columns` — without `--strict` —
and **no workflow invoked it**. Nothing has ever compared the tree to
the baseline. Every "the baseline ratchet holds them" assumption in this
program rested on a check that does not run.

That explains both classes of hole:

**1. Three PRs lowered counts without re-recording,** leaving allowances
the deleted guards could return through while every check stayed green.
I've tightened them across #2593 and earlier PRs, but nothing stops the
next one.

**2. #2621 GREW the count while its own title claimed "count 0 → 0".**
It added `column === "triage"` and `column === "todo"` at
`register-task-workflow-routes.ts:2681`, taking that file to **23
against an allowance of 22**. It landed unchallenged. This is the
failure mode the ratchet exists to prevent, and it happened *inside this
program*, in a PR that asserted the opposite.

## The change

Adds `check:lifecycle-columns` (the census with `--strict`) to the
`pr-checks.yml` lint job, next to `check:changesets` and
`check:routes-modular` — the established pattern. **~1.8s over ~1950
files**, so this is not a slow-test addition.

## Proven to fail, in both directions

A guard that reports success without checking anything is worse than no
guard, so:

| injected defect | result |
|---|---|
| `const __probe = (c: string) => c === "triage"` added to `moves.ts` |
`count ROSE — moves.ts: 39 -> 40`, exit 1 |
| run against main's current baseline | exit 1 on
`mission-feature-sync.ts: allows 5, tree has 0` |

Both reverted; exit 0 restored. Note the second row: **this check is RED
on main right now**, which is the point.

## Merge order

**Stacked on #2593**, which carries the `DELIBERATE-LITERAL` marker for
the #2621 site (a v1 IR declares no roles, so no trait can answer that
question) plus the baseline re-record. Standalone on main this PR is red
— correctly. **Merge #2593 first**, then this.

I stacked rather than duplicating those two edits because I already
caused one conflict today by appending related content from two
branches, and #2651 merged a correction ahead of the section it
corrected. Same-content edits in two PRs is the same mistake.

## Census

Unchanged by this PR: **776 total, triage 5, reviewed 16** — it adds no
guards and converts none. It only makes the numbers enforceable.

## For the fleet

This should land before the 776-guard fleet launches. The brief says
"the baseline ratchet must shrink by exactly the converted count" —
until now nothing verified that claim, so a batch worker could report a
shrink that did not happen, or grow the count while converting, and CI
would agree.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:20:31 -07:00
gsxdsm
642a4fa264 consolidate/u12 — U12 consolidation: 4 live defects, the AST ratchet fail-closed, and the moves.ts flag scoped (#2647)
One branch, one PR, per the consolidation directive. Contents
file-by-file below.

**Supersedes #2625** (its overlapping conversions landed via U11's
#2624/#2626/#2636; only the parts nobody else did are folded here).
**#2630 and #2639 stay open** — both green with zero threads, per rule
3.

## Four live defects, each measured

**1. Every planning card renders an actions menu.**
`TaskContextMenu.tsx` still had `shouldShowActionsMenu: task.column !==
"triage"` on main *after* the rest of that file was converted. Since
#2515 removed the id, the condition is TRUE for every card, so the
suppression stopped applying anywhere — including on cards whose menu is
empty, the orphaned click target the Surface Enumeration rule exists to
catch.

Found **twice independently**: by reading the guard, and again by the
invariance test below, which failed on main with `shouldShowActionsMenu`
true on one lineage and false on another. That is the argument for an
invariance property over per-site conversion — the file had already been
converted "2 → 1" and the survivor was the live one.

**2. Worktree upcoming-work list empty on renamed boards.**
`groupByWorktree` filtered `t.column === "todo"`. On the default board
the id and the role coincide so every existing test passed; renamed, it
matched nothing and a whole panel read as idle.

**3. Hold-lane FIFO ordering lost on renamed boards.**
`sortTasksForDisplayColumn` gated priority-then-FIFO on `column ===
"todo"`, degrading to the generic id-ordered sort elsewhere. Cards
simply appear in the wrong order, silently.

**4. The AST ratchet still failed open** — fourth time in that file,
third found by review. `receiverName` understood only one-level property
access and bare identifiers, so `task["column"]`, `metadataColumn(entry,
"to")`, ternaries, `(task!.column)` and backtick literals were dropped.
**Measured on main: `in-progress` 196 → 197, `in-review` 211 → 213** —
three real guards nobody counted, including `metadataColumn(entry, "to")
=== "in-review"` in `reliability-metrics.ts`. Now walks wrappers,
resolves calls to the callee name, and emits a `<SyntaxKind>`
**sentinel** for anything unnameable: counted *and* trips the
classification guard, so a human judges it instead of it vanishing.

## Per-file guard counts

| file | before | after |
|---|---:|---:|
| `app/components/TaskContextMenu.tsx` | 1 | **0** |
| `app/utils/worktreeGrouping.ts` | 1 | **0** |
| `app/components/taskSorting.ts` | 1 | **0** |

The other dashboard files I had converted reached 0 via U11's PRs; where
our work overlapped I took theirs during the rebase, including two
places where theirs was **stronger** than mine — they deleted Column's
unreachable quick-create arm outright (with fixtures migrated) where I
had converted it, and they verified the same `isPreExecutionHoldColumn`
degraded-set asymmetry I did, independently.

## Flip precondition: the moves.ts flag is scoped, not flipped

`move-target-declared-census.test.ts` answers precondition 2 with
measurement. 41 engine `moveTask` calls have literal targets — `todo`
27, `in-progress` 7, `done` 6, `archived` 1 — and **all four are
declared by the default lineage**, so the default board is not the
exposure. `triage` appears only in a comment noting `replan-target.ts`
used to hardcode it. My own grep had said `todo=29`; the AST says 27,
because grep counts comments.

The exposure is **custom** lineages: 20 of the 41 carry no
`recoveryRehome` and would reject with unknown-column post-flip; 21 are
exempt via the #1411 carve-out, which makes that carve-out load-bearing.

I did not flip the flag. It is six seams, not the `789`/`837` pair every
summary including mine described, and seam 2 turns on *new refusals*
rather than swapping equivalent implementations — a green suite says
nothing about that. #2639 pins the blast radius.

## Tests

- `column-role-id-invariance.test.tsx` — hold traits fixed, vary only
the column id across MERGED / LEGACY / RENAMED; every decision must
agree. Drives the real consumers, so a component keeping an inline
comparison fails it. Includes a unanimous-and-**false** case so it can't
be satisfied by a predicate hardwired to true. **This is the test that
caught defect 1 on main.**
- `worktreeGrouping.test.ts` — includes two cards both in a column named
`staging`, one hold and one not, asserting opposite answers. That
assertion is impossible under a board-wide column-id set, which is why
hold resolution is keyed per task via `getEffectiveTaskWorkflowId`
(#2625 review).
- `taskSorting.test.ts` — discriminates on the **tiebreak**, not
priority: both branches sort by priority, so my first version passed for
the wrong reason. Equal-priority cards whose `createdAt` order disagrees
with their id order.
- `no-hardcoded-lifecycle-columns.test.ts` — 16 detector cases: 11
shapes counted, 4 legitimate ignored, one asserting the sentinel path.

Revert checks, all run: menu suppression → diff names the field;
worktree → `expected [] to include 'FN-50'`; sort → `FN-2, FN-9` instead
of `FN-9, FN-2`; ratchet → the 3 recovered guards disappear.

## One site that should never be converted

`MissionControlPanel.tsx:46` — `{ id: "triage", match: (c) => c ===
"triage" || c === "signal" || c === "backlog" }` is a deliberate
name-similarity heuristic for the SDLC funnel; it matches synonyms and
folds unknown columns into an "other" bucket so custom columns still
contribute. Converting it changes what the funnel displays. Like the
`live-agent-count` fallbacks, it belongs in a documented floor — **the
ratchet's target is that floor, not zero.**

`DocumentsView.tsx:73` is convertible but the file has no column flags
at all, so a real fix means plumbing board-workflow metadata into a view
that doesn't fetch it — its own unit of work.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 482 / 71). `tsc -p
packages/dashboard/tsconfig.app.json` and `packages/core/tsconfig.json`
clean. Core ratchet + seam suites 24/24. Dashboard target suites 37/38 —
the one failure is the pre-existing `"Back to In Progress"` label
casing, confirmed identical on the base.

---

## Added after the initial push

**5. `TaskCard` lost inline editing on renamed boards; `TaskDetailModal`
kept it.** Still live on main: the modal resolved field editability from
traits in U10/R8, the card used a hardcoded `{triage, todo}` set with
**no trait path at all** — even though `taskColumnFlags` was already in
scope. On a renamed board the title was editable in the modal and the
pencil was missing from the card. Body moved unchanged into
`isFieldEditableColumnRole` so the two surfaces cannot drift again.

The veto traits are the substance: a column can legally carry `hold`
**and** a WIP or review trait, and a plain `intake || hold` check would
let an operator rewrite a description while a session executes against
it.

Coverage gap **measured, not assumed**: mutating `canEdit` back to the
hardcoded set left `TaskCard*` at the same failure count as the
unmutated run — nothing caught it. The four render cases assert the real
`aria-label`; that mutation now fails with `Unable to find an accessible
element ... name 'Edit task'`.

**6. The ratchet's target is a documented FLOOR, not zero** — and this
changes the completion bar.

Zero is not reachable, and chasing it means breaking working code. Two
categories are permanent, now protected as positive assertions so a
future sweep cannot "finish the job" by deleting them:

- `MissionControlPanel.tsx`'s `FUNNEL_STAGES` is a deliberate
**name-similarity** heuristic — it matches `signal`, `backlog`, `to-do`,
`ready`, `shipped` and folds unrecognised columns into an "other" bucket
so a custom board still contributes counts. It is not asking whether a
column has the intake trait; it buckets arbitrary column *names* for
display. Asserted on the **synonym list**, because the synonyms are what
prove it is name matching — if they disappear the site has changed
character and the exemption stops applying.
- `live-agent-count.ts`'s no-flags arm is reachable (a remote store is
deliberately given an empty flag map; a card in an undeclared column has
no flags at all) and deleting the literal makes such a card match **no**
arm, so the queued total silently under-reports a stranded card.

A count with an undocumented floor invites someone to drive it to zero.

**Not done, and why:** `DocumentsView.tsx:73` is convertible but that
file has no column flags anywhere, so a real fix means plumbing
board-workflow metadata into a view that does not fetch it — its own
unit of work, not something to smuggle into a conversion.

**Re-verified after these commits:** `pnpm lint` clean, `pnpm test:gate`
green (10 / 482 / 71), `tsc` clean on core and `tsconfig.app.json`, core
ratchet suite 26/26, `columnRoles` 10/10, `TaskCard.test.tsx` 384/386
(the 2 are pre-existing CSS assertions). `TaskDetail*` is 130 failed /
551 passed **both with and without** this change — verified by stashing,
so pre-existing and unrelated.

---

## Flag resolution: preconditions 1 and 2 are now DISCHARGED.
Precondition 3 is blocked, and by evidence.

**Precondition 1 — the side-effect equivalence proof — done.**
`moves-flag-equivalence.test.ts` runs the same journey under both flag
states against live PG and diffs the persisted row. **Result:
identical** — whole-row equality across 128 fields plus an equal timing
shape, over `todo → in-progress → in-review → todo → in-progress`.

That test was **wrong twice** before it meant anything, and both times
it was passing:

1. **It proved nothing.** `experimentalFeatures` is **global-only**, and
`moves.ts` reads `getSettingsFast()`, which filters global-only keys out
of the project layer. My `updateSettings` write was silently discarded,
`useWorkflow` was false in *both* runs, and the "proof" compared the
legacy path against itself. Found by stamping the flag-ON branch and
observing the test still passed. Now written via `updateGlobalSettings`,
and the helper **asserts the flag took effect** before the journey runs.
2. **The journey was forward-only**, so it never reached the reopen
hook's field resets (`status`, `error`, `blockedBy`, pause clearing) — a
mutation there passed. Extended with a backward move and a re-entry.

Mutation-verified after both fixes: stamping seam 3, and diverging the
reopen hook, each fail the comparison.

**Precondition 2 — done, and its answer is a blocker.** The census says
the default board is safe: all 41 literal engine move targets are
declared by the default lineage. But **20 of those 41 carry no
`recoveryRehome`**, so on a custom lineage that does not declare `todo`
/ `in-progress` / `done`, seam 2 would start rejecting them with
unknown-column. That is a user-facing break on custom boards, not a
theoretical one, and it is not fixed by the equivalence proof — seam 2
adds *new refusals* rather than swapping implementations.

**So the flip is one step away, and the step is not mine to take
alone:** those 20 call sites need to resolve their target from the
task's workflow (or justify `recoveryRehome`), and they live across
engine lanes in `moves.ts` caller territory — U2b/MAIN. Flipping before
that trades a dormant flag for broken custom boards.

What remains for precondition 3 once those land: flip both readers
**atomically** (`moves.ts` + `workflow-task-create-ops.ts`, since the
latter computes the preflight the former consumes), delete the flag-OFF
branch with its guards, and drop the settings key.

---

## CORRECTION: seam 2 is not a blocker. My earlier claim was wrong.

I stated in #2639 and above that "with the flag off there is **no**
target-column validation on the move path", so flipping would introduce
new refusals. **That is not what happens.** Reproduced against live PG:
the identical custom-lineage move rejects with the flag **OFF** as well
—

```
Error: Invalid transition: 'backlog' -> 'todo'. Valid targets: building
```

Transition validation is already in force on the flag-OFF path. So for
the shape in question — an engine move to a column the task's own
workflow does not declare — **the move already fails today**, and seam 2
introduces no new break for it. The 20 census sites lacking
`recoveryRehome` are broken on a custom lineage *now*, not broken by the
flip.

I found this because the discriminator I added to prove "the flag is the
cause" failed. Had I written the test to my assumption it would have
passed and the false claim would have shipped — the same way the
equivalence test passed while proving nothing until I tried to make it
fail.

**Revised precondition status:**

| precondition | status |
|---|---|
| 1 — side-effect equivalence | **discharged** — identical rows,
mutation-verified both directions |
| 2 — seam-2 exposure census | **discharged, and it is not a blocker** —
the rejection predates the flag |
| 3 — flip both readers atomically, delete the flag-OFF branch, drop the
settings key | **the remaining work** |

So the flip is no longer gated on fixing 20 engine call sites. What it
is still gated on is precondition 3 being done atomically across
`moves.ts` and `workflow-task-create-ops.ts` (the latter computes the
preflight the former consumes), which is `moves.ts` caller territory.

Three cases now cover seam 2: the flag-ON rejection, the flag-OFF
rejection (asserting the error *message*, so a change in which guard
rejects stays visible rather than reading as agreement), and the #1411
`recoveryRehome` carve-out succeeding — pinning why that carve-out is
load-bearing and must not be tidied away.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:13:58 -07:00
gsxdsm
3aa942ee5f capacity: spawned agents count against the project agent count (#2579)
Two configurable numbers per project. `maxSpawnedAgentsPerParent` (5)
and `maxSpawnedAgentsGlobal` (20) were a **third and fourth** limiter
with private budgets invisible to both.

## This closes a hole, not just knobs

A spawned child **is** an agent and gets **its own git worktree**
(branched from the parent’s — the tool’s own description says so), but
children were counted by **neither** capacity gate. A fan-out could put
up to 20 extra worktrees on disk while the scheduler believed the
project was at its configured limit. The operator’s two numbers were
simply wrong about what was running.

## The old caps also measured the wrong thing

`totalSpawnedCount` decrements on child cleanup, but the per-parent
**set** is cleared only when the **parent task** ends. So
`maxSpawnedAgentsPerParent` throttled *cumulative* spawns across a
task’s life rather than *concurrent* ones — a long-running task could
exhaust its budget with five children that had all long since finished,
and the operator had no way to see why.

## Fix

`fn_spawn_agent` gates on the same project agent count every other lane
uses (`computeTopLevelConcurrencyClaimedFromStore`) plus live children.
One number, one answer, no private budget that can disagree with the
board.

The refusal names **Max Concurrent Tasks** — a control the operator
actually has. The old messages pointed at settings that no longer exist,
which is worse than no message: it sends someone hunting for a knob that
is not there.

## Verification

**Revert-proof, measured:** restoring the private budgets turns **3 of
the 4** new cases red — a project at 1/1 could still spawn, which is
precisely the hole. `executor.ts` restored byte-identical.

`pnpm lint` clean · core + engine `tsc` clean · `pnpm test:gate` green
(414 + 10 + 71) · new suite 4/4 · `settings-default-descriptions` 4/4.

There was no spawn-capacity test before this; the file is new.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Spawned agents now count toward the project’s **Max Concurrent Tasks**
capacity.
* Agent spawning is blocked when capacity is reached, including
concurrent spawn attempts.
* **Bug Fixes**
  * Prevented over-allocation during simultaneous agent spawns.
  * Restored available capacity when agent creation fails.
* **Changes**
  * Removed separate per-parent and global spawned-agent limits.
  * Updated settings to reflect the revised capacity controls.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:41 -07:00
Phil Larson
15b21dead1 fix(dashboard): reconcile task state through live API (#2595)
## Summary

- add a project-scoped live API route for updating individual task
checklist steps
- add an atomic live API route for resolving stale durable wedge
episodes
- prevent operator repair tooling from opening a second embedded store
that can diverge from the running dashboard backend

## Why

Legacy graph-native workflow runs can retain successful
`workflowStepResults` while their narrative checklist remains at 0/N.
The existing `fn task update` fallback may open a separate embedded
store, producing split-brain writes that do not accumulate in the live
dashboard backend. There was also no API surface for the existing atomic
wedge-episode resolver.

## Verification

- `pnpm exec vitest run
src/routes/__tests__/register-task-workflow-routes.step-update.test.ts`
— 5/5 passing
- `pnpm build` in `packages/dashboard` — passing
- full managed runtime workspace build — passing
- deployed to the managed local runtime and used to reconcile six legacy
review-deadlock tasks
- live board audit: zero `in-review-stall-deadlock` paused reasons
- exact local and Tailscale dashboard roots: HTTP 200 with 16,926-byte
bodies


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added live API endpoints to update individual task checklist steps
with validation (step index and allowed status values).
* Added an endpoint to reconcile/resolve stale task “wedge” episodes,
resolving only the matching active episode and returning conflicts on
mismatches.
* **Tests**
* Expanded route tests for step updates and wedge resolution, including
consistent 404 behavior for soft-deleted and missing tasks, plus
conflict and invalid-input cases.
* Expanded PostgreSQL coverage for wedge resolution persistence and
concurrent episode replacement scenarios.
* **Bug Fixes**
* Improved task-lookup error handling so soft-deleted tasks are
consistently treated as “not found” (HTTP 404).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 23:27:32 -07:00
gsxdsm
f91b8a4178 TAKING mission-feature-sync.ts: roadmap reconciliation resolves lifecycle roles (unowned drift site) (#2602)
> **Taking `packages/engine/src/mission-feature-sync.ts`** from the
shared backlog — announced in the title per the collision protocol.
Based on `main`, no dependencies.

It is in **no unit's file list**: absent from the plan's per-file census
*and* from the drift review's ownership split (self-healing, dashboard,
triage/replan-target, core, executor). It is a planning-lane reader.

## What was broken

`reconcileMissionFeatureState` maps a task's lifecycle **position** onto
its mission feature's roadmap status, and read five column literals:
`done`, `archived`, `in-progress`, `in-review`, `triage`/`todo`.

On a renamed workflow **every branch answers "no"**, so the function
collapses to a permanent `noop`.

**What an operator sees:** a mission roadmap frozen at whatever status
it last held, while the tasks underneath it run to completion. Nothing
errors, nothing retries. Worse than a wrong status, because a stale
roadmap reads as a stable one.

## Guard counts (per the reporting requirement)

| Metric | Before | After |
|---|---:|---:|
| `column === / !== "triage"` in this file | **1** | **1** |
| role comparisons converted | — | **5** |

**The metric does not move here, and I am not claiming it does.** The
five role comparisons are converted; the one literal that remains is the
deliberate scoped migration acceptance this change *adds*. That is the
third time on this program the real fix has been invisible to the
convergence count — the count finds the site, it does not define done.
Worth knowing while the shared backlog is being tracked by that number:
repo-wide it currently reads **29** triage comparisons (including 4 in
`plugins/`, which are also unowned).

## Fallback direction matters

**Unresolvable workflow falls back to the legacy ids, not to `noop`.** A
mission whose workflow cannot be read should keep tracking on the
default vocabulary rather than go silent — going silent *is* the failure
being fixed, so the fallback must not reproduce it.

**The planner-lane branch also accepts an orphaned legacy id.** A
pre-existing test asserted a card in `triage` returns its feature to
`triaged`; that stopped holding for the default lineage after #2515 —
the migration-window population again. Accepting `triage`/`todo`
additively keeps those rows tracked, **scoped to ids the workflow does
not declare**, for the reason greptile gave on #2593: a custom workflow
may legitimately name its **review** lane `triage`, and mapping a card
there to `triaged` would walk the roadmap backwards while the task is
awaiting merge.

## A test of mine that proved nothing until fixed

The scoping case first used a `triaged` feature. The planner-lane branch
only fires for an **in-progress** feature, so the fixture fell through
to the review branch and **passed under both implementations**. It
discriminates only once the feature status lets the wrong branch win —
verified by reverting the scoping and watching exactly that case fail.

## Revert proofs

| Reverted | Result |
|---|---|
| all five literals restored | **5 of 15 fail** — every renamed case;
every default case passes |
| legacy acceptance unscoped | **1 of 15 fails** — the
custom-`triage`-as-review case |

## Verification

| Check | Result |
|---|---|
| new suite | 15/15 |
| pre-existing mission-feature-sync + mission-autopilot +
scheduler-trait-dispatch | 94/94, **no expectation edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (482 + 10 + 71) |
| `pnpm check:changesets` | clean |

## Next from the shared backlog

Taking
`plugins/fusion-plugin-even-realities-glasses/src/agent-actions.ts` (3)
and `plugins/fusion-plugin-dependency-graph/src/GraphTaskNode.tsx` (1)
next — 4 sites in `plugins/`, which no unit owns and which the #2587
ratchet now scans. Shout if anyone is already there.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:06:55 -07:00
gsxdsm
50ebf3c543 TAKING cli/project.ts (fn project reported 0 running agents) + two test fixes — dashboard conversions WITHDRAWN in favour of #2626 and #2636 (#2631)
Three app-cluster conversions plus the evidence that they behave on a
renamed AND a merged board.

## Per-file guard counts

| file | before | after | note |
|---|---|---|---|
| `packages/cli/src/commands/project.ts` | 0 | 0 | not a comparison site
— see below |
| `packages/dashboard/app/components/TaskContextMenu.tsx` | 2 | 2 |
**count does not move — deliberate, see below** |
| `packages/dashboard/app/components/Column.tsx` | 2 | 2 | **count does
not move — deliberate, see below** |

**Read this before scoring the PR against the bar.** You said a claim
that does not move your number is not done, so I am telling you up front
that *this PR does not move it*, and why.

Both dashboard conversions are **fallback-preserving**:

```ts
const isIntakeColumn = columnFlags ? columnFlags.intake === true : column === "triage";
```

The literal survives as the no-flags branch, so the grep still counts
it. That is the shape the sibling code already uses
(`isPreExecutionHoldColumn`, same file, converted earlier in the
program), and dropping the fallback would make an unresolved-column
render *lose* the affordance a second way. What changes is the
**behaviour when flags exist** — which is what the mutation results
below measure.

If you want these to zero out the count, the fallback has to go, and
that is a separate decision about whether an unresolved column should
fail open or closed. Say the word and I will do it as a follow-up; I did
not make that call unilaterally because it is not reversible from a
rendering standpoint.

`cli/project.ts` was never a comparison site at all — it fed **raw
rows** to `isRunningAgentTaskShape`, so the helper's own internal legacy
fallback kicked in and `fn project` reported **0 running agents** on any
renamed board. Fixed by resolving the IR per task before counting.
Nothing to subtract.

## Two of the three had a test that looked like coverage and was not

- **`Column.tsx`** — the quick-create gate is `workflowMode ||
isIntakeColumn`. Every pre-existing intake case in `Column.test.tsx`
*also* passes `workflowMode`, so the `||` short-circuited and **none of
them ever reached the trait lookup**. Added cases that omit
`workflowMode`, the only path where the conversion changes the answer.
- **`TaskContextMenu.tsx`** — the intake suppression was asserted only
for the legacy `triage` id, the one board shape where a broken
conversion still returns the right answer.

Mutation-verified rather than asserted:

| mutation | result |
|---|---|
| `isIntakeColumn` → `column === "triage"` | **2 of 88 fail** (exactly
the renamed and merged cases) |
| menu suppression → `task.column !== "triage"` | **1 of 12 fail** |

## A pre-existing red I fixed on the way past

`uses VALID_TRANSITIONS and in-review back-to-progress labels` was
**already failing on origin/main**. #2521 correctly moved the "Back to
X" label onto the host's `columnLabel` function; this file's stub is
`(column) => column`, so the hardcoded `"Back to In Progress"`
expectation was left over from the pre-#2521 hardcode and nothing had
updated it.

Matching the raw id would have made it pass while proving nothing, so
instead that one case gets a display-like label function — the assertion
now fails both if the "Back to" prefix regresses **and** if the label
stops routing through `columnLabel`. Strengthened, not relaxed. Counts
against completion criterion #2.

## Verification

- `Column.test.tsx` + `TaskContextMenu.test.tsx`: **100 passed**
- `tsc -p tsconfig.app.json` (the root config does not cover `app/`) and
the CLI typecheck: clean
- `pnpm test:gate`: green

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:52:30 -07:00
Phil Larson
a54d60ee70 fix(core): restore standalone central backend initialization (#2596)
## Summary

- restore the owned PostgreSQL backend bootstrap for layer-less
`CentralCore.init()` callers
- fix node, mesh, and project CLI commands returning empty state and
logging `backendHandle is only available in backend mode` during cleanup
- add a hermetic regression test for standalone backend ownership and
shutdown

PR #2454 accidentally added an unconditional early return immediately
before the existing standalone bootstrap. Runtime pool sharing remains
unchanged: `attachBackendLayer()` releases the central-only connections
before adopting the project store layer.

## Verification

- RED: regression failed because `createCentralBackendLayer` had zero
calls
- GREEN: focused regression passes
- `pnpm --filter @fusion/core typecheck`
- `pnpm --filter @fusion/core build`
- `pnpm test` with `FUSION_PG_TEST_SKIP=1`: 482 engine + 132 Core gate +
71 CLI shape + changed regression passed; isolation clean
- changeset format check passed

The local PostgreSQL merge-gate harness is unavailable without
credentials (`empty password returned by client`), so its 10 tests were
explicitly skipped rather than misreported.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Restored PostgreSQL central registry access for standalone
Node/mesh/project CLI commands without marking the host offline on
shutdown.
* Improved CentralCore lifecycle handling: concurrent `init()`
coalesces, and operations are blocked once `close()` is requested/in
progress.
* Refined embedded PostgreSQL runtime shutdown: owner stop is
coordinated with lease release, registrations are rejected while
stopping, and shutdown/teardown uses lease lifecycle consistently.
Embedded start failures now treat stopping as retryable.
* **Tests**
* Expanded coverage for CentralCore close/init/attach races and embedded
PostgreSQL lease/shutdown coordination scenarios.
* **Documentation**
  * Updated Changeset notes to clarify CLI and shutdown semantics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 22:41:17 -07:00
gsxdsm
76b513e028 comments-ops.ts: user comments stopped invalidating spec approval (guards 3 → 0) (#2606)
Taking **`packages/core/src/task-store/comments-ops.ts`** from the
shared 48-guard backlog.

| file | before | after |
|---|---:|---:|
| `packages/core/src/task-store/comments-ops.ts` | 3 | **0** |

## One of the three was a live defect

The awaiting-approval branch read:

```ts
task.column === "triage" && task.status === "awaiting-approval"
```

#2515 merged the two pre-implementation columns into one with id `todo`,
so a card awaiting spec approval now sits in `todo` and **that condition
can never match**. A user comment on such a card silently stopped
invalidating the approval — the operator types a correction, the spec
stays approved, and the task proceeds on the very spec they were
correcting.

No error, no log line, nothing to notice. This is exactly the failure
mode the census exists to eliminate, and it is user-visible: the
operator’s correction is accepted into the comment thread and then
ignored by the pipeline.

The other two guards survived by luck — their `column === "todo"` arm
still matched the merged column, so only the dead `triage` arm was
inert.

## Fix

All three resolve the **intake/hold roles** from the task’s own
workflow. Unresolvable workflows fall back to the legacy pair: this is a
best-effort re-triage path whose failure mode is a *missed* re-spec, so
degrading to the old vocabulary beats dropping the card out of the
branch entirely.

## Verification

Regression test drives the **real store** on the merged column and
asserts the approval is invalidated.

**Revert-proof, measured:** restoring the `triage` literal fails with
`expected awaiting-approval not to be awaiting-approval`.
`comments-ops.ts` restored byte-identical.

`pnpm lint` clean · core `tsc` clean · `pnpm test:gate` green (482 +
132) · `store-comments` 15/15.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:40:15 -07:00
gsxdsm
f14059e8d3 Retry still refuses cards parked mid-planning on 5 builtins (survives #2614; count 0 → 0, defect-only) (#2621)
**Rebased onto main after #2614 landed this file.** That PR's conversion
already took the tracked count for `register-task-workflow-routes.ts` to
**0**, so this PR does **not** move your number and I am not claiming it
does.

| file | before | after |
|---|---:|---:|
| `packages/dashboard/src/routes/register-task-workflow-routes.ts` | 0
(post-#2614) | 0 |

What it fixes is a **live 400** that #2614 left in place. Measured on
current main: **9 of this file's 14 retry tests fail** without the
change below.

## The defect

`POST /api/tasks/:id/retry` must answer *"does this card sit where its
workflow **plans**?"*, because the yes-branch is **destructive** — it
stamps `needs-replan` **and deletes PROMPT.md**. Two predicates stood in
for that question and neither answered it:

- **#2614** resolved the **intake** column. Correct for the merged
lineage; wrong wherever intake and the planning column differ.
- The older arm asked `!workflowHasColumn(ir, "triage")`.

**Measured across all 12 builtins:** *not one* plans in `triage`, while
**seven** still declare that column. So for the five that declare
`triage` **and** run every plan node in `todo` — `quick-fix`,
`review-heavy`, `compound-engineering`, `design`, `legacy-coding` — the
predicate is `false` and a `planning`/`needs-replan` card sitting in
**its own planning column** is refused outright:

```
400 — "Task is not in a retryable state (current status: needs-replan)"
```

The operator has no button at all on a card parked mid-planning. The
mirror-image fault is destructive rather than obstructive: a workflow
that plans anywhere other than `todo` had a `todo` card's PROMPT.md
deleted for a re-plan nobody asked for.

## Fix

`workflowPlansInColumn` asks the graph. Planning nodes are recognised by
the **semantic markers** the builtins carry — `config.seam ===
"planning"` and an **exact** `workflowAction` set (measured vocabulary:
`plan-replan`, `code-review`, `pre-merge-remediation`) — with node ids
as a backstop.

Deliberately **not** a `startsWith("plan")` prefix. That was my first
attempt and greptile was right to kill it: it matched in the
**destructive** direction, classifying a custom `plan-execute` column as
a planning column, which deletes a specification. An unlisted planning
action costs a replan (recoverable, card stays retryable); a
wrongly-listed one costs a spec (not). Hence opt-in.

### Second concern, split out

Narrowing the destructive branch must not narrow **retryability** —
those were one boolean and are two questions. A card parked outside its
planning column would otherwise fail the gate and answer 400: that
trades *a card which loses its spec* for *a card nothing can rescue*. It
stays retryable via the non-destructive branch, scoped to pre-WIP
columns so no `in-progress`/`in-review` status gains a path it lacked.

A **v1 IR** declares neither columns nor nodes, so placement is
**unanswerable** rather than answered "no".
`workflowDeclaresColumnModel` distinguishes the two — reading that
silence as "past planning" is exactly what 400'd a v1 planning card.

## Verification

All three greptile P1s on the earlier revision were real and are fixed
with revert-proof tests (bespoke planning-node ids; the v1 regression I
introduced; my own loose action prefix).

`pnpm lint` clean · dashboard `tsc` clean · `pnpm test:gate` green (132
+ 10 + 482 + 71) · core 11/11 · dashboard 119/119
(`retry-planning-column` + `stale-merge-status` + `routes-tasks`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:39:33 -07:00
gsxdsm
31e49b684a TAKING default-workflow-hooks.ts + executor.ts + live-agent-count.ts + 6 dashboard files: reopen semantics by role, and the census's blind spot in both directions (13 sites) (#2628)
Batched conversion of every lifecycle-column guard I hold, plus the
three the census could not see. **Six files to zero, repo-wide 60 → 49
by a comment-stripped unanchored sweep.** Each conversion has an
isolated revert proof and a paired negative case, and the one code move
is a separate commit from the behavior changes.

## Per-file before → after

Counts from a comment-stripped, unanchored `(===|!==) ["']triage["']`
sweep over `packages/*/src` + `plugins/*/src`, excluding tests.

| file | before | after | note |
|---|---:|---:|---|
| `core/default-workflow-hooks.ts` | 4 | **0** | |
| `core/task-store/moves.ts` | 5 | **4** | only the flag-ON mirror
converted; the flag-OFF inline block is the parity reference and stays |
| `engine/executor.ts` | 3 | **0** | **absent from the 45-guard list** —
see below |
| `core/live-agent-count.ts` | 2 | **0** | duplication removed; answer
deliberately unchanged |
| `engine/replan-target.ts` | 2 | **0** | both were comment prose, not
guards |
| `core/agent-prompts.ts` | 3 | **0** | ROLE comparisons, never column
guards |
| `engine/usage-limit-detector.ts` | 2 | **0** | ROLE comparisons |
| `dashboard/app/components/DocumentsView.tsx` | 1 | **0** | real column
guard |
| `dashboard/app/components/TaskChatTab.tsx` | 2 | **0** | ROLE |
| `dashboard/app/components/AgentLogViewer.tsx` | 1 | **0** | ROLE |
| `dashboard/app/components/effective-model-resolution.ts` | 1 | **0** |
ROLE |
| `dashboard/app/hooks/useTasks.ts` | 1 | **0** | ROLE |
| `dashboard/…/command-center/MissionControlPanel.tsx` | 1 | 1 | alias
table, marked `DELIBERATE-LITERAL` with its reason |

## The census errs in BOTH directions

This is the finding I would most like carried into the remaining work.

- It **flagged 10 sites that were never column guards.** `role ===
"triage"` / `agentType === "triage"` compare an **AGENT ROLE**. The
planner *lane* is named `triage` and keeps that name — U11 removed the
*column*. Worse than noise: the obvious "finish the migration" edit is
to rename the role, and that silently empties the planner's prompt
template and mis-binds its model markers. `PLANNER_AGENT_ROLE` now names
it, so the two vocabularies are distinguishable by grep and a rename
fails loudly (revert proof: 4 tests, two of them pre-existing).
- It **missed 3 real guards in `executor.ts`**, because the pattern
matches `column`/`toColumn`/`fromColumn` and those locals are named
`from` and `originColumn`. A census keyed on variable names will keep
missing guards wherever a local was named for its role in the function.

## Two real defects, not tidying

**1. A renamed board could merge with its re-review never run.**
`default-workflow-hooks.ts` is named for the default workflow, but the
store runs it on the flag-ON path for *every* workflow — the trait
registry resolves hooks by trait id, not by workflow. Its reopen
predicates listed the default lineage's column names, so on a renamed
board **no reopen effect fired at all**. One of them clears
`workflowStepResults`, which `getTaskMergeBlocker` reads: a card bounced
out of review carried its old `passed` result back in, and that
satisfies the merge gate. Same regression the graph-owned-crossing
carve-out exists to prevent, arriving through the other door. (Two
smaller ones rode along: failure state never cleared on a renamed
reopen, and an operator dragging a card back to the queue never parked
it, so the scheduler re-dispatched what they had just pulled back.)

**I forgot the carve-out on my first pass, and that was worse than not
converting.** A role-resolved clear plus a *name*-matched exemption
means a renamed board takes the clear and never the exemption,
destroying the remediation input the graph had just written. My own
paired negative test caught it.

**2. The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** In `recoverCompletedTask`,
`promotedFromPlannerColumn` was false on a renamed board, so finished
work resting in the planning lane was never promoted — the code fell
through to `handoffTaskToReview` straight from the planning column, and
role adjacency has no planning → review edge, so the handoff was
rejected and the card stayed stuck with its work complete. I converted
the promotion **target** too: resolving the lane and then moving to a
literal `in-progress` is the half-conversion I have already been burned
by twice this program, where the guard starts admitting cards and the
move then sends them to a column the board does not declare.

## E2E evidence

`renamed-board-reopen.pg.test.ts` drives a **real PostgreSQL store** and
a real `moveTask` on a workflow whose columns carry the standard traits
under non-default names. The unit tests cannot show this: if `moves.ts`
passed `undefined`, every unit case still passes via the no-basis
fallback while the real board keeps the old behavior. **Proof it is
load-bearing: forcing `moveLifecycleColumns` to `undefined` fails 2 of
3.** The executor suite covers both the split-role and the MERGED
post-U11 shape.

## Revert proofs, isolated per site

| change reverted | result |
|---|---|
| reopen predicate → literal names | 4 of 10 fail |
| reopen field clears → literal names | 2 of 10 fail |
| `userPaused` hold lane → literal `todo` | 1 of 10 fail |
| graph carve-out → literal names | 1 of 10 fail |
| store passes `undefined` lifecycle columns | 2 of 3 fail (real PG) |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| two-hop condition → `=== "triage"` | 1 of 7 fails |
| promotion target → `"in-progress"` | 3 of 7 fail |
| `isPlannerColumnFor` → literals | 1 of 7 fails |
| live-agent-count: one arm dropped | 2 of 11 fail |
| DocumentsView: trait branch removed | 3 of 7 fail |
| planner role renamed to `"planner"` | 4 fail (2 pre-existing) |

Every conversion is paired with a negative case (a forward move, a
not-a-planner-lane card, a default-lineage card, a renamed column with
no traits), so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Deliberately NOT converted, with reasons

- **`moves.ts` flag-OFF inline block (4).** That branch *is* the legacy
path, kept verbatim so the two can be parity-checked. Converting it
erases the reference implementation.
- **`live-agent-count.ts`'s no-flags fallback.** Reachable, and there is
nothing to resolve from — `enrich…FromFlags` exists for callers with
board flags rather than an IR, so a column missing from that map is the
renamed case. "Not intake" is as much a guess as "todo is intake", and
Running/Waiting are complements, so a card matching neither arm is
reported as neither and the footer's queued total under-reports it. The
real fix is at the caller; four new cases pin that flags override the
legacy answer **in both directions**. What did change is the
duplication: two hand-written copies of one rule now call one named
function.
- **`MissionControlPanel`'s `FUNNEL_STAGES`.** An alias table of column
*names* where `triage` sits beside `signal` and `backlog`. Command
Center aggregates across projects, so there is no single workflow to
resolve traits from — the honest conversion is a data change, not a
predicate change.
- **`DocumentsView` with no traits.** Same no-basis rule; the documents
list is full of historical columns absent from the current board. A case
asserts a renamed column with no traits still reads as "working",
documenting the gap rather than hiding it.

## Fixture findings

Each cost a red run that looked like the code under test:

- a `merge-blocker` column needs a reachable merge-class node, or
`parseWorkflowIr` rejects the workflow;
- a back-edge must be `kind: "rework"`, and a rework edge is legal only
**into** a node with `config.reworkRegion: true`;
- a workflow gets role-level transitions only when it declares wip +
review + complete + **archived** plus a planning lane — without the
archived column, adjacency falls back to order-derived neighbours and
`checking -> queued` is not a legal move at all;
- `recoverCompletedTask` only *reaches* the promotion seam when nothing
is left to gate; without passed `plan-review`/`code-review` rows it
re-enters the workflow graph and returns first, so a naive fixture
silently tests the wrong branch and every assertion reads "no moves
happened" for an unrelated reason.

## Verification

- `pnpm test:gate` **71/71**
- new suites: 10/10 reopen-semantics, 3/3 renamed-board-reopen (real
PG), 7/7 executor-planner-lanes, 7/7 documents-status-dot, 4/4
planner-role-is-not-a-column
- neighbours: 132 + 10 + 482 (gate shards), 350/351 engine
planning/replan suites, 64/64 agent-prompts, 51/51 usage-limit-detector,
11/11 live-agent-count, 11/11 dashboard hook/log suites
- the single engine failure (`executor-fast-mode-workflows.test.ts` ›
"raw fast mode still invokes non-executable review seam nodes")
**reproduces with my changes stashed** — pre-existing on `origin/main`
- typechecks clean for core, engine, and dashboard-app
(`tsconfig.app.json`; `tsconfig.json` checks nothing under `app/`);
`pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:39:14 -07:00
gsxdsm
d438cd1d13 U12 drift: register-task-workflow-routes.ts — resolve the intake column (7 -> 1) (#2614)
**File claimed:
`packages/dashboard/src/routes/register-task-workflow-routes.ts`.**
Per-file lifecycle-column guard count: **7 → 1**, and the 1 is comment
prose (line 3758), so this file is done for completion-bar item 1.

## The bug this fixes

`retrySpecification` decided "this Retry is a re-plan, not a generic
retry" with `task.column === "triage"`. `status: "planning"` is
retryable **only** through that flag — it is not in the generic
`failed`/`stuck-killed` set. So on any lineage whose intake column is
not literally named `triage`, a card visibly sitting in planning got
`400 Task is not in a retryable state`. The operator's Retry button did
nothing, with no error to explain why.

Post-#2515 that includes the **default** workflow:
`columnsWithFlag(resolveDefaultWorkflowIr(), "intake")` is `["todo"]`
and the default's columns are `[todo, in-progress, in-review, done,
archived]` — `triage` is not declared at all. The pre-existing `todo`
fallback below it papered over the default case (it fires when the
workflow has no `triage`), which is why this did not show up as a total
outage; custom and renamed lineages had no such cover.

Now: `const retryIntakeColumn = await
resolveIntakeColumnForTask(scopedStore, task.id)`.

## Red-green, measured

`packages/dashboard/src/__tests__/plan-approval-intake-column.test.ts` —
new case, custom lineage with intake `backlog`, card in `backlog` with
`status: "planning"`:

- with the change: `200`
- with `task.column === "triage"` restored: **`AssertionError: expected
400 to be 200`**

The fixture uses `planning` deliberately. A `failed` fixture would pass
either way through the generic retryable set and prove nothing.

## What is NOT tested, and why not

This PR also removes four `&& task.column !== "triage"` disjuncts I
added earlier while widening the P0 approve/reject guard. **Those are
untestable by construction** and I am not claiming coverage for them:
removing an extra acceptance only shrinks what the guard accepts, and no
case can feed these routes a `triage` card now that no shipped lineage
declares one. I re-widened one guard and confirmed the suite stays green
— i.e. nothing depends on the disjunct in either direction. That is the
honest result, not a passing test.

## Three fixtures updated, not guards re-widened

`stranded-refinements-routes.test.ts` failed with three `expected 400 to
be 200` — the same failures that made me widen in the first place. This
time I probed instead: `BASE_TASK` had `column: "triage"`, a column the
default workflow no longer declares, so a 400 is **correct** and the
fixtures were pre-merge artifacts describing a board shape the product
stopped shipping. Changed to `column: "todo"` with the resolver output
recorded in the file.

## Observation, deliberately not fixed here

The `todo` fallback at ~2642 (`retrySpecification =
!workflowHasColumn(workflowIr, "triage")`) is now near-dead: for the
merged default the first branch already fires. It survives only for a
lineage that has a `todo` column, no `triage`, and some *other* intake
column — where treating a `todo` card as planning is arguably wrong.
Deleting it is a behaviour change with its own blast radius, so it does
not ride along in a conversion commit.

## Verification

`pnpm lint` clean. `pnpm test:gate` green (10 / 482 / 71). Target
suites: `plan-approval-intake-column.test.ts` 8/8,
`stranded-refinements-routes.test.ts` 12/12.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:34:31 -07:00
gsxdsm
89d6d76d60 Unowned: the R7 sweep guessed with another workflow's columns — its "do not guess" guard was unreachable dead code (#2600)
## Unowned: the R7 sweep's "do not guess a column" guard could not fire

Picked up from my own #2543 finding. Independent of my other PRs.

### The guard existed in comment form only

`reconcileUndeclaredTaskColumns` wraps IR resolution in a try/catch
whose comment reads:

> An unresolvable workflow is its own fault path; do not guess a column.

But `resolveWorkflowIrById` catches **every** failure and returns
`defaultCodingWorkflowIr()`, and `resolveWorkflowIrForTask` does the
same for a failed selection read. The resolver never rejects, so that
catch is **dead code**.

What actually happened to a card whose workflow could not be loaded: it
was judged against the **default** workflow, and if its column was not
one the default declares, the sweep re-homed it to the **default's**
rebound target. It guessed, using a workflow that is not the card's own
— the precise outcome the guard was written to prevent, in a **startup
recovery path that runs against every task**.

### How it was found, which is the part worth keeping

By being **unable to make a test of the guard fail**. Three separate
mutations all passed — deleting the `continue`, deleting the try/catch,
and simulating a whole-sweep abort at that very catch. I had written
that off once as "this case pins the outcome, not the mechanism". The
inability was the signal, not a limitation of the assertion: the branch
is unreachable.

This is the seventh instance of the program's core shape, and the first
I found in a guard I had just finished writing coverage for.

### The fix

The sweep now **proves the resolved IR belongs to the task** before
moving its card: it reads the task's workflow selection and confirms
that id resolves to a real definition (built-in or stored).

- A task with **no** selection legitimately resolves to the default
workflow — not treated as unresolvable.
- An unreadable selection **read** is itself grounds not to guess.

Placed at the **move site**, not at resolution, deliberately: it costs
one definition read only for a card already about to be moved — a
healthy board reaches that line for nobody — and it keeps the fix inside
the sweep instead of changing a resolver whose soft-failure many other
callers depend on. Changing `resolveWorkflowIrById` to reject would have
been the tidier-looking fix and a much wider blast radius.

### Revert-proof, both directions

- Remove the proof → the case fails `expected 2 to be 1`: the unloadable
card is re-homed on a guess.
- The same case asserts the neighbour **is** still repaired, so the fix
cannot be mistaken for letting one bad card disable the sweep for
everyone else. That is the per-task isolation property, and a
single-task fixture cannot distinguish it from a whole-sweep abort —
verified by injecting a throw at the loop head (`expected 0 to be 2`).

### Verification

`pnpm test:gate` (482 + 10 + 71), `pnpm lint`, engine typecheck green.
Sweep suite + `legacy-tombstones`: 13 passed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Prevented startup recovery from moving cards into incorrect columns
when their workflow cannot be loaded or resolved.
* Cards with unreadable workflow information now remain in place, while
other recoverable cards continue to be repaired correctly.
* Added safeguards to avoid guessing a fallback workflow during column
reconciliation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:21:04 -07:00
gsxdsm
3f763cba87 U8: the graph owns the pending-review park — ownership ledger 28 → 27 (#2590)
The routing move this unit has been building toward, landing on the path
the engine actually runs. **Includes #2578's commit** (the live-path fix
it depends on) — merge that first, or this supersedes it.

## What changes

Three things together, because a half-routed move is a card that
silently does not advance:

1. The **live** implementation primitive (`runCodingSession`) returns
`{outcome: "failure", value: "review-pending"}` for that ending.
2. The primitive step handler stops flattening every ending to
`step-done`/`step-failed`, so the value survives the foreach —
`runForeach` propagates a failing instance's value as the node's own —
and reaches an edge.
3. The inline `handoffTaskToReview` in `runImplementation` is
**deleted**. The phase reports and stops, which is all an implementation
phase should do.

Built-in workflows route to the `review-pending-handoff` node added in
#2519/#2546, which performs the handoff and ends the run: the same two
effects in the same order, with the graph as the owner.

## Proof, end to end

FN-5436 — the test that blocked this move twice and was right both times
— now passes, with a **stronger** assertion than it had:

```ts
expect(store.moveTask).toHaveBeenCalledWith("FN-5436-B", "in-review",
  expect.objectContaining({
    workflowMoveSource: "workflow-graph",
    workflowMoveMetadata: expect.objectContaining({ nodeId: "review-pending-handoff" }),
  }));
```

The old two-argument `moveTask(id, "in-review")` could not distinguish a
graph-owned park from an out-of-band one — which is the entire
distinction this unit exists to make. The invariant (park in review,
never `failed`) is unchanged; the owner is now proven.

## Every ratchet fired, and each records a real change

| Ratchet | Before | After | Why |
|---|---|---|---|
| Ownership ledger — `runImplementation` review handoffs | 3 | **2** |
the handoff left the phase |
| Ownership ledger — `handleGraphFailure` | 0 | **1** | the named compat
classifier |
| Ledger headline — executor-owned dispositions | 28 | **27** | first
decrement of the unit |
| Out-of-band exit list | 2 | **1** | pending-review is graph-owned now
|
| Primitive routing pin | "must not reroute" | routes *only* the moved
ending | declared, not discovered |

None was relaxed. The `handleGraphFailure` 0 → 1 is the honest one: for
a user-authored graph without the edge this is a **relocation, not an
elimination** — the transition is still executor-performed, but from one
named classifier in the failure ladder rather than a call buried two
thousand lines into a session loop. The ledger says so rather than
letting the headline number imply more progress than there is.

## Why it took four attempts

Recorded because the reason is reusable: the value was being produced on
`createAuthoritativeWorkflowSeams`, a handler that never runs (#2578).
Every earlier attempt was correct code on a dead path, and the only
thing that showed it was instrumenting until a negative result was
proven observable rather than assumed.

## Verification

- step-session + exit-events + primitive-exit-events + ownership ledger
+ graph-requeue-gate + task-done-blocked — **83 tests green**
- `pnpm test:gate` green (10 / 482 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved handling of tasks awaiting review so they are correctly
routed to the review workflow.
* Tasks now remain in review instead of being marked as failed when no
follow-up review route is configured.
* Review handoffs now include workflow ownership and provenance details.
* Preserved standard failure handling for tasks that are not awaiting
review.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:59:06 -07:00
gsxdsm
131feb243c U8: the exit announcement was on a dead code path — move it to the handler the engine actually runs (#2578)
A merged behavior of mine has never executed. This fixes it and adds the
ratchet that would have caught it.

## The finding

`createDefaultNodeHandlers` chooses the prompt-node handler like this:

```ts
const promptLike = deps?.primitives
  ? createPrimitivePromptLikeHandler(deps.primitives, runCustomNode)
  : createPromptLikeHandler(seams, runCustomNode);
```

`executeWorkflowGraph` always passes `primitives:
this.createAuthoritativeWorkflowPrimitives(settings)`
(`executor.ts:6051`). **So `createPromptLikeHandler` — and with it every
`execute` / `step-execute` function in
`createAuthoritativeWorkflowSeams` — is unreachable for prompt nodes.**
Both objects are passed to the graph executor and only one is consulted.

The `NodeCompleted.exit` announcement added in #2507 was wired into that
seam. It type-checks, its tests pass (they call the seam object
directly), and it has never run in production. `runCodingSession` in the
primitives is the live twin, and that is where it emits now.

## How it was found — and why the negative is trustworthy

Instrumenting `createAuthoritativeWorkflowSeams.stepExecute` produced no
output for a run that demonstrably visits `steps#0:step-execute`. So did
instrumenting `createPromptLikeHandler`'s dispatch. A negative result
from instrumentation is worthless until the instrumentation is shown to
be observable, so: a `process.stderr.write` at module load of the same
file **did** appear, exactly once, in the same run. The two negatives
were real, not swallowed output.

This is also the answer to the open question I left in #2546 — the
pending-review routing move kept failing because the seam value it
depends on is never produced. **That move is still not landed here.**
This commit only relocates the announcement, so it stays small and
separately revertable; the routing move follows once its value
originates on the live path.

## The ratchet

A source assertion pins the dispatch rule: `deps?.primitives ?
createPrimitivePromptLikeHandler` and the executor's wiring of
`primitives`. Inverting or conditionalising that preference would
silently disable every behavior attached to the primitives path — the
same failure in the other direction — and **a seam-level unit test
cannot tell the two apart**, which is precisely how this survived review
twice.

## Red-green

Removing the emit fails 2 of the 4 new tests (`Tests 2 failed | 2 passed
(4)`). The other two are the regression floor: an ordinary completion
emits `success` with no `exit`, and the returned routing outcome is
unchanged — announcing must not reroute.

## Scope note

I did **not** delete the now-known-dead seam wiring in this PR.
`createAuthoritativeWorkflowSeams` is still passed to the graph executor
and its non-prompt entries (`stepReview`, `merge`) are reached through
other handlers, so deciding what is genuinely dead there is a deletion
audit of its own — and this program's rule is that deletions never ride
along with behavior changes. Filed as the next slice.

## Verification

- 4 new tests + exit-events + step-session + triage audit + ownership
ledger — **54 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:54 -07:00
gsxdsm
a68785a41d P0: two silent triage guards in the executor's ownership — one strands a card with nothing to rescue it (#2572)
P0 audit of the executor's assigned `triage` sites after the
Planning-column merge. **One of them can strand a card**, so leading
with that.

## The stall — `handleDepAbortCleanup`

`executor.ts` moved a dependency-aborted task to the **literal**
`triage`. The default coding lineage no longer declares that column.

A card that gains a dependency mid-execution has its work discarded and
is then parked in a column its own workflow does not define. Nothing in
the graph routes a card out of an undeclared column. The only rescue is
`reconcileUndeclaredTaskColumns`, which runs on the **next engine
start** — so between the abort and a restart the card is stalled with no
automatic recovery. It does not throw, so it would have surfaced as a
user report, not a red test.

Fixed to `resolveReboundColumnFor`, the helper the other ~16 executor
rebounds already use.

## The silent skip — `UsageLimitPauser.taskUsesProvider`

The planning lane was identified by the same literal. For a default card
the lane resolved to **no providers**, so when a provider hit a usage
limit during a *planning* session, the fan-out that pauses peers on that
provider skipped every default-workflow card and they kept hammering the
rate-limited provider.

Not a stall: the triggering task is still paused by the explicit
fallback below the filter. What was lost is blast-radius containment. A
planning session runs while the card is pre-implementation, and the
caller has already excluded `done`/`archived`, so that is exactly "not
the implementation column and not the review column" — which matches
`todo`, `triage`, `ideas`, and a renamed planner alike.

## Full audit table for my assigned sites

| Site | (a) Still fires for a default card? | (b) What silently stops |
(c) Action |
|---|---|---|---|
| `executor.ts:16395` `moveTask(id, "triage")` | **No** — writes an
undeclared column | Card parked where nothing routes it; rescue only at
next engine start | **Fixed** — `resolveReboundColumnFor` |
| `usage-limit-detector.ts:126` `column === "triage"` | **No** |
Usage-limit fan-out skips every default card; peers keep hitting the
limited provider | **Fixed** — pre-implementation predicate |
| `executor.ts:3409` `from === "todo" \|\| from === "triage"` | **Yes**,
via the `todo` arm | — | Unchanged; `triage` arm still live for
legacy-coding |
| `executor.ts:4951` `originColumn === "todo" \|\| === "triage"` |
**Yes**, via the `todo` arm | — | Unchanged |
| `executor.ts:4963` `originColumn === "triage"` double-hop | No, and
correctly so | Nothing — the extra hop exists only for shapes that
declare `triage` | Unchanged; still required by legacy-coding |
| `executor.ts:1110` `Type.Literal("triage")` | n/a | — | **Not a
column** — an agent ROLE in `spawnAgentParams` |

Counts for my ownership: **6 sites audited, 2 defects, 2 fixed, 3
correct as-is, 1 false positive.**

## Red-green

Reverting each fix fails its own test:

```
Tests  2 failed | 2 passed (4)
  × dependency-abort cleanup requeues to a DECLARED column
  × usage-limit fan-out … pauses a peer card sitting in the merged Planning column (id `todo`)
```

The other two are the regression floor and pass both ways by design: a
legacy workflow that **does** declare `triage` still fans out, and an
in-progress card is still **not** swept into the planning lane (the
guard must stay narrow — "any non-wip column" would have been the easy
wrong fix).

## Verification

- New audit suite + graph-boundary + step-session + ownership ledger —
**45 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 19:03:10 -07:00
gsxdsm
cf7b1a3d46 Drift review (unowned): gridlock detection + autopilot retries resolve the hold column — main 103→101 (#2561)
> **Based on `main`, not on my U7 stack** — merges in any order, no
dependency on #2517.

My assigned files (`triage.ts`, `replan-target.ts`) are at zero, so this
picks up two lifecycle-column literals **no unit's file list claims**.
Both ask *"is this card in the hold column?"* by the id `todo`, and both
are broken **today** for any workflow that renamed it.

## gridlock-detector — the worse of the two

`column !== "todo"` decides which cards count as **schedulable**, and an
empty schedulable set is an **early return**. On a renamed board the
detector concluded *"no gridlock"* at exactly the moment a real one
would be visible.

> A detector that goes quiet on the boards it cannot parse is worse than
one that is absent, because its silence reads as health.

**Converting only the `todo` half would have shipped a still-broken
detector**, and the test caught it. The `active` filter is equally
literal (`in-progress` / `in-review`) — and an empty active set is
*also* an early return. Two literals, one silence.

The `in-progress` half sits **outside the drift review's `todo|triage`
pattern**, which is precisely why a count-driven sweep would have left
it behind and declared the file done. Converted here rather than
deferred as out of scope. Worth flagging to the other workers: the
convergence metric is a good *tracker* but a bad *definition of done* —
an adjacent literal in the same predicate can preserve the whole bug at
a lower score.

## mission-autopilot

The retry compared against `todo` **and moved to the literal `todo`** —
so on a renamed workflow it relocated the card into a column the
workflow may not declare (R7) on **every retry**. Now resolves the hold
role; when the workflow declares none it leaves the card in place and
says so, because the error/status clear still runs, so the retry is not
lost — the card just stays in its own lane.

## Two fixture defects of my own, both caught by the tests failing
wrongly

**My first autopilot tests re-implemented the decision** and asserted on
the copy — proving only that the copy works. That is the anti-pattern
named in
`docs/solutions/store-fake-defects-that-masquerade-as-production-bugs.md`
(#2534) and in the #2527 ratchet review, and I had no excuse: the
constructor takes two stores and `handleTaskFailure` is public.
Rewritten to drive the real method.

**My first gridlock fixture failed on both vocabularies** — the detector
needs three preconditions and I supplied one. A test that fails on its
*no-regression* half is a broken fixture, not a discovered bug. The
"both halves failed" heuristic from that same doc is what flagged it.

That is eight fixture defects across this unit, every one caught by
reading *why* a test failed rather than making it pass.

## Revert proofs, each isolated to one literal

| Restored | Result |
|---|---|
| gridlock hold filter | **1 of 5 fails** (renamed case) |
| autopilot move target | **1 of 5 fails** (renamed case) |

Default-vocabulary halves pass either way — the correct signature for
conversions that change no existing behavior.

## Convergence

Measured against `origin/main` with a comment-stripped scan of `column
=== / !== "todo" | "triage"` in `packages/*/src`, excluding tests:

**103 → 101.**

(The gridlock `active` filter is a third site fixed here that this
pattern does not count.)

## Verification

| Check | Result |
|---|---|
| new suite | 5/5 |
| pre-existing gridlock + autopilot suites | 85/85, **no expectation
edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (414 + 10 + 71) |
| `pnpm check:changesets` | clean |

## Still unowned after this

`mission-feature-sync.ts` (1: a planning-lane check) and
`auto-claim-snapshot.ts` (1: `isRunnableAutoClaimCandidate`, a **pure
sync** predicate that needs the injected-lane pattern from #2551, not a
resolve). `notification-service.ts` has one more with a different
semantic — *"has progressed past"* — which needs its own thinking rather
than a mechanical swap. I will take these next unless someone claims
them.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:54:35 -07:00
gsxdsm
969c2cdf1d capacity part 4: drop the central global_concurrency table (migration 0037) (#2555)
Final piece of the cross-project cap removal. Enforcement (#2509),
settings/API/UI (#2529) are merged; this removes the storage.

Nothing read the table. `global_max_concurrent` held the deleted
machine-wide cap; `currently_active`/`queued_count` were written only by
`acquireGlobalSlot`/`releaseGlobalSlot`, measured earlier in this
program to have **no production caller**, so those counters were
fiction. Live “N running (all projects)” telemetry comes from
`CentralCore.getLiveRunningAgentCounts` and is unaffected.

Dropped rather than left unread: a lingering table with
plausible-looking counters invites a future reader to trust it — the
same trap as a readable-but-ignored settings key.

## The trap this hit, because the first attempt looked correct

`schema-applier.ts` warns that *“migrations are registered here
explicitly (not auto-discovered from the migrations dir), so a new .sql
file that is not wired through a version constant + bookkeeping check
silently never runs.”*

My first pass added the `.sql`, updated the drizzle model and bumped the
baseline — **and the table was still present in a fresh database**. It
was caught only because the test asserts the table is *gone*
(`to_regclass(...) IS NULL`) rather than merely unreferenced; an
absence-of-reference assertion would have passed while the table
survived.

Now registered properly: `DROP_GLOBAL_CONCURRENCY_VERSION = "0037"`,
explicit path constant, applied-check, bookkeeping insert.

The historical `0000` baseline is deliberately **not** rewritten — a
fresh database CREATEs the table then drops it, converging with upgraded
databases without editing history, which is how every prior migration
here behaves.

Also removed: the drizzle model, the `centralTableNames` entry, and the
`replacesCentralSeed` special case in the SQLite migrator (a legacy
SQLite `globalConcurrency` table now has no destination and is simply
not migrated — correct, since its cap is deleted and its counters were
never written).

## Verification

`pnpm lint` clean · core `tsc` clean · `pnpm test:gate` green (414 + 10
+ 71) · `schema-applier` 75/75 · `sqlite-migrator` 43/43 · full core PG
suite **1044 passed / 3 failed** — the same 3 pre-existing
(`central-archive-secrets` log-prefix,
`workflow-settings-project-identity` legacy fallback ×2).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:15:42 -07:00
gsxdsm
41031dbe2c Drift review (unowned): auto-claim candidacy resolves hold + completion roles — three literals, two opposite failures (#2565)
> **Based on `main`** — independent of my U7 stack and of #2561; merges
in any order.

Third unowned drift-review site. `isRunnableAutoClaimCandidate` is the
single source of truth for *"may an agent claim this task?"* (FN-6873),
and it carried **three** lifecycle literals that fail in **opposite
directions**.

## The two failures

**`column === "todo"` gated candidacy** on the hold role. Keyed on the
literal, a renamed workflow's candidate set was **permanently empty** —
agents were never offered its work, and nothing anywhere reported it.
Silence, not an error.

**`dependency?.column === "done" || "archived"` gated dependency
satisfaction**, and this is the more dangerous half: a dependency that
finished in a renamed **complete** column was never recognised as done,
so the dependent stayed **blocked forever**.

One makes work invisible; the other makes it permanently ineligible.
Both are silent.

## Roles resolve per task, not per pass

The non-obvious part: **a dependency may sit on a different workflow
from the claimant.** A single per-pass answer is wrong for one of them
on any mixed board — so the map is keyed by task id, and the dependency
check reads the *dependency's* roles, not the claimant's.

Asserted directly: a dependency completed in `done` (default vocabulary)
satisfying a claimant waiting in `drafting` (renamed).

## Shape

Both callers already have the store and are async, so they resolve for
real rather than taking the injected-lane fallback the *synchronous*
predicates needed (#2551). The predicate itself stays synchronous — a
resolved-roles map is passed in — because it runs inside two
`filter`/`flatMap` bodies.

Tasks absent from the map keep the legacy ids, so a partially-resolvable
board degrades to today's behavior instead of silently emptying the
candidate set.

**Type narrowing preserved.** The two callers take `Pick<TaskStore,
"listTasks">`, which is what makes them testable without a real store.
Rather than widening to the whole `TaskStore`, they now take
`Pick<TaskStore, "listTasks"> & WorkflowIrResolverStore` — the minimal
additional shape resolution needs.

## Revert proofs, isolated per literal

| Restored | Result |
|---|---|
| hold literal only | **3 of 6 fail** |
| dependency-completion literals only | **1 of 6 fails** |

The three default-vocabulary cases pass under both. Splitting the proof
matters here: it confirms the two halves are **independently**
load-bearing rather than one masking the other — a single combined
revert would have shown 3 failures and told me nothing about the
dependency half.

## Convergence

Measured on `main`, comment-stripped scan of `column === / !== "todo" |
"triage"` in `packages/*/src` excluding tests:

- this file alone: **103 → 102**
- with #2561: **103 → 100**

The `done` / `archived` literals fixed here sit outside that pattern and
are not counted — same caveat as #2561's gridlock `active` filter. Two
PRs now where the real fix is larger than the metric shows.

## Verification

| Check | Result |
|---|---|
| new suite | 6/6 |
| pre-existing auto-claim suite | 17/17, **no expectation edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (414 + 10 + 71) |
| `pnpm check:changesets` | clean |

## Remaining unowned in my area

`mission-feature-sync.ts` (1, a planning-lane check) and
`notification-service.ts` (1, *"has progressed past"* — a different
semantic needing its own thinking, not a mechanical swap). Taking those
next unless claimed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 10:01:14 -07:00
gsxdsm
8578a1d27d U8 PR5: thread the implementation exit to the step seam, and declare the stepwise pending-review park (inert) (#2546)
Follows **#2519** (U8 PR4). Both halves are inert — **no behavior
change** — and this removes the blocker PR4 documented.

## What was blocking

PR4 could only land its IR half because the pending-review ending could
not reach a graph edge on the **default** workflow. Three links in the
chain:

| Link | Problem |
|---|---|
| `runGraphTaskStep` | awaited the memoized implementation pass and
**discarded** its result |
| `RunTaskStepResult` / `RunSingleStep` | had nowhere to carry an exit |
| `stepExecute` seam | flattened every ending to `step-done` /
`step-failed` |

All three are fixed. The outcome stays `failure` (the step genuinely did
not complete) while the **value** now names the ending — which is what
`runForeach` propagates upward, since it returns a failing instance's
value as the foreach node's own. Every other ending keeps `step-failed`
byte-identically.

One design note: the exit is a property of the **pass**, not of a step.
A single memoized pass serves every foreach instance, so all instances
report the same ending — correct, because the ending is what stopped the
whole session.

With the value surviving, the stepwise IR declares the same
`review-handoff` park node and `steps --outcome:review-pending-->
review-pending-handoff --success--> end` edge the plain-`execute` shape
got in PR4, inherited by the final-review and Ideas variants that clone
it.

## A bug my own threading introduced, and what caught it

The first threading commit covered **one of the two** paths out of
`runProjectedGraphTaskStep`. The early-return branch carried the exit;
the main path goes through `runTaskStep` in `step-runner.ts`, which
builds its own result and dropped it — i.e. it worked on the path I
happened to read, and not on the path the default workflow actually
takes.

**FN-5436's regression test caught it, not code review.** That is the
second time this test has stood between this unit and a silent
regression, which is worth recording somewhere durable:
`executor-step-session.test.ts > FN-5436: pending-review skip on
no-fn_task_done exit` is the load-bearing test for this area.

## Why the seam flip is still not here

With the threading complete I applied the behavior half again — flip the
execute seam to return `review-pending`, delete the inline
`handoffTaskToReview`, add a named compat classifier for user-authored
graphs. **FN-5436 still failed**: the card did not reach `in-review`, so
something between the seam value and the park node is not routing under
that harness. I have not isolated whether that is the mock store's IR
resolution (it exposes no `getWorkflowDefinition`, so the run resolves
the built-in through a different path), a foreach aggregation detail, or
the park node's own seam.

I stopped rather than keep guessing, and reverted the behavior edits so
this lands green and inert. Shipping a half-routed move is exactly the
failure this unit exists to remove — a lifecycle transition that
silently does not happen. The alternative on offer was to relax
FN-5436's assertion, which would have been appeasing a test that is
telling the truth.

### What the instrumentation showed (done after opening this PR)

I ran the bounded next step rather than leaving it as a note. Two facts,
both measured:

1. **The IR is correct.** Resolving
`BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR` at runtime shows the
node and the edge survive the final-review variant's edge rewiring:

```
EDGES [{"from":"steps","to":"browser-verification","condition":"success"},
       {"from":"steps","to":"review-pending-handoff","condition":"outcome:review-pending"},
       {"from":"steps","to":"end","condition":"failure"}]
HAS NODE true
```

That matters because the variant does `template.edges = [ ... ]` (a
wholesale replacement) and filters outer edges touching `review` —
`review-pending-handoff` is not `review`, so it survives. Worth knowing
before anyone adds another node near it.

2. **The `stepExecute` seam is never invoked in that harness**, even
though the run terminates at `steps#0:step-execute` and the
implementation session demonstrably runs (`"Agent finished without
calling fn_task_done but Step 0 is blocked on pending review"` is in the
task log). A `console.log` at the seam's value computation produced no
output. So the exit is threaded correctly and the IR can route it, but
under this harness the value never originates.

3. **Nor is `createPromptLikeHandler`'s returned handler.**
Instrumenting its dispatch (`node.id` + resolved seam) produced nothing
either — so the node is not reaching the prompt-like path at all.

**Control experiment, because a negative result from instrumentation is
worthless until you prove the instrumentation is observable.** A
`process.stderr.write` at module load of the same file appears exactly
once in the same run, so writes from that module *are* captured under
this harness and the two negatives above are real, not artifacts of
swallowed output.

That narrows the remaining work to one question — what actually drives
`steps#0:step-execute` in this run, if neither the prompt-like handler
nor the `stepExecute` seam does — and rules out the IR, the foreach
propagation, the threading, and the instrumentation as suspects.

**Next step, now much narrower:** find the handler registration this run
resolves for a foreach instance node (the graph executor's handler map,
not the seam table), then flip the seam, delete the inline handoff, and
update the three ratchets that will correctly fire — PR3's routing pin,
the out-of-band adjacency check, and PR1's ownership ledger
(`runImplementation` 3 → 2; `handleGraphFailure` 0 → 1 for custom graphs
only).

## Verification

- `executor-step-session` + exit-events + ownership ledger +
graph-boundary — **56 tests green**
- `builtin-workflows` + `builtin-coding-workflow-ir` — green. The
layout-completeness contract required a layout entry for the new node in
all four stepwise-derived workflows; placed off the main line, because a
park is an exit and not a stage.
- `pnpm test:gate` green (10 / 309 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:54:39 -07:00
Phil Larson
01a75f9edc fix(desktop): typecheck streamed model downloads (#2493)
## Summary
- make the fetch response-body cast explicit across DOM and desktop
TypeScript library definitions
- preserve the existing async byte-stream runtime behavior

## Test plan
- `pnpm --filter @fusion/dashboard exec vitest run
src/stt/__tests__/model-manager.test.ts --reporter=dot`
- `pnpm --filter @fusion/desktop typecheck`
- `pnpm --filter @fusion/dashboard typecheck`
- `node scripts/check-changeset-format.mjs`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved desktop build typechecking for streamed speech-model
downloads by refining how streamed response bodies are interpreted for
TypeScript.
* Preserved runtime behavior, including streaming, integrity/hash
checking, file writing, and cancellation handling.
* **Maintenance**
* Updated the release metadata so this fix is published with the correct
patch classification.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 09:15:28 -07:00
gsxdsm
6721bdc652 U12 part 7: the List view never self-healed a card's workflow — extract Board's FN-7591 refetch and wire it up (#2530)
## U12 part 7 — the List view never self-healed a card's workflow

**Stacks on #2528.** Merge that first.

Paying off something I owed on #2525: greptile pointed out that a task
whose `taskWorkflowIds` entry is absent — or present but resolving to a
workflow that does not declare the task's stored column — gets no
per-workflow move metadata, so its menu falls back to the neighbour
approximation and **stays there until some unrelated refresh happens**.

Board has forced one board-workflows refetch for exactly this since
FN-7591. List had none. So the degraded state persisted longest
precisely where it is most likely: a **just-created card**, which is
when a workflow was actually chosen.

I said there that porting the self-heal deserved its own change rather
than riding along in a move-menu fix. This is it.

### Two commits, deliberately separable

**1. Extraction — move only.** Board's ~55 lines (refs, suspect-mapping
predicate, signature guard, deferred macrotask) become
`useUnmappedWorkflowRefetch`. Copying them into ListView would have
created a second copy of subtle race-avoidance logic to keep in sync.

Evidence it is a move: with comments and the new wrapper signature
stripped, the hook's **41 body lines** and the **42 removed from Board**
differ by exactly one line — the `}` that closed Board's enclosing
scope. Nothing added, removed or reordered. The original FNXC notes
travel with the code, since they are the reason each line exists.
Board's suite is green with no expectation edits.

**2. Wiring — behaviour change.** ListView calls the hook.

### Revert-proof

Remove the hook call from ListView and the new case fails:
`fetchBoardWorkflows` is never called a second time, so the mapping
never resolves. A companion case pins the other half — a fully-mapped
board must **not** refetch, so the signature guard cannot turn a healthy
list into a loop. It measures calls made *after* the initial load
settles, because mount fetch and switcher-open legitimately call the
fetcher and counting from zero would measure those instead.

### Two existing tests needed fixture corrections — neither a regression

Both because the self-heal now fires **correctly** where the fixture did
not expect a fetch:

- `refreshes workflow columns when workflow metadata SSE arrives`
chained two `mockResolvedValueOnce` payloads. The file-level cache seed
maps no tasks, so first paint saw FN-001 as unmapped and the repair
fetch ate the payload the test asserts on. Seeded that test's own
first-paint cache, and added a trailing default — the SSE swap
(`backlog` → `ready`) leaves FN-001 in a column its workflow no longer
declares, so a repair fetch there is right, and without a fallback it
resolved `undefined` and wiped the payload.

Worth stating plainly: both fixtures had quietly depended on List
*never* self-healing. That dependency is what the change removes.

### Verification

`pnpm test:gate` (309 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (18
steps), dashboard typecheck green. ListView + Board suites: **320
passed, 0 failed, 0 skipped**.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* List and Board views now self-recover when task-to-workflow mappings
are missing or incorrect, avoiding degraded workflow UI until a later
refresh.
* Workflow recovery retries are more robust and coordinated to handle
delayed/failed refreshes.
* Recovery behavior correctly stops/reset when switching projects or
unmounting.

* **Tests**
* Added comprehensive ListView coverage for unmapped-workflow self-heal,
including retry timing, StrictMode effect replay, SSE refresh
interactions, and mapped-vs-unmapped scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:04:58 -07:00
gsxdsm
7fd1c7f124 P0 fix: stop reaping worktrees out from under live planners (FN-6756) (#2531)
User-reported: worktrees deleted while a planning agent was still
working in them. Small, isolated, ahead of all remaining capacity work.

## Mechanism

`clearPhantomExecutorBinding` is documented as *"the last line of
defense against pulling a worktree out from under a running agent"*. It
computed liveness from four sets — `activeSessions`,
`activeStepExecutors`, `activeWorkflowStepSessions`,
`activeCliTaskSessions` — **all TaskExecutor-owned**. A triage PLANNING
session is owned by `TriageProcessor`, lives in *its own*
`activeSessions` map, and registers in the module-level
`activeSessionRegistry`. It matched none of the four.

Worse: the method **writes** to that registry (unregistering the task’s
paths) but never **read** it as a liveness signal. It destroyed the very
evidence that proved the planner alive.

Under plan-in-place a card is specified while it sits in
`todo`/`triage`, and `reapLeakedConcurrencySlots` treats both as
reapable on a rationale written *before* planning moved there (“a task
waiting to run must not pin a worktree”). Every gate ahead of the last
one passes for a planner:

| Gate | Saves a planner? |
|---|---|
| in `listWorktreeHolders()`? | **No** — `ensureTaskWorktreeForPlanning`
→ `ensureGraphCustomNodeWorktree` → `addActiveWorktree`
(`executor.ts:8581`) |
| reapable column? | **No** — plan-in-place keeps the card in
`todo`/`triage` |
| in the executor’s `executing` set? | **No** — a planner is
triage-owned |
| 60 s `LEAKED_WORKTREE_SLOT_GRACE_MS` | **No** — keyed on
`columnMovedAt`, and planning routinely runs for minutes |

So the broken guard decided alone.

## This is FN-8600 recurring through a second sweep

That fix registered planning paths in the registry and taught the
**self-owned-branch reclaim** sweep to consult `isPathActive`. The
leaked-slot reaper never got the same signal — fixed at one surface, not
enumerated across all. Exactly what the AGENTS.md Surface Enumeration
rule exists to prevent.

## Fix

The refusal now also fires when
`activeSessionRegistry.pathsForTask(taskId)` is non-empty. Keyed on
**any** registered path rather than on kind: the point is that a
registered surface of any kind means someone is working in that
worktree.

## Enumeration — the part that stops a third recurrence

The guard is a **chokepoint**, so this covers every caller rather than
just the reported one:

- `reapLeakedConcurrencySlots` — the reported path
- `recoverPausedAbortFailures` — **had the identical executor-only
pre-gate**
- the `preserveWorktrees: true` reclaim

Audited the rest of self-healing’s liveness gates: the self-owned-branch
reclaim, worktree-metadata reconcile and PR-branch sweeps already
consult `isPathActive`/`lookupByPath`. The three that read only
`getExecutingTaskIds` — `checkStuckBudget`, `recoverCompletedTasks`,
`recoverStrandedCompletedTodoTasks` — move columns and never destroy a
worktree, so they are noted rather than changed.

## Trade-off, stated plainly

A leaked registry entry now blocks this sweep instead of a live planner
losing its worktree. That is the strictly safer failure and the one the
“last line of defense” wording already promises. The registry is
process-local and in-memory, so a leak cannot outlive the process, and
stale entries have their own reconciler. **A test pins that a genuine
phantom — no executor surface AND no registration — still clears**, so
this is not a blanket refusal that would trade this bug for a wedged
queue.

**The 60 s grace is deliberately unchanged.** Raising it would only make
the bug rarer and harder to reproduce; the liveness gate was the defect.

## Verification

Revert-proof, measured: removing the registry term turns **3 of the 4**
new tests red, including the end-to-end sweep case (card in `triage`,
past the grace, executor sets empty → asserts the slot is not reaped and
the worktree survives). The 4th stays green both ways *by design* — it
is the anti-overcorrection guard.

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (309 +
10 + 71) · new suite 4/4. The 2 failures in `self-healing.test.ts` /
`-completion-fanout.test.ts` are **pre-existing** — identical with this
change stashed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Prevented active planning worktrees from being mistakenly deleted or
reclaimed while related planning sessions are still active.
* Enhanced session liveness checks so phantom executor bindings are not
cleared when a live session is registered.
* Updated paused abort recovery to defer or abort safely when a live
planning session is detected, avoiding unintended task/worktree
mutations.
* **Tests**
* Added regression coverage for leaked-slot reaping, paused abort
recovery behavior, phantom binding refusal, and end-to-end sweep
outcomes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:04:51 -07:00
Phil Larson
b85a5d4531 fix(core): bound compound engineering review remediation (#2532)
## Summary

- cap Compound Engineering Code Review remediation at two Execute→Review
repair passes
- enable no-progress detection for the built-in CE workflow
- preserve explicit project/workflow overrides while making the authored
CE default visible in settings and docs
- update stale IR/changeset language that still described Code Review as
unbounded when unset

## Why

The previous CE default was effectively unbounded. A reviewer that
repeatedly returned `REVISE` could consume thousands of remediation
cycles without terminally parking the task. The built-in workflow should
fail closed after a small, explicit budget while still allowing
operators to author a different numeric cap.

## Verification

- `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/core
exec vitest run src/__tests__/builtin-workflows.test.ts` — 46 passed, 17
skipped
- `corepack pnpm@10.33.0 --filter @fusion/core typecheck`
- `corepack pnpm@10.33.0 --filter @fusion/dashboard exec vitest run
app/components/__tests__/WorkflowSettingsPanel.test.tsx
app/components/__tests__/workflow-setting-display.test.ts` — 33 passed
- `corepack pnpm@10.33.0 --filter @fusion/dashboard typecheck`
- `corepack pnpm@10.33.0 changeset status --since=origin/main`
- `git diff --check origin/main...HEAD`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Improvements**
- Compound Engineering Code Review now caps remediation attempts at 2;
after two unsuccessful attempts, the process parks instead of retrying
indefinitely.
- Post-restart review recovery now completes in a single maintenance
cycle to reduce delays.
  - Default post-review fix budget increased from 3 to 10.
- Review revision limits now consistently honor workflow-authored
defaults when settings are left empty, and `0` disables automatic
remediation.

- **Documentation**
- Updated the workflow editor, settings reference, workflow steps, and
operator panel text to clarify cap/default/disable semantics (including
CE: 2).

- **Tests**
- Added/updated unit tests to validate the new bounded remediation
behavior and messaging.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:05:29 -07:00
Phil Larson
72391c90b2 fix(engine): route workflow reviews through validator models (#2533)
## Summary

- classify review-type workflow steps with the existing review-step
classifier
- resolve their primary, fallback, and thinking-level settings from the
validator model lane
- retain per-step model overrides and executor-purpose workflow-step
tooling
- keep ordinary workflow steps on the execution lane
- make missing-fallback diagnostics identify the correct lane

## Why

Code Review, Plan Review, verification, and inline-review gates were
executed through the implementation model lane merely because they run
inside `executeWorkflowStep()`. That defeats configured reviewer-model
separation and can make the same model implement and validate its own
work.

This changes model selection—not the workflow-step session/tooling
contract—so review steps remain executor-purpose sessions while using
validator lane models.

## Verification

- `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/engine
exec vitest run src/__tests__/executor-workflow-step-model.test.ts` — 14
passed
- `corepack pnpm@10.33.0 --filter @fusion/engine typecheck`
- `corepack pnpm@10.33.0 changeset status --since=origin/main`
- `git diff --check origin/main...HEAD`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Review-type workflow steps now route through the configured validator
model lane (instead of the execution lane).
* Validator primary/fallback and thinking-level settings are applied
correctly for review steps.
  * Step/task overrides still take priority over lane-based resolution.
* Fallback retry sessions now use the appropriate validator/executor
configuration, with lane-specific fallback guidance when fallback
settings are missing.
* **Tests**
* Expanded executor workflow-step model resolution and routing/fallback
precedence assertions for validator-lane behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:05:04 -07:00
Phil Larson
9a8fc409ff fix: persist manual task pauses (#2536)
## Summary

- persist an explicit `userPaused` latch when operators pause tasks
through CLI, MCP, dashboard task routes, or mission stop
- keep automatic/internal pauses distinct (`userPaused` remains false
unless explicitly requested)
- clear the latch on unpause
- route the flag through in-memory and PostgreSQL task stores
- add contract coverage across core, CLI, MCP, dashboard task routes,
and mission stop

## Why

A manually paused task could lose the reason for its pause across
dashboard/runtime restart. Startup recovery then treated it like an
internally interrupted task and reclaimed it, restarting automation
against the operator’s intent. Manual pauses must survive restart and
remain non-runnable until explicitly unpaused.

## Verification

- core pause durability tests: 2 passed
- CLI task/extension tests: 150 passed; PostgreSQL integration lane
remains active in CI
- dashboard route tests: 261 passed
- `@fusion/core`, `@runfusion/fusion`, and `@fusion/dashboard`
typechecks passed
- full workspace build passed with pnpm 10.33.0
- changeset validation and `git diff --check` passed
- live aggregate runtime verification also confirmed
`paused=true,userPaused=true` survived a normal dashboard restart with
zero active tasks


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Manual task pauses now persist across application restarts and
recovery.
- Pauses initiated via the CLI, dashboard, MCP tools, and mission stop
controls are recorded as explicit user actions.
  - Automatically paused tasks remain eligible for recovery.
  - Unpausing clears the durable manual-pause state.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:04:28 -07:00
gsxdsm
3ff98aae56 U12 part 8: delete the lossy normalizeColumn + behaviour ratchet — and the definitive answer on the raw flag (2 reads left, both U2b's) (#2535)
## U12 part 8 — deletes the lossy `normalizeColumn`, and ratchets it
shut

Independent of the #2525 → #2528 → #2530 stack; touches only
`@fusion/core` exports.

This closes **one of the two `@deprecated (workflowColumns, U12)`
markers** the unit was named for.

### The hazard

`normalizeColumn` coerced an arbitrary value to a **legacy** column,
rewriting every workflow-defined custom id to `triage`. Silent data loss
for any project whose workflow declares a column outside the six
built-ins — and it sat one line away from `normalizeColumnId`, which
sanitises structurally and passes real ids through.

The dashboard picked the wrong one for its entire task-ingest path until
that was diagnosed; `useTasks.ts` and `routes-trait-rekey.test.ts` still
carry the notes from that fix. So this is not a hypothetical footgun —
it already fired once, on the surface where it mattered most.

Deleted rather than left deprecated because it has **zero callers
anywhere in the workspace**. It was pure exported hazard: a lossy
coercion next to its safe twin, waiting to be picked again.

### The ratchet is the point

`no-lossy-column-coercion-export.test.ts` bans the **behaviour, not the
identifier**: it walks every exported single-argument function whose
name mentions "column" and fails if one maps a valid custom id onto a
different legacy id. Re-adding `normalizeColumn` under any name trips
it.

Verified by actually reintroducing the function — **two of the three
cases fail, including the name-agnostic one**. That last detail is what
stops it being a guard that checks nothing.

Coverage stated plainly: deleting an unused export has no behaviour to
revert-check. The compile is the proof it had no callers; the ratchet is
the proof it cannot return.

---

## Answering the standing question: does anything still read the raw
`workflowColumns` flag?

**Yes. Exactly two sites, and both are U2b's.** I am not able to close
this out, and here is the complete list rather than a summary:

```
packages/core/src/store.ts:38,43                                  ← the definition
packages/core/src/task-store/moves.ts:9,363                       ← `useWorkflow`
packages/core/src/task-store/workflow-task-create-ops.ts:11,351   ← move-policy preflight
```

That is the whole list in production code. Everything else that greps is
a comment, a test that writes the flag deliberately to exercise the dead
path, or the unrelated `workflowColumns.*` i18n namespace for the
Columns editor panel.

**Why I have not deleted the settings key.** It cannot go while those
two read it — the key is what they read. And the two are not separable
from each other: `workflow-task-create-ops.ts:351` computes the
`movePolicyPreflight` that `moves.ts` consumes and validates, and
un-gating the preflight alone would start evaluating workflow move
policies (with their plugin-gate side effects) while the branch that
consumes the result stays off. That is a behaviour change with no
consumer, which is worse than either state.

**Status of the blocker.** U2b has not landed. `main` at `919f68f9b`
still has both reads; the program's merged history goes `#2466 → #2467 →
#2468 (characterisation only) → #2469 → #2479 → #2500 → #2512 → #2513`,
with no convergence PR. PR #2468 was Phase A2 **steps 1–2 only** — the
differential characterisation — and the convergence that deletes one of
the two move paths was never merged.

So the honest state of the unit: everything U12 owns is done except the
two reads that U2b owns, and the settings key that cannot be deleted
until they are gone. If you want me to take U2b itself, say so — I have
the inventory and the divergence list, and I would want the current U2b
worker stood down from `moves.ts` first.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:25 -07:00
gsxdsm
18d654a5ff capacity, part 3: delete the globalMaxConcurrent setting, API and UI (#2529)
Part 3 of the capacity simplification, and the half that removes the
**knob**. Enforcement (shared semaphore, runtime wiring) went in #2509;
this removes everything an operator or API client can still see, so
nothing is left readable-but-ignored.

## Deleted

Settings key + schema default · CentralCore’s
`getGlobalConcurrencyState` / `updateGlobalConcurrency` /
`acquireGlobalSlot` / `releaseGlobalSlot` and the `concurrency:changed`
event · the whole Global Concurrency block in `async-central-core` ·
`PUT /api/global-concurrency` · the Scheduling · Global settings section
· the footer and Command Center global sliders · the dead
`getGlobalConcurrencyLimit` reader whose only caller went in #2509.

## Kept, deliberately

**`GET /api/global-concurrency` survives as telemetry only** — live
`currentlyActive` / `projectsActive` from CentralCore’s side-effect-safe
source. “How busy is this machine?” is still a real question once the
cap that used to answer it is gone. It no longer reports
`globalMaxConcurrent`/`queuedCount`: those came from the deleted cap and
from slot bookkeeping production code never incremented, so publishing
them was publishing zeros dressed as state.

**`useGlobalConcurrency` becomes read-only.** Everything that existed to
*persist* went with the cap — the 500 ms debounce, the save-state
machine, the commit-on-close/unmount flush, the slider clamp, the
`interactive` gate. The module-level shared store is **kept**: its
original justification (two mounted consumers drift apart with private
copies) holds for a polled read exactly as it did for a cap, and one
fetch now serves both.

The live “N running (all projects)” readout survives in both surfaces,
moved onto the per-project row.

## Two sections become one

Scheduling · Global existed to host exactly one control. With it deleted
the section renders an empty pane, so the Global/Project pair merges
back into **“Scheduling”**. An empty nav entry is a promise of settings
that are not there.

## One real fix found on the way

`SchedulingSection`’s `concurrencyLoading` gated the **project**
concurrency inputs on the **global**-concurrency fetch — never the right
source, since `maxConcurrent` and `maxWorktrees` come from the settings
form. It is repointed at the form’s own load, preserving the invariant
it existed for: a concurrency input stays disabled until its live value
arrives, so an operator cannot overwrite a resolved limit with a blank
fallback.

## Migration

A stored `globalMaxConcurrent` is **ignored** — it is a project-blob key
nothing reads, so dropping it needs no schema change. The
`central.global_concurrency` **table** is dropped in a follow-up; this
slice stops seeding and reading it first, so that drop has no live
writer to race.

## Verification, and how the wider suite was controlled

`pnpm lint` clean · core/engine/dashboard `tsc` clean · `pnpm test:gate`
green (309 + 10 + 71) · dashboard settings/footer/command-center/hooks
**2237/2237** · core `central-core-backend` 9/9.

The broader dashboard suite shows failures, and I checked rather than
assumed: running the suspect files on **clean main** reproduces
`api-git` (49), `TaskDetailModal.rendering` (28) and `settings-mobile`
(17) identically. Two were genuinely mine —
`SettingsModal.scheduling-merge` (0 on main, 17 on this branch: my nav
rename) and one `settings-mobile` picker case asserting `scheduling` is
a scoped pair — and both are fixed.

Tests for deleted behaviour are removed with it (footer
confirm/cancel/flush/dedupe, global marker geometry, the hook’s PUT
case, the CentralCore slot cases), each carrying a note on what it
guarded and where the surviving **project-side** equivalent lives.
Fixture-only references were updated, not deleted.

Nothing booted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:13 -07:00
gsxdsm
7003dc9803 U12 part 6: Board re-rendered every column on every state change — one inline arrow, measured with a memo-comparator probe (#2528)
## U12 part 6 — Board re-rendered every column on every state change

**Stacks on #2525.** Merge that first.

`canDropTask` was allocated as a fresh inline arrow, per column, per
render:

```tsx
canDropTask={(taskId) => canDropTask(taskId, columnDef.id, selectedWorkflow.id)}
```

`Column` is `React.memo`, and a new function identity on any prop
defeats that entirely. So **any** Board state change — collapsing
Archived, changing Done sort, opening the workflow switcher —
re-rendered every column and every card beneath it, not just the
affected one.

Bound through a `useMemo` cache keyed by lane + column. After the fix,
collapsing Archived re-renders exactly one column: `archived`.

### Measured, not guessed

I instrumented `React.memo`'s comparator to print which props actually
change identity on a collapse toggle. For every unaffected column the
answer was exactly one:

```
PROBE todo         changed: canDropTask
PROBE in-progress  changed: canDropTask
PROBE in-review    changed: canDropTask
PROBE done         changed: canDropTask
PROBE archived     changed: canDropTask,collapsed     <- the one that should re-render
```

After:

```
PROBE archived     changed: collapsed
```

### Why this hid, and why my first attempt failed

Two things worth recording, because both were mistakes I made in this
program:

**The test was pointed at dead code.** "keeps unaffected columns stable"
measured the **legacy single-lane board**, whose props were all stable —
so it passed for a long time while covering nothing operators use.
Deleting that board in part 1 repointed it at the real board, where it
failed 3-vs-2. I skipped it then rather than weaken it to the observed
number, and said it needed its own investigation. This is that
investigation.

**My first fix was wrong and I was right to revert it.** In part 1 I
tried a `useRef` cache invalidated by `useEffect`, it did not fix the
test, and I reverted it as unproven rather than ship it. The reason is
now clear: the effect runs *after* the render that populated the cache,
so it wipes the very bindings that render created and the next render
allocates fresh ones — the invalidation defeated the cache. `useMemo`
keyed on the resolver has no such window; the map lives exactly as long
as the closure owning it.

### Revert-proof

The test is un-skipped **with the fix, not with a new expected number**.
Restore the inline arrow at either call site and it fails 3-vs-2 again.

### Verification

`pnpm test:gate` (309 + 10 + 71), `pnpm lint`, dashboard typecheck
green. Board, Board.canDropTask, workflow-resolved-columns and
board-no-legacy-flash: 132 passed, 0 failed, **0 skipped** — the skip
introduced in part 1 is gone.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Move menus now show exactly the destinations permitted by each custom
workflow, including non-adjacent moves.
- Invalid or hidden destination columns are excluded from move options.
  - Older workflow data continues to use a compatible fallback behavior.

- **Performance**
- Improved board responsiveness by preventing unaffected columns and
cards from re-rendering when archived sections collapse or Done sorting
changes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:07 -07:00
gsxdsm
da0351857e U12 part 5: put real workflow adjacency on the wire — custom-workflow move menus were guessing (measured), and the VALID_TRANSITIONS shortcut is gone (#2525)
## U12 part 5 — the move menu was guessing; now it asks the graph

**Stacks on #2521** (same file). Merge that first.

The context menu had **no adjacency data at all**, so it did two wrong
things at once: it approximated move targets from a column's
**neighbours in declared order**, and — because that approximation is
strictly weaker than the real graph — it kept a `VALID_TRANSITIONS`
shortcut for any workflow whose column-id set matched the six built-ins.

Measured, the approximation loses real operator moves:

| current | workflow graph | neighbour approximation |
|---|---|---|
| `in-progress` | in-review, todo, triage, done | todo, in-review |
| `todo` | in-progress, triage, archived | triage, in-progress |
| `done` | todo, triage, archived | in-review, archived |

So **every custom workflow has been offering a guess**: menu entries the
store would reject, and legal moves it never offered. The built-ins were
fine only because the shortcut bypassed the guess entirely.

### The fix

`BoardWorkflowColumn` gains `moveTargets`, resolved by
`resolveAllowedColumns` — *the same resolver `moveTaskInternal`
validates against*. The menu now offers exactly what the store will
accept, for any workflow. Threaded through all four metadata builders
(Board, Lane, ListView, TaskDetailModal).

Optional on the wire, deliberately: a client older than this field keeps
the neighbour fallback rather than losing its move menu mid-upgrade.

### Why deleting the legacy shortcut is safe

Not an assertion — a measurement, then a pin.
`resolveAllowedColumns(BUILTIN_CODING_WORKFLOW_IR, c)` is **identical to
`VALID_TRANSITIONS[c]` for all six columns, order included**:

```
triage       ["todo","archived"]                     == VALID  SAME
todo         ["in-progress","triage","archived"]     == VALID  SAME
in-progress  ["in-review","todo","triage","done"]    == VALID  SAME
in-review    ["done","in-progress","todo","triage"]  == VALID  SAME
done         ["todo","triage","archived"]            == VALID  SAME
archived     ["done"]                                == VALID  SAME
```

`builtin-adjacency-matches-legacy-transitions.test.ts` pins it so the
equivalence cannot drift silently — if the built-in workflow's edges
change without `VALID_TRANSITIONS` following, default menus change shape
and that test fails first. It compares **order** too, since the menu
renders targets in the order it receives them, so a reorder is
operator-visible.

Default-workflow menus are therefore byte-identical. Custom ones stop
guessing.

### What's left of the legacy vocabulary here

`COLUMNS` is gone from `TaskContextMenu` — deleting the shortcut removed
its last use. `VALID_TRANSITIONS` survives for exactly one thing: the
**no-metadata load window**, documented at the site. I measured removing
that in #2521 and it left Task Detail with no move options during load,
which is a regression rather than a cleanup. It retires when the load
window does.

### Revert-proof, two ways

- Drop the `declaredTargets` branch → the custom-workflow case fails:
the neighbour fallback returns `["backlog","building"]`, missing the
legal `shipped` jump **and** offering `backlog`, which that graph
forbids. That is exactly the defect class shipped to every custom
workflow today.
- A second case pins that an adjacency edge into a column the board
cannot show is **dropped**, not rendered as a dead menu entry.

### Verification

`pnpm test:gate` (309 + 10 + 71), `pnpm lint`, `pnpm verify:fast`, core
+ dashboard typechecks green.

**No new test failures**: five suites report 31 failures with and
without the change — an identical, pre-existing set, verified by diffing
failing test *names* against a stashed clean tree, not by comparing
counts.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Move menus for custom workflows now show only the destinations
permitted by that workflow.
* Task-specific workflow rules are applied consistently across boards,
lists, lanes, and task details.
  * Invalid or unavailable destinations are excluded from move options.
* Existing clients remain supported when workflow destination data is
unavailable.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:00 -07:00
gsxdsm
ebc89310bc U12 part 4: derive the move menu's "Back to" label from workflow traits (plus two legacy reads I did NOT delete, with measurements) (#2521)
## U12 part 4 — the move menu's "Back to" label followed hardcoded
column ids

`getTaskMoveTransitions` is shared by Board cards, List rows and Task
Detail. It labelled a backwards move with:

```ts
column === "in-progress" && task.column === "in-review"
  ? t("taskDetail.move.backToInProgress", "Back to In Progress")
```

Two hardcoded lifecycle ids **and** a hardcoded English column name. On
a workflow that renames those lanes the condition never matched, so the
affordance silently vanished — and had it matched, it would have
announced "In Progress", a column absent from that board. Same
legacy-vocabulary class U10 removed from Board and U12 removed from
ListView, surviving in the context menu all three surfaces render.

Now keyed on the traits it was approximating: the **current** column
carries `mergeBlocker`, the **target** carries `countsTowardWip`, and
the label interpolates the column's own name through a new
`taskDetail.move.backTo` key (added to all six locales).

### Scope I deliberately held back

**The set of moves labelled "Back to" is unchanged.** For
`builtin:coding` the traits resolve to exactly `in-review` and
`in-progress`.

I first generalised this to "any target earlier in the workflow's
declared order" — arguably nicer, and I had it working. Then I measured
it: it relabels moves this change never set out to touch. **18 assertion
sites across three suites** flip from "Move to" to "Back to" (e.g. a
card in In progress gets "Back to Todo", "Back to Planning").
Same-set-different-derivation is the honest scope here; widening which
moves read as backwards is a separate, visible product decision, not a
side effect of a vocabulary fix.

### Two things I chose not to delete, and why

Both are still-live `VALID_TRANSITIONS` reads in this file. Neither is
removable today, and the reason is the same missing wire field —
documented at both sites rather than left as a puzzle.

**1. The default-column-set shortcut.** `TaskContextMenuColumnMetadata`
carries id/label/flags but **no adjacency**, so the workflow branch can
only guess targets from a column's neighbours in declared order.
Measured against the real graph that is a strict loss:

| current | `VALID_TRANSITIONS` | neighbour-derived |
|---|---|---|
| `in-progress` | in-review, todo, triage, done (4) | todo, in-review
(2) |
| `todo` | in-progress, triage, archived (3) | triage, in-progress (2) |
| `done` | todo, triage, archived (3) | in-review, archived (2) |

Deleting that read is not a cleanup — it drops real operator moves
(archive from Todo, straight-to-Done from In progress). Note the guard
keys on the column **id set**, so a workflow that merely renames the six
built-ins still takes this path and still gets correct targets; only
reordering or replacing them falls through to the weaker logic.

**2. The no-metadata fallback.** I removed it first, on principle, and
measured the result: `workflowMoveColumns` is optional at both call
sites (`workflowMoveMetadata?.moveColumns`, `taskMoveColumns`) and
genuinely undefined until board-workflows resolves, so dropping it left
Task Detail with **no move options during load**. That is a live surface
degraded to satisfy a purity rule, so it is not shipped. Unlike Board
and ListView — where the legacy path was provably unreachable — this one
is reachable and useful.

Both retire the same way: put each column's allowed targets on the
board-workflows payload so the load window has real data instead of a
guess. That is a server + wire + client change and belongs in its own
slice.

### Revert-proof

The renamed-workflow fixture declares `signoff` (mergeBlocker) and
`building` (countsTowardWip). Restore the id literals and the new case
fails — `Move to Building` instead of `Back to Building` — which no
relabelling of the old hardcoded string could satisfy, since that string
names a column absent from the board. The same case asserts the forward
move keeps "Move to Shipped", so the rule stays a distinction rather
than a blanket relabel.

### Verification

`pnpm test:gate` (309 + 10 + 71), `pnpm lint`, `pnpm verify:fast`,
dashboard typecheck green.

**No new test failures**, established properly: the three suites this
touches report 30 failures both with and without the change, and I
diffed the failing test *names* against a stashed clean tree rather than
comparing counts — the sets are identical. (An earlier count-only
comparison had me chasing two failures that turned out to be my own new
assertions.)

Also regenerates `packages/i18n/src/resources.d.ts` via `pnpm
i18n:types`. That picks up **~45 lines of pre-existing drift** from
earlier merges that did not regenerate it; the file is generated, and
leaving it stale would omit the new key from the types. Flagged so the
extra lines are not mistaken for scope creep. Note
`packages/dashboard/app/locales/` is gitignored (copied from
`packages/i18n/locales/`), so only the canonical locales are committed.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:19:42 -07:00
gsxdsm
5de083ef08 U8 PR4: declare the pending-review park as a graph node (inert) — and why the behavior move is blocked on the step-session chain (#2519)
Fourth PR of **U8 — the graph owns execution**. This is the IR half of
the pending-review routing move. **Inert: no behavior change.** The
behavior half is deliberately NOT in this PR, for a measured reason
below.

## What lands

A `review-handoff` seam node (`review-pending-handoff`, column
`in-review`) in `BUILTIN_CODING_WORKFLOW_IR`, with:

```
execute --outcome:review-pending--> review-pending-handoff --success--> end
```

An implementation session can end because a step is blocked on a pending
review: the agent cannot continue, and the card belongs in review rather
than in an error bucket (`status: failed` on an `in-review` row
deadlocks the merge queue). Today the **executor** performs that
transition inline, mid-session, and the graph finds out afterwards —
which is why `handleGraphFailure` carries `alreadyFinalizedToReview`, a
classifier whose only job is recognising a move the graph did not make.

Two design points worth recording, both verified against the interpreter
rather than assumed:

- **The edge goes to `end`, not to `review`.** Routing to the ordinary
`review` node would have continued the run into `merge-gate` and
`merge-attempt` on work whose steps are incomplete. "Hand off and stop"
is what the inline handoff does; the edge to `end` is what preserves it.
- **`outcome:` edges match on the node's VALUE and take priority over
generic `success`/`failure` edges** (`shouldTraverseEdge` /
`traverseChildren`). So this claims only the pending-review ending, and
a workflow that does not declare the edge falls through to its generic
`failure` edge — exactly today's behavior. That is what makes the
eventual move safe for user-authored graphs.

## Why the behavior half is not here — a measured finding

I implemented it, and backed it out. The record matters more than the
diff:

1. **`BUILTIN_CODING_WORKFLOW_IR` is not the default workflow.** It
backs `builtin:legacy-coding`; `builtin:coding` uses the
*stepwise-final-review* IR, which has no `execute` node — its
implementation runs as a `foreach` of `step-execute`.
2. **The foreach mechanism would work.** `runForeach` propagates a
failing instance's `value` up as the foreach node's own value, so a
`steps` node could carry an `outcome:review-pending` edge.
3. **But `stepExecute` flattens it first.** The seam returns `value:
result.outcome === "success" ? "step-done" : "step-failed"`, discarding
the exit before it can reach any edge.

So on the default workflow the exit cannot reach an edge, and a compat
classifier in `handleGraphFailure` keyed on the failure value cannot see
it either. **Removing the inline handoff therefore regressed the default
path**: the card stopped reaching `in-review` at all.
`executor-step-session.test.ts`'s FN-5436 case caught it —

```
FAIL  FN-5436: pending-review skip on no-fn_task_done exit
      > parks in-review when review request has no subsequent verdict
      expected "moveTask" to be called with [ 'FN-5436-B', 'in-review' ]
      Number of calls: 0
```

I could have made that green by relaxing the assertion. That would have
been appeasement of a test that was telling the truth, so the behavior
commit came out instead.

**Also caught, and worth noting as the ratchets earning their keep:**
the PR1 ownership ledger flagged the change as `runImplementation` 3 → 2
review handoffs and `handleGraphFailure` 0 → 1 — i.e. a *relocation*,
not an elimination, for every non-plain-`execute` shape. That number is
what turned "this move is good" into "this move is only good for one
workflow shape". And PR3's routing-unchanged pin plus its out-of-band
adjacency ratchet both fired, forcing the routing change to be declared
rather than slipping in.

## PR5

Thread the implementation exit through the step-session chain
(`runImplementationPhase` → `graphStepRunOnce` → `runGraphTaskStep` →
`runProjectedGraphTaskStep` → `stepExecute`) so the seam can return
`review-pending` instead of flattening to `step-failed`; add the node +
edge to the stepwise IRs; then flip the execute seam and delete the
inline handoff **in one correct step** for every built-in shape at once.
The compat path for user-authored graphs is then a single named
classifier rather than a call buried two thousand lines into a session
loop.

## Verification

- `builtin-coding-workflow-ir` + `builtin-workflows` — 76 tests green
(the layout-completeness contract required a layout entry for the new
node; it is placed off the main line because the park is an exit, not a
stage)
- `executor-step-session` + ownership ledger + exit events — 50 tests
green, unchanged
- `pnpm test:gate` green (309/10/71); `pnpm lint` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:19:28 -07:00
gsxdsm
063978c289 U12 part 3: make the v1-IR persistence unconditional — after this, every raw-flag read is on the move path (U2b) (#2513)
## U12 part 3 — every remaining raw-flag read is now on the move path

**Stacks on #2512** (shares a line in `workflow-ops.ts`). Merge that
first.

**Behaviour-preserving. Not a single persisted byte changes.**

### What changed

The three v1-IR rollback-compat persist sites (#1405) all read `flagOn ?
ir : downgradeIrToV1IfPure(ir)`, where `flagOn` came from the retired
raw `experimentalFeatures.workflowColumns` key. No production writer
sets it, so **every real project has always taken the downgrade arm**.
Removing the branch is a runtime no-op; it deletes three flag reads.

Sites: `createWorkflowDefinitionImpl`, `updateWorkflowDefinitionImpl`,
and `insertWorkflowDefinitionSyncImpl` — whose `flagOn` *parameter* is
gone too, along with the plumbing that resolved it in
`migrateLegacyWorkflowStepsImpl`.

With those gone, **`TaskStore.workflowColumnsFlagOn()` has no callers
and is deleted.** Its six readers were the three U5 guards (part 2) and
these three persist sites.

### The decision I made, and why I went the other way

I had this slice scoped as "retire the v1 downgrade." **I rejected
that.** It is a compatibility affordance, not cutover machinery: it
fires only for a graph exactly equivalent to pure v1 (default columns,
default placements, no v2-only features), and `upgradeV1ToV2` re-reads
it into an identical v2 graph, so the runtime never sees a difference.
Retiring it would break a binary downgrade for zero benefit — and stale
binaries opening these databases is an **observed event** in this
project, not a hypothetical.

So the slice became the strictly better version of itself: same three
flag reads removed, no compat surface touched.

### Why this matters for sequencing

`isWorkflowColumnsCompatibilityFlagEnabled` survives. It is still read
by `moves.ts:363` and by `workflow-task-create-ops.ts:351`'s move-policy
preflight that feeds it. Removing those reads **is** the U2b move-path
convergence with its equivalence-proof obligation.

The point of deleting the wrapper is that it makes the remainder
enumerable:

```
$ grep -rn isWorkflowColumnsCompatibilityFlagEnabled --include=*.ts packages/ | grep -v __tests__
packages/core/src/store.ts:38                      <- the definition
packages/core/src/task-store/moves.ts:9,363        <- U2b
packages/core/src/task-store/workflow-task-create-ops.ts:11,351  <- U2b (feeds moves.ts)
```

**Every surviving read is on the move path.** U2b deletes the definition
and the unit closes.

### On coverage — stated honestly

This change is behaviour-preserving, so it has **no revert-proof test**,
and I am not going to claim one. `flagOn ? ir : downgrade(ir)` with an
always-false flag *is* `downgrade(ir)`.

What needed a guard is the next edit someone is tempted to make —
deleting `downgradeIrToV1IfPure` as dead cutover machinery. New
`workflow-ir-v1-rollback-persistence.test.ts` fails if it is removed,
and pins the exact boundary: the built-in coding workflow (named columns
+ traits) stays v2; a pure-v1-equivalent graph stores as v1 without the
synthesized `columns`; a downgraded graph re-parses to an **identical**
runtime graph (the property that makes unconditional application safe);
a graph with a custom column stays v2.

### Verification

`pnpm test:gate` (307 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (17
steps), typecheck green. Core workflow-named suites: 383 passed, 1
failed — `workflow-ir-settings.test.ts > moved-key catalog ...`
(`expected 10 to strictly equal 3`), which I confirmed fails identically
on a stashed clean tree. Pre-existing, unrelated. No Fusion instance
booted.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved workflow persistence compatibility by consistently storing
pure v1-equivalent workflows in the compatible format.
* Preserved v2 workflows and custom column information when they are not
v1-equivalent.
* Retired obsolete feature-flag checks without changing stored workflow
or board behavior.

* **Tests**
* Added coverage for workflow version preservation, rollback-compatible
serialization, and custom columns.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 19:09:06 -07:00
gsxdsm
3badc244a7 U12 part 2: bind the three U5 reconciliation guards — USER-VISIBLE (and one path that couldn't run under PostgreSQL at all) (#2512)
## U12 part 2 — the three U5 reconciliation guards now actually fire

USER-VISIBLE. Taken on standing authority; here is exactly what changed
for operators.

All three read the RAW `experimentalFeatures.workflowColumns` key via
`store.workflowColumnsFlagOn()`. Nothing in production writes it, so all
three have been inert since the workflow-columns cutover.

| Guard | Before (every real project) | After |
|---|---|---|
| Workflow edit removing an **occupied** column | Save succeeded; cards
left in a column the workflow no longer declares | Save fails with
`OccupiedColumnsError` unless `rehomeTo` is supplied |
| Workflow **delete** | Occupant capture returned `[]`; cards sat in the
deleted workflow's columns until the next engine start | Cards move to
the default workflow's entry column as part of the delete |
| Workflow **switch** | Never reconciled; the `reconciliation` field in
the declared return type was never populated | Card in an undeclared
column moves to the resolved target; a declared column is preserved |

Both consumers already handle the new outcomes and needed no change:
`register-workflow-routes.ts` maps `OccupiedColumnsError` to a
structured 409 carrying per-column occupant counts, and
`fn_workflow_update` returns a retryable structured result. The
dashboard editor's `rehomeTo` retry flow becomes reachable for the first
time. I only updated two stale "flag-ON" comments there — that code was
correct all along and simply never fired.

### What an operator actually sees (USER-VISIBLE — read this bit)

Four changes to what the board and the API do. Nothing here is silent.

1. **Editing a workflow to remove a column that has cards in it now
FAILS.** Previously the save succeeded and the cards were left in a
column their workflow no longer declared. The dashboard shows the
existing 409 with per-column occupant counts and prompts for a re-home
target; retrying with `rehomeTo` moves the cards and saves. Removing an
EMPTY column is unaffected.
2. **Deleting a workflow moves its cards immediately** to the default
workflow's entry column, instead of leaving them until the next engine
start.
3. **Switching a task's workflow moves the card** when the new workflow
does not declare its current column. A card whose column IS declared
stays exactly where it is. The API response now carries the
`reconciliation` summary it always promised.
4. **A switch whose re-home would be REJECTED is now refused before
anything is written.** If the destination column is at its WIP limit,
the switch fails with a structured 409 (`workflow-switch-rehome-failed`)
naming the task, both columns and the reason — and **nothing changes**:
the task keeps its current workflow AND its current column. Retry after
making room. Previously this combination committed the selection and
then silently reported a move that never happened, leaving selection and
column disagreeing.

**Can a torn card still happen? Yes, in one narrow case, and here is how
you recover.** If the destination fills in the window between the
pre-flight and the move, the selection is already committed and the card
ends up in a column its new workflow does not declare. That case is not
silent: it writes a `task:workflow-switch-torn` run-audit row, and the
error carries `selectionCommitted: true` with both columns. Recovery:
make room in the destination and move the card there, or switch the task
back — and if neither happens, the R7 startup sweep
`reconcileUndeclaredTaskColumns` re-homes it on the next engine start.
The card is never lost; it is visible in a lane the board may not draw
until one of those runs.

The one thing to watch after merge: (1) converts a previously-silent
success into a visible failure, so an operator mid-edit on a busy
workflow will start seeing a 409 they never saw before. That is the
point — the alternative was stranding their cards — but it is the change
most likely to generate a "this used to work" report.

### The thing that made this more than a gate removal

Un-gating the switch guard surfaced that
`selectTaskWorkflowAndReconcileImpl` read the task through
`store.readTaskFromDb` — the **synchronous SQLite** reader, which throws
under PostgreSQL:

```
TaskStore.db: SQLite Database is not available in backend mode
```

The flag returned before that line, so the gate was hiding a path that
**could not execute at all in the production backend**, not merely a
disabled feature. Ported to the async `readTaskRow`. Found by the new
tests, not by reading the code.

### Review round 2 (both findings real, both fixed)

**Torn write with no alarm — fixed by ORDERING, not by a louder
message.** My first attempt only made the error loud, which left the
torn state intact. The real fix is that the deterministic rejection
cause (destination at its WIP limit) is now checked BEFORE
`selectTaskWorkflow` commits, by resolving the target IR straight from
`workflowId` instead of through the task's selection. Nothing commits on
that path.

For the residual race the failure is loud AND recorded: `rehomeOccupant`
now returns `{ moved, error? }` (additive; sweep callers ignore it), the
switch writes a `task:workflow-switch-torn` run-audit row, and throws
`WorkflowSwitchRehomeFailedError` with `committed: true`. Consumers
translate it: the dashboard route returns a structured 409 with
`selectionCommitted`, and `fn_task_set_workflow` returns the same fields
— no more generic "something went wrong".

**Fabricated column for a deleted task.** My first fix fell back to
`fromColumn` when the final read found no row, so a task soft-deleted
mid-switch was reported as having its old column *preserved*. Absent now
reads as absent (the optional `reconciliation` is omitted). Extracted as
the pure `buildSwitchReconciliation` seam because the window is not
reachable through the public call — `selectTaskWorkflow` rejects an
already-deleted task up front — so it is a genuine race, and I test the
decision directly rather than asserting it from reading the code.

### Revert-proof, measured

New `workflow-reconciliation-production-shape.pg.test.ts` — 6 cases,
with the flag **never written**, which is the configuration every real
project has. Each flip reverted individually:

- re-gate the edit guard → **2 failures** (OccupiedColumnsError case;
rehomeTo re-home case)
- re-gate the delete capture → **1 failure** (card stays in
`custom-hold`)
- restore the switch early return → **2 failures** (`reconciliation`
undefined; card does not move)
- all three in place → **6/6 green**

Round-2 fixes, also measured:
- restore the `fromColumn` fallback → the "row is gone" case fails
(reports `preserved: true` for a deleted task)
- drop the `!outcome.moved` throw → the capacity-blocked case fails
(resolves instead of raising)
- **move the capacity pre-flight back AFTER the commit → the case fails
on the SELECTION assertion** (expected `WF-002`, received `WF-001`),
i.e. it proves the ordering, not the wording

The pre-existing coverage in `workflow-authoritative-reads.pg.test.ts`
reached the occupied-column guard by **writing the flag ON itself** —
same pattern as the ListView/Board suites in part 1. Its flag write is
removed; it now runs in the production shape.

### Where I nearly got this wrong

My first revert harness was buggy and I briefly concluded the delete
re-home was **redundant** — I had probed the stored column and seen
`triage` with what I thought was the flip reverted. It wasn't.
`workflow-ops.ts` contains two identical `const occupantTaskIds = await
store.listWorkflowOccupantTaskIds(id, false)` lines (field-reconcile
block, delete path), so my first-match edit reverted the wrong one.
Re-run anchored on surrounding context, the delete case fails as
predicted. Recorded in the test header as a caution. I also chased and
**refuted** a scarier hypothesis along the way — that an unrelated
`updateTask` coerces a custom column back to `triage`. It does not; the
column survives.

### Deliberately NOT in this PR

The v1-IR rollback-compat persistence (`downgradeIrToV1IfPure`) on the
workflow UPDATE path. It shared the same `flagOn` variable, which is how
it surfaced: **one flag read was feeding two unrelated decisions, so the
flag has more decision sites than call sites** — my earlier 9-site
inventory undercounted. It chooses the stored *shape* of the graph
rather than gating a guard, so it is a persistence-format change with a
different blast radius. It now reads the flag explicitly, behaviour
unchanged, for a follow-up.

The `moves.ts` group remains U2b's.

### Verification

`pnpm test:gate` (307 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (17
steps), both typechecks green. Full `packages/core` PostgreSQL suite:
**1042 passed, 3 failed** — `central-archive-secrets.test.ts`
(log-prefix assertion) and
`workflow-settings-project-identity.pg.test.ts` (×2, project-id
resolution). I confirmed the identical 3 failures on a stashed clean
tree: pre-existing, unrelated. No Fusion instance booted.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Workflow edits now prevent removal of occupied columns unless cards
are moved to a specified destination.
* Cards are automatically re-homed when workflows are deleted or
switched.
* Workflow switches now check destination capacity before committing and
provide clear conflict details when re-homing fails.
* Reconciliation results now indicate whether cards were moved or
preserved.


<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:46 -07:00
gsxdsm
743df98aa4 capacity, part 1: merge pinned at 1, worktrees-off mode, and one dead knob deleted (#2502)
First slice of the capacity simplification. Operator: *"just have two
capacity — overall per project agent count and max worktrees. Remove all
other capacities and counts."* Plus two later additions: **merge is
always 1, fixed**, and **worktrees off ⇒ limit by total agents only**.

Three independently revertable commits. No limiter is added anywhere;
one is deleted, one is made structurally absent, and one is pinned.

---

## 1. Merge concurrency ratcheted at 1 (test-only)

I was asked to add a limiter if merge concurrency could be raised. **It
cannot** — there is no setting, workflow property, pool or trait config
anywhere that raises it, so this adds no code and pins what already
holds.

Serialization lives in the **pump**: `drainMergeQueue`’s `mergeRunning`
re-entrancy latch, `activeMergeTaskId` as a single-slot identity, the
`mergeBodyInFlight` next-generation latch, and one `ProjectEngine` per
projectId.

**Not** in the merge-queue lease, which is a per-task ROW (`primaryKey
[projectId, taskId]`) — two tasks can hold leases simultaneously by
construction, and it has exactly one caller (the worktree-reuse
handoff). Ordinary merges never take it. A lease-level test would have
been describing an invariant that layer has never held.

The second half guards the other direction: a merge-concurrency
*setting* would not fail the pump ratchet — it would sit unread until
someone wired it up.

**Revert-proof:** deleting the latch → `expected 1 times, but got 2
times`; deleting the `finally` → latch-stuck; injecting
`maxConcurrentMerges: 2` → fails naming the key; injecting a
`maxParallelLanes` merge-trait field → fails naming the field. Sources
restored byte-identical after each injection.

## 2. `worktreesEnabled` — off means the worktree limit cannot bind

No worktrees-off mode existed (no
`worktreesEnabled`/`useWorktrees`/`worktreeMode` anywhere — only
worktree *configuration*).

**Why not `maxWorktrees: 0`, which needs no new key:** it deadlocks. `??
4` keeps `0` (not nullish), the gate is `used >= limit`, so `0 >= 0`
holds **on an empty board** and nothing ever dispatches — while the
operator-visible reason reads `gate=maxWorktrees; used=0/0`, a limiter
that looks like it is working while the board is dead. It also needs the
Command Center `{min:1}` clamp relaxed. So `0` costs the gate rewrite
*and* the clamp change *and* encodes a mode as a magic value.

**Off is absence, not a big number.** `resolveWorktreeCapacityLimit`
returns `number | null`; `ConcurrencyGateDiagnostic.maxWorktreesGate` is
now optional, so consulting a worktree limit in OFF mode does not
type-check. A gate holding `Infinity` can start binding again the moment
someone "fixes" a comparison; an absent gate cannot.

That paid for itself immediately: making it nullable surfaced a
**second, independent** worktree gate (`activeWorktrees >= maxWorktrees`
early-return) that a skip-by-convention approach would have missed
silently.

**Scope, deliberately:** this is a statement about *counting*, not
isolation. It does not make concurrent agents safe to share one checkout
and builds nothing toward that — the non-worktree paths that exist today
are fallbacks to the operator’s own tree, one of which caused FN-8600.

**Revert-proof:** a resolver ignoring the flag turns both OFF scheduler
tests red while every ON test stays green — they reuse the *same*
fixture (5 in-progress, limit 4) that pre-existing tests prove blocks,
so the pair moves in opposite directions. Removing `disabled:` reddens
the UI test.

## 3. `maxTriageConcurrent` deleted — it controlled nothing

**Measured: zero enforcement reads.** The only `.maxTriageConcurrent`
reference in the repo was a route echoing it back in `/config`. FN-8453
removed the pool it gated and left the knob shipping in
`DEFAULT_SETTINGS`, the settings type, the section registry, the API
response and six i18n catalogs, doing nothing, for releases.

Historical FNXC comments are **updated, not deleted** — they explain a
real past incident; they now say "planning admission slot" so they stop
implying a live setting. Tombstoned so it cannot return.

`/config` loses a field; safe in-repo since `fetchConfig`’s own return
type never declared it.

---

## Two corrections worth recording

- I earlier reported `maxWorktrees` had **no** Settings UI. Wrong —
`WorktreesSection.tsx:47`; my grep was truncated by `head`. It changed
the placement (toggle beside it, rather than a duplicate key in
Scheduling).
- I planned to assert the queued-reason string is rewritten in OFF mode.
Measured that it is **unreachable**: when `maxConcurrent` binds, the
sweep bails before the per-task reason and logs nothing. The test
asserts absence instead.

Two near-misses caught before commit: a pre-existing FN-7505 guard
caught my *new* key missing a description mapping; and editing i18n via
`json.load/dump` silently dropped unrelated duplicate keys
(`autoUpdateAndRestart` in `fr`) — Python keeps only the last of a
duplicated key. Redone textually, every catalog re-validated.

## Verification

`pnpm lint` clean · core/engine/dashboard/i18n typecheck clean · `pnpm
test:gate` green (309 + 10 + 71) · capacity/worktree suites 11/11 ·
engine merge-invariant + scheduler 45/45 · dashboard settings 114/114.
Rebased onto current main and re-verified.

Nothing was booted at any point.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a project setting to enable or disable running tasks in
worktrees.
* Disabling worktrees removes worktree capacity limits from task
scheduling.
* The “Max Worktrees” setting is disabled when worktree execution is
turned off.

* **Changes**
* Removed the unused triage concurrency setting from configuration and
dashboard responses.
* Updated scheduling diagnostics and queue messages to reflect disabled
worktree capacity limits.
  * Added localized labels and help text for the new setting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:31 -07:00
gsxdsm
35b0df1838 U11 PR2: entry contract under the merged column + a real intake-column bug the audit surfaced (#2503)
Second small PR for **U11**. Two commits: a tests-only entry-contract
pin, then a **real present-day bug fix** the audit surfaced.

## The audit you asked for, finished — no design fork

You named four surfaces as the remaining risk. All four can take a
combined `intake` + `hold` column. One needed a code change; here it is.

| Surface | Verdict | Evidence |
|---|---|---|
| `isUnplannedForExecution` | Safe | PR1 (#2495) — passed unmodified; a
mutation now fails exactly the merged-column test |
| Capacity hold / release | Safe | PR1 — `hold-release.ts:260` already
accepts intake **or** hold |
| `start`'s column / entry contract | Safe | commit 1 — all 6 assertions
passed unmodified |
| `createTask` intake wiring | **Broken today** | commit 2 — fixed,
revert-proven |
| *(also found)* triage auto-discovery | Needs conversion |
`triage.ts:1382` — deferred to PR3, see below |

## Commit 1 — entry contract under the merged column (tests only)

All 6 new assertions passed on the first run. **Regression floor, not
evidence of a fix** — I could not make them fail and am not claiming
otherwise.

They pin one real behavioral **difference** rather than asserting
sameness everywhere: the merged shape answers `start` where the split
shape answers `plan`, because `start` becomes the first node in that
column once the columns collapse. That is equivalent *only* because
`start` reaches the specification node by a single unconditional success
edge — asserted, so if a node is ever inserted between them this fails
instead of silently admitting an unspecified card into implementation.

Also pinned: past planning both shapes agree exactly; a card past the
merged column still never resumes at a planning node (the backward drag
that fires `abort-on-exit`); and a row persisted in the **deleted**
`triage` column resolves to `undefined`, safe only while the executor's
start-node fallback exists.

## Commit 2 — a real bug, found by the audit

The intake column was resolved **only** as a by-product of materializing
workflow steps. A create supplying `enabledWorkflowSteps` without an
explicit `workflowId` takes **neither** materialization branch, so
`resolvedEntryColumn` stays `undefined` and `column:` falls through to
the hard-coded `|| "triage"`.

Today, on Coding (Ideas), that lands the card in `triage` — **a column
that workflow does not declare.** Created straight into a phantom lane.
Measured: the new test fails `expected 'triage' to be 'ideas'` against
unmodified sources.

**Why it blocks U11.** Once `triage` leaves the coding IRs this stops
being an Ideas edge case and becomes the default workflow's behavior for
every create down this path: the card lands in an undeclared column
**and** — because `isIntakeColumn` keys on the same `"triage"` literal —
gets `generateSpecifiedPrompt` instead of the bootstrap seed. Triage
admits a card for planning only when its `PROMPT.md` reads as a seed, so
a placeholder spec is classified "already planned" and never planned.
The card sits in Planning forever with no log line in any lane —
**FN-8587's exact failure mode, promoted from one edge case to every new
card.**

The fix resolves the intake column **side-effect-free** (read the IR,
ask which column carries `intake`). It deliberately does *not* call
`materializeDefaultWorkflowSteps`, which would persist step rows the
caller explicitly opted out of by supplying its own toggles.
Unresolvable workflow returns `undefined` and each call site keeps its
legacy fallback, so no path loses behavior when the IR cannot be read.

Applied to both create paths. Branch ordering preserved in both — the
explicit empty-toggle case (`length === 0` hydrating back as `[]`) still
runs, now nested rather than sequential.

**Revert check:** with `task-creation.ts` reverted, *"lands a Coding
(Ideas) task in ideas even when enabledWorkflowSteps is supplied"* fails
`expected 'triage' to be 'ideas'`. The companion bootstrap-`PROMPT.md`
assertion passes either way today — it is correct **by accident of the
`"triage"` literal** — and is kept precisely because that accident
disappears with U11.

## Verification

37 tests green across the three intake/create suites; 119 across the
entry-contract, merged-column and lifecycle suites; `pnpm test:gate`
green (307 + 10 + 71); lint and core typecheck clean. Changeset added.

## Deferred to PR3, with the line numbers

`discoverReadyPlanningTasks` has two hardcoded branches:

```ts
(t) => t.column === "triage" && isTaskStillInPlanningStage(t)   // triage.ts:1382
(t) => t.column === "todo"   && !this.processing.has(t.id) …    // triage.ts:1389
```

Delete `triage` and branch 1 matches nothing for coding cards; branch 2
then does all the work and is **narrower** (it admits only
`needs-replan` or bootstrap-stub cards). Commit 2 is what makes branch 2
sufficient — every new card now gets a real bootstrap seed. They cannot
double-fire: a card is in `todo` xor `triage`.

Two adjacent sites are already merged-shape-ready: `triage.ts:3899`
skips the redundant same-column move for a plan-in-place card, and
`triage.ts:753`'s stale-status sweep already scans both columns.

Then the ~10-line IR change, then the migration proof.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:14 -07:00
gsxdsm
9d3e53d0c5 U8 PR3: the implementation phase announces HOW it ended — including when the executor moved the card itself (#2507)
Third PR of **U8 — the graph owns execution**. Independent of everything
merged so far; small, green, revertable on its own.

## The problem this makes visible

`result.taskDone` is the entire language the execute seam has for
talking to the graph:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
return { outcome: "failure", value: paused ? "implementation-paused" : "implementation-incomplete" };
```

The endings that one bit cannot express are exactly the ones the
implementation phase **transitions itself**:

- a session that paused *after* the work was already complete →
finalizes to review inline;
- a session that stopped because a step is blocked on a pending review →
hands off to review inline (a pending-review block is a wait, not a
failure; marking it failed deadlocks a row that is both `in-review` and
`failed`).

The graph then sees `taskDone === false`, reports
`implementation-incomplete`, and `handleGraphFailure` compensates with
`alreadyFinalizedToReview` / `completionFinalized` — classifiers whose
entire job is recognising a move the graph did not make.

**That was invisible.** An out-of-band transition and a genuine
implementation failure were indistinguishable in logs, in events, and in
tests. You cannot remove a transition you cannot see, and you cannot
prove you removed it either.

## What lands

A closed `ImplementationExit` enum
(`engine/executor/implementation-exit.ts`) reported from six
completion-adjacent exits in `runImplementation`, announced by the
execute seam as `NodeCompleted.exit` on the U3 lifecycle bus. Two ids
are flagged as out-of-band — the ones where the executor, not the graph,
performs the transition.

**Routing is unchanged, and that is the point.** The seam returns
byte-identically what it returned before for every exit, so this PR
cannot move a card. The routing move needs new IR edges and lands
separately; splitting them is what keeps both independently revertable.
Per R5 an exit id is a **reaction** — nothing branches on one, and
dropping every subscriber must change no outcome (a named U8 test
scenario, asserted here).

`NodeCompleted.exit` is added to the event key allow-list deliberately —
which is exactly what that allow-list is for — and carries closed enum
ids only, never prose.

## Revert-proofs, each observed failing

| Injected change | Result |
|---|---|
| Remove the emit entirely | **6 failures** |
| Let an exit change the returned outcome | **2 failures** (the
routing-unchanged pins) |
| Delete one `reportImplementationExit(...)` call site | **1 failure**
(the wiring ratchet) |

**The third proof exists because of a hole I found in my own tests.**
These tests stub `runImplementationPhase` — the only way to reach all
six exits deterministically — which means deleting a real call site left
the entire file **green**. A stubbed seam can only prove the seam. I'd
also written "every exit is reported — the signal is real, not a
placeholder" in the header, which the tests did not support. Both are
fixed: there is now a ratchet asserting every enum id is wired at a real
call site and that each out-of-band id sits adjacent to the handoff it
describes, and the header says what the tests actually prove.

## Scope

**6 of `runImplementation`'s ~28 dispositions** (per the ownership
ledger merged in #2490), chosen as the ones the routing move needs. The
remaining ~22 report nothing yet — the ledger, not this enum, stays the
record of that gap, and the module says so.

## Verification

- 15 new tests + ledger + graph-boundary + task-done-blocked +
graph-requeue-gate + step-session + review-verdicts + tool-failure-retry
— **9 files, 115 tests green**
- `@fusion/core` `workflow-events` — 20 tests green (allow-list change
covered)
- `pnpm test:gate` green (17/307, 2/10, 1/71); `pnpm lint` clean; `tsc
--noEmit` clean on both packages
- Changeset included (`patch`, `internal`), passes `check:changesets`

## Next

PR4 is the routing move itself: `review-handoff-pending-review` becomes
a graph outcome with its own IR edge, and `alreadyFinalizedToReview`
becomes provably unreachable for that path. The IR edge change will be
its own commit, separate from the seam change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:07 -07:00
gsxdsm
8288e4a8ab U7 PR3: the specification reaction acts on what finalize DID, not on the fact that planning stopped (#2506)
Completes the pair started in #2498. That landed the outcome; this makes
the engine's reaction consume it.

## The bug

`onSpecifyComplete` fired on **every** finished specification, because
the seam announcing it fired unconditionally. So a card parked at the
manual plan-approval gate — finalize writes `status:
"awaiting-approval"` and **returns early**, before the release move —
was logged as `Specified X → todo` and had a Plan Review run armed for a
plan the operator had not approved.

#2491 stopped the **seeder** from acting on that, defensively, at the
seeder. This removes the reason it was ever asked. Both layers are
deliberate and neither is redundant:

- the seeder guard covers **every caller**, including self-healing's
re-seed;
- this one stops the engine doing work nobody asked for, and stops it
telling the operator something false about their own board.

`released` is the only outcome that licenses arming a run — the only one
meaning the card crossed into the hold column (or was already resting
there, plan-in-place) and is the graph's now. `parked` belongs to a
human; `withheld` belongs to the caller's retry budget.

## The event still fires on every outcome

Deliberately. Dropping the reaction for a non-release would also drop
the runtime's `recordActivity()` idle signal, and a reaction that
silently does not happen is harder to reason about than one that happens
with an accurate payload. R5's division of labour: **the seam announces,
the subscriber decides what a given outcome licenses.**

## Why there is a new extracted function

`reactToSpecificationComplete` is pulled out of the inline
`InProcessRuntime` callback for the same reason the continuation drain
was in #2491: the callback is built inside a class whose construction
attaches to the real central project registry, so no test could
distinguish *"the reaction respects the outcome"* from *"the reaction
ignores it"*.

**Revert proof:** with the outcome gate removed from the reaction, **5
of 8 fail**.

## Two call-site decisions worth naming

**`tryFinalizeExplicitDuplicateMarker` reports through a mutable ref,
not a widened return type.** Its boolean answers a *different* question
— "was this a duplicate marker at all?" — and 16 existing tests assert
it directly. I tried the widened return first and it turned all 16 red.
Expectation edits are exactly how a behavior change travels disguised as
churn, so I backed it out. **This diff touches zero existing test
expectations.**

**A duplicate-marker redirect reports `parked`**, which is accurate: it
deletes, flags, or clears the marker; it never releases the card into
the hold column.

## A fixture note — third of this shape on the program

My "task vanished between release and reaction" case passed `undefined`,
which triggered the harness **default parameter** and silently handed
the reaction a live task — making it a duplicate of the control rather
than the case it claimed to be. It now passes `null`, with a comment
saying why.

Running tally of near-false-greens on this unit, all the same family: a
fake that ignores its predicate (#2491), a stub that ignores its
callback (#2498), a default parameter that swallows the interesting
input (here). Each was caught by the test failing for the *wrong reason*
and being read rather than fixed.

## Verification

| Check | Result |
|---|---|
| new suite | 8/8 |
| 15 triage / planning / continuation suites | 361/361, **no expectation
edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (307 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:15:29 -07:00
gsxdsm
46f35323cf fix(core): make the capacity gate actually bind for real projects (R2) — USER-VISIBLE (#2499)
Follow-up to #2488 (merged). **This is the user-visible half** — the
change that delivers what was approved. #2488 alone is latent.

## One line

`workflow-capacity.ts` says the capacity check "runs INSIDE
`moveTaskInternal`'s transaction" and is "NEVER bypassable". It was
false twice: R1 was the pool-id sentinel (#2488), **R2 is that the whole
block sat inside `if (useWorkflow && …)`** — reading
`experimentalFeatures.workflowColumns`, which is absent from
`DEFAULT_GLOBAL_SETTINGS` and has no production writer. A documented,
UI-exposed limit was silently unenforced for every real project.

**Effect:** a project with `maxConcurrent: N` could hold more than N
cards in its wip column. Now the move is refused with
`capacity-exhausted`.

## Scope is deliberately narrow

**Only the capacity check is un-gated.** `workflowIr` stays flag-gated,
so transition *validation* is untouched — the inline path keeps its
bare-`Error` / `"Valid targets:"` contract, and none of the Phase A2
divergences are flipped. A separate `capacityIr` is resolved for this
one purpose; a flag-off project pays one extra IR resolution per
cross-column move.

## The release path already expected this

`hold-release`'s own docstring:

> the in-txn capacity check is **NOT a guard — it still runs** (KTD-10),
so two holds racing into one slot serialize: exactly one commits, the
other rejects with `capacity-exhausted` and retries next sweep

and it reserves worktree + semaphore slots *before* issuing a move
specifically so it can release them on that rejection. **That handler
was dead code.** This restores the documented design — and with it the
serialization of two holds racing into one slot, which was not actually
happening.

## Measured blast radius — not estimated

| suite | with R2 | baseline | new failures |
|---|---|---|---|
| core PG (real store) | 1037 passed / 3 failed | 1037 passed / 3 failed
| **0** |
| engine-default | 279 failed / 9167 | 279 failed | **0**
(failing-file-set diff) |

The three core-PG failures are the same pre-existing ones that reproduce
with everything stashed. Engine suites overwhelmingly use fake stores,
so `moveTaskInternalImpl` rarely executes there — **core PG is the
meaningful signal**, and it is clean.

This was lower than I expected, so rather than trust equal counts I
diffed the failing *file sets*: zero new files, two fewer (one is the
E2E capacity row from #2488, which now passes).

## Acceptance

Flipped exactly as Phase A3 specified: `DEFECT (R2, STILL LIVE)` →
`FIXED (R2)`, and move-path-equivalence's capacity `DIVERGENCE` →
`CONVERGED`. **Both fail with this change reverted** (verified: 2 failed
/ 12 passed).

## Why I proceeded without a decision

I had escalated R2 and had no answer. Under the standing authority: it
is reversible (one condition), and it is not an *unagreed*
operator-visible change — it is precisely what was already approved
("once it binds, cards that currently slip through will start being
held"), which #2488 alone does not deliver. My recommendation was option
B and I acted on it. Revert is one PR.

Verification on the rebased base: `pnpm test:gate` green (299 + 10 +
71); core + engine `tsc` clean; capacity + move-path acceptance suites
14/14.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Column WIP limits are now enforced when moving tasks into full
columns.
- Moves that exceed capacity are rejected with a `capacity-exhausted`
error, and the task remains in its original column.
- Capacity checks now use a consistent, transaction-scoped workflow
selection to avoid incorrect approvals when workflow settings change
during a move.
- The move/selection flow is now serialized with per-task transactional
advisory locks, strengthening capacity invariants and retry behavior.
  - Existing transition validation behavior remains unchanged.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 16:28:00 -07:00
gsxdsm
eaea082259 U8 PR2: the execution-policy ladder resolves its own workflow's columns (the wip literal made retry, escalation and loop protection unreachable) (#2497)
Second PR of **U8 — the graph owns execution**, independent of
[#2490](https://github.com/Runfusion/Fusion/pull/2490) and of every
other unit. Small, green, independently revertable.

## The defect

`handleGraphFailure`'s execution-policy ladder — FN-7863/FN-7926
dispatch-loop terminalization, FN-7996 tool-failure retry, FN-7998
escalation — decided a task's own lifecycle by naming `"todo"` and
`"in-progress"` **literally, at 9 sites**. U5b converted the executor's
*rebounds* to `resolveReboundColumnFor`; these were left behind, each
sitting somewhere an awaited resolver could not reach: inside
synchronous `updateTaskAtomic` mutators, inside fire-and-forget resume
closures, and in conditions evaluated before any resolution happened.

**The severe one is the wip gate, and it fails silently in the worst
direction:**

```ts
if (live.column !== "in-progress") {
  // "Workflow graph run ended after task already advanced — no further action needed"
  return;
}
```

Under a workflow that renames the implementation column, that is true of
a card sitting in **its own wip column**. So the graph failure was
swallowed whole — no terminal park, no status, no error, nothing on the
board — and the scheduler re-dispatched the same doomed run. Every later
branch sits behind that gate, which is why the retry budgets, the
escalation, and the bounded terminalization were **unreachable rather
than mistargeted**.

This is precisely the failure the program's problem frame predicts: *a
guard that stops matching disables a recovery path invisibly and the
suite stays green.* I found it because my first renamed-column test for
the escalation site could not reach the escalation code at all.

Two further sites misbehave once the gate is passable:

- **FN-7998 node escalation** wrote `column: "todo"` inside the atomic
claim — parking the card where no workflow declares it, which is on the
plan's **"Stop implementation if"** list and what R7 exists to clean up
after. The scheduler's effective-node resolution, the entire point of a
node escalation, never runs.
- **FN-7863/FN-7926's `live.column === "todo"` arm** is the classic
guard that stops matching. In-process the `executeNodeSelfRequeued`
marker covers the same case, so this degrades only on the **durable**
arm — after a restart, or for a second `TaskExecutor` instance in the
process, where the column read is the only evidence the inner executor
requeued. A progressing card then falls through to the terminal sink and
is parked `failed`.

## The fix

Resolve hold and wip **once per graph failure** through U1's
`resolveTaskLifecycleColumns` and thread the pair through the ladder.
Both fall back to the legacy literal when the workflow cannot be
resolved, so an unresolvable workflow keeps exactly its pre-conversion
behavior rather than guessing. One IR read on a terminal recovery path —
not an enumeration loop.

## Red-green, measured

**3 of the 8 new tests fail with this commit's executor change
reverted:**

```
FAIL  FN-7998 … > requeues a node escalation to the RENAMED hold column, not the literal todo
FAIL  FN-7998 … > still does not move the card for a MODEL-target escalation
FAIL  FN-7863/FN-7926 … > recognises an inner-executor requeue that landed in the RENAMED hold column
      Tests  3 failed | 5 passed (8)     ← reverted
      Tests  8 passed (8)                ← with the fix
```

The other **5 pass both ways by design**, and I am not claiming them as
red-green — they are the regression floor:

- default coding workflow still resolves hold → `todo`, wip →
`in-progress` (byte-identical);
- an unresolvable workflow still uses the legacy literals;
- the in-process self-requeue marker still works when no workflow
resolves;
- and a **negative case** proving the dispatch-loop gate stays narrow —
a card still in its wip column with no marker is a genuine execute
failure and must NOT be swallowed as a benign recovery. Widening that
gate to "any column" would have been the easy wrong fix.

## Scope

Deliberately the execution-policy ladder only. **20 further column
literals remain in the same method's pause-abort, merge, and in-review
regions** — they belong to U5's executor slice (B4, not started) and
U9's merge lane, and are untouched here. Flagging the overlap: this PR
edits `executor.ts`, so whoever takes U5-B4 should rebase onto it rather
than converting these 9 sites again.

## Verification

- 8 new tests + the preserved-behavior suites
(`executor-tool-failure-retry`, `executor-graph-requeue-gate`,
`executor-task-done-blocked`, `executor-graph-boundary`,
`executor-stuck-requeue-preserve-progress`,
`executor-paused-abort-todo-benign`, `executor-abort-provenance`) — **9
files, 112 tests, green**
- `pnpm test:gate` — green (2/10, 16/299, 1/71); `pnpm lint` clean; `tsc
--noEmit` on `@fusion/engine` clean
- Changeset included (`patch`, category `fix`), passes `pnpm
check:changesets`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed execution recovery for workflows with renamed lifecycle columns
so retry, escalation, and loop-protection behaviors correctly follow the
workflow’s declared hold/WIP columns.
* Preserved legacy behavior for default workflows and continued safe
handling when lifecycle columns can’t be resolved.
* **Tests**
* Added a Vitest suite validating execution-policy “ladder” behavior for
renamed columns, including node escalation, dispatch-loop gating, and
fail-closed scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:47 -07:00
gsxdsm
fd6d005333 U12 part 1: delete the legacy board path (262 ListView + 39 Board tests were measuring it; 9-site flag inventory, moves.ts group blocked on U2b) (#2500)
## U12, part 1 of 2 — and one blocker you need to route

The unit's headline deletion
(`isWorkflowColumnsCompatibilityFlagEnabled`) is **blocked by U2b** and
is not in this PR. What is here is everything that could be deleted
without making a convergence decision that belongs to another unit.

### The blocker

PR #2468 landed as `b941d3cba` — but that was **Phase A2 steps 1–2 only:
the differential characterization**. The convergence (pick a path,
delete the other, delete the flag) has not landed;
`feature/workflow-move-path-convergence` is still live.

Deleting the raw flag **is** that convergence.
`move-path-equivalence.pg.test.ts` says so in its own header, and its
second `describe` is literally *"the flag gates MORE than side
effects"*. The plan makes this a blocking unit with an equivalence
*proof obligation* and an explicit "stop and escalate rather than
reconcile silently" note. So I stopped.

### Inventory: every read of the raw flag, with a verdict

Nine sites. All false in production because nothing writes
`experimentalFeatures.workflowColumns`.

**Blocked on U2b — one branch, not separable:**

| Site | Silently disabled today | Visible if flipped |
|---|---|---|
| `moves.ts:312` `useWorkflow` | typed `TransitionRejectionError`,
workflow adjacency, the shared transition invariants (merge-blocker
*trait* generalization), plugin column gates, the `transitionPending`
marker, `workflowId` in `task:move` run-audit, and the trait-hook
side-effect path | Yes — rejections change **type and message** |
| `moves.ts:931` | the in-transaction capacity gate.
`resolveColumnCapacity` never runs | Yes — WIP limits begin binding |
| `workflow-task-create-ops.ts:351` |
`prepareWorkflowMovePolicyPreflight` returns `undefined` unconditionally
→ **workflow/plugin move policies have never been evaluated** | Yes —
new rejections |

On #2488: the pool-id sentinel fix is correct *and* still inert. Two
dead layers stacked — the gate it fixed is inside `if (useWorkflow &&
…)`.

**Not blocked, but each moves operators' cards — deferred to PR 2 per
your call:**

| Site | Silently disabled today |
|---|---|
| `workflow-ops.ts:183` | `OccupiedColumnsError` + `rehomeTo` when a
workflow edit removes an **occupied** column. Today the save succeeds
and strands the cards |
| `workflow-ops.ts:344` | occupant re-home on workflow **delete** |
| `workflow-definitions.ts:700` | workflow-**switch** reconciliation,
and the `reconciliation` field in the API response |

I verified these three are **not** coupled to `moves.ts`:
`rehomeOccupant` reaches a custom target via the
`isWorkflowDeclaredRecoveryRehome` carve-out (`moves.ts:641`), which
exists because the repair "silently no-oped on every store open" before
it.

**Not blocked, no behaviour change for current binaries** (also PR 2):
`project-store-ops.ts:687` + `lifecycle-ops.ts:1119` —
`downgradeIrToV1IfPure` on persist, for *binary-downgrade* rollback.
Needs a round-trip test, not an assumption.

### What this PR deletes

**Dashboard.** `workflowColumnsEnabled` was a literal `true` at all
three `MainContent` call sites; the server hardcodes `flagEnabled:
true`. Gone: Board's legacy single-lane board (55 lines mapping the
hardcoded `COLUMNS` enum — the last board surface deriving columns from
the legacy vocabulary, an R8 violation that survived U10);
`tasksByColumn` and its cache ref, orphaned with it; ListView's
`LEGACY_LIST_COLUMNS` (the ListView copy of the synthesized-trait-flags
defect U10 fixed in Board); both props; the `shouldHydrateCache` gate;
TaskDetailModal's `flagEnabled` early return. **Neither Board nor
ListView imports the legacy column enum any more.**

**Core.** `evacuateCustomColumnsToLegacy` (#1409) — both triggers
require the previous settings to have the flag ON, which no writer
produces. `runWorkflowColumnsIntegrityPass` — no caller anywhere,
superseded by `reconcileUndeclaredTaskColumns` (registered in startup
recovery), and it read through the sync SQLite handle, so invoking it
under PostgreSQL would have thrown rather than reconciled.

**Migration answer:** a project with `workflowColumns: false` persisted
needs no migration and no read-time drop. Nothing in this PR reads the
key, and it stays in `HIDDEN_EXPERIMENTAL_FEATURE_KEYS` so Settings
still suppresses it rather than resurrecting it as an unknown setting.
Proven by tests, no instance booted.

**`flagEnabled` stays on the wire** as a constant. Removing it changes
the response shape, and a browser tab outliving a server upgrade would
read the missing field as "off" and degrade. One boolean, no client
branches on it, droppable a release later.

### Measured

- Production sources: **-332 / +131** (net **-201**). Additions are
almost entirely FNXC comments recording why each branch was unreachable.
- Dashboard production only: -168 / +93.
- Core: -164 / +38.

### The finding I'd actually flag

`Board.test.tsx` and `ListView.test.tsx` both left
`workflowColumnsEnabled` unset and stubbed `fetchBoardWorkflows` with a
**never-resolving promise**. Under the old gate that rendered the
**legacy** board — so **262 ListView tests and 39 Board tests were
asserting against a configuration production never reached**, and a real
regression in the workflow board or list would not have failed either
file. Same shape as the other four: looked enforced, wasn't.

Both now seed the first-paint lane cache with the default workflow's
**real** columns (ids and names copied from
`BUILTIN_CODING_WORKFLOW_IR`) — the same seam production uses.
Repointing them surfaced assertions that encoded legacy-only values:
`"In Progress"`/`"In Review"` (real IR names are `"In progress"`/`"In
review"`), and Planning Mode asserted to receive `null` as the workflow
id, which is only what `getTaskPlanningWorkflowId` returns when
`workflowMode` is false.

`"Back to In Progress"` is **not** one of those — it is a hardcoded i18n
string in `TaskContextMenu:210`, not derived from the column name. Left
alone, and flagged: it will not follow a renamed column. That's U11
vocabulary territory.

**One test is SKIPPED, not weakened** — "keeps unaffected columns stable
when archived collapse toggles". Pointed at the real board the invariant
is **false**: toggling the archived column re-renders unaffected columns
(measured: todo renders 3×, not 2×). Pre-existing production behaviour
this deletion exposed, never covered because the test measured the dead
path. I ruled out the obvious causes (every callback prop is
`useCallback`; the per-column task memo's deps exclude
`archivedCollapsed`; memoizing the inline `canDropTask` binding did
**not** close it — I wrote that fix, could not prove it with a failing
test, and **reverted it**). The reason is recorded at the test: un-skip
with a fix, never with a new expected number.

### Verification

`pnpm test:gate` (299 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (17
steps), and both package typechecks green. `settings-defaults.test.ts >
warns once per process for legacy cwd-main mode` fails —
**pre-existing**, confirmed by stashing my changes and re-running. No
Fusion instance was booted.

### Routing request

Per your call: the `moves.ts` group and the final removal of
`isWorkflowColumnsCompatibilityFlagEnabled` go to **U2b**, inside the
convergence PR where the equivalence proof already lives. The
divergences their characterization suite does **not** yet cover: plugin
column gates, the `transitionPending` marker, `workflowId` in
`task:move` run-audit, and move-policy preflight.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:40 -07:00
gsxdsm
2934cccad8 U7 PR2: finalize reports what it did with the card — a refused planning handoff is retried, not counted as recovered (#2498)
## The bug

`finalizeApprovedTask` has ~25 exit points and returned `void`, so no
caller could tell *"the card was handed off"* from *"finalize gave up"*.
Both callers assumed success.

`recoverApprovedTask` returned `true` **unconditionally** after
finalize, and `handleStuckAbortRequeue` treats `true` as "recovery done,
stop here". So when the release move was **refused by the planning-stage
guard** (FN-8361), or the store could not perform the move at all,
recovery reported success and the card's stuck-retry budget was skipped
— nothing re-planned it, nothing escalated it, and it sat in the planner
column holding a finished spec.

The refusal was already logged loudly by FN-8596's visibility work. The
return value was the part still lying.

## Three states, not a boolean

This is the load-bearing decision in the PR:

| Outcome | Meaning | Retry? |
|---|---|---|
| `released` | crossed into the hold column, or already resting there
(plan-in-place) | n/a — handed off |
| `parked` | deliberate, terminal-for-now: awaiting manual plan
approval, duplicate decision, operator pause, deleted duplicate | **no**
— a human owns it |
| `withheld` | finalize could not complete the handoff, nothing waiting
on a human | **yes** — caller's budget owns it |

`recoverApprovedTask` returns `outcome !== "withheld"`, so **`parked`
still returns `true`**. Narrowing to `=== "released"` is the tempting
simplification and it is wrong: it would send the stuck handler down its
draft path and stamp `needs-replan` over a plan a human is mid-review on
— a worse bug than the one being fixed. That is asserted, and the
assertion fails under exactly that narrowing.

## Why a mutable report, not a return at each exit

Threading a return through 25 exits is 25 chances to mis-classify a
branch, and mis-classifying turns a truthfulness fix into a lifecycle
bug. The report defaults to `parked`, which is equivalent to today's
observable behavior at every exit — so the plumbing is **inert
everywhere except the three sites explicitly classified**. Adding a
state to an exit is then a deliberate, reviewable act rather than a
diff-wide judgement call.

Only **two** exits are marked `withheld`, both in the release block,
both already warning loudly. Deliberately *not* marked:

- the `updatePlanningStateIfStillCurrent` guard — FN-8024 says a normal
scheduler advance legitimately lands there; the card has moved on, so a
retry would be wrong.
- `recoverMissingPromptBeforeRelease` — it owns its own recovery budget;
retrying would double up.

## Revert proofs (measured)

| Reverted | Result |
|---|---|
| `recoverApprovedTask` back to unconditional `true` | `Tests 2 failed
\| 3 passed (5)` |
| narrowed to `outcome === "released"` | `Tests 1 failed \| 4 passed
(5)` — the approval-park control |

The second row is the point: the park case is load-bearing, not
decoration.

## A fixture note that nearly produced a false green

A `vi.fn()` stub for `updateTaskAtomic` that ignores its callback makes
**every** finalize report "no longer in the planning stage" and return
before the release — silently collapsing every case into the same
uninteresting early exit. My first run was 3 failures for that reason,
not the reason I expected. The fake now applies the patch, and the
control asserts `moveTaskIf` was actually reached. Same class as the
`moveTaskIf` fake caught on #2491; recording it so the next person
recognises the shape.

## Scope

The other caller — `specifyTask`'s unconditional `onSpecifyComplete` —
is **not** gated here. Reaching it needs a live planning session, so
gating it without first extracting the reaction would be a change I
cannot prove, which is exactly the finding review caught on #2491's
deferral. That lands next, on this plumbing.

## Verification

| Check | Result |
|---|---|
| new suite | 5/5 |
| 11 triage/planning suites (triage, finalize-duplicate-lineage,
stuck-requeue-preserve-draft, explicit-duplicate-marker, preflight,
plan-artifact-writeback, refinement-routing, planning-wake,
planning-evacuation, …) | 326/326 |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (299 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:36:20 -07:00
gsxdsm
a271f1868f U10: dashboard renders workflow-resolved columns (6 legacy-vocabulary defects, incl. a silently-disabled open-PR guard) (#2492)
Phase D / **U10** of the workflow-owned-lifecycle program (**R8**). This
unit **blocks U11** (merge Todo into Planning) — the board must render
IR-resolved columns before the column shape can change.

## What was wrong

Six dashboard surfaces answered a column question from the legacy
`COLUMNS` / `VALID_TRANSITIONS` vocabulary rather than the card's own
workflow IR. Each is a defect today, and each is a way U11 would ship
visibly broken.

| # | Surface | Defect |
|---|---|---|
| 1 | Board — All workflows | Appended **every** legacy column id to the
lane union with synthesised flags → a phantom lane for a column no
workflow declares, labelled with the raw id, ordered by the enum index
with an alphabetical tie-break that scrambled a custom workflow's
declared order |
| 2 | ListView | `if (groups[column])` **silently dropped** a row whose
stored column the workflow no longer declares — no lane, no row, no
error |
| 3 | Move menu | A card stranded in an undeclared column got an **empty
move list** — the one surface that could rescue it offered nothing |
| 4 | Task Detail | Header badge rendered the raw stored id;
title/description editing gated on the literal `{triage, todo}` — a
renamed planning lane lost Edit with nothing on screen to explain it |
| 5 | `board-workflows` | The built-in lifecycle label map was an
**override**, not a fallback, so it replaced a name a built-in
deliberately chose |
| 6 | `POST /tasks/:id/move` | The open-PR backward guard used
`COLUMNS.indexOf(...)` → **-1 on any renamed board**, and the guard
treats a negative index as "allow" |

**#6 is the one worth reading twice.** The guard did not start rejecting
the wrong things — it stopped existing. On a renamed board an operator
could drag a card backward out of review with an open GitHub PR,
orphaning it, and nothing failed. This is precisely the "a converted
guard silently stops firing" row in the plan's risk table, reached
through a rename rather than a conversion.

The re-engage copy of that same guard is **deliberately left on the
legacy enum**, with a comment saying why: it is gated on literal
`in-review` / `in-progress` end to end, so converting only its indices
would make it *weaker* (a workflow declaring `in-review` but not
`in-progress` would score -1 and disable it). U5 owns that lane.

## Evidence

**Every fix has a test that fails when the fix is reverted.** With the
six production files stashed and the tests kept, **10 of the 28 tests
fail**:

- Board aggregate: phantom lane present (2)
- ListView: stranded card dropped, desktop **and** mobile (2)
- Move menu: empty move list for a stranded card (1)
- Task Detail: badge shows `staging`, Edit missing in a renamed intake
**and** hold lane (3)
- `board-workflows`: `builtin:lead-generation`'s `triage` renders as
"Planning" (1)
- Move route: backward move between renamed columns **allowed** with an
open PR (1)

The other 18 are regression pins on behaviour that must not change
(default-workflow lane order and labels, legacy `in-review →
in-progress` block, legacy editable columns, forward moves, terminal
PRs).

**Measured, not estimated.** The label-map clobber was quantified
against the built-in IRs actually in tree: **4 column names replaced — 3
case-only variants ("In progress" → "In Progress"), 1 genuine semantic
rename.** Only the rename is a user-visible defect; the fix preserves
the case normalisation rather than churning the default board.

## Surface enumeration (AGENTS.md)

Desktop **and** mobile — the breakpoint is `(max-width: 768px),
(max-height: 480px)`, so landscape phones exceed 768 wide and match on
height. Column states: empty, populated, duplicate id across two
workflows, and a column no workflow declares. Views: single-workflow
lane, All-workflows aggregate, list, move menu, task detail, move route.

## Regression check

Full dashboard suite, both sides of the change:

| | Test Files | Tests |
|---|---|---|
| Before | 41 failed / 1068 | **296 failed** / 21195 |
| After | 42 failed / 1071 | **297 failed** / 21223 |

`+28` total is exactly the tests this change adds. The single failure
delta is `register-model-routes-kimi-k3-supplemental`, which **fails
identically on this branch's base when run in isolation** — shard-order
dependent, unrelated to columns. **Zero regressions attributable to
U10.** The ~296 pre-existing dashboard failures are inherited from main
and are flagged to the coordinator, not touched here.

`pnpm test:gate`, `pnpm lint`, both dashboard typechecks
(`tsconfig.json` and `tsconfig.app.json`), `pnpm smoke:boot`, and `pnpm
check:changesets` are green.

## Not in this unit

- `Board`'s legacy single-lane `COLUMNS.map` fallback still exists.
`MainContent` passes `workflowColumnsEnabled` unconditionally, so it is
unreachable in-app, but proving that is a deletion argument and this is
not a deletion unit — flagging rather than removing.
- `ListView`'s `LEGACY_LIST_COLUMNS` fallback, same reasoning.
- `flagEnabled` on the wire (U2 noted U10 retires it once no client
reads it) — four clients still branch on it; retiring it is a
client-shape change that belongs with U11's shape work.
- `GET /api/tasks?column=` still validates against `COLUMNS`, rejecting
a workflow-declared custom column as a list filter. Server-side filter
surface, not a rendering decision.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:48:08 -07:00
gsxdsm
fbe7eb5c5a U7 PR1: the manual plan-approval gate was bypassable (3 planning-lane surfaces, 8/13 revert-proof) (#2491)
## What this is

The first slice of **U7 — the graph owns planning**. Characterizing the
planning lane's dual ownership turned up a live defect in the exact seam
the unit exists to remove, so this PR fixes that first and reports the
measured map of what U7 still has to move.

## The defect

The manual plan-approval gate parks a card by writing `status:
"awaiting-approval"` and **returning early** from `finalizeApprovedTask`
— before the release move. `specifyTask` then calls `onSpecifyComplete`
**unconditionally** afterwards. Three automated surfaces went on to
advance the parked card, each having re-derived its own weaker "may I
advance this?" check from `paused`/`userPaused` alone.

`isTaskBlockedOnApproval` (`packages/core/src/task-merge.ts`) already
declares itself *"the single shared predicate core and engine code must
consult before rebounding, requeuing, resuming, re-planning, or
otherwise advancing a task"*. **Measured: it had exactly one production
consumer** (`overseer-human-control-policy.ts`). Now four.

Reachable end to end for a **plan-in-place** card — one whose column
already equals the plan-review node's column (Coding (Ideas), or any
`needs-replan` revision resting in the default workflow's `todo`):

```
park at awaiting-approval
  → onSpecifyComplete fires anyway
  → a runnable plan-review continuation is seeded
  → the drain dispatches it
  → Plan Review runs on a plan the operator never approved
  → its evidence satisfies isUnplannedForExecution
  → the capacity sweep releases the card into In progress
```

Blast radius: projects that have manual plan approval switched on.
`planApprovalMode` defaults to auto-approve (FN-7557), so unset projects
have no gate to skip — but the operator who turns it on is precisely the
one who cares.

## Surface enumeration

Per AGENTS.md — fix the invariant, not the repro.

| # | Surface | Fix |
|---|---|---|
| 1 | `issueRelease` — the choke point for the sweep, `promoteHeldTask`,
`releaseHeldTaskByEvent`, and the scheduler's `reserveSlot` guard |
Guarded there rather than inside `isUnplannedForExecution`, because an
approval-held card is not "unplanned". Guarded **again** inside the
`moveTaskIf` predicate so a park landing mid-sweep cannot lose the race
(R6 — only the in-txn check is authoritative). Operator force-promote
(`allowUnplanned`) still waives it: that *is* a human decision about
this card. |
| 2 | **Both** continuation seeders —
`seedPreReleasePlanReviewContinuation` (normal completion) and
`evaluateStrandedHoldContinuation` (FN-8592 self-healing re-seed) |
Guard at the seam, not in the callers: the seeder itself checked
nothing, and its two callers each pre-checked a different subset. |
| 3 | `resolvePlanningContinuationCandidate` (drain classifier) |
**Skip, never orphan.** Cancelling terminalizes the item, so an approval
landing a minute later would have nothing left to resume and would need
a second repair to come back. |

## Measured, not assumed

The two hold shapes `isTaskBlockedOnApproval` accepts were **not equally
broken**. The `paused` + `pausedReason` shape was already refused by the
sweep and the drain — they happen to test `paused` — so it was refused
*for the wrong stated reason*, not advanced. Every genuine advance gap
is on the **status-only** shape, which is exactly what the gate writes.
Both are covered anyway, plus an `ORDINARY_PAUSE` counter-case so the
new check cannot quietly become a catch-all for every operator park.

## Revert proof

With the three production files reverted: **8 of 13 tests fail.** The 5
that still pass are the 3 controls and the 2 pause-shape rows the
pre-existing `paused` checks already covered.

```
·x··xxxxx·xx·      → Tests 8 failed | 5 passed (13)
```

## Verification

| Check | Result |
|---|---|
| new suite | 13/13 |
| hold-release (×2) + plan-review (×3) + pre-release-plan-review +
promote-force-unplanned | 43/43 |
| stranded-hold-continuation (×2) + continuation-selection +
planning-finished-wake + planning-service | 27/27 |
| scheduler-trait-dispatch | 9/9 |
| `pnpm --filter @fusion/engine exec tsc --noEmit` | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green |
| `pnpm check:changesets` | clean |

## Two findings for the coordinator

**1. `triage.ts` is absent from the Phase B census.** The plan's
per-file table (535 sites) covers `self-healing.ts` (U4), the
executor/scheduler cluster (U5), and the core policy modules (U6).
`triage.ts` appears in none of them, so its lifecycle-column literals
are unowned scope — U7 absorbs them.

Measured with the plan's own methodology (block and line comments
stripped, code lines only): a naive quoted-literal grep of `triage.ts`
reports **50** sites, but **35 of those are the agent *role* string
`"triage"`**, not the column. The genuine lifecycle-column surface is
**15 sites**, of which 12 are planning-lane and 3 are `column !==
"done"` in duplicate search. The 50 figure would over-count by 3.3×.

**2. The graph's planning seam is a rubber stamp, in triplicate.**
`createAuthoritativeWorkflowSeams().planning` returns `{ outcome:
"success", value: "pre-specified" }`;
`WorkflowPlanningService.runPlanningSession` returns the same;
`createNoopLegacySeams().planning` is a bare success. The real
specification is ~1,000 lines of `triage.specifyTask`, entirely outside
the graph. That is the flip U7's remaining slices have to make, and it
is the reason the planning lane has two owners at all.

## Deliberately not in this PR

Triage's unconditional `onSpecifyComplete` call. That is the
**ownership** half — `finalizeApprovedTask` must report whether it
released, and the reaction must key on that outcome — and it belongs
with the seam flip, where finalize's outcome becomes the graph's edge
condition anyway, rather than as a half-measure now. With the three
guards above in place, the downstream damage is already contained; what
remains is a reaction firing for a non-event and an operator-visible log
line (`Specified X → todo`) that is untrue for a parked card.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks awaiting manual plan approval are no longer automatically
planned, reviewed, started, or released into active work.
* Approval-held items are consistently skipped across planning
continuations and related workflows.
* Approval-held due work is deferred to prevent starvation while
waiting, and operator force-promotion still bypasses the gate.
* **Tests**
* Added regression coverage to ensure the manual approval hold behavior
remains invariant across multiple continuation scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:39:21 -07:00
gsxdsm
7871b28766 fix(core): bind the in-transaction capacity gate — one shared pool-id convention (NOT user-visible yet — see R2) (#2488)
## The bug

`moves.ts` asked `countActiveInCapacitySlotAsync` for occupants of pool
`"builtin:coding"`, while the counter buckets selection-less rows under
`DEFAULT_WORKFLOW_POOL_ID` (`"__default-workflow__"`). Nothing ever
landed in the pool being asked about, so the count came back **0** and a
finite limit could never bind.

## Root fix, not a literal swap

A shared *constant* would not have prevented this:
**`DEFAULT_WORKFLOW_ID` was already imported in `moves.ts` and the code
still wrote a literal.** So both sides now call a shared **function**,
`resolveCapacityPoolId` — "which pool does a selection-less task belong
to" has exactly one answer and no call site is in a position to disagree
with it.

The one variable serving two masters is split: a capacity **pool key**
(a bucketing sentinel that must not collide with a workflow id) and a
**workflow id** (telemetry, must stay a real id). The emitted
`TaskTransitioned` payload is byte-identical.

## Checked, not assumed: no second copy

`scheduler.ts:2514` and `:2536` do carry `?? "builtin:coding"` — but as
an **IR resolution key** (`resolveWorkflowIrById`), where a real
workflow id is required and the pool sentinel would not resolve at all.
Same literal, different concept, correctly used. A blanket replace would
have broken it.

## Something did depend on the gate being dead — exactly one thing

`move-path-equivalence.pg.test.ts` → *"UNPROVEN: in-transaction column
capacity did NOT reject on EITHER path in this fixture"*. It left the
cause open —

> something further in (`resolveColumnCapacity`'s limit resolution, or
what `countActiveInCapacitySlotAsync` counts as an occupant — a task
with no session/agent may not count) keeps the check from firing … This
suite does not establish which.

— and predicted its own obsolescence (*"if a future change makes this
reject, that is the capacity gate coming alive"*). **Neither guess was
right; it was the pool id.** Updated to assert the divergence with the
answer recorded — **not weakened**. Its fixture also had to start each
phase from an empty wip column: once the gate binds, the inline phase's
leftovers trip the cap on the *holder* move before the contended move
under test runs.

`schema-applier.test.ts` failed only in the full-suite run and passes in
isolation both with and without the fix — cross-file contamination, not
mine.

## Before / after — measured, both directions

`maxConcurrent: 1`, real PG store, real `moveTask`:

| | flagOFF / no selection | flagOFF / selection | flagON / no selection
| flagON / selection |
|---|---|---|---|---|
| **before** | ADMITTED | ADMITTED | **ADMITTED** ← the bug | REJECTED |
| **after** | ADMITTED | ADMITTED | **REJECTED** | REJECTED |

The E2E acceptance row asserts **held at cap 1 and admitted at cap 2 on
the same fixture**, so it cannot pass by simply never admitting
anything. **With the fix reverted that row fails**; the `admitted` case
still passes, as it should. The Phase A3 ratchet's two flipped
assertions also fail with the fix reverted.

Ratchet flipped exactly as its author specified: `DEFECT (R1)` becomes a
rejection, and `it.fails` on the invariant becomes a plain `it`.

## ⚠️ This is NOT user-visible yet — please read before merging

The premise this was approved on ("once it binds, cards that currently
slip through will start being held") **does not hold for this change
alone.** The whole capacity block sits inside `if (useWorkflow &&
workflowIr && fromColumn !== toColumn)`, and `useWorkflow` is
`experimentalFeatures.workflowColumns === true` — absent from
`DEFAULT_GLOBAL_SETTINGS`, with **no writer anywhere outside tests**.
That is Phase A3's R2, still live and now retitled `DEFECT (R2, STILL
LIVE)` with the measured matrix recorded in it.

So on merge: nothing changes for any real project. Making it actually
bind means **also** removing the `useWorkflow` condition — a materially
larger, genuinely user-visible change that I have not made unilaterally.
Escalated for a decision; if that lands, the changeset here should be
re-categorised.


## Review follow-up (48e79ffd9): the convention was still duplicated —
swept and ratcheted

The first pass added the resolver and routed the transactional gate +
counters, but **hold-release still derived the pool independently**.
Swept the repo: six sites name the sentinel, **five derive the
convention** and now call `resolveCapacityPoolId`
(`hold-release.ts:116/118/442/576`, `task-store-helpers.ts:290`). The
sixth, `scheduler.ts:1558`, names the default pool as a literal in a
capacity *diagnostic* — no selection input, nothing to disagree with —
so it keeps the constant.

**Does this change hold-release behavior? No, and it was never releasing
against the wrong pool.** hold-release computed `x ??
DEFAULT_WORKFLOW_POOL_ID`, which is exactly what the counter buckets
under; `moves.ts` (`?? "builtin:coding"`) was the sole disagreeing site,
and the first commit moved *it* into agreement with hold-release, not
the reverse. `resolveCapacityPoolId(x)` **is** `x ??
DEFAULT_WORKFLOW_POOL_ID`, so every routed site computes an identical
value for every input. **No second user-visible change rides along with
this PR** — the only behavior delta remains the gate binding on the
flag-ON path, which per R2 is still not the path production takes.
Evidence: hold-release + capacity suites **43/43 identical before and
after**.

**The resolver is now the only way to compute a pool id, not merely the
newest way.** `scripts/check-capacity-pool-id.mjs` fails on any inline
`?? DEFAULT_WORKFLOW_POOL_ID` outside `workflow-capacity.ts`, wired into
**both `pretest` and the blocking `test:gate`**. A review note would not
have sufficed: the original defect landed in a file that *already
imported* the canonical constant. Verified both ways — clean run scans
1124 files and passes; reintroducing the old hold-release expression
exits 1 and names the line.


## Review follow-up (a5b675503): the ratchet was rebuilt because it
would not have caught the bug

The first ratchet matched one spelling (`?? DEFAULT_WORKFLOW_POOL_ID`)
and the real defect used another (`?? "builtin:coding"`). **Verified:
reintroducing the original defect and running the old checker exits 0.**
A guard that reports success without checking is worse than no guard —
it stops anyone looking.

Rebuilt on the TypeScript AST with two rules. **Rule 1 (sink):** a value
reaching a capacity counter's `workflowId` must come from
`resolveCapacityPoolId`, or a local initialized from it — so it fires on
the original defect regardless of which literal was used, on one line or
twenty. **Rule 2 (sentinel):** no `??` onto the sentinel at any
qualification depth or as its raw value; multiline is one AST node and
caught by construction. `?? "builtin:coding"` is deliberately *not*
banned outright — it is the legitimate default for a *workflow* id in ~8
places, and is only a bug when it reaches a capacity pool.

**Fails closed three ways** that previously reported success without
inspecting: unreadable file, unparseable file, and an empty file listing
(the old script would have printed a green tick off a broken glob).

**Acceptance was not "passes on main".** Each form was reintroduced into
the real source and confirmed to fail: the original defect in
`moves.ts`, a multiline fallback, and a deeply qualified sentinel. All
are pinned in `capacity-pool-id-check.test.ts` (12 cases: 7 must-catch
starting with the reduced actual pre-fix `moves.ts`, 4 must-not-flag, 1
fail-closed) so the guard cannot silently narrow again.

Also added to `pretest:full`, which had omitted it.


### Follow-up (0be8df6ea): a dead rule found by fixing a test title

Splitting the mislabelled fail-closed test surfaced more than a
mislabel: **`ts.createSourceFile` is error-tolerant and does not throw
on malformed syntax**, so the `try/catch` behind the `unparseable` rule
was unreachable and that rule could never fire. The earlier "fails
closed three ways" claim was overstated — the guard advertised a
capability it did not have. Detection now reads `sf.parseDiagnostics`; a
partial AST can silently lack the `??` nodes and sink calls the rules
look for, so "did not parse" must not read as "inspected and clean".
Mutation-verified: reverting the detection fails that case and only that
case.

Test-file exclusion also moved to the repo's `{test,spec}.{ts,tsx}`
guideline shape — a `.spec.ts` under `packages/<pkg>/src/` was being
scanned as production source. Verified both ways: the `.spec.ts` is
skipped, and the identical content in a non-test file is still caught,
so the exclusion is scoped rather than a hole.

## Verification

- engine + core `tsc --noEmit` clean
- `pnpm test:gate` green (299 + 10 + 71)
- E2E 20/20; capacity + move-path suites 14/14
- full core PG: **1037 passed / 3 failed** — all three reproduce with
the fix stashed (pre-existing)
- engine-default: **279 failed** vs **280 at baseline** with the fix
stashed — pre-existing red lane, no regression
- hold-release + capacity suites: **43/43 identical before and after**
the resolver routing
- `check-capacity-pool-id` ratchet: 14/14 regression cases; clean over
1124 files; exits 1 on the original defect, a multiline fallback, and a
deeply qualified sentinel reintroduced into real source

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed capacity-limit accounting when workflow selection is missing by
consistently deriving the correct capacity pool id.
* Made capacity enforcement align across move and hold/release paths,
rejecting over-limit moves with `capacity-exhausted`.
* **Tests**
* Updated PostgreSQL and added an E2E scenario to verify the corrected
in-transaction gating behavior at `maxConcurrent` limits of 1 and 2.
* **Chores**
* Added an automated guard to detect inconsistent capacity pool id
fallback patterns in code.
* **Public API**
* Exposed `resolveCapacityPoolId` for consistent capacity pool id
derivation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:51 -07:00
gsxdsm
4158cf1ab7 Phase A: workflow-owned lifecycle foundation (U1, U2, U3) (#2467)
Phase A (Foundation) of
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`.
Three units, one commit each. No operator-visible behavior change.

## U1 — Lifecycle-column resolution seam

`resolveLifecycleColumns(ir)` returns `{ intake, hold, wip, review,
complete, archived }` — the first column carrying each trait,
`undefined` for a role no column carries.
`resolveTaskLifecycleColumns(store, taskId, cache?)` is the store-aware
form; the cache is caller-owned so a sweep reads one IR per workflow
rather than one per card.

A v1/column-less IR resolves to `undefined` for the **whole struct**
rather than a struct of undefined roles. A caller must be able to
distinguish "this workflow declares no hold column" (a real shape to
honor) from "no column vocabulary at all" (skip and log) — only the
second licenses conservative fallback.

Nothing consumes the seam yet; Phases B–D convert the ~207 hardcoded
column literals onto it.

## U2 — Delete the pre-cutover parity machinery (delete-only)

**`workflow-columns-settings.ts`** — `isWorkflowColumnsEnabled` had the
body `return true`. Six live call sites branched on it, so every
flag-OFF arm was dead code that read as a supported configuration.
Deleted; surviving side inlined at self-healing's transitionPending
sweep, the scheduler's per-column capacity diagnostic, merge-trait's
policy resolver, the board-workflows payload, two task-workflow routes,
and the CLI TUI's column enrichment.

**`workflow-parity.ts`** — asserted the default workflow's adjacency
*equals* the legacy `VALID_TRANSITIONS`. U11 deliberately breaks that
equality by merging Todo into Planning, so this is not a stale assertion
to update; it is a contract against the target state. Its emitter
(`workflow-parity-observer.ts`) is already a tombstone, so
`getWorkflowParitySummary` and `computeWorkflowColumnsGraduationReport`
aggregated run-audit rows nothing writes and had no caller outside
`TaskStore`. Both store methods go with it.

`flagEnabled` stays on the board-workflows **wire** as a constant `true`
— shipped dashboard clients still branch on it, and changing the
response shape is not a deletion. U10 retires the field once no client
reads it.

The `legacy-tombstones` ratchet is extended to both files plus seven
symbols, each with the reason it is gone.

### ⚠️ Finding: the third listed deletion was NOT dead

The plan also lists "the flag-off inline move path" in
`task-store/moves.ts`. It is **not** deleted, per U2's execution note
("any behavior change found while removing a branch means the branch was
not dead").

That path is gated on `isWorkflowColumnsCompatibilityFlagEnabled`
(`store.ts:38`) — a **different** function from the always-true public
helper. It reads the raw `experimentalFeatures.workflowColumns` setting,
which nothing in production sets (`settings-schema.ts:396` — "no default
flags are emitted"; zero non-test writers; the operator's own
`~/.fusion/settings.json` has no such key). So `useWorkflow` is false
for effectively every real project: the flag-OFF inline side effects are
the **live** default move path and the flag-ON `default-workflow-hooks`
path is the dead one. The code says so itself at `moves.ts:638`.

Deleting that branch would swap every project onto an untravelled code
path — a behavior change, not a deletion.

**Carry this into Phases B and C, stated plainly so the plan's error is
not repeated:**

> **The inline move path in `moves.ts` is LIVE.
`default-workflow-hooks.ts` (the trait-hook path) is DEAD.** KTD-6
asserted the inverse. Until the convergence unit lands, **nothing may
assume trait hooks run** — a guard, sweep, or subscriber written against
`applyDefaultWorkflowMoveEffects` would never fire in production and
would still pass its tests.

Convergence is **not** attempted here. It is its own unit (Phase A2)
with a proper equivalence proof, per operator decision.

### U3's emit point is on the LIVE path — the seam is not born dead

Worth stating explicitly because it is the failure mode that would make
every later subscriber silently never fire: the `TaskTransitioned` emit
is **not** inside the `if (useWorkflow)` branch. That block closes at
`moves.ts:1212`; the emit sits at `:1214`, beside the existing
`store.emit("task:moved", …)`, on the unconditional post-commit path. It
therefore fires on **both** the live inline path and the dead hooks
path, and the convergence unit inherits the obligation to keep it firing
on whichever path survives — same events, same order, same payloads.

The graph-side emitters (`NodeEntered`, `RunSuspended`) carry the same
risk from a different direction: the bus refuses an invalid payload
*silently* by design, so an emitter regression would stop the event with
no test failure. They are asserted end-to-end through the real bus —
"did a subscriber actually receive it", not "was emit called" — because
a spy passes on a refused payload. The `moveTaskInternalImpl` emit does
**not** yet have that end-to-end assertion against a real store move;
that proof belongs to the convergence unit, which has to build the
both-paths fixture anyway.

## U3 — Post-commit event seam with a transactional outbox

**The bus is not a queue, not a transaction participant, and not a
delivery guarantee.** Durable follow-on work uses the transactional
outbox — a `workflow_work_items` row written *inside* the transition
transaction (the shape `createCompletionHandoffWorkflowWork` already
uses). "Emit after commit, let a subscriber enqueue the work" has a
crash window where a process dies between commit and subscriber, leaving
no event *and* no work-item row, so required work is skipped permanently
with nothing to recover from. Post-commit subscribers therefore carry
only losable reactions.

Emission is consequently lossy and isolated by design: a throwing or
rejecting subscriber is caught and logged, cannot roll back the
transition, and cannot stop the others. Deliveries append to one serial
chain, so two transitions on a task deliver in commit order.

The ids/outcomes-only rule is **mechanised, not documented** —
run-audit's equivalent lives only in prose and has been violated
repeatedly. A payload carrying an object body or a prose string is
refused at the emit boundary and never reaches a subscriber or log sink.
It degrades rather than throws: the emitter is post-commit, so a shape
bug must not become a lifecycle failure.

Emit points: `TaskTransitioned` from the single post-commit point in
`moveTaskInternalImpl`; `NodeEntered` and `RunSuspended` from the graph
column boundary, the latter *after* the durable continuation is
persisted so an observed suspension implies a resumable run.

`registerWorkflowEventSubscribers` (engine) is empty on purpose —
U7/U8/U10 move real reactions onto it, each with the characterization
test proving the reaction was non-authoritative first.

## Verification

- `pnpm test:gate` — green (2/10, 16/299, 1/71).
- `pnpm lint`, `pnpm build`, `tsc --noEmit` on core and engine — green.
- U1: 20 tests in `workflow-lifecycle-traits.test.ts`, including the
fully-renamed-workflow case (fails if the resolver falls back to a
literal) and a shared-cache read-count assertion.
- U2: `legacy-tombstones.test.ts` green with the extended ratchet;
`board-workflows`, `merge-trait`, `workflow-graph-executor-parity`, and
move-hook suites green with no expectation edits.
- U3: 20 bus-invariant unit tests (isolation, ordering, the allowed-key
and required-key halves of the ids-only rule, lossiness) plus 3
end-to-end emitter-delivery tests; 5 outbox tests against a **real
PostgreSQL** work-item table (crash survival, rollback, at-least-once
redelivery on lease expiry, idempotent handler → one effect,
dropped-subscriber vs. durable work). A hand-written fake of the lease
predicate would only prove the fake redelivers.

**Not verified:** the `moveTaskInternalImpl` emit is confirmed on the
unconditional post-commit path by structure and by the surrounding
tests, but is *not* yet asserted end-to-end against a real store move on
both flag settings — that is Phase A2's fixture. The engine subscriber
registry ships empty by design, so no production subscriber exercises
the bus end-to-end yet. `settings-defaults.test.ts` has one pre-existing
failure on `main` (a logger-prefix mismatch in the
`mergeIntegrationWorktree=cwd-main` warning) — confirmed present on a
clean tree, unrelated to this branch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Workflow lifecycle columns are now derived from workflow definitions,
supporting renamed and custom workflows.
* Added post-commit lifecycle events for task transitions, node entry,
and run suspend/resume with validated payloads.
* Follow-on processing for lifecycle emissions is now more robust
(rollback-safe, at-least-once delivery, idempotent handling).
* **Bug Fixes**
* Workflow board responses, task enrichment, and promotion no longer
depend on workflow-columns feature-flag gating.
  * Subscriber failures no longer impact committed workflow transitions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 13:30:13 -07:00
gsxdsm
8b039a543e fix(desktop): advance Pi runtime pin to 0.82.1 for packaging PR lane (#2465)
## Summary
- Advance the matched Pi runtime pin (`pi-ai`, `pi-coding-agent`,
`pi-agent-core`, `pi-tui`) from **0.82.0 → 0.82.1** so
electron-builder's production-dependency walk accepts `pi-agent-core`'s
`pi-ai@^0.82.1` requirement.
- Fixes the Desktop packaging PR-lane failure:
`Production dependency @earendil-works/pi-ai not found for package
@earendil-works/pi-agent-core` (required `^0.82.1`).
- Keep the workspace override guard; update pin-policy fixtures and CLI
package-config expectations.
- Tighten the advisory packaging step-order test so it asserts against
the real `electron-builder --dir` step (not a missing release-only step
name that previously passed via `indexOf === -1`).
- Run `pnpm dedupe` so the packaging lane's lockfile dedupe
early-warning is clean.

## Context
#2439 pinned the full Pi closure at 0.82.0 and made recent main-based
packaging runs green. This advances to the current upstream patch so
deploy + electron-builder stay aligned with `pi-agent-core@0.82.1`'s
declared dependency range.

## Test plan
- [x] `node scripts/check-pi-versions-pinned.mjs`
- [x] `node --test scripts/__tests__/check-pi-versions-pinned.test.mjs`
- [x] `pnpm --filter @runfusion/fusion exec vitest run
src/__tests__/package-config.test.ts`
- [x] `pnpm --filter @fusion/desktop exec vitest run
src/__tests__/release-workflow.test.ts`
- [x] `pnpm dedupe --check`
- [ ] GitHub: Desktop packaging (should run full packaging walk —
lockfile/package.json touched)
- [ ] GitHub: PR Checks (Lint, Typecheck, Build, Gate)
2026-07-26 23:47:49 -07:00
gsxdsm
99c9f14ee0 feat: run Plan Review in the planning lane with a Plan Review badge (#2462)
## What

Plan Review, planning, and the replan loop move from the implementation
column into the **planning lane** (`todo`), so a task under
specification never holds a WIP slot. The card crosses into
`in-progress` exactly once, at `parse`, released by the scheduler.

Operators also finally see a **Plan Review** badge while the gate runs —
it was previously invisible on the default workflow.

## The part that made it possible

Moving the node is ten lines. It was attempted three times and reverted
each time, because a graph run with no durable continuation replayed
from `start` and dragged an in-progress card *backward* out of the WIP
column, firing `abort-on-exit` and stranding it in a pre-WIP column with
no releaser.

So this PR adds the graph **entry contract** —
`resolveColumnResumeNode`:

| Card is in | Resumes at |
|---|---|
| `triage` | `start` |
| `todo` | `plan` |
| `in-progress` | `parse` — never re-plans, never moves backward |
| `in-review` | first review node — gates are not skipped |

`ir.columns` is ordered and that order is the lifecycle order; rework
and failure edges are excluded so the entry point is always the main
path. The proof it's the right fix: **`executor-task-done-invariant`
passes unmodified** after failing every previous attempt.

## Also in here

- **Release gate narrowed twice.** `isUnplannedForExecution` applies its
pre-release plan-review gate only when the node's column equals the
card's column *and* the group is enabled for the task. The enablement
check fixes a real deadlock — a task with Plan Review toggled off was
held forever waiting for evidence nothing would ever write.
- **Badge cleanup.** Gate badge reads "Plan Review" instead of the
ambiguous "Reviewing" and no longer hides behind a lane restriction; the
status badge stops duplicating it; `planning` renders as "Planning"
instead of the raw engine token.
- **Coding (Ideas)** renames its planner column to "Planning" (id `todo`
unchanged) and loses its private planning-node re-home — the graph it
clones is already plan-in-place.
- **New sweep** `reconcileUndeclaredTaskColumns` re-homes a row whose
column its workflow no longer declares. Written for a follow-up, kept
because it makes any column edit survivable.

## Test changes

Scheduler and release fixtures now model a card whose Plan Review passed
— the state every real card is in when the capacity sweep sees it. A
held unreviewed card is the gate working, and that path stays owned by
`pre-release-plan-review.test.ts`.

New `workflow-graph-entry-contract.test.ts` covers the invariant at
every lifecycle position, plus the gap-column and remediation-node
cases.

## Verification

Gate 299 + 70 + 10, dashboard badge suites 672, engine
workflow/entry/executor suites 147, core 122. Lint and typecheck clean.
Full engine suite sits at the pre-existing baseline (notifier /
plugin-runner / notification-service, untouched by this).

## Follow-up

Removing the Todo column entirely is a separate ~207-site
lifecycle-vocabulary refactor — planned in
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`
(companion docs PR).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Plan Review now runs in the Planning lane before implementation
begins.
* Cards resume from their current workflow column without replaying
earlier steps.
* Added automatic recovery for cards stranded in outdated workflow
columns.
* **Improvements**
  * Renamed the Coding (Ideas) planner column to “Planning.”
* Refined Plan Review gating to respect enabled settings and the card’s
current column.
* Updated planning and Plan Review badges for clearer, consistent labels
across cards and lists.
* **Bug Fixes**
* Improved workflow transitions and release behavior around planning,
review, and execution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 23:42:46 -07:00