Commit Graph

3577 Commits

Author SHA1 Message Date
Phil Larson
72391c90b2 fix(engine): route workflow reviews through validator models (#2533)
## Summary

- classify review-type workflow steps with the existing review-step
classifier
- resolve their primary, fallback, and thinking-level settings from the
validator model lane
- retain per-step model overrides and executor-purpose workflow-step
tooling
- keep ordinary workflow steps on the execution lane
- make missing-fallback diagnostics identify the correct lane

## Why

Code Review, Plan Review, verification, and inline-review gates were
executed through the implementation model lane merely because they run
inside `executeWorkflowStep()`. That defeats configured reviewer-model
separation and can make the same model implement and validate its own
work.

This changes model selection—not the workflow-step session/tooling
contract—so review steps remain executor-purpose sessions while using
validator lane models.

## Verification

- `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/engine
exec vitest run src/__tests__/executor-workflow-step-model.test.ts` — 14
passed
- `corepack pnpm@10.33.0 --filter @fusion/engine typecheck`
- `corepack pnpm@10.33.0 changeset status --since=origin/main`
- `git diff --check origin/main...HEAD`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Review-type workflow steps now route through the configured validator
model lane (instead of the execution lane).
* Validator primary/fallback and thinking-level settings are applied
correctly for review steps.
  * Step/task overrides still take priority over lane-based resolution.
* Fallback retry sessions now use the appropriate validator/executor
configuration, with lane-specific fallback guidance when fallback
settings are missing.
* **Tests**
* Expanded executor workflow-step model resolution and routing/fallback
precedence assertions for validator-lane behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:05:04 -07:00
gsxdsm
18d654a5ff capacity, part 3: delete the globalMaxConcurrent setting, API and UI (#2529)
Part 3 of the capacity simplification, and the half that removes the
**knob**. Enforcement (shared semaphore, runtime wiring) went in #2509;
this removes everything an operator or API client can still see, so
nothing is left readable-but-ignored.

## Deleted

Settings key + schema default · CentralCore’s
`getGlobalConcurrencyState` / `updateGlobalConcurrency` /
`acquireGlobalSlot` / `releaseGlobalSlot` and the `concurrency:changed`
event · the whole Global Concurrency block in `async-central-core` ·
`PUT /api/global-concurrency` · the Scheduling · Global settings section
· the footer and Command Center global sliders · the dead
`getGlobalConcurrencyLimit` reader whose only caller went in #2509.

## Kept, deliberately

**`GET /api/global-concurrency` survives as telemetry only** — live
`currentlyActive` / `projectsActive` from CentralCore’s side-effect-safe
source. “How busy is this machine?” is still a real question once the
cap that used to answer it is gone. It no longer reports
`globalMaxConcurrent`/`queuedCount`: those came from the deleted cap and
from slot bookkeeping production code never incremented, so publishing
them was publishing zeros dressed as state.

**`useGlobalConcurrency` becomes read-only.** Everything that existed to
*persist* went with the cap — the 500 ms debounce, the save-state
machine, the commit-on-close/unmount flush, the slider clamp, the
`interactive` gate. The module-level shared store is **kept**: its
original justification (two mounted consumers drift apart with private
copies) holds for a polled read exactly as it did for a cap, and one
fetch now serves both.

The live “N running (all projects)” readout survives in both surfaces,
moved onto the per-project row.

## Two sections become one

Scheduling · Global existed to host exactly one control. With it deleted
the section renders an empty pane, so the Global/Project pair merges
back into **“Scheduling”**. An empty nav entry is a promise of settings
that are not there.

## One real fix found on the way

`SchedulingSection`’s `concurrencyLoading` gated the **project**
concurrency inputs on the **global**-concurrency fetch — never the right
source, since `maxConcurrent` and `maxWorktrees` come from the settings
form. It is repointed at the form’s own load, preserving the invariant
it existed for: a concurrency input stays disabled until its live value
arrives, so an operator cannot overwrite a resolved limit with a blank
fallback.

## Migration

A stored `globalMaxConcurrent` is **ignored** — it is a project-blob key
nothing reads, so dropping it needs no schema change. The
`central.global_concurrency` **table** is dropped in a follow-up; this
slice stops seeding and reading it first, so that drop has no live
writer to race.

## Verification, and how the wider suite was controlled

`pnpm lint` clean · core/engine/dashboard `tsc` clean · `pnpm test:gate`
green (309 + 10 + 71) · dashboard settings/footer/command-center/hooks
**2237/2237** · core `central-core-backend` 9/9.

The broader dashboard suite shows failures, and I checked rather than
assumed: running the suspect files on **clean main** reproduces
`api-git` (49), `TaskDetailModal.rendering` (28) and `settings-mobile`
(17) identically. Two were genuinely mine —
`SettingsModal.scheduling-merge` (0 on main, 17 on this branch: my nav
rename) and one `settings-mobile` picker case asserting `scheduling` is
a scoped pair — and both are fixed.

Tests for deleted behaviour are removed with it (footer
confirm/cancel/flush/dedupe, global marker geometry, the hook’s PUT
case, the CentralCore slot cases), each carrying a note on what it
guarded and where the surviving **project-side** equivalent lives.
Fixture-only references were updated, not deleted.

Nothing booted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 23:23:13 -07:00
gsxdsm
919f68f9bc test(U9): cover the two unguarded merge safeguards and admit them to the gate (#2526)
**U9, PR5.** Closes the gap #2520 measured. Tests + gate config only; no
production behavior change.

## The gap

#2520 found that safeguards **1 (user pause)** and **4 (capacity
single-flight)** had **zero test coverage**. Deleting either guard
produced no new failure anywhere in the merge, project-engine,
self-healing, or concurrency suites. Both guards work correctly today —
nothing would have noticed if they stopped. U9 moves merge behind graph
nodes, so this is exactly the state not to convert on top of.

## Two tests

- **`merge admission excludes a user-paused card`** — safeguard 1, the
pause invariant re-ratified in #2486. Without the `paused || userPaused`
filter, the admission provider offers a user-paused card to the merge
pump.
- **`drainMergeQueue is single-flight`** — safeguard 4. Asserted via
`reconcileStaleMergeActive`, the first statement *inside* the guard, so
the probe isolates the guard rather than dispatching a real merge.
(Driving a real drain crashed the vitest worker; probing the guard
directly is both safer and more precise.)

**Both are two-sided** — they assert the guard blocks *and* permits. A
one-sided test would still pass against a guard that rejects everything,
which is a real failure mode for a filter.

## Proven by mutation delta

Baseline fail-set vs mutated fail-set on the identical selection, NEW
failures only:

| Mutation | NEW failures |
|---|---|
| remove the pause filter | **1** — the pause test, and only it |
| remove the single-flight guard | **1** — the single-flight test, and
only it |
| filter rejects *everything* | **1** — proves not one-sided |
| drain *always* refuses | **1** — proves not one-sided |

## Gate admission

`project-engine.test.ts` joins the `engine-core` allow-list. **One file
proves five safeguards** — user pause, `autoMerge:false`, capacity
single-flight, the pre-enqueue merge-proof consult, and at-most-once
enqueue.

Before this, **none of the six safeguards was defended by blocking CI**.
A regression surfaced only in non-blocking full-suite, after the merge.

Measured, not assumed:

| | Files | Tests | Wall (3 runs) |
|---|---|---|---|
| before | 17 | 309 | 5.19 / 5.51 / 5.19s |
| after | 18 | 412 | 6.19 / 6.24 / 6.21s |

**+~1.0s against a ~60s ceiling.**

**Verified the gate fires**, rather than assuming the allow-list edit
took — the failure mode greptile caught in #2494:

- remove safeguard 1 → `pnpm test:gate` **exits 1** (1 failed / 411
passed)
- remove safeguard 4 → **exits 1** likewise
- restored → **exits 0**

Deterministic: store, runtime, merger and notifier all mocked; no real
git, no network, no real timers in these two cases.

## Reversible calls I made rather than asking

- **Added to `project-engine.test.ts` rather than a new file.** A
dedicated file would need ~200 lines of duplicated `vi.mock`
scaffolding; reusing the existing harness also means one gate admission
covers five safeguards instead of two.
- **Did not wait for U8.** These guard code that exists today and the
conversion needs them in place first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:20:04 -07:00
gsxdsm
e9bfd0d313 test(engine): prove the admission-control agent count on a renamed board — and two of my own cases were vacuous (#2516)
Test-only. Fifth E2E family. Closes three of the four `live-agent-count`
classifications.

## Why this one is a scheduler bug, not a display bug

`live-agent-count.ts` classifies a card's column, and
`persistedTopLevelAgentSlotsFromStore` turns that into **the number
admission control compares against the cap**. So a mis-classified column
fails in whichever direction hurts:

| mis-classification | consequence |
|---|---|
| wip column not recognised | under-count → **over-admits past the
operator's cap** |
| complete column not recognised | a finished card counts forever →
**board silently stalls** |

Both are silent, and both land only on a renamed board.

Everything in the path is real: PostgreSQL store, real persisted
workflows, cards walked through the real transition policy, and the real
counting function resolving each card's own IR. Nothing about counting
is reimplemented here.

## Two of my own cases were vacuous — mutation-testing caught it

This is the more useful half of the PR.

**1. "does not count a card in the COMPLETE column" passed with the
terminal classification hardcoded to `done`.** `isRunningAgentTask`
rejects that card at the *wip* check anyway, so the test was really
asserting "shipped isn't a wip column". `terminalKind` short-circuits
**first**, so it only changes the answer for a card whose status would
otherwise make it count. Now covered by a complete card carrying a
live-looking `planning` status — the state a crashed run leaves behind,
which on a renamed board consumes a slot forever.

**2. The review/merge lane had no case at all.** A review status is
deliberately *not* globally live (a stale `fixing` in wip must not
consume capacity), so it is gated on `columnIsReviewOrMerge`. If the
renamed review lane isn't recognised, a genuinely-active reviewer stops
counting and admission control lets another agent in over the cap. Now
covered by a review card with an active merge-pipeline status.

Both new cases assert their fixture took effect first, so they can't
degrade back into the weaker version silently.

## Mutation-verified independently

| classification | mutation | result |
|---|---|---|
| `countsTowardWip` | → `"in-progress"` | 3 renamed cases fail |
| `complete` | → `id === "done"` | exactly the new terminal case fails |
| `mergeBlocker` | → `id === "in-review"` | exactly the new review case
fails |

## What this does NOT cover, stated plainly

The fourth classification, `columnIsIntakeOrHold`, is read only by the
**waiting** predicate, which the admission count never calls. It stays
in the ledger as unproven rather than being claimed by proximity — the
mistake I made last slice with `resolveMergeOrchestrationColumn`.

## Verification

- five live-E2E suites green together: **52/52**
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (309 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added end-to-end coverage for live agent-count admission across
workflow lanes and lifecycle states.
* Verified slot handling for active, completed, held, and mid-review
tasks, including stale statuses.
* Confirmed mixed-lane counts and renamed board vocabularies produce
consistent results.
* **Documentation**
* Expanded coverage notes for live agent-count classifications and
waiting-state behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 22:19:22 -07:00
gsxdsm
2e39763930 test(engine): prove agent-link hygiene on a renamed board — a leaked agent slot, not a stale link (#2514)
Stacked on #2510. Test-only. Closes the `task-agent-sync` ledger entry.

## The defect this reproduces, in the code's own words

`task-agent-sync.ts`'s conversion note:

> a move into a renamed terminal column matched nothing and this handler
returned early — so the agent kept a `taskId` pointing at a finished
card and stayed `running`, **with no error and no failing test**.

"No error and no failing test" is the whole problem — and the cost is
not a stale link. **The scheduler counts `running` agents against its
cap**, so on a renamed board every completed task permanently consumes
an agent slot until a human notices. A board would just get slower and
slower.

## Everything in the path is real

Real PostgreSQL `TaskStore`, real `AgentStore`, the real
`attachAgentLinkSync` subscribed to the store's real `task:moved` event
(the same call `in-process-runtime` makes), and a real `moveTask` to
trigger it. Assertions read the **agent row** back out of the store —
never "the handler was called".

## Mutation-verified

Forcing the legacy literal sets (the pre-conversion behavior) fails
**exactly the two renamed cases**, leaving the default-vocabulary floor
and both negatives green. So this reproduces the original defect rather
than merely covering the file.

## Negative half

An ordinary mid-lifecycle move (`wip → review`) must **not** release the
agent — given the same time to run as the positive case. "Clear the link
whenever the card moves" would drop the binding the moment work started,
a louder failure than the leak it fixes.

## Two anti-flake, anti-vacuity details

- **Async delivery.** `task:moved` is a plain EventEmitter and the
handler is async, so the assertions **poll the persisted row** to a
bounded deadline and fail with the row's actual contents. A fixed sleep
would flake in both directions.
- **The fixture asserts itself.** The link and `running` state are
verified *before* the move, so an agent that was never linked cannot
make this pass for the wrong reason — the failure mode I hit twice
already in this program.

## Verification

- four live-E2E suites green together: **39/39**
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Improved agent-link cleanup when tasks reach workflow completion.
* Ensured completed task links are released even when workflow columns
have been renamed.
* Preserved active agent links when tasks move through non-terminal
workflow stages.
* Improved reporting of link cleanup outcomes and handling of
synchronization errors.

* **Tests**
* Added live PostgreSQL end-to-end coverage for completion, in-progress
moves, and renamed-column scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:59 -07:00
gsxdsm
3badc244a7 U12 part 2: bind the three U5 reconciliation guards — USER-VISIBLE (and one path that couldn't run under PostgreSQL at all) (#2512)
## U12 part 2 — the three U5 reconciliation guards now actually fire

USER-VISIBLE. Taken on standing authority; here is exactly what changed
for operators.

All three read the RAW `experimentalFeatures.workflowColumns` key via
`store.workflowColumnsFlagOn()`. Nothing in production writes it, so all
three have been inert since the workflow-columns cutover.

| Guard | Before (every real project) | After |
|---|---|---|
| Workflow edit removing an **occupied** column | Save succeeded; cards
left in a column the workflow no longer declares | Save fails with
`OccupiedColumnsError` unless `rehomeTo` is supplied |
| Workflow **delete** | Occupant capture returned `[]`; cards sat in the
deleted workflow's columns until the next engine start | Cards move to
the default workflow's entry column as part of the delete |
| Workflow **switch** | Never reconciled; the `reconciliation` field in
the declared return type was never populated | Card in an undeclared
column moves to the resolved target; a declared column is preserved |

Both consumers already handle the new outcomes and needed no change:
`register-workflow-routes.ts` maps `OccupiedColumnsError` to a
structured 409 carrying per-column occupant counts, and
`fn_workflow_update` returns a retryable structured result. The
dashboard editor's `rehomeTo` retry flow becomes reachable for the first
time. I only updated two stale "flag-ON" comments there — that code was
correct all along and simply never fired.

### What an operator actually sees (USER-VISIBLE — read this bit)

Four changes to what the board and the API do. Nothing here is silent.

1. **Editing a workflow to remove a column that has cards in it now
FAILS.** Previously the save succeeded and the cards were left in a
column their workflow no longer declared. The dashboard shows the
existing 409 with per-column occupant counts and prompts for a re-home
target; retrying with `rehomeTo` moves the cards and saves. Removing an
EMPTY column is unaffected.
2. **Deleting a workflow moves its cards immediately** to the default
workflow's entry column, instead of leaving them until the next engine
start.
3. **Switching a task's workflow moves the card** when the new workflow
does not declare its current column. A card whose column IS declared
stays exactly where it is. The API response now carries the
`reconciliation` summary it always promised.
4. **A switch whose re-home would be REJECTED is now refused before
anything is written.** If the destination column is at its WIP limit,
the switch fails with a structured 409 (`workflow-switch-rehome-failed`)
naming the task, both columns and the reason — and **nothing changes**:
the task keeps its current workflow AND its current column. Retry after
making room. Previously this combination committed the selection and
then silently reported a move that never happened, leaving selection and
column disagreeing.

**Can a torn card still happen? Yes, in one narrow case, and here is how
you recover.** If the destination fills in the window between the
pre-flight and the move, the selection is already committed and the card
ends up in a column its new workflow does not declare. That case is not
silent: it writes a `task:workflow-switch-torn` run-audit row, and the
error carries `selectionCommitted: true` with both columns. Recovery:
make room in the destination and move the card there, or switch the task
back — and if neither happens, the R7 startup sweep
`reconcileUndeclaredTaskColumns` re-homes it on the next engine start.
The card is never lost; it is visible in a lane the board may not draw
until one of those runs.

The one thing to watch after merge: (1) converts a previously-silent
success into a visible failure, so an operator mid-edit on a busy
workflow will start seeing a 409 they never saw before. That is the
point — the alternative was stranding their cards — but it is the change
most likely to generate a "this used to work" report.

### The thing that made this more than a gate removal

Un-gating the switch guard surfaced that
`selectTaskWorkflowAndReconcileImpl` read the task through
`store.readTaskFromDb` — the **synchronous SQLite** reader, which throws
under PostgreSQL:

```
TaskStore.db: SQLite Database is not available in backend mode
```

The flag returned before that line, so the gate was hiding a path that
**could not execute at all in the production backend**, not merely a
disabled feature. Ported to the async `readTaskRow`. Found by the new
tests, not by reading the code.

### Review round 2 (both findings real, both fixed)

**Torn write with no alarm — fixed by ORDERING, not by a louder
message.** My first attempt only made the error loud, which left the
torn state intact. The real fix is that the deterministic rejection
cause (destination at its WIP limit) is now checked BEFORE
`selectTaskWorkflow` commits, by resolving the target IR straight from
`workflowId` instead of through the task's selection. Nothing commits on
that path.

For the residual race the failure is loud AND recorded: `rehomeOccupant`
now returns `{ moved, error? }` (additive; sweep callers ignore it), the
switch writes a `task:workflow-switch-torn` run-audit row, and throws
`WorkflowSwitchRehomeFailedError` with `committed: true`. Consumers
translate it: the dashboard route returns a structured 409 with
`selectionCommitted`, and `fn_task_set_workflow` returns the same fields
— no more generic "something went wrong".

**Fabricated column for a deleted task.** My first fix fell back to
`fromColumn` when the final read found no row, so a task soft-deleted
mid-switch was reported as having its old column *preserved*. Absent now
reads as absent (the optional `reconciliation` is omitted). Extracted as
the pure `buildSwitchReconciliation` seam because the window is not
reachable through the public call — `selectTaskWorkflow` rejects an
already-deleted task up front — so it is a genuine race, and I test the
decision directly rather than asserting it from reading the code.

### Revert-proof, measured

New `workflow-reconciliation-production-shape.pg.test.ts` — 6 cases,
with the flag **never written**, which is the configuration every real
project has. Each flip reverted individually:

- re-gate the edit guard → **2 failures** (OccupiedColumnsError case;
rehomeTo re-home case)
- re-gate the delete capture → **1 failure** (card stays in
`custom-hold`)
- restore the switch early return → **2 failures** (`reconciliation`
undefined; card does not move)
- all three in place → **6/6 green**

Round-2 fixes, also measured:
- restore the `fromColumn` fallback → the "row is gone" case fails
(reports `preserved: true` for a deleted task)
- drop the `!outcome.moved` throw → the capacity-blocked case fails
(resolves instead of raising)
- **move the capacity pre-flight back AFTER the commit → the case fails
on the SELECTION assertion** (expected `WF-002`, received `WF-001`),
i.e. it proves the ordering, not the wording

The pre-existing coverage in `workflow-authoritative-reads.pg.test.ts`
reached the occupied-column guard by **writing the flag ON itself** —
same pattern as the ListView/Board suites in part 1. Its flag write is
removed; it now runs in the production shape.

### Where I nearly got this wrong

My first revert harness was buggy and I briefly concluded the delete
re-home was **redundant** — I had probed the stored column and seen
`triage` with what I thought was the flip reverted. It wasn't.
`workflow-ops.ts` contains two identical `const occupantTaskIds = await
store.listWorkflowOccupantTaskIds(id, false)` lines (field-reconcile
block, delete path), so my first-match edit reverted the wrong one.
Re-run anchored on surrounding context, the delete case fails as
predicted. Recorded in the test header as a caution. I also chased and
**refuted** a scarier hypothesis along the way — that an unrelated
`updateTask` coerces a custom column back to `triage`. It does not; the
column survives.

### Deliberately NOT in this PR

The v1-IR rollback-compat persistence (`downgradeIrToV1IfPure`) on the
workflow UPDATE path. It shared the same `flagOn` variable, which is how
it surfaced: **one flag read was feeding two unrelated decisions, so the
flag has more decision sites than call sites** — my earlier 9-site
inventory undercounted. It chooses the stored *shape* of the graph
rather than gating a guard, so it is a persistence-format change with a
different blast radius. It now reads the flag explicitly, behaviour
unchanged, for a follow-up.

The `moves.ts` group remains U2b's.

### Verification

`pnpm test:gate` (307 + 10 + 71), `pnpm lint`, `pnpm verify:fast` (17
steps), both typechecks green. Full `packages/core` PostgreSQL suite:
**1042 passed, 3 failed** — `central-archive-secrets.test.ts`
(log-prefix assertion) and
`workflow-settings-project-identity.pg.test.ts` (×2, project-id
resolution). I confirmed the identical 3 failures on a stashed clean
tree: pre-existing, unrelated. No Fusion instance booted.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Workflow edits now prevent removal of occupied columns unless cards
are moved to a specified destination.
* Cards are automatically re-homed when workflows are deleted or
switched.
* Workflow switches now check destination capacity before committing and
provide clear conflict details when re-homing fails.
* Reconciliation results now indicate whether cards were moved or
preserved.


<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:46 -07:00
gsxdsm
3bbb6ffc6b capacity, part 2: delete the cross-project concurrency cap (enforcement half) (#2509)
Stacked on #2502 — review that first; this branch contains its three
commits.

Operator: two capacities **per project**. `globalMaxConcurrent` is a
machine-wide *third* limiter kept in a separate authority (a central-DB
singleton row) that every runtime had to subscribe to and periodically
re-reconcile. It goes.

This slice removes **enforcement and wiring only**. The setting key,
central DB state, API route and Settings UI come out in part 3, so each
half lands green and independently revertable.

**Deleted:** the shared `AgentSemaphore` instance in `ProjectManager`
and `ProjectEngineManager`; the per-project `ScopedAgentSemaphore` in
`InProcessRuntime`; the `globalSemaphore` runtime-config field; both
`concurrency:changed` subscriptions; ProjectManager’s 30s limit-refresh
poll; the residual-slot return on project stop.

The scheduler/triage semaphore gate is now simply **absent** — same
shape as the worktrees-off gate in #2502. `semaphoreGate?` was already
optional, so no gate object is constructed rather than one holding an
infinite limit. Absence cannot start binding again by accident.

---

## Two findings that changed the shape of this slice

**1. `AgentSemaphore` the class stays — my earlier estimate was wrong
and I withdraw it.**

I previously told the coordinator that ~75% of `concurrency.ts` (≈662 of
886 lines) was semaphore machinery that could go with this cap. That was
line-range arithmetic, and it was wrong. `AgentSemaphore` is a general
primitive with four consumers unrelated to the global cap:

| Consumer | Governs |
|---|---|
| `verification-concurrency.ts` | `maxConcurrentVerifications` |
| `research-orchestrator.ts` | research `maxConcurrentRuns` |
| `experiment-executor.ts` | `maxConcurrentExperiments` — **a knob
absent from my original inventory** |
| `step-session-executor.ts` | parallel workflow steps |

What goes is the global **instance** and its wiring, not the class. I
will report the measured `concurrency.ts` delta after part 3 rather than
repeat an estimate.

**2. `acquireGlobalSlot` / `releaseGlobalSlot` had no production callers
— only tests.**

So the cross-project cap had *two* mechanisms: the in-memory semaphore
(live) and a durable central-DB `currentlyActive` counter (dead — never
incremented by real work). Both deleted, along with the tests that
pinned the dead passthrough.

## The regression this almost introduced

`runWithMergeAdmission` in `project-engine.ts` opened with:

```ts
if (!semaphore) return await start();
```

Unreachable while a global semaphore always existed. With the semaphore
gone it would have fired on **every** merge and skipped
`projectAdmissionCoordinator.admitOldest` entirely — silently stopping
merges from counting against the **per-project** agent count.

That is the opposite of the intent: a merge *is* an agent and still
consumes one of the project’s slots; it just no longer consumes a
machine-wide one. So the early return is **deleted rather than left to
fire**. `admitOldest` already declares `semaphore` as optional and
enforces `maxConcurrent` independently of it (`claimed() + reservations
>= maxConcurrent`), so dropping the argument preserves per-project
admission and oldest-first fairness exactly.

Worth flagging as a pattern: this is the third time in this unit that a
branch which was *unreachable* became *always-taken* once a limiter was
removed. The type system caught the worktree one; this one was only
visible by reading the branch, because the semaphore was reached through
an `any` cast.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (309 +
10 + 71) · project-manager + hybrid-executor + merge-single-flight +
scheduler 93/93.

Nothing booted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:39 -07:00
gsxdsm
743df98aa4 capacity, part 1: merge pinned at 1, worktrees-off mode, and one dead knob deleted (#2502)
First slice of the capacity simplification. Operator: *"just have two
capacity — overall per project agent count and max worktrees. Remove all
other capacities and counts."* Plus two later additions: **merge is
always 1, fixed**, and **worktrees off ⇒ limit by total agents only**.

Three independently revertable commits. No limiter is added anywhere;
one is deleted, one is made structurally absent, and one is pinned.

---

## 1. Merge concurrency ratcheted at 1 (test-only)

I was asked to add a limiter if merge concurrency could be raised. **It
cannot** — there is no setting, workflow property, pool or trait config
anywhere that raises it, so this adds no code and pins what already
holds.

Serialization lives in the **pump**: `drainMergeQueue`’s `mergeRunning`
re-entrancy latch, `activeMergeTaskId` as a single-slot identity, the
`mergeBodyInFlight` next-generation latch, and one `ProjectEngine` per
projectId.

**Not** in the merge-queue lease, which is a per-task ROW (`primaryKey
[projectId, taskId]`) — two tasks can hold leases simultaneously by
construction, and it has exactly one caller (the worktree-reuse
handoff). Ordinary merges never take it. A lease-level test would have
been describing an invariant that layer has never held.

The second half guards the other direction: a merge-concurrency
*setting* would not fail the pump ratchet — it would sit unread until
someone wired it up.

**Revert-proof:** deleting the latch → `expected 1 times, but got 2
times`; deleting the `finally` → latch-stuck; injecting
`maxConcurrentMerges: 2` → fails naming the key; injecting a
`maxParallelLanes` merge-trait field → fails naming the field. Sources
restored byte-identical after each injection.

## 2. `worktreesEnabled` — off means the worktree limit cannot bind

No worktrees-off mode existed (no
`worktreesEnabled`/`useWorktrees`/`worktreeMode` anywhere — only
worktree *configuration*).

**Why not `maxWorktrees: 0`, which needs no new key:** it deadlocks. `??
4` keeps `0` (not nullish), the gate is `used >= limit`, so `0 >= 0`
holds **on an empty board** and nothing ever dispatches — while the
operator-visible reason reads `gate=maxWorktrees; used=0/0`, a limiter
that looks like it is working while the board is dead. It also needs the
Command Center `{min:1}` clamp relaxed. So `0` costs the gate rewrite
*and* the clamp change *and* encodes a mode as a magic value.

**Off is absence, not a big number.** `resolveWorktreeCapacityLimit`
returns `number | null`; `ConcurrencyGateDiagnostic.maxWorktreesGate` is
now optional, so consulting a worktree limit in OFF mode does not
type-check. A gate holding `Infinity` can start binding again the moment
someone "fixes" a comparison; an absent gate cannot.

That paid for itself immediately: making it nullable surfaced a
**second, independent** worktree gate (`activeWorktrees >= maxWorktrees`
early-return) that a skip-by-convention approach would have missed
silently.

**Scope, deliberately:** this is a statement about *counting*, not
isolation. It does not make concurrent agents safe to share one checkout
and builds nothing toward that — the non-worktree paths that exist today
are fallbacks to the operator’s own tree, one of which caused FN-8600.

**Revert-proof:** a resolver ignoring the flag turns both OFF scheduler
tests red while every ON test stays green — they reuse the *same*
fixture (5 in-progress, limit 4) that pre-existing tests prove blocks,
so the pair moves in opposite directions. Removing `disabled:` reddens
the UI test.

## 3. `maxTriageConcurrent` deleted — it controlled nothing

**Measured: zero enforcement reads.** The only `.maxTriageConcurrent`
reference in the repo was a route echoing it back in `/config`. FN-8453
removed the pool it gated and left the knob shipping in
`DEFAULT_SETTINGS`, the settings type, the section registry, the API
response and six i18n catalogs, doing nothing, for releases.

Historical FNXC comments are **updated, not deleted** — they explain a
real past incident; they now say "planning admission slot" so they stop
implying a live setting. Tombstoned so it cannot return.

`/config` loses a field; safe in-repo since `fetchConfig`’s own return
type never declared it.

---

## Two corrections worth recording

- I earlier reported `maxWorktrees` had **no** Settings UI. Wrong —
`WorktreesSection.tsx:47`; my grep was truncated by `head`. It changed
the placement (toggle beside it, rather than a duplicate key in
Scheduling).
- I planned to assert the queued-reason string is rewritten in OFF mode.
Measured that it is **unreachable**: when `maxConcurrent` binds, the
sweep bails before the per-task reason and logs nothing. The test
asserts absence instead.

Two near-misses caught before commit: a pre-existing FN-7505 guard
caught my *new* key missing a description mapping; and editing i18n via
`json.load/dump` silently dropped unrelated duplicate keys
(`autoUpdateAndRestart` in `fr`) — Python keeps only the last of a
duplicated key. Redone textually, every catalog re-validated.

## Verification

`pnpm lint` clean · core/engine/dashboard/i18n typecheck clean · `pnpm
test:gate` green (309 + 10 + 71) · capacity/worktree suites 11/11 ·
engine merge-invariant + scheduler 45/45 · dashboard settings 114/114.
Rebased onto current main and re-verified.

Nothing was booted at any point.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a project setting to enable or disable running tasks in
worktrees.
* Disabling worktrees removes worktree capacity limits from task
scheduling.
* The “Max Worktrees” setting is disabled when worktree execution is
turned off.

* **Changes**
* Removed the unused triage concurrency setting from configuration and
dashboard responses.
* Updated scheduling diagnostics and queue messages to reflect disabled
worktree capacity limits.
  * Added localized labels and help text for the new setting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:31 -07:00
gsxdsm
35b0df1838 U11 PR2: entry contract under the merged column + a real intake-column bug the audit surfaced (#2503)
Second small PR for **U11**. Two commits: a tests-only entry-contract
pin, then a **real present-day bug fix** the audit surfaced.

## The audit you asked for, finished — no design fork

You named four surfaces as the remaining risk. All four can take a
combined `intake` + `hold` column. One needed a code change; here it is.

| Surface | Verdict | Evidence |
|---|---|---|
| `isUnplannedForExecution` | Safe | PR1 (#2495) — passed unmodified; a
mutation now fails exactly the merged-column test |
| Capacity hold / release | Safe | PR1 — `hold-release.ts:260` already
accepts intake **or** hold |
| `start`'s column / entry contract | Safe | commit 1 — all 6 assertions
passed unmodified |
| `createTask` intake wiring | **Broken today** | commit 2 — fixed,
revert-proven |
| *(also found)* triage auto-discovery | Needs conversion |
`triage.ts:1382` — deferred to PR3, see below |

## Commit 1 — entry contract under the merged column (tests only)

All 6 new assertions passed on the first run. **Regression floor, not
evidence of a fix** — I could not make them fail and am not claiming
otherwise.

They pin one real behavioral **difference** rather than asserting
sameness everywhere: the merged shape answers `start` where the split
shape answers `plan`, because `start` becomes the first node in that
column once the columns collapse. That is equivalent *only* because
`start` reaches the specification node by a single unconditional success
edge — asserted, so if a node is ever inserted between them this fails
instead of silently admitting an unspecified card into implementation.

Also pinned: past planning both shapes agree exactly; a card past the
merged column still never resumes at a planning node (the backward drag
that fires `abort-on-exit`); and a row persisted in the **deleted**
`triage` column resolves to `undefined`, safe only while the executor's
start-node fallback exists.

## Commit 2 — a real bug, found by the audit

The intake column was resolved **only** as a by-product of materializing
workflow steps. A create supplying `enabledWorkflowSteps` without an
explicit `workflowId` takes **neither** materialization branch, so
`resolvedEntryColumn` stays `undefined` and `column:` falls through to
the hard-coded `|| "triage"`.

Today, on Coding (Ideas), that lands the card in `triage` — **a column
that workflow does not declare.** Created straight into a phantom lane.
Measured: the new test fails `expected 'triage' to be 'ideas'` against
unmodified sources.

**Why it blocks U11.** Once `triage` leaves the coding IRs this stops
being an Ideas edge case and becomes the default workflow's behavior for
every create down this path: the card lands in an undeclared column
**and** — because `isIntakeColumn` keys on the same `"triage"` literal —
gets `generateSpecifiedPrompt` instead of the bootstrap seed. Triage
admits a card for planning only when its `PROMPT.md` reads as a seed, so
a placeholder spec is classified "already planned" and never planned.
The card sits in Planning forever with no log line in any lane —
**FN-8587's exact failure mode, promoted from one edge case to every new
card.**

The fix resolves the intake column **side-effect-free** (read the IR,
ask which column carries `intake`). It deliberately does *not* call
`materializeDefaultWorkflowSteps`, which would persist step rows the
caller explicitly opted out of by supplying its own toggles.
Unresolvable workflow returns `undefined` and each call site keeps its
legacy fallback, so no path loses behavior when the IR cannot be read.

Applied to both create paths. Branch ordering preserved in both — the
explicit empty-toggle case (`length === 0` hydrating back as `[]`) still
runs, now nested rather than sequential.

**Revert check:** with `task-creation.ts` reverted, *"lands a Coding
(Ideas) task in ideas even when enabledWorkflowSteps is supplied"* fails
`expected 'triage' to be 'ideas'`. The companion bootstrap-`PROMPT.md`
assertion passes either way today — it is correct **by accident of the
`"triage"` literal** — and is kept precisely because that accident
disappears with U11.

## Verification

37 tests green across the three intake/create suites; 119 across the
entry-contract, merged-column and lifecycle suites; `pnpm test:gate`
green (307 + 10 + 71); lint and core typecheck clean. Changeset added.

## Deferred to PR3, with the line numbers

`discoverReadyPlanningTasks` has two hardcoded branches:

```ts
(t) => t.column === "triage" && isTaskStillInPlanningStage(t)   // triage.ts:1382
(t) => t.column === "todo"   && !this.processing.has(t.id) …    // triage.ts:1389
```

Delete `triage` and branch 1 matches nothing for coding cards; branch 2
then does all the work and is **narrower** (it admits only
`needs-replan` or bootstrap-stub cards). Commit 2 is what makes branch 2
sufficient — every new card now gets a real bootstrap seed. They cannot
double-fire: a card is in `todo` xor `triage`.

Two adjacent sites are already merged-shape-ready: `triage.ts:3899`
skips the redundant same-column move for a plan-in-place card, and
`triage.ts:753`'s stale-status sweep already scans both columns.

Then the ~10-line IR change, then the migration proof.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:14 -07:00
gsxdsm
9d3e53d0c5 U8 PR3: the implementation phase announces HOW it ended — including when the executor moved the card itself (#2507)
Third PR of **U8 — the graph owns execution**. Independent of everything
merged so far; small, green, revertable on its own.

## The problem this makes visible

`result.taskDone` is the entire language the execute seam has for
talking to the graph:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
return { outcome: "failure", value: paused ? "implementation-paused" : "implementation-incomplete" };
```

The endings that one bit cannot express are exactly the ones the
implementation phase **transitions itself**:

- a session that paused *after* the work was already complete →
finalizes to review inline;
- a session that stopped because a step is blocked on a pending review →
hands off to review inline (a pending-review block is a wait, not a
failure; marking it failed deadlocks a row that is both `in-review` and
`failed`).

The graph then sees `taskDone === false`, reports
`implementation-incomplete`, and `handleGraphFailure` compensates with
`alreadyFinalizedToReview` / `completionFinalized` — classifiers whose
entire job is recognising a move the graph did not make.

**That was invisible.** An out-of-band transition and a genuine
implementation failure were indistinguishable in logs, in events, and in
tests. You cannot remove a transition you cannot see, and you cannot
prove you removed it either.

## What lands

A closed `ImplementationExit` enum
(`engine/executor/implementation-exit.ts`) reported from six
completion-adjacent exits in `runImplementation`, announced by the
execute seam as `NodeCompleted.exit` on the U3 lifecycle bus. Two ids
are flagged as out-of-band — the ones where the executor, not the graph,
performs the transition.

**Routing is unchanged, and that is the point.** The seam returns
byte-identically what it returned before for every exit, so this PR
cannot move a card. The routing move needs new IR edges and lands
separately; splitting them is what keeps both independently revertable.
Per R5 an exit id is a **reaction** — nothing branches on one, and
dropping every subscriber must change no outcome (a named U8 test
scenario, asserted here).

`NodeCompleted.exit` is added to the event key allow-list deliberately —
which is exactly what that allow-list is for — and carries closed enum
ids only, never prose.

## Revert-proofs, each observed failing

| Injected change | Result |
|---|---|
| Remove the emit entirely | **6 failures** |
| Let an exit change the returned outcome | **2 failures** (the
routing-unchanged pins) |
| Delete one `reportImplementationExit(...)` call site | **1 failure**
(the wiring ratchet) |

**The third proof exists because of a hole I found in my own tests.**
These tests stub `runImplementationPhase` — the only way to reach all
six exits deterministically — which means deleting a real call site left
the entire file **green**. A stubbed seam can only prove the seam. I'd
also written "every exit is reported — the signal is real, not a
placeholder" in the header, which the tests did not support. Both are
fixed: there is now a ratchet asserting every enum id is wired at a real
call site and that each out-of-band id sits adjacent to the handoff it
describes, and the header says what the tests actually prove.

## Scope

**6 of `runImplementation`'s ~28 dispositions** (per the ownership
ledger merged in #2490), chosen as the ones the routing move needs. The
remaining ~22 report nothing yet — the ledger, not this enum, stays the
record of that gap, and the module says so.

## Verification

- 15 new tests + ledger + graph-boundary + task-done-blocked +
graph-requeue-gate + step-session + review-verdicts + tool-failure-retry
— **9 files, 115 tests green**
- `@fusion/core` `workflow-events` — 20 tests green (allow-list change
covered)
- `pnpm test:gate` green (17/307, 2/10, 1/71); `pnpm lint` clean; `tsc
--noEmit` clean on both packages
- Changeset included (`patch`, `internal`), passes `check:changesets`

## Next

PR4 is the routing move itself: `review-handoff-pending-review` becomes
a graph outcome with its own IR edge, and `alreadyFinalizedToReview`
becomes provably unreachable for that path. The IR edge change will be
its own commit, separate from the seam change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:07 -07:00
gsxdsm
d5030c55ea test(engine): prove BOTH rebound paths on a renamed board — 2 more ledger entries closed (#2510)
Stacked on #2508. Test-only.

## Why this site matters more than most

`resolveReboundTarget` answers one question: **where does a recovered
card go back to?**

Keyed on the literal `todo`, a recovered card on a renamed board is
requeued to a column that board **does not declare**. That is not
cosmetic — an undeclared column carries no trait flags, so `findColumn`
returns undefined and the card becomes invisible to every trait-driven
sweep: nothing schedules it, nothing releases it, the board does not
draw the column. **The "recovery" strands the card harder than the
failure it was recovering from.**

One of the two covered paths, `reconcileUndeclaredTaskColumns`, exists
*specifically* to repair that state — which makes it the worst possible
place for this bug to live.

## Covered, each mutation-verified independently

| site | mutation | result |
|---|---|---|
| `reconcileUndeclaredTaskColumns` | target → `"todo"` | exactly the 2
renamed cases fail |
| `autoRecoverWorktreeSessionStartFailure` | rebound → `"todo"` |
exactly the renamed requeue fails |

Neither needs git — the corrected ledger lens from #2508 (*what the
function touches*, not *what family it sits in*) made that obvious
rather than assumed.

## Both negatives included

"Re-home anything whose column looks wrong" would be a louder failure
than the strand it repairs, so: a card whose column **is** declared is
left alone, and an operator `userPaused` park is never undone.

## Fixture finding, kept in-file

`updateTask({ userPaused: true })` leaves the field `undefined` on both
`getTask` and `listTasks({slim:true})`. Seeding it that way produced a
card the sweep **correctly** saw as unpaused — a broken fixture that
would have read as a broken guard, and would have looked like a real
safety hole in the paused-park protection.

Found by probing the persisted row rather than trusting the write. Now
seeded through the integer column directly, and the test asserts the
seed took effect *before* exercising the sweep, so this cannot silently
regress into a vacuous pass.

## Verification

- three live-E2E suites green together (lifecycle, merge-family,
rebound-family)
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:01 -07:00
gsxdsm
4eaa509024 test(engine): prove merge finalization on a renamed board — 3 ledger entries closed, and the ledger itself corrected (#2508)
Test-only. Closes three `auto-merge-finalization` entries from the
unproven-sites ledger.

## Why this one first

It is the **last move a card makes**. Keyed on the literal `done`, a
renamed board's proven-merged card is moved to a column its own workflow
does not declare — or refused and left stranded in review **with the
work already landed**. That is the most expensive failure shape in the
lifecycle, and nothing had run it against a renamed workflow.

## The ledger was wrong, and is corrected in this PR

My own ledger said this family *"needs a REAL git worktree, branch, and
squash … an engine-slow real-git lane, not another table row"*.

`finalizeProvenAutoMergeTask` **needs no git at all** — the merge proof
is a field on the row. It was reachable the whole time. The inference
came from *the family the code sits in* rather than from what the
function actually touches, and it parked reachable coverage for a slice.
The correction is written into the ledger so the remaining entries get
re-checked the same way rather than inheriting the assumption.

## A second correction, from mutation-testing rather than reading

I first claimed `resolveMergeOrchestrationColumn` as covered because it
sits in the same resolver as the other two. **All cases passed with it
hardcoded.** It changes only whether finalization records a
column-mismatch *repair* — never where the card lands, which is why the
other cases are blind to it.

It got its own case. Keyed on `in-review`, a renamed board's card
resting in `checking` compares unequal, so **every ordinary finalization
would be audited as repairing a mismatch that never existed** — a
healthy board reads as one constantly self-healing, and the audit trail
operators use to spot real strandings fills with false positives.

Sitting next to covered code is not coverage.

## Mutation-verified independently

| mutation | result |
|---|---|
| `completeColumn` → `"done"` | 3 fail — both renamed cases + the
differential |
| `isCompleteColumn` → `id === "done"` | exactly the already-done case
fails |
| `mergeColumn` → `"in-review"` | exactly the new audit case fails |

## Shared fixture extracted (pure move)

The vocabulary + IR builder moved to `_workflow-vocabulary-fixture.ts`
so the two suites cannot drift into testing different workflows — two
copies of a differential fixture is precisely how a renamed-workflow
test starts passing for reasons unrelated to the code under test. The
lifecycle suite is unchanged: **20/20 before and after**. The
`mergeOrchestration` trait is an opt-in option so the existing suite's
IR stays byte-identical.

## Fixture note worth keeping

Seeding needed **completed steps**: task creation parses three pending
steps out of the bootstrap PROMPT even with `applyDefaultWorkflowSteps:
false`, and `getTaskHardMergeBlocker` refuses on them (`"task has
incomplete steps"`). Found by the suite blocking on **both**
vocabularies — the signature of a broken fixture rather than a broken
guard.

## Verification

- 51/51 across the lifecycle, merge-family, ratchet and hold-release
suites
- engine `tsc --noEmit` clean
- `pnpm test:gate` green (307 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:15:35 -07:00
gsxdsm
8288e4a8ab U7 PR3: the specification reaction acts on what finalize DID, not on the fact that planning stopped (#2506)
Completes the pair started in #2498. That landed the outcome; this makes
the engine's reaction consume it.

## The bug

`onSpecifyComplete` fired on **every** finished specification, because
the seam announcing it fired unconditionally. So a card parked at the
manual plan-approval gate — finalize writes `status:
"awaiting-approval"` and **returns early**, before the release move —
was logged as `Specified X → todo` and had a Plan Review run armed for a
plan the operator had not approved.

#2491 stopped the **seeder** from acting on that, defensively, at the
seeder. This removes the reason it was ever asked. Both layers are
deliberate and neither is redundant:

- the seeder guard covers **every caller**, including self-healing's
re-seed;
- this one stops the engine doing work nobody asked for, and stops it
telling the operator something false about their own board.

`released` is the only outcome that licenses arming a run — the only one
meaning the card crossed into the hold column (or was already resting
there, plan-in-place) and is the graph's now. `parked` belongs to a
human; `withheld` belongs to the caller's retry budget.

## The event still fires on every outcome

Deliberately. Dropping the reaction for a non-release would also drop
the runtime's `recordActivity()` idle signal, and a reaction that
silently does not happen is harder to reason about than one that happens
with an accurate payload. R5's division of labour: **the seam announces,
the subscriber decides what a given outcome licenses.**

## Why there is a new extracted function

`reactToSpecificationComplete` is pulled out of the inline
`InProcessRuntime` callback for the same reason the continuation drain
was in #2491: the callback is built inside a class whose construction
attaches to the real central project registry, so no test could
distinguish *"the reaction respects the outcome"* from *"the reaction
ignores it"*.

**Revert proof:** with the outcome gate removed from the reaction, **5
of 8 fail**.

## Two call-site decisions worth naming

**`tryFinalizeExplicitDuplicateMarker` reports through a mutable ref,
not a widened return type.** Its boolean answers a *different* question
— "was this a duplicate marker at all?" — and 16 existing tests assert
it directly. I tried the widened return first and it turned all 16 red.
Expectation edits are exactly how a behavior change travels disguised as
churn, so I backed it out. **This diff touches zero existing test
expectations.**

**A duplicate-marker redirect reports `parked`**, which is accurate: it
deletes, flags, or clears the marker; it never releases the card into
the hold column.

## A fixture note — third of this shape on the program

My "task vanished between release and reaction" case passed `undefined`,
which triggered the harness **default parameter** and silently handed
the reaction a live task — making it a duplicate of the control rather
than the case it claimed to be. It now passes `null`, with a comment
saying why.

Running tally of near-false-greens on this unit, all the same family: a
fake that ignores its predicate (#2491), a stub that ignores its
callback (#2498), a default parameter that swallows the interesting
input (here). Each was caught by the test failing for the *wrong reason*
and being read rather than fixed.

## Verification

| Check | Result |
|---|---|
| new suite | 8/8 |
| 15 triage / planning / continuation suites | 361/361, **no expectation
edits** |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (307 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:15:29 -07:00
gsxdsm
a2b4ca76ac U11: delete the unreachable legacy todo dispatcher from scheduler.schedule() (-929 lines, pure deletion) (#2505)
Based on `main`. **Pure deletion — no behavior change**, because the
deleted code cannot execute.

## Found while trying to convert it

This started as a U11 slice to make the scheduler's dispatch path
resolve its column by trait. Per the lesson from the dependency-blocked
feature I checked reachability *before* converting:

```ts
function shouldRunWorkflowColumnScheduler(_settings: Settings): boolean {
  return true;                       // parameter UNUSED, body a literal
}
...
if (shouldRunWorkflowColumnScheduler(settings)) {
  await this.runHoldReleaseSweepPass(tasks, settings);
  ...
  return;                            // UNCONDITIONAL, at the block's own depth
}
<929 lines of legacy pull-from-todo dispatcher>   // unreachable
```

The guard takes an **unused** parameter and returns a **literal**, so
the branch is statically always taken, and it ends in an **unconditional
`return`**. Everything after it in `schedule()` is unreachable.

`tsc` doesn't flag it because the condition is a function call rather
than a literal — which is exactly why 929 lines survived the U6 cutover.
The replacement was added *in front of* the old dispatcher rather than
*instead of* it, and the in-file comment says so outright:

> the hold/release sweep owns todo→in-progress pickup, so do not fall
through into the legacy pull-from-todo dispatcher after the sweep runs

## Why this matters beyond line count

**4 of the 15 `"todo"` literals in `scheduler.ts` live in this dead
region.** Converting them would have been pure waste — and worse, it
would have reported progress against the U11 critical path while
changing nothing. 11 live sites remain and are the real work.

## Corroborating evidence

Six imports became unused and are removed with it:
`resolveDependencyOrder`, `sortTasksByPriorityFanoutThenAgeAndId`,
`buildUnblockWeightMap`, `TransitionRejectionError`,
`isUnplannedSeedPrompt`, `DEFAULT_WORKFLOW_POOL_ID`.

That the dead region was their **only** consumer in this file is itself
evidence: a live dispatcher would still need dependency ordering and
priority sorting.

## Why no new test

The proof here is **static, not behavioral** — an unconditional `return`
before the code. A test cannot demonstrate absence of execution more
strongly than the control flow already does, and one that passed both
before and after would be theatre.

The evidence that nothing depended on it: **all 100 scheduler tests and
the full merge gate pass unchanged.**

## Measured

`scheduler.ts` **3,726 → 2,797 = −929 lines.**

Unlike every consolidation in this program, this is a **genuine net
reduction** — nothing was moved elsewhere.

## Verification

100 scheduler tests green across all 7 scheduler suites; merge gate
green (307 + 10 + 71); tsc clean; lint clean.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-28 17:15:23 -07:00
gsxdsm
e4004c8694 U9 baseline: pin merge-region IR config as a dead policy authority (test-only) (#2494)
**U9, PR1 of several.** Test-only, no production code touched. This is
the characterization baseline the plan's Execution note asks for before
the merge lane converts.

## The finding

`builtin-coding-workflow-ir.ts` declares merge-region policy that **no
engine code reads**:

| IR declaration | Consumed by |
|---|---|
| `merge-retry` → `{ policy: "merge", maxAttempts: 3 }` | nothing —
`retry-backoff` handler is `async () => ({ outcome: "success" })`
(`workflow-node-handlers.ts:728`) |
| `merge-manual-hold` → `{ release: "manual" }` | nothing — returns a
constant `manual-required` |
| `branch-group-*` → `{ maxReworkCycles: 3 }` | nothing — returns a
constant `success` |

Live merge policy authority is elsewhere, on two separate axes:
- **conflict** retries — `settings.maxAutoMergeRetries` (default 3),
already covered by `auto-merge-retry-cap-settings.test.ts`
- **transient** retries —
`ProjectEngine.MAX_AUTO_MERGE_TRANSIENT_RETRIES = 5`
(`project-engine.ts:545`)

So the IR is a **third, dead authority**. These are different axes, not
a same-axis contradiction — but a reader looking at the IR would
reasonably take the declared numbers as live, and nothing currently says
otherwise. U9's acceptance criterion is "merge policy changes via IR
config alone, with no code change"; that fails today and this pins why.

## Why characterization rather than a fix

Making these handlers config-driven is a **merge behavior change**, and
the `requestMerge` primitive it routes through lives at
`executor.ts:7383` — inside U8's blast radius. U9 is sequenced behind U8
precisely so the merge lane converts onto an executor that is already
substrate. Landing the behavior change now would change merge semantics
on an executor about to be reshaped. It lands inside U9 proper.

When U9 wires a node kind onto its IR config, the matching case here
goes **red** and the U9 commit must move that kind out of
`CONFIG_BLIND_MERGE_REGION_KINDS`. That is the ratchet working.

## Proof it fails when reverted

A test that passes with the change reverted is not a test. The "change"
here is the test itself, so the honest analogue is mutating the
characterized production behavior. Three independent mutations, each
reverted after measuring:

| Mutation | Result |
|---|---|
| `retry-backoff` honours `config.maxAttempts` (what U9 will do) | **2
failed** / 6 passed |
| `manual-merge-hold` honours `config.release === "external-event"` |
**2 failed** / 6 passed |
| IR declaration drift: `maxAttempts: 3` → `7` | **1 failed** / 7 passed
|

Measured: 8 tests, 4.16s. `pnpm lint` clean. Tree restored to clean
after each mutation.

The assertions are behavioral, not string matches: each handler is
invoked with two contradictory configs (opposite budgets, opposite
release modes, disjoint surfaces) and asserted to return deep-equal
results.

## Six safeguards

This PR changes no production behavior, so no safeguard is altered by
it. The full six-row table with test attribution is the required
artifact for the **conversion** PR, not this one. Baseline located so
far, to be completed and verified by mutation before any conversion
lands:

| # | Safeguard | Consulted at (today) | Test attribution |
|---|---|---|---|
| 1 | user pause | `project-engine.ts:645` (`task.paused \|\|
task.userPaused`) | not yet verified |
| 2 | `autoMerge:false` | `allowsAutoMergeProcessing` —
`project-engine.ts:2797`, `merger.ts:7178` | not yet verified |
| 3 | dependency gating | not yet located | not yet verified |
| 4 | capacity | not yet located |
`workflow-column-boundary-capacity.test.ts` (unverified) |
| 5 | merge-proof | `getTaskMergeBlocker` — `project-engine.ts:2609` |
`merger-file-scope-invariant.test.ts`,
`merger-diff-volume-gate.slow.test.ts` (unverified) |
| 6 | at-most-once merge | `activeMergeTaskId` single-flight —
`project-engine.ts:693`/`:2729` | not yet verified |

Rows 3, 4 and all attributions are honestly incomplete rather than
asserted — I will not present a table I have not earned.

## Also found, for the coordinator

- **Slice statuses are stale.** S02/S03/S04 in
`docs/plans/workflow-owned-merge-stack/` are all marked
`draft-stack-handoff` but S04 has **landed** (the merge-region IR nodes
above), S03's `claimDueWorkflowWorkItem` is implemented and wired via
`workflow-work-processor.ts`, and S02's
`projectMergeRequestToWorkflowWorkItem` is implemented with **zero
production callers**. S06/S07/S08 are genuinely not started. Doc
correction coming as its own small PR.
- **S1 prerequisite verified present, not assumed** — all four store
methods live in `store.ts`, migration `0031` in tree. No S1-completion
gap.
- **Second control plane into the merge lane:** `self-healing.ts:3198`
and `:7200` call `enqueueMerge` directly, bypassing the graph. That
needs to become a recovery-fact/wake (the stack's R6) during U9.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added a new test suite to cover U9 merge-region behavior across
supported workflow node types.
* Verified merge-region results are consistent across built-in,
contradictory, and missing configuration inputs.
* Documented current behavior for retry backoff (always succeeds) and
manual merge hold (fails as manual-required).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:55 -07:00
gsxdsm
eaea082259 U8 PR2: the execution-policy ladder resolves its own workflow's columns (the wip literal made retry, escalation and loop protection unreachable) (#2497)
Second PR of **U8 — the graph owns execution**, independent of
[#2490](https://github.com/Runfusion/Fusion/pull/2490) and of every
other unit. Small, green, independently revertable.

## The defect

`handleGraphFailure`'s execution-policy ladder — FN-7863/FN-7926
dispatch-loop terminalization, FN-7996 tool-failure retry, FN-7998
escalation — decided a task's own lifecycle by naming `"todo"` and
`"in-progress"` **literally, at 9 sites**. U5b converted the executor's
*rebounds* to `resolveReboundColumnFor`; these were left behind, each
sitting somewhere an awaited resolver could not reach: inside
synchronous `updateTaskAtomic` mutators, inside fire-and-forget resume
closures, and in conditions evaluated before any resolution happened.

**The severe one is the wip gate, and it fails silently in the worst
direction:**

```ts
if (live.column !== "in-progress") {
  // "Workflow graph run ended after task already advanced — no further action needed"
  return;
}
```

Under a workflow that renames the implementation column, that is true of
a card sitting in **its own wip column**. So the graph failure was
swallowed whole — no terminal park, no status, no error, nothing on the
board — and the scheduler re-dispatched the same doomed run. Every later
branch sits behind that gate, which is why the retry budgets, the
escalation, and the bounded terminalization were **unreachable rather
than mistargeted**.

This is precisely the failure the program's problem frame predicts: *a
guard that stops matching disables a recovery path invisibly and the
suite stays green.* I found it because my first renamed-column test for
the escalation site could not reach the escalation code at all.

Two further sites misbehave once the gate is passable:

- **FN-7998 node escalation** wrote `column: "todo"` inside the atomic
claim — parking the card where no workflow declares it, which is on the
plan's **"Stop implementation if"** list and what R7 exists to clean up
after. The scheduler's effective-node resolution, the entire point of a
node escalation, never runs.
- **FN-7863/FN-7926's `live.column === "todo"` arm** is the classic
guard that stops matching. In-process the `executeNodeSelfRequeued`
marker covers the same case, so this degrades only on the **durable**
arm — after a restart, or for a second `TaskExecutor` instance in the
process, where the column read is the only evidence the inner executor
requeued. A progressing card then falls through to the terminal sink and
is parked `failed`.

## The fix

Resolve hold and wip **once per graph failure** through U1's
`resolveTaskLifecycleColumns` and thread the pair through the ladder.
Both fall back to the legacy literal when the workflow cannot be
resolved, so an unresolvable workflow keeps exactly its pre-conversion
behavior rather than guessing. One IR read on a terminal recovery path —
not an enumeration loop.

## Red-green, measured

**3 of the 8 new tests fail with this commit's executor change
reverted:**

```
FAIL  FN-7998 … > requeues a node escalation to the RENAMED hold column, not the literal todo
FAIL  FN-7998 … > still does not move the card for a MODEL-target escalation
FAIL  FN-7863/FN-7926 … > recognises an inner-executor requeue that landed in the RENAMED hold column
      Tests  3 failed | 5 passed (8)     ← reverted
      Tests  8 passed (8)                ← with the fix
```

The other **5 pass both ways by design**, and I am not claiming them as
red-green — they are the regression floor:

- default coding workflow still resolves hold → `todo`, wip →
`in-progress` (byte-identical);
- an unresolvable workflow still uses the legacy literals;
- the in-process self-requeue marker still works when no workflow
resolves;
- and a **negative case** proving the dispatch-loop gate stays narrow —
a card still in its wip column with no marker is a genuine execute
failure and must NOT be swallowed as a benign recovery. Widening that
gate to "any column" would have been the easy wrong fix.

## Scope

Deliberately the execution-policy ladder only. **20 further column
literals remain in the same method's pause-abort, merge, and in-review
regions** — they belong to U5's executor slice (B4, not started) and
U9's merge lane, and are untouched here. Flagging the overlap: this PR
edits `executor.ts`, so whoever takes U5-B4 should rebase onto it rather
than converting these 9 sites again.

## Verification

- 8 new tests + the preserved-behavior suites
(`executor-tool-failure-retry`, `executor-graph-requeue-gate`,
`executor-task-done-blocked`, `executor-graph-boundary`,
`executor-stuck-requeue-preserve-progress`,
`executor-paused-abort-todo-benign`, `executor-abort-provenance`) — **9
files, 112 tests, green**
- `pnpm test:gate` — green (2/10, 16/299, 1/71); `pnpm lint` clean; `tsc
--noEmit` on `@fusion/engine` clean
- Changeset included (`patch`, category `fix`), passes `pnpm
check:changesets`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed execution recovery for workflows with renamed lifecycle columns
so retry, escalation, and loop-protection behaviors correctly follow the
workflow’s declared hold/WIP columns.
* Preserved legacy behavior for default workflows and continued safe
handling when lifecycle columns can’t be resolved.
* **Tests**
* Added a Vitest suite validating execution-policy “ladder” behavior for
renamed columns, including node escalation, dispatch-loop gating, and
fail-closed scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:47 -07:00
gsxdsm
2934cccad8 U7 PR2: finalize reports what it did with the card — a refused planning handoff is retried, not counted as recovered (#2498)
## The bug

`finalizeApprovedTask` has ~25 exit points and returned `void`, so no
caller could tell *"the card was handed off"* from *"finalize gave up"*.
Both callers assumed success.

`recoverApprovedTask` returned `true` **unconditionally** after
finalize, and `handleStuckAbortRequeue` treats `true` as "recovery done,
stop here". So when the release move was **refused by the planning-stage
guard** (FN-8361), or the store could not perform the move at all,
recovery reported success and the card's stuck-retry budget was skipped
— nothing re-planned it, nothing escalated it, and it sat in the planner
column holding a finished spec.

The refusal was already logged loudly by FN-8596's visibility work. The
return value was the part still lying.

## Three states, not a boolean

This is the load-bearing decision in the PR:

| Outcome | Meaning | Retry? |
|---|---|---|
| `released` | crossed into the hold column, or already resting there
(plan-in-place) | n/a — handed off |
| `parked` | deliberate, terminal-for-now: awaiting manual plan
approval, duplicate decision, operator pause, deleted duplicate | **no**
— a human owns it |
| `withheld` | finalize could not complete the handoff, nothing waiting
on a human | **yes** — caller's budget owns it |

`recoverApprovedTask` returns `outcome !== "withheld"`, so **`parked`
still returns `true`**. Narrowing to `=== "released"` is the tempting
simplification and it is wrong: it would send the stuck handler down its
draft path and stamp `needs-replan` over a plan a human is mid-review on
— a worse bug than the one being fixed. That is asserted, and the
assertion fails under exactly that narrowing.

## Why a mutable report, not a return at each exit

Threading a return through 25 exits is 25 chances to mis-classify a
branch, and mis-classifying turns a truthfulness fix into a lifecycle
bug. The report defaults to `parked`, which is equivalent to today's
observable behavior at every exit — so the plumbing is **inert
everywhere except the three sites explicitly classified**. Adding a
state to an exit is then a deliberate, reviewable act rather than a
diff-wide judgement call.

Only **two** exits are marked `withheld`, both in the release block,
both already warning loudly. Deliberately *not* marked:

- the `updatePlanningStateIfStillCurrent` guard — FN-8024 says a normal
scheduler advance legitimately lands there; the card has moved on, so a
retry would be wrong.
- `recoverMissingPromptBeforeRelease` — it owns its own recovery budget;
retrying would double up.

## Revert proofs (measured)

| Reverted | Result |
|---|---|
| `recoverApprovedTask` back to unconditional `true` | `Tests 2 failed
\| 3 passed (5)` |
| narrowed to `outcome === "released"` | `Tests 1 failed \| 4 passed
(5)` — the approval-park control |

The second row is the point: the park case is load-bearing, not
decoration.

## A fixture note that nearly produced a false green

A `vi.fn()` stub for `updateTaskAtomic` that ignores its callback makes
**every** finalize report "no longer in the planning stage" and return
before the release — silently collapsing every case into the same
uninteresting early exit. My first run was 3 failures for that reason,
not the reason I expected. The fake now applies the patch, and the
control asserts `moveTaskIf` was actually reached. Same class as the
`moveTaskIf` fake caught on #2491; recording it so the next person
recognises the shape.

## Scope

The other caller — `specifyTask`'s unconditional `onSpecifyComplete` —
is **not** gated here. Reaching it needs a live planning session, so
gating it without first extracting the reaction would be a change I
cannot prove, which is exactly the finding review caught on #2491's
deferral. That lands next, on this plumbing.

## Verification

| Check | Result |
|---|---|
| new suite | 5/5 |
| 11 triage/planning suites (triage, finalize-duplicate-lineage,
stuck-requeue-preserve-draft, explicit-duplicate-marker, preflight,
plan-artifact-writeback, refinement-routing, planning-wake,
planning-evacuation, …) | 326/326 |
| `tsc --noEmit` (engine) | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green (299 + 10 + 71) |
| `pnpm check:changesets` | clean |

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:36:20 -07:00
gsxdsm
8aba310d78 U11: resolve the worktree-acquisition requeue column by trait (2 sites, both branches) (#2496)
Based on `main`. First of my U11 conversion PRs — small, green,
independently revertable.

Both heartbeat worktree-acquisition requeue sites hardcoded `"todo"`.

## Why this is critical path, not a renamed-workflow nicety

**U11 deletes the `todo` column from the builtin workflows.** After
that, these two sites would requeue every acquisition-failed card into a
column that no longer exists.

## Both sites converted together

They are different branches of the same failure:
- the **bounded-retry** requeue, and
- the **retry-cap-exhausted** terminal park.

Converting one and not the other would leave the rarer path — which
fires only after three consecutive failures, so it's the one least
likely to be noticed — still writing the literal.

Target is the KTD-10 ordering via `resolveReboundTarget` (hold → intake
→ first column): the same helper `self-healing` and `mesh-lease-manager`
already use for "requeue a recovered card", so the recovery paths cannot
drift apart.

## What is deliberately untouched

`preserveStatus: true` on the exhausted path. It exists because
reopen-to-todo semantics would otherwise wipe the `status: "failed"`
written immediately before (FN-7721) — changing the column must not
disturb that flag. A test asserts the full options object, not just the
column.

Fail-soft to the legacy id: a requeue must not be abandoned because a
workflow lookup failed, or the card is left holding a worktree it could
not acquire. Covered by a regression-floor test.

## Verification

- **Mutation-verified:** restoring the literal fails 2 of the 3 new
tests
- 7 tests green (3 new + the 4 pre-existing worktree tests, unchanged)
- tsc clean, lint clean, merge gate green (299 + 10 + 71)

## Measured progress

**2 of the 74** code-level `"todo"` sites in my unit (engine
recovery/scheduling core) are now trait-resolved.

Remaining in-unit: `self-healing` 48, `scheduler` 15, `triage` 8,
`replan-target` 1.
`stuck-task-detector` needs **no work** — all 4 of its occurrences are
comments, not code.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Tasks now return to the workflow’s configured hold column when
heartbeat worktree acquisition fails, including workflows that use a
renamed hold column.
* Retry and retry-limit handling now preserves task progress and, when
applicable, status.
* Added a safe fallback to the default “todo” column when workflow
details cannot be resolved.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-28 11:36:13 -07:00
gsxdsm
8492278fdd U11 PR1: pin the merged intake+hold column contract before the IR moves (a mutation proved the first 7 tests insufficient) (#2495)
First of several small PRs for **U11** (merge Todo into Planning).
**Tests only — no production change.** It lands the precondition so the
IR edit arrives on proven substrate instead of an assumption.

## Decision taken (reversible, proceeding on it)

**The surviving Planning column keeps the id `todo`; `triage` is
deleted.** Same board the operator asked for — one column labelled
"Planning", no "Todo" — via the cheaper and safer half.

Measured, comments excluded, non-test, `packages/*/src` +
`dashboard/app`:

| | guards | writes | fallbacks | total |
|---|---:|---:|---:|---:|
| `"todo"` | 121 | 68 | 9 | 323 |
| `"triage"` | 90 | 30 | 21 | 304 |

Deleting `triage` instead of `todo` also means **no data migration**
(every live card in `todo` is already in the surviving column) and **no
guard changes meaning** (`column === "todo"` still denotes the hold
column). Under the plan's letter the opposite is true, and worse than
"dead": because Coding (Ideas) keeps `todo` per R10/R11, a surviving
`column === "todo"` guard would stay live for Ideas cards while silently
never matching for Coding cards — workflow-dependent, not dead.

This is also a proven in-tree pattern rather than a new idea:
**`builtin:coding-ideas` already ships this exact merge** — id `todo`,
display name "Planning", `hold(capacity)` + `reset-on-entry`,
plan-in-place.

Consequence worth flagging: **U11 no longer waits on Phase B.** The 121
`todo` guards keep their meaning, so converting them becomes U12 cleanup
rather than a U11 blocker.

## What this PR pins

Nothing in tree has ever carried `intake` and `hold` on one column.
Every built-in splits them. KTD-1 asserts the merged shape works; that
assertion was untested.

## The result, reported as found

**All 11 assertions passed on the first run against unmodified
sources.** The merged column is already supported by trait resolution,
the capacity sweep, and the release gate. **I could not make the first
seven fail**, so they are a regression floor — not evidence of a fix,
and I am not claiming them as one.

What makes them worth keeping is that they are *differential*: the same
scenario runs against the split-role vocabulary and the merged one and
asserts the role-level outcomes are **equal**, so a literal creeping
into any path fails the merged half while the split half stays green.

## The finding

**The first seven tests were not enough, and proving that is the point
of this PR.**

A mutation encoding the plausible-but-wrong belief *"an intake column
has no releaser"*:

```diff
- if (currentFlags.intake !== true && currentFlags.hold !== true) return false;
+ if (currentFlags.intake === true) return false;
+ if (currentFlags.hold !== true) return false;
```

left **all seven green**.

That belief is not hypothetical — it is stated verbatim in
`builtin-plan-review-group.ts`'s own FNXC comment as the reason Plan
Review lives in `todo` rather than `triage` today. Under U11 the
planning column **is** an intake column, so any code encoding it
silently stops holding unplanned cards and they release into
implementation with a bootstrap stub for a spec.

The gap: nothing reached `isUnplannedForExecution`. The mock store had
no `getTasksDir`, so both halves of the gate returned early — the tests
were exercising less than they appeared to. The fourth block drives it
with a real temp dir and a real bootstrap `PROMPT.md`.

**Re-running the same mutation now fails exactly one test — the
merged-column one — while its split-shape twin stays green.** That
discrimination is what the suite is for.

## Verification

26 tests green across this file plus `hold-release-renamed-columns`,
`hold-release-instrumentation`, and `pre-release-plan-review`. Lint
clean. No production file touched, so there is nothing to regress.

## Next PRs in this unit

1. Entry-contract test for `start` in a hold-carrying column (the
specific interaction the earlier, reverted attempt got wrong).
2. The ~10-line IR change itself — deliberately last, per KTD-7.
3. The intake-lane `triage` conversion: the 21 `?? "triage"` creation
defaults are the dangerous ones, since they would silently create cards
into a column that no longer exists.

`self-healing.ts` (11 triage guards + 4 writes) is the main worker's
file — I am not touching it and will hand over the line list rather than
race them.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added coverage for merged Planning column behavior across split and
merged workflow configurations.
* Verified intake and hold resolution, rebound targeting, capacity
hold/release outcomes, and execution gating.
* Confirmed planned cards are released appropriately while cards already
in progress are not unnecessarily held.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 11:36:05 -07:00
gsxdsm
fbe7eb5c5a U7 PR1: the manual plan-approval gate was bypassable (3 planning-lane surfaces, 8/13 revert-proof) (#2491)
## What this is

The first slice of **U7 — the graph owns planning**. Characterizing the
planning lane's dual ownership turned up a live defect in the exact seam
the unit exists to remove, so this PR fixes that first and reports the
measured map of what U7 still has to move.

## The defect

The manual plan-approval gate parks a card by writing `status:
"awaiting-approval"` and **returning early** from `finalizeApprovedTask`
— before the release move. `specifyTask` then calls `onSpecifyComplete`
**unconditionally** afterwards. Three automated surfaces went on to
advance the parked card, each having re-derived its own weaker "may I
advance this?" check from `paused`/`userPaused` alone.

`isTaskBlockedOnApproval` (`packages/core/src/task-merge.ts`) already
declares itself *"the single shared predicate core and engine code must
consult before rebounding, requeuing, resuming, re-planning, or
otherwise advancing a task"*. **Measured: it had exactly one production
consumer** (`overseer-human-control-policy.ts`). Now four.

Reachable end to end for a **plan-in-place** card — one whose column
already equals the plan-review node's column (Coding (Ideas), or any
`needs-replan` revision resting in the default workflow's `todo`):

```
park at awaiting-approval
  → onSpecifyComplete fires anyway
  → a runnable plan-review continuation is seeded
  → the drain dispatches it
  → Plan Review runs on a plan the operator never approved
  → its evidence satisfies isUnplannedForExecution
  → the capacity sweep releases the card into In progress
```

Blast radius: projects that have manual plan approval switched on.
`planApprovalMode` defaults to auto-approve (FN-7557), so unset projects
have no gate to skip — but the operator who turns it on is precisely the
one who cares.

## Surface enumeration

Per AGENTS.md — fix the invariant, not the repro.

| # | Surface | Fix |
|---|---|---|
| 1 | `issueRelease` — the choke point for the sweep, `promoteHeldTask`,
`releaseHeldTaskByEvent`, and the scheduler's `reserveSlot` guard |
Guarded there rather than inside `isUnplannedForExecution`, because an
approval-held card is not "unplanned". Guarded **again** inside the
`moveTaskIf` predicate so a park landing mid-sweep cannot lose the race
(R6 — only the in-txn check is authoritative). Operator force-promote
(`allowUnplanned`) still waives it: that *is* a human decision about
this card. |
| 2 | **Both** continuation seeders —
`seedPreReleasePlanReviewContinuation` (normal completion) and
`evaluateStrandedHoldContinuation` (FN-8592 self-healing re-seed) |
Guard at the seam, not in the callers: the seeder itself checked
nothing, and its two callers each pre-checked a different subset. |
| 3 | `resolvePlanningContinuationCandidate` (drain classifier) |
**Skip, never orphan.** Cancelling terminalizes the item, so an approval
landing a minute later would have nothing left to resume and would need
a second repair to come back. |

## Measured, not assumed

The two hold shapes `isTaskBlockedOnApproval` accepts were **not equally
broken**. The `paused` + `pausedReason` shape was already refused by the
sweep and the drain — they happen to test `paused` — so it was refused
*for the wrong stated reason*, not advanced. Every genuine advance gap
is on the **status-only** shape, which is exactly what the gate writes.
Both are covered anyway, plus an `ORDINARY_PAUSE` counter-case so the
new check cannot quietly become a catch-all for every operator park.

## Revert proof

With the three production files reverted: **8 of 13 tests fail.** The 5
that still pass are the 3 controls and the 2 pause-shape rows the
pre-existing `paused` checks already covered.

```
·x··xxxxx·xx·      → Tests 8 failed | 5 passed (13)
```

## Verification

| Check | Result |
|---|---|
| new suite | 13/13 |
| hold-release (×2) + plan-review (×3) + pre-release-plan-review +
promote-force-unplanned | 43/43 |
| stranded-hold-continuation (×2) + continuation-selection +
planning-finished-wake + planning-service | 27/27 |
| scheduler-trait-dispatch | 9/9 |
| `pnpm --filter @fusion/engine exec tsc --noEmit` | clean |
| `pnpm lint` | clean |
| `pnpm test:gate` | green |
| `pnpm check:changesets` | clean |

## Two findings for the coordinator

**1. `triage.ts` is absent from the Phase B census.** The plan's
per-file table (535 sites) covers `self-healing.ts` (U4), the
executor/scheduler cluster (U5), and the core policy modules (U6).
`triage.ts` appears in none of them, so its lifecycle-column literals
are unowned scope — U7 absorbs them.

Measured with the plan's own methodology (block and line comments
stripped, code lines only): a naive quoted-literal grep of `triage.ts`
reports **50** sites, but **35 of those are the agent *role* string
`"triage"`**, not the column. The genuine lifecycle-column surface is
**15 sites**, of which 12 are planning-lane and 3 are `column !==
"done"` in duplicate search. The 50 figure would over-count by 3.3×.

**2. The graph's planning seam is a rubber stamp, in triplicate.**
`createAuthoritativeWorkflowSeams().planning` returns `{ outcome:
"success", value: "pre-specified" }`;
`WorkflowPlanningService.runPlanningSession` returns the same;
`createNoopLegacySeams().planning` is a bare success. The real
specification is ~1,000 lines of `triage.specifyTask`, entirely outside
the graph. That is the flip U7's remaining slices have to make, and it
is the reason the planning lane has two owners at all.

## Deliberately not in this PR

Triage's unconditional `onSpecifyComplete` call. That is the
**ownership** half — `finalizeApprovedTask` must report whether it
released, and the reaction must key on that outcome — and it belongs
with the seam flip, where finalize's outcome becomes the graph's edge
condition anyway, rather than as a half-measure now. With the three
guards above in place, the downstream damage is already contained; what
remains is a reaction firing for a non-event and an operator-visible log
line (`Specified X → todo`) that is untrue for a parked card.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Tasks awaiting manual plan approval are no longer automatically
planned, reviewed, started, or released into active work.
* Approval-held items are consistently skipped across planning
continuations and related workflows.
* Approval-held due work is deferred to prevent starvation while
waiting, and operator force-promotion still bypasses the gate.
* **Tests**
* Added regression coverage to ensure the manual approval hold behavior
remains invariant across multiple continuation scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:39:21 -07:00
gsxdsm
319e051c65 U8 PR1: pin the execution-lifecycle ownership ledger (measured: 28 executor-owned dispositions vs 3 graph handbacks) (#2490)
First PR of **U8 — the graph owns execution** (plan
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`,
line ~436). The plan states this unit "is expected to land as several
commits; it must not be attempted as one sweep", and
Execution-note-first: **characterization before ownership moves**. This
is that floor. **No behavior change.**

## Why a ledger and not a refactor

U8's goal is "the executor stops deciding *what happens next*" — and
that had no measurable form.

- **Executor line count does not measure it.** A 3,178-line
`runImplementation` can shrink substantially with every lifecycle
decision still exactly where it was.
- **A green suite measures it least of all.** Every disposition counted
below already has passing tests, because each one was *correct behavior*
when it was written. What is wrong is the **owner**, not the behavior.

So the unit needs a number, and the number has to exist *before* the
migration — a ratchet written afterwards cannot prove the migration
happened.

## The measured baseline

Counted from source, comments stripped, method bodies extracted by brace
matching:

| Method | `store.moveTask` | `handoffTaskToReview` | terminal
`status:"failed"` | `graphCompletion` handbacks |
|---|---:|---:|---:|---:|
| `runImplementation` (3,178 lines) | 16 | 3 | 9 | **3** |
| `handleGraphFailure` (~930 lines) | 0 | 0 | 7 | — |

**The implementation phase decides its own lifecycle 28 times and asks
the graph 3 times.**

These are measured, not estimated. My first `handleGraphFailure`
estimate was **wrong** (2 moves / 4 parks); the extractor corrected it
to 0 / 7 — the `moveTask` calls that read as belonging to that method
sit past its closing brace, in the recovery helpers below it. The
correction is in the ledger comment so the next reader does not repeat
the misread.

## The finding this makes concrete

`createAuthoritativeWorkflowSeams.execute` collapses that entire
implementation phase to one boolean:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
```

The graph has no vocabulary for *"the agent stopped because a step is
blocked on a pending review"* or *"the session paused after the work was
already complete"*. So the implementation phase performs those
transitions itself (`executor-exit-while-review-pending`,
`paused-after-completion`) and the graph finds out afterwards.

That is why `handleGraphFailure` carries `alreadyFinalizedToReview` /
`completionFinalized` — **classifiers whose entire job is to recognise a
move the graph did not make.** They are compensation for dual ownership,
and they are U8's acceptance test: they become unreachable, and then
deletable, exactly when the last out-of-band transition is gone. This PR
records that contract in source at the seam (FNXC comment), which is
where the next PR starts.

## Proof the guard fails on the defect

A ratchet that reports success without checking anything is worse than
no ratchet. Both failure modes were injected and observed:

1. **The defect it exists to catch** — injected one `await
this.store.moveTask(task.id, "in-review", {})` into
`runImplementation`'s completion path → ledger fails, `16 -> 17`.
2. **A broken guard** — injected a string literal containing `}` so
naive brace matching ends the body early → the size self-check fails at
**13 lines**, instead of silently reporting a comfortable zero for every
count.

Both injections were reverted; `git diff` against the pre-injection copy
is empty.

## Direction of travel

Executor-owned counts may only go **down**, and a decrement must land
with the disposition visible as a **graph outcome** — not merely
deleted. An increment is a new out-of-graph lifecycle decision and needs
a stated justification in its PR, not a quiet edit to the constant.

This is the precursor to U12's planned
`no-out-of-graph-lifecycle-writes.test.ts`; when the counts reach their
floor the assertion becomes "zero, outside the allowlist", and this file
is where that allowlist grows up.

## Preserved behaviors

Untouched, and re-run green as the regression floor for everything that
follows: FN-8141 honest-blocked exit
(`executor-task-done-blocked.test.ts`), FN-7996/FN-7998 tool-failure
retry + escalation (`executor-tool-failure-retry.test.ts`), FN-7863
dispatch-loop terminalization and FN-7926 completed-blocked parking
(`executor-graph-requeue-gate.test.ts`).

## Verification

- `pnpm --filter @fusion/engine exec vitest run` on the ledger + the
four preserved-behavior suites + `legacy-tombstones` — **6 files, 49
tests, green**
- `pnpm test:gate` — **green** (2/10, 16/299, 1/71)
- `pnpm lint` — clean; `tsc --noEmit` on `@fusion/engine` — clean

No changeset: test-only plus a source comment, no `@runfusion/fusion`
behavior change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added a lifecycle-ownership “source-scanning” test that analyzes the
executor’s task disposition patterns to ensure counts remain consistent
across execution and graph-failure flows.
  * Added safeguards to catch unintended changes to lifecycle handling.

* **Documentation**
* Documented the lifecycle-ownership boundary for task disposition
handling, including how completion and failure transitions are
consolidated and how related failure classifiers are affected.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:59 -07:00
gsxdsm
7871b28766 fix(core): bind the in-transaction capacity gate — one shared pool-id convention (NOT user-visible yet — see R2) (#2488)
## The bug

`moves.ts` asked `countActiveInCapacitySlotAsync` for occupants of pool
`"builtin:coding"`, while the counter buckets selection-less rows under
`DEFAULT_WORKFLOW_POOL_ID` (`"__default-workflow__"`). Nothing ever
landed in the pool being asked about, so the count came back **0** and a
finite limit could never bind.

## Root fix, not a literal swap

A shared *constant* would not have prevented this:
**`DEFAULT_WORKFLOW_ID` was already imported in `moves.ts` and the code
still wrote a literal.** So both sides now call a shared **function**,
`resolveCapacityPoolId` — "which pool does a selection-less task belong
to" has exactly one answer and no call site is in a position to disagree
with it.

The one variable serving two masters is split: a capacity **pool key**
(a bucketing sentinel that must not collide with a workflow id) and a
**workflow id** (telemetry, must stay a real id). The emitted
`TaskTransitioned` payload is byte-identical.

## Checked, not assumed: no second copy

`scheduler.ts:2514` and `:2536` do carry `?? "builtin:coding"` — but as
an **IR resolution key** (`resolveWorkflowIrById`), where a real
workflow id is required and the pool sentinel would not resolve at all.
Same literal, different concept, correctly used. A blanket replace would
have broken it.

## Something did depend on the gate being dead — exactly one thing

`move-path-equivalence.pg.test.ts` → *"UNPROVEN: in-transaction column
capacity did NOT reject on EITHER path in this fixture"*. It left the
cause open —

> something further in (`resolveColumnCapacity`'s limit resolution, or
what `countActiveInCapacitySlotAsync` counts as an occupant — a task
with no session/agent may not count) keeps the check from firing … This
suite does not establish which.

— and predicted its own obsolescence (*"if a future change makes this
reject, that is the capacity gate coming alive"*). **Neither guess was
right; it was the pool id.** Updated to assert the divergence with the
answer recorded — **not weakened**. Its fixture also had to start each
phase from an empty wip column: once the gate binds, the inline phase's
leftovers trip the cap on the *holder* move before the contended move
under test runs.

`schema-applier.test.ts` failed only in the full-suite run and passes in
isolation both with and without the fix — cross-file contamination, not
mine.

## Before / after — measured, both directions

`maxConcurrent: 1`, real PG store, real `moveTask`:

| | flagOFF / no selection | flagOFF / selection | flagON / no selection
| flagON / selection |
|---|---|---|---|---|
| **before** | ADMITTED | ADMITTED | **ADMITTED** ← the bug | REJECTED |
| **after** | ADMITTED | ADMITTED | **REJECTED** | REJECTED |

The E2E acceptance row asserts **held at cap 1 and admitted at cap 2 on
the same fixture**, so it cannot pass by simply never admitting
anything. **With the fix reverted that row fails**; the `admitted` case
still passes, as it should. The Phase A3 ratchet's two flipped
assertions also fail with the fix reverted.

Ratchet flipped exactly as its author specified: `DEFECT (R1)` becomes a
rejection, and `it.fails` on the invariant becomes a plain `it`.

## ⚠️ This is NOT user-visible yet — please read before merging

The premise this was approved on ("once it binds, cards that currently
slip through will start being held") **does not hold for this change
alone.** The whole capacity block sits inside `if (useWorkflow &&
workflowIr && fromColumn !== toColumn)`, and `useWorkflow` is
`experimentalFeatures.workflowColumns === true` — absent from
`DEFAULT_GLOBAL_SETTINGS`, with **no writer anywhere outside tests**.
That is Phase A3's R2, still live and now retitled `DEFECT (R2, STILL
LIVE)` with the measured matrix recorded in it.

So on merge: nothing changes for any real project. Making it actually
bind means **also** removing the `useWorkflow` condition — a materially
larger, genuinely user-visible change that I have not made unilaterally.
Escalated for a decision; if that lands, the changeset here should be
re-categorised.


## Review follow-up (48e79ffd9): the convention was still duplicated —
swept and ratcheted

The first pass added the resolver and routed the transactional gate +
counters, but **hold-release still derived the pool independently**.
Swept the repo: six sites name the sentinel, **five derive the
convention** and now call `resolveCapacityPoolId`
(`hold-release.ts:116/118/442/576`, `task-store-helpers.ts:290`). The
sixth, `scheduler.ts:1558`, names the default pool as a literal in a
capacity *diagnostic* — no selection input, nothing to disagree with —
so it keeps the constant.

**Does this change hold-release behavior? No, and it was never releasing
against the wrong pool.** hold-release computed `x ??
DEFAULT_WORKFLOW_POOL_ID`, which is exactly what the counter buckets
under; `moves.ts` (`?? "builtin:coding"`) was the sole disagreeing site,
and the first commit moved *it* into agreement with hold-release, not
the reverse. `resolveCapacityPoolId(x)` **is** `x ??
DEFAULT_WORKFLOW_POOL_ID`, so every routed site computes an identical
value for every input. **No second user-visible change rides along with
this PR** — the only behavior delta remains the gate binding on the
flag-ON path, which per R2 is still not the path production takes.
Evidence: hold-release + capacity suites **43/43 identical before and
after**.

**The resolver is now the only way to compute a pool id, not merely the
newest way.** `scripts/check-capacity-pool-id.mjs` fails on any inline
`?? DEFAULT_WORKFLOW_POOL_ID` outside `workflow-capacity.ts`, wired into
**both `pretest` and the blocking `test:gate`**. A review note would not
have sufficed: the original defect landed in a file that *already
imported* the canonical constant. Verified both ways — clean run scans
1124 files and passes; reintroducing the old hold-release expression
exits 1 and names the line.


## Review follow-up (a5b675503): the ratchet was rebuilt because it
would not have caught the bug

The first ratchet matched one spelling (`?? DEFAULT_WORKFLOW_POOL_ID`)
and the real defect used another (`?? "builtin:coding"`). **Verified:
reintroducing the original defect and running the old checker exits 0.**
A guard that reports success without checking is worse than no guard —
it stops anyone looking.

Rebuilt on the TypeScript AST with two rules. **Rule 1 (sink):** a value
reaching a capacity counter's `workflowId` must come from
`resolveCapacityPoolId`, or a local initialized from it — so it fires on
the original defect regardless of which literal was used, on one line or
twenty. **Rule 2 (sentinel):** no `??` onto the sentinel at any
qualification depth or as its raw value; multiline is one AST node and
caught by construction. `?? "builtin:coding"` is deliberately *not*
banned outright — it is the legitimate default for a *workflow* id in ~8
places, and is only a bug when it reaches a capacity pool.

**Fails closed three ways** that previously reported success without
inspecting: unreadable file, unparseable file, and an empty file listing
(the old script would have printed a green tick off a broken glob).

**Acceptance was not "passes on main".** Each form was reintroduced into
the real source and confirmed to fail: the original defect in
`moves.ts`, a multiline fallback, and a deeply qualified sentinel. All
are pinned in `capacity-pool-id-check.test.ts` (12 cases: 7 must-catch
starting with the reduced actual pre-fix `moves.ts`, 4 must-not-flag, 1
fail-closed) so the guard cannot silently narrow again.

Also added to `pretest:full`, which had omitted it.


### Follow-up (0be8df6ea): a dead rule found by fixing a test title

Splitting the mislabelled fail-closed test surfaced more than a
mislabel: **`ts.createSourceFile` is error-tolerant and does not throw
on malformed syntax**, so the `try/catch` behind the `unparseable` rule
was unreachable and that rule could never fire. The earlier "fails
closed three ways" claim was overstated — the guard advertised a
capability it did not have. Detection now reads `sf.parseDiagnostics`; a
partial AST can silently lack the `??` nodes and sink calls the rules
look for, so "did not parse" must not read as "inspected and clean".
Mutation-verified: reverting the detection fails that case and only that
case.

Test-file exclusion also moved to the repo's `{test,spec}.{ts,tsx}`
guideline shape — a `.spec.ts` under `packages/<pkg>/src/` was being
scanned as production source. Verified both ways: the `.spec.ts` is
skipped, and the identical content in a non-test file is still caught,
so the exclusion is scoped rather than a hole.

## Verification

- engine + core `tsc --noEmit` clean
- `pnpm test:gate` green (299 + 10 + 71)
- E2E 20/20; capacity + move-path suites 14/14
- full core PG: **1037 passed / 3 failed** — all three reproduce with
the fix stashed (pre-existing)
- engine-default: **279 failed** vs **280 at baseline** with the fix
stashed — pre-existing red lane, no regression
- hold-release + capacity suites: **43/43 identical before and after**
the resolver routing
- `check-capacity-pool-id` ratchet: 14/14 regression cases; clean over
1124 files; exits 1 on the original defect, a multiline fallback, and a
deeply qualified sentinel reintroduced into real source

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed capacity-limit accounting when workflow selection is missing by
consistently deriving the correct capacity pool id.
* Made capacity enforcement align across move and hold/release paths,
rejecting over-limit moves with `capacity-exhausted`.
* **Tests**
* Updated PostgreSQL and added an E2E scenario to verify the corrected
in-transaction gating behavior at `maxConcurrent` limits of 1 and 2.
* **Chores**
* Added an automated guard to detect inconsistent capacity pool id
fallback patterns in code.
* **Public API**
* Exposed `resolveCapacityPoolId` for consistent capacity pool id
derivation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:51 -07:00
gsxdsm
387e836432 U4 substrate PR1: extract the git-evidence readers (19/19 bodies byte-identical, self-healing.ts -449) (#2489)
Based on `main`. **PR1 of the substrate decomposition** — deliberately
small and boring.

## Mechanical proof (the point of this PR)

Every moved body diffed against its pre-move text:

```
BLOCK                            KIND     DIFF-vs-ORIGINAL
  LandedTaskCommit               iface    EMPTY
  commitOwnedByTask              helper   EMPTY
  escapeRegex                    helper   EMPTY
  shellQuote                     helper   EMPTY
  parseShortstat                 helper   EMPTY
  findLandedTaskCommit           method   EMPTY
  findAlreadyMergedTaskCommit    method   EMPTY
  refreshRemoteBaseRef           method   EMPTY
  readCommitTaskOwnership        method   EMPTY
  branchHasNoUniqueDiff          method   EMPTY
  baseHasExplicitTaskOwnership   method   EMPTY
  foreignTipRejection            method   EMPTY
  branchTipForeignOwnership      method   EMPTY
  isCommitReachableFromBranch    method   EMPTY
  findWorktreePathForBranch      method   EMPTY
  repoBranchExists               method   EMPTY
  readShortstatForSha            method   EMPTY
  readLandedFilesForSha          method   EMPTY
  isBranchTipMisboundToTask      method   EMPTY

RESULT: 19/19 BYTE-IDENTICAL modulo the enumerated deviations; 0 differ.
```

No condition reordered, no signature changed, no rename, no inlining, no
"while I am here" cleanup.

## The premise needed correcting — the cheap lesson this cut was meant
to buy

The brief described **13 pure functions**. They are 13 `private async`
**methods** closing over `this.options`, so nothing here is
byte-identical in the strict sense. Three deviations were structurally
unavoidable:

1. **`private` → `protected`** on the 14 methods and the `options` field
— a subclass cannot call a `private` base member. One token per
declaration.
2. **Two type annotations** rewritten:
`SelfHealingManager["readCommitTaskOwnership"]` → the base class, since
the original would be a circular import.
3. **Four module-level helpers moved along** and re-exported. Forced by
direction — the new module must not import `self-healing.ts`, so
everything the bodies call has to live beside them. `execAsync` and
`shellQuote` are imported back because call sites remain.

## The set is 14, not 13

`foreignTipRejection` had to come too, and it's what **closes** the
cluster — it depends only on `baseHasExplicitTaskOwnership` and
`branchHasNoUniqueDiff`, both already in the set. Without it the
extraction isn't self-contained and `this.options` isn't the only
external dependency.

## Base class, not free functions — deliberately

Converting to free functions would change all 14 signatures: a
behavior-adjacent edit riding inside a file move, which is exactly the
combination that hid the last four safeguard regressions. An abstract
base keeps `this` semantics, so every call site stays
`this.<method>(...)` and every body is unchanged text.

## Two pre-existing collisions, preserved exactly

- The detector module already exports a **free**
`findAlreadyMergedTaskCommit` sharing a name with the protected method;
inside the class body the bare identifier resolves to the import.
- `SelfHealingOptions` is declared in `self-healing.ts`, imported here
**type-only** so it's erased at runtime and creates no module cycle.

## Measured line delta — not an estimate

| | lines |
|---|---:|
| `self-healing.ts` | 13,394 → 12,945 = **−449** |
| new module | **+527** (454 moved verbatim, 73 scaffold) |
| **NET** | **+78** |

Same shape as every consolidation in this program: the target file
shrinks, the total grows slightly. At −449 for the first and safest cut,
the remaining substrate cuts plausibly take `self-healing.ts` under 11k
— but that is **file-size reduction, not code reduction**.

## Verification

464 passed across three engine suites with the single known
**pre-existing** `archiveStaleDoneTasks` failure; tsc clean; lint clean;
merge gate green (299 + 10 + 71).

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 20:15:46 -07:00
gsxdsm
0021bd363e U4: retire the surfacing family onto one policy-driven runner (measured: trim is ~34% of the estimate) (#2487)
Stacked on #2486. Base is `feature/u4-safeguard-scoped-to-mutation` — do
not merge before it.

Retires the surfacing family — `surfaceStalePausedTodos`,
`surfaceStalePausedReviews`, `surfaceInReviewStalled` — onto one
policy-driven runner. Migrated **together**, because they were three
copies of one skeleton and a fix applied to one had to be remembered for
the other two.

#2484's characterization suite is the regression floor and **passes
unchanged**.

## Two unconverted sites found in core

Without these the migration would have been **cosmetic for two of the
three**:

`getStalePausedReviewSignal` and `getInReviewStalledSignal` both
hardcoded `task.column !== "in-review"`. `getStalePausedTodoSignal`
gained the equivalent `holdColumn` parameter back in B1 — its two
siblings were missed, so they silently stopped matching for any workflow
that renames its review column. Both now take `reviewColumn`, defaulting
to the legacy id.

## Three bugs this work introduced, and the tests caught

Each was silent in the diff:

1. **I dropped the engine activation floor** from all three signal
calls. Wall-clock the engine wasn't running for is not quiet time, so
every sweep would have reported cards as stale purely because the engine
restarted. Caught by the **pre-existing** suite — not by my own
characterization floor, which is exactly why that floor wasn't
sufficient alone.
2. **I passed `task.column` as the role column**, making the signal's
own column check compare a value against itself — tautologically true,
so the check was silently deleted. The *resolved* role column is now
passed.
3. **The role gate was conditional on the role resolving**, so it was
silently absent for exactly the workflows whose role failed to resolve.
Now unconditional, with the legacy id as fallback.

## The shared test is a table

One row per sweep; every invariant asserted for all three — threshold
inheritance, declared-policy override, renamed role column, the negative
case outside the role column, at-most-once, activation floor, both pause
gates, non-positive-threshold disable, soft-delete. **Adding a fourth
surfacing sweep means adding a row.**

Its log mock **appends** to the card's log, because the at-most-once
dedup reads that log — a call-recording mock cannot observe suppression
at all.

## Measured line delta — this corrects the survey estimate downward

| | lines |
|---|---:|
| `self-healing.ts` | **−192 / +95 = −97** |
| new runner file | **+209** |
| core signal conversions | **+23** |
| **NET** | **+135** |

Three sweeps migrated **increased** total lines by 135 while shrinking
`self-healing.ts` by 97. Per sweep: **−32** in `self-healing.ts`.
**Break-even is 6.5 sweeps.**

Extrapolated to all 34 POLICY sweeps: **−1,099** in `self-healing.ts`,
**−890 net** once the runner is amortized.

My survey estimated **−2,612** for that bucket. The measured figure is
**34% of it**. The reconciler's size was never the issue — the estimate
assumed per-sweep bodies collapse to almost nothing, and they do not:
each keeps a real eligibility predicate, signal call, and operator
message.

## Verification

- 33 shared-family tests green
- 471 passed across five engine suites, with the single known
**pre-existing** `archiveStaleDoneTasks` failure
- 44 core signal tests green
- tsc clean in core and engine; lint clean; merge gate green (299 + 10 +
71)

No changeset: `@fusion/core` and `@fusion/engine` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added support for overriding which workflow column is treated as the
relevant “in review” column for stale-paused and stalled-review
surfacing.

- **Bug Fixes**
- Tightened safety safeguards so only lifecycle-mutating recovery
actions are blocked when a card is user-paused; observational surfacing
remains allowed.
- Improved surfacing consistency across stale paused todos, stale paused
reviews, and in-review stalled cases, including stronger deduping and
cycle-aware behavior.

- **Tests**
- Expanded and reworked safety and surfacing “family” coverage to verify
the new invariants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 19:49:05 -07:00
dependabot[bot]
af3d78d2cd chore(deps-dev): bump vitest from 4.1.8 to 4.1.10 (#2446)
Bumps
[vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest)
from 4.1.8 to 4.1.10.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/vitest-dev/vitest/releases">vitest's
releases</a>.</em></p>
<blockquote>
<h2>v4.1.10</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li><strong>browser</strong>: Check fs access in builtin commands
[backport to v4]  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>OpenCode
(claude-opus-4-8)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10680">vitest-dev/vitest#10680</a>
<a href="https://github.com/vitest-dev/vitest/commit/5c18dd267"><!-- raw
HTML omitted -->(5c18d)<!-- raw HTML omitted --></a></li>
<li><strong>vm</strong>: Fix external module resolve error with deps
optimizer query for encoded URI [backport to v4]  -  by <a
href="https://github.com/SveLil"><code>@​SveLil</code></a> and <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10661">vitest-dev/vitest#10661</a>
<a href="https://github.com/vitest-dev/vitest/commit/bae52b511"><!-- raw
HTML omitted -->(bae52)<!-- raw HTML omitted --></a></li>
</ul>
<h5>    <a
href="https://github.com/vitest-dev/vitest/compare/v4.1.9...v4.1.10">View
changes on GitHub</a></h5>
<h2>v4.1.9</h2>
<h3>🐞 Bug Fixes</h3>
<ul>
<li>Fix <code>importOriginal</code> with optimizer and query import
[backport to v4] - by <strong>Hiroshi Ogawa</strong>, <strong>David
Harris</strong>, <strong>Codex</strong>and <strong>Vladimir</strong> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/10546">vitest-dev/vitest#10546</a>
<a href="https://github.com/vitest-dev/vitest/commit/a5180190c"><!-- raw
HTML omitted -->(a5180)<!-- raw HTML omitted --></a></li>
<li><strong>browser</strong>:
<ul>
<li>Wait for orchestrator readiness before resolving browser sessions
[backport to v4] - by <strong>Vladimir</strong> and <strong>Séamus
O'Connor</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10555">vitest-dev/vitest#10555</a>
<a href="https://github.com/vitest-dev/vitest/commit/7fb29651a"><!-- raw
HTML omitted -->(7fb29)<!-- raw HTML omitted --></a></li>
<li>Wait for iframe tester readiness before preparing [backport to v4] -
by <strong>Vladimir</strong> and <strong>Séamus O'Connor</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10497">vitest-dev/vitest#10497</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10556">vitest-dev/vitest#10556</a>
<a href="https://github.com/vitest-dev/vitest/commit/fbc626c40"><!-- raw
HTML omitted -->(fbc62)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>mocker</strong>:
<ul>
<li>Hoist vi.mock() for vite-plus/test imports [backport to v4] - by
<strong>Hiroshi Ogawa</strong>, <strong>LongYinan</strong>,
<strong>Claude Opus 4.8</strong> and <strong>Vladimir</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10548">vitest-dev/vitest#10548</a>
<a href="https://github.com/vitest-dev/vitest/commit/2c9559c02"><!-- raw
HTML omitted -->(2c955)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>pool</strong>:
<ul>
<li>Prevent test run hang on worker crash [backport to v4] - by
<strong>Ari Perkkiö</strong> and <strong>Jattioui Ismail</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10543">vitest-dev/vitest#10543</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10564">vitest-dev/vitest#10564</a>
<a href="https://github.com/vitest-dev/vitest/commit/934b0f587"><!-- raw
HTML omitted -->(934b0)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<h5><a
href="https://github.com/vitest-dev/vitest/compare/v4.1.8...v4.1.9">View
changes on GitHub</a></h5>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="db616d227b"><code>db616d2</code></a>
chore: release v4.1.10 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10718">#10718</a>)</li>
<li><a
href="bae52b5112"><code>bae52b5</code></a>
fix(vm): fix external module resolve error with deps optimizer query for
enco...</li>
<li><a
href="a7a61e78c7"><code>a7a61e7</code></a>
chore: release v4.1.9 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10598">#10598</a>)</li>
<li><a
href="934b0f587c"><code>934b0f5</code></a>
fix(pool): prevent test run hang on worker crash (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/10543">#10543</a>)
[backport to v4] (#...</li>
<li><a
href="7fb29651af"><code>7fb2965</code></a>
fix(browser): wait for orchestrator readiness before resolving browser
sessio...</li>
<li><a
href="a5180190c1"><code>a518019</code></a>
fix: fix <code>importOriginal</code> with optimizer and query import
[backport to v4] (#...</li>
<li>See full diff in <a
href="https://github.com/vitest-dev/vitest/commits/v4.1.10/packages/vitest">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
2026-07-27 19:40:08 -07:00
dependabot[bot]
ff58df22e4 chore(deps-dev): bump @vitest/coverage-v8 from 4.1.8 to 4.1.10 (#2445)
Bumps
[@vitest/coverage-v8](https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8)
from 4.1.8 to 4.1.10.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/vitest-dev/vitest/releases">@​vitest/coverage-v8's
releases</a>.</em></p>
<blockquote>
<h2>v4.1.10</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li><strong>browser</strong>: Check fs access in builtin commands
[backport to v4]  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>OpenCode
(claude-opus-4-8)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10680">vitest-dev/vitest#10680</a>
<a href="https://github.com/vitest-dev/vitest/commit/5c18dd267"><!-- raw
HTML omitted -->(5c18d)<!-- raw HTML omitted --></a></li>
<li><strong>vm</strong>: Fix external module resolve error with deps
optimizer query for encoded URI [backport to v4]  -  by <a
href="https://github.com/SveLil"><code>@​SveLil</code></a> and <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10661">vitest-dev/vitest#10661</a>
<a href="https://github.com/vitest-dev/vitest/commit/bae52b511"><!-- raw
HTML omitted -->(bae52)<!-- raw HTML omitted --></a></li>
</ul>
<h5>    <a
href="https://github.com/vitest-dev/vitest/compare/v4.1.9...v4.1.10">View
changes on GitHub</a></h5>
<h2>v4.1.9</h2>
<h3>🐞 Bug Fixes</h3>
<ul>
<li>Fix <code>importOriginal</code> with optimizer and query import
[backport to v4] - by <strong>Hiroshi Ogawa</strong>, <strong>David
Harris</strong>, <strong>Codex</strong>and <strong>Vladimir</strong> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/10546">vitest-dev/vitest#10546</a>
<a href="https://github.com/vitest-dev/vitest/commit/a5180190c"><!-- raw
HTML omitted -->(a5180)<!-- raw HTML omitted --></a></li>
<li><strong>browser</strong>:
<ul>
<li>Wait for orchestrator readiness before resolving browser sessions
[backport to v4] - by <strong>Vladimir</strong> and <strong>Séamus
O'Connor</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10555">vitest-dev/vitest#10555</a>
<a href="https://github.com/vitest-dev/vitest/commit/7fb29651a"><!-- raw
HTML omitted -->(7fb29)<!-- raw HTML omitted --></a></li>
<li>Wait for iframe tester readiness before preparing [backport to v4] -
by <strong>Vladimir</strong> and <strong>Séamus O'Connor</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10497">vitest-dev/vitest#10497</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10556">vitest-dev/vitest#10556</a>
<a href="https://github.com/vitest-dev/vitest/commit/fbc626c40"><!-- raw
HTML omitted -->(fbc62)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>mocker</strong>:
<ul>
<li>Hoist vi.mock() for vite-plus/test imports [backport to v4] - by
<strong>Hiroshi Ogawa</strong>, <strong>LongYinan</strong>,
<strong>Claude Opus 4.8</strong> and <strong>Vladimir</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10548">vitest-dev/vitest#10548</a>
<a href="https://github.com/vitest-dev/vitest/commit/2c9559c02"><!-- raw
HTML omitted -->(2c955)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>pool</strong>:
<ul>
<li>Prevent test run hang on worker crash [backport to v4] - by
<strong>Ari Perkkiö</strong> and <strong>Jattioui Ismail</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10543">vitest-dev/vitest#10543</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/10564">vitest-dev/vitest#10564</a>
<a href="https://github.com/vitest-dev/vitest/commit/934b0f587"><!-- raw
HTML omitted -->(934b0)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<h5><a
href="https://github.com/vitest-dev/vitest/compare/v4.1.8...v4.1.9">View
changes on GitHub</a></h5>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="db616d227b"><code>db616d2</code></a>
chore: release v4.1.10 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8/issues/10718">#10718</a>)</li>
<li><a
href="a7a61e78c7"><code>a7a61e7</code></a>
chore: release v4.1.9 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/coverage-v8/issues/10598">#10598</a>)</li>
<li>See full diff in <a
href="https://github.com/vitest-dev/vitest/commits/v4.1.10/packages/coverage-v8">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
2026-07-27 19:21:32 -07:00
Phil Larson
3822e1081c test(engine): remove orphan dependency reporter coverage (#2483)
## Summary
- remove the remaining dependency-blocked reporter test after the
reporter was deleted in #2477
- prevent Vitest from failing during module collection on a deleted
import

## Test plan
- `pnpm --filter @fusion/engine typecheck`
- `pnpm --filter @fusion/engine build`
2026-07-27 18:36:51 -07:00
gsxdsm
17fdf2f0ae U4: scope the user-pause safeguard to lifecycle MUTATION, not observation (re-ratified) (#2486)
Stacked on #2482. Base is
`feature/workflow-vocabulary-u4-override-layer` — do not merge before
it.

Implements the coordinator's **re-ratification** of the user-pause
safeguard with a narrower definition.

## The invariant, written into the code

> The user-pause safeguard means **NEVER MUTATE LIFECYCLE STATE** of a
user-paused card. It does **NOT** mean never observe one.

Respecting a pause exists to stop the engine acting on a card *behind*
the operator who paused it — moving, rebounding, archiving, resuming. A
read-only diagnostic does the opposite: it tells that same operator what
their paused card is doing. Blinding them to their own paused work is
not safety; it is the engine deciding they should not be told.

The sentence is in the source, because the distinction is the whole
point and a future reader will otherwise re-broaden it.

## Why it needed narrowing

Ratified broadly first, that reading was caught suppressing the very
sweeps it was meant to protect. `surfaceStalePausedTodos` exists to
report cards that have sat paused too long — routing it through a
reconciler that suppresses paused cards turns a diagnostic into one that
**silently reports nothing**. Measured, not argued: #2484 proves that
sweep surfaces user-paused cards on current main.

## Scoped by action, never by sweep

`OBSERVATIONAL_ACTIONS` is an **allow-list**, so a newly added mutating
action is suppressed by default — the scoping fails closed. A sweep
cannot opt itself out.

`RecoveryActionKind` deliberately names actions the policy vocabulary
cannot yet author (`rebound`, `archive`, `requeue`, `resume`). The
scoping is only testable if those exist as values, and **a rule that
cannot be tested is a rule that erodes**. `parseWorkflowIr` keeps a
closed action list, so nothing becomes authorable by being named.

## Which field — chosen, not inherited

The gate reads `userPaused`, **not** `paused`. The two diverge
(`branch-group-ops.ts:128` says so outright), and `paused` also covers
engine-authored automation pauses like dispatch-storm, which carry no
operator intent to respect — gating on it would suppress recovery from
the engine's own throttles. The safeguard defers to a **human**
decision, so it keys on the field that records one.

## Both halves kept

The broad case is **narrowed, not deleted**: one test proves mutation is
still suppressed, one proves observation is now permitted. A future
reader must be able to tell the scoping was *deliberate* rather than
eroded by someone who found the broad rule inconvenient.

**Mutation-verified in both directions**, since either error is silent:

| mutation | result |
|---|---|
| re-broaden (suppress observation) | **2 tests fail** |
| over-narrow (`rebound` treated observational) | **3 tests fail** |

## Verification

46 tests green (28 safety + 18 inheritance); tsc clean; lint clean;
merge gate green (299+10+71).

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 17:37:20 -07:00
gsxdsm
89284df85e E2E: table-driven converted-sweep coverage, two new sites, and an honest unproven-sites ledger (#2485)
Follow-up to #2475 (merged). Test-only, plus one test-utility seam.

## Why a table

#2475 proved one converted sweep. The count has since gone to **three**,
twice while this work was open — `surfaceStalePausedTodos` appeared
during #2475's review, and #2478 landed `recovery-reconciler.ts` while
this branch was open. A suite with a bespoke `describe` per sweep is a
coverage claim that quietly becomes false.

Replaced with a table of `(seed, run, acted, roles, observability)`. The
driver derives four assertions per entry:

| | positive | negative |
|---|---|---|
| **renamed vocabulary** | acts on the card | inert in a non-target
column |
| **default vocabulary** | acts (regression floor) | inert |

Adding a converted sweep is **one entry** — the #2478 site proved that
in practice, not in principle. `actsOnRole`/`inertRole` are keys of
`Vocabulary`, not column strings, so an entry cannot hardcode `todo` and
pass for the wrong reason.

## Two findings, both from mutation rather than reading

**1. The census was wrong about `recovery-reconciler.ts:198.`** It was
flagged as a `resolveLifecycleColumns` site, so the row was first
labelled as covering it. **Destroying that role resolution leaves all 18
tests green** — `decideRecovery` looks policy up by *column id* and
never consults a role. The row is relabelled to what it actually proves,
and mutation-verified against that instead: keying the reconciler's
policy lookup on the `todo` literal fails exactly its renamed test.

**2. `resolveRoleRecovery` is an unreachable export.** It is the only
use of `resolveLifecycleColumns` in that file and has **no production
caller anywhere** in engine, core, or dashboard. So that census line is
not a live converted site — it is a helper written ahead of its
consumer. **Not fixed here:** it is production code owned by the U4
slice, and whether the consumer is still to land or it should be deleted
is its author's call.

## Observability is now explicit in the type

`persisted-row` is the strong form. `returned-decision` is recorded as
**weaker evidence** and the reconciler row uses it, because
`reconcileRecovery` decides and does not apply — there is no row to
read. Naming it in the type is what stops a return-value assertion from
quietly passing as observed state, and it is what keeps the ledger
truthful per site.

## Harness seam

`PgTestHarness` now exposes its raw admin SQL client. The store
**stamps** `updatedAt`/`columnMovedAt` on every write, so `updateTask`
cannot express an aged row at all — the patch is accepted and the value
silently replaced with `now`. **Found by the new case failing on BOTH
vocabularies**, which is what distinguishes a broken fixture from a
broken guard. Seeding only; assertions still read back through the real
`getTask` path.

## Mutation verification

| Mutation | Result |
|---|---|
| revert **only** `recoverStrandedCompletedTodoTasks`'s resolution |
**exactly** that row's renamed test fails |
| revert **only** `surfaceStalePausedTodos`'s resolution | **exactly**
its own renamed test fails |
| reconciler policy lookup keyed on `todo` | exactly the reconciler
row's renamed test fails |
| `resolveRoleRecovery` role resolution destroyed | **nothing fails** →
finding #2 |
| `hold-release` `isHeldTask` keyed on `todo` | 5 of 18 fail; default
spine survives |
| `markMoveInFlight` dropped | both spine tests fail |

Per-site verification matters here: three rows could all be riding one
guard. They are not.

## The honest number

**Proven end to end: 5** (two self-healing sweeps, the reconciler's
policy lookup at the weaker observability, hold-release's capacity
release, and the graph boundary + `moveTask` + post-commit bus).

**Not proven: 11 call sites** — `merger.ts:324-326`,
`merger-ai.ts:1022,1039`, `auto-merge-finalization.ts:20-22`,
`executor.ts:1763,6339,6341`, `self-healing.ts:713,6732`,
`mesh-lease-manager.ts:61`, `task-agent-sync.ts:59`,
`core/task-store/reads.ts:130`, `core/live-agent-count.ts:63-75`, and
four dashboard route sites.

The ledger lives in the file, not just here, so it stays with the code.

## Where the table does not fit — reported, not papered over

The **merge/rebound family** cannot be a table row: those sweeps have no
observable persisted effect without a real git repository, so `acted`
cannot be written against the row at all. They need an engine-slow
real-git lane. The dashboard sites need an HTTP route test with a live
store. Both are different lanes, not missing entries.

## Verification

- 18/18 green; engine + core `tsc --noEmit` clean; `pnpm test:gate`
green (299 + 10 + 71)
- full core PG suite run (the harness is shared): 1036 passed, 3 failed
in `central-archive-secrets` and `workflow-settings-project-identity` —
**reproduce identically with this change stashed**, pre-existing

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:25:17 -07:00
gsxdsm
2dfac47917 U4: recovery policy as an OVERRIDE LAYER (unset defers to the operator setting) (#2482)
Stacked on #2478. Base is
`feature/workflow-vocabulary-u4-reconciler-slice` — do not merge before
it.

Implements the ratified precedence: **declared explicitly → policy wins;
left unset → defer to the project/global setting, exactly as today.**

## The design

`resolveEffectiveRecovery(declared, inherited)` composes the two **per
field**, so a workflow may declare a threshold while inheriting the
action. It mirrors the two-tier merge `effective-settings.ts` already
implements for workflow settings (a stored value overrides the base; a
declaration default only fills an absent key) rather than inventing a
fourth precedence system beside model selection, project settings, and
workflow settings.

**Absence stays absent.** `??` treats an explicitly-`undefined` field as
unset, so a policy is never normalized into a built-in default. The
distinction a naive implementation gets wrong:

> **equal-to-default is not the same as unset**

A declaration whose value happens to equal the legacy literal is a
*deliberate choice* and must still override a customized operator
setting. Only true absence defers.

An effective policy requires **both** halves — a threshold with no
action never fires, an action with no threshold has nothing to fire on —
so a half-resolved policy yields `undefined` rather than something
present but inert.

## The test that matters was written first, and failed

> a project with a CUSTOMIZED threshold and the policy key UNSET must
observe the customized value

This is where a green suite lies. "Read the policy, else use the
built-in default" passes every obvious test while silently resetting an
operator who tuned `stalePausedTodoThresholdMs` — no error, nothing in
any diff, the sweep just starts firing on a schedule nobody chose.

**Mutation-verified in both directions:**

| mutation | result |
|---|---|
| substitute a built-in default for the inherited setting | **5 tests
fail** |
| invert precedence (inherited beats declared) | **3 tests fail** |

## Upgrade guarantee

Asserted as a property over several operator values: an undeclared
workflow observes *exactly* the operator's value. That is what makes
landing the policy table a zero-behavior-change upgrade that touches no
project.

## What is NOT here

**`surfaceStalePausedTodos` is not retired.** Migrating it surfaced a
safeguard-semantics collision I escalated rather than resolved
unilaterally: the sweep exists to surface cards that have been
**paused** too long, but the reconciler's ratified user-pause safeguard
suppresses `surface` on user-paused cards — so migrating it as-is would
suppress a large part of what the sweep is for. `paused` and
`userPaused` are distinct fields that can diverge (see
`branch-group-ops.ts:128`). The sweep is untouched pending that
decision.

## Verification

- 34 tests green (10 new inheritance + 24 safety)
- `tsc --noEmit` clean, `pnpm lint` clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Recovery decisions now correctly inherit operator settings when a
workflow does not specify a recovery policy.
- Workflow-specific recovery settings override inherited values,
including when matching built-in defaults.
- Recovery settings can now be applied independently by field, allowing
thresholds and stale-item actions to inherit separately.
- Recovery is suppressed safely when no complete policy is available,
preventing unintended recovery actions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 15:24:19 -07:00
gsxdsm
3578903b16 U4: characterize surfaceStalePausedTodos before the policy migration (regression floor + safeguard evidence) (#2484)
Based on `main`, independent of #2482 — mergeable on its own.

Pins what `surfaceStalePausedTodos` does **today**, so the first real
policy migration is judged against *observed* behavior rather than
against what the sweep looks like it should do. **No production code
changed.** Passes against current main; must still pass after the
migration.

## 1. Threshold source — the case you asked to see proven first

The sweep reads `settings.stalePausedTodoThresholdMs` directly, so under
the override layer an undeclared workflow must keep observing exactly
that value. A migration that reaches for a declaration default instead
finds the card fresh and **silently stops surfacing it** — resetting a
deliberate operator choice with no error and nothing in any diff.

Both directions are asserted, deliberately:
- a customized **1h** threshold surfaces a 2h-old card;
- the **same card is NOT surfaced** under the 24h built-in.

Without that discriminator, "always surface" would pass the first test
and prove nothing about which threshold was used.

## 2. User-paused cards — evidence for the open safeguard question

This replaces my argument with a measurement.

`getStalePausedTodoSignal` gates on `paused === true` and **never
consults `userPaused`**, so a user-paused card that has sat too long
**is surfaced today**. That test passes on current main.

This is what blocks the migration: the reconciler's ratified user-pause
safeguard suppresses the `surface` action for `userPaused` cards, so
routing this sweep through it unchanged would **stop surfacing them** —
a silent behavior change dropping a diagnostic operators rely on.

`paused` and `userPaused` are distinct fields that can diverge (see the
comment at `branch-group-ops.ts:128` about `userPaused` remaining true
while legacy `paused` is false), so these are genuinely separable states
rather than one condition spelled two ways.

**The decision needed:** does the user-pause safeguard mean *"never act
on a user-paused card"* or *"never mutate lifecycle state of one"*? It
generalizes — `surfaceStalePausedReviews` and `surfaceInReviewStalls`
are the same shape.

## 3. Things easiest to lose in a rewrite

Also pinned: inert under `globalPause` and `enginePaused`; a
non-positive threshold disables it entirely; an unpaused card is never
surfaced; automation-paused and user-paused cards are not distinguished
today.

## Verification

8 tests green against unmodified main; lint clean.

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 15:20:22 -07:00
gsxdsm
b133d521c4 U4 vertical slice: recovery-policy reconciler + ratified safety invariant (measured: engine ~780, real cost is a settings migration) (#2478)
Stacked on #2477. Base is
`feature/workflow-vocabulary-u4-delete-dep-blocked` — do not merge
before it.

The smallest end-to-end slice of the U4 reshape, built to **measure**
the real cost before committing to the full policy table. The survey's
~900-line reconciler figure was reasoned, not prototyped; this replaces
it with numbers.

**Everything here is additive and unwired. No behavior changes.**

## What lands

| | lines | what |
|---|---:|---|
| `WorkflowColumnRecovery` (IR) | 42 | one key — `stalenessMs` +
`onStale`. Optional and omitted when unset, so existing workflows
serialize byte-identically. |
| `recovery-reconciler.ts` | 176 | one engine: walks live cards,
resolves each card's policy from **its own** workflow (per task, shared
`irCache` — a 400-card board across three workflows reads three IRs),
returns decisions. Decision and application are separate so the safety
boundary is assertable without running an engine. |
| `recovery-policy-safety.test.ts` | 156 | one-time. The **ratified
invariant**. |

## Measured cost vs the ~900 estimate

**Engine + IR types = 222 lines** for one action (`surface`) and one
safeguard.

Extrapolating the rest — `rebound` (target resolution, attempt budgets,
backward-move proof, five more safeguards) ≈ +350, `archive` ≈ +50, the
`budgets`/`dependencies` keys ≈ +150 — lands near **780**.

So **~900 was a good estimate for the engine**, and the vertical slice
does not move it much. That is the answer to the question asked.

## But the estimate's real miss is not lines

**16 of the 34 POLICY sweeps read an operator setting today** — ~17
distinct policy-threshold keys, including `stalePausedTodoThresholdMs`,
`inReviewStalledThresholdMs`, `taskStuckTimeoutMs`,
`doneAutoArchiveDays`, `maxPostReviewFixes`.

Moving those sweeps into workflow policy is **not a code refactor — it
is a settings migration with operator-visible blast radius**, and it
needs three decisions the line estimate never surfaced:

1. Does workflow policy **override** the global setting, or defer to it?
2. What happens to **existing projects** that already configured those
settings?
3. Does an **unset** policy inherit the setting, or the built-in
default?

That is the gating question for the full table — not the reconciler's
size.

## Why the sweep is not retired here

Retiring `surfaceStalePausedTodos` requires builtin:coding to declare
the policy **and** `stalePausedTodoThresholdMs` to migrate — or the
behavior silently disappears for every existing project. That is the
settings migration above, and it belongs behind its own decision rather
than smuggled into a measurement slice.

The reconciler is therefore **unwired — deliberately dead code**, for
exactly as long as it takes to get that decision.

## The ratified safety invariant

The six safeguards (user pause, `autoMerge:false`, dependency, capacity,
merge-proof, at-most-once) live **outside** the policy table. A workflow
must never be able to author a safety invariant away.

Encoded two ways, because either alone is defeatable:
- **structural** — the policy exposes only an allow-listed key set;
adding a key requires editing the test and re-stating the safety
argument (the friction is the point);
- **behavioral** — a policy attempting every spelling of "ignore the
user pause" has no effect.

**Both halves mutation-verified**, because a safety test that cannot
fail is worse than none:
- making the reconciler honor a policy field that disables the
user-pause safeguard → **fails**
- adding an unreviewed key to the policy schema → **fails**

A third test asserts the reconciler still **acts** on an unpaused card,
so a reconciler that suppressed everything cannot pass by doing nothing.

## Scope limits stated rather than implied

Only the `surface` action is implemented, so only its relevant safeguard
is wired. `surface` mutates no lifecycle state; the other five gate
lifecycle-**mutating** actions that do not exist yet, and wiring them
now would be untestable dead code. A test records this so the absence
reads as deliberate and must be updated when `rebound` lands.

## Verification

- `tsc --noEmit` clean in core and engine; `pnpm lint` clean
- merge gate green (299 + 10 + 71)
- 23 safety tests green; `workflow-lifecycle-traits` green

No changeset: `@fusion/core` and `@fusion/engine` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 15:06:19 -07:00
gsxdsm
2dce642ccc E2E validation: run a RENAMED-column workflow against a live engine (real graph + real PostgreSQL) (#2475)
Stacked on #2472 (`feature/workflow-vocabulary-b3-stranded-todo`).

Test-only. No production file is touched.

## Why

Every slice of this program has closed with the same caveat: *no renamed
workflow was run against a live engine; all evidence is unit-level*.
That caveat is load-bearing — eight times this session a test passed
without exercising its subject. This PR removes it for the lifecycle
spine.

## What actually runs

`packages/engine/src/__tests__/workflow-lifecycle-live-e2e.pg.test.ts`
drives the REAL pieces:

- a **real PostgreSQL `TaskStore`** on a throwaway per-file database
(shared PG harness; never the operator's DB, never port 4040),
- the **real graph interpreter** (`WorkflowGraphTaskRunner`) with the
**real column-boundary controller** wired to the **real
`store.moveTask`** — all of its guards, traits, capacity reservation,
and post-commit emission,
- the **real scheduler release** (`runHoldReleaseSweep`),
- the **real post-commit lifecycle bus** (`getWorkflowEventBus`),
- the **real converted self-healing sweep**
(`SelfHealingManager.recoverStrandedCompletedTodoTasks`, slice B3.1).

Only the AI **seams** are scripted — the same boundary `testMode`/`mock`
draws in production.

**Assertion rule:** every lifecycle claim is asserted on **persisted
state** (a fresh `getTask` with the store's task cache defeated,
`run_audit_events` rows, `workflow_work_items` rows), never on "a
function was called". The one spy — the event-bus subscriber — is
asserted on the **received payload**, because the bus silently drops
events that fail its shape check, so "emit was called" proves nothing.

**Differential design:** the default-vocabulary
(`todo`/`in-progress`/`in-review`/`done`) and renamed-vocabulary
(`backlog`/`building`/`checking`/`shipped`) workflows come from ONE
builder and differ ONLY in their four column ids. Any behavioral delta
is attributable to the vocabulary alone.

## Coverage (9 tests, all green)

| Scenario | What is proven |
|---|---|
| Default vocabulary, full spine | planning runs in the hold column, the
card parks (graph does not self-promote), the **scheduler** performs
hold→wip, the resumed run walks exec → review → merge-gate → end,
persisted column is `done` |
| **Renamed vocabulary, full spine** | identical, and no leg of the run
touches any legacy column id |
| Audit differential | the graph-owned boundary crossings are the same
crossings node-for-node on both vocabularies; no legacy id appears in
the renamed trail |
| Event seam | a real subscriber **receives** a well-formed
`TaskTransitioned` for the renamed `backlog`→`building` release and for
the terminal move; `NodeEntered` arrives for every traversed node
including `end` |
| Crash / restart | exactly one durable continuation row at `exec`; a
brand-new runner resumes from the row and the already-completed
`planning` seam does **not** re-run; no duplicate continuation |
| Converted sweep (B3.1) | a completed card in a **renamed** hold column
is promoted (asserted on its persisted column), a card in the renamed
**wip** column is not, and the default `todo` case still works |

## Mutation verification (both directions)

Green suites are not evidence in this codebase, so both halves were
falsified:

1. Keying `hold-release`'s `isHeldTask` on the `todo` literal → **5 of 6
spine tests fail, and the one that survives is the default-vocabulary
one.** That is the exact signature the conversion program cares about.
2. Reverting slice B3.1's per-task hold-column resolution to the literal
→ **only the renamed stranded-todo test fails**; the default regression
floor stays green.

## Findings surfaced by running it

1. **The IR validator refuses a `merge-blocker` column with no reachable
merge-class node** ("the gate can never clear without one"). Kept rather
than worked around — it means the review column here is genuinely gated.
2. **Entry into the merge region collapses to the legacy `merge` seam**
(`MERGE_REGION_KINDS`), so a `merge-gate` node reaches the merge lane.
Documented in the fixture.
3. **The transition policy refuses a direct hold → review move**, and it
refuses it *workflow-resolved*: on the renamed board the only legal
target is its own `building`, not `in-progress`. The recovery callback
therefore promotes hold → wip → review rather than bypassing the policy.
4. **`moves.ts` still special-cases the `done` literal** (`if (toColumn
=== "done") clearNearDuplicateReferencesTo...`) after the post-commit
emit. Not converted here and not in this PR's scope — flagged for the
Phase B owner.

## Not driven end to end (stated plainly)

- **Triage / specification.** The lifecycle starts from a task already
bound to a workflow; `triage.ts` was not driven. The `planning` seam is
scripted.
- **Real merge.** No git worktree, no branch, no squash. `merge-gate` is
pure policy; the `merge` seam is scripted.
- **Lightweight / self-healing-off workflow.** The Tier 1 policy keys do
not exist on this tip — there is no `policies` surface on the IR to set.
Not drivable; not substituted with a unit test.
- **Process-level crash.** The restart is an in-process one: a brand-new
runner resuming from the persisted `workflow_work_items` row with no
carried-over memory. No OS process was killed, so this proves
durable-state resumption, not signal handling.

## Lane

`.pg.test.ts` under the engine-default include glob, gated by
`pgDescribe` so it skips cleanly with no PostgreSQL. The merge gate is
untouched. Engine `tsc --noEmit` is clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added comprehensive live PostgreSQL workflow lifecycle coverage,
including graph execution, suspension and resume, scheduler capacity
release, crash recovery, and durable continuation.
* Added validation for renamed workflow column configurations and
columnless task movements.
* Added event delivery checks for task transitions and node entry
events.
* Added self-healing recovery for stranded completed tasks in valid hold
columns.

* **Refactor**
* Centralized workflow boundary handling, including task moves,
continuation state, audit events, and diagnostics.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:05:05 -07:00
gsxdsm
710d56b2db U4 trim: delete the dependency-blocked-todo feature (unreachable in production) and revert 5a2de7d (#2477)
Stacked on #2474. Base is `feature/workflow-vocabulary-u4-dead-code` —
do not merge before it.

Deletes an **entire feature that has never executed in production**, and
reverts `5a2de7d`, which only threaded resolved lifecycle columns
through it.

## Reachability evidence — the whole basis for this

```
surfaceDependencyBlockedTodos          ← in NEITHER sweep registry; no caller in
  └─ getDependencyBlockedTodoReporter()      engine/dashboard/cli — only tests
      └─ engine/dependency-blocked-todo-reporter.ts   ← sole caller of ↓
          └─ core/computeDependencyBlockedTodoReport
```

self-healing owns two name-based sweep registries (`runStartupRecovery`,
58 entries; `runMaintenance`, 76). `surfaceDependencyBlockedTodos` is in
**neither**, so nothing ever invoked the chain below it. Its four tests
passed while proving nothing about production.

## Why delete rather than wire it up

Wiring was the tempting option and is the riskier one. Switching on a
450-line path that has never run — whose tests therefore establish
nothing about its behavior against real data — is a **behavior change
with unquantified blast radius**. This program already refused exactly
that move for the **pool-id sentinel**, a one-line change that would
switch on dormant enforcement across every project. This is the same
class of move at ~450× the size.

Deleting is also the recoverable direction: git keeps the feature, and
it can be resurrected deliberately — with tests that prove it *runs* —
if dependency-blocked reporting is actually wanted.

## The settings keys go with it

`dependencyBlockedTodoReportEnabled` defaulted `true` while driving
nothing. A schema/API-visible switch that lies about what the system
does is worse than no switch. (It had no dashboard UI field — the
dashboard test allowlist already recorded it as *"no UI field"*.) Four
sibling tuning keys are removed with it.

## Against my own earlier work

`5a2de7d` threaded resolved lifecycle roles into
`computeDependencyBlockedTodoReport` and its reporter, answering a
review finding I confirmed as real. **The code was correct; the impact
claim was not**, because the path never executes. Neither the reviewer
nor I checked *reachability* before agreeing the defect mattered — only
correctness. A correction is posted on that thread in #2470.

**Scope limit on that admission:** the same finding also described
*incorrect scheduler ordering*. That half runs through
`buildUnblockWeightMap` in `task-priority.ts`, which is **live** and was
already threading `terminalColumns` (B1, `434b385`). Scheduler ordering
was never affected, before or after.

## What survives

`blocker-fanout.ts` **stays** — it is live via `task-priority.ts`. Only
the plural `holdColumns` option added by `5a2de7d` is reverted, since
the deleted report was its sole consumer. `holdColumn` (singular, from
B1) remains.

## Net

**1,244 deletions / 5 insertions across 15 files** — ~450 production
lines, ~684 test lines, 5 settings keys.

## Verification

- `tsc --noEmit` clean in **core, engine, and dashboard-app**; `pnpm
lint` clean
- merge gate green (299 + 10 + 71)
- self-healing suite: 411 passed, 1 **pre-existing** failure
(`archiveStaleDoneTasks`)
- dashboard settings-descriptions suite green
- `settings-parity.test.ts` has one **pre-existing** failure
(`agentToolOutputMaxChars` overlap) that fails identically with these
changes stashed — unrelated to this deletion

No changeset: `@fusion/core` and `@fusion/engine` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added quiet-window backlog health diagnostics for stalled items in
review, with repeat-alert suppression.
  * Added default thresholds for backlog-pressure alerts.

* **Changes**
  * Removed dependency-blocked todo reporting and related alerts.
* Removed the dependency-blocked todo enable/disable setting; remaining
tuning options are no longer active.
* Updated the workflow hold classification to use a single todo column.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:54:32 -07:00
gsxdsm
923a7c0fc9 U4 trim: delete the dead resetStepsIfWorkLost duplicate in self-healing (#2474)
Stacked on #2472 (Phase B slice B3.1). Base is
`feature/workflow-vocabulary-b3-stranded-todo` — do not merge before it.

First output of the **U4 reshape survey**: dead code removed with
reachability evidence, not a redundancy argument.

## What is deleted

**`SelfHealingManager.resetStepsIfWorkLost`** — 50 lines including its
docblock and a section header left with no other member.

Evidence:
- `private`, with **zero in-file references** beyond its own
declaration.
- TypeScript already reported it as `declared but its value is never
read` — the compiler has been flagging this.
- The **live** implementation is `executor.ts:20080`, an independent
copy that `executor.ts` actually calls (12667, 14664). The self-healing
copy is an orphaned duplicate of it.
- Not in either sweep registry, no test of its own, no caller anywhere
in engine / dashboard / cli.

Safe against every U4 constraint: it enforces none of user pause,
`autoMerge:false`, dependency, capacity, merge-proof, or an at-most-once
safeguard.

## Why the second approved deletion is NOT here

`surfaceDependencyBlockedTodos` was approved alongside this one as a
28-line orphan. It is not. On inspection it is the **tip of an entire
unreachable feature**:

```
surfaceDependencyBlockedTodos   (in NEITHER registry, no production caller)
  └─ getDependencyBlockedTodoReporter()        ← sole caller
      └─ engine/dependency-blocked-todo-reporter.ts   223 lines  ← sole caller of ↓
          └─ core/dependency-blocked-todo-report.ts   184 lines
```

≈ **450 production lines across three files, plus 684 lines of tests in
four files.** The operator-visible setting
`dependencyBlockedTodoReportEnabled` defaults `true` in
`settings-schema.ts` and drives nothing.

Deleting only the approved 28-line tip would be **strictly worse than
leaving it** — it orphans the getter and field and strands 407 lines of
module with no remaining reference to explain why. Escalated for a
decision (delete the subtree / wire the feature up / leave it) rather
than resolved unilaterally.

## Verification

- `tsc --noEmit` clean, `pnpm lint` clean
- self-healing suite: **415 passed**, 1 failure **pre-existing**
(`archiveStaleDoneTasks` — fails identically before this change)

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Behavior Changes**
* Removed automatic detection and reset of task steps when no unique
work is found on a task branch.
* Tasks with completed or in-progress steps will no longer be
automatically returned to pending based on this condition.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:44:28 -07:00
gsxdsm
a4dee1162e Phase B slice B3.1 (U4): resolve the hold column in recoverStrandedCompletedTodoTasks — query and guard together (#2472)
Stacked on #2471 (Phase B slice B2). Base is
`feature/workflow-vocabulary-b2` — do not merge before it.

**First landable slice of U4 (self-healing.ts).** One sweep, one PR, per
the phase's sub-split rule.

## The finding: the guard and the query must convert together

`recoverStrandedCompletedTodoTasks` promotes a card whose steps are all
done/skipped but which is still sitting in the hold column — finished
work that never handed off to review. It decided *"is this card in the
hold column?"* **twice**, and both were literal:

| | was |
|---|---|
| the QUERY | `listTasks({ column: "todo", slim: true })` |
| the GUARD | `task.column !== "todo"` |

**Either half alone is a green diff with zero behavior change.** A
correct guard behind a literal query never runs; a converted query
behind a literal guard rejects every row it just fetched. This is the
shape that made B1's stale-paused-todo fix cosmetic, and the phase brief
predicted more of it here — correctly.

I proved it rather than asserting it:

- literal **QUERY** restored (converted guard kept) → **3 tests fail**
- literal **GUARD** restored (converted query kept) → **2 tests fail**

Neither half passes the suite alone.

## Falsification came first

Per the brief I tried to prove the work unnecessary before doing it. It
is necessary, and the evidence is empirical, not assumed: the 7 tests
were written against unmodified code and 3 failed. Unlike B2's
hold-release — which turned out already converted — **self-healing is
uniformly unconverted at the query level**: 53 of its sweeps carry a
hardcoded `column:` filter (survey in the worker report).

## Negative half, per the brief

A completed card resting in a WIP or review column is **not** promoted.
Dropping a column filter without a per-task hold check would promote
finished cards out of every column — laundering work past review, a
louder bug than the silent one being fixed.

## Test-harness hazard (will recur in every remaining U4 slice)

The pre-existing self-healing store mock returns its fixture from
`listTasks` **regardless of arguments**. A renamed-hold test on that
harness passes while the query stays hardcoded, because the mock hands
the sweep rows the real store never would. The new harness **honors**
the column filter, and one test asserts the query is no longer scoped to
the literal. This is documented in the new file's header for whoever
writes the next slice.

## Cost

The column filter is gone, so the cheap non-column rejections (paused /
executing / incomplete steps / errored / no-commits / skip-bypass taint)
run **first and synchronously**; only survivors pay an IR resolution,
shared through an `irCache`. A board spanning three workflows resolves
three IRs regardless of card count. `includeArchived: false` preserves
what the column filter did implicitly. The hold column resolves **per
task** — a board spans workflows, and a card in *another* workflow's
hold column must not be promoted.

## One pre-existing assertion changed, deliberately

`self-healing.test.ts` pinned `listTasks` being called with `{ column:
"todo", slim: true }`. That query shape changed on purpose; the
assertion now pins the new one. The behavioral assertions either side of
it (one qualifying card, promoted exactly once) are untouched and still
pass.

## Carried A3 questions — both answered

**Q1 — does the sync/SQLite counter have the same pool-id mismatch?**
**Not applicable: there is no sync counter.**
`occupantsByColumnForWorkflowImpl` and `listWorkflowOccupantTaskIds` are
async/PG-only and throw without an initialized `AsyncDataLayer`; the
sync twin went with the PG cutover. There is no second counter that
could mismatch. The surviving pool-id sentinel sites are
`project-store-ops.ts:767/819` and `moves.ts` — both parked by operator
decision, untouched here.

**Q2 — are custom workflows with an explicit numeric limit affected?**
**No, by design.** `resolveColumnCapacity` gives `config.limit` top
precedence (`configLimit` → `limitSetting` → default-workflow
read-through → `Infinity`), and `resolveWipBudgetColumns` documents that
a column with an explicit numeric limit is **independent — its budget is
itself alone**. Such a column never pools, so there is no pool id to
mismatch. Read-only analysis; no code changed for either question.

## Remaining U4 scope (not in this PR)

214 literal occurrences across ~70 methods; **53 sweeps carry a
query-level column filter**. Hold-gated sweeps still to convert:
`clearStaleBlockedBy`, `reclaimSelfOwnedBranchConflicts`,
`reconcileCompletedTask`, `recoverMergedReviewTasks`,
`recoverStuckMergeDeadlocks`, plus non-query `todo` guards in
`recoverPausedAbortFailures`, `reconcileDependencyBlockingLeases`, and
others. `surfaceStalePausedTodos` was already converted (B1 follow-up)
and is verified intact on this branch.

## Verification

- 7 new tests green; **both mutations kill the suite**
- self-healing suite: 415 passed, **1 failure pre-existing**
(`archiveStaleDoneTasks` — confirmed identical by stashing my changes)
- merge gate green (299 + 10 + 71)
- `tsc --noEmit` clean, `pnpm lint` clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved recovery of completed tasks stranded in workflow-specific
hold columns, including renamed hold columns.
* Preserved recovery for built-in workflows while correctly handling
boards with mixed workflow configurations.
* Prevented recovery for tasks in non-hold columns or with paused,
incomplete, or errored states.
  * Added fallback handling when workflow details cannot be resolved.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:36:32 -07:00
gsxdsm
02b0f4f860 Phase B slice B2: U5 small movers — 12 literal sites converted, plus a negative result on hold-release (#2471)
Stacked on #2470 (Phase B slice B1). Base is
`feature/workflow-vocabulary-conversion` — do not merge before it.

## What this is

Phase B slice B2 — the U5 small movers. **12 literal sites converted,
plus one negative result.**

The plan estimated 36 sites. A survey found 12 genuinely-convertible
ones, and separately found that the plan's headline hold-release
scenario **was already fixed**. Both are reported below rather than
padded into a bigger-looking diff.

## The negative result (commit 1)

The plan named hold-release.ts as a target on the scenario *"release
readiness must hold and release identically for a RENAMED hold column."*
I wrote that test first, to prove it broken.

**It is not broken.** All five assertions passed against unmodified
`hold-release.ts`. U6/KTD-5 had already converted the module —
`isHeldTask`, `resolveReleaseTarget`, and `dependencySatisfied` each
resolve the task's IR. **`hold-release.ts` has no production change in
this PR.**

The tests are kept as a regression floor: the invariant rests on three
independent trait resolutions any of which could be "simplified" back to
a literal, and nothing else covered a renamed vocabulary end-to-end
through the sweep.

**I verified the tests can actually fail.** Mutating `isHeldTask` back
to `task.column === "todo"` kills all five. Without that check, a green
run against unmodified code is indistinguishable from a test asserting
something trivially true.

Two drafting notes kept in the file: the renamed ids deliberately avoid
colliding with any legacy literal, and the first draft's two dependency
tests used a `capacity` hold — which never consults dependencies at all,
so one passed **vacuously**. Both now use a `dependency` hold.

## The 12 conversions

Each was red-green: the renamed-workflow test written first and
**observed failing**, then made to pass.

| Site | Was | Now |
|---|---|---|
| `task-agent-sync` CLEAR_COLUMNS | `{done,archived,todo,triage}` |
resolved complete+archived+hold+intake |
| `task-agent-sync` isParkedTaskColumn | `{todo,triage}` |
`parkedColumns` param (hold+intake) |
| `task-agent-sync` handler branch | `to === "todo" \|\| "triage"` |
resolved parked set |
| `mesh-lease` parked guard | `task.column !== "todo"` | resolved
rebound column |
| `mesh-lease` rebound move | `moveTask(id,"todo")` | resolved rebound
column |
| `mesh-lease` audit decisionPath | `=== "todo" ? … : …` | same resolved
column |
| `mesh-lease` audit newColumn | `… : "todo"` | same resolved column |
| `merger-ai` already-finalized | `=== "done" \|\| "archived"` |
resolved complete+archived |
| `merger-ai` ×4 rebounds | `moveTask(id,"todo")` | shared
`resolveFinalizeReboundColumn` |

Rebound targets all use KTD-10 `resolveReboundTarget` (hold → intake →
first column), the helper `self-healing.ts:714` already uses — reused,
not invented.

## Three findings worth reading

**1. The mesh-lease bug was in the AUDIT, not the move.** The guard and
the audit were *independent* `=== "todo"` comparisons, so `newColumn`
asserted the card landed in `todo` regardless of what the move did. For
a workflow with no `todo` column that produced a lease-recovery trail
naming a nonexistent column — and run-audit is the only post-hoc record
of a lease recovery. Now resolved once and threaded to both, so they are
structurally incapable of disagreeing.

**2. The merger-ai failure mode was not what I predicted.** I expected
the already-finalized guard to fail open and re-merge a finished card.
The red run showed it actually throws `Cannot merge FN-1: task is in
'shipped', must be in 'in-review'` — a hard error blaming the column, on
a task whose real state is "already done". The thing preventing the
re-merge is *itself* a literal in core's `getTaskMergeBlocker`, outside
this slice. Two bugs coinciding, not a design.

**3. A fourth site had to move that wasn't on the list.**
`evaluateParkedAgentTaskLink` calls `isParkedTaskColumn` internally.
Converting only the handler would have left the preservation branch on
legacy ids after the caller resolved a renamed workflow — trading a
stale-link bug for a **worse** dropped-link bug (a live agent's link
cleared mid-run).

## Deliberately NOT converted

Both keep their literals with the reason recorded at the site under a
greppable `DELIBERATE-LITERAL` tag:

- **`hold-release.ts:326` `legacyDependencySatisfied`** — the FN-5719
dual-accept half. Converting makes both halves compute the same answer,
deleting the compatibility signal *and* its divergence detector while
looking like a cleanup.
- **`replan-target.ts` final fallback** — its value is precisely that it
is *not* trait-resolved; resolving it against the workflow is the
stranded-card bug it was written to fix.

⚠️ **The U12 literal ratchet does not exist in the tree yet.** The brief
assumed an allowlist to add entries to; there is none. `grep -rn
DELIBERATE-LITERAL packages/*/src` enumerates the sites it must admit.

## What I could NOT verify

- **One of the four merger-ai rebound sites is untested.** The
`landWorkspaceTask` rebound is verified by inspection and the shared
resolver's unit tests only — `landWorkspaceTask` is only ever *mocked*
(project-engine.test.ts), never executed. Covering it needs a multi-repo
git fixture and a full land run. The **other three are genuinely
exercised** by pre-existing merger-ai.test.ts (lines 676/716/777/895
assert `moveTask("FN-1","todo",…)` through a real git repo) and pass
unchanged — real wiring proof for those.
- **3 of the 9 task-agent-sync tests passed before the conversion too**,
vacuously — the literal handler early-returned and cleared nothing. They
assert nothing about the old code; they are guardrails against the
conversion over-clearing.
- **No renamed workflow was run against a live engine.** All evidence is
unit-level.

## Call sites outside this slice — NOT converted, byte-identical

They keep the legacy defaults: `scheduler.ts:1273`,
`agent-heartbeat.ts:1169/3642`, `self-healing.ts:11600/11665` (all
`evaluateParkedAgentTaskLink`), and `merger.ts:6585` (the sibling
terminal guard). Each is its own Phase C/D surface.

## Behavior changes (not a pure refactor)

For a **renamed** workflow: agent links now actually get cleared on
terminal moves (they never were); lease rebounds land in the resolved
hold column; finalize-blocked rebounds land in the resolved hold column
and their operator-facing task-log lines name the real column;
already-finalized cards short-circuit cleanly instead of throwing.

For **builtin:coding** and any unresolvable workflow: byte-identical.
Every new parameter defaults to the legacy set, and both merger-ai
resolvers fail *soft* to legacy ids in opposite directions — the
terminal guard keeps `done`/`archived` (losing it sends a finished card
into the merge path), the rebound keeps `todo` (abandoning it strands
the card in the merge lane with no owner).

## Verification

- Merge gate **green**: 299 + 10 + 71 tests
- Slice suites **green**: 100 tests across 8 files (new + all
pre-existing neighbours)
- Existing merger suites **green**: 82 tests across 5 files, unchanged
- `tsc --noEmit` clean, `pnpm lint` clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 14:28:53 -07:00
gsxdsm
5d0f1ef631 Phase B slice B1: lifecycle column roles in the U6 policy modules (4 guards, red-green) (#2479)
**Stacked on #2469** → #2468 → #2467. Base is
`feature/workflow-capacity-ground-truth`.

This is **slice B1 of Phase B, not all of Phase B.** Sizing escalation
sent separately; the census is below.

## Why this is a slice

Measured census of code lines referencing a lifecycle column literal
(comments excluded):

| Unit | Files | Sites |
|---|---|---:|
| U4 | `self-healing.ts` | 203 |
| U5 | `executor.ts` 171, `scheduler.ts` 55, `replan-target.ts` 20,
`merger-ai.ts` 5, `hold-release.ts` 4, `mesh-lease-manager.ts` 4,
`task-agent-sync.ts` 3 | 262 |
| U6 | `moves.ts` 34, `default-workflow-hooks.ts` 13, `board-config.ts`
9, `blocker-fanout.ts` 6, `task-priority.ts` 5,
`dependency-blocked-todo-report.ts` 2, `stale-paused-todo.ts` 1 | 70 |
| | **Total** | **535** |

The plan's "~207" counts the guard category only. Under the phase's
non-negotiable rule — a test that **fails before** conversion, per guard
— that is ~200 red-green cycles. Doing it as one sweep would reproduce
exactly the failure this phase exists to prevent: converted guards
nobody proved still fire.

`moves.ts` and `default-workflow-hooks.ts` stay **parked** per the
dispatch constraint (move-path convergence and the pool-id sentinel are
on an operator decision).

## Guards converted (4), each red-green

Every case below was written **first** and observed failing against the
literal implementation.

| Module | Guard | Before → After |
|---|---|---|
| `stale-paused-todo.ts` | stall detection | `column !== "todo"` →
resolved **hold** column |
| `blocker-fanout.ts` | active | `ACTIVE_COLUMNS.has(col)` →
`!terminalColumns.has(col)` |
| `blocker-fanout.ts` | hold-wait metric | `col === "todo"` → resolved
**hold** column |
| `task-priority.ts` | unblock active | `UNBLOCK_ACTIVE_COLUMNS`
**deleted**, folded into the terminal set |

Three of the seven new cases are **regression floors** that pass before
and after. One of them earned its keep immediately: it failed on my own
fixture (`activeCount` vs the public `totalCount`), catching a bad test
rather than bad code — which is the point of asserting the default path
alongside the renamed one.

### The `task-priority` finding

`UNBLOCK_ACTIVE_COLUMNS` and `DONE_COLUMNS` encoded **one concept
twice**, two lines apart, and disagreed for any custom column:
dependency counting treated a `drafting` card as unmet (correct) while
the active check treated it as inactive (wrong), zeroing the blocker's
unblock weight. The enumeration wasn't just legacy-shaped — it
contradicted its own neighbour.

## ⚠️ Behavior change, not a pure refactor

Inverting active from enumeration to exclusion means **a card in a
column that is neither terminal nor in the legacy enum now counts as
active where it previously did not.** That is the plan's stated intent,
but it is a real change for any project already using a custom column —
**Coding (Ideas)' `ideas` column is the in-tree case.** Fan-out counts
and unblock weights for such cards will rise.

## Verification

- Four affected suites green (45 tests), each conversion observed
red→green.
- `pnpm lint`, `tsc --noEmit` (core) green.

**Not verified / not done, stated plainly:**

- **Call sites are not wired.** These modules now *accept* resolved
roles; every parameter still defaults to the legacy set, so at the call
sites the vocabulary is unchanged. A caller that cannot resolve a
workflow keeps literal behavior. Threading `resolveLifecycleColumns`
through `reads.ts` and `self-healing.ts` is follow-on work — until then
the guards are *convertible*, not *converted end-to-end*.
- `dependency-blocked-todo-report.ts` and `board-config.ts` are
untouched in this slice.
- 19 core-suite failures exist on this branch; all confirmed
**pre-existing** by stashing and re-running on a clean tree
(`duplicate-guard`, `log-severity-spam-contract`, `settings-parity`,
`task-delete-caller-attribution`, `settings-defaults`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

**Supersedes #2470**, which GitHub force-closed when its base branch was
deleted by the merge of #2469 and refuses to reopen. Same head branch,
same commits (rebased onto `main`), now based on `main` directly. The
two P1 review threads on #2470 were resolved there — one of them with a
correction noting the threading half landed in code that was
subsequently deleted as a dead feature in #2477.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Dependency and blocker reports now correctly recognize custom hold,
active, and terminal workflow columns.
* Blockers in renamed terminal columns are no longer incorrectly
reported as active.
* Stale paused-task badges and self-healing now work with
workflow-specific hold columns.
* Mixed boards with different workflow column names are handled
consistently.
* Existing default workflow behavior remains compatible, including
fallback handling when workflow details cannot be resolved.

* **Enhancements**
* Reporting and task-priority calculations now support configurable
single or multiple hold and terminal columns.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:19:32 -07:00
gsxdsm
4158cf1ab7 Phase A: workflow-owned lifecycle foundation (U1, U2, U3) (#2467)
Phase A (Foundation) of
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`.
Three units, one commit each. No operator-visible behavior change.

## U1 — Lifecycle-column resolution seam

`resolveLifecycleColumns(ir)` returns `{ intake, hold, wip, review,
complete, archived }` — the first column carrying each trait,
`undefined` for a role no column carries.
`resolveTaskLifecycleColumns(store, taskId, cache?)` is the store-aware
form; the cache is caller-owned so a sweep reads one IR per workflow
rather than one per card.

A v1/column-less IR resolves to `undefined` for the **whole struct**
rather than a struct of undefined roles. A caller must be able to
distinguish "this workflow declares no hold column" (a real shape to
honor) from "no column vocabulary at all" (skip and log) — only the
second licenses conservative fallback.

Nothing consumes the seam yet; Phases B–D convert the ~207 hardcoded
column literals onto it.

## U2 — Delete the pre-cutover parity machinery (delete-only)

**`workflow-columns-settings.ts`** — `isWorkflowColumnsEnabled` had the
body `return true`. Six live call sites branched on it, so every
flag-OFF arm was dead code that read as a supported configuration.
Deleted; surviving side inlined at self-healing's transitionPending
sweep, the scheduler's per-column capacity diagnostic, merge-trait's
policy resolver, the board-workflows payload, two task-workflow routes,
and the CLI TUI's column enrichment.

**`workflow-parity.ts`** — asserted the default workflow's adjacency
*equals* the legacy `VALID_TRANSITIONS`. U11 deliberately breaks that
equality by merging Todo into Planning, so this is not a stale assertion
to update; it is a contract against the target state. Its emitter
(`workflow-parity-observer.ts`) is already a tombstone, so
`getWorkflowParitySummary` and `computeWorkflowColumnsGraduationReport`
aggregated run-audit rows nothing writes and had no caller outside
`TaskStore`. Both store methods go with it.

`flagEnabled` stays on the board-workflows **wire** as a constant `true`
— shipped dashboard clients still branch on it, and changing the
response shape is not a deletion. U10 retires the field once no client
reads it.

The `legacy-tombstones` ratchet is extended to both files plus seven
symbols, each with the reason it is gone.

### ⚠️ Finding: the third listed deletion was NOT dead

The plan also lists "the flag-off inline move path" in
`task-store/moves.ts`. It is **not** deleted, per U2's execution note
("any behavior change found while removing a branch means the branch was
not dead").

That path is gated on `isWorkflowColumnsCompatibilityFlagEnabled`
(`store.ts:38`) — a **different** function from the always-true public
helper. It reads the raw `experimentalFeatures.workflowColumns` setting,
which nothing in production sets (`settings-schema.ts:396` — "no default
flags are emitted"; zero non-test writers; the operator's own
`~/.fusion/settings.json` has no such key). So `useWorkflow` is false
for effectively every real project: the flag-OFF inline side effects are
the **live** default move path and the flag-ON `default-workflow-hooks`
path is the dead one. The code says so itself at `moves.ts:638`.

Deleting that branch would swap every project onto an untravelled code
path — a behavior change, not a deletion.

**Carry this into Phases B and C, stated plainly so the plan's error is
not repeated:**

> **The inline move path in `moves.ts` is LIVE.
`default-workflow-hooks.ts` (the trait-hook path) is DEAD.** KTD-6
asserted the inverse. Until the convergence unit lands, **nothing may
assume trait hooks run** — a guard, sweep, or subscriber written against
`applyDefaultWorkflowMoveEffects` would never fire in production and
would still pass its tests.

Convergence is **not** attempted here. It is its own unit (Phase A2)
with a proper equivalence proof, per operator decision.

### U3's emit point is on the LIVE path — the seam is not born dead

Worth stating explicitly because it is the failure mode that would make
every later subscriber silently never fire: the `TaskTransitioned` emit
is **not** inside the `if (useWorkflow)` branch. That block closes at
`moves.ts:1212`; the emit sits at `:1214`, beside the existing
`store.emit("task:moved", …)`, on the unconditional post-commit path. It
therefore fires on **both** the live inline path and the dead hooks
path, and the convergence unit inherits the obligation to keep it firing
on whichever path survives — same events, same order, same payloads.

The graph-side emitters (`NodeEntered`, `RunSuspended`) carry the same
risk from a different direction: the bus refuses an invalid payload
*silently* by design, so an emitter regression would stop the event with
no test failure. They are asserted end-to-end through the real bus —
"did a subscriber actually receive it", not "was emit called" — because
a spy passes on a refused payload. The `moveTaskInternalImpl` emit does
**not** yet have that end-to-end assertion against a real store move;
that proof belongs to the convergence unit, which has to build the
both-paths fixture anyway.

## U3 — Post-commit event seam with a transactional outbox

**The bus is not a queue, not a transaction participant, and not a
delivery guarantee.** Durable follow-on work uses the transactional
outbox — a `workflow_work_items` row written *inside* the transition
transaction (the shape `createCompletionHandoffWorkflowWork` already
uses). "Emit after commit, let a subscriber enqueue the work" has a
crash window where a process dies between commit and subscriber, leaving
no event *and* no work-item row, so required work is skipped permanently
with nothing to recover from. Post-commit subscribers therefore carry
only losable reactions.

Emission is consequently lossy and isolated by design: a throwing or
rejecting subscriber is caught and logged, cannot roll back the
transition, and cannot stop the others. Deliveries append to one serial
chain, so two transitions on a task deliver in commit order.

The ids/outcomes-only rule is **mechanised, not documented** —
run-audit's equivalent lives only in prose and has been violated
repeatedly. A payload carrying an object body or a prose string is
refused at the emit boundary and never reaches a subscriber or log sink.
It degrades rather than throws: the emitter is post-commit, so a shape
bug must not become a lifecycle failure.

Emit points: `TaskTransitioned` from the single post-commit point in
`moveTaskInternalImpl`; `NodeEntered` and `RunSuspended` from the graph
column boundary, the latter *after* the durable continuation is
persisted so an observed suspension implies a resumable run.

`registerWorkflowEventSubscribers` (engine) is empty on purpose —
U7/U8/U10 move real reactions onto it, each with the characterization
test proving the reaction was non-authoritative first.

## Verification

- `pnpm test:gate` — green (2/10, 16/299, 1/71).
- `pnpm lint`, `pnpm build`, `tsc --noEmit` on core and engine — green.
- U1: 20 tests in `workflow-lifecycle-traits.test.ts`, including the
fully-renamed-workflow case (fails if the resolver falls back to a
literal) and a shared-cache read-count assertion.
- U2: `legacy-tombstones.test.ts` green with the extended ratchet;
`board-workflows`, `merge-trait`, `workflow-graph-executor-parity`, and
move-hook suites green with no expectation edits.
- U3: 20 bus-invariant unit tests (isolation, ordering, the allowed-key
and required-key halves of the ids-only rule, lossiness) plus 3
end-to-end emitter-delivery tests; 5 outbox tests against a **real
PostgreSQL** work-item table (crash survival, rollback, at-least-once
redelivery on lease expiry, idempotent handler → one effect,
dropped-subscriber vs. durable work). A hand-written fake of the lease
predicate would only prove the fake redelivers.

**Not verified:** the `moveTaskInternalImpl` emit is confirmed on the
unconditional post-commit path by structure and by the surrounding
tests, but is *not* yet asserted end-to-end against a real store move on
both flag settings — that is Phase A2's fixture. The engine subscriber
registry ships empty by design, so no production subscriber exercises
the bus end-to-end yet. `settings-defaults.test.ts` has one pre-existing
failure on `main` (a logger-prefix mismatch in the
`mergeIntegrationWorktree=cwd-main` warning) — confirmed present on a
clean tree, unrelated to this branch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Workflow lifecycle columns are now derived from workflow definitions,
supporting renamed and custom workflows.
* Added post-commit lifecycle events for task transitions, node entry,
and run suspend/resume with validated payloads.
* Follow-on processing for lifecycle emissions is now more robust
(rollback-safe, at-least-once delivery, idempotent handling).
* **Bug Fixes**
* Workflow board responses, task enrichment, and promotion no longer
depend on workflow-columns feature-flag gating.
  * Subscriber failures no longer impact committed workflow transitions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 13:30:13 -07:00
gsxdsm
8b039a543e fix(desktop): advance Pi runtime pin to 0.82.1 for packaging PR lane (#2465)
## Summary
- Advance the matched Pi runtime pin (`pi-ai`, `pi-coding-agent`,
`pi-agent-core`, `pi-tui`) from **0.82.0 → 0.82.1** so
electron-builder's production-dependency walk accepts `pi-agent-core`'s
`pi-ai@^0.82.1` requirement.
- Fixes the Desktop packaging PR-lane failure:
`Production dependency @earendil-works/pi-ai not found for package
@earendil-works/pi-agent-core` (required `^0.82.1`).
- Keep the workspace override guard; update pin-policy fixtures and CLI
package-config expectations.
- Tighten the advisory packaging step-order test so it asserts against
the real `electron-builder --dir` step (not a missing release-only step
name that previously passed via `indexOf === -1`).
- Run `pnpm dedupe` so the packaging lane's lockfile dedupe
early-warning is clean.

## Context
#2439 pinned the full Pi closure at 0.82.0 and made recent main-based
packaging runs green. This advances to the current upstream patch so
deploy + electron-builder stay aligned with `pi-agent-core@0.82.1`'s
declared dependency range.

## Test plan
- [x] `node scripts/check-pi-versions-pinned.mjs`
- [x] `node --test scripts/__tests__/check-pi-versions-pinned.test.mjs`
- [x] `pnpm --filter @runfusion/fusion exec vitest run
src/__tests__/package-config.test.ts`
- [x] `pnpm --filter @fusion/desktop exec vitest run
src/__tests__/release-workflow.test.ts`
- [x] `pnpm dedupe --check`
- [ ] GitHub: Desktop packaging (should run full packaging walk —
lockfile/package.json touched)
- [ ] GitHub: PR Checks (Lint, Typecheck, Build, Gate)
2026-07-26 23:47:49 -07:00
gsxdsm
99c9f14ee0 feat: run Plan Review in the planning lane with a Plan Review badge (#2462)
## What

Plan Review, planning, and the replan loop move from the implementation
column into the **planning lane** (`todo`), so a task under
specification never holds a WIP slot. The card crosses into
`in-progress` exactly once, at `parse`, released by the scheduler.

Operators also finally see a **Plan Review** badge while the gate runs —
it was previously invisible on the default workflow.

## The part that made it possible

Moving the node is ten lines. It was attempted three times and reverted
each time, because a graph run with no durable continuation replayed
from `start` and dragged an in-progress card *backward* out of the WIP
column, firing `abort-on-exit` and stranding it in a pre-WIP column with
no releaser.

So this PR adds the graph **entry contract** —
`resolveColumnResumeNode`:

| Card is in | Resumes at |
|---|---|
| `triage` | `start` |
| `todo` | `plan` |
| `in-progress` | `parse` — never re-plans, never moves backward |
| `in-review` | first review node — gates are not skipped |

`ir.columns` is ordered and that order is the lifecycle order; rework
and failure edges are excluded so the entry point is always the main
path. The proof it's the right fix: **`executor-task-done-invariant`
passes unmodified** after failing every previous attempt.

## Also in here

- **Release gate narrowed twice.** `isUnplannedForExecution` applies its
pre-release plan-review gate only when the node's column equals the
card's column *and* the group is enabled for the task. The enablement
check fixes a real deadlock — a task with Plan Review toggled off was
held forever waiting for evidence nothing would ever write.
- **Badge cleanup.** Gate badge reads "Plan Review" instead of the
ambiguous "Reviewing" and no longer hides behind a lane restriction; the
status badge stops duplicating it; `planning` renders as "Planning"
instead of the raw engine token.
- **Coding (Ideas)** renames its planner column to "Planning" (id `todo`
unchanged) and loses its private planning-node re-home — the graph it
clones is already plan-in-place.
- **New sweep** `reconcileUndeclaredTaskColumns` re-homes a row whose
column its workflow no longer declares. Written for a follow-up, kept
because it makes any column edit survivable.

## Test changes

Scheduler and release fixtures now model a card whose Plan Review passed
— the state every real card is in when the capacity sweep sees it. A
held unreviewed card is the gate working, and that path stays owned by
`pre-release-plan-review.test.ts`.

New `workflow-graph-entry-contract.test.ts` covers the invariant at
every lifecycle position, plus the gap-column and remediation-node
cases.

## Verification

Gate 299 + 70 + 10, dashboard badge suites 672, engine
workflow/entry/executor suites 147, core 122. Lint and typecheck clean.
Full engine suite sits at the pre-existing baseline (notifier /
plugin-runner / notification-service, untouched by this).

## Follow-up

Removing the Todo column entirely is a separate ~207-site
lifecycle-vocabulary refactor — planned in
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`
(companion docs PR).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Plan Review now runs in the Planning lane before implementation
begins.
* Cards resume from their current workflow column without replaying
earlier steps.
* Added automatic recovery for cards stranded in outdated workflow
columns.
* **Improvements**
  * Renamed the Coding (Ideas) planner column to “Planning.”
* Refined Plan Review gating to respect enabled settings and the card’s
current column.
* Updated planning and Plan Review badges for clearer, consistent labels
across cards and lists.
* **Bug Fixes**
* Improved workflow transitions and release behavior around planning,
review, and execution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 23:42:46 -07:00
gsxdsm
5ae6332563 refactor: collapse dead SQLite dual-path code; keep migration-only readers (#2454)
# Remove dead SQLite dual-path code; keep migration-only readers

## Summary
PostgreSQL cutover left hundreds of production dual-path branches
(`backendMode ? PG : SQLite/store.db`) whose SQLite arms only hit
throwing `Database`/`ArchiveDatabase`/`CentralDatabase` stubs. This
change mechanically collapses those unreachable arms so production
authority is AsyncDataLayer/PostgreSQL only, while preserving the six
authorized read-only migration/recovery `DatabaseSync` seams.

## Dual-path mass removed
| Metric | Before | After |
|---|---|---|
| `if (…backendMode)` (non-test) | ~328 | ~70 |
| `store.db` / `this.db` refs in core (non-test) | ~570+ | ~375 (mostly
pure legacy MissionStore/eval/insight SQLite classes + thin getters) |
| Net diff | — | **~6.7k lines removed** across 41 files |

Remaining `backendMode` checks are intentional (incomplete-PG sync
safe-defaults, settings-sync disabled-on-PG, symbol-lock PG-only gates,
“requires PostgreSQL” config versioning throws), not live SQLite
authority.

## Subsystems cleaned
- **Core TaskStore / task-store/***: collapsed if/else and early-return
dual-path across reads, moves, lifecycle, mutations, workflow, archive,
branch/PR, artifacts, comments, audit, project ops, etc. `initImpl` is
PostgreSQL-only (SQLite startup tail deleted).
- **Satellite stores**: automation, agent, routine, plugin, secrets,
approval-request, central-core dual-path arms collapsed.
- **Plugins**: reports async methods, compound-engineering pipeline +
session stores, CLI Printing Press store — SQLite fallbacks removed; PG
required.
- **Engine**: no functional dual-path change beyond whitespace
(settings-sync / peer-exchange PG-disabled behavior kept).

## Six migration-only readers retained (allowlist unchanged)
1. `packages/core/src/postgres/sqlite-migrator.ts`
2. `packages/core/src/project-identity.ts`
3. `packages/core/src/sqlite-validation.ts`
4. `packages/core/src/postgres/startup-factory.ts`
5. `packages/cli/src/commands/db.ts`
6. `scripts/lib/start-local-project.mjs`

Plus low-level `sqlite-adapter` and migrator/startup-import tests.
Inventory ratchet still requires exactly these six `new DatabaseSync(`
production sites, all `readOnly: true`.

## Not treated as SQLite
- `.fusion/project.json`, `task.json`, `agent-log.jsonl` file storage
- AsyncDataLayer / Drizzle PG paths
- Incomplete-PG sync safe-default stubs (still return empty/false/null
under backend without consulting SQLite)

## Verification
- `sqlite-production-reader-inventory.test.ts` — 15/15 pass
- `incomplete-pg-ports.pg.test.ts` — 6/6 pass
- Targeted PG tests (create-task, move, handoff, runtime-persistence,
agent, mission, insight, central-core) — green
- `tsc --noEmit` for `@fusion/core`, `@fusion/engine`,
`@fusion/dashboard` — green
- `scripts/check-no-getdatabase.mjs` — clean

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Improved end-to-end consistency by making PostgreSQL/async persistence
the standard across core task/workflow, automation, agents, plugins,
routines, secrets, approvals, central operations, and session storage.
* Unified scheduling, settings, configuration revision writes,
run/workflow selection, queues/leases/transitions, and audit/lifecycle
updates around consistent async transaction behavior.
* **Bug Fixes**
* Fixed edge cases for archived/deleted reads, unarchive/recovery flows,
not-found handling, and task/artifact/document/log/comment operations,
including more reliable emissions and hydration across search/list and
lifecycle operations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 23:28:42 -07:00
Phil Larson
52d64fa66e fix(engine): project CE steps after review handoff (#2464)
## Summary

- reconcile successful graph-native workflow results with pending task
checklist steps even when review handoff already moved the card into the
merge column
- preserve terminal, paused, and no-redundant-move behavior
- cover the real Compound Engineering post-review-handoff state with a
regression test

## Root cause

Compound Engineering runs `review-handoff` before `merge`. Review
handoff moves the task to `in-review`, which is also the merge column.
`ensureWorkflowMergeBoundaryTask()` returned immediately for cards
already in that column, before projecting successful
`workflowStepResults` onto legacy `Task.steps[]`. The merger then saw
`0/N` and rejected approved work with `task has incomplete steps`.

## Verification

- RED: regression test failed before the fix because `store.updateTask`
was never called
- GREEN: `executor-graph-boundary.test.ts` — 6 passed
- relevant non-PostgreSQL set — 31 passed, 5 PostgreSQL tests explicitly
skipped
- `@fusion/engine` typecheck passed
- changeset format passed
- `git diff --check` passed

## Baseline note

`ce-workflow-step-executor.test.ts` currently has three failures on
clean `origin/main` after FN-8601 foreach-proof hardening. The same
failures reproduce without this patch and are not regressions from this
change.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved reconciliation after review handoff by projecting completed
step results onto the legacy checklist when reaching the merge column.
* Prevented tasks from being marked approved with incomplete step counts
(including “0/N” style states).
* Reduced unnecessary merge failures and deadlock/pause scenarios when
merge-column progress was already recorded.
* **Tests**
* Added coverage for execute-and-merge workflows, ensuring
merge-boundary resolution updates pending steps without moving the task.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 23:13:21 -07:00
gsxdsm
0e3d2a2265 refactor: delete meta-task auto-archive and automated recovery follow-ups (#2461)
Deletes two pieces of automated "meta" machinery that filed and
garbage-collected cards restating state already on the task that failed.
Net **-1015 lines**.

## Why

**Automated recovery follow-ups.** `createAutomatedFollowup` and its
dedup engine (289 lines of signature matching, 1h recurrence
rate-limiting, 24h supersedes windows) existed to file recovery cards
for verification-cap and merge-conflict give-ups. In both cases the
parent is *already* parked `failed` with a descriptive `error` and a log
entry carrying the failing command, branch, and output — the card was a
second copy of that.

**Meta-task auto-archive.** The sweeps that garbage-collected those
cards were worse than redundant: the regex classifier matched ordinary
feature work, and its positional fallback bound cards to unrelated
tasks, so **live work could be archived**.

They are removed together, because the auto-archive sweeps only existed
to clean up after the follow-up engine.

## What changed

### Deleted
- `packages/engine/src/verification-followup-dedup.ts` in full —
`createAutomatedFollowup`, `decideAutomatedFollowup`,
`AutomatedFollowupKind`, `computeVerificationFailureSignature`,
`extractFailingTestFiles`.
- `findActiveRecoveryFollowUp` — dead code, defined and never called
(`tsc` independently flagged it `6133 declared but its value is never
read`).
- The meta-task auto-archive sweeps `autoArchiveResolvedMetaTasks` /
`autoArchiveStalledMetaTasks` and helpers `classifyMetaTask` /
`resolveMetaTargetTaskId` / `computeMetaChainDepth` / `archiveMetaTask`
/ `evaluateMetaAutoArchiveGuards`, plus settings
`metaTaskStallAutoCloseMs` and `metaTaskActiveExecutionGraceMs`.
- Run-audit types `task:auto-archived-meta-resolved`,
`task:auto-archived-meta-stalled`,
`task:auto-archive-meta-resolved-skipped`,
`task:auto-archive-meta-stalled-skipped`,
`verification:followup-created`, `verification:followup-deduped`.

The two signature helpers were **deleted rather than relocated** — once
the three call sites went they were provably unreachable:
`buildVerificationFailureSignature` had exactly one caller, and it was
the only caller of `extractFailingTestFiles`.

### Call sites 1 and 2 — park kept, card dropped
Verification-cap and merge-conflict give-ups keep their park, audit
event, operator comment, and log entry. Site 1's `error` string was
reworded off `"See follow-up task for investigation."` (no follow-up
will exist) to carry the guidance itself. `autoResolveDisabled` was
**kept** — it still drives the outer park guard and the `reason` string;
only the inner branch that guarded card creation is gone.

### Call site 3 — autostash orphan, replaced not deleted
This one is a genuine data-loss guard, so it keeps a durable trail. A
`live`-classified orphan is a merger stash holding **real uncommitted
work**, and unlike sites 1–2 there is no parked parent — the parent may
already be `done` and merged, so nothing else on the board would ever
mention the stash.

The card is replaced by a `logEntry` **and** an `addTaskComment` on the
parent, preserving every fact the old description carried: the sha,
`record.label` (the handle `git stash` recovery needs),
`record.detectedByTaskId`, and `sourcePhase`. New truthful run-audit
event `task:autostash-orphan-live-detected` replaces the borrowed
`verification:followup-*` name, with ids/outcomes-only metadata per
AGENTS.md.

### Kept unchanged: the two real product features
Eval follow-ups (`eval-followups.ts`) and PR-comment follow-ups
(`pr-comment-handler.ts`) only borrowed the shared engine for its dedup
pass. Both keep their exact behavior, column, priority, `sourceType`,
and log lines, with dedup inlined as a `listTasks` scan on
`suggestionId` / `prNumber` respectively. Both fail open (create) if the
listing throws, matching the old engine.

## Test changes — read this one

Two tests asserted the *deleted* engine's rate-limited `"[verification
recurrence]"` logEntry. Those assertions were removed, **not loosened**:
both tests still assert no duplicate card is created, and the eval test
still asserts the existing id is reported back. No coverage of surviving
behavior was weakened. The three `meta-*` test files were deleted along
with the sweeps they covered.

## Verification

```
$ pnpm test:gate
 Test Files  2 passed (2)     Tests   10 passed (10)    # core
 Test Files  16 passed (16)   Tests  299 passed (299)   # engine-core
 Test Files  1 passed (1)     Tests   70 passed (70)    # ci-shape
GATE_EXIT=0

$ pnpm --filter @fusion/engine --filter @fusion/core exec tsc --noEmit -p tsconfig.json
TSC_EXIT=0   (no output)
```

Plus a file-scoped run over the touched surfaces (`eval-followups`,
`pr-comment-handler`, `merger-autostash-orphan-surface`,
`merger-autostash-cleanup`, `run-audit`, `run-audit-secret-taxonomy`,
`project-engine`, `project-engine-manager`): **213/213 passed**.

A repo-wide grep confirms no surviving references to any deleted symbol,
module, or audit event.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Failed tasks now retain recovery and verification details directly on
the original task instead of generating separate follow-up cards.
* Live autostash issues now preserve stash information in task comments
and activity logs.
* Existing evaluation and pull-request follow-ups continue to be reused
when appropriate.

* **Changes**
  * Removed automatic archival of meta-tasks.
  * Removed obsolete meta-task timing settings.

* **Documentation**
* Updated architecture and settings documentation to reflect these
workflow changes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:38:58 -07:00
Phil Larson
232e17d1cb test(engine): complete runtime logger mock (#2458)
## Summary
- add the missing `debug` method to the runtime-resolution logger mock
- prevent logger calls from short-circuiting runtime selection
assertions

## Test plan
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/runtime-resolution.test.ts`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm check:changesets`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Updated the runtime-resolution test suite’s mocked logger to also
support debug-level messages, alongside existing log, warn, and error
handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 22:35:28 -07:00
Phil Larson
00778cb10f test(engine): refresh shellout allowlist (#2451)
## Summary
- refresh the engine synchronous-shellout allowlist after recent
self-healing and executor source additions shifted audited call sites
- keep the guard's path, primitive, and signature checks unchanged

## Test plan
- `corepack pnpm --filter @fusion/engine exec vitest run
src/__tests__/engine-no-blocking-shellout.test.ts --silent=passed-only
--reporter=dot`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Updated static validation allowlists for audited synchronous shell
command operations so matching stays accurate with the latest call-site
locations.
* Kept safeguards that prevent unapproved blocking shell commands from
passing validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
2026-07-26 22:33:21 -07:00
gsxdsm
256c64a7bd chore(release): v0.74.0-beta.5
Version bump via changesets.
2026-07-26 18:11:47 -07:00
gsxdsm
beebd270bd fix: make Queued to plan / Ready badges agree with the planning lane
TaskCard inferred "unplanned" from steps.length === 0 while triage's
todo-discovery and the scheduler's dispatch filter both decide from
PROMPT.md seed-ness, so the badges disagreed with the engine in both
directions: a real spec that parsed to zero steps read as "Queued to
plan" while the scheduler already treated it as a WIP-slot candidate, and
a re-seeded card still carrying old steps read as "Ready" while triage
was about to plan it. Either way the badge sent operators to the wrong
cap.

Adds the shared isTaskAwaitingPlanning predicate (replan park, missing
spec, seed-vs-real content) used by both triage's discovery and a new
best-effort `awaitingPlanning` enrichment on GET /api/tasks. TaskCard
derives both badges from that one value — strict complements — and keeps
the step count only as a fallback for SSE payloads and older servers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:48:46 -07:00
gsxdsm
5ea98f7d4b fix: release admission claims when triage evicts a hung planner
evictStaleProcessing cleared `processing` but left the task in
`coordinatorAdmittedTaskIds`, which is only cleared by specifyTask's
finally — the path a hung promise never reaches. The card stayed
eligible (so the throttle branch never logged or emitted
`task:plan-admission-throttled`) while admitOldest's refresh filtered it
out, leaving it on the "Queued to plan" badge with free slots and no
diagnostic until engine restart. Also drop an untransferred pre-held host
slot, which otherwise waits out the 600s stale-excess valve.

Regression tests assert the invariant on the real production candidate
source: an evicted card is re-offered and its host slot returned, while a
still-live stale task keeps both claims.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:33:35 -07:00
gsxdsm
0022621d22 chore(release): v0.74.0-beta.4
Version bump via changesets.
2026-07-26 17:00:59 -07:00