Commit Graph

835 Commits

Author SHA1 Message Date
gsxdsm
632d10a9b4 fix(engine): a completed-blocked guard was inert on renamed boards — plus one owner for the terminal pair (#2568)
Two commits: a behaviour-preserving extraction, then the behaviour
change.

## ⚠️ Stack note worth acting on

**#2550 and #2554 both report MERGED, but their content is not on
`main`.** They merged into their *base branches*, and the bottom of that
stack (**#2544**) is still open. Nothing in this chain has reached
`main` yet.

Nothing is lost — everything is in
`origin/feature/workflow-e2e-merge-rebound`, which is why this PR
targets it. But "merged" reads as "landed" and here it doesn't.
**Merging #2544 flows the whole chain down.**

## The bug

`parkCompletedBlockedTask` opens with *"is this card already finished?"*
and answered it with:

```ts
if (task.column === "done" || task.column === "archived") return false;
```

On a renamed board neither matches, so **the guard was inert** — and the
very next branch (`if (task.column !== "todo")`) would then have **moved
a completed card back out of its own terminal column**.

A guard that never fires does not fail a test. This one was found by
tracing the last ledger site, not by anything going red.

## Why a shared owner, not a local fix

`merger-ai`'s `isAlreadyFinalizedColumn` held the **only** copy of the
per-role terminal-pair rule — a P1 learned the hard way (PR #2471
review): a per-**set** fallback collapses to one element for a workflow
declaring `complete` but no `archived`, silently dropping the archived
half of every already-finished check.

Executor's guard was the raw literal pair, so **whoever converted it
next would have re-made exactly that mistake** — the lesson lived in a
comment in another file. Hence `resolveTerminalColumns(ir)` in core: one
owner, one place for the rule.

## Evidence, and its limits

**Commit 1 (extraction) is proven behaviour-preserving**:
`workflow-already-finalized-live-e2e` is unchanged and green through the
delegation, and the per-set mutation **still fails** through the shared
helper.

**Commit 2 (the fix) is unproven at the call site, and I'm labelling it
rather than implying otherwise.** `parkCompletedBlockedTask` is private
and reached only from inside executor dispatch — I could not drive it
end to end. So the shared helper gets its **own** tests, in both
partial-role directions, precisely because its other consumer can't
vouch for it. The call site is a one-line delegation to a tested
function.

Weaker evidence than the rest of this unit's work. Saying so, because
quietly counting it as proven is the exact failure this unit exists to
catch.

## Census

417 → 416. That ratchet (#2557) is a **ceiling**, so it stays green
without coordination; lower the pin when convenient.

## Verification

- E2E suites 10/10; helper unit tests 5/5
- core + engine `tsc --noEmit` clean
- `pnpm test:gate` green (414 + 10 + 71)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

## Note on the conflict status (2026-07-31)

GitHub reports this PR `CONFLICTING / DIRTY`. **It is not.** Three
independent checks:

- `git rebase origin/main` on the pushed head reports *"up to date"* and
leaves the SHA unchanged — the branch is already on top of main.
- `git merge-tree` against the merge base produces **zero** conflict
markers.
- `origin/main` is unchanged at the commit this was rebased onto.

The remote SHA matches the local head, so the push landed. The
`mergeable` field is a **stale computation** — it goes stale after a
force-push and doesn't always recompute.

This branch has now been rebased and force-pushed four times against
that cached value. Worth guarding at the source: the auto-retry treats
`mergeable` as ground truth, so a stale value generates conflict notices
indefinitely. Confirming with a trial rebase or `git merge-tree` before
dispatching distinguishes "actually conflicting" from "GitHub hasn't
recomputed" — one command, and it ends the loop.

Verification on the current head: merge gate green (487 + 158 + 10),
engine + core tsc clean, lint clean, 20 tests in the affected suite,
zero unresolved threads.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 01:33:03 -07:00
gsxdsm
dca20496f4 consolidate/u7: plugins to zero + 8 executor rebound guards + resume lanes (supersedes #2607, #2635, #2640) (#2644)
Consolidation branch for U7, per the new one-branch working mode.
**Supersedes #2607, #2635, #2640** — the three of my PRs that were stuck
on review threads. My other seven (#2602, #2605, #2606, #2611, #2621,
#2628, #2633) are green with **zero unresolved threads** and are
deliberately left alone for the merge sweep.

## What is in here, file by file

| file | change | guards before → after |
|---|---|---|
| `plugins/…/glasses/src/agent-actions.ts` | gates, destinations and
degraded-resolution refusal all resolve from the task's own workflow | 2
→ 0 |
| `plugins/…/glasses/src/quick-capture.ts` | accepted capture columns
come from the board; default no longer names the deleted column | 1 → 0
|
| `plugins/…/glasses/src/settings.ts` | quick-capture default was
`triage`, the column #2515 removed | (assignment, uncounted) |
| `plugins/…/dependency-graph/src/GraphTaskNode.tsx` | redundant column
condition deleted | 1 → 0 |
| `packages/engine/src/executor.ts` | 8 rebound guards compare the
resolved column; 4 resume-eligibility literals share one resolver | 151
→ 143 (+4 off-bar) |
| `packages/engine/src/__tests__/` | 4 new suites, 26 cases | — |

`plugins/` reaches **zero** column guards with this branch.

## The three threads it closes

**#2607 — five findings, all mine, all the same rule.** I kept
*qualifying* a legacy-id fallback instead of removing it:

| attempt | rule | hole review found |
|---|---|---|
| 1 | fall back to `todo` when the role is missing | moved cards to
phantom columns |
| 2 | …only if the workflow **declares** `todo` | aliased **review**
lane named `todo` |
| 3 | …and only if no other role is assigned to it | **traitless**
parking column named `todo` |

The qualifications were the mistake. Once `resolveLanes` returns a lane
set the workflow *has* a column vocabulary, so "no column carries the
hold trait" is a complete answer — refuse. `destination()` is two lines
now, with no aliasing surface left to qualify.

Plus a sixth, which is a genuinely different state: **degraded
resolution is indistinguishable from the default board.**
`resolveWorkflowIrForTask` is total by design — a missing definition
silently returns the *default* coding IR — so a card on a custom board
whose definition could not be read resolved to `todo`/`in-progress`.
`undefined` lanes cannot express that (it means "no workflow at all",
where the legacy ids *are* the answer). The actions now refuse with 409.
#2618 would replace this check with resolver provenance; it is not
merged, so this does not depend on it.

**#2635 — "seven rebound sites remain untested."** Fair; my "same shape"
note was an assertion, not coverage. Seven of the eight need a live
graph run to reach, so the *shape* is pinned instead: a static check
that no guard in front of a rebound move compares against a column
literal, with a vacuity case (the same detection run against the
original shape) and a match-count floor (≥8), because a guard reporting
success on zero matches is worse than no guard.

**#2640 — duplicate workflow resolution.** Framed as I/O; it is also a
correctness bug. Eligibility and re-entry are two halves of one decision
and resolved the workflow separately, so a workflow edit landing between
them has the halves reading *different boards*. Now one caller-owned
memo per decision — caller-owned because a process-lifetime cache would
have to guess when a mid-flight workflow edit invalidates it.

## Behavioural findings, not tidying

- **The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** `promotedFromPlannerColumn` was false
on a renamed board, so finished work resting in planning was never
promoted; the code fell through to a review handoff that role adjacency
rejects, and the card stayed stuck with its work complete.
- **Rebound guards could not see the column their own move targeted.**
U5b converted the move target; the eight `column !== "todo"` checks in
front of it were left literal, so on a renamed board the engine moved a
card into the column it was already in — and `moveTaskInternal` runs
reset-on-entry on every real move, so at the `preserveProgress: false`
site it reset step progress a second time.
- **The FN-1404 `task:move` audit row was lying**, recording `to:
"todo"` while the move target was resolved. A run-audit trail that
disagrees with the move it describes is worse than none. Not a
comparison, so no census counts it.
- **A task interrupted by an engine pause never resumed on a renamed
board** (off-bar, `in-review`/`in-progress` literals): four comparisons
decided one question and had to agree; two of them disagreed on a
renamed board, so re-entry silently never fired.

## Revert proofs, isolated per site

| reverted | result |
|---|---|
| `destination()` back to attempt 3 | 3 of 38 fail |
| degraded-resolution refusals removed | 2 of 42 fail |
| capture set back to the legacy five | 2 of 3 fail (renamed-board
suite) |
| forward exclusions → literals | 1 of 14 fails |
| missing-wip refusal removed | 2 of 14 fail |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| promotion target → `"in-progress"` | 3 of 7 fail |
| one rebound guard → `!== "todo"` | 1 of 3 fails (static shape) |
| resume lanes → legacy trio | 1 of 5 fails |

Every conversion is paired with a negative — a forward move, a
not-a-planner-lane card, a default-lineage card, an unresolvable
workflow — so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Commit discipline

Twelve commits, each one thing: the code move (`resolvePlannerLanes` out
of `triage.ts`) is separate from every behavior change, and each review
fix is its own commit with its own revert proof.

## Verification

- `pnpm test:gate` **71/71**
- 162/162 across the glasses plugin's 19 files; 26/26 across the four
new engine suites
- engine + glasses typecheck clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Engine recovery and retries now work correctly with renamed or
customized workflow columns.
  * Tasks in manual-intake columns are no longer automatically planned.
* Agent actions and quick capture now respect each board’s declared
columns and lifecycle stages.
* Awaiting-approval tasks are recognized regardless of their current
column.
* Command Center SDLC funnel stages now accurately reflect customized
workflows.

* **Documentation**
* Added guidance for safely changing workflow-column logic and
interpreting lifecycle-column checks.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 00:52:55 -07:00
gsxdsm
6ed284f36a drop the dead semaphore parameter from dropPreHeldExecutorSlot (#2574)
Small follow-on to the cross-project cap removal.

`dropPreHeldExecutorSlot(taskId, semaphore?)` released a cross-project
semaphore slot. That semaphore is deleted, and **all 16 production call
sites passed `this.options.semaphore`**, which nothing wires any more —
so the release was a no-op on an always-undefined value: an optional
parameter that reads as if it does something.

## What is *not* deleted

Pre-held slots are **dual-purpose**: a cross-project semaphore slot
**and** the FN-8453 per-project coordinator reservation. Only the first
is gone. The reservation is the half that matters — every rejection path
funnels through this helper so an early scheduler/triage return cannot
permanently consume a project slot — and it stays. That is why this is a
parameter change, not a helper deletion.

Sites that still hold a semaphore reference release it **explicitly**
next to their drop, so behaviour is unchanged for any caller that
supplies one. Nothing wires one in production today, but silently
leaking a slot for a caller that does is not a trade a cleanup is
allowed to make.

## One real leak fixed — found by a failing test, not by reading

`ProjectAdmissionCoordinator.admitOldest`’s release lambda took the
pre-held branch and **returned**, relying on the deleted parameter to
hand the host slot back. With the parameter gone, that branch unwound
the registration and the reservation while **leaking the host slot** the
attempt had acquired. The release is now unconditional across both
branches.

Worth noting how it surfaced: the test that caught it (`drops a declined
candidate’s pre-held executor slot`) asserted `semaphore.activeCount`,
which I had initially assumed was just coupling to the deleted half. It
was not — it was pinning a real invariant.

## Tests

Five cases in `concurrency.test.ts` pinned `sem.activeCount` through a
drop. Each is re-pointed at the surviving contract — registration and
reservation unwound, nothing left for a later pass to “take” — with the
semaphore assertions moved to the sites that now own the release.

## Verification

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green ·
`concurrency.test.ts` **56/56**. The 8 `triage.test.ts` failures are
**pre-existing** — reproduced identically with this branch’s `triage.ts`
replaced by main’s.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:52 -07:00
gsxdsm
3aa942ee5f capacity: spawned agents count against the project agent count (#2579)
Two configurable numbers per project. `maxSpawnedAgentsPerParent` (5)
and `maxSpawnedAgentsGlobal` (20) were a **third and fourth** limiter
with private budgets invisible to both.

## This closes a hole, not just knobs

A spawned child **is** an agent and gets **its own git worktree**
(branched from the parent’s — the tool’s own description says so), but
children were counted by **neither** capacity gate. A fan-out could put
up to 20 extra worktrees on disk while the scheduler believed the
project was at its configured limit. The operator’s two numbers were
simply wrong about what was running.

## The old caps also measured the wrong thing

`totalSpawnedCount` decrements on child cleanup, but the per-parent
**set** is cleared only when the **parent task** ends. So
`maxSpawnedAgentsPerParent` throttled *cumulative* spawns across a
task’s life rather than *concurrent* ones — a long-running task could
exhaust its budget with five children that had all long since finished,
and the operator had no way to see why.

## Fix

`fn_spawn_agent` gates on the same project agent count every other lane
uses (`computeTopLevelConcurrencyClaimedFromStore`) plus live children.
One number, one answer, no private budget that can disagree with the
board.

The refusal names **Max Concurrent Tasks** — a control the operator
actually has. The old messages pointed at settings that no longer exist,
which is worse than no message: it sends someone hunting for a knob that
is not there.

## Verification

**Revert-proof, measured:** restoring the private budgets turns **3 of
the 4** new cases red — a project at 1/1 could still spawn, which is
precisely the hole. `executor.ts` restored byte-identical.

`pnpm lint` clean · core + engine `tsc` clean · `pnpm test:gate` green
(414 + 10 + 71) · new suite 4/4 · `settings-default-descriptions` 4/4.

There was no spawn-capacity test before this; the file is new.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Spawned agents now count toward the project’s **Max Concurrent Tasks**
capacity.
* Agent spawning is blocked when capacity is reached, including
concurrent spawn attempts.
* **Bug Fixes**
  * Prevented over-allocation during simultaneous agent spawns.
  * Restored available capacity when agent creation fails.
* **Changes**
  * Removed separate per-parent and global spawned-agent limits.
  * Updated settings to reflect the revised capacity controls.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 23:34:41 -07:00
gsxdsm
31e49b684a TAKING default-workflow-hooks.ts + executor.ts + live-agent-count.ts + 6 dashboard files: reopen semantics by role, and the census's blind spot in both directions (13 sites) (#2628)
Batched conversion of every lifecycle-column guard I hold, plus the
three the census could not see. **Six files to zero, repo-wide 60 → 49
by a comment-stripped unanchored sweep.** Each conversion has an
isolated revert proof and a paired negative case, and the one code move
is a separate commit from the behavior changes.

## Per-file before → after

Counts from a comment-stripped, unanchored `(===|!==) ["']triage["']`
sweep over `packages/*/src` + `plugins/*/src`, excluding tests.

| file | before | after | note |
|---|---:|---:|---|
| `core/default-workflow-hooks.ts` | 4 | **0** | |
| `core/task-store/moves.ts` | 5 | **4** | only the flag-ON mirror
converted; the flag-OFF inline block is the parity reference and stays |
| `engine/executor.ts` | 3 | **0** | **absent from the 45-guard list** —
see below |
| `core/live-agent-count.ts` | 2 | **0** | duplication removed; answer
deliberately unchanged |
| `engine/replan-target.ts` | 2 | **0** | both were comment prose, not
guards |
| `core/agent-prompts.ts` | 3 | **0** | ROLE comparisons, never column
guards |
| `engine/usage-limit-detector.ts` | 2 | **0** | ROLE comparisons |
| `dashboard/app/components/DocumentsView.tsx` | 1 | **0** | real column
guard |
| `dashboard/app/components/TaskChatTab.tsx` | 2 | **0** | ROLE |
| `dashboard/app/components/AgentLogViewer.tsx` | 1 | **0** | ROLE |
| `dashboard/app/components/effective-model-resolution.ts` | 1 | **0** |
ROLE |
| `dashboard/app/hooks/useTasks.ts` | 1 | **0** | ROLE |
| `dashboard/…/command-center/MissionControlPanel.tsx` | 1 | 1 | alias
table, marked `DELIBERATE-LITERAL` with its reason |

## The census errs in BOTH directions

This is the finding I would most like carried into the remaining work.

- It **flagged 10 sites that were never column guards.** `role ===
"triage"` / `agentType === "triage"` compare an **AGENT ROLE**. The
planner *lane* is named `triage` and keeps that name — U11 removed the
*column*. Worse than noise: the obvious "finish the migration" edit is
to rename the role, and that silently empties the planner's prompt
template and mis-binds its model markers. `PLANNER_AGENT_ROLE` now names
it, so the two vocabularies are distinguishable by grep and a rename
fails loudly (revert proof: 4 tests, two of them pre-existing).
- It **missed 3 real guards in `executor.ts`**, because the pattern
matches `column`/`toColumn`/`fromColumn` and those locals are named
`from` and `originColumn`. A census keyed on variable names will keep
missing guards wherever a local was named for its role in the function.

## Two real defects, not tidying

**1. A renamed board could merge with its re-review never run.**
`default-workflow-hooks.ts` is named for the default workflow, but the
store runs it on the flag-ON path for *every* workflow — the trait
registry resolves hooks by trait id, not by workflow. Its reopen
predicates listed the default lineage's column names, so on a renamed
board **no reopen effect fired at all**. One of them clears
`workflowStepResults`, which `getTaskMergeBlocker` reads: a card bounced
out of review carried its old `passed` result back in, and that
satisfies the merge gate. Same regression the graph-owned-crossing
carve-out exists to prevent, arriving through the other door. (Two
smaller ones rode along: failure state never cleared on a renamed
reopen, and an operator dragging a card back to the queue never parked
it, so the scheduler re-dispatched what they had just pulled back.)

**I forgot the carve-out on my first pass, and that was worse than not
converting.** A role-resolved clear plus a *name*-matched exemption
means a renamed board takes the clear and never the exemption,
destroying the remediation input the graph had just written. My own
paired negative test caught it.

**2. The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** In `recoverCompletedTask`,
`promotedFromPlannerColumn` was false on a renamed board, so finished
work resting in the planning lane was never promoted — the code fell
through to `handoffTaskToReview` straight from the planning column, and
role adjacency has no planning → review edge, so the handoff was
rejected and the card stayed stuck with its work complete. I converted
the promotion **target** too: resolving the lane and then moving to a
literal `in-progress` is the half-conversion I have already been burned
by twice this program, where the guard starts admitting cards and the
move then sends them to a column the board does not declare.

## E2E evidence

`renamed-board-reopen.pg.test.ts` drives a **real PostgreSQL store** and
a real `moveTask` on a workflow whose columns carry the standard traits
under non-default names. The unit tests cannot show this: if `moves.ts`
passed `undefined`, every unit case still passes via the no-basis
fallback while the real board keeps the old behavior. **Proof it is
load-bearing: forcing `moveLifecycleColumns` to `undefined` fails 2 of
3.** The executor suite covers both the split-role and the MERGED
post-U11 shape.

## Revert proofs, isolated per site

| change reverted | result |
|---|---|
| reopen predicate → literal names | 4 of 10 fail |
| reopen field clears → literal names | 2 of 10 fail |
| `userPaused` hold lane → literal `todo` | 1 of 10 fail |
| graph carve-out → literal names | 1 of 10 fail |
| store passes `undefined` lifecycle columns | 2 of 3 fail (real PG) |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| two-hop condition → `=== "triage"` | 1 of 7 fails |
| promotion target → `"in-progress"` | 3 of 7 fail |
| `isPlannerColumnFor` → literals | 1 of 7 fails |
| live-agent-count: one arm dropped | 2 of 11 fail |
| DocumentsView: trait branch removed | 3 of 7 fail |
| planner role renamed to `"planner"` | 4 fail (2 pre-existing) |

Every conversion is paired with a negative case (a forward move, a
not-a-planner-lane card, a default-lineage card, a renamed column with
no traits), so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Deliberately NOT converted, with reasons

- **`moves.ts` flag-OFF inline block (4).** That branch *is* the legacy
path, kept verbatim so the two can be parity-checked. Converting it
erases the reference implementation.
- **`live-agent-count.ts`'s no-flags fallback.** Reachable, and there is
nothing to resolve from — `enrich…FromFlags` exists for callers with
board flags rather than an IR, so a column missing from that map is the
renamed case. "Not intake" is as much a guess as "todo is intake", and
Running/Waiting are complements, so a card matching neither arm is
reported as neither and the footer's queued total under-reports it. The
real fix is at the caller; four new cases pin that flags override the
legacy answer **in both directions**. What did change is the
duplication: two hand-written copies of one rule now call one named
function.
- **`MissionControlPanel`'s `FUNNEL_STAGES`.** An alias table of column
*names* where `triage` sits beside `signal` and `backlog`. Command
Center aggregates across projects, so there is no single workflow to
resolve traits from — the honest conversion is a data change, not a
predicate change.
- **`DocumentsView` with no traits.** Same no-basis rule; the documents
list is full of historical columns absent from the current board. A case
asserts a renamed column with no traits still reads as "working",
documenting the gap rather than hiding it.

## Fixture findings

Each cost a red run that looked like the code under test:

- a `merge-blocker` column needs a reachable merge-class node, or
`parseWorkflowIr` rejects the workflow;
- a back-edge must be `kind: "rework"`, and a rework edge is legal only
**into** a node with `config.reworkRegion: true`;
- a workflow gets role-level transitions only when it declares wip +
review + complete + **archived** plus a planning lane — without the
archived column, adjacency falls back to order-derived neighbours and
`checking -> queued` is not a legal move at all;
- `recoverCompletedTask` only *reaches* the promotion seam when nothing
is left to gate; without passed `plan-review`/`code-review` rows it
re-enters the workflow graph and returns first, so a naive fixture
silently tests the wrong branch and every assertion reads "no moves
happened" for an unrelated reason.

## Verification

- `pnpm test:gate` **71/71**
- new suites: 10/10 reopen-semantics, 3/3 renamed-board-reopen (real
PG), 7/7 executor-planner-lanes, 7/7 documents-status-dot, 4/4
planner-role-is-not-a-column
- neighbours: 132 + 10 + 482 (gate shards), 350/351 engine
planning/replan suites, 64/64 agent-prompts, 51/51 usage-limit-detector,
11/11 live-agent-count, 11/11 dashboard hook/log suites
- the single engine failure (`executor-fast-mode-workflows.test.ts` ›
"raw fast mode still invokes non-executable review seam nodes")
**reproduces with my changes stashed** — pre-existing on `origin/main`
- typechecks clean for core, engine, and dashboard-app
(`tsconfig.app.json`; `tsconfig.json` checks nothing under `app/`);
`pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:39:14 -07:00
gsxdsm
592fd5c0c6 U11 [mission-feature-sync + spec-staleness]: convert the last two planner-lane guards (48 -> 46) (#2610)
**Taking: `engine/mission-feature-sync.ts`, `engine/spec-staleness.ts`**
— the last two planner-lane guards in my area.

## Census (comment-stripped, `=== "triage"` / `!== "triage"` in
`packages/*/src`, tests excluded)

| file | before | after |
|---|---:|---:|
| `packages/engine/src/mission-feature-sync.ts` | 1 | **0** |
| `packages/engine/src/spec-staleness.ts` | 1 | **0** |
| **repo total** | **48** | **46** |

## Both are real conversions, not seams

Each guard takes its vocabulary from the **caller**, which holds the
store — so unlike a defaulted parameter nothing passes, these can
actually be driven.

**`reconcileMissionFeatureState`** — a card back in a planner lane
returns the mission feature to `triaged`. Keyed on literals, a renamed
workflow left the feature reading `in-progress` forever: the roadmap
claims work is underway while the card waits to be re-planned. Nothing
errors; the rollup is just wrong. The vocabulary arrives via
`MissionFeatureSyncContext` rather than by widening this module's
deliberately narrowed `Pick<TaskStore, "getTask">`.

**`shouldSkipSpecStalenessForPreservedProgress`** — returning `false`
for a planner-lane card is what *keeps* staleness evaluation on. Miss
the lane and it falls through to the preserved-progress branch, so a
card with progress skips staleness and keeps a spec that should have
been re-validated.

## The two take different defaults — and I got it wrong first

I defaulted **both** to the `triage`/`todo` pair and broke the
pre-existing U11 proof in `spec-staleness.test.ts`, which states the
reason exactly:

> same column, different status, opposite correct answer

- **mission-feature-sync → the PAIR.** It asks "is this card waiting to
be planned?", true in either lane.
- **spec-staleness → the DEDICATED planner column only.** On a merged
lineage `todo` is *also* the hold lane, so the planner distinction there
is carried by **status** (`planning` / `needs-replan`), not by the
column. Treating the merged column as a planner lane stops a parked card
with preserved progress from skipping staleness. Its default is now the
single legacy id — byte-identical to the literal it replaced.

That asymmetry is now pinned by its own test rather than left for the
next reader to rediscover.

## Findings on the remaining census, from measuring it

Two of the 46 are **not lifecycle-column guards** and converting them
would be wrong:

- `tool-availability.ts:32` — `surface === "triage"` where `surface:
"triage" | "executor"` is an **agent lane**, not a column.
- `skill-resolver.ts:432` — `sessionPurpose === "triage"`, a **session
purpose**.

Also worth noting for the count: `replan-target.ts` reads as 2 in a raw
grep but is **0** — both hits are inside comments. `board-workflows.ts`
(2) and `archive-planning.ts` (1) are likewise comment-only. A raw grep
says 52; comment-stripped says 46.

## Not wired at the call sites yet

`scheduler.ts` / `mission-autopilot.ts` (mission sync) and `executor.ts`
/ `scheduler.ts` (staleness) still omit the new option, so behaviour is
byte-identical today. Deliberate: `executor.ts` belongs to u8's active
slice and I would rather not create a textual collision for a
pass-through. The seam is proven by tests and the count is real; wiring
is a follow-up.

## Verification

- **Mutation-verified:** restoring either literal fails a test
- 35 tests green across the three suites, merge gate green (482 + 132 +
10), tsc clean, lint clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-29 21:20:09 -07:00
gsxdsm
45e8b5f7ac U8: pin the completion-finalize ordering invariant before moving the last out-of-band exit (#2599)
Groundwork for moving `paused-after-completion`, the **last**
out-of-band exit. Stacked on #2590.

## What lands

1. **An indentation defect I introduced.** My bulk edit when the exit
vocabulary landed left the second `paused-after-completion` site
mis-indented inside a `finally` block. Cosmetic, but misleading
indentation in a `finally` is how a future reader misjudges scope.

2. **The adjacency ratchet now requires `markCompletionFinalized` before
the handoff, at every reporting site.** It previously checked only the
first occurrence, and only for the handoff itself.

That ordering is the invariant `handleGraphFailure` depends on and
**cannot check for itself**: `alreadyFinalizedToReview` /
`completionFinalized` exist to recognise this out-of-band move when a
later teardown re-marks the abort as `hard-cancel`. Without the durable
marker set first, a completed no-commit task is re-parked `failed` —
FN-6644/FN-6641.

It is asserted **structurally, and labelled as such in the test**. Both
call sites sit in pause and `finally` paths that cannot be driven
without mocking an entire agent session; presenting a source assertion
as behavioural coverage would repeat the overclaim I have been correctly
pulled up on twice in this unit.

Red-green: removing `markCompletionFinalized` from either site fails the
ratchet.

## Why the move itself is not in this PR

`paused-after-completion` is structurally harder than the pending-review
ending that #2590 moved, and the difference is worth recording before
someone assumes it is a copy-paste:

- it does **four** things, not one — `markCompletionFinalized`,
`handoffTaskToReview`,
`clearCompletedTaskWatchdog`/`signalTaskComplete`. Only the handoff is
lifecycle; the rest is substrate that must stay put.
- one of the two sites is inside a **`finally`**. Moving a transition
out of a `finally` is not the same operation as moving one out of a
branch: the graph may already be unwinding, so "report and let the graph
route" needs a defined answer for a run that is already ending.
- there is **no behavioural coverage of either site today** — the
closest tests only exercise the exit vocabulary. The pending-review move
succeeded on the fourth attempt precisely because FN-5436 existed to
catch each wrong version; this exit has no equivalent, so the move needs
that floor built first, and building it means real session mocking
rather than a shortcut.

## Verification

- exit-events + primitive-exit-events + step-session + ownership ledger
— green
- `pnpm lint` clean; `tsc --noEmit` clean
- No user-facing behaviour change, so no changeset

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
  - Improved handling of workflow steps that pause for review.
- Tasks now remain in review when a review request has no subsequent
decision.
  - Added clearer completion events for primitive prompt steps.
  - Preserved correct failure handling when later workflow steps fail.

- **Workflow Improvements**
- Built-in workflows now route pending reviews through a dedicated
review handoff.
- User-authored workflows retain compatible review parking behavior when
routing is unavailable.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 21:04:39 -07:00
gsxdsm
3f763cba87 U8: the graph owns the pending-review park — ownership ledger 28 → 27 (#2590)
The routing move this unit has been building toward, landing on the path
the engine actually runs. **Includes #2578's commit** (the live-path fix
it depends on) — merge that first, or this supersedes it.

## What changes

Three things together, because a half-routed move is a card that
silently does not advance:

1. The **live** implementation primitive (`runCodingSession`) returns
`{outcome: "failure", value: "review-pending"}` for that ending.
2. The primitive step handler stops flattening every ending to
`step-done`/`step-failed`, so the value survives the foreach —
`runForeach` propagates a failing instance's value as the node's own —
and reaches an edge.
3. The inline `handoffTaskToReview` in `runImplementation` is
**deleted**. The phase reports and stops, which is all an implementation
phase should do.

Built-in workflows route to the `review-pending-handoff` node added in
#2519/#2546, which performs the handoff and ends the run: the same two
effects in the same order, with the graph as the owner.

## Proof, end to end

FN-5436 — the test that blocked this move twice and was right both times
— now passes, with a **stronger** assertion than it had:

```ts
expect(store.moveTask).toHaveBeenCalledWith("FN-5436-B", "in-review",
  expect.objectContaining({
    workflowMoveSource: "workflow-graph",
    workflowMoveMetadata: expect.objectContaining({ nodeId: "review-pending-handoff" }),
  }));
```

The old two-argument `moveTask(id, "in-review")` could not distinguish a
graph-owned park from an out-of-band one — which is the entire
distinction this unit exists to make. The invariant (park in review,
never `failed`) is unchanged; the owner is now proven.

## Every ratchet fired, and each records a real change

| Ratchet | Before | After | Why |
|---|---|---|---|
| Ownership ledger — `runImplementation` review handoffs | 3 | **2** |
the handoff left the phase |
| Ownership ledger — `handleGraphFailure` | 0 | **1** | the named compat
classifier |
| Ledger headline — executor-owned dispositions | 28 | **27** | first
decrement of the unit |
| Out-of-band exit list | 2 | **1** | pending-review is graph-owned now
|
| Primitive routing pin | "must not reroute" | routes *only* the moved
ending | declared, not discovered |

None was relaxed. The `handleGraphFailure` 0 → 1 is the honest one: for
a user-authored graph without the edge this is a **relocation, not an
elimination** — the transition is still executor-performed, but from one
named classifier in the failure ladder rather than a call buried two
thousand lines into a session loop. The ledger says so rather than
letting the headline number imply more progress than there is.

## Why it took four attempts

Recorded because the reason is reusable: the value was being produced on
`createAuthoritativeWorkflowSeams`, a handler that never runs (#2578).
Every earlier attempt was correct code on a dead path, and the only
thing that showed it was instrumenting until a negative result was
proven observable rather than assumed.

## Verification

- step-session + exit-events + primitive-exit-events + ownership ledger
+ graph-requeue-gate + task-done-blocked — **83 tests green**
- `pnpm test:gate` green (10 / 482 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved handling of tasks awaiting review so they are correctly
routed to the review workflow.
* Tasks now remain in review instead of being marked as failed when no
follow-up review route is configured.
* Review handoffs now include workflow ownership and provenance details.
* Preserved standard failure handling for tasks that are not awaiting
review.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:59:06 -07:00
gsxdsm
131feb243c U8: the exit announcement was on a dead code path — move it to the handler the engine actually runs (#2578)
A merged behavior of mine has never executed. This fixes it and adds the
ratchet that would have caught it.

## The finding

`createDefaultNodeHandlers` chooses the prompt-node handler like this:

```ts
const promptLike = deps?.primitives
  ? createPrimitivePromptLikeHandler(deps.primitives, runCustomNode)
  : createPromptLikeHandler(seams, runCustomNode);
```

`executeWorkflowGraph` always passes `primitives:
this.createAuthoritativeWorkflowPrimitives(settings)`
(`executor.ts:6051`). **So `createPromptLikeHandler` — and with it every
`execute` / `step-execute` function in
`createAuthoritativeWorkflowSeams` — is unreachable for prompt nodes.**
Both objects are passed to the graph executor and only one is consulted.

The `NodeCompleted.exit` announcement added in #2507 was wired into that
seam. It type-checks, its tests pass (they call the seam object
directly), and it has never run in production. `runCodingSession` in the
primitives is the live twin, and that is where it emits now.

## How it was found — and why the negative is trustworthy

Instrumenting `createAuthoritativeWorkflowSeams.stepExecute` produced no
output for a run that demonstrably visits `steps#0:step-execute`. So did
instrumenting `createPromptLikeHandler`'s dispatch. A negative result
from instrumentation is worthless until the instrumentation is shown to
be observable, so: a `process.stderr.write` at module load of the same
file **did** appear, exactly once, in the same run. The two negatives
were real, not swallowed output.

This is also the answer to the open question I left in #2546 — the
pending-review routing move kept failing because the seam value it
depends on is never produced. **That move is still not landed here.**
This commit only relocates the announcement, so it stays small and
separately revertable; the routing move follows once its value
originates on the live path.

## The ratchet

A source assertion pins the dispatch rule: `deps?.primitives ?
createPrimitivePromptLikeHandler` and the executor's wiring of
`primitives`. Inverting or conditionalising that preference would
silently disable every behavior attached to the primitives path — the
same failure in the other direction — and **a seam-level unit test
cannot tell the two apart**, which is precisely how this survived review
twice.

## Red-green

Removing the emit fails 2 of the 4 new tests (`Tests 2 failed | 2 passed
(4)`). The other two are the regression floor: an ordinary completion
emits `success` with no `exit`, and the returned routing outcome is
unchanged — announcing must not reroute.

## Scope note

I did **not** delete the now-known-dead seam wiring in this PR.
`createAuthoritativeWorkflowSeams` is still passed to the graph executor
and its non-prompt entries (`stepReview`, `merge`) are reached through
other handlers, so deciding what is genuinely dead there is a deletion
audit of its own — and this program's rule is that deletions never ride
along with behavior changes. Filed as the next slice.

## Verification

- 4 new tests + exit-events + step-session + triage audit + ownership
ledger — **54 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 20:30:54 -07:00
gsxdsm
a68785a41d P0: two silent triage guards in the executor's ownership — one strands a card with nothing to rescue it (#2572)
P0 audit of the executor's assigned `triage` sites after the
Planning-column merge. **One of them can strand a card**, so leading
with that.

## The stall — `handleDepAbortCleanup`

`executor.ts` moved a dependency-aborted task to the **literal**
`triage`. The default coding lineage no longer declares that column.

A card that gains a dependency mid-execution has its work discarded and
is then parked in a column its own workflow does not define. Nothing in
the graph routes a card out of an undeclared column. The only rescue is
`reconcileUndeclaredTaskColumns`, which runs on the **next engine
start** — so between the abort and a restart the card is stalled with no
automatic recovery. It does not throw, so it would have surfaced as a
user report, not a red test.

Fixed to `resolveReboundColumnFor`, the helper the other ~16 executor
rebounds already use.

## The silent skip — `UsageLimitPauser.taskUsesProvider`

The planning lane was identified by the same literal. For a default card
the lane resolved to **no providers**, so when a provider hit a usage
limit during a *planning* session, the fan-out that pauses peers on that
provider skipped every default-workflow card and they kept hammering the
rate-limited provider.

Not a stall: the triggering task is still paused by the explicit
fallback below the filter. What was lost is blast-radius containment. A
planning session runs while the card is pre-implementation, and the
caller has already excluded `done`/`archived`, so that is exactly "not
the implementation column and not the review column" — which matches
`todo`, `triage`, `ideas`, and a renamed planner alike.

## Full audit table for my assigned sites

| Site | (a) Still fires for a default card? | (b) What silently stops |
(c) Action |
|---|---|---|---|
| `executor.ts:16395` `moveTask(id, "triage")` | **No** — writes an
undeclared column | Card parked where nothing routes it; rescue only at
next engine start | **Fixed** — `resolveReboundColumnFor` |
| `usage-limit-detector.ts:126` `column === "triage"` | **No** |
Usage-limit fan-out skips every default card; peers keep hitting the
limited provider | **Fixed** — pre-implementation predicate |
| `executor.ts:3409` `from === "todo" \|\| from === "triage"` | **Yes**,
via the `todo` arm | — | Unchanged; `triage` arm still live for
legacy-coding |
| `executor.ts:4951` `originColumn === "todo" \|\| === "triage"` |
**Yes**, via the `todo` arm | — | Unchanged |
| `executor.ts:4963` `originColumn === "triage"` double-hop | No, and
correctly so | Nothing — the extra hop exists only for shapes that
declare `triage` | Unchanged; still required by legacy-coding |
| `executor.ts:1110` `Type.Literal("triage")` | n/a | — | **Not a
column** — an agent ROLE in `spawnAgentParams` |

Counts for my ownership: **6 sites audited, 2 defects, 2 fixed, 3
correct as-is, 1 false positive.**

## Red-green

Reverting each fix fails its own test:

```
Tests  2 failed | 2 passed (4)
  × dependency-abort cleanup requeues to a DECLARED column
  × usage-limit fan-out … pauses a peer card sitting in the merged Planning column (id `todo`)
```

The other two are the regression floor and pass both ways by design: a
legacy workflow that **does** declare `triage` still fans out, and an
in-progress card is still **not** swept into the planning lane (the
guard must stay narrow — "any non-wip column" would have been the easy
wrong fix).

## Verification

- New audit suite + graph-boundary + step-session + ownership ledger —
**45 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 19:03:10 -07:00
gsxdsm
8578a1d27d U8 PR5: thread the implementation exit to the step seam, and declare the stepwise pending-review park (inert) (#2546)
Follows **#2519** (U8 PR4). Both halves are inert — **no behavior
change** — and this removes the blocker PR4 documented.

## What was blocking

PR4 could only land its IR half because the pending-review ending could
not reach a graph edge on the **default** workflow. Three links in the
chain:

| Link | Problem |
|---|---|
| `runGraphTaskStep` | awaited the memoized implementation pass and
**discarded** its result |
| `RunTaskStepResult` / `RunSingleStep` | had nowhere to carry an exit |
| `stepExecute` seam | flattened every ending to `step-done` /
`step-failed` |

All three are fixed. The outcome stays `failure` (the step genuinely did
not complete) while the **value** now names the ending — which is what
`runForeach` propagates upward, since it returns a failing instance's
value as the foreach node's own. Every other ending keeps `step-failed`
byte-identically.

One design note: the exit is a property of the **pass**, not of a step.
A single memoized pass serves every foreach instance, so all instances
report the same ending — correct, because the ending is what stopped the
whole session.

With the value surviving, the stepwise IR declares the same
`review-handoff` park node and `steps --outcome:review-pending-->
review-pending-handoff --success--> end` edge the plain-`execute` shape
got in PR4, inherited by the final-review and Ideas variants that clone
it.

## A bug my own threading introduced, and what caught it

The first threading commit covered **one of the two** paths out of
`runProjectedGraphTaskStep`. The early-return branch carried the exit;
the main path goes through `runTaskStep` in `step-runner.ts`, which
builds its own result and dropped it — i.e. it worked on the path I
happened to read, and not on the path the default workflow actually
takes.

**FN-5436's regression test caught it, not code review.** That is the
second time this test has stood between this unit and a silent
regression, which is worth recording somewhere durable:
`executor-step-session.test.ts > FN-5436: pending-review skip on
no-fn_task_done exit` is the load-bearing test for this area.

## Why the seam flip is still not here

With the threading complete I applied the behavior half again — flip the
execute seam to return `review-pending`, delete the inline
`handoffTaskToReview`, add a named compat classifier for user-authored
graphs. **FN-5436 still failed**: the card did not reach `in-review`, so
something between the seam value and the park node is not routing under
that harness. I have not isolated whether that is the mock store's IR
resolution (it exposes no `getWorkflowDefinition`, so the run resolves
the built-in through a different path), a foreach aggregation detail, or
the park node's own seam.

I stopped rather than keep guessing, and reverted the behavior edits so
this lands green and inert. Shipping a half-routed move is exactly the
failure this unit exists to remove — a lifecycle transition that
silently does not happen. The alternative on offer was to relax
FN-5436's assertion, which would have been appeasing a test that is
telling the truth.

### What the instrumentation showed (done after opening this PR)

I ran the bounded next step rather than leaving it as a note. Two facts,
both measured:

1. **The IR is correct.** Resolving
`BUILTIN_STEPWISE_FINAL_REVIEW_CODING_WORKFLOW_IR` at runtime shows the
node and the edge survive the final-review variant's edge rewiring:

```
EDGES [{"from":"steps","to":"browser-verification","condition":"success"},
       {"from":"steps","to":"review-pending-handoff","condition":"outcome:review-pending"},
       {"from":"steps","to":"end","condition":"failure"}]
HAS NODE true
```

That matters because the variant does `template.edges = [ ... ]` (a
wholesale replacement) and filters outer edges touching `review` —
`review-pending-handoff` is not `review`, so it survives. Worth knowing
before anyone adds another node near it.

2. **The `stepExecute` seam is never invoked in that harness**, even
though the run terminates at `steps#0:step-execute` and the
implementation session demonstrably runs (`"Agent finished without
calling fn_task_done but Step 0 is blocked on pending review"` is in the
task log). A `console.log` at the seam's value computation produced no
output. So the exit is threaded correctly and the IR can route it, but
under this harness the value never originates.

3. **Nor is `createPromptLikeHandler`'s returned handler.**
Instrumenting its dispatch (`node.id` + resolved seam) produced nothing
either — so the node is not reaching the prompt-like path at all.

**Control experiment, because a negative result from instrumentation is
worthless until you prove the instrumentation is observable.** A
`process.stderr.write` at module load of the same file appears exactly
once in the same run, so writes from that module *are* captured under
this harness and the two negatives above are real, not artifacts of
swallowed output.

That narrows the remaining work to one question — what actually drives
`steps#0:step-execute` in this run, if neither the prompt-like handler
nor the `stepExecute` seam does — and rules out the IR, the foreach
propagation, the threading, and the instrumentation as suspects.

**Next step, now much narrower:** find the handler registration this run
resolves for a foreach instance node (the graph executor's handler map,
not the seam table), then flip the seam, delete the inline handoff, and
update the three ratchets that will correctly fire — PR3's routing pin,
the out-of-band adjacency check, and PR1's ownership ledger
(`runImplementation` 3 → 2; `handleGraphFailure` 0 → 1 for custom graphs
only).

## Verification

- `executor-step-session` + exit-events + ownership ledger +
graph-boundary — **56 tests green**
- `builtin-workflows` + `builtin-coding-workflow-ir` — green. The
layout-completeness contract required a layout entry for the new node in
all four stepwise-derived workflows; placed off the main line, because a
park is an exit and not a stage.
- `pnpm test:gate` green (10 / 309 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `internal`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:54:39 -07:00
gsxdsm
7fd1c7f124 P0 fix: stop reaping worktrees out from under live planners (FN-6756) (#2531)
User-reported: worktrees deleted while a planning agent was still
working in them. Small, isolated, ahead of all remaining capacity work.

## Mechanism

`clearPhantomExecutorBinding` is documented as *"the last line of
defense against pulling a worktree out from under a running agent"*. It
computed liveness from four sets — `activeSessions`,
`activeStepExecutors`, `activeWorkflowStepSessions`,
`activeCliTaskSessions` — **all TaskExecutor-owned**. A triage PLANNING
session is owned by `TriageProcessor`, lives in *its own*
`activeSessions` map, and registers in the module-level
`activeSessionRegistry`. It matched none of the four.

Worse: the method **writes** to that registry (unregistering the task’s
paths) but never **read** it as a liveness signal. It destroyed the very
evidence that proved the planner alive.

Under plan-in-place a card is specified while it sits in
`todo`/`triage`, and `reapLeakedConcurrencySlots` treats both as
reapable on a rationale written *before* planning moved there (“a task
waiting to run must not pin a worktree”). Every gate ahead of the last
one passes for a planner:

| Gate | Saves a planner? |
|---|---|
| in `listWorktreeHolders()`? | **No** — `ensureTaskWorktreeForPlanning`
→ `ensureGraphCustomNodeWorktree` → `addActiveWorktree`
(`executor.ts:8581`) |
| reapable column? | **No** — plan-in-place keeps the card in
`todo`/`triage` |
| in the executor’s `executing` set? | **No** — a planner is
triage-owned |
| 60 s `LEAKED_WORKTREE_SLOT_GRACE_MS` | **No** — keyed on
`columnMovedAt`, and planning routinely runs for minutes |

So the broken guard decided alone.

## This is FN-8600 recurring through a second sweep

That fix registered planning paths in the registry and taught the
**self-owned-branch reclaim** sweep to consult `isPathActive`. The
leaked-slot reaper never got the same signal — fixed at one surface, not
enumerated across all. Exactly what the AGENTS.md Surface Enumeration
rule exists to prevent.

## Fix

The refusal now also fires when
`activeSessionRegistry.pathsForTask(taskId)` is non-empty. Keyed on
**any** registered path rather than on kind: the point is that a
registered surface of any kind means someone is working in that
worktree.

## Enumeration — the part that stops a third recurrence

The guard is a **chokepoint**, so this covers every caller rather than
just the reported one:

- `reapLeakedConcurrencySlots` — the reported path
- `recoverPausedAbortFailures` — **had the identical executor-only
pre-gate**
- the `preserveWorktrees: true` reclaim

Audited the rest of self-healing’s liveness gates: the self-owned-branch
reclaim, worktree-metadata reconcile and PR-branch sweeps already
consult `isPathActive`/`lookupByPath`. The three that read only
`getExecutingTaskIds` — `checkStuckBudget`, `recoverCompletedTasks`,
`recoverStrandedCompletedTodoTasks` — move columns and never destroy a
worktree, so they are noted rather than changed.

## Trade-off, stated plainly

A leaked registry entry now blocks this sweep instead of a live planner
losing its worktree. That is the strictly safer failure and the one the
“last line of defense” wording already promises. The registry is
process-local and in-memory, so a leak cannot outlive the process, and
stale entries have their own reconciler. **A test pins that a genuine
phantom — no executor surface AND no registration — still clears**, so
this is not a blanket refusal that would trade this bug for a wedged
queue.

**The 60 s grace is deliberately unchanged.** Raising it would only make
the bug rarer and harder to reproduce; the liveness gate was the defect.

## Verification

Revert-proof, measured: removing the registry term turns **3 of the 4**
new tests red, including the end-to-end sweep case (card in `triage`,
past the grace, executor sets empty → asserts the slot is not reaped and
the worktree survives). The 4th stays green both ways *by design* — it
is the anti-overcorrection guard.

`pnpm lint` clean · engine `tsc` clean · `pnpm test:gate` green (309 +
10 + 71) · new suite 4/4. The 2 failures in `self-healing.test.ts` /
`-completion-fanout.test.ts` are **pre-existing** — identical with this
change stashed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Prevented active planning worktrees from being mistakenly deleted or
reclaimed while related planning sessions are still active.
* Enhanced session liveness checks so phantom executor bindings are not
cleared when a live session is registered.
* Updated paused abort recovery to defer or abort safely when a live
planning session is detected, avoiding unintended task/worktree
mutations.
* **Tests**
* Added regression coverage for leaked-slot reaping, paused abort
recovery behavior, phantom binding refusal, and end-to-end sweep
outcomes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:04:51 -07:00
Phil Larson
72391c90b2 fix(engine): route workflow reviews through validator models (#2533)
## Summary

- classify review-type workflow steps with the existing review-step
classifier
- resolve their primary, fallback, and thinking-level settings from the
validator model lane
- retain per-step model overrides and executor-purpose workflow-step
tooling
- keep ordinary workflow steps on the execution lane
- make missing-fallback diagnostics identify the correct lane

## Why

Code Review, Plan Review, verification, and inline-review gates were
executed through the implementation model lane merely because they run
inside `executeWorkflowStep()`. That defeats configured reviewer-model
separation and can make the same model implement and validate its own
work.

This changes model selection—not the workflow-step session/tooling
contract—so review steps remain executor-purpose sessions while using
validator lane models.

## Verification

- `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/engine
exec vitest run src/__tests__/executor-workflow-step-model.test.ts` — 14
passed
- `corepack pnpm@10.33.0 --filter @fusion/engine typecheck`
- `corepack pnpm@10.33.0 changeset status --since=origin/main`
- `git diff --check origin/main...HEAD`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Review-type workflow steps now route through the configured validator
model lane (instead of the execution lane).
* Validator primary/fallback and thinking-level settings are applied
correctly for review steps.
  * Step/task overrides still take priority over lane-based resolution.
* Fallback retry sessions now use the appropriate validator/executor
configuration, with lane-specific fallback guidance when fallback
settings are missing.
* **Tests**
* Expanded executor workflow-step model resolution and routing/fallback
precedence assertions for validator-lane behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-29 00:05:04 -07:00
gsxdsm
9d3e53d0c5 U8 PR3: the implementation phase announces HOW it ended — including when the executor moved the card itself (#2507)
Third PR of **U8 — the graph owns execution**. Independent of everything
merged so far; small, green, revertable on its own.

## The problem this makes visible

`result.taskDone` is the entire language the execute seam has for
talking to the graph:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
return { outcome: "failure", value: paused ? "implementation-paused" : "implementation-incomplete" };
```

The endings that one bit cannot express are exactly the ones the
implementation phase **transitions itself**:

- a session that paused *after* the work was already complete →
finalizes to review inline;
- a session that stopped because a step is blocked on a pending review →
hands off to review inline (a pending-review block is a wait, not a
failure; marking it failed deadlocks a row that is both `in-review` and
`failed`).

The graph then sees `taskDone === false`, reports
`implementation-incomplete`, and `handleGraphFailure` compensates with
`alreadyFinalizedToReview` / `completionFinalized` — classifiers whose
entire job is recognising a move the graph did not make.

**That was invisible.** An out-of-band transition and a genuine
implementation failure were indistinguishable in logs, in events, and in
tests. You cannot remove a transition you cannot see, and you cannot
prove you removed it either.

## What lands

A closed `ImplementationExit` enum
(`engine/executor/implementation-exit.ts`) reported from six
completion-adjacent exits in `runImplementation`, announced by the
execute seam as `NodeCompleted.exit` on the U3 lifecycle bus. Two ids
are flagged as out-of-band — the ones where the executor, not the graph,
performs the transition.

**Routing is unchanged, and that is the point.** The seam returns
byte-identically what it returned before for every exit, so this PR
cannot move a card. The routing move needs new IR edges and lands
separately; splitting them is what keeps both independently revertable.
Per R5 an exit id is a **reaction** — nothing branches on one, and
dropping every subscriber must change no outcome (a named U8 test
scenario, asserted here).

`NodeCompleted.exit` is added to the event key allow-list deliberately —
which is exactly what that allow-list is for — and carries closed enum
ids only, never prose.

## Revert-proofs, each observed failing

| Injected change | Result |
|---|---|
| Remove the emit entirely | **6 failures** |
| Let an exit change the returned outcome | **2 failures** (the
routing-unchanged pins) |
| Delete one `reportImplementationExit(...)` call site | **1 failure**
(the wiring ratchet) |

**The third proof exists because of a hole I found in my own tests.**
These tests stub `runImplementationPhase` — the only way to reach all
six exits deterministically — which means deleting a real call site left
the entire file **green**. A stubbed seam can only prove the seam. I'd
also written "every exit is reported — the signal is real, not a
placeholder" in the header, which the tests did not support. Both are
fixed: there is now a ratchet asserting every enum id is wired at a real
call site and that each out-of-band id sits adjacent to the handoff it
describes, and the header says what the tests actually prove.

## Scope

**6 of `runImplementation`'s ~28 dispositions** (per the ownership
ledger merged in #2490), chosen as the ones the routing move needs. The
remaining ~22 report nothing yet — the ledger, not this enum, stays the
record of that gap, and the module says so.

## Verification

- 15 new tests + ledger + graph-boundary + task-done-blocked +
graph-requeue-gate + step-session + review-verdicts + tool-failure-retry
— **9 files, 115 tests green**
- `@fusion/core` `workflow-events` — 20 tests green (allow-list change
covered)
- `pnpm test:gate` green (17/307, 2/10, 1/71); `pnpm lint` clean; `tsc
--noEmit` clean on both packages
- Changeset included (`patch`, `internal`), passes `check:changesets`

## Next

PR4 is the routing move itself: `review-handoff-pending-review` becomes
a graph outcome with its own IR edge, and `alreadyFinalizedToReview`
becomes provably unreachable for that path. The IR edge change will be
its own commit, separate from the seam change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 17:35:07 -07:00
gsxdsm
eaea082259 U8 PR2: the execution-policy ladder resolves its own workflow's columns (the wip literal made retry, escalation and loop protection unreachable) (#2497)
Second PR of **U8 — the graph owns execution**, independent of
[#2490](https://github.com/Runfusion/Fusion/pull/2490) and of every
other unit. Small, green, independently revertable.

## The defect

`handleGraphFailure`'s execution-policy ladder — FN-7863/FN-7926
dispatch-loop terminalization, FN-7996 tool-failure retry, FN-7998
escalation — decided a task's own lifecycle by naming `"todo"` and
`"in-progress"` **literally, at 9 sites**. U5b converted the executor's
*rebounds* to `resolveReboundColumnFor`; these were left behind, each
sitting somewhere an awaited resolver could not reach: inside
synchronous `updateTaskAtomic` mutators, inside fire-and-forget resume
closures, and in conditions evaluated before any resolution happened.

**The severe one is the wip gate, and it fails silently in the worst
direction:**

```ts
if (live.column !== "in-progress") {
  // "Workflow graph run ended after task already advanced — no further action needed"
  return;
}
```

Under a workflow that renames the implementation column, that is true of
a card sitting in **its own wip column**. So the graph failure was
swallowed whole — no terminal park, no status, no error, nothing on the
board — and the scheduler re-dispatched the same doomed run. Every later
branch sits behind that gate, which is why the retry budgets, the
escalation, and the bounded terminalization were **unreachable rather
than mistargeted**.

This is precisely the failure the program's problem frame predicts: *a
guard that stops matching disables a recovery path invisibly and the
suite stays green.* I found it because my first renamed-column test for
the escalation site could not reach the escalation code at all.

Two further sites misbehave once the gate is passable:

- **FN-7998 node escalation** wrote `column: "todo"` inside the atomic
claim — parking the card where no workflow declares it, which is on the
plan's **"Stop implementation if"** list and what R7 exists to clean up
after. The scheduler's effective-node resolution, the entire point of a
node escalation, never runs.
- **FN-7863/FN-7926's `live.column === "todo"` arm** is the classic
guard that stops matching. In-process the `executeNodeSelfRequeued`
marker covers the same case, so this degrades only on the **durable**
arm — after a restart, or for a second `TaskExecutor` instance in the
process, where the column read is the only evidence the inner executor
requeued. A progressing card then falls through to the terminal sink and
is parked `failed`.

## The fix

Resolve hold and wip **once per graph failure** through U1's
`resolveTaskLifecycleColumns` and thread the pair through the ladder.
Both fall back to the legacy literal when the workflow cannot be
resolved, so an unresolvable workflow keeps exactly its pre-conversion
behavior rather than guessing. One IR read on a terminal recovery path —
not an enumeration loop.

## Red-green, measured

**3 of the 8 new tests fail with this commit's executor change
reverted:**

```
FAIL  FN-7998 … > requeues a node escalation to the RENAMED hold column, not the literal todo
FAIL  FN-7998 … > still does not move the card for a MODEL-target escalation
FAIL  FN-7863/FN-7926 … > recognises an inner-executor requeue that landed in the RENAMED hold column
      Tests  3 failed | 5 passed (8)     ← reverted
      Tests  8 passed (8)                ← with the fix
```

The other **5 pass both ways by design**, and I am not claiming them as
red-green — they are the regression floor:

- default coding workflow still resolves hold → `todo`, wip →
`in-progress` (byte-identical);
- an unresolvable workflow still uses the legacy literals;
- the in-process self-requeue marker still works when no workflow
resolves;
- and a **negative case** proving the dispatch-loop gate stays narrow —
a card still in its wip column with no marker is a genuine execute
failure and must NOT be swallowed as a benign recovery. Widening that
gate to "any column" would have been the easy wrong fix.

## Scope

Deliberately the execution-policy ladder only. **20 further column
literals remain in the same method's pause-abort, merge, and in-review
regions** — they belong to U5's executor slice (B4, not started) and
U9's merge lane, and are untouched here. Flagging the overlap: this PR
edits `executor.ts`, so whoever takes U5-B4 should rebase onto it rather
than converting these 9 sites again.

## Verification

- 8 new tests + the preserved-behavior suites
(`executor-tool-failure-retry`, `executor-graph-requeue-gate`,
`executor-task-done-blocked`, `executor-graph-boundary`,
`executor-stuck-requeue-preserve-progress`,
`executor-paused-abort-todo-benign`, `executor-abort-provenance`) — **9
files, 112 tests, green**
- `pnpm test:gate` — green (2/10, 16/299, 1/71); `pnpm lint` clean; `tsc
--noEmit` on `@fusion/engine` clean
- Changeset included (`patch`, category `fix`), passes `pnpm
check:changesets`

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed execution recovery for workflows with renamed lifecycle columns
so retry, escalation, and loop-protection behaviors correctly follow the
workflow’s declared hold/WIP columns.
* Preserved legacy behavior for default workflows and continued safe
handling when lifecycle columns can’t be resolved.
* **Tests**
* Added a Vitest suite validating execution-policy “ladder” behavior for
renamed columns, including node escalation, dispatch-loop gating, and
fail-closed scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 15:45:47 -07:00
gsxdsm
319e051c65 U8 PR1: pin the execution-lifecycle ownership ledger (measured: 28 executor-owned dispositions vs 3 graph handbacks) (#2490)
First PR of **U8 — the graph owns execution** (plan
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`,
line ~436). The plan states this unit "is expected to land as several
commits; it must not be attempted as one sweep", and
Execution-note-first: **characterization before ownership moves**. This
is that floor. **No behavior change.**

## Why a ledger and not a refactor

U8's goal is "the executor stops deciding *what happens next*" — and
that had no measurable form.

- **Executor line count does not measure it.** A 3,178-line
`runImplementation` can shrink substantially with every lifecycle
decision still exactly where it was.
- **A green suite measures it least of all.** Every disposition counted
below already has passing tests, because each one was *correct behavior*
when it was written. What is wrong is the **owner**, not the behavior.

So the unit needs a number, and the number has to exist *before* the
migration — a ratchet written afterwards cannot prove the migration
happened.

## The measured baseline

Counted from source, comments stripped, method bodies extracted by brace
matching:

| Method | `store.moveTask` | `handoffTaskToReview` | terminal
`status:"failed"` | `graphCompletion` handbacks |
|---|---:|---:|---:|---:|
| `runImplementation` (3,178 lines) | 16 | 3 | 9 | **3** |
| `handleGraphFailure` (~930 lines) | 0 | 0 | 7 | — |

**The implementation phase decides its own lifecycle 28 times and asks
the graph 3 times.**

These are measured, not estimated. My first `handleGraphFailure`
estimate was **wrong** (2 moves / 4 parks); the extractor corrected it
to 0 / 7 — the `moveTask` calls that read as belonging to that method
sit past its closing brace, in the recovery helpers below it. The
correction is in the ledger comment so the next reader does not repeat
the misread.

## The finding this makes concrete

`createAuthoritativeWorkflowSeams.execute` collapses that entire
implementation phase to one boolean:

```ts
if (result.taskDone) return { outcome: "success", value: "implemented" };
```

The graph has no vocabulary for *"the agent stopped because a step is
blocked on a pending review"* or *"the session paused after the work was
already complete"*. So the implementation phase performs those
transitions itself (`executor-exit-while-review-pending`,
`paused-after-completion`) and the graph finds out afterwards.

That is why `handleGraphFailure` carries `alreadyFinalizedToReview` /
`completionFinalized` — **classifiers whose entire job is to recognise a
move the graph did not make.** They are compensation for dual ownership,
and they are U8's acceptance test: they become unreachable, and then
deletable, exactly when the last out-of-band transition is gone. This PR
records that contract in source at the seam (FNXC comment), which is
where the next PR starts.

## Proof the guard fails on the defect

A ratchet that reports success without checking anything is worse than
no ratchet. Both failure modes were injected and observed:

1. **The defect it exists to catch** — injected one `await
this.store.moveTask(task.id, "in-review", {})` into
`runImplementation`'s completion path → ledger fails, `16 -> 17`.
2. **A broken guard** — injected a string literal containing `}` so
naive brace matching ends the body early → the size self-check fails at
**13 lines**, instead of silently reporting a comfortable zero for every
count.

Both injections were reverted; `git diff` against the pre-injection copy
is empty.

## Direction of travel

Executor-owned counts may only go **down**, and a decrement must land
with the disposition visible as a **graph outcome** — not merely
deleted. An increment is a new out-of-graph lifecycle decision and needs
a stated justification in its PR, not a quiet edit to the constant.

This is the precursor to U12's planned
`no-out-of-graph-lifecycle-writes.test.ts`; when the counts reach their
floor the assertion becomes "zero, outside the allowlist", and this file
is where that allowlist grows up.

## Preserved behaviors

Untouched, and re-run green as the regression floor for everything that
follows: FN-8141 honest-blocked exit
(`executor-task-done-blocked.test.ts`), FN-7996/FN-7998 tool-failure
retry + escalation (`executor-tool-failure-retry.test.ts`), FN-7863
dispatch-loop terminalization and FN-7926 completed-blocked parking
(`executor-graph-requeue-gate.test.ts`).

## Verification

- `pnpm --filter @fusion/engine exec vitest run` on the ledger + the
four preserved-behavior suites + `legacy-tombstones` — **6 files, 49
tests, green**
- `pnpm test:gate` — **green** (2/10, 16/299, 1/71)
- `pnpm lint` — clean; `tsc --noEmit` on `@fusion/engine` — clean

No changeset: test-only plus a source comment, no `@runfusion/fusion`
behavior change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added a lifecycle-ownership “source-scanning” test that analyzes the
executor’s task disposition patterns to ensure counts remain consistent
across execution and graph-failure flows.
  * Added safeguards to catch unintended changes to lifecycle handling.

* **Documentation**
* Documented the lifecycle-ownership boundary for task disposition
handling, including how completion and failure transitions are
consolidated and how related failure classifiers are affected.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 21:09:59 -07:00
gsxdsm
2dce642ccc E2E validation: run a RENAMED-column workflow against a live engine (real graph + real PostgreSQL) (#2475)
Stacked on #2472 (`feature/workflow-vocabulary-b3-stranded-todo`).

Test-only. No production file is touched.

## Why

Every slice of this program has closed with the same caveat: *no renamed
workflow was run against a live engine; all evidence is unit-level*.
That caveat is load-bearing — eight times this session a test passed
without exercising its subject. This PR removes it for the lifecycle
spine.

## What actually runs

`packages/engine/src/__tests__/workflow-lifecycle-live-e2e.pg.test.ts`
drives the REAL pieces:

- a **real PostgreSQL `TaskStore`** on a throwaway per-file database
(shared PG harness; never the operator's DB, never port 4040),
- the **real graph interpreter** (`WorkflowGraphTaskRunner`) with the
**real column-boundary controller** wired to the **real
`store.moveTask`** — all of its guards, traits, capacity reservation,
and post-commit emission,
- the **real scheduler release** (`runHoldReleaseSweep`),
- the **real post-commit lifecycle bus** (`getWorkflowEventBus`),
- the **real converted self-healing sweep**
(`SelfHealingManager.recoverStrandedCompletedTodoTasks`, slice B3.1).

Only the AI **seams** are scripted — the same boundary `testMode`/`mock`
draws in production.

**Assertion rule:** every lifecycle claim is asserted on **persisted
state** (a fresh `getTask` with the store's task cache defeated,
`run_audit_events` rows, `workflow_work_items` rows), never on "a
function was called". The one spy — the event-bus subscriber — is
asserted on the **received payload**, because the bus silently drops
events that fail its shape check, so "emit was called" proves nothing.

**Differential design:** the default-vocabulary
(`todo`/`in-progress`/`in-review`/`done`) and renamed-vocabulary
(`backlog`/`building`/`checking`/`shipped`) workflows come from ONE
builder and differ ONLY in their four column ids. Any behavioral delta
is attributable to the vocabulary alone.

## Coverage (9 tests, all green)

| Scenario | What is proven |
|---|---|
| Default vocabulary, full spine | planning runs in the hold column, the
card parks (graph does not self-promote), the **scheduler** performs
hold→wip, the resumed run walks exec → review → merge-gate → end,
persisted column is `done` |
| **Renamed vocabulary, full spine** | identical, and no leg of the run
touches any legacy column id |
| Audit differential | the graph-owned boundary crossings are the same
crossings node-for-node on both vocabularies; no legacy id appears in
the renamed trail |
| Event seam | a real subscriber **receives** a well-formed
`TaskTransitioned` for the renamed `backlog`→`building` release and for
the terminal move; `NodeEntered` arrives for every traversed node
including `end` |
| Crash / restart | exactly one durable continuation row at `exec`; a
brand-new runner resumes from the row and the already-completed
`planning` seam does **not** re-run; no duplicate continuation |
| Converted sweep (B3.1) | a completed card in a **renamed** hold column
is promoted (asserted on its persisted column), a card in the renamed
**wip** column is not, and the default `todo` case still works |

## Mutation verification (both directions)

Green suites are not evidence in this codebase, so both halves were
falsified:

1. Keying `hold-release`'s `isHeldTask` on the `todo` literal → **5 of 6
spine tests fail, and the one that survives is the default-vocabulary
one.** That is the exact signature the conversion program cares about.
2. Reverting slice B3.1's per-task hold-column resolution to the literal
→ **only the renamed stranded-todo test fails**; the default regression
floor stays green.

## Findings surfaced by running it

1. **The IR validator refuses a `merge-blocker` column with no reachable
merge-class node** ("the gate can never clear without one"). Kept rather
than worked around — it means the review column here is genuinely gated.
2. **Entry into the merge region collapses to the legacy `merge` seam**
(`MERGE_REGION_KINDS`), so a `merge-gate` node reaches the merge lane.
Documented in the fixture.
3. **The transition policy refuses a direct hold → review move**, and it
refuses it *workflow-resolved*: on the renamed board the only legal
target is its own `building`, not `in-progress`. The recovery callback
therefore promotes hold → wip → review rather than bypassing the policy.
4. **`moves.ts` still special-cases the `done` literal** (`if (toColumn
=== "done") clearNearDuplicateReferencesTo...`) after the post-commit
emit. Not converted here and not in this PR's scope — flagged for the
Phase B owner.

## Not driven end to end (stated plainly)

- **Triage / specification.** The lifecycle starts from a task already
bound to a workflow; `triage.ts` was not driven. The `planning` seam is
scripted.
- **Real merge.** No git worktree, no branch, no squash. `merge-gate` is
pure policy; the `merge` seam is scripted.
- **Lightweight / self-healing-off workflow.** The Tier 1 policy keys do
not exist on this tip — there is no `policies` surface on the IR to set.
Not drivable; not substituted with a unit test.
- **Process-level crash.** The restart is an in-process one: a brand-new
runner resuming from the persisted `workflow_work_items` row with no
carried-over memory. No OS process was killed, so this proves
durable-state resumption, not signal handling.

## Lane

`.pg.test.ts` under the engine-default include glob, gated by
`pgDescribe` so it skips cleanly with no PostgreSQL. The merge gate is
untouched. Engine `tsc --noEmit` is clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added comprehensive live PostgreSQL workflow lifecycle coverage,
including graph execution, suspension and resume, scheduler capacity
release, crash recovery, and durable continuation.
* Added validation for renamed workflow column configurations and
columnless task movements.
* Added event delivery checks for task transitions and node entry
events.
* Added self-healing recovery for stranded completed tasks in valid hold
columns.

* **Refactor**
* Centralized workflow boundary handling, including task moves,
continuation state, audit events, and diagnostics.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:05:05 -07:00
gsxdsm
5ae6332563 refactor: collapse dead SQLite dual-path code; keep migration-only readers (#2454)
# Remove dead SQLite dual-path code; keep migration-only readers

## Summary
PostgreSQL cutover left hundreds of production dual-path branches
(`backendMode ? PG : SQLite/store.db`) whose SQLite arms only hit
throwing `Database`/`ArchiveDatabase`/`CentralDatabase` stubs. This
change mechanically collapses those unreachable arms so production
authority is AsyncDataLayer/PostgreSQL only, while preserving the six
authorized read-only migration/recovery `DatabaseSync` seams.

## Dual-path mass removed
| Metric | Before | After |
|---|---|---|
| `if (…backendMode)` (non-test) | ~328 | ~70 |
| `store.db` / `this.db` refs in core (non-test) | ~570+ | ~375 (mostly
pure legacy MissionStore/eval/insight SQLite classes + thin getters) |
| Net diff | — | **~6.7k lines removed** across 41 files |

Remaining `backendMode` checks are intentional (incomplete-PG sync
safe-defaults, settings-sync disabled-on-PG, symbol-lock PG-only gates,
“requires PostgreSQL” config versioning throws), not live SQLite
authority.

## Subsystems cleaned
- **Core TaskStore / task-store/***: collapsed if/else and early-return
dual-path across reads, moves, lifecycle, mutations, workflow, archive,
branch/PR, artifacts, comments, audit, project ops, etc. `initImpl` is
PostgreSQL-only (SQLite startup tail deleted).
- **Satellite stores**: automation, agent, routine, plugin, secrets,
approval-request, central-core dual-path arms collapsed.
- **Plugins**: reports async methods, compound-engineering pipeline +
session stores, CLI Printing Press store — SQLite fallbacks removed; PG
required.
- **Engine**: no functional dual-path change beyond whitespace
(settings-sync / peer-exchange PG-disabled behavior kept).

## Six migration-only readers retained (allowlist unchanged)
1. `packages/core/src/postgres/sqlite-migrator.ts`
2. `packages/core/src/project-identity.ts`
3. `packages/core/src/sqlite-validation.ts`
4. `packages/core/src/postgres/startup-factory.ts`
5. `packages/cli/src/commands/db.ts`
6. `scripts/lib/start-local-project.mjs`

Plus low-level `sqlite-adapter` and migrator/startup-import tests.
Inventory ratchet still requires exactly these six `new DatabaseSync(`
production sites, all `readOnly: true`.

## Not treated as SQLite
- `.fusion/project.json`, `task.json`, `agent-log.jsonl` file storage
- AsyncDataLayer / Drizzle PG paths
- Incomplete-PG sync safe-default stubs (still return empty/false/null
under backend without consulting SQLite)

## Verification
- `sqlite-production-reader-inventory.test.ts` — 15/15 pass
- `incomplete-pg-ports.pg.test.ts` — 6/6 pass
- Targeted PG tests (create-task, move, handoff, runtime-persistence,
agent, mission, insight, central-core) — green
- `tsc --noEmit` for `@fusion/core`, `@fusion/engine`,
`@fusion/dashboard` — green
- `scripts/check-no-getdatabase.mjs` — clean

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Improved end-to-end consistency by making PostgreSQL/async persistence
the standard across core task/workflow, automation, agents, plugins,
routines, secrets, approvals, central operations, and session storage.
* Unified scheduling, settings, configuration revision writes,
run/workflow selection, queues/leases/transitions, and audit/lifecycle
updates around consistent async transaction behavior.
* **Bug Fixes**
* Fixed edge cases for archived/deleted reads, unarchive/recovery flows,
not-found handling, and task/artifact/document/log/comment operations,
including more reliable emissions and hydration across search/list and
lifecycle operations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 23:28:42 -07:00
Phil Larson
52d64fa66e fix(engine): project CE steps after review handoff (#2464)
## Summary

- reconcile successful graph-native workflow results with pending task
checklist steps even when review handoff already moved the card into the
merge column
- preserve terminal, paused, and no-redundant-move behavior
- cover the real Compound Engineering post-review-handoff state with a
regression test

## Root cause

Compound Engineering runs `review-handoff` before `merge`. Review
handoff moves the task to `in-review`, which is also the merge column.
`ensureWorkflowMergeBoundaryTask()` returned immediately for cards
already in that column, before projecting successful
`workflowStepResults` onto legacy `Task.steps[]`. The merger then saw
`0/N` and rejected approved work with `task has incomplete steps`.

## Verification

- RED: regression test failed before the fix because `store.updateTask`
was never called
- GREEN: `executor-graph-boundary.test.ts` — 6 passed
- relevant non-PostgreSQL set — 31 passed, 5 PostgreSQL tests explicitly
skipped
- `@fusion/engine` typecheck passed
- changeset format passed
- `git diff --check` passed

## Baseline note

`ce-workflow-step-executor.test.ts` currently has three failures on
clean `origin/main` after FN-8601 foreach-proof hardening. The same
failures reproduce without this patch and are not regressions from this
change.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved reconciliation after review handoff by projecting completed
step results onto the legacy checklist when reaching the merge column.
* Prevented tasks from being marked approved with incomplete step counts
(including “0/N” style states).
* Reduced unnecessary merge failures and deadlock/pause scenarios when
merge-column progress was already recorded.
* **Tests**
* Added coverage for execute-and-merge workflows, ensuring
merge-boundary resolution updates pending steps without moving the task.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 23:13:21 -07:00
gsxdsm
f01461a70e feat(engine): thread the node id onto review-gate leases, activating pre-boot reclaim
Completes 3b83282273. The classifier and the self-healing reader landed there but
nothing stamped `leaseNodeId`, so the pre-boot reclaim path was unreachable.

Wiring: InProcessRuntime -> TaskExecutor -> WorkflowGraphTaskRunner ->
WorkflowGraphExecutor, which writes the field onto the pending lease.

The executor takes `getLocalNodeId`, a GETTER rather than a value, because the
runtime resolves the node id asynchronously (a CentralCore read) partway through
start() while `executorOptions` is built earlier in the same method. A snapshot
taken at construction would freeze `undefined` and silently disable attribution
forever -- the failure mode where the feature looks wired, typechecks, and never
fires. Reading it at runner-construction time picks up the resolved id.

With this, a review gate whose session dies to an engine restart is reclaimed on
the next self-healing pass instead of waiting out the 15-minute staleness floor.
Peer-owned and legacy unattributed leases still take the floor, so the
double-dispatch protection multi-node depends on is unchanged.

Adds five classifier cases: own-node pre-boot reclaims; peer-node, unattributed,
own-node-post-boot, and no-identity-supplied all still adopt. Verified the first
is not vacuous -- disabling the branch fails exactly that case (1 failed / 13
passed) and no other.

Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green
(299 + 10 + 70), plan-review-lease + plan-review-single-owner +
self-healing-orphaned-pending-step-results green (27).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 12:29:54 -07:00
gsxdsm
00011b0113 fix(engine): recover restart-orphaned review steps in one cycle, raise fix budget
FN-8603 sat in-review for ~36 minutes after an engine restart killed its Code
Review session 34 seconds in. It did recover on its own; the cost was latency,
not a terminal park.

Sweep ordering. reconcile-orphaned-pending-step-results PRODUCES the failed
results that recover-failed-pre-merge-steps CONSUMES, but in the periodic
maintenance list it ran ~15 entries after it. A step orphaned in cycle N was
therefore rewritten to failed only after recovery had already scanned, so
nothing re-ran it until cycle N+1. Moved it immediately before its consumer and
removed the now-duplicated later entry. Startup recovery already ordered the two
correctly.

Post-review fix budget. Default raised 3 -> 10 per operator request. Three
passes is below the observed convergence length for the gates this fallback
actually governs -- Browser Verification and custom optional gates -- since Plan
Review and Code Review already resolve to "unbounded" when unset, and exhausting
the budget parks the card for a human. The declaration default and five inline
`settings.maxPostReviewFixes ?? 3` call sites in executor.ts/self-healing.ts had
drifted into separate literals, so raising one alone would have left every
unset-settings path on the old value; they now share the exported
DEFAULT_MAX_POST_REVIEW_FIXES.

Not done, and why. Re-dispatching a restart-orphaned lease immediately at
startup is the change that would close the remaining ~14-minute wait, but it is
unsound as specified: liveness is judged by a 15-minute lease-staleness floor
because leases carry no node attribution, so treating a pre-boot lease as dead
would let one node orphan another node's genuinely running review. Needs a node
id on the lease record first. Left the floor intact.

Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green,
self-healing orphaned-pending-step-results and optional-step-revision suites
green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 12:08:03 -07:00
gsxdsm
a00f2633ce fix(engine): demote more TUI chatter across merger, self-heal, and ntfy
Route foreach/merger/worktree/self-healing skips, ntfy send bookkeeping, session-purpose runtime picks, planning using-model, and checkpoint rewind lines to debug so recoveries and failures stay visible in the operator log.
2026-07-26 10:15:18 -07:00
gsxdsm
9bad0e1233 fix(engine): demote high-frequency TUI log spam to debug
Route process spawn/exit, verification success paths, MCP connect, skill info listings, createFnAgent/session bookkeeping, and executor dispatch chatter through FUSION_DEBUG so the operator log pane keeps real lifecycle outcomes.
2026-07-26 09:50:43 -07:00
gsxdsm
ae512aec2b FN-8601: enforce foreach merge proof
Require complete foreach execution evidence before workflow merge review.

- Add reusable foreach instance coverage proof evaluation.
- Block checklist projection and merge admission on incomplete or failed node results.
- Cover core proof logic and PostgreSQL merge-boundary behavior.
- Add a patch changeset for the merge safeguard.

Files changed:
 .changeset/fn-8601-foreach-merge-proof.md          |   7 ++
 .../src/__tests__/workflow-merge-proof.test.ts     |  43 ++++++++
 packages/core/src/index.gate.ts                    |   2 +
 packages/core/src/index.ts                         |   2 +
 packages/core/src/workflow-merge-proof.ts          |  74 +++++++++++++
 ...xecutor-merge-boundary-foreach-proof.pg.test.ts | 111 +++++++++++++++++++
 packages/engine/src/executor.ts                    | 117 +++++++++++++--------
 7 files changed, 314 insertions(+), 42 deletions(-)

Fusion-Task-Id: FN-8601

Fusion-Task-Lineage: 40578171-0b13-4538-8f38-3948ed1e92c0

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 09:40:37 -07:00
gsxdsm
795a38c018 fix(engine): quiet graph review-entry audits and label engine aborts truthfully
Recognise workflow-graph moves into in-review so gate entry no longer emits handoff-invariant violations, and split pause-abort provenance so engine teardowns are engine-abort instead of hard-cancel.
2026-07-26 08:56:58 -07:00
gsxdsm
f005cee885 fix(engine): surface silent stalls and add stalled-card watchdog
Make planning-guard and remediation no-ops emit warnings, and detect idle non-terminal cards with no session or continuation so FN-8596-class strands show up in logs and run-audit.
2026-07-26 07:49:05 -07:00
gsxdsm
106c61e6ee fix(agent-tools): close the fn_delegate_task Deny bypass and the store's window clamp
Follow-up to 13a2b2a9d, from a multi-agent review of that commit. Three of its
claims did not hold.

1. fn_delegate_task bypassed the gate entirely (P0). It reaches the same
   createAgentTask primitive, was registered unconditionally in both session
   lanes, and validated only that the TARGET agent is non-ephemeral — never the
   caller. Under Deny an ephemeral worker could enumerate agents and delegate
   unlimited tasks. It is now withheld under Deny, and also under
   upon_validation: delegation has no proposal channel, so leaving it available
   would launder a create past the operator review that policy requires.

2. The widened dedupe window was capped at 5 minutes. The store query in
   branch-and-pr-entities.ts carried its own independent `?? 60_000` /
   `min(300_000, …)` pair, so widening only duplicate-guard.ts under-delivered
   and made the new ceiling unreachable. Both sites now share
   FINGERPRINT_WINDOW_DEFAULT_MS / FINGERPRINT_WINDOW_MAX_MS.

3. The pi-extension gate does not fire at all. pi's ExtensionContext carries no
   agentId — the read is a speculative cast and only tests supply one, so every
   real call short-circuits as a human caller. The fail-closed direction is kept
   for the day an identity signal exists, but the limitation is now documented
   instead of implied to be enforcement.

Also: the session prompt now states when creation is disabled and names
fn_task_log as the fallback (the base prompt still taught fn_task_create, which
is the same instruction/capability mismatch that fed the retry storm);
suppression emits an `agent:task-create-withheld` run-audit event; and the two
source-text ratchet tests are replaced with behavioral assertions on the tool
list the executor actually hands the model — verified to fail when the guard is
broken, which the string assertions did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:44:56 -07:00
gsxdsm
0c85613313 fix(engine): address code-review findings on the planner/worktree recovery fixes
Review of 2dbfe3d31 + 05b704dc6 surfaced real defects in both fixes:

- The unusable-worktree probe composed two helpers across an unnecessary
  self-healing -> step-runner import edge, and the directory check added no
  discriminating power over the `.git` probe. Replaced with one canonical
  hasUsableWorktreeShape beside classifyTaskWorktree, which also applies the
  repo-root gate (FN-6861) when a rootDir is available; both call sites pass one.
  Its narrower guarantee vs the canonical classifier is now documented and
  pinned by tests, including the de-registered shape it cannot see.
- REPLAN_PARK_STATUSES is derived from PLANNING_STAGE_STATUSES instead of
  re-listed, so a new durable park status cannot be added to one set only.
- The preserve/clear decision no longer pretends to steer `worktree`: the rebound
  is a reopen move, which clears it regardless. Documented, and the test now
  asserts the durable row rather than only the updateTask argument.
- `branch` is cleared only when it is the re-derivable canonical fusion/<id>;
  a non-canonical branch survives so a card's only commit pointer is not dropped.
- The recovery log named the recorded worktree even when the session had targeted
  an AI-merge clean room. It now names the refused path and says whether the
  recorded worktree was gone too.
- Added task:auto-recover-worktree-session-metadata so the decision is legible to
  agents, not only in human log prose.
- isTaskStillInPlanningStage's parameter type now includes the execution stamps
  its implementation reads.
- Test hygiene: real-fs fixtures wrapped in try/finally; changeset dev note
  corrected; FN-8361 asserted at the discovery surface, not only in the guard
  table.

Also captures the shared bug class in docs/solutions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:24:10 -07:00
gsxdsm
13a2b2a9da fix(agent-tools): hide fn_task_create under Deny and widen the dedupe window
Operator report: with project policy "Ephemeral agent follow-up tasks = Deny",
an executing agent filed ten follow-up tasks — five parallel fn_task_create
calls it reported as timed out, then five sequential retries.

Two defects:

1. Deny was advisory. fn_task_create was registered for every session and only
   refused inside execute(), so the model still saw the tool, planned around it,
   and retried it. The pi extension's isEphemeralCallerAgent also failed OPEN
   whenever the caller id did not resolve to an agent row — which is the normal
   shape of an ephemeral task-worker — so on that lane Deny was a no-op.

2. The deterministic content-fingerprint duplicate window was 60s, which only
   covered concurrent in-flight creates. A retry two minutes later saw nothing
   and filed a second task.

Fixes: isAgentTaskCreateToolAvailable() withholds the tool from ephemeral
sessions under Deny in both engine lanes (outer execution session, per-step
workflow session); isEphemeralCallerAgent fails closed on an unresolvable
caller id; the fingerprint window goes 60s -> 10m (clamp ceiling 5m -> 1h).
upon_validation keeps the tool, and permanent-agent and human/chat callers are
unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:03:46 -07:00
gsxdsm
d10d91bae5 fix(engine): stop planning when a card is withdrawn; sweep stale pre-execution worktrees
Withdrawing a card from planning (todo -> Ideas) now stops the work:
- triage aborts and disposes the planning session through the same path
  pause/delete already use, and clears status:"planning" so the planning badge
  goes away and the card reads as a plain idea again;
- the executor aborts in-flight graph work on any backward move out of
  todo/triage, so a Plan Review does not keep streaming against a card the
  operator pulled back;
- moving it back to todo needs no new code: the existing column wake fires and,
  with the status cleared, the card is an ordinary planning candidate again.

Pre-execution worktrees (planning acquires one now) are reclaimed two ways: an
immediate release on an explicit withdrawal, and a self-healing sweep
`reconcile-pre-execution-worktrees`. The sweep is deliberately timid — 30 days
of complete inactivity, and it skips anything active or waiting (todo,
executing, in-review, done, paused, carrying any status, blocked, or scheduled
for recovery). Every real safety condition lives in the executor: never
executed, no live session, clean branch, nothing uncommitted.

hasAdvancedPastPlanning no longer reads a worktree as execution evidence.
Planning owns a worktree now, so that signal would have made every planning
write skip; execution timestamps carry the meaning instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 15:20:17 -07:00
gsxdsm
168819b35d fix(engine): run every lane in the task worktree; contention is a wait, not a failure
Contention prevention (why tasks shared a path at all):
- Planning ran `tools: "coding"` at the repo root, so every planner had write
  tools in the operator's checkout and all planners shared one path. Planning
  now acquires the task's own worktree (TriageProcessor.acquirePlanningWorktree
  -> TaskExecutor.ensureTaskWorktreeForPlanning).
- Graph nodes with no worktree acquired one instead of falling back to rootDir,
  so Plan Review / Code Review / custom gates all run isolated. Plan Review
  re-acquires when its recorded worktree is gone, replacing FN-7996's
  run-from-the-repo-root degrade. Workspace projects are unchanged.
- Registration goes through acquireActiveSessionPath, which reclaims a leaked
  entry whose holder is provably dead and aged past the FN-5256 floor. A live
  holder still contends — real serialization is never clobbered.

Classification (the reported symptom):
- A lease held by another task is no longer a provider failure. It carries
  SESSION_CONTENTION_HOLD_VALUE, classifies transient, is excluded from
  isNonPlanDefectPlanReviewFailure, and stops burning the node's fast retries.
- The executor waits it out on a 10-attempt 5s->60s ladder and then leaves the
  task cleanly queued. There is no terminal branch: contention always ends, so
  parking would only ask a human to press Retry on a condition that fixed
  itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 14:59:53 -07:00
gsxdsm
e6b2da6cae fix(engine): let concurrent tasks run Plan Review on the shared repo root
Plan Review needs no worktree, so it runs rooted at the project root. The
activeSessionRegistry key was the bare root path, so the second task to reach
Plan Review hit ActiveSessionPathHeldByForeignTaskError ("path ... is held by
task FN-1398; task FN-1403 may not overwrite it"). That surfaced as a Plan
Review provider failure, burned the in-place retry budget against a hold no
retry could clear, and left the task parked.

Task-scope the registry key for any session rooted at rootDir, in every project
mode — the workspace fix already did this for the shared browse-root. Root
exclusivity protects nothing here: write-capable nodes are refused at the root
outright, and every isPathActive consumer guards removable worktree paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 13:55:54 -07:00
gsxdsm
e5caea542a FN-8544: gate mission remediation behind autopilot
Keep mission validation report-only until an operator explicitly enables autopilot.

- Gate validator-created remediation features and task dispatch behind mission autopilot.
- Audit attributed status and autopilot transitions atomically across mission stores.
- Expose mission autonomy controls and document the opt-in lifecycle.

Files changed:
 .changeset/fn-8544-mission-autonomy-audit.md       |   7 ++
 docs/missions.md                                   |  10 +-
 packages/cli/src/extension.ts                      |  26 ++++-
 .../__tests__/postgres/mission-store.pg.test.ts    |  27 +++++
 packages/core/src/async-mission-store.ts           |  70 +++++++++++--
 packages/core/src/index.gate.ts                    |   3 +
 packages/core/src/index.ts                         |   3 +
 packages/core/src/mission-store.ts                 | 113 +++++++++++++--------
 packages/core/src/mission-types.ts                 |  22 ++++
 .../dashboard/app/components/MissionManager.tsx    |   1 +
 packages/dashboard/src/mission-routes.ts           |  25 +++--
 .../src/__tests__/agent-mission-tools.test.ts      |  17 +++-
 .../src/__tests__/mission-execution-loop.test.ts   |  21 ++++
 packages/engine/src/agent-heartbeat.ts             |   4 +-
 packages/engine/src/agent-tools.ts                 |  30 +++++-
 packages/engine/src/executor.ts                    |   5 +-
 packages/engine/src/mission-autopilot.ts           |  51 +++++-----
 packages/engine/src/mission-execution-loop.ts      |  49 ++++++---
 packages/engine/src/triage.ts                      |   5 +-
 19 files changed, 377 insertions(+), 112 deletions(-)

Fusion-Task-Id: FN-8544

Fusion-Task-Lineage: 23a69923-5a19-407e-9fe2-8973c166ee9a

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-23 15:25:09 -07:00
gsxdsm
227281dc32 FN-8503: preserve unbounded Code Review retries
Keep Code Review remediation retry policies accurate across graph execution and recovery.

- Preserve unlimited retry presentation when Code Review has no configured cap
- Enforce finite Code Review caps during failed-step recovery
- Validate non-negative revision settings and document the active retry policy

Files changed:
 .../fn-8503-unbounded-code-review-retries.md       |   7 ++
 docs/workflow-steps.md                             |   2 +-
 .../core/src/__tests__/builtin-workflows.test.ts   |   8 +-
 packages/core/src/builtin-workflow-settings.ts     |   4 +
 .../workflow-graph-optional-step-fix.test.ts       | 135 +++++++++++++++++++++
 packages/engine/src/executor.ts                    |  51 ++++++--
 6 files changed, 193 insertions(+), 14 deletions(-)

Fusion-Task-Id: FN-8503

Fusion-Task-Lineage: 7bd555d1-23e5-42ea-b6f5-0b9fe4da7f94

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-22 18:32:36 -07:00
gsxdsm
53e3063e9f FN-8490: load skills for foreach step-execute sessions
Honor skill-executor configuration for implementation sessions created by foreach templates.

- Propagate validated step-execute skill names through workflow seam context.
- Load namespaced and bare skills with configured discovery paths for pinned step sessions.
- Add regression coverage, workflow documentation, and a minor changeset.

Files changed:
 .changeset/fn-8490-step-execute-skill.md           |   7 ++
 docs/workflow-steps.md                             |   4 +-
 .../__tests__/step-execute-skill-loading.test.ts   | 128 +++++++++++++++++++++
 packages/engine/src/executor.ts                    |  62 +++++++++-
 packages/engine/src/workflow-node-handlers.ts      |  21 ++++
 5 files changed, 219 insertions(+), 3 deletions(-)

Fusion-Task-Id: FN-8490

Fusion-Task-Lineage: aa1ff02d-3139-45f2-8853-f53c0aef0f2f

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-22 15:07:59 -07:00
gsxdsm
56efd7488e fix(engine): stop false-positive stuck loop kills on iterative work (#2404)
## Summary

- Fix a false-positive in `StuckTaskDetector` where legitimate long
single-step work (E2E debugging, iterative fix/test cycles) was
classified as a loop and kill/requeued.
- Root cause: loop meant “no step status transition for
`taskStuckTimeoutMs` + high activity volume,” conflating **step
progress** with **actual activity**. Agents can stay productively busy
on one step for 10+ minutes with zero repetition.
- Loop now requires thrash evidence on top of volume + no step progress:
- **repetitive tool fingerprints** (`toolName` + primary-arg detail in a
sliding window), or
  - **elevated ignored step-update rebuffs** (≥ 10)
- Wire tool name/detail from `AgentLogger` → executor / step-session
into `recordActivity(...)` so novelty is measurable.
- Document the thrash-evidence rule in `docs/architecture.md`.

## Test plan

- [x] `pnpm --filter @fusion/engine exec vitest run
src/__tests__/stuck-task-detector.test.ts
src/__tests__/reliability-interactions/non-progress-churn.test.ts`
- [x] Regression: high-volume **diverse** iterative activity (174
events) does **not** classify as loop
- [x] High bare text/heartbeat volume without tools does **not**
classify as loop
- [x] Repetitive identical tool fingerprint + timeout **does** classify
as loop
- [x] Ignored step-update thrash (≥10) with volume **does** classify as
loop
- [x] Existing FN-5168 no-progress-churn + FN-6598 verification
suppression paths still pass
- [ ] CI gate green

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved stuck/loop classification by requiring explicit “thrash
evidence” (repetitive tool fingerprints and/or elevated ignored progress
rebuffs), reducing false positives for busy but diverse work.
* Updated loop evidence tracking to incorporate tool name plus
summarized tool-argument detail.
* Cleared loop evidence appropriately after verification, progress
updates, and task resumption.
* Extended tool-start telemetry/callbacks to include optional tool
detail.
* **Documentation**
* Refined loop-classification criteria to match the new evidence gates.
* **Tests**
* Updated/expanded stuck/loop and churn scenarios to validate the
evidence-based behavior and callback ordering.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-22 13:43:55 -07:00
gsxdsm
d194290a75 fix(engine): reject out-of-order step starts (#2403)
## Summary

Ordered task steps can no longer appear active ahead of unfinished
predecessors. Step starts now use the same dependency-aware ordering
guard as completions, while steps explicitly declared independent remain
parallelizable. Rejected executor updates explain that the lifecycle
transition was suppressed instead of implying completed work was
overwritten.

## Validation

- Reproduced the FN-8490 concurrent update sequence and verified later
steps remain pending.
- Passed 15 PostgreSQL step-order tests, the focused executor response
test, core and engine typechecks, changeset validation, and `pnpm
verify:fast` including boot smoke.
- The full `executor-prompt.test.ts` run retains five pause-behavior
expectation failures that reproduce unchanged on `origin/main`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Enhanced the step start hook to support an awaited “pre-start
projection” that can reject startup via `false` (sync or async),
preventing step-session creation/completion.
* Added a step-start “verdict” so steps can be started or blocked
deterministically (including “resumed” behavior).
* **Bug Fixes**
* Prevented ordered/dependency steps from transitioning out-of-order by
enforcing guards for both in-progress and done transitions, including
concurrent update attempts.
* Improved integrity/out-of-order warning behavior and suppression
details when persisted status doesn’t match expectations.
* **Tests**
* Added/updated PostgreSQL and engine regression coverage for
blocked/resumed start and start-rejection control flow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-22 11:02:57 -07:00
gsxdsm
6422cb93a4 fix(engine): stop overseer hard-cancel thrash on live step sessions (#2393)
## Summary

Prevents the FN-8471 failure mode where planner overseer `retry_step`
bounced `in-progress → todo` while a live step-execute session was still
coding, hard-cancelling the agent up to three times until recovery
budget exhausted.

Also closes concurrent resume races after plan-review release that
parked `status=failed` on a losing graph while a peer session still
owned work.

### Changes
- **Overseer live gate:** `retryStep` skips the hard-cancel bounce when
`isTaskLiveForOverseerRetry` is true; returns `false` so attempt budget
is not burned; durable skip log is deduped per task/stage.
- **Single-flight graph dispatch:** `executeCore` claims `graphRouting`
before any await; `executeWorkflowGraph({ alreadyClaimed })` owns
release.
- **Single-flight unpause resume:** claim `resumingUnpaused` before
await; treat existing graph claim as already-owned; clear claim before
completed-work recovery.
- **No false park:** execute-family graph endings with a peer live
session no longer stamp `status=failed` (merge-region failures still
park).

### Tests
- `executor-live-overseer-retry-gate.test.ts` — live probe matrix,
execute-family preserve, merge still parks
- `planner-overseer-intervention-wiring.test.ts` — live skip keeps
column in-progress and `getAttemptCount === 0`

## Test plan
- [x] `vitest run` scoped to the two new/updated test files (16 passed)
- [ ] CI gate (lint/typecheck/build/test:gate)
- [ ] Optional manual: fail a raced graph with a live step session and
confirm overseer does not bounce to todo

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved overseer-retry “live session” gating to avoid interrupting
active work, covering more live surfaces and preventing multi-resume
races.
- Updated failure handling so execute-family failures can be preserved
when another live session is still running, while merge-attempt failures
are still marked failed.
- Added deduping for “retry skipped due to live session” logs so they’re
emitted only once per task stage, and ensured the recovery attempt
budget isn’t consumed when intentionally skipped.
- **Tests**
- Added coverage for live-gating, retry-skip/budget behavior, and the
revised failure-parking rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-21 23:32:18 -07:00
gsxdsm
1e05793876 fix(ci): green full-suite bookkeeping after origin/main cutover (#2392)
## Summary

Restores green merge-gate and package-default suites after repeated
`origin/main` merges brought workflow-graph ownership cutover drift into
CI.

- Align engine/dashboard/core tests with post-cutover contracts
(`moveTaskIf`/`deleteTaskIf`, graph handoff, worktree-pool reclaim via
`removeWorktree` + `RemovalReason`, multi-step RESUMING parse,
soft-pause merge requester, graph-terminal failure surfaces).
- Small product fixes needed for real regressions uncovered by the
suite: soft-delete refuse before graph routing, skip DUPLICATE
step-heading withhold when an explicit marker is present, PG schema
applier guards, and related bookkeeping (research promote tool inventory
/ migration seed, stop shell `psql` in PG admin DDL).
- Quarantine/ledger hygiene only where required by standing rules; no
timeout/worker appeasement.

## Verification

- `pnpm test:gate` ×2 green
- `@fusion/engine` full package suite green (~9083 tests)
- Targeted core/dashboard clusters green (schema applier, agent-runs UI,
settings descriptions, mobile close)

## Test plan

- [x] `pnpm test:gate` (twice)
- [x] `pnpm --filter @fusion/engine test`
- [ ] CI full suite / PR checks on this branch
- [ ] Confirm no unrelated product behavior changes beyond the listed
regression fixes

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for `roadmap-item` native structure kinds, including
native structure embeds and metadata validation.
  * Added Stable and Beta release channel options in General settings.
* Added per-action reporting target configuration with clearer “unset”
guidance.

* **Bug Fixes**
  * Improved heartbeat/prompt behavior when patrol is disabled.
  * Prevented deleted tasks from continuing through execution.
  * Made recovery for explicit duplicate redirects more permissive.
* Hardened database migration and test database cleanup to reduce flaky
failures.

* **Documentation**
* Updated settings text for release channels, reporting targets, and
inheritance/unset behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-21 23:09:30 -07:00
gsxdsm
080a8e7134 FN-8464: guard baseline capture against invalid worktrees
Prevent baseline Git probes from using stale or non-directory task worktrees.

- Gate baseline capture on an existing worktree directory
- Defer graph step projection until worktree acquisition completes
- Cover missing, non-directory, and filesystem-race worktree paths
- Add a patch changeset for the operator-facing fix

Files changed:
 .changeset/fn-8464-baseline-cwd.md                 |   7 ++
 .../__tests__/executor-fast-mode-workflows.test.ts | 100 ++++++++++++++++++++-
 .../engine/src/__tests__/executor-test-helpers.ts  |   6 +-
 packages/engine/src/__tests__/step-runner.test.ts  |  62 +++++++++++++
 packages/engine/src/executor.ts                    |  16 +++-
 packages/engine/src/step-runner.ts                 |  24 +++++
 6 files changed, 210 insertions(+), 5 deletions(-)

Fusion-Task-Id: FN-8464

Fusion-Task-Lineage: e4116dd0-decd-4f9d-87f0-e695cc7f182b

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-21 19:40:36 -07:00
gsxdsm
648971634a FN-8461: suppress spurious workflow skill-load warnings
Prevent optional CE configuration from producing warnings when a requested plugin skill is discoverable.

- Merge plugin skill body directories with the optional CE discovery root
- Warn only when the named workflow skill lacks every viable discovery source
- Cover plugin, CE-namespaced, and unrelated-skill discovery cases

Files changed:
 .changeset/fn-8461-skill-load-warning.md           |   7 +
 docs/workflow-steps.md                             |   8 +-
 .../__tests__/ce-workflow-step-executor.test.ts    | 149 ++++++++++++++++++++-
 .../engine/src/__tests__/executor-test-helpers.ts  |   1 +
 packages/engine/src/executor.ts                    |  53 ++++++--
 5 files changed, 200 insertions(+), 18 deletions(-)

Fusion-Task-Id: FN-8461

Fusion-Task-Lineage: ef743df4-8bd2-44e6-9498-f6448738d6dc

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-21 19:05:48 -07:00
flexi767
c71a9545b0 fix(engine): isolate provider rate-limit pauses (#2339)
## What changed

- Construct one `UsageLimitPauser` per project runtime and wire it into
both executor and triage.
- Replace the project-wide emergency stop for 429/quota failures with
provider-scoped task parking.
- Resolve execution, planning, validator, and merger providers for
active tasks; park only tasks routed through the unavailable provider.
- Preserve the actual reviewer provider on `ReviewerProviderError`, so a
Claude Plan Review 429 does not stop Codex work.
- Record `provider-rate-limit:<provider>` pause provenance without
storing provider response bodies in pause metadata.
- Run one daemon-owned provider-health monitor that probes only
providers with persisted rate-limit parks.
- Resume exact matching provider parks across every project only after
the existing authenticated usage probe succeeds and all reported
capacity windows are usable.
- Probe at five-minute intervals for the first five checks, then back
off independently per provider to 10/20/40/60 minutes with a one-hour
cap.

## Root cause and impact

The runtime refactor left `usageLimitPauser` undefined for
`TriageProcessor`. In the observed FN-922 incident, Claude Plan Review
returned four explicit 429 responses; Fusion backed off for roughly
60/120/240 seconds and then failed the task, but never invoked its pause
coordinator. The older coordinator also used `globalPause`, which would
terminate healthy sessions on every other provider.

After this change, active tasks using the unavailable provider are
parked while work routed exclusively through healthy providers
continues. Recovery is a provider-health state transition: the daemon
checks Claude/Codex authentication and metered capacity independently of
task execution, including after restart, and clears only exact
`provider-rate-limit:<provider>` parks. Logged-out, errored, exhausted,
manually paused, user-paused, and other-provider tasks remain parked.
Explicit global/engine pause controls remain unchanged.

## Surface enumeration

- executor usage-limit catches
- triage planner and Plan Review catches
- reviewer provider-error propagation
- merger usage-limit catches
- per-project runtime construction and wiring
- task model overrides plus project/global execution, planning,
validator, and merger resolution
- daemon startup/listen and shutdown lifecycle
- multi-project provider-probe deduplication
- Claude and Codex authenticated usage/capacity probes
- done/archived/already-paused task exclusions
- manual, user, generic, and other-provider pause provenance

## Symptom verification

**Original symptom:** Anthropic/Claude 429s retried and failed FN-922
without pausing Claude-routed work; a functioning global pauser would
also have stopped Codex, and provider parks had no positive-health
recovery path.

**Exact reproduction:** Raise `ReviewerProviderError("429
overloaded_error", "usage-limit", { provider: "anthropic" })` during
Plan Review with Anthropic and Codex tasks present, then return
logged-out/error/exhausted and finally healthy Claude usage responses
from the daemon probe.

**Assertion it is gone:** Anthropic-routed active tasks receive
`provider-rate-limit:anthropic`; Codex-only tasks are not paused and
`globalPause` is never changed. Unhealthy probes leave the Anthropic
tasks parked; a positive authenticated response with remaining capacity
resumes only exact Anthropic provider parks without executing a model
call as a probe.

## Validation

- `packages/engine/src/__tests__/usage-limit-detector.test.ts`: 49
passed
- `packages/dashboard/src/__tests__/provider-health-monitor.test.ts`: 8
passed
- Engine TypeScript check passed
- Dashboard server and app TypeScript checks passed
- Scoped ESLint passed
- Changeset strict format check passed
- Reapply script passed `bash -n`, two consecutive fixture applications,
and `node --check`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Tasks paused due to a provider’s rate limits can now automatically
resume when capacity returns.
- Provider health is monitored in the background, including retry
backoff for unavailable providers.

- **Bug Fixes**
- Rate-limit issues now pause only affected provider-routed tasks
instead of stopping unrelated work.
- Provider failures are handled separately from invalid review results,
improving recovery behavior.
- Healthy providers remain available while another provider is
rate-limited.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: v <v@v.speedport.ip>
2026-07-21 17:08:44 -07:00
gsxdsm
de2cad7535 fix(workflows): reject missing plan review artifacts (#2390)
## Summary

Workflows could reach Plan Review without an authoritative PROMPT.md,
producing misleading approvals or stranding the task. Planning now
verifies durable prompt persistence before releasing the card, and every
workflow entry/review surface fails closed when its required plan is
absent. Confirmed absence triggers bounded automatic replanning;
TaskStore read outages retry in place; exhausted recovery parks visibly
without consuming review-fix budget or overriding pause, manual-review,
terminal, or merge-confirmed state.

Related: FN-8455

## Validation

- Focused workflow-artifact, graph-recovery, review, writer, and triage
regression suites pass.
- @fusion/engine typecheck passes.
- Repository lint, changeset validation, and diff checks pass.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Plan Review now fails closed when `PROMPT.md` is missing or blank,
returning a revision request with a typed `failureValue`.
* Required workflow artifacts are treated as missing unless they exist
with non-empty content; read failures are handled separately.
* Recovery now deterministically chooses replan vs “park-failed” with
bounded retries, and records a `task:required-artifact-missing` audit
event.

* **Workflow Improvements**
* Triage and approval now persist `PROMPT.md` through the dedicated
prompt-write flow and verify it was stored exactly.
* Optional-group remediation preserves typed required-artifact missing
failures for pre-merge fixes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-21 17:06:26 -07:00
gsxdsm
dc834e582e fix(workflows): address lifecycle review follow-ups (#2380)
## Summary

- preserve workflow IR hashes in production column-transition audit
metadata
- centralize active workflow-continuation states across release,
runtime, and executor paths
- extract and test actionable planning-continuation selection
- expand Coding (Ideas) remapping/removal coverage and add required
lifecycle decision records

Follow-up to the review body on #2378 after that PR was merged.

## Validation

- `pnpm lint`
- 123 focused core/engine tests
- `pnpm verify:fast`
- `pnpm test:gate` (487 tests)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved workflow continuation handling by centralizing
“active/continuation-eligible” state selection across executor,
hold/release logic, and in-process runtime.
- Persisted richer task column-transition metadata (including `irHash`)
to preserve workflow provenance.
- Ensured planning continuations exclude paused/missing/invalid tasks
and that task resolution failures surface instead of being ignored.
- Corrected fresh-worktree step execution ordering to return expected
`baselineSha`/`checkpointId` behavior.

- **New Features**
- Added and exposed `ACTIVE_WORKFLOW_WORK_ITEM_STATES` for consistent
work-item “active” semantics.
- Introduced a shared planning-continuation candidate selector to
standardize dispatchable planning work filtering.

- **Documentation**
- Clarified the small coding-ideas workflow preset omits verification
while preserving a continuous executable path.

- **Tests**
- Added coverage for planning continuation filtering and fresh-worktree
ordering behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-21 13:36:05 -07:00
gsxdsm
83209e64dc fix(workflows): align stages with board columns (#2378)
## Summary

The Coding (Ideas) workflow now behaves like the board it presents:
Ideas stays inert, Todo owns planning and plan review, In progress owns
implementation, and In review owns code review and merge. The restored
preset is intentionally limited to that five-stage path, while the
existing Coding workflow remains unchanged.

Workflow execution now suspends at Todo→In progress instead of running
the implementation node early. A durable, single-owner continuation
records the exact resume node and survives process restarts; the
scheduler remains the only component allowed to admit the task into WIP.
Disabled optional review groups traverse the same boundary without
invoking a reviewer, avoiding the prior stuck-task behavior.

Workflow validation also rejects capacity holds with no reachable WIP
destination, so deterministic lifecycle deadlocks fail at authoring time
rather than after a task is running.

Session-settled decisions carried from planning: columns are execution
invariants, scheduler-owned WIP admission is preserved, the existing
Coding (Ideas) preset is restored and simplified, and invalid release
topology is rejected (user-approved).

## Validation

- `pnpm lint`
- `pnpm verify:fast`
- `pnpm test:gate` (296 engine, 128 PostgreSQL core, and 63 CI-shape
tests)
- Focused workflow lifecycle tests (106 assertions)
- PostgreSQL regression coverage proves atomic continuation replacement
and database rejection of a second active owner


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added durable, resumable workflow execution across capacity boundaries
(including explicit suspend/resume at the correct node).
* Introduced Todo “plan review” workflow continuations and automated
planning/capacity draining.
* Restored Coding (Ideas) as a selectable built-in and updated its lane
placement; improved optional-step group enablement support.
* **Bug Fixes**
  * User moves back to Todo now cancels active workflow continuations.
* Rejected workflow boundary transitions now surface as errors (instead
of silently continuing).
* Workflows with undriveable capacity-hold configurations are now
rejected.
* **Tests / Data**
* Expanded coverage for workflow suspension, continuations, and
continuation replacement; updated database schema to persist
continuation metadata and enforce single active continuation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-21 12:17:47 -07:00
gsxdsm
4c0dfbcfd6 fix(engine): preserve workflow completion summaries
Keep approved-contract retry instructions scoped to review nodes so advisory and completion-summary agents can produce their intended output.
2026-07-20 15:26:23 -07:00
gsxdsm
1d4e8afa7b FN-8444: include planning time in task metrics
Track active planning time alongside execution time for costs, analytics, and task displays.

- Persist planning timing state across task lifecycle transitions and recovery
- Include planning activity in token cost, analytics, and dashboard timing displays
- Add PostgreSQL migration support using the configured migration directory

Files changed:
 .changeset/fn-8444-planning-time-cost.md           |  7 +++
 docs/dashboard-guide.md                            |  3 ++
 docs/task-management.md                            |  5 ++
 packages/core/src/index.ts                         |  1 +
 .../migrations/0029_planning_active_timing.sql     |  3 ++
 packages/core/src/postgres/schema-applier.ts       | 14 ++++-
 packages/core/src/postgres/schema/project.ts       |  2 +
 packages/core/src/productivity-analytics.ts        | 29 +++++-----
 packages/core/src/store.ts                         |  2 +-
 .../core/src/task-store/archive-lifecycle-2.ts     |  2 +
 packages/core/src/task-store/moves.ts              |  7 +++
 packages/core/src/task-store/persistence.ts        |  4 ++
 packages/core/src/task-store/remaining-ops-2.ts    |  2 +-
 packages/core/src/task-store/serialization.ts      |  7 +++
 packages/core/src/task-store/task-row-mappers.ts   |  2 +-
 packages/core/src/task-store/task-update.ts        | 10 ++++
 packages/core/src/task-timing.ts                   | 35 ++++++++++++
 packages/core/src/types.ts                         | 12 +++++
 packages/dashboard/app/components/TaskCard.tsx     | 13 ++---
 .../app/components/TaskTokenStatsPanel.tsx         |  6 ++-
 .../app/components/__tests__/TaskCard.test.tsx     | 17 ++++++
 .../app/utils/__tests__/taskTiming.test.ts         |  9 +++-
 packages/dashboard/app/utils/taskTiming.ts         | 14 +++++
 packages/dashboard/app/utils/taskTokenCost.ts      |  2 +
 .../dashboard/src/task-planner-chat-metrics.ts     | 14 ++++-
 packages/engine/src/__tests__/self-healing.test.ts | 61 +++++++++++++++++++++
 packages/engine/src/executor.ts                    | 50 +++++++++++++++++
 packages/engine/src/runtimes/in-process-runtime.ts |  3 ++
 packages/engine/src/self-healing.ts                | 62 ++++++++++++++++++++++
 packages/engine/src/triage.ts                      | 10 ++++
 packages/i18n/locales/en/app.json                  |  2 +-
 packages/i18n/locales/es/app.json                  |  2 +-
 packages/i18n/locales/fr/app.json                  |  2 +-
 packages/i18n/locales/ko/app.json                  |  2 +-
 packages/i18n/locales/zh-CN/app.json               |  2 +-
 packages/i18n/locales/zh-TW/app.json               |  2 +-
 36 files changed, 384 insertions(+), 36 deletions(-)

Fusion-Task-Id: FN-8444

Fusion-Task-Lineage: 0178e0a7-3018-4ef4-be9b-6de5f964fb58

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-20 13:40:50 -07:00
gsxdsm
1b8b7f617e fix(FN-8426): wait for answers to agent questions
Convert supported runtime question-tool calls into Fusion's durable awaiting-user-input contract so workflow execution cannot continue while the operator question is unanswered.

Fusion-Task-Id: FN-8426
2026-07-20 13:12:22 -07:00
gsxdsm
71c0d0a970 fix(engine): recover completed triage tasks
Preserve workflow ownership across Plan Review replan moves and route advanced completed triage rows through legal lifecycle transitions before review. Clear only stale same-task session claims after live executor, planner, and merger ownership checks.
2026-07-20 08:46:34 -07:00
gsxdsm
876a278afd fix(engine): prevent workflow boundary restart loops (#2360)
## Summary

Workflow tasks no longer restart or become stranded in Planning when the
graph moves through replan and review boundaries. The executor now
distinguishes its own synchronous column transition from an external
cancellation, while preserving the existing hard-cancel behavior for
user and unrelated engine moves.

Existing advanced tasks left in Planning are recovered from durable
worktree and graph-pin evidence: completed work advances through the
normal review handoff, and incomplete remediation resumes at its pinned
execution column. A shared synchronous reservation keeps Planning and
recovery mutually exclusive, and Planning excludes advanced rows so they
cannot consume capacity in a repeated claim/skip loop.

## Validation

- 238 affected engine tests passed, including graph-boundary
cancellation, planner eligibility, ownership races, and advanced-task
recovery coverage.
- `pnpm --filter @fusion/engine typecheck`
- `pnpm verify:fast` — workspace build, CLI bundle, and real
`/api/health` boot smoke passed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Workflow graph tasks now continue running correctly when crossing
workflow column boundaries.
* Improved recovery of interrupted advanced-triage tasks, including
completed and in-progress work.
* Prevented duplicate triage dispatches and protected tasks from
competing recovery and planning actions.
* Added safeguards for task state changes during recovery and
maintenance operations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-20 00:18:06 -07:00