Commit Graph

3612 Commits

Author SHA1 Message Date
Phil Larson
920d68e10f fix(dashboard): expose column roles to browser bundle (#3151)
## Summary
- export the browser-safe `@fusion/core/column-roles` subpath
- keep Vite/Vitest aliases ahead of broad `@fusion/core` aliases
- restore production dashboard builds after task undo classification
adopted shared column-role helpers

## Test plan
- `node scripts/check-no-node-only-core-imports-in-dashboard.mjs`
- `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest
run app/utils/__tests__/taskRevert.test.ts --pool=threads
--maxWorkers=1`
- `pnpm --filter @fusion/core typecheck`
- `pnpm --filter @fusion/dashboard typecheck`
- `CI=true pnpm check:changesets`
- `pnpm build`


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dashboard build compatibility for browser-based environments.
* Improved reliability when importing column role functionality across
supported application components.

* **Refactor**
* Made column role utilities available through a dedicated browser-safe
entry point.

* **Chores**
* Updated development and test configurations to consistently resolve
the new entry point.
* Documented the browser-safe module classification and recorded the
release patch.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 06:46:19 -07:00
gsxdsm
ada62a7c4a census: --claims shows which remaining files an open PR already holds (two duplicate claims today) (#3124)
The census says **where** the work is but not **who has it**, and
duplicate claims are now the dominant coordination cost of this phase.
This adds an opt-in `--claims` report mapping each remaining file to the
open PRs already touching it.

## The problem is measured, not suspected

- **`self-healing.ts` took three overlapping conversions** from
different lanes while one branch was open (#3049, #3075, #3078). Each
forced a full rebuild of #3094, and every conflict was the same shape:
*same guard, two spellings, different variable names*. That PR's body
asks, in as many words, for one lane to own the file.
- **`executor.ts` took two independent conversions today** — #3112 and
#3118 — same four literals, same payload-lanes fix, two branches. Two
workers each read the census, saw the top cluster, and started. Neither
could see the other; I only caught it because both appeared in one `gh
pr list`.

The census is what sends everyone to the same file, so the claim signal
belongs here rather than in a side channel nobody reads. `--triage`
(#3097) already measured the underlying fact — 53 of 88 guards sat
inside an open PR — one step short of being actionable.

## Measured on current main (29 guards)

```
  CLAIMED by an open PR: 6 files holding 15 guards
       6  packages/engine/src/self-healing.ts  ← #3121 #3116
       4  packages/engine/src/executor.ts  ← #3118 #3112
       2  packages/engine/src/auto-merge-finalization.ts  ← #3107
       1  packages/core/src/task-store/task-artifacts-ops.ts  ← #3120 #3119 #3091
       …
  UNCLAIMED: 12 files holding 14 guards — start here
       2  packages/dashboard/app/utils/taskRevert.ts
       2  packages/engine/src/scheduler.ts
       …
```

It independently reproduces **both** collisions I found by hand today,
which is the strongest evidence I can offer that it works: `executor.ts
← #3118 #3112` and `self-healing.ts ← #3121 #3116`.

It also answers the standing fleet instruction empirically. "Claim the
largest unclaimed cluster" currently resolves to **12 files holding 14
guards, none larger than 2** — and one of those two (`scheduler.ts`) is
in the SYNC-RESOLVED list, where conversion is inert. That is a
materially different picture from the headline `29`.

## Design decisions

**Report-only and fail-soft**, on the same terms as `--triage`: opt-in,
printed beside the totals, changes no count and no exit code. It shells
to `gh`, so it is unavailable offline, in CI without a token, and in
sandboxes — all of which print a notice and continue. A gate must not
depend on network state; this is a work-selection aid, not a gate.

**The fail-soft path is loud on purpose**, and it is the case I care
most about. A claim report that silently degrades to "nothing is
claimed" is *worse than no report*, because it actively sends the reader
into work another lane holds — the exact failure the flag exists to
prevent. So when `gh` cannot answer it prints `POSSIBLY CLAIMED` and
suppresses the start-here list entirely rather than rendering it empty.

**Heuristic, and says so.** A PR touching a file is not proof it
converts *that file's* guards — it may edit an unrelated function. It
over-reports rather than misses, which is the safe direction: a false
claim costs one comment asking, a missed one costs a rebuilt branch.

**One bulk `gh pr list` call**, not a request per PR — the per-PR shape
was too slow to become habitual, and a report nobody runs is not a fix.

## Verification

- `lifecycle-column-census.test.ts` — **42 passed** (was 40)
- Differential: disabling the flag gives **2 failed | 40 passed**. Both
new tests fail on the defect they were written for.
- `--strict` and `check-fnxc-future-dates` — exit 0
- Tests stub `gh` on PATH, so no network call and no dependency on the
live PR list. The fixture reads the census's **own current top file**
rather than a hardcoded path, so it cannot rot as the backlog shrinks
(same self-maintaining discipline as #3106).

## What this does not do

It does not reserve anything — there is no lock, and two workers who
both run it can still collide if they start simultaneously. It reports
what is already visible in the PR list, which is enough to catch the
every-case-so-far pattern of *starting work on a file someone has held
for hours*. A real reservation would need shared mutable state, and I
would not add that without an owner asking for it.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:50:13 -07:00
gsxdsm
d09a856941 test(engine): pin the FN-5256 liveness guard on a renamed board (the sweep that clears a live task's worktree) (#3132)
Top item from the verified coverage map on #3115.
`reconcileTaskWorktreeMetadata` had **three** uncovered resolvers — the
most of any sweep in the file — and it is the one that nulls
`worktree`/`branch`/`sessionFile` on a live row.

## Why this sweep first

Its own header names FN-5256: the incident where clearing worktree
metadata yanked a checkout out from under a running shell. The guard
that prevents it is `scopeOverrideMergeActiveSafe`, and that guard is
exactly what the wip/review resolvers feed.

The existing guard test uses `column: "in-progress"` — **the literal**.
So blinding `worktreeReconcileWipColumns` back to `["in-progress"]`
leaves all 825 self-healing tests green. The guard is converted; nothing
in the suite could tell.

On a renamed board the pre-conversion form matched nothing,
`scopeOverrideMergeActiveSafe` became true for a card an executor was
actively running, and the sweep cleared its metadata.

## The case

The renamed twin of the existing FN-5256 test: a `scopeOverride` task
live in a **renamed wip lane** keeps its metadata. Same shape, same
assertions, different vocabulary — which is the whole point, since the
original passes either way.

**Measured:** 414 pass; blinding `worktreeReconcileWipColumns` to the
legacy id fails **exactly this test**.

## Remaining from the map

25 uncovered resolvers left. Next by risk: `reclaimStaleActiveBranches`
(deletes branches) and `reconcileInReviewBranchRebind` (rebinds branches
of live cards) — both need a git-shelling harness, so they are slower to
pin than this one was. Then the two `reconcileDependencyBlockingLeases`
resolvers.

I will keep working down that list. The map is on #3115 with verified
names; anyone can pick an entry and check it the same way — blind one
resolver, run `vitest run src/__tests__/self-healing`, and if it stays
green that conversion has nothing behind it.

## Verification

`self-healing.test.ts` **414 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.
2026-07-31 05:49:59 -07:00
gsxdsm
ce84aa48d0 test(self-healing): cover the renamed-board starved-refinement wake that main's conversion lacked (#3116)
**Rebased onto current `main`, and it shrank to one test.** Was
"self-healing consolidated (45 → 39)".

## What happened

**Every code change in this PR landed independently from other workers**
while it was open, and in each case theirs is equal or better. I took
theirs and dropped mine:

| My change | Landed on `main` as |
|---|---|
| pre-execution worktree seizure | `preExecLiveColumns` — same
"dangerous direction" reasoning |
| FN-5256 liveness cluster | `worktreeReconcileWipColumns` /
`worktreeReconcileReviewColumns` |
| agent-link membership | `agentLinkLiveColumns` /
`agentLinkTerminalColumns` |
| starved-refinement peer progress | `starvedWaitingColumns` — a project
union covering both duplicated sites |

Resolving the rebase by taking `main` left two orphaned declarations
(`activeOrQueuedColumns`, `holdPeerIds`) that nothing referenced. `tsc`
doesn't flag unused locals here, so I checked references by hand and
removed them rather than ship dead code that reads as converted.

## What's worth landing

**Their starved-refinement conversion has no renamed-board test — the
suite had zero.** This adds one.

A candidate resting in a renamed **intake** lane, with its peers in a
renamed **hold** lane, must still escalate. The two are deliberately
distinct columns so a wrong role set resolves no peers and escalates
nothing; a fixture where they coincide would pass either way.

The fake needed `listWorkflowDefinitions` — `starvedWaitingColumns` is a
**project union**, so per-task selection readers alone leave it
resolving nothing and the test would pass for the wrong reason. That
mismatch is how I found the gap: my original test failed against their
implementation.

## Verification

- Green against **their** code
- **Revert-proof against theirs:** restoring the literal fails it — 0
escalations against 1 expected
- 8 tests in the suite green

## Note for the fleet

This is the second PR of mine to shrink to a test on rebase (#3096 was
the first). Both times the duplicated work was real and mine was the
later arrival. The pattern is worth acting on at the coordination level,
not by me working faster.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:41:32 -07:00
gsxdsm
6483f9ce2b fix(scheduler): resolve task:updated / task:deleted lanes asynchronously (scheduler inert 5 → 0) (#3128)
The last inert guards in `scheduler.ts`. Independent of my other
branches.

## Inert-guard ratchet

| Scope | Before | After |
|---|---:|---:|
| `scheduler.ts` | 5 | **0** |
| total | 12 | **7** (triage.ts 8 → other worker; executor.ts 4 → #3112)
|

## The live bug

These read `resolveTaskParkedColumnsSync`, which answers with the
**default** workflow in production. On a renamed board the scheduler
**never woke** on unpause or planning-finish, and a **deleted blocker
never unblocked its dependents** — the card sat behind a task that no
longer existed.

## The criterion, restated because I got it wrong before

**What blocks a guard is whether its answer is consumed synchronously —
not whether the enclosing listener is declared sync.** I assumed the
latter earlier in this program and reverted for it.

All three fail that test: two only gate `schedule()`, which is itself
`async`, fire-and-forget and re-entrance-guarded; the third already sits
below an `await getSettings()`. The edge-trigger bookkeeping
(`planningTaskIds.delete`) **stays synchronous** on purpose — deferring
*that* would let a second update re-enter the branch.

## The union is load-bearing, not defensive

Post-U11 the default lineage has no `triage` column, so a **resolved**
answer returns `intake: "todo"` where the inert path fell back to
`"triage"`. Converting without unioning the legacy ids silently
**narrowed** the wake set and stopped waking cards in a legacy-named
lane — caught by *"schedules when planning clears in triage"*.

**A resolved conversion must be a superset of what it replaces, or it is
a behaviour change wearing a vocabulary change's clothes.** That's the
reusable lesson here.

## Tests

- Drained with the repo's existing **`flushAsyncHandlers`** helper —
written for exactly this fire-and-forget shape — rather than loosening
any assertion.
- **The characterization test flipped, as designed.**
`workflow-scheduler-parked-columns-live-e2e.pg.test.ts` asserted *"a
dependent in a RENAMED hold column is NEVER unblocked"*, with its author
noting: *"expected to flip to null the moment the resolver is fixed —
and that flip is the whole point of writing it down."* It flipped.
Inverted to a REGRESSION case so the assertion holds the fix rather than
the defect; it now matches its own CONTROL arm, which still guards
against a vacuous pass.

## Verification

- 21 scheduler suites — **361 green**, including the live PostgreSQL e2e
- **`pnpm test:gate` green**; eslint and `tsc` clean
- Changeset added; `check:changesets` passes

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:41:16 -07:00
gsxdsm
ad5172afd5 fix(engine): main is red on check:inert-sync-lanes — #3114's triage conversion is inert, revert the arm (#3126)
## `main` is red on `check:inert-sync-lanes` right now

```
inert-sync-lane: NEW inert conversions — a lane guard now reads a sync resolver
that always answers with the DEFAULT board.
  packages/engine/src/triage.ts: 7 -> 8
```

Verified on a clean `origin/main` checkout, not on my branch. #3114
converted this guard's third arm to `disposeLanes.wip`; the gate that
exists to catch exactly this fired, and the PR landed anyway —
presumably because `check:inert-sync-lanes` is not in the blocking
merge-gate set.

## The change did not change behaviour

`disposeLanes` comes from `resolvePlannerLanes`, which resolves through
`resolveTaskWorkflowIrSync` — inert under PostgreSQL for two independent
reasons (#3103). So `disposeLanes.wip` evaluates to `in-progress`: **the
same value as the literal it replaced.**

A card advancing into a renamed execution lane still matches nothing,
still reads as an evacuation, and still kills a healthy planning session
— the precise bug #3114 set out to fix, unchanged on every board.

So the arm goes back to the literal. The gate's own failure text rules
out the alternative:

> Do NOT re-record the baseline to clear this — that is the same false
green one layer up.

## #3114's analysis is kept — only the code reverts

Its behavioural description is **correct** and is the clearest statement
of this bug anywhere in the file. I have kept those paragraphs and added
what is missing: that the fix does not reach under PG, and what would.

Whoever supplies a lane answer that is not sync-resolved should make
this line read `disposeLanes.wip` and delete the note. The specification
is sitting right there for them.

## It also reconciles two contradictory notes, one of them mine

My #3108 flag said converting the third arm this way adds an inert
comparison and removes a census entry that is telling the truth. #3114
then converted it and added a note saying it fixes the bug. **Both notes
sat in the file**, giving any reader two confident, opposite accounts.
They are now one account with the evidence attached.

## Read this file's census count carefully

#3114 took it to **0** while the inert count went to **8**. The census's
own `--triage` output warns about exactly this shape:

> for a sync-resolved file, a count of 0 is the WORST case, not the best
— the file reads as fully converted

Reverting restores it to 1, which is the honest signal.

## Census

| | before | after |
|---|---|---|
| `triage.ts` | 0 | **1** |
| repo backlog | 26 | **27** |

**The number going up is the point.** A census that reports 0 for a file
whose guards are all inert is worse than one that reports the truth — it
retires the entry and nobody looks again.

## Measured

- `check-inert-sync-lane-conversions`: **exits 1 on `main`, 0 here** (8
→ 7).
- `src/__tests__/triage*` — **25 files / 374 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-fnxc-future-dates` clean.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:35:28 -07:00
gsxdsm
4b61170a51 fix(executor): read task:moved lanes from the payload (executor.ts 4 → 0) (#3112)
**Stacked on #3109** — merge that first; this is its first consumer.

## Census

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 47 | **43** |
| `executor.ts` | 4 | **0** |

`executor.ts` is off the census top-files list.

## Why these four could not be converted in place

This listener is synchronous and its branches **start execution**,
dispose worktrees and release sessions. An await ahead of them defers
the `execute()` dispatch itself. The sync IR resolver isn't an option
either — it answers with the default workflow under PostgreSQL, so a
guard written through it is inert.

Reading the lanes the emitter already resolved costs nothing and leaves
the prologue synchronous. This listener is the reason #3109 has the
shape it does.

## The archive branch is the one with teeth

`to === "archived"` matched nothing on a board with a renamed terminal
lane, so **archiving never released the task's active-session registry
entry** — and that entry is what blocks a **successor** task from
acquiring the same path. Not cosmetic: the next task wanting that path
fails to register.

## Verification

- **Revert-proof:** the new case drives a `shipped` terminal lane
(matching no legacy id) and asserts the release. Reverting the branch to
the literal leaves the entry held — `expected [Array(1)] to have a
length of 0`.
- 43 executor suites — **483 green**
- **`pnpm test:gate` green**; eslint clean

## Note on shape

Lanes are read as **single ids, not sets**, because each branch here is
a lane-identity test on one column — exactly what the literals were.
Widening to membership would change behaviour, not just vocabulary.
Fail-soft to the legacy ids when the emit path could not resolve,
matching every other consumer of this payload.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:15:38 -07:00
gsxdsm
218086bea2 fleet(engine): self-healing 6 → 1 — the board-stall counter, the last guard that needed a sync answer (#3121)
The last fan-out guard, and the one I explicitly said needed a
synchronous answer. #3109 made that answer available without an await,
so the flag comes off.

## Why this one was last

The other two guards in this listener gated work the listener **already
`void`s**, so they moved onto the async resolver in #3094. This one
increments in-memory state **in the handler's own tick**, so it
genuinely needed a synchronous answer.

The sync IR path was never that answer: `resolveTaskWorkflowIrSync`
cannot resolve a **custom** workflow at all — two independent blockers,
#3103 — which is why I wrote that conversion, measured it, and withdrew
it.

#3109's emitter-carried `lanes` removes the dilemma rather than trading
one horn for the other: reading them needs **no await**, so the
increment stays in the same tick *and* the guard becomes correct.

## What it fixes

On a renamed board this counter read **zero**. The board-stall watchdog
was blind to a board whose cards were moving out of implementation the
whole time — the signal it exists to raise was never raised.

## Census

| | before | after |
|---|---|---|
| `self-healing.ts` | 6 | **1** |
| repo backlog | 29 | **24** |

The remaining 1 is the log-dedup closure — a pre-existing flag whose
degraded answer costs a duplicate log line, not a lifecycle decision.

## Measured

- 3 new cases; `self-healing-completion-fanout.test.ts` **13/13 pass**.
- **MUTATION**: restoring the literal pair fails the renamed case.
- **The paired negative is the load-bearing one.** The guard means
*"left implementation for somewhere that is not implementation"*, so a
move **between two non-wip lanes** must not count. Without that case, a
conversion that counted every move would pass the positive and inflate
the watchdog's denominator — breaking it in the opposite direction,
which is harder to notice than a zero.
- A **fail-soft** case pins that an emit carrying no `lanes` still
counts on the legacy ids.
- **Asserted through the counter itself**, not a downstream alert. The
increment *is* what this guard decides; routing the assertion through
the watchdog would let an unrelated threshold change mask a regression
here.
- `src/__tests__/self-healing*` + `task-agent*` — **42 files / 848 tests
pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## On the withdrawal this reverses

#3094 withdrew a sync-IR conversion of this listener and recorded why,
precisely. That record is what made this cheap: I could tell in one read
that #3109 addressed the *specific* obstacle rather than a general
"async is hard". A flag that names its blocker exactly is a flag that
can be retired the day the blocker goes.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:15:14 -07:00
gsxdsm
56b5cdfeed test(notifications): cover the wedge-episode renamed-lane clear that main's conversion lacked (#3096)
**Rebased onto current `main`, and it shrank to a test.** Was "close the
wedge-episode race, then resolve its lanes."

## What happened

Another worker landed **both halves of this PR independently** while it
was open. Rebasing showed their versions are better, so I took theirs
and dropped mine:

- **The serialisation** — theirs is `enqueueWedgeHandling(taskId, run)`,
a general callback; mine was wedge-specific.
- **The conversion** — theirs is **project-union membership** over the
four roles; mine was first-match-per-role via `resolveLifecycleColumns`.
Membership is correct: more than one lane can fill a role on a renamed
board, and first-match silently ignores the rest.

My rebased branch initially compiled to a **duplicate
`wedgeHandlingChains` field and duplicate method** — caught by `tsc`,
removed. Nothing of my implementation survives, and it shouldn't.

## What's left is worth landing

Their conversion has **no renamed-board test**. This adds one.

A card recovering into a renamed hold lane must **clear** its episode.
Asserted through the *second* notification, because a stale active
episode also **refuses the next genuine wedge its claim** — so the
visible symptom is a real wedge going unannounced, not merely a stale
alert.

The fixture needed `listWorkflowDefinitions`:
`resolveProjectColumnsForRoles` unions across the project's workflows,
so the per-task selection readers alone leave it resolving nothing and
the test would pass for the wrong reason. That's how I found the
mismatch — my original test failed against their implementation.

## Verification

- Green as written against **their** implementation
- **Revert-proof against theirs:** restoring the four literals fails it
— 1 delivered, 2 expected
- 7 notification suites — **80 green**; `tsc` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed wedge notifications so they can trigger again after a task
recovers into a renamed workflow’s hold lane.

* **Tests**
* Added regression coverage confirming that recovered tasks correctly
clear their wedge state and support subsequent notifications.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:09:36 -07:00
gsxdsm
6050d6eb83 chore(engine): mark the auto-merge-finalization reviewed literals DELIBERATE (census 47→45) (#3107)
Fleet phase. `packages/engine/src/auto-merge-finalization.ts` was the
last census file with no branch, worktree, or open PR against it. Claim
published by pushing the branch before starting.

## Census before / after

| | total | this file |
|---|---|---|
| before | **47** | 2 |
| after | **45** | 0 |

`--strict` exits 0, baseline re-recorded. **Reclassification, not
conversion** — both lines are unchanged.

## Both sites were already reasoned, in a note that calls them
non-defects

- **Line 30** is the resolver's **degraded fallback arm**, inside
`catch`. The live arm two lines up calls `columnHasFlag(ir, columnId,
"complete")`. The literal is reached only when IR resolution throws,
where the legacy id is the only answer left — removing it would make a
failed resolve return nothing.
- **Line 99** picks an **error string**. The note above it works through
threading `isCompleteColumn` in and concludes the signature widening
costs more than the sharper diagnostic buys.

I did not revisit either judgement. The gap was mechanical: prose the
census cannot read, so both stayed in `byFile` as apparent debt for the
next pass to re-derive.

## This is the fourth, and it closes the set

With #3056, #3060, and #3063, **every census file that was unclaimed
during this phase has now been examined, and not one needed a
conversion.** Each site was a three-state fallback arm, or a site a
prior pass had already reviewed and kept.

The corollary is the finding I would most want carried forward: the
remaining count is not a work queue. A worker told to "claim the largest
cluster" reads the number, finds most of it already reasoned, and
reaches for whatever moves it — which is how three PRs converted guards
to a synchronous resolver that is inert under PostgreSQL.

One exception worth preserving: **`taskRevert.ts` should stay counted.**
I claimed, inspected, and released it without marking. Converting it
would classify a *neighbour* row using the modal task's flags — wrong on
data, not merely stale on vocabulary — and its note correctly calls the
entry **accurate debt** blocked on a per-neighbour flag map. Fallback
arms and dead paths → mark. Placeholders awaiting a capability → leave
counted.

## Verification

- `census --strict` exit 0; `tsc --noEmit` (engine) **0 errors**
- No dedicated test file for this module (`vitest` reports none), so no
suite to run — comment-only diff, no behaviour change

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:09:23 -07:00
gsxdsm
89b21e2906 fleet: triage's planning-evacuation check uses the resolved wip lane (census 45 → 44) (#3114)
## Census

| | column guards |
|---|---|
| before | **45** |
| after | **44** |

## What changed

```ts
if (task.column === disposeLanes.hold || task.column === disposeLanes.intake
    || task.column === "in-progress") return;
```

Two role questions and one id question on the same line.
`resolvePlannerLanes` is **already called immediately above**, and its
result carries `wip` — so this needs no new resolution and no new await.
The literal just stops being the odd one out among its neighbours.

## What it cost on a renamed board

This handler aborts a planning session when a card leaves the planner
lanes. `in-progress` is excluded because *a card advancing into
execution is not an evacuation* — that's stated in the note directly
above it.

Against the literal, that exclusion **never matched** on a board whose
execution lane is renamed. So a legitimate advance into execution read
as an evacuation and **killed a healthy planning session** — precisely
the case the comment says must not abort.

`wip` is optional by design (PR #2628: a missing role stays `undefined`
so callers refuse rather than invent a column). Undefined here means the
board declares no execution lane, so there's no advance-into-execution
to exclude and the comparison is correctly false.

## Not addressed, and pre-existing

This line resolves through `resolvePlannerLanes` — the **sync** twin,
which returns the default workflow's lanes under PostgreSQL. That
affects all three lanes on the line equally and predates this change:
the handler is `(task: Task) => {}` with no await available, so fixing
it needs the same emitter-side change as #3082.

Making the third lane consistent with the other two doesn't deepen that,
and it leaves **one** shape to fix there rather than two.

## Measured

| check | result |
|---|---|
| triage / evacuation / planner-lane suites | **453 tests green** |
| four gates + strict census | green |
| engine `tsc` | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 05:06:08 -07:00
gsxdsm
6eeeb43d4b test(engine): pin #3047's archive-sweep conversion — measured uncovered (825 tests passed against the reverted fix) (#3115)
Not a conversion — the fleet's conversions are landing faster than their
coverage, and this is the audit that shows which ones actually have any.

## Method

For each of today's fleet commits to `self-healing.ts`: revert that
single commit, re-run the file's suites, see whether anything fails. If
nothing fails, the conversion has no regression protection and the "N
tests passed" cited on its PR was measuring something else.

| commit | reverted | verdict |
|---|---|---|
| #3075 pause-abort recovery | suites **fail** | covered |
| **#3047 archiveStaleDoneTasks** | **825 tests all pass** |
**uncovered** |
| #3078 (mine) | 204 tests all passed | was uncovered — closed by #3090,
#3102 |

Every fixture in the `archiveStaleDoneTasks` describe block uses the id
`done`, where the literal is correct, so none of them could see the
conversion at all.

## What the literal cost

The sweep's dependent scan skips tasks in terminal lanes. On a renamed
board **nothing matched `done`/`archived`, so every task read as
active** — which means every archive candidate looked like it had active
dependents, and the sweep archived **nothing**. The board quietly stops
auto-archiving: no error, no log line, no failing test.

## The case covers both halves of #3047

- a stale card in a **renamed complete lane** is archived — the
`complete` role
- a card whose dependent is still live in the **renamed wip lane** is
**not** archived, and a dependent already in the **renamed archive
lane** does not count as live — the `terminal` role

That second assertion is the one that matters: it stops the fix from
degenerating into "archive everything", which is the failure mode a
one-sided test would miss.

## Measured

**413 pass** on current main; reverting #3047 fails **exactly this
test**.

## Remaining audit

I have now audited 3 of ~13 fleet conversions to this file this way. The
method is cheap (one revert, one 17s suite run) and I will keep working
through the rest unless someone else picks it up. #3049 could not be
auto-reverted — later commits overlap its hunks — so it needs a manual
read rather than a mechanical revert.

## Verification

`self-healing.test.ts` **413 passed** · `pnpm test:gate` 13 + 161 + 487
+ 71 · lint — green.
2026-07-31 04:59:44 -07:00
gsxdsm
ec2921b958 fleet(engine): self-healing 23 → 6 — async-reachable guards, plus 3 of 4 fan-out guards the sync path could not serve (#3094)
**Replaces #3093, which I am closing.** Third rebuild of this work.

## A coordination note first, because it is costing more than the code

`self-healing.ts` has had **three** overlapping conversions land from
other lanes while my branch was open — #3049, #3075, #3078. Every time,
replaying my commits produced conflicts that were all the same shape:
*same guard, two spellings, different variable names*. Each rebuild is a
full cycle spent on merge mechanics rather than on lanes.

I have rebuilt against `main`'s own census each time rather than argue
about whose spelling wins, and this PR contains only what `main` (23)
does not have. But if this file is going to keep receiving concurrent
fleet passes, one lane should own it — otherwise the next PR pays the
same tax again.

## Converted

| site | note |
|---|---|
| `isPhantomExecutorBinding` | caller resolved
`lanesOfReclaim(task.id).wip` **three lines above the call**, then
passed a task whose column the predicate compared against `in-progress`
|
| `isWorkspaceOwnerLive` | required `completeColumns` |
| `recoverPausedAbortFailures` **body** | #3075 converted this sweep's
*router* and left three body guards comparing ids |
| `reconcilePreExecutionWorktrees` | a four-id literal in a sweep that
**removes worktrees** |
| `recoverStarvedRefinementTriageTasks` | 2 peer counts that read zero,
so escalation never fired |
| `evaluateParkedAgentTaskLink` | the omitted `parkedColumns` argument |

All through the **async** `resolveProjectColumnsForRoles`, whose only
store read is `listWorkflowDefinitions()` — answerable under PostgreSQL.
That is what separates these from the inert kind.

**The half-converted sweep is the important one.** A router that
resolves correctly feeding a body that compares ids is worse than
converting neither: the route now fires on a renamed board and the body
then acts on the wrong lane. The `moveTask` **target** is the sharp end
— an undeclared target is rejected *except* under `recoveryRehome` with
a legacy id (`moves.ts:570`, the #1411 escape hatch), so a converted
route feeding the literal `"todo"` rehomes the card into a column its
workflow does not declare, which is the state other reconcilers exist to
repair.

**A real dropped-behaviour bug**: `evaluateParkedAgentTaskLink` was
called without `parkedColumns`, falling back to `LEGACY_PARKED_COLUMNS`.
A live durable agent linked to a card resting in a renamed hold lane
read as not-parked, so the safeguard preserving its task link never
applied.

## Withdrawn: the `task:moved` fan-out

I wrote the sync-IR conversion, measured it, removed it.
`getTaskWorkflowSelectionImpl` returns `undefined` **unconditionally**
under PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR and `columnsWithFlag` on it yields exactly the legacy
ids — inert on every board.

Worse than the literal, because **the literal is counted**. My own test
passed only because its store mock supplied a renamed IR: it pinned the
helper's shape, not production behaviour. The refutation is recorded in
place, and `check-inert-sync-lane-conversions` exits 0 on this branch.

## One question, one answer

An earlier pass of this work mapped the notification-attach guard onto a
wider `activeWork` set, and `self-healing-paused-abort-recovery >
"rehomes an in-progress pause-abort park back to todo"` caught it — an
in-progress park attached a transition notification it should not have.

The fix is not a narrower set. Both guards ask **one** question — *"is
the card already at the requeue target?"* — which the literal happened
to spell twice as `=== "todo"`. The target now resolves once, before the
write, and both read it. Deriving one question two ways is exactly how a
converted guard and an unconverted target drift apart.

## Census

| | before | after |
|---|---|---|
| `self-healing.ts` | 23 | **11** |
| repo backlog | 53 | **41** |

## Measured

- `src/__tests__/self-healing*` + `task-agent*` — **42 files / 836 tests
pass**
- `tsc --noEmit -p packages/engine` clean;
**`check-inert-sync-lane-conversions` exits 0**; census `--strict`,
`check-lane-wiring`, `check-fnxc-future-dates` clean

## The remaining 11, flagged not guessed

- **4** — the fan-out, withdrawn above; blocked on a sync-capable
selection reader.
- **1** — the log-dedup closure: pre-existing flag; it sits before the
lane prefetch it needs, and the degraded answer costs a duplicate log
line, not a lifecycle decision.
- **1** — the synthetic `{ column: "todo" }` for a *missing* task:
deliberate, and now correct rather than unconverted, because
`parkedColumns` is legacy-seeded.
- The rest are status/deliberate classifications the census counts but
that are not lane guards.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved workflow automation for boards using renamed or customized
workflow lanes.
* Fixed task completion fan-out, branch rebinding, recovery, and
stalled-task detection across custom lifecycle columns.
* Prevented completion actions from triggering when tasks move back to
the work-in-progress lane.
* Improved cleanup and pause recovery behavior for customized workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:56:49 -07:00
gsxdsm
41cdcc741e fix(events): carry resolved lanes on task:moved so listener guards stop being inert (#3109)
Removes the **inert-guard class at its source** instead of one call site
at a time. Independent of my other branches.

## The problem

`task:moved` listeners run synchronously, so a listener needing a lane
answer had to resolve one synchronously — and
`resolveTaskWorkflowIrSync` returns the **default** workflow under
PostgreSQL, the shipped backend. Every such guard behaved exactly as the
literal it replaced, while the census scored it as converted.

**Resolving asynchronously inside the listener is not available**, and
that is measured rather than assumed. The scheduler's
`snapshotManager.invalidate` is asserted to run in the listener's
**synchronous prologue**; putting an await ahead of it produced **3
failures across 21 scheduler suites**.

## The fix

The emitter carries the answer, which removes the dilemma rather than
trading one horn for the other. `moves.ts` is already async and already
post-commit, so it resolves the moving task's lanes **once** and hands
them to every listener. The guard becomes correct **and** the prologue
stays synchronous.

This is the file's own recorded preferred fix — *"having the emitter
carry the resolved lanes on the event payload so no listener resolves at
all"* — now that the audit it was waiting on is done and came back as
**one** prologue-dependent consumer, not a class.

## Design choices

- **`lanes` is optional and fail-soft to `undefined`** — "unknown",
never "legacy". Some emit paths fire from sync contexts or a cached row
mid-teardown. Listeners keep their existing fallback, so those paths are
no better than before but **no worse**, and they become the exception
rather than the rule.
- **`mergeParkedColumns` overlays only fields the emitter actually
resolved**, so a partial payload cannot blank a lane back to a wrong
answer.
- **The sync resolver stays** as that fallback. Deleting it would strand
the emit paths that cannot resolve.

## Verification

- **Revert-proof and it pins the prologue:** the new case asserts
invalidation on a **renamed** hold lane with **no `waitFor`**. Ignoring
the payload gives **0 calls**.
- 21 scheduler suites — **361 green**
- self-healing + notification suites — **491 green**
- core moves + the `sync-workflow-ir-callsite-allowlist` ratchet — green
- **`pnpm test:gate` green** (71)
- Changeset added; `check:changesets` passes

## What it unblocks

`scheduler.ts`'s 10 allow-listed guards now resolve correctly for every
move that goes through `moves.ts` — the path real moves take. Those were
already absent from the backlog, so **the census number does not move**;
what changes is that they now do what the number claimed.
`executor.ts`'s 4 remaining sites can follow the same pattern in a
separate PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:47:09 -07:00
gsxdsm
15a664a8f5 docs(engine): flag executor's four task:moved literals — the obvious conversion is provably inert (#3104)
The largest unclaimed census cluster. **Nothing in this file said why
the sync-lane pass skipped it**, and that silence is the hazard: the
obvious next move is to convert these the way `scheduler.ts`'s ten were
converted, which would make them **inert rather than fixed**.

## The literals are genuinely wrong — this is not a "non-issue" flag

All four sit in one synchronous `task:moved` listener, and on a renamed
board:

- execution **never starts** on a move into the board's own wip lane;
- terminal session release **never runs** on a move into its archive
lane;
- both `from` guards never fire, so **in-flight work is not aborted**
when a card leaves implementation.

Nothing errors. The engine simply stops reacting.

## Why the obvious fix is inert — proved, not argued

`task:moved` is emitted synchronously, so an `await` here reorders this
handler against every other subscriber. That points at the sync IR path,
which cannot answer for a renamed board for **two independent reasons**
(`sync-workflow-ir-second-blocker.test.ts`, #3103):

1. `getTaskWorkflowSelectionImpl` returns `undefined` unconditionally
under PostgreSQL, so `resolveTaskWorkflowIrSync` always takes its
`!workflowId` branch.
2. Even **with** a selection, the custom-workflow branch loads its IR
through `store.db`, whose implementation is an **unconditional throw** —
so it falls into the catch and returns the default IR anyway.

**A renamed lane is a custom workflow, so (2) alone is decisive.** The
sync path can never serve this listener's case, whatever the selection
reader is fixed to do. That is the part the existing notes across this
repo miss, and it is why flagging beats attempting here.

`check-inert-sync-lane-conversions` already baselines **twenty** guards
in exactly that state in `scheduler.ts`. These four must not join them.

## Census

**Unchanged at 4, deliberately.**

Marking them DELIBERATE-LITERAL would buy a smaller number by asserting
the code is *fine*. It is not fine — it is *blocked*. Those are
different claims with different expiries, and the census should keep
pointing here until the block is lifted. An unconverted literal is
visible; an inert conversion leaves the backlog and takes the evidence
with it.

## Measured

- Comment-only change.
- `src/__tests__/executor*` — **84 files / 853 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` clean.

## Unblocking, for whoever takes it

Either an async listener contract — a behaviour change to handler
ordering, not a column conversion — or a sync reader that answers for
**custom** workflows *and* survives a writer on another node. All three
constraints are written up in `sync-workflow-ir-second-blocker.test.ts`
(#3103).

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:40:42 -07:00
gsxdsm
f5926d3b54 docs(engine): flag triage's evacuation guard — it looks two-thirds converted and is fully literal (#3108)
Completes the sync-listener audit across the three files holding the
remaining blocked guards — `executor.ts` (#3104), `scheduler.ts`
(#3100), and this one. Triage is the most misleading of the three.

## The shape lies

The guard reads as **two resolved arms and one literal**:

> `task.column === disposeLanes.hold || task.column ===
disposeLanes.intake || task.column === "in-progress"`

So the obvious next move is to convert the third arm with the same
helper. That is wrong twice:

**1. The two "resolved" arms are not resolved.** `resolvePlannerLanes`
goes through `resolveTaskWorkflowIrSync`, which cannot answer for a
**custom** workflow — the sync selection reader returns `undefined`
unconditionally, *and* the custom-workflow IR read goes through
`store.db`, whose implementation is an unconditional throw (#3103). So
`disposeLanes.hold` / `.intake` are `todo` / `triage` on every board.
**All three arms are literal in effect.** Converting the third the same
way adds a third inert comparison and retires a census entry that is
currently telling the truth.

**2. The guard's answer is consumed synchronously** — the criterion I
had to correct in #3104. Below it, `pauseAborted.add`,
`session.dispose()` and `activeSessions.delete` mutate in-memory state
in this tick, and other paths read those maps. Contrast
`self-healing.ts`'s fan-out, where three of four guards only gated work
the listener already `void`s and so *were* convertible via the async
resolver (#3094).

## What it costs, and the obvious reading is backwards

An evacuation **into** a renamed destination still falls through and
disposes correctly — no bug there.

The failure is the other direction: on a board whose **hold or intake**
lane is renamed, arms 1 and 2 stop matching, so a card **sitting still
in its own planning lane** is treated as evacuated and its live triage
session is aborted mid-run.

I state it that way because "renamed board → guard misses → nothing
happens" is the pattern everywhere else in this program, and here it
inverts.

## Census

**Unchanged at 1**, deliberately. Blocked, and now documented as *fully
literal* rather than part-converted — which is the fact a future pass
needs in order not to make it worse.

## Measured

- Comment-only.
- `src/__tests__/triage*` — **25 files / 374 tests pass**.
- `tsc --noEmit -p packages/engine` clean; `check-fnxc-future-dates`
clean.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:40:21 -07:00
gsxdsm
0e4a559a0a test(census): make the tighten fixture self-maintaining instead of pinned to committed state (#3106)
## What broke, and why it will break again

The two cases in this block assert the CLI tightens an inflated
allowance **by exactly the inflation**. That arithmetic only held while
the *committed* baseline matched the tree — so it broke the moment a
fleet PR took `self-healing.ts` from 26 to 22 without re-recording. The
CLI correctly tightened to 22 while the fixture expected 26, and both
cases went red for a reason that had nothing to do with the code under
test.

#3101 fixed that instance by committing the number. **This fixes the
class.**

## Why it recurs

The census **exits 0 on a drop** — deliberately, so one worker's merge
can't redden the gate for everyone else. The cost is that the committed
baseline goes stale *silently*: every run rewrites the file, prints
`COMMIT IT`, and exits 0. This fixture is what eventually trips over it.

With a fleet actively converting the largest file (eight open PRs
against `self-healing.ts` as I write this), that's a recurring red, not
a one-off.

## The change

The fixture syncs its temp copy to the tree with `--strict
--update-baseline` **before** inflating. The assertion is then about the
CLI's behaviour rather than about what happens to be recorded on disk.

## Differential proof, both directions

Against an artificially staled baseline (22 → 26):

| | result |
|---|---|
| with this change | **40 passed** |
| without it | **2 failed / 38 passed** |

So the fixture now tolerates drift it previously broke on — and still
fails if the CLI stops tightening, which is the property it was written
to guard. That second half matters: a fixture made tolerant of
everything would be worse than the flake.

The CLI invocation is extracted to a `runCli` helper so the sync run and
the assertion run share one path. No behaviour rides on that extraction.

## Measured

| check | result |
|---|---|
| census suite, clean tree | 40/40 |
| census suite, staled baseline | 40/40 |
| `--strict` | exits 0, no residual drift |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:33:37 -07:00
gsxdsm
0da19f7963 fix(core): a renamed archive lane was recorded as done in the eval corpus; flag the scheduler's two honest literals (#3100)
Two pieces, both about the same distinction: which literals are worth
**converting** and which are worth **naming**.

## Converted — the eval corpus was mislabelling renamed archive lanes

`collectDeterministicSignals` writes `column` as a two-value eval-record
field. Against the `archived` literal, a card resting in a renamed
archive lane was recorded as `"done"`.

No crash, no lifecycle decision — a **mislabelled row in the eval
corpus**, which is a dataset every later comparison reads. That is the
expensive kind of quiet: nothing fails, the numbers just drift.

The collector is sync and pure (no store, no workflow), so the lane
answer arrives as an optional parameter.
`HybridEvaluatorService.evaluateTask` is async and already holds an
optional store, which is where the resolution is paid; a store-less
evaluator degrades to the legacy literal rather than failing.

**Only the archived arm was ever wrong.** A renamed *complete* lane was,
and remains, recorded as `"done"` — which is correct. So only that
answer is resolved, and a third case pins that the widening did not turn
every renamed lane into `"archived"`.

## Flagged, not converted — the scheduler's two honest literals

These are the two `scheduler.ts` literals the sync-lane pass did not
take, and **nothing in the file said why**. That silence is the problem:
the obvious next move is to "finish the job" the way the other ten were
converted, and that would make them **inert, not fixed**.

`getTaskWorkflowSelectionImpl` returns `undefined` unconditionally under
PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR — proved in
`postgres/sync-workflow-ir-is-always-default.pg.test.ts`, and
`check-inert-sync-lane-conversions` already baselines **twenty** guards
in that state in this same file.

They stay literal and **counted**, which is the honest state. An
unconverted literal is visible to the census; an inert conversion leaves
the backlog and takes the evidence with it. The note names the real
blocker — a sync-capable workflow-selection reader — so the next pass
does not spend a cycle discovering this the way I did.

## Measured

- 3 new cases in `eval-signal-collector.test.ts` — file **5/5 pass**.
- **MUTATION**: restoring the `archived` literal fails the renamed case
and leaves **both** the legacy control and the renamed-complete negative
green. The negative matters here: the fix must not turn every renamed
lane into `"archived"`.
- core eval suites — **4 files / 20 tests**; engine scheduler +
evaluator — **14 files / 143 tests**.
- `tsc --noEmit` clean in both packages; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

Both files keep their counts, deliberately:

- `eval-signal-collector.ts` — the remaining entry is the new
parameter's documented default, which is the fallback doing its job.
- `scheduler.ts` — the two literals this PR deliberately leaves visible.

A census that fell here would mean the flags had been marked exempt,
which would assert the code is fine. It is not fine; it is blocked, and
those are different claims with different expiries.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:29:21 -07:00
gsxdsm
f1e96f7a17 test(engine): pin the agent-link-drift terminal check on a renamed board (an agent stayed linked to finished work) (#3102)
Second of the two uncovered sweeps I flagged when #3078 merged. Not a
conversion — the conversion is already on main
(`driftedTerminalColumns`, landed by another fleet PR). This is the
coverage it shipped without.

## What was unprotected

Every existing case in this file uses `done` or `archived`, where the
literal is correct. So the terminal check had **no renamed-board case at
all**, and the same measurement that caught #3078 applies: a green file
proves nothing about a conversion whose fixtures can't express the
failure.

What the literal cost on a renamed board: **a durable agent stayed
linked to a finished task forever.** A linked agent is not free to pick
up new work, so the drift this sweep exists to clear is exactly the
drift it stopped clearing.

## Measured, both directions

| case | result |
|---|---|
| agent linked to a task in a RENAMED complete lane is cleared | **fails
on revert** — `taskId` still `"FN-9"`, agent pinned to finished work |
| agent linked to a task still in a RENAMED wip lane keeps its link |
passes either way — the sweep must narrow, not widen |

14 pass on current main; reverting the terminal check to the id pair
fails exactly one.

## Note on the fleet

This sweep was converted by someone else's PR while I was writing the
test for it — I found out because my revert probe hit
`driftedTerminalColumns`, a name I did not write. That is the collision
pattern working in a *useful* direction for once: their conversion, my
coverage, no duplicated code.

It also means the two of us independently chose the same sweep from a
7-PR pileup on this file. Assigning files from the census list would
still be cheaper than discovering the overlap in a test harness.

## Verification

`self-healing-agent-link-drift` **14 passed** · `pnpm test:gate` 13 +
161 + 487 + 71 · lint — green.
2026-07-31 04:25:59 -07:00
gsxdsm
d448ab6951 fix(engine): the merge-refusal reason was classified by a column id, and it lands in run-audit (#3098)
Claimed `auto-merge-finalization.ts` — and this one is a **reversal of
an earlier audit in the same file**, which is the interesting part.

## The earlier note said "diagnostic only". It was wrong about the
consequence

`validateWorkflowDoneMergeProof` picks between two refusal reasons with
`task.column === "done"`. Both arms return `{ ok: false }`, so this
never changed which branch ran — and on that basis a prior pass recorded
it as *"REAL but DIAGNOSTIC-ONLY"* and declined it, reasoning that
widening a signature to improve an error string is a poor trade.

**The reason is not an error string.** It is written to run-audit
metadata alongside `previousColumn` — `merger-merge-lifecycle.test.ts`
asserts exactly that — and that row is what an operator reads to find
out why a merge was refused.

So on a board whose complete lane is not called `done`, a card resting
in that lane was refused with the generic `missing-merge-confirmation`:
the classification for a card that is **not in the complete lane at
all**. The audit trail recorded the opposite of what happened. A wrong
record is worse than a vague one, because it gets acted on.

## The trade was also cheaper than the note claimed

The function is **already async** and **already takes an options bag**.
`resolveFinalizationColumns`, two functions up in the same file,
**already builds this exact predicate** for its own guard.

Nothing new is resolved. The answer that existed is handed down instead
of being re-asked with an id — the half-conversion shape this program
keeps finding, here inside a single file, one line apart: the caller
guards on the resolved `isCompleteColumn(latest.column)`, then calls a
validator that re-asked the same question with the literal.

`isCompleteColumn` is **optional with the legacy literal as its
default** — the same default-to-legacy contract the lane-parameter
vocabulary uses elsewhere — and `check-lane-wiring` watches the
parameter, so the two call sites cannot silently stop passing it.

## Measured

- New `merge-proof-reason-renamed-complete-lane.test.ts` — **2 pass**.
- **MUTATION**: dropping the parameter fails the renamed case and leaves
the legacy **control** green. The control earns its place: a failure now
means *"renamed board"*, not *"the refusal stopped working"*.
- **Driven through `finalizeProvenAutoMergeTask`**, not by calling the
validator with the new argument. The contract under test is the
**wiring** — a test that passed the argument directly would assert my
own parameter works and prove nothing about the seam that was broken.
- **The audit row is asserted, not just the return value.** The return
value alone is not the contract that failed here.
- merger / auto-merge suites — **5 files / 159 tests pass**.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

`auto-merge-finalization.ts` stays at **2**, deliberately. Both
remaining entries are now documented **degraded-fallback arms** — the
resolver's `catch` and this parameter's default — which is the right
kind of literal rather than a missed conversion. Converting a fallback
to a resolution would defeat its purpose.

## A note on the FNXC gate

My first stamps were dated `2026-08-01` while local today is
`2026-07-31`. `check-fnxc-future-dates` caught it and I re-stamped.
Worth mentioning because it is the second time this session that a
date-only local-calendar comparison has caught a stamp written near
midnight — the gate is doing real work, not ceremony.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:22:33 -07:00
gsxdsm
a7b2a757fa fix(engine): serialise wedge handling per task, then convert the lane guards it was blocking (5 → 1) (#3087)
The largest unclaimed census cluster, and the one two earlier fleet
passes explicitly declined.

## The standing blocker, taken on

Both passes converted these four ids and reverted, each time after the
same test went red:

```
task-wedge-notification.test.ts > sends one actionable push and mailbox message per active terminal episode
  expected 2 calls, got 1
```

Their diagnosis was right and I have kept it: this branch **resolves** a
wedge episode, `handleTaskUpdated` starts it fire-and-forget from a
synchronous `(task) => void` listener, and **any** await introduced
before the resolve lets a re-wedge arriving close behind reach `claim`
while the previous episode is still active — `claimed: false`, second
operator notification silently dropped. Column resolution needs an
await, so the conversion could not be made safe from inside the branch.

Both notes named the fix and left it for "whoever owns the wedge episode
contract": *serialise wedge handling per task*. This PR does that, then
takes the conversion.

## 1. Serialisation

`enqueueWedgeHandling` chains handling per task id, so
resolve-then-claim keeps its order however many awaits either branch
acquires. Details that matter:

- **Keyed by task, not global** — different tasks stay concurrent, so
this is not a throughput regression on a busy board.
- **The map entry is dropped when its chain drains**, and only if no
later link was appended while it ran, so it does not grow with the task
table.
- **Links never reject.** `maybeNotifyTaskWedge` already owns its error
handling; a rejected link would poison every later notification for that
task.

## 2. The conversion it was blocking

The four ids are an enumeration of *"every lane except review"* — the
lanes whose occupancy proves a wedged card's lifecycle has visibly
resumed. On a renamed board none of them matched, so a recovered card's
episode never resolved. Two consequences, and the second is worse than
the first:

1. the operator keeps an open "needs operator action" alert for work
that has moved on;
2. an active episode **suppresses re-claim**, so the *next* genuine
wedge on that task is never delivered.

Membership over the four roles, legacy-seeded, so an unconverted board
resolves exactly the four ids it used to compare.

## Measured

**The acceptance test the earlier notes named is the gate on both
halves.** With the conversion and *without* the serialisation, "sends
one actionable push and mailbox message per active terminal episode"
fails exactly as they reported. With the serialisation, green. I
reproduced their finding rather than taking it on trust — it is the
evidence that the serialisation is load-bearing and not incidental
refactoring.

| | result |
|---|---|
| `task-wedge-notification.test.ts` | **15/15** (2 new) |
| notification suites | **11 files / 234 tests pass** |
| `tsc --noEmit -p packages/engine` | clean |
| census `--strict`, `check-lane-wiring`,
`check-inert-sync-lane-conversions`, `check-fnxc-future-dates` | clean |

**MUTATION**: restoring the four literals fails the renamed-recovery
case and leaves its paired negative green.

**A vacuity I caught and fixed, worth stating plainly.** My first
version of the renamed case recovered the card with `status: "queued"`.
`hasProgressed` is an OR whose other arm is *"status is a non-failed
string"* — so that arm answered true and the column comparison never
ran. The mutation did not fail it. The case now clears `status` and
`error` together, which makes column membership the only thing that can
resolve the episode, and the paired negative uses the identical shape so
only the lane differs.

## Census

| | before | after |
|---|---|---|
| `notification-service.ts` | 5 | **1** |
| repo backlog | 71 | **67** |

## The remaining 1, flagged not guessed

`isManualMergeHold` (`task.column !== "in-review"`) is sync, and so is
its only caller `classifyWorkflowTransitionNotification`, reached from
the same `handleTaskUpdated` listener. Converting it means making that
whole chain async — a change to notification *classification ordering*
against every other `task:updated` handler, which is a different
contract from the episode one this PR owns. The serialisation added here
does not cover it: it wraps wedge handling, not transition
classification. Threading a pre-resolved `LifecycleColumns` in as a
parameter is the likely fix, and it wants the same gate-placement
judgement applied deliberately rather than swept in behind this.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:10:03 -07:00
gsxdsm
701677a2e5 test(engine): pin #3078's executor-owned skip — it merged without coverage (204 tests passed against the reverted fix) (#3090)
#3078 merged its conversion of the orphaned-pending-step-results sweep
**before this test landed**, so that sweep is on main with no coverage.
This closes the gap.

## The gap was measured, not assumed

With all three of #3078's conversions reverted, **all 204 self-healing
tests still passed**. I had cited that number as verification when I
opened it. It was meaningless for that change: every existing test in
the file uses `in-review` / `in-progress`, where the literal is correct,
so none of them could see the defect.

This is the same "a green suite is not coverage" failure I flagged in
other PRs today — in my own work, twice. The only reason I caught it is
that I finally ran the revert check on myself.

## Two cases

- **An executor-owned card in a renamed wip lane is SKIPPED.** Against
the pre-#3078 sweep this fails: the sweep reaches a card an executor is
actively running and rewrites its `pending` step results to `failed` —
the one thing that file's header says it must never do. The liveness
triple does not cover it; those legs prove an *in-process* session, and
an executor on another node or between session handles is exactly what
the column skip is for.
- **A genuine orphan on that same renamed board is still recovered** —
the skip must narrow, not disable. Passes either way, deliberately.

## What it pins, precisely

The **invariant**, not a line. Reverting either single guard still
passes, because the page-snapshot check and the fresh-row re-read
protect independently. What fails is reverting the sweep's column
handling as a whole — which is the condition worth pinning, and matches
the project's "fix the invariant, not the repro" rule.

## Still uncovered, said plainly

#3078's other two sweeps — worktree-metadata liveness and agent-link
drift — have no dedicated case. The orphaned-step-results sweep got the
test first because it is the one that can corrupt a live executor's
state. The other two remain honest debt rather than implied coverage.

## Verification

`self-healing-orphaned-pending-step-results` **10 passed** on current
main · full self-healing suites 204 · `pnpm test:gate` 161 + 13 + 487 +
71 · lint — green.
2026-07-31 04:09:48 -07:00
gsxdsm
9e242ea294 fix(engine): backlog pressure called every dependency unfinished on a renamed board (#3081)
## The third lane question

This reporter had **three** lane questions. Two were resolved when the
file's query-blindness was fixed — hold and wip, both through
`resolveProjectColumnsForRoles`. The third sat one method down and was
never touched:

```ts
if (dependency.column !== "done") return false;
```

One board, two lane answers.

## What it cost

On a renamed board every dependency reads unfinished, so
`isRunnableCandidate` rejects every card that has one. The
backlog-pressure alert then names **only dependency-free cards** as the
runnable ones.

The failure mode is the quiet kind: the report still renders, the counts
are right, and the candidate list looks plausible. The operator is told
the queue is blocked on nothing in particular. No default-board test can
see it — which is exactly why the earlier conversion of this same file,
which fixed its reads, left this behind.

## Fix

`finishedColumns` (complete ∪ archived) resolved once by the async
caller alongside hold and wip, then passed into the sync predicate.

- **Required parameter, not optional-with-a-literal-default.** An
optional parameter leaves `done` in the file as a silent fallback and
the next caller gets pre-conversion behaviour by writing nothing.
- **Archived is included** because a dependency that has been archived
is finished too — and this reporter already reads with `includeArchived:
true` precisely so archived blockers resolve.
- **Async resolution.** `resolveProjectColumnsForRoles`' only store read
is `listWorkflowDefinitions()`, a project-wide async read that works
under PostgreSQL. That is the line between a real conversion and the
inert sync-IR kind (#3058), and the new test supplies its board through
that same reader so it exercises the production path.

## Census

| | before | after |
|---|---|---|
| `backlog-pressure-reporter.ts` | 1 | **0** |

## Measured

- One new case; file **11/11 pass**.
- **MUTATION**: restoring `dependency.column !== "done"` fails it.
- The case asserts **both directions in one test** — a dependency
resting in the board's own complete lane makes its card runnable, *and*
a dependency still in the hold lane still blocks it. Asserting only the
first would pass against a predicate that had simply stopped checking
dependencies.
- The file already had a `RENAMED_IR` scoped to its second describe;
mine is a distinct `RENAMED_DEPENDENCY_IR` with different lane names. I
hit the shadowing first and the test failed as `under-threshold` — worth
noting because a same-named fixture that silently resolves to the
*other* board is precisely how a renamed-lane test goes vacuous.
- `tsc --noEmit -p packages/engine` clean; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Flagged, not guessed

Adjacent census entries I looked at and deliberately left:

- **`executor.ts` (4)** — all inside a sync `task:moved` listener.
Converting via `resolveTaskWorkflowIrSync` would be inert for #3058's
reason, and making the listener async reorders it against every other
subscriber. Correctly out of scope, as #3048 judged.
- **`triage.ts:724`** — half-converted in the same shape:
`disposeLanes.hold`/`.intake` come from a sync resolver, so the resolved
arms are themselves inert and "finishing" the guard would add a third
inert comparison.
- **`auto-merge-finalization.ts` (2)** — one is the resolver's
documented degraded fallback (the live arm calls `columnHasFlag`), the
other is already recorded as a deferred signature-widening whose cost
exceeds the error string it sharpens.
- **`in-review-stall.ts:196`** — an explicitly marked DELIBERATE-LITERAL
no-metadata fallback.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:51:43 -07:00
gsxdsm
3c531d984c fix(engine): self-healing lane cluster round 2 — 38 → 26 (two sweeps could disturb live work) (#3078)
The largest census cluster became **unclaimed again** when #3055 closed
conflicting. I had closed my own #3050 an hour earlier expecting #3055
to land, so this re-applies the conversions #3047 and #3049 did not
cover.

**Re-applied from current main rather than rebasing the closed branch.**
The conversions are small; the conflict archaeology is what went wrong
last time — nine conflicts against #3049, several on variable names
identical to mine, and my mechanical fixup corrupted the file badly
enough that I aborted. Starting from main cost less than resolving that
and carries no risk of resurrecting a stale line.

## Census

| | before | after |
|---|---|---|
| `packages/engine/src/self-healing.ts` | **38** | **26** |
| repo-wide column guards | 84 | **72** |

## Three sweeps, existing role helpers only

| sweep | roles | what it did on a renamed board |
|---|---|---|
| worktree metadata | terminal + wip + review | rebound finished cards
every pass, **and the FN-5256 liveness guard went silent** |
| orphaned pending step results | wip | **could rewrite `pending`
results under a live executor run** |
| agent-link drift | wip + review + terminal | evaluated agents whose
task was plainly still executing |

Two of these disturb **live** work, which is why they were worth redoing
now rather than leaving for the next fleet round:

- The worktree-metadata sweep clears `worktree`/`branch` metadata. Its
liveness guard is the thing standing between that and a running shell
(FN-5256). Keyed on ids, it matched nothing on a renamed board. The
scope-override safety condition beside it now reads the **same resolved
sets**, so the two cannot disagree about which lanes are live —
previously they were two independent literal lists.
- The orphaned-step-results sweep's own header says it must never touch
an executor-owned row. The id-keyed skip made it do exactly that.
Resolved once per sweep, outside the paging loop, so a large board still
pays one resolve.

## Flagged, not guessed — the 26 that remain

Unchanged from my earlier audit and re-verified on this base:

- **Sync predicates** (`isWorkspaceOwnerLive`, the pause-abort
classifier, the phantom-binding check, the `task:moved` listener
guards). No store handle; converting means a signature change or making
a synchronous event listener async, which reorders handlers against a
synchronous emitter.
- **Already-converted fallbacks** — `own.length > 0 ? own.includes(...)
: task.column === "in-review"`. The resolved answer wins; the literal is
the documented no-metadata path.
- **The notification-route `fresh.column === "todo"` sites** — measured
previously: any `await` before the wedge resolve drops an operator
notification. Needs the wedge-episode contract, not a column pass.

## Verification

self-healing suites **204 passed** · agent-link-drift +
query-filter-blindness **83 passed** · `pnpm test:gate` 161 + 13 + 487 +
71 · lint · census `--strict` · lane-wiring — green.
2026-07-31 03:48:44 -07:00
gsxdsm
58791fac88 fix(self-healing): route pause-abort recovery on resolved columns (self-healing 38 → 34) (#3075)
First real cut into `self-healing.ts`, the last large cluster. Converts
the **pause-abort recovery router** — a coherent unit with one owner,
rather than a scattered pass.

## Census before/after

| Metric | Before | After |
|---|---:|---:|
| COLUMN guards (backlog) | 86 | **82** |
| `self-healing.ts` | 38 | **34** |

## The bug this was hiding

The router keyed on three literals — `in-review` twice (review progress,
manual merge hold) and `todo || in-progress` (active work). On a renamed
board **all three stop matching**, so a parked card in a renamed lane
falls through to `no-action` and is never recovered — silently, no log
line, and with every existing test still green because they all use
legacy ids.

## Conversion

Used the file's own resolvers. `resolveReviewColumnsFor` already
existed; added `resolveActiveWorkColumnsFor` as its sibling from the
same `columnsWithFlag` / `resolveLifecycleColumns` helpers — no new
vocabulary.

**ACTIVE WORK is hold + `countsTowardWip`, deliberately NOT the intake +
hold that the neighbouring `resolvePreWipColumns` returns.** The router
asks *"is this card mid-flight, so a requeue is right?"* — an intake
lane is not mid-flight; a WIP lane is. Reusing the intake-shaped helper
would have widened the requeue to triage rows and dropped in-progress
ones. Because the two sets **overlap on `hold`**, that mistake looks
correct in every legacy-id test. This is the trap worth knowing about
for the rest of the file: it has several resolvers, and picking the
nearest one is not the same as picking the right one.

**`columns` is required, not optional-with-a-fallback.** The router has
exactly two callers — the candidate filter and the post-re-read
re-verify — and they must agree. An optional parameter lets one resolve
and the other default, and that divergence surfaces as a sweep that
selects a card and then declines to act on it, writing nothing. Required
makes it a compile error. (Same reasoning as #3059; safe here because
both resolvers union the legacy ids internally, so "required" never
means callers invent a column set.)

**On the filter/re-verify hazard I flagged earlier:** both call sites
are the *same function*, so converting it once keeps them consistent by
construction — no split-CAS risk. Per-task resolution can't hoist out of
the loop but must not read an IR per row, so the marker test
(column-independent) stays a cheap sync prefilter and the IR is read
only for rows that pass it, over a shared cache. `parked` keeps its
exact former membership, so the log count still means what it said.

## Verification

- **Revert-proof:** reverting the three guards to literals fails 2 of 3
routing cases. The third exercises the new resolver directly, so it
cannot fail on revert — stated rather than counted as evidence.
- 4 self-healing suites, **437 tests green**
- `tsc --noEmit` clean; eslint clean
- Degraded-resolution case included, since the legacy-id union is what
keeps recovery alive on an unreadable workflow

## Still flagged in this file (not guessed)

34 remain. They are not one batch: the sweeps around
L2855/3297/8877/9018 depend on the unowned `listTasks({ column })`
decision, and L5330/5359 sit under the FN-5256 liveness guard whose own
comment says those columns can be live when the heuristic calls them
stale.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:45:52 -07:00
gsxdsm
8eef8852a0 fleet: 4 long-tail fallback arms become named sets (census 101 → 97) (#3064)
## Census

| | column guards |
|---|---|
| before | **101** |
| after | **97** |

The single-guard long tail is **19 files**. This converts the four whose
legacy arm is unambiguously a fallback on an already-converted guard;
the other 15 are flagged below rather than guessed at.

## Two shapes

**`in-review-stall.ts`, `stalled-review-detector.ts`** — the resolved
answer with an inline legacy arm:

```ts
reviewColumns ? reviewColumns.has(col) : col === "in-review"
→ (reviewColumns ?? LEGACY_REVIEW_LANES).has(col)
```

**`merger.ts`, `in-process-runtime.ts`** — belt-and-braces:

```ts
col !== (lifecycle?.complete ?? "done") && col !== "done"
```

That accepted the resolved lane **or** the legacy id, stated twice. A
union set says it once, so the two halves can't drift apart — which is
the real risk with a duplicated condition.

## A finding for anyone else marking fallbacks

`in-review-stall.ts` **already carried a `DELIBERATE-LITERAL` marker**
on that arm and was counted anyway. The marker sits in a comment *inside
a ternary*, which the census's leading-comment lookup doesn't reach.

So: **naming the set works, marking it does not.** Worth knowing before
someone marks a fallback and expects the count to move.

## No behaviour change

`new Set(["in-review"]).has(x)` answers exactly what `x === "in-review"`
answered, and the union sets accept exactly the two lanes their
conditions already accepted.

## Flagged, not converted

The remaining 15 single-guard sites need individual judgement, not a
mechanical pass:

- **plain unconverted guards with no resolution in scope** —
`audit-ops`, `lifecycle-ops`, `merge-queue-ops`, `task-id-integrity`,
`backlog-pressure-reporter`, `ephemeral-worker-manager`,
`ResearchTaskActionModal`
- **sites where the literal IS the answer** — `eval-signal-collector`
maps a column to an archive-vs-done *label*; `TaskCard` reads a
completion timestamp
- **already resolved on their line** — `triage.ts`,
`restart-recovery-coordinator.ts`, both covered by open PRs

## Measured

| check | result |
|---|---|
| core stall suites | 4 files, **85 tests green** |
| engine merger/runtime suites | **1044 tests green** |
| five gates + strict census | green |
| `tsc` (core, engine) | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:37:18 -07:00
gsxdsm
c220455e3a fleet: 10 inline fallback arms become named sets (census 102 → 92) (#3061)
## Census

| | column guards |
|---|---|
| before | **102** |
| after | **92** |

Five files drop to **0** guards each. Baseline re-recorded in the same
commit.

## A cluster the census could not distinguish from real debt

**Every site here is already converted.** Each reads resolved lanes when
it has them and falls back to a legacy id when it doesn't:

```ts
reviewColumns ? reviewColumns.has(task.column) : task.column === "in-review"
```

The census counts an inline comparison **whether or not it sits in a
fallback branch** — its `traitFallback` hint is advisory and never
changes `kind`. So ten correctly-converted guards sat on the backlog
permanently, and the number stopped distinguishing *work still to do*
from *documented degraded answers*.

Naming the fallback set fixes the bookkeeping without touching
behaviour: `new Set(["in-review"]).has(x)` answers exactly what `x ===
"in-review"` answered.

## Files

| file | sites | what they gate |
|---|---|---|
| `restart-recovery-coordinator.ts` | 4 | three shared review gates +
one `??` default |
| `github-tracking-state.ts` | 2 | complete / archived lane predicates |
| `planner-overseer.ts` | 2 | wip / review classification |
| `async-mission-store-queries.ts` | 2 | terminal complete / archived |
| `register-task-workflow-routes.ts` | 2 | wip promotion target,
archived respecify guard |

**No behaviour change is claimed and none is intended** — that's the
point. These were already right; only the accounting was wrong.

## Worth the fleet's attention

Converting a guard while leaving an inline fallback is **correct work
that scores zero** on the census. My own first pass at `reads.ts` did
exactly that — behaviourally correct, census unmoved. Anyone converting
this way is doing real work the number won't credit, and the backlog
will look stuck.

## Measured

| check | result |
|---|---|
| engine suites | **173 tests green** |
| core mission suites | **70 tests green** |
| dashboard route suites | **211 tests green** |
| five gates + strict census | green |
| `tsc` (core, engine, dashboard) | clean |

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Standardized fallback handling for workflow stages, including
in-progress, review, completed, and archived states.
* Preserved existing behavior when explicit workflow column settings are
available or unavailable.
* Improved consistency across task tracking, planning, and recovery
workflows.

* **Chores**
* Updated lifecycle tracking baselines to reflect current source-file
coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:34:27 -07:00
gsxdsm
18fab9b8e5 fleet: name the WIP half of an existing fallback (census 88 → 87) (#3070)
## Census

| | column guards |
|---|---|
| before | **88** |
| after | **87** |

## What

`ephemeral-worker-manager.ts` answers its unresolvable-workflow default
two ways, two lines apart:

```ts
if (TERMINAL_TASK_COLUMNS.has(task.column)) return true;   // named set — not counted
return task.column !== "in-progress";                       // inline — counted
```

Both are the **same documented fallback** — the block carries one
`DELIBERATE-LITERAL` marker covering both — but only the inline one was
on the backlog, because the census reads comparisons regardless of which
branch they sit in while a set is a definition.

Naming it makes the pair consistent and stops the site reading as
unconverted debt.

## Correction to my own flag in #3064

I listed `ephemeral-worker-manager`, `backlog-pressure-reporter`,
`merge-queue-ops` and `lifecycle-ops` as *"plain unconverted guards with
no resolution in scope."*

**That was wrong for all four.** Each already imports the resolvers — 7,
4, 3 and 2 references respectively. I wrote the flag without checking,
which is the same mistake as an untested deferral rationale, just inside
a PR body instead of an issue.

Re-examined, the other three are genuinely harder rather than unresolved
— and these are the real reasons:

- **`backlog-pressure-reporter:197`** classifies a **dependency**, a
different row from the one the caller resolved. Per-dependency
resolution is needed or it repeats the wrong-row shape that `taskRevert`
is blocked on.
- **`lifecycle-ops:655`** guards an emit whose **target** is also a
literal (`to: "archived"`). Converting the guard alone leaves the pair
inconsistent — the move-target half is invisible to this census.
- **`merge-queue-ops:352`** is an early return on an already-complete
task inside a merge path that resolves lanes elsewhere; the placement
needs its own judgement about which resolution it should share.

They stay flagged, now with the real reason rather than an unchecked
one.

## Measured

| check | result |
|---|---|
| ephemeral-worker suites | green |
| four gates + strict census | green |
| engine `tsc` | clean |

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:31:03 -07:00
gsxdsm
21ef60047e fix(engine): a second complete lane is terminal too — restore the scheduler's dependency reconciliation (main is red) (#3065)
## main is red, and this is the fix

```
FAIL src/__tests__/scheduler-renamed-hold-events.test.ts
 > dependency unblocking (failure mode is a card that waits forever)
 > finds dependents resting in the renamed hold column when a blocker completes
 AssertionError: expected [] to include 'drafting'
```

Not in the thin merge gate's `engine-core` allow-list, so CI stayed
green and only the non-blocking full suite sees it. The test file is
unchanged since #2518; #3051 converted the guard underneath it.

## What broke

#3051 turned the guard into `to === parked.complete || to ===
parked.archived`. `resolveLifecycleColumns` answers **first match per
role** — the right shape for a move *target*, the wrong shape for *"did
this card just reach a finished lane"*, which is a membership question.

Two consequences, both silent:

1. **A board with more than one complete-trait column reconciles
nothing** when a blocker finishes in the second one. The test file's own
header flags this path specifically: *"This one is NOT latency: a
dependent never gets unblocked, so it waits on a blocker that is already
done."*
2. The legacy `done`/`archived` ids stopped matching at all — the
failing assertion.

## Fix

`resolveTaskParkedColumnsSync` gains `terminal`, a membership set:
legacy `done`/`archived` seeded, then **every** complete- and
archived-trait column from the task's own IR.

Seeding legacy ids is safe in the direction that matters here. This is
an **inclusion**: a superset makes the reconciliation run on a move it
would otherwise ignore — one extra query, and it cannot wrongly withhold
work. Seeding a **refusal** is the bug (`node-override-guard.ts`
documents that one); this is not that.

Same sync IR path and same fail-soft legacy default as the single-column
answers, so event ordering and unresolvable-workflow behaviour are
unchanged — the constraint the sync resolver's own header sets.

The two sibling guards in the same listener (dispatch-oscillation reset
at what is now line 1097, and the scheduling wake at 1114) had the
identical arity defect and convert with it.

## Measured

| | result |
|---|---|
| before | `scheduler-renamed-hold-events`: **1 failed / 9 passed** |
| after | **11 passed** (one new case) |
| `src/__tests__/scheduler*` | **14 files / 143 tests pass** |
| `tsc --noEmit -p packages/engine` | clean |
| census `--strict` / `check-lane-wiring` / `check-fnxc-future-dates` |
clean, no baseline movement |

**Proved it is main's red, not my branch's:** I checked out
`origin/main:packages/engine/src/self-healing.ts` over my unrelated
fleet branch and re-ran — identical failure. Then branched this fix
straight off `origin/main`.

**Mutation-tested.** Restoring `to === parked.complete || to ===
parked.archived` fails **both** the pre-existing case and the new
second-complete-lane case. The new test is not vacuous: the second
complete lane is invisible to first-match resolution, so it cannot pass
against the old guard.

## Not done here

I did not convert the remaining `to === parked.review` / `from ===
parked.wip` single-column comparisons in this listener. Review is
genuinely two roles (`mergeBlocker` + `humanReview`) and wip has its own
limit-setting semantics — both want the same membership-vs-target
judgement applied deliberately rather than swept in behind a red-fix.
Flagged, not guessed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 03:12:53 -07:00
gsxdsm
af470f7c05 convert(engine): self-healing lane cluster 56 -> 38 guards (repo 126 -> 108) (#3049)
## Census before / after

```
                                    before    after
self-healing.ts column guards          56        38
repo-wide COLUMN guards (backlog)     126       108
```

`self-healing.ts` was the largest single cluster by a wide margin — 56
guards against 12 in the next file. Baseline re-recorded in the same
commit; `--strict` green.

## Converted: 15 guards across 11 sweeps

Existing helpers only — `resolveProjectColumnsForRoles` with
`TERMINAL_ROLES` / `REVIEW_ROLES` / `countsTowardWip` / `hold` /
`archived`, the same shape this file already uses. No new helper, no new
resolution pattern.

What each was silently doing on a renamed board:

| sweep | behaviour before |
| --- | --- |
| `archiveStaleDoneTasks` | **both** guards inert, so every card counted
as an active dependent and the sweep archived **nothing at all** |
| `reconcileDependencyBlockingLeases` | no holder matched, so a stale
file-scope lease blocking an unmet dependency was never cleared |
| `reconcileCompletedBlockedTasks` | work whose blocker had cleared
stayed parked instead of advancing |
| `reconcileInReviewUnmetDependencies` | a card sat in review with unmet
dependencies and no rebound |
| `reclaimStaleActiveBranches` | archived cards were eligible for branch
reclaim |
| `reconcileInReviewBranchRebind` | the rebind list was empty |
| `autoReboundPausedScopeDecayDetailed` | no card was ever seen as
executing |
| `detectStalledCards` | finished cards counted as stall candidates |
| `recoverApprovedStrandedAiMergeCommit`,
`recoverDriftedAgentTaskLinks`, `cleanupStaleTempMergeWorktrees` | same
shape |

**Reused rather than duplicated:** `recoverWedgedActiveMerge` already
resolves `wedgedReviewColumns` via `resolveReviewColumnsFor` three lines
above the site I was converting, so the site now uses it instead of a
second resolution of the same question.

## One site I converted and then reverted

`clearStaleBlockedBy`'s memo closure carries an FNXC note stating the
literal is **deliberate**: the closure only decides whether to re-log an
already-logged blocker, so a renamed board costs a duplicate log line —
not a wrong lifecycle decision — and restructuring a sweep's control
flow to convert a logging decision is the wrong trade.

I read that note *after* editing the line. Restored.

Worth flagging separately: **it has the reasoning but no
`DELIBERATE-LITERAL` marker**, so the census keeps counting it and it
re-appears in the backlog as if unexamined. That is a marker gap, not a
conversion gap — the next person will make the same mistake I did.

## Not converted — flagged, not guessed

Ten of the twenty-four remaining sites are in **sync predicates with no
resolution seam**:

- `classifyPausedAbortWorkflowRecovery` (3)
- the `start()` task-moved listener (5) — compares event `from`/`to`
columns inside a sync callback
- `isWorkspaceOwnerLive` (1)
- `isPhantomExecutorBinding` (1)

Converting these means threading a flags parameter down from every
caller — precisely the unwired-optional-parameter shape this program
keeps finding inert (five were live on `main` at once per
`unwired-lane-parameter-guard`). They need a decision about *where the
resolution lives*, not a guess from me.

The other **14** are in async sweeps with a seam available and are
ordinary follow-on work in this same file.

## Verification (measured)

- self-healing suites — **816 passed / 41 files**
- `tsc --noEmit`, `eslint` — clean
- `pnpm test:gate` — green
- `lifecycle-column-census --strict`, `check-lane-wiring`,
`check-sql-column-literals`, `check-inert-flag-seams`,
`check-fnxc-future-dates` — green

No changeset: `@fusion/engine` is private.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Self-healing workflows now continue functioning when workflow columns
are renamed.
* Improved recovery for stalled, blocked, paused, or disconnected
workflow states while preserving existing filters and actions.
* Temporary merge worktrees and drifted agent links are cleaned up more
reliably.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-31 02:54:39 -07:00
gsxdsm
0fd3e38628 test(engine): PR #3051's scheduler conversion is inert — live-PG refutation (#3058)
## Escalation — a conversion on main changed nothing

**#3051 ("scheduler.ts 12 → 2 lifecycle-column guards") is inert.** It
widened `resolveTaskParkedColumnsSync` from `{hold,intake}` to the full
role set and replaced ten handler literals with `parked.review` /
`parked.wip` / `parked.complete` / `parked.archived`. The census fell by
ten. The behaviour did not change, on any board.

Everything rests on one line in that helper:

```ts
const l = resolveLifecycleColumns(store.resolveTaskWorkflowIrSync(taskId));
```

`resolveTaskWorkflowIrSync` resolves through the sync workflow
**selection** reader, which answers `undefined` for every task under
PostgreSQL — the shipped backend. The resolver takes its `!workflowId`
branch and returns the **default builtin IR**.

Note the shape precisely, because the obvious reading is wrong and this
PR corrected itself on it mid-run: the helper does **not** get
`undefined` and fall through to `?? legacy.review`. It gets a **real IR
that resolves real traits** — the default board's. So `parked.review` is
`"in-review"` for every card on every board, the `?? legacy` arms are
dead code, and the helper answers with full confidence. It looks
resolved at every level except the one that decides the answer.

## Evidence (live PostgreSQL, 3/3 passing)

For a card bound to a **stored** renamed workflow and sitting in that
board's review column (`checking`, carrying `human-review` +
`merge-blocker` + `merge`):

| | sync path (what every converted arm uses) | async resolver | board
actually declares |
|---|---|---|---|
| review | `in-review` | `checking` | `checking` |
| wip | `in-progress` | — | `building` |
| complete | `done` | — | `shipped` |

The async arm is in the test on purpose: it attributes the failure to
the **sync path** and nothing else.

The **control** is the point of the whole thing — on the default board
the sync answer is *correct*, by coincidence rather than resolution.
That is why every default-board scheduler test passes either way, and
how ten inert conversions read as a fix.

## Why this is worse than leaving the literals

A conversion that changes nothing is worse than an unconverted literal,
because **the literal was counted and this is not.** Ten guards left the
backlog, the file now reads as converted, and the next reader has no
reason to look again.

Two further signals that this was not a deliberate trade-off:

1. The FNXC block still standing directly above the handler (unmodified
by #3051) **contradicts the code beneath it** — it says these ten arms
cannot be converted this way and names `resolveTaskParkedColumnsSync` as
the hazard-avoidance device, not the fix.
2. #3051's own added note asserts the fix as fact: *"so on a renamed
board PR monitoring never started or stopped, failure bookkeeping never
recorded, and terminal cleanup never ran."* Those failures are real.
This conversion does not fix them.

## Scope — driven vs argued, stated in the file

- **Driven:** the roles the sync path yields for a real card on a real
stored renamed board, against a real PostgreSQL store, versus the async
resolver on the same card.
- **Not driven, and the file says so rather than substituting a spy:**
the `parked.review` arm's own side effect. Its only outputs are four
dispatch-oscillation fields that do not round-trip through `updateTask`
on this store (measured: writing `dispatchStormCount: 3` reads back
`undefined`), so there is no persisted observable. The behavioural half
is carried by the sibling
`workflow-scheduler-parked-columns-live-e2e.pg.test.ts`, which drives
the **same helper** on the hold role through to persisted state.

## The real unblock

Unchanged from the note already in the file: carry the resolved lanes
**on the `task:moved` payload**, so no listener resolves at all. That
removes the class rather than one instance, and it is the only option
that survives the synchronous-prologue constraint — these listeners run
in the same tick as a synchronous emitter, which is why an `await`
cannot simply be added.

## Verification

`test:gate` exit 0 · live-PG E2E surface **174/174** · census exit 0 ·
`pnpm lint` clean. Test-only; no production file touched.

## Recommendation

Do not revert #3051 — the widened helper is harmless and the note it
added is useful once the resolution is real. **Restore the ten guards to
the census**, or land the payload change. Either way the count must not
read as paid.

Related: **#3055, #3050 and #3049 are all converting `self-healing.ts`
concurrently** — three PRs, one file, 51 guards. Worth de-conflicting
before any of them merges.
2026-07-31 02:51:25 -07:00
gsxdsm
740fea38c2 fleet: restart-recovery-coordinator.ts 4 → 1 (dead fallbacks deleted, not converted) (#3059)
Claiming `packages/engine/src/restart-recovery-coordinator.ts`.

## Census before/after

| File | Before | After |
|---|---:|---:|
| `packages/engine/src/restart-recovery-coordinator.ts` | 4 | **1** |

## These were deletions, not conversions

All three sites were fail-soft fallbacks behind an **optional**
`reviewColumns` parameter:

```ts
return (reviewColumns ? reviewColumns.has(task.column) : task.column === "in-review")
```

Production never took that branch — `self-healing.ts:13646-13649`
supplies the resolved set at every call site. So the correct change is
to make the parameter required and delete the literal, not to swap it
for a role lookup.

## The trap this hit, which would have shipped a crash

**Making the parameter required produced ZERO tsc errors.** That looked
like proof the fallback was unreachable. It is not: the engine
`tsconfig` covers `src` and not `__tests__`, so the type-checker cannot
see the callers that actually relied on the default. Running the tests
surfaced them immediately as `TypeError: Cannot read properties of
undefined (reading 'has')`.

This is the same class as finding 2 in
`docs/solutions/best-practices/proving-a-code-path-actually-runs.md` — a
negative result from a checker that cannot see the thing it is being
asked about. Anyone converting a `src`-only-typechecked package should
assume tsc is blind to test call sites.

The blast radius was also one site larger than grep suggested: the
`isRecoverableMissingWorktreeReviewFailure` **combiner** threads the set
to all three inner predicates. Its own comment already names why — *"a
caller cannot convert the outer question and leave one of the three
inner ones on the legacy id — the half-conversion shape this program
keeps finding."*

Tests now pass the set production always passes, preserving exactly what
each case asserted.

## Remaining 1, flagged not guessed

`L149` uses a different shape (`isReviewColumn ?? task.column ===
"in-review"`) whose callers I did not establish. Absence from grep is
not proof of no caller, so it stays counted.

## Verification

- census: 4 → 1
- `restart-recovery-coordinator` + `self-healing` — **424 tests green**
- `tsc --noEmit` clean; `pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:51:14 -07:00
gsxdsm
06717ac3fa refactor(engine): resolve replan-target's advancement test by role (fleet, 4 sites) (#3052)
## Census

| | column guards |
|---|---|
| before | **126** |
| after | **122** |

`replan-target.ts`: **4 → 0**, and it drops out of the top-files list.
Baseline re-recorded in the same PR, as the ratchet requires.

## What changed

`hasAdvancedPastPlanning` asked "has this card moved past planning" as
four literal comparisons — `in-progress`, `in-review`, `done`,
`archived`. It now asks the same question in roles, from lanes the
**caller** resolves.

## Caller-resolved is the whole point

The module's sync twin `resolvePlannerLanes` reads
`store.resolveTaskWorkflowIrSync`, which returns the **default workflow
IR for every task under PostgreSQL**. Converting through it would have
improved the census while answering about a board the card isn't on —
the second failure shape in the learnings doc, already proven at this
exact seam by
`workflow-planner-lanes-sync-vs-async-live-e2e.pg.test.ts`.

The only production caller is `async`, so it uses
`resolvePlannerLanesForTaskAsync`.

**The caller's own inert resolution is fixed too**, not just the four
arms: `releasedToTodo` compared against `resolvePlannerLanes(...).hold`
— the sync twin — so it read `todo` on every board regardless of
vocabulary. One async resolution now supplies the planner column, the
merged-planning column and the forward lanes.

## Flagged, not guessed

The archive lane is a **separate argument** rather than a fifth
`PlannerLanes` role. Adding the field surfaced a genuine divergence
between the sync and async twins — `_workflow-vocabulary-fixture` models
no archive lane, so they disagree there — and that fixture backs **37
test files**. That divergence deserves its own change with its own
evidence; forcing it through a conversion PR would have meant editing a
37-file fixture to make my own change pass.

## Two larger clusters I did NOT claim, with reasons

I went by census size first and verified before writing:

- **`self-healing.ts` (56 guards, 44% of the backlog)** — already
claimed. Three branches hold it, one checked out in another worktree
(`convert/self-healing-lane-cluster-u7`). I'd drafted four sibling role
helpers before checking; reverted rather than collide.
- **`scheduler.ts` (12 guards)** — blocked by design and already
documented at line 907 by a prior fleet worker. The `task:moved` handler
is `async` but its **prologue is not**: no `await` between entry and the
terminal-blocker branch ~55 lines down, so hoisting a resolution turns
the prologue into a microtask and reorders this listener against every
other synchronous subscriber ("verified, not assumed"). Lazy resolution
doesn't help — the *condition* needs the lanes. Unblocking needs the
emitter to carry resolved lanes on the payload, which is a design change
rather than a conversion.

`restart-recovery-coordinator.ts`'s 4 sites are the trait-fallback arms
the census already counts as converted — converting those would delete
the legacy fallback, not add resolution.

## Measured

| check | result |
|---|---|
| replan + planner-lane suites | 11 files, **102 tests green** |
| triage suites | **374 tests green** |
| five gates + strict census | green; `tsc` clean |
| unconverted callers | byte-identical — absent lanes fall back to
`LEGACY_PLANNER_LANES`, absent `archivedColumn` keeps the legacy id |

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:39:45 -07:00
gsxdsm
24f5ffaffa fleet: scheduler.ts 12 → 2 lifecycle-column guards (#3051)
Claiming `packages/engine/src/scheduler.ts` from the census work order.

## Census before/after

| File | Before | After |
|---|---:|---:|
| `packages/engine/src/scheduler.ts` | 12 | **2** |

Measured with `scripts/lifecycle-column-census.mjs` (kind `column`
only).

## The file already had the right shape — it just under-answered

`resolveTaskParkedColumnsSync` already resolves a task's lanes from its
own workflow, **synchronously on purpose**: these run inside
`task:moved` / `task:updated` listeners, and its own comment records why
an `await` is forbidden there — it would defer everything after it to a
microtask and reorder handlers relative to a synchronous emitter. It
also already fails soft to the legacy ids.

But it only returned `{hold, intake}`, so every *other* lane question in
the same listeners was still asked with a literal. Widening it to the
full role set converted ten sites with no new abstraction, no new
resolution per site, and no change to the event-ordering contract.

## What was silently broken on a renamed board

- **PR monitoring never started** (`to === "in-review"`) and **never
stopped** (`from === "in-review"`) — a card's PR either untracked, or
tracked forever with its buffered comments never drained.
- **Terminal cleanup never ran** (`to === "done" || "archived"`).
- **The wip → hold failure bookkeeping never recorded** (`from ===
"in-progress"`).

None of these throw. They just stop happening — which is why the census,
not a red test, is what found them.

## Remaining 2, deliberately not converted

`L1097` (`task.column === "in-progress"`) and `L1171` (`task.column !==
"in-review"`) sit outside the listener where `parked` is in scope. They
need their own resolution, and resolving per call there is a different
cost profile than one-per-event; I flagged rather than guessed, per the
fleet rule.

## Verification

- census: `scheduler.ts` 12 → 2
- `scheduler-workflow-cutover` + `scheduler` — 42 tests green
- `tsc --noEmit` on `@fusion/engine` clean; `pnpm lint` clean

Behaviour on an unresolvable workflow is unchanged: the widened helper
keeps the same fail-soft legacy defaults the narrow one had.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:33:28 -07:00
gsxdsm
c9516dbd09 fix(engine): resolve archiveStaleDoneTasks lane guards by role (fleet: self-healing 56→51) (#3047)
Fleet phase. Claimed **`packages/engine/src/self-healing.ts`** — the
largest cluster at **56 of 126** total sites. Verified unclaimed first:
no open PR touches the file and no active worktree held a branch on it.

## Census before / after

| | total | self-healing.ts |
|---|---|---|
| before | **126** | **56** |
| after | **121** | **51** |

`census --strict` exits 0; baseline re-recorded in this commit so the
retired allowances cannot be regrown into.

## What converted, and why each role

`archiveStaleDoneTasks` asked "has this card finished?" by comparing
column ids, so on a renamed board it treated every finished card as live
and archived nothing — the sweep was inert on exactly the boards this
program exists to support.

- **active-dependents scan** and **temp-worktree age gate** →
`TERMINAL_ROLES` (complete ∪ archived): both ask "is this card done
with, in any sense?"
- **staleness filter** → `complete` **alone**: this sweep *archives*
finished cards, so an already-archived card is not a candidate. Using
the terminal pair here would have made the sweep consider its own
output.

**Union, not per-task, deliberately.** Over-inclusion is free at these
sites because the per-card check still discards, and the union needs no
per-task workflow selection — the failure mode
`resolveWorkflowIrForTask` has, where a card with no recorded selection
silently resolves to the built-in board. Recorded in
`docs/solutions/workflow-learnings/project-union-versus-per-task-lanes.md`.

## The half-converted state is the interesting part

Converting only the first two guards made `archiveStaleDoneTasks`
**register as a converted sweep** — the existing ratchet suite grew from
**36 to 38 tests** — and it then failed for still carrying `t.column !==
"done"`.

That is the failure mode worth naming: a partial conversion is worse
than none, because the function now *looks* converted (it calls the
resolver, it reads as role-aware) while one guard still pins it to the
legacy vocabulary. Finishing the function turned it green. I would not
have caught it from the diff.

## Verification

- `self-healing` suites — **807 pass** (41 files)
- `tsc --noEmit` — **0 errors**
- `census --strict`, `check:lane-wiring`, `check:fnxc-future-dates`,
`check:inert-flag-seams`, `check:sql-column-literals` — all exit 0

## Flagged, not guessed — the remaining 51

Deliberately left, each for a stated reason rather than an omission:

1. **Move-transition matrices** (~1489–1504): `from`/`to` pairs encoding
a legal-transition graph (`in-progress → todo|in-review|done|archived`).
These are the *shape* of the lifecycle, not a lane lookup; converting
them needs a transition-role model that does not exist yet. Guessing
here would encode a wrong graph.
2. **`getLiveTaskColumn` comparisons** (~1398, 5313–5342): compared
against a normalizing accessor that manufactures `"archived"` for
soft-deleted rows. Those are protocol values, not column ids —
converting them changes what the sentinel means.
3. **Sites without store access** in scope (several module-level
predicates): need the resolved set threaded in as a parameter, which is
a seam change per call site, not a substitution.
4. **`todo` requeue targets** (1927, 6181, 6303, 12060–12064): these
pick a destination, so they want the single `intake`/`hold` answer from
`resolveLifecycleColumns`, not a set — different arity, and several are
inside sweeps whose rebound semantics I would be changing rather than
preserving.

Each is a real conversion; none is a one-line substitution, and doing
them blind is how a guard count drops while behaviour gets worse.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:28:01 -07:00
gsxdsm
eb0ee4ae98 fleet: executor.ts 7 → 4 lifecycle-column guards (3 converted, 4 flagged out of scope) (#3048)
Claiming `packages/engine/src/executor.ts` from the census work order.

## Census before/after

| File | Before | After |
|---|---:|---:|
| `packages/engine/src/executor.ts` | 7 | **4** |

Measured with `scripts/lifecycle-column-census.mjs` (kind `column`
only), not grep.

## Converted (3)

**L17258 — the completed-task watchdog never armed on a renamed board.**
It required the card to sit in a literal `in-progress`. This does not
error; the watchdog simply never fires, which is the silent-guard class
this program exists to remove. The branch immediately above already
resolves the same lane through `resolveWipTargetForTask`, and there is
even an FNXC note there saying `latestColumn` must come from that
resolved value — so the comparison now asks the same resolver rather
than an id.

**L14940 (×2) — the duplicate-handoff finalize never ran on a renamed
review lane.** `fromColumn`/`toColumn` are parsed out of the store's
rejection message (`Invalid transition: 'X' → 'Y'`), so they carry
whatever ids that workflow declares. Comparing them to the literal
`in-review` meant a renamed lane never matched and
`finalizeAlreadyReviewedTask` was skipped, leaving the card
mid-transition with nothing to complete it. Now resolves the task's own
review role, falling back to the legacy literal when the workflow cannot
be read — so behaviour is unchanged wherever the vocabulary is
unreadable.

## Flagged, not converted (4) — per the fleet rule that behavior changes
are out of scope

**L3557 / L3581 / L3632 / L3642** are branch conditions inside the
**synchronous** `store.on("task:moved")` listener. Resolving a task's
workflow requires an `await`, which is not available in a sync
listener's condition. Moving the test into the deferred body would widen
the branch to every non-forward move and then re-narrow it — a
**behaviour change to the planning-evacuation path**, not a vocabulary
conversion. Converting them properly means making the listener async,
which wants its own commit and its own test.

I flagged rather than guessed, which is why this is 7 → 4 and not 7 → 0.

## On test coverage, stated plainly

Both converted sites are pure resolver swaps in `async` contexts,
verified by tsc, the census delta, and the existing executor suites (48
tests green). I did **not** add new fixtures: this is the file where I
twice wrote tests that passed against the *unconverted* code —
`recoverCompletedTask`'s seven early-return guards make negative
assertions succeed trivially — and reverted both times rather than claim
coverage I did not have. A fixture that genuinely drives L17258 needs a
satisfied `workflowStepResults` so the run does not divert into graph
re-entry; that is worth doing, and it is worth doing honestly rather
than as a green-looking placeholder.

## Verification

- census: `executor.ts` 7 → 4
- `tsc --noEmit` on `@fusion/engine` clean; `pnpm lint` clean
- `executor-graph-boundary`, `executor-task-done-summary`,
`executor-triage-column-audit`, `executor-step-session` — 48 tests green

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 02:27:49 -07:00
gsxdsm
d4add985fe test(engine): the unwired-parameter guard has been red on main — its list is stale by one (#3033)
## The unwired-parameter guard has been red on `main`

```
A parameter LEAVING this list is the goal; one arriving is a regression — update the list only to shorten it.
  expected [ …(16) ] to deeply equal [ …(17) ]
```

Reproduced on clean `origin/main`, so it is not a branch artifact.

`packages/engine/src/scheduler.ts isWipColumn` is **supplied at both
production call sites now** — `self-healing.ts:4754` and `:5825` pass
`isWipColumn: completedWipColumns.has(blocker.column)` and the
blocked-lane equivalent, wired by #2975/#2987. The list was not
shortened in the same change.

So this is the good direction: a parameter got wired. The assertion just
was not told.

## Shortened, not re-recorded

I diffed the computed set against the recorded one rather than
regenerating:

```
DEPARTED (wired since):  packages/engine/src/scheduler.ts isWipColumn
ARRIVED (new):           (none)
```

Exactly one departure, nothing arrived — so removing that single line is
the whole fix, and it follows the file's own instruction (*"update the
list only to shorten it"*). Re-recording wholesale would have silently
absorbed any arrival too, which is the one thing this ratchet must not
do.

## How it was found

While measuring an unrelated change to the sibling census. **This suite
is outside the merge gate**, which is why a red assertion sat unnoticed
— the same reason #2969's 15 red agent-action tests survived, and worth
noting as a pattern rather than a one-off.

## What I abandoned to get here, and why it belongs in this PR's story

I was trying to remove two false positives from the sibling
`check-lane-wiring` census — `bucketForTask(task: TaskItem)` and
`otherBucketSecondaryLabel(task: TaskItem)`, both flagged only because
`TaskItem` declares `columnFlags?`, both reading it off the entity
internally.

The rule I tried was the sibling guard's own documented one: a
**required** parameter is enforced by the compiler, so it is not this
census's question. It measured perfectly — 15 sites → 13, removing
exactly those two and retaining every genuine entry.

Then it failed `lane-wiring-census-named-types.test.ts`:

```ts
export type MergeContext = { completeColumns?: ReadonlySet<string> };
export function canMerge(task: string, context: MergeContext): string { … }
```

A **required** parameter with a named options type is a shape that
census deliberately covers — `canMerge(task, {})` really can omit the
lane member. My "exact" rule was exact only against the current tree,
and it broke a tested contract. I dropped it rather than edit their test
to match my change.

The two false positives therefore stay baselined, and the cost stands as
previously recorded: a genuinely new unwired call in those two TUI files
would be masked. I do not have a rule I can prove safe, and three
attempts at this class have now traded false positives for worse false
negatives.

## Verification (measured)

- both guard suites — **18 passed / 0 failed** (was 1 failed)
- `check-lane-wiring`, `lifecycle-column-census --strict`,
`check-fnxc-future-dates` — green
- `eslint` — clean (one pre-existing warning, no errors)

Test-only; no product file touched. No changeset.
2026-07-31 01:44:11 -07:00
gsxdsm
fb53a96eaa test(engine): the lease-seam alarm fired downward — re-point it, and close two ways it could pass without a fix (#2987)
## What happened

Both of my source-level audits went red on the advance that landed
#2975. They are exact counters, not floors, so this is the alarm working
**downward** — the direction it was written for. #2975 converted the two
self-healing `shouldHoldActiveFileScopeLease` call sites and closed the
2-of-4 seam that
`workflow-file-scope-lease-caller-gap-live-e2e.pg.test.ts` was
measuring.

I judged the conversion real before updating anything. Both sites now
derive their answers from `resolveProjectColumnsForRoles(...)` sets the
sweep had already resolved a few lines above (`self-healing.ts:4754`,
`:5825`) — trait membership, not a literal. So the numbers moved on
purpose.

## The part that is not bookkeeping

Re-pointing an audit to whatever the code now says is how a guard goes
dead. Both assertions could have been satisfied by something that is
**not** a fix, so both were tightened:

| Way it could pass without a fix | Old assertion | Now |
|---|---|---|
| `isWipColumn: true` hardcoded at a self-healing site — the original
defect wearing the converted call shape, answering "yes" for a blocker
resting anywhere | only checked the key was *absent* | requires
resolved-set membership: `/isWipColumn:\s*\w+\.has\(\w+\.column\)/` |
| a site answering one of the two independent role questions and not the
other | `includes(a) \|\| includes(b)` counted it as converted | `&&` |

The scheduler's own two sites *do* pass literal `true`, correctly — they
have already filtered to a role-resolved bucket, so there the answer is
a fact about the loop, not about the card. The form check is scoped to
`self-healing.ts` for that reason.

## Mutation evidence

Not reasoned — measured. Each mutant applied to `self-healing.ts`, suite
re-run, file restored:

| Mutant | Result |
|---|---|
| baseline | 8 passed |
| M1 — hardcode `isWipColumn: true` at one site | **1 failed** |
| M2 — drop `isReviewColumn` at one site (half-converted) | **2 failed**
|
| M3 — revert both sites to pre-#2975 | **2 failed** |

M2 is the one that justifies the `&&`: re-running it with the counter
reverted to `||` leaves the file **green (4/4 passing)**. The old
counter provably could not distinguish a closed seam from a half-closed
one — the exact blind spot this file exists to remove.

## Verification

`test:gate` exit 0 · live-PG E2E surface **171/171** · lifecycle-column
census exit 0 · FNXC date ratchet exit 0 · `pnpm lint` clean. Production
files untouched (`git status` clean on `self-healing.ts` after every
mutant).

## Scope

The other measured seam, `evaluateParkedAgentTaskLink`, is **unchanged
at 2-of-6** — four callers still omit the resolved columns, so the class
is not closed, only one of its two instances is. That remains
characterized, not fixed, in the same file; converting those four is the
capacity worker's file, not mine.

The call-site facts are still asserted against source text rather than
driven through the self-healing sweep, which would need the full
dependency-lease reconcile harness. That limit was stated in the
original file and still is.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Updated end-to-end workflow checks to verify resolved workflow-role
values at active lease call sites.
  * Strengthened assertions for WIP and review column detection.
* Updated audit coverage to reflect conversion of all active-file-scope
lease callers.
  * Preserved tracking for the remaining parked-link integration points.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 23:38:10 -07:00
gsxdsm
16921fc518 fix(engine,core): role resolution was half-done in two shared lifecycle predicates (surfacing family + file-scope leases) (#2975)
The three surfacing sweeps stopped reporting anything for a card resting
in a board's **second** review or hold column.

A lifecycle role is a **trait**, and any number of columns may carry it.
The shared runner resolved it with `resolveLifecycleColumns()[role]` —
**first match** — then gated on it:

```ts
const roleColumn = lifecycle?.[spec.role];        // FIRST column carrying the trait
if (task.column !== resolved.roleColumn) continue; // everything else dropped
```

A workflow that splits human sign-off from the merge lane has two review
columns; one that parks dependency-blocked cards separately has two hold
columns. Cards in the second got **no stale-paused-todo, no
stale-paused-review, no in-review-stalled** diagnostic — silently, with
no error, on all three sweeps at once.

## The second bug hiding inside the fix for the first

Resolving membership but still reading `roleColumns[0]`'s declared
`recovery` applies the **merge lane's** threshold to a card sitting in
the **sign-off** lane. Each card's policy now comes from its own column,
and one of the new cases fails if it doesn't: the first role column
declares a policy that suppresses the signal, the card's own column
declares one that fires.

## Reverted

| | |
|---|---|
| **6 of 12** new cases fail | `fires for a card in the SECOND column
carrying its role` and `reads the recovery policy of the card's OWN role
column` — × 3 sweeps |
| the other 6 pass either way | non-regression halves: still fires for
the FIRST role column, still does **not** fire for a card outside every
role column. Membership must widen the gate, not move it. |

The pre-existing 45 cases were all green throughout — the
single-role-column fixture could not express the case, which is why the
table-driven file that exists to stop these three sweeps drifting apart
never caught it.

## Verification

`pnpm test:gate` 161 + 13 + 487 + 71 · surfacing family 57 · core
stale-paused 20 · lint · census `--strict` · sql-literals · fnxc-dates ·
lane-wiring · changesets — all green.

## Note

`holdColumns` was missing from the lane-wiring vocabulary, so the gate
could not see that argument dropped. Added in the same commit.

While reviewing, I found and measured **two problems in #2974** (comment
posted there): six of its newly-visible sites are `satisfies`-wrapped
false positives, and baselining them means deleting a real
`reviewColumns` argument keeps the count unchanged and the gate green;
and its baseline predates #2970, re-opening the slot that PR closed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved stale-card detection across all applicable review and hold
columns.
* Cards are now surfaced using the policies configured for their
specific lifecycle column.
* Cards outside matching lifecycle columns are no longer incorrectly
surfaced.
* Preserved existing fallback behavior when no lifecycle columns are
configured.

* **Tests**
  * Added coverage for workflows with split review and hold columns.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

## Second commit: the same predicate, half-converted
(`shouldHoldActiveFileScopeLease`)

Folded in here rather than stacked — same file, same class, and a
stacked PR on an unmerged base is not mergeable. Reversible; say the
word and I'll split it.

`shouldHoldActiveFileScopeLease` is the **scheduler's** lease predicate,
shared with the self-healing repair paths deliberately so the two cannot
disagree about who holds a file-scope lease. Its two role answers are
optional parameters defaulting to the legacy ids. The scheduler's own
call sites were converted to pass resolved answers; self-healing's two
were not:

```ts
const isWipColumn    = options?.isWipColumn    ?? task.column === "in-progress";
const isReviewColumn = options?.isReviewColumn ?? task.column === "in-review";
```

On a renamed board neither branch matches, so the predicate returns
`false` for every card. The scheduler kept the lease; self-healing saw
none, cleared `overlapBlockedBy`, and **released a dependent to edit
files another agent still holds** — the outcome `groupOverlappingFiles`
exists to prevent.

Membership comes from the wip/review sets each sweep already resolved a
few lines above, so this adds no reads.

**Reverted:** both new cases fail with `overlapBlockedBy` = `null` — the
release itself, not a proxy. The pre-existing legacy-column case in the
same file passes either way, because `in-progress` satisfies the literal
default; that is exactly why it never caught this.

Lane-wiring baseline re-recorded `9 -> 7` in the same commit (the
ratchet refused a stale allowance, as intended).

**Verification:** gate 161 + 13 + 487 + 71 · surfacing 57 · overlap-seam
+ scheduler-lease + query-blindness 79 · core stale-paused 20 · lint ·
census `--strict` · sql-literals · fnxc-dates · changesets — green.
2026-07-30 23:13:34 -07:00
gsxdsm
41af5e5dbd fix(gate): the lane-wiring census could not see its own motivating case (#2956) (#2974)
#2966 shipped a gate that **cannot detect the defect named first in its
own header.**

`findLaneAcceptingFunctions` matched a lane parameter only when
`param.type` was a `TypeLiteralNode` — an inline `{ reviewColumns?: …
}`. But the real code declares these as interfaces:

```ts
export function getInReviewStallReason(
  task: Pick<Task, …>,
  context: InReviewStallContext = {},   // TypeReference — invisible
): InReviewStallSignal | undefined
```

so the function never entered `accepting` and none of its call sites
were examined.

### Measured, both directions

| | before | after |
|---|---|---|
| lane-accepting functions detected | 20 | **30** |
| `getInReviewStallReason` detected | no | **yes** |
| re-introduce #2956 (drop `reviewColumns` from one call site) | `none
added` — **passes** | **fails**: `reads.ts: 7 unwired now, baseline
allows 6` |

The gate now catches the thing it was built for.

### The baseline moves 10 → 24, and that number needs context

`10 unwired call site(s) across 8 files` → `24 across 15`. **No entry
was removed** — every previously-recorded file kept its count and 14
sites became visible for the first time:

```
core/task-store/reads.ts                        0 -> 6
engine/self-healing.ts                          2 -> 4
core/task-store/branch-and-pr-entities.ts       0 -> 1
core/task-store/task-update.ts                  0 -> 1
engine/scheduler.ts                             0 -> 1
dashboard/routes/register-task-workflow-routes  0 -> 1
cli/commands/dashboard-tui/bucket-mapping.ts    0 -> 1
cli/extension.ts                                0 -> 1
```

**These are newly VISIBLE, not newly broken** — they have been unwired
all along. I have **not** audited them, and recording them in the
baseline is not a claim that they are fine; it is the ratchet doing what
its header describes, since the census's own note says roughly half of
the original hits were legitimately unwired (identity proven by a
stronger means, sentinel columns, dead exports). Someone should walk the
14. Two stand out as worth a look first: **`reads.ts` at 6** is the file
#2956 was about, and **`scheduler.ts`** is a dispatch path.

Flagging rather than fixing, because wiring a call site that should not
be wired is its own defect and each needs the judgement call the census
header describes.

### Regression test

`packages/engine/src/__tests__/lane-wiring-census-named-types.test.ts`
pins the detector's shape — named interface, type alias, inline literal,
positional — against fixtures rather than live counts, so it does not
churn when someone legitimately wires a call site. Plus one anti-vacuity
case asserting the named-type arm is still load-bearing on real source
(`getInReviewStallReason` resolves in the live tree), so the fixtures
cannot pass while the tool has quietly stopped applying here.

**Mutation:** removing the `TypeReference` arm fails **4 of 5**.

### Also worth knowing

`findLaneAcceptingFunctions` still only visits
`ts.isFunctionDeclaration` at top level, so `export const fn = (ctx) =>
…` remains invisible. I checked — no exported arrow function currently
takes a lane argument, so nothing is missed today, and I left it rather
than widen the surface in the same change.

Resolved by **name across the corpus** instead of a type-checker
`Program`: these are plain source scans and a checker would cost a full
type-resolution pass for one lookup. Two same-named types merge, which
only ever widens what counts as wired — safe for a ratchet.

**Verified:** 5/5 new tests, `check-lane-wiring` clean at the new
baseline, lint clean, FNXC gate exit 0. Core suite on main is green
(4923 passed / 0 failed) — unrelated, but I had it running.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Improved lane-wiring analysis to recognize named interfaces and type
aliases.
* Added support for wrapped configuration expressions and positional
parameters when detecting lane information.

* **Tests**
* Added comprehensive coverage for lane-wiring detection, including
named contexts and live-tree validation.

* **Chores**
* Updated baseline counts to reflect newly recognized application areas
and improved self-healing detection.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 22:54:31 -07:00
gsxdsm
126cee7e6d engine: finalization parked ALREADY-MERGED work as failed on a renamed board (#2964)
**The worst symptom in this family: the branch landed, and the board
says the task failed.**

`project-engine`'s merge-confirmed finalization spread the task's
**real** column into `getTaskHardMergeBlocker` with no `reviewColumns`,
so the identity check ran against the literal `in-review`. On a renamed
board it returned `task is in 'signoff', must be in 'in-review'`, and
the caller parked the card:

```
status: "failed"
error:  "Merge confirmed but finalization blocked: task is in 'signoff', must be in 'in-review'"
```

For work that had already merged.

## Its sibling had already solved this

`auto-merge-finalization.ts` passes the **review-eligible sentinel**
instead of the card's own column, with the reasoning recorded at that
site: `getTaskHardMergeBlocker` asks *"is this card blocked by anything
other than where it sits?"*, and its callers are recovery paths for
landed work that a graph crash can leave resting in any column.
`project-engine` simply never got the same treatment.

## One name instead of two spellings

Rather than write the sentinel a second time, it is exported once as
`REVIEW_ELIGIBLE_SENTINEL_COLUMN` next to the helper whose contract
gives it meaning, and both recovery paths use it. **Two sites
independently spelling a magic value is how one of them came to be
missing it** — that is the actual root cause here, not the literal
itself.

This also answers the census, which flagged the new literal — correctly.
Its guidance (which I wrote, in #2909) is to hoist a deliberate literal
into a *declaration*, where a `DELIBERATE-LITERAL` marker actually
attaches, instead of leaving it mid-expression where the marker is
silently ignored. The shared constant is exactly that, and it lowers
`auto-merge-finalization`'s literal count too.

## Revert result

| | reverted → |
| --- | --- |
| sentinel replaced by the card's own renamed column | reproduces the
shipped string |

The middle test asserts that string deliberately — it is what landed in
`task.error`, so a regression reports what the operator would actually
have seen. A third case checks the sentinel does **not** suppress
genuine blockers: incomplete steps still block finalization in any lane.

These drive the helper directly; reaching `project-engine`'s
finalization end to end needs a live engine, a merge run and a real
repo, while the defect is entirely in *what the blocker is asked*.

## Verification

`pnpm test:gate` 161 + 487 + 13 + 71; `project-engine` +
`auto-merge-finalization` + the new suite, 207; `tsc` clean on core and
engine; lint, census `--strict`, FNXC gate, changesets all clean.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed merge-confirmed tasks being finalized correctly when boards use
renamed workflow columns.
* Prevented already-merged tasks from being incorrectly marked as failed
due to custom review-column names.
  * Preserved enforcement of genuine incomplete-step blockers.

* **Tests**
* Added coverage for finalization on renamed lanes and legitimate merge
blockers.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 22:06:24 -07:00
gsxdsm
8e0219d573 fix(merger-ai): gate the no-commits dep-sync skip on the branch diff — the P1 #2501 shipped without (#2958)
## #2501 merged without its P1 fix; this is that fix, alone

#2501 has landed. Its review threads were resolved — I judged and fixed
them — but its head was a fork branch I could not push to, so **the
fixes were never in it**. Confirmed on `main` at `c1c1b964af`:

```
merger-ai.ts:902:      if (ctx.noCommitsExpected === true) {      ← bare flag, no diff gate
merge-dependency-sync.ts: export const LOCKFILE_CANDIDATES → 0 matches
```

Rebasing dropped this PR's five duplicated base commits, so it is now
**one commit**: the review fix and its regression.

## The defect on main

The dep-sync skip trusts `ctx.noCommitsExpected` alone, and **only ever
runs on a branch that has commits** — the `rev-list --count`
short-circuit ~50 lines above returns `outcome: "empty"` at zero ahead,
so control reaches it only when the branch is AHEAD.

Nothing revalidates the flag. Both downstream empty-lane guards carve
no-commits tasks out explicitly — `merger-ai.ts:1372` (#2259
already-landed proof) and `:1994` (FN-8141 executor veto) — and both
guard the *opposite* direction: commit-expected task, empty branch. The
inverse has no check.

So a task marked no-commits whose executor committed a manifest or
lockfile change gets its dependency install **and** its frozen-lockfile
validation skipped, and the change lands unvalidated.

## The fix

The flag says *look*; the branch diff decides. A `main...branch` diff
touching `package.json` or any `LOCKFILE_CANDIDATES` entry falls through
to the normal sync and emits an audit row with `skipOverridden: true`.
An unreadable diff **also** syncs — matching the hard-fail contract
documented directly above that block, rather than treating absence of
evidence as evidence of safety.

`LOCKFILE_CANDIDATES` is exported instead of duplicated, so the skip and
the installer cannot drift on what counts as a dependency change.

**Mutation-verified:** reverting to trust-the-flag fails exactly the new
case and nothing else. The existing *"lands successfully with
noCommitsExpected: true and actual changes"* case is untouched and still
passes — `feature.txt` is not a dependency file, so an ordinary source
change on a no-commits task still skips. The new case differs only in
*which* file the branch touches.

## Also carried over from the #2501 review

**coderabbit's env nit** — `process.env.X = undefined` stores the string
`"undefined"`, leaving a previously-absent var truthy and leaking into
later tests. `restoreEnv` applied at both sites.

**Both entry paths** — deferred with reasons:
`runAiMerge`/`landWorkspaceTask` sit behind real worktrees, sessions and
a merge agent, and the cheap version is a mirrored-implementation test
that cannot fail on a revert (this repo has deleted two of those). The
fix above also means propagation is no longer the only thing between a
stale flag and an unvalidated lockfile.

## A correction to my own work

My first version of the regression committed the lockfile while the
fixture had left the tree on `main`, so the `main...branch` diff could
not see it and the case **passed for the wrong reason**. Corrected, with
the reason recorded in the test.

## Verification

- `merger-ai-no-commits-deps-skip` — **5/5**, mutation-verified
- `merge-dependency-sync-lockfile-heal` — **10/10**
- engine typecheck — clean
- `pnpm lint` — clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 22:05:04 -07:00
Drew Donaldson
c1c1b964af fix(dashboard): expose full permission-mapped task toolset in chat sessions (#2376)
## Bug
Chat-session tool surface missing task-mutation tools that exist outside
of chat, even when the agent's permission record grants them.\n\nRepro:
agent-09dcf8b2 (role: custom, CEO) in NextGenEHS has tasks:archive /
tasks:delete / tasks:merge / tasks:retry / tasks:update true in its
permission record with permissionPolicy.presetId = unrestricted and
task_agent_mutation = allow. Calling fn_task_archive / fn_task_delete /
fn_task_merge in chat returns: Tool fn_task_* not found.\n\nRoot cause:
packages/dashboard/src/chat.ts createChatFusionToolset() built a
hardcoded narrow chat-only allowlist while heartbeat registered the
complete lifecycle surface unconditionally.\n\nFix:\n- Add exported
factories in packages/engine/src/agent-tools.ts for missing lifecycle
tools: fn_task_archive, fn_task_unarchive, fn_task_delete,
fn_task_retry, fn_task_pause, fn_task_unpause, fn_task_duplicate,
fn_task_merge, fn_task_update, fn_task_add_dep, fn_task_promote,
fn_trait_list, fn_ask_question, fn_reflect_on_performance,
fn_read_evaluations, fn_update_identity, fn_send_message,
fn_read_messages.\n- Wire those factories into
createChatFusionToolset(). Mission/ideation mutations stay behind
missionMutationGated. Agent-scoped tools still require agentId.\n-
Re-export from packages/engine/src/index.ts.\n- Regression test:
packages/dashboard/src/__tests__/chat-toolset-permissions.test.ts (3/3
passing). Existing chat.test.ts (14/14 passing).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Chat now exposes task lifecycle actions—including archive, retry,
pause, duplicate, merge, and dependency updates—when permitted by the
agent’s action controls.
- Added support for identity updates and evaluation viewing in
agent-linked chats.
- Existing read-only tools remain available, while restricted actions
stay hidden when authorization is unavailable.
- **Tests**
- Added regression coverage for authorized and unauthorized chat tool
surfaces, including preservation of read-only capabilities.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 21:51:09 -07:00
gsxdsm
189f237a07 fix(security): make the approval and permission gates actually enforce (#2457)
## Why

An AI agent autonomously deleted a live task (FN-8609) via
`fn_task_delete`. Investigating that led to two audits, which found the
gate machinery is well-built and its *enforcement* well-tested — but the
*decision boundary* was unguarded and the tools that matter never
reached the gate at all.

The headline defect: all ~76 tools in `packages/cli/src/extension.ts`
are delivered through `hostExtensionPaths` → `DefaultResourceLoader`,
which never passes through `wrapToolsWithActionGate`. `fn_task_delete`
*was* classified as gated in `gating-classifications.ts` — that
classification was dead code. Roughly 15 destructive tools were ungated
even for an operator who had locked their policy all the way down.

Second: `POST /api/approvals/:id/decision` took the approver identity
from the request body, validated for shape only, with no check that the
decider was human or differed from the requester. Combined with a bash
tool whose command string was never inspected (only `params.cwd`), an
agent could read the daemon token from `~/.fusion/settings.json`, find
its own pending request, and approve it.

## Scope

This does **not** tighten defaults. The `unrestricted` preset is
byte-for-byte unchanged — `git diff` on `agent-permission-policy.ts` is
empty — and regression tests assert that an out-of-the-box install
behaves exactly as before. The bug was never "the default is too
permissive"; it was "strict policy doesn't enforce." This makes turning
security up actually work.

The one deliberate exception: the containment that stops an agent
escalating its *own* privileges (reading the daemon token / credentials,
calling the approvals API to self-approve) applies at every preset
including `unrestricted`. That is a privilege-escalation boundary rather
than a permission preference — if it only engaged under strict policy it
would not have prevented the incident that prompted this.

## What changed

8 bisectable commits:

- **Approval lifecycle** — self-approval blocked via server-derived
deciders; same-verdict replay 409s; decide re-reads and re-validates
inside the transaction; expiry TTLs; `markCompleted` ownership check;
session identity registry in core.
- **Engine gates enforce for real** — unclassified tools resolve to a
policy-governed category instead of hardcoded `allow`; missing-policy
fail-open closed; bash containment floor + exact-command approval
binding.
- **Dashboard decision routes** — stop trusting client-supplied actors
(decision, bypass-review, worktrunk → 403 on forged actors).
- **`fn serve` authenticated by default** — auto-mints a token following
the existing `fn dashboard` precedent; `--no-auth` opts out.
- **Sibling entry points closed** — user-sourced hard-cancel moves, ACP
execute-once approvals, plugin task-store gating.
- **pi-extension principal resolution** — the extension resolves the
acting principal and can withhold or policy-gate the previously ungated
destructive tools.
- **Root-cause bonus fix** — `findLatestByDedupeKey` was broken in
PostgreSQL backend mode (already-parsed jsonb fed through a string-only
parser), so approved-grant redemption **never matched in production**,
minting duplicate requests. This explains the live DB state of 17
approved / 0 completed. *(Also cherry-picked to `main` as `a9b30013bb`,
since it is an active production defect on its own.)*
- **Review follow-ups** (`627f1b1fa8`) — operator-configured
provisioning privilege and a configurable grant TTL; see below.

## Review follow-ups

**Provisioning privilege is operator-configured, not role-derived.**
`isCallerPrivileged` had gone from `caller.reportsTo == null` (every
top-level agent privileged — permanent escalation by creating a
manager-less agent) to `caller.role === "ceo"`, which swapped an
implicit rule for a magic string: any agent config can claim that role,
while an operator who genuinely wants a privileged agent had no
supported way to say so. Privilege now derives solely from
`agentProvisioning.trustedAgentIds` / `trustedRoles` and fails closed
when settings are unresolvable.

It is also no longer forwarded to `resolveAgentProvisioningPolicy` as
`isPrivileged`, because that flag short-circuits ahead of
`alwaysApproveDelete` — a trusted caller was bypassing delete approval
entirely. The policy applies the same trusted rules itself, in the right
order. The function now governs only the org-chart escape hatch (acting
outside your own direct reports).

**Grant TTL defaults to 1 hour and is configurable.** Approval →
redemption is not instantaneous: an operator approving from their phone,
an engine restart, a queued lane, or a task waiting on a worktree all
routinely exceeded 15 minutes, after which the grant expired and the
agent silently re-requested. One hour remains far short of the
"redeemable forever" hazard the TTL exists to bound. Override via
`FUSION_APPROVAL_GRANT_TTL_MS` or `configureApprovalRequestTtls()`;
invalid overrides are ignored rather than widening the window to
infinity or collapsing it to zero.

## Behavior changes requiring operator review before rollout

1. `fn serve` requires a bearer token by default (`--no-auth` opts out);
unauthenticated clients get 401.
2. Agents can no longer run withheld destructive tools
(`fn_task_delete`, `fn_task_bypass_review`,
mission/milestone/slice/feature/workflow deletes, `experiment_finalize`,
`skills_install`). Operators keep them via CLI/dashboard. **This is the
incident fix.**
3. Agents get provisioning privilege only when the operator lists them
in `agentProvisioning.trustedAgentIds` / `trustedRoles`; the
provisioning gate is now live in production. Previously-implicit
privilege (top-level position, or a `ceo` role) no longer grants
anything on its own.
4. Decision replay 409s (was 200); pending approvals expire after 24h,
approved grants after 1h (configurable); bash approvals bind per exact
command.
5. Forged/body actors on decision, bypass-review, worktrunk routes →
403; `archive-all-done` requires `{confirm:true}` (external scripts
affected).
6. `fn_secret_get` approvals grant exactly one reveal (previously
granted nothing and looped forever); ACP approvals are execute-once
(previously infinite reuse).
7. Bash containment denies token/credential/approvals-API commands in
all agent sessions at every preset.

## Verification

Independently re-run against the branch, not just self-reported:

- 5 typechecks (core, engine, cli, dashboard `tsconfig.json` +
`tsconfig.app.json`) — clean
- `pnpm lint` — clean
- `pnpm test:gate` — 379 passed
- `pnpm build --force` — green (a plain `pnpm build` skips packages as
unchanged and does **not** compile the branch)
- `pnpm check:changesets` — clean
- ~650 file-scoped tests including new negative-path suites for the
decision boundary, which previously had **zero** test coverage

`packages/engine/src/__tests__/plugin-runner.test.ts` fails 56/80 —
**verified pre-existing**, reproducing identically at base commit
`93a403af67` on `main`. Not in the merge gate.

### A mutation check that failed to fail

Worth recording, because it nearly shipped an untested security fix. The
first mutation check on the provisioning change reintroduced the `ceo`
hardcode and **all 17 tests still passed** — the tests asserted through
the policy path, which can no longer observe `isCallerPrivileged` at
all, precisely because `isPrivileged` is no longer forwarded there.
Org-chart cases that do exercise the function were added; the hardcode
now fails exactly 1 of 19, and restoring is green. A green mutation run
is only meaningful if the test can actually see the code under test.

## Known limitations (stated, not papered over)

- The bash containment floor is string-matching: a cost-raiser, not a
sandbox. Quoting, encoding, `$HOME`, symlinks, or an interpreter
one-liner can evade it. The durable protection is the decision route
refusing agent-originated deciders — the filter is the belt, not the
braces.
- Approval expiry is lazy (evaluated at decide/complete/redeem), not
swept, so an expired pending row stays visible in lists until touched.
- The extension's require-approval path returns a pending message but
cannot suspend a pi session mid-turn; engine-side pause hooks cover
engine lanes only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Security**
* Hardened approval and permission gating with server-side decider
attribution, self-approval blocking, ownership checks, replay/race
protection, and status/TTL enforcement.
* Added fail-closed behavior for sensitive/unclassified tools and
sandbox provisioning approvals.
* Blocked credential/approval access via bash containment; plugin
destructive task operations now require explicit permission.
* **New Features**
* `fn serve` now defaults to bearer-token auth, with `--no-auth` as the
explicit opt-out.
* **Bug Fixes**
* Improved task move-source attribution (`moveSource: "user"`) and
tightened dashboard archive/bypass confirmation and operator attribution
behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-30 21:50:37 -07:00
gsxdsm
7712e0ada2 engine: merging was broken outright on a board with a renamed review lane (#2963)
**Not a degraded message — no task on such a board could be merged at
all.**

`getTaskMergeBlocker`'s column-identity check *returns a blocker* when
the task's column is not a review lane. Both merge entry points called
it without `reviewColumns`, so the check ran against the literal
`in-review`:

```
Cannot merge FN-1: task is in 'signoff', must be in 'in-review'
```

`aiMergeTask` (`merger.ts`) and `runAiMerge` (`merger-ai.ts`) turn that
into a thrown error. Every merge on a renamed board fails, with a
message naming a column the board does not have.

## This exact defect was already found once

The helper's own FNXC comment records it, in `moves.ts`:

> *"so on a renamed board that move threw `Cannot move FN-1 to done:
task is in 'signoff', must be in 'in-review'` even though the transition
had just been validated as legal. A half-conversion, where the outer
question is resolved and the inner one is not."*

That fix added the `reviewColumns` option and wired `moves.ts`. **These
two callers were missed** — same shape, one layer out. A fix that adds
an optional parameter is only as good as the call-site sweep that
follows it.

## How it was found

By enumerating the call sites of every lane-taking helper, rather than
trusting the `unwired-lane-parameter` guard. That guard is deliberately
conservative — a mention of the parameter *anywhere* satisfies it — so
**partial** wiring is invisible to it, and `reviewColumns` is mentioned
plentifully elsewhere. This is the method #2956 used on a sibling
defect, applied to every seam I have touched.

## Two sites deliberately unchanged

- **`moves.ts`** passes `skipColumnIdentityCheck: true`. It has already
proven lane identity from resolved IR traits, so supplying lanes *as
well* would be contradictory rather than additive — the helper's comment
is explicit that the two options answer different questions.
- **`isTaskReadyForMerge`** has **zero** production callers. Adding a
parameter there is precisely the unwired-parameter anti-pattern this
program keeps removing.

## Revert result

| | reverted → |
| --- | --- |
| `reviewColumns` at either call | reproduces the shipped string exactly
|

The middle test pins that string deliberately: it is the
operator-visible failure, so if the wiring regresses the test says what
the operator would have seen. A third case checks that supplying lanes
does **not** switch the identity check off — a card in the wip lane is
still blocked, and the message names the resolved lanes rather than a
column the board lacks.

The cases drive `getTaskMergeBlocker` directly: reaching it through the
merge entry points needs a real repo, worktree and merge run, while the
defect is entirely in *which columns the blocker is asked about*. The
wiring itself is covered by tsc and the guard.

## Verification

`pnpm test:gate` 161 + 487 + 13 + 71; `merger` + `merger-ai` +
`self-healing` suites 461; `tsc` engine clean; lint, census `--strict`,
FNXC gate, changesets all clean.
2026-07-30 21:50:32 -07:00
ischindl
8d6acf1314 fix(RUFU-018): add noCommitsExpected dep-sync skip and corepack/pnpm env passthrough (#2501)
Manually land RUFU-018 fix bypassing the AI merge pipeline.

## Summary
- Add `noCommitsExpected` flag to `LandRepoContext`; skip dependency
sync when set
- Forward `COREPACK_HOME`/`PNPM_HOME`/`npm_config_registry` in
`installWorktreeDependencies`
- Add comprehensive tests for both changes

This unblocks all downstream RUFU audit tasks.

## Surface Enumeration
- Providers/bridges: `installWorktreeDependencies` called from
`landOneRepo` (AI merge) and legacy `merger.ts`; `landOneRepo` called
from `runAiMerge` and `landWorkspaceTask`
- Data states: `noCommitsExpected` can be `true`, `false`, or
`undefined` — both callers use `=== true` strict check

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Improved support for tasks that do not produce commits by skipping
unnecessary dependency installation during merges.
- Preserved normal merge and review behavior when dependency
installation is skipped.

- **Bug Fixes**
- Dependency installation now correctly preserves relevant
package-manager and system environment settings.
- Reduced installation failures caused by missing or unavailable
package-manager configuration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Fusion <noreply@runfusion.ai>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
2026-07-30 21:50:25 -07:00
gsxdsm
01f081e8aa engine: restore the stall-signal lane wiring #2951 dropped (and the test that proved it) (#2961)
**My defect, shipped in #2951 — and the same family as the one #2956
just fixed.** Found by auditing my own seams after that, not by a
failing check.

## What is on `main` right now

`surfaceInReviewStalls` reads the project's review columns (converted in
#2951), then calls `getInReviewStallReason` **without** `reviewColumns`.
The classifier falls back to the literal `in-review`, returns no signal
for a renamed-lane card, and the sweep surfaces nothing.

That is the textbook **missed pair** this program has a ratchet for: a
widened read handing every renamed-board card to a literal classifier.
The resolve work happens and is then discarded. On a renamed board an
operator sees no stall warnings at all.

#2951's conflict resolution dropped two things together:
- the per-card `stallLanes` map and the `reviewColumns` argument
- **the test that proved the wiring**

## Why nothing caught it

**A deleted test cannot fail.** I verified that rebase by comparing the
68 conflict *hunks* — stripping FNXC stamps, confirming 0 of 68 had real
content differences — and then ran the gate. The gate passed precisely
because the proving test had gone with the code it proved.

I verified the conflicts. I did not verify the outcome. Those are
different things, and the difference is invisible when the evidence
disappears alongside the feature.

The `unwired-lane-parameter` guard cannot catch this either, by design:
it is deliberately conservative — a mention of the parameter *anywhere*
satisfies it — so **partial** wiring is outside its reach.
`reviewColumns` is mentioned plenty in `reads.ts`, so the guard is green
while this call site goes unwired.

## How I found it

The check #2956 used on the sibling defect, applied to every lane seam I
have touched: enumerate each function's **call sites** and confirm each
one carries the parameter. That enumeration also flags several other
call sites without `reviewColumns`/lane arguments (`merger.ts`,
`moves.ts`, `auto-merge-finalization.ts`, `merger-ai.ts`,
`project-engine.ts`) — I have **not** touched those here; they need
per-site judgement about whether the lane answer is even available, and
that is a separate change rather than a sweep.

## Revert result

| | reverted → |
| --- | --- |
| `reviewColumns` at the call (i.e. exactly what #2951 shipped) | fails
the restored test |

## Verification

`pnpm test:gate` 161 + 487 + 13 + 71; blindness suite 71;
`self-healing.test.ts` 412; `tsc` engine clean; lint, census `--strict`,
FNXC gate, changesets all clean.
2026-07-30 21:45:17 -07:00
gsxdsm
fd795883c5 feat(missions): per-mission taskPrefix override for triaged task ids (#2347)
## Summary
Maintainer re-land of
[#2334](https://github.com/Runfusion/Fusion/pull/2334) (fork
`flexi767:feat/per-mission-task-prefix`) after resolving merge conflicts
with current `main`.

Fork push was unavailable despite `maintainerCanModify`, so this branch
carries the conflict resolution.

### Feature
- Optional per-mission `taskPrefix` for triaged task ids (inherits
project prefix when unset)
- Dashboard MissionManager + routes + store/triage plumbing
- Postgres migration for `project.missions.task_prefix`

### Conflict resolution
- Main claimed migration **0026** (bigint counters) and **0027**
(workflow IR pin)
- Mission task-prefix migration renumbered **0026 → 0028**
- Baseline `0000_initial.sql` includes `task_prefix` on missions
- `legacy.ts` keeps code-org re-exports; `missions.ts` carries
`taskPrefix` on create/update types

## Test plan
- [ ] CI green (lint/typecheck/build/gate)
- [ ] Create mission with custom prefix; triage feature → task ids use
that prefix
- [ ] Clear mission prefix via PATCH null; new tasks inherit project
prefix

Closes / supersedes #2334 once this lands (or re-point the fork PR).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Missions can now set an optional per-mission task ID prefix
(overriding the project default).
* Added task prefix support to mission create/edit UI and dashboard
APIs, including normalized uppercase values and validation.
* **Bug Fixes**
* Improved commit hook generation for custom prefixes and special
characters, with safer shell handling to prevent unsafe interpretation.
* **Chores**
* Added PostgreSQL migration and schema-applier support to persist and
propagate mission task prefixes, including upgrade/backfill coverage.
* **Tests**
* Added backend and UI/API test coverage for task-prefix creation,
clearing, and ID minting behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 21:35:23 -07:00
gsxdsm
e6e70a2562 test(engine): delete two temp-cleanup mechanisms guarding a leak the harness already prevents (#2960)
`scheduler-paused-dispatch-refusal.test.ts` carried **two** tracking
arrays and **two** `afterEach` hooks, both collecting the same
`mkdtempSync` path and removing it twice. One was added per review round
on #2779 — I wrote both, and neither round noticed the other.

The obvious fix is to merge them into one. **I checked whether the leak
was real first, and it isn't.**

### Measured

`packages/core/src/__test-utils__/vitest-setup.ts` **redirects
`os.tmpdir()`** to a per-worker sink and sweeps it by owning pid. So
`tmpdir()` inside a test does not resolve to the real temp root at all.
Probing the paths this file actually creates:

```
/var/folders/.../T/fusion-test-workers-8Tv8um/redir-5845/fusion-paused-dispatch-ZUhUVL
```

| run | fixtures created | left behind |
|---|---|---|
| cleanup as shipped | 4 | 0 |
| **cleanup disabled** | 4 | **0** |

The sink is reclaimed either way. Both mechanisms were appeasing a
review comment about a problem that could not occur.

### Why deleted rather than merged

A cleanup that cannot be observed to clean anything is not a cheap
safety net — it is a claim the file cannot back, and it misreports which
layer owns temp lifetime. Keeping one "just in case" would leave the
next reader believing this file manages its own fixtures. If the
redirect is ever removed, cleanup belongs in the shared setup for
**every** test, not re-added file by file. An FNXC note records the
measurement and says exactly that, so a third round doesn't re-add a
third copy.

### A note on my own measurement

My first check was `ls $TMPDIR/fusion-paused-dispatch-*` before and
after — it reported zero leaked with cleanup **on**, which I nearly took
as "cleanup works." It also reported zero with cleanup **off**. That
contradiction is the only reason I looked further; the glob was
measuring a directory the fixtures never reach. The before/after count
would have "confirmed" a working cleanup just as readily as a redundant
one.

**Verified:** 4/4 pass, `tsc` 0 errors, lint clean, FNXC gate exit 0.
Test-only, no product change, no changeset.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 21:29:39 -07:00
gsxdsm
8e5e1147d2 core,engine: the last literal lifecycle query — and the three stall signals that disagreed (#2951)
**This is the last one.** `surfaceInReviewStalls` was the final literal
`listTasks({ column })` in production — I verified it by direct scan,
not by census arithmetic: **1 remaining before this, 0 after.**

It tells an operator that a card is stalled in review. On a renamed
board the stall was real and the board simply never said so.

## It came last on purpose

Converting the read alone would have been **worse than leaving it**.
`getInReviewStallReason` gated on the literal `in-review` itself, so a
widened read hands every renamed-board card to a classifier that drops
it — the missed-pair class, wearing the shape of a clean one-line
conversion.

## What was actually there

Three sibling signals decorate the same row, and they **disagreed about
which lane it is in**:

| signal | before |
| --- | --- |
| `getInReviewStalledSignal` | singular `reviewColumn` — resolved, but
**first-per-role** |
| `getStalePausedReviewSignal` | singular `reviewColumn` — same |
| `getInReviewStallReason` | **no seam at all** — literal |

So one row could be judged in-review by one signal and not by another.
And the singular ones are the **arity trap**:
`resolveLifecycleColumns().review` is the *first* column carrying a
review role, so a board with a separate merge lane beside its
human-review lane had a second review column matching none of them.

All three now take `reviewColumns` (membership), resolved **once per
row** through `resolveReviewColumns` — the union of the three review
roles — so they cannot disagree by construction. The singular/literal
paths remain as the no-metadata fallback, so a caller passing nothing is
byte-identical to today. Ten call sites in `reads.ts` wired from that
one answer; the singular resolver is deleted.

## Revert results

Each applied alone and re-run:

| conversion | reverted → |
| --- | --- |
| the resolved read | fails — the card is never listed |
| `reviewColumns` at the call | fails — the classifier drops the renamed
card the widened read just found |

That second row is the whole point: it proves the pair had to move
together, which is the thing I got wrong twice earlier in this series.

## Second commit: a red on `main`, not from this branch

`check-fnxc-future-dates` landed and **`main` fails it** — verified by
running the script on a clean `origin/main` checkout rather than
inferring. Nine files carry stamps dated after today, so every worker's
gate fails on a check none of their changes caused. Several are mine: I
had been stamping tomorrow's date across this whole series, which is
precisely the out-of-order record the check exists to prevent.

Scope held deliberately: a repo-wide sweep touched **266 files** across
docs, scripts and every package. I ran it, backed it out, and limited
this to the nine files the check actually flags — a mechanical rewrite
that size during a queue freeze would conflict with every in-flight
branch, which is worse than the red it fixes.

## Verification

`pnpm test:gate` 161 + 487 + 13 + 71 (green **only** with the stamp
commit); `@fusion/core` full suite **4810 passed**; engine self-healing
+ blindness + both ratchets **758 passed**; `tsc` clean on core and
engine; `pnpm lint`, `check:changesets`, `lifecycle-column-census
--strict`, `check-sql-column-literals` and `check-fnxc-future-dates` all
clean.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Review-stall detection now recognizes renamed and multiple review
columns while retaining support for the legacy review column.
  - Paused tasks continue to be excluded from stall detection.
- Self-healing review-stall sweeps now search all configured review
lanes and avoid duplicate task results.

- **Tests**
- Added regression coverage for renamed and legacy review-lane queries.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-30 20:30:18 -07:00