Commit Graph

2600 Commits

Author SHA1 Message Date
gsxdsm
89284df85e E2E: table-driven converted-sweep coverage, two new sites, and an honest unproven-sites ledger (#2485)
Follow-up to #2475 (merged). Test-only, plus one test-utility seam.

## Why a table

#2475 proved one converted sweep. The count has since gone to **three**,
twice while this work was open — `surfaceStalePausedTodos` appeared
during #2475's review, and #2478 landed `recovery-reconciler.ts` while
this branch was open. A suite with a bespoke `describe` per sweep is a
coverage claim that quietly becomes false.

Replaced with a table of `(seed, run, acted, roles, observability)`. The
driver derives four assertions per entry:

| | positive | negative |
|---|---|---|
| **renamed vocabulary** | acts on the card | inert in a non-target
column |
| **default vocabulary** | acts (regression floor) | inert |

Adding a converted sweep is **one entry** — the #2478 site proved that
in practice, not in principle. `actsOnRole`/`inertRole` are keys of
`Vocabulary`, not column strings, so an entry cannot hardcode `todo` and
pass for the wrong reason.

## Two findings, both from mutation rather than reading

**1. The census was wrong about `recovery-reconciler.ts:198.`** It was
flagged as a `resolveLifecycleColumns` site, so the row was first
labelled as covering it. **Destroying that role resolution leaves all 18
tests green** — `decideRecovery` looks policy up by *column id* and
never consults a role. The row is relabelled to what it actually proves,
and mutation-verified against that instead: keying the reconciler's
policy lookup on the `todo` literal fails exactly its renamed test.

**2. `resolveRoleRecovery` is an unreachable export.** It is the only
use of `resolveLifecycleColumns` in that file and has **no production
caller anywhere** in engine, core, or dashboard. So that census line is
not a live converted site — it is a helper written ahead of its
consumer. **Not fixed here:** it is production code owned by the U4
slice, and whether the consumer is still to land or it should be deleted
is its author's call.

## Observability is now explicit in the type

`persisted-row` is the strong form. `returned-decision` is recorded as
**weaker evidence** and the reconciler row uses it, because
`reconcileRecovery` decides and does not apply — there is no row to
read. Naming it in the type is what stops a return-value assertion from
quietly passing as observed state, and it is what keeps the ledger
truthful per site.

## Harness seam

`PgTestHarness` now exposes its raw admin SQL client. The store
**stamps** `updatedAt`/`columnMovedAt` on every write, so `updateTask`
cannot express an aged row at all — the patch is accepted and the value
silently replaced with `now`. **Found by the new case failing on BOTH
vocabularies**, which is what distinguishes a broken fixture from a
broken guard. Seeding only; assertions still read back through the real
`getTask` path.

## Mutation verification

| Mutation | Result |
|---|---|
| revert **only** `recoverStrandedCompletedTodoTasks`'s resolution |
**exactly** that row's renamed test fails |
| revert **only** `surfaceStalePausedTodos`'s resolution | **exactly**
its own renamed test fails |
| reconciler policy lookup keyed on `todo` | exactly the reconciler
row's renamed test fails |
| `resolveRoleRecovery` role resolution destroyed | **nothing fails** →
finding #2 |
| `hold-release` `isHeldTask` keyed on `todo` | 5 of 18 fail; default
spine survives |
| `markMoveInFlight` dropped | both spine tests fail |

Per-site verification matters here: three rows could all be riding one
guard. They are not.

## The honest number

**Proven end to end: 5** (two self-healing sweeps, the reconciler's
policy lookup at the weaker observability, hold-release's capacity
release, and the graph boundary + `moveTask` + post-commit bus).

**Not proven: 11 call sites** — `merger.ts:324-326`,
`merger-ai.ts:1022,1039`, `auto-merge-finalization.ts:20-22`,
`executor.ts:1763,6339,6341`, `self-healing.ts:713,6732`,
`mesh-lease-manager.ts:61`, `task-agent-sync.ts:59`,
`core/task-store/reads.ts:130`, `core/live-agent-count.ts:63-75`, and
four dashboard route sites.

The ledger lives in the file, not just here, so it stays with the code.

## Where the table does not fit — reported, not papered over

The **merge/rebound family** cannot be a table row: those sweeps have no
observable persisted effect without a real git repository, so `acted`
cannot be written against the row at all. They need an engine-slow
real-git lane. The dashboard sites need an HTTP route test with a live
store. Both are different lanes, not missing entries.

## Verification

- 18/18 green; engine + core `tsc --noEmit` clean; `pnpm test:gate`
green (299 + 10 + 71)
- full core PG suite run (the harness is shared): 1036 passed, 3 failed
in `central-archive-secrets` and `workflow-settings-project-identity` —
**reproduce identically with this change stashed**, pre-existing

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:25:17 -07:00
gsxdsm
b133d521c4 U4 vertical slice: recovery-policy reconciler + ratified safety invariant (measured: engine ~780, real cost is a settings migration) (#2478)
Stacked on #2477. Base is
`feature/workflow-vocabulary-u4-delete-dep-blocked` — do not merge
before it.

The smallest end-to-end slice of the U4 reshape, built to **measure**
the real cost before committing to the full policy table. The survey's
~900-line reconciler figure was reasoned, not prototyped; this replaces
it with numbers.

**Everything here is additive and unwired. No behavior changes.**

## What lands

| | lines | what |
|---|---:|---|
| `WorkflowColumnRecovery` (IR) | 42 | one key — `stalenessMs` +
`onStale`. Optional and omitted when unset, so existing workflows
serialize byte-identically. |
| `recovery-reconciler.ts` | 176 | one engine: walks live cards,
resolves each card's policy from **its own** workflow (per task, shared
`irCache` — a 400-card board across three workflows reads three IRs),
returns decisions. Decision and application are separate so the safety
boundary is assertable without running an engine. |
| `recovery-policy-safety.test.ts` | 156 | one-time. The **ratified
invariant**. |

## Measured cost vs the ~900 estimate

**Engine + IR types = 222 lines** for one action (`surface`) and one
safeguard.

Extrapolating the rest — `rebound` (target resolution, attempt budgets,
backward-move proof, five more safeguards) ≈ +350, `archive` ≈ +50, the
`budgets`/`dependencies` keys ≈ +150 — lands near **780**.

So **~900 was a good estimate for the engine**, and the vertical slice
does not move it much. That is the answer to the question asked.

## But the estimate's real miss is not lines

**16 of the 34 POLICY sweeps read an operator setting today** — ~17
distinct policy-threshold keys, including `stalePausedTodoThresholdMs`,
`inReviewStalledThresholdMs`, `taskStuckTimeoutMs`,
`doneAutoArchiveDays`, `maxPostReviewFixes`.

Moving those sweeps into workflow policy is **not a code refactor — it
is a settings migration with operator-visible blast radius**, and it
needs three decisions the line estimate never surfaced:

1. Does workflow policy **override** the global setting, or defer to it?
2. What happens to **existing projects** that already configured those
settings?
3. Does an **unset** policy inherit the setting, or the built-in
default?

That is the gating question for the full table — not the reconciler's
size.

## Why the sweep is not retired here

Retiring `surfaceStalePausedTodos` requires builtin:coding to declare
the policy **and** `stalePausedTodoThresholdMs` to migrate — or the
behavior silently disappears for every existing project. That is the
settings migration above, and it belongs behind its own decision rather
than smuggled into a measurement slice.

The reconciler is therefore **unwired — deliberately dead code**, for
exactly as long as it takes to get that decision.

## The ratified safety invariant

The six safeguards (user pause, `autoMerge:false`, dependency, capacity,
merge-proof, at-most-once) live **outside** the policy table. A workflow
must never be able to author a safety invariant away.

Encoded two ways, because either alone is defeatable:
- **structural** — the policy exposes only an allow-listed key set;
adding a key requires editing the test and re-stating the safety
argument (the friction is the point);
- **behavioral** — a policy attempting every spelling of "ignore the
user pause" has no effect.

**Both halves mutation-verified**, because a safety test that cannot
fail is worse than none:
- making the reconciler honor a policy field that disables the
user-pause safeguard → **fails**
- adding an unreviewed key to the policy schema → **fails**

A third test asserts the reconciler still **acts** on an unpaused card,
so a reconciler that suppressed everything cannot pass by doing nothing.

## Scope limits stated rather than implied

Only the `surface` action is implemented, so only its relevant safeguard
is wired. `surface` mutates no lifecycle state; the other five gate
lifecycle-**mutating** actions that do not exist yet, and wiring them
now would be untestable dead code. A test records this so the absence
reads as deliberate and must be updated when `rebound` lands.

## Verification

- `tsc --noEmit` clean in core and engine; `pnpm lint` clean
- merge gate green (299 + 10 + 71)
- 23 safety tests green; `workflow-lifecycle-traits` green

No changeset: `@fusion/core` and `@fusion/engine` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-27 15:06:19 -07:00
gsxdsm
710d56b2db U4 trim: delete the dependency-blocked-todo feature (unreachable in production) and revert 5a2de7d (#2477)
Stacked on #2474. Base is `feature/workflow-vocabulary-u4-dead-code` —
do not merge before it.

Deletes an **entire feature that has never executed in production**, and
reverts `5a2de7d`, which only threaded resolved lifecycle columns
through it.

## Reachability evidence — the whole basis for this

```
surfaceDependencyBlockedTodos          ← in NEITHER sweep registry; no caller in
  └─ getDependencyBlockedTodoReporter()      engine/dashboard/cli — only tests
      └─ engine/dependency-blocked-todo-reporter.ts   ← sole caller of ↓
          └─ core/computeDependencyBlockedTodoReport
```

self-healing owns two name-based sweep registries (`runStartupRecovery`,
58 entries; `runMaintenance`, 76). `surfaceDependencyBlockedTodos` is in
**neither**, so nothing ever invoked the chain below it. Its four tests
passed while proving nothing about production.

## Why delete rather than wire it up

Wiring was the tempting option and is the riskier one. Switching on a
450-line path that has never run — whose tests therefore establish
nothing about its behavior against real data — is a **behavior change
with unquantified blast radius**. This program already refused exactly
that move for the **pool-id sentinel**, a one-line change that would
switch on dormant enforcement across every project. This is the same
class of move at ~450× the size.

Deleting is also the recoverable direction: git keeps the feature, and
it can be resurrected deliberately — with tests that prove it *runs* —
if dependency-blocked reporting is actually wanted.

## The settings keys go with it

`dependencyBlockedTodoReportEnabled` defaulted `true` while driving
nothing. A schema/API-visible switch that lies about what the system
does is worse than no switch. (It had no dashboard UI field — the
dashboard test allowlist already recorded it as *"no UI field"*.) Four
sibling tuning keys are removed with it.

## Against my own earlier work

`5a2de7d` threaded resolved lifecycle roles into
`computeDependencyBlockedTodoReport` and its reporter, answering a
review finding I confirmed as real. **The code was correct; the impact
claim was not**, because the path never executes. Neither the reviewer
nor I checked *reachability* before agreeing the defect mattered — only
correctness. A correction is posted on that thread in #2470.

**Scope limit on that admission:** the same finding also described
*incorrect scheduler ordering*. That half runs through
`buildUnblockWeightMap` in `task-priority.ts`, which is **live** and was
already threading `terminalColumns` (B1, `434b385`). Scheduler ordering
was never affected, before or after.

## What survives

`blocker-fanout.ts` **stays** — it is live via `task-priority.ts`. Only
the plural `holdColumns` option added by `5a2de7d` is reverted, since
the deleted report was its sole consumer. `holdColumn` (singular, from
B1) remains.

## Net

**1,244 deletions / 5 insertions across 15 files** — ~450 production
lines, ~684 test lines, 5 settings keys.

## Verification

- `tsc --noEmit` clean in **core, engine, and dashboard-app**; `pnpm
lint` clean
- merge gate green (299 + 10 + 71)
- self-healing suite: 411 passed, 1 **pre-existing** failure
(`archiveStaleDoneTasks`)
- dashboard settings-descriptions suite green
- `settings-parity.test.ts` has one **pre-existing** failure
(`agentToolOutputMaxChars` overlap) that fails identically with these
changes stashed — unrelated to this deletion

No changeset: `@fusion/core` and `@fusion/engine` are private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added quiet-window backlog health diagnostics for stalled items in
review, with repeat-alert suppression.
  * Added default thresholds for backlog-pressure alerts.

* **Changes**
  * Removed dependency-blocked todo reporting and related alerts.
* Removed the dependency-blocked todo enable/disable setting; remaining
tuning options are no longer active.
* Updated the workflow hold classification to use a single todo column.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:54:32 -07:00
gsxdsm
5d0f1ef631 Phase B slice B1: lifecycle column roles in the U6 policy modules (4 guards, red-green) (#2479)
**Stacked on #2469** → #2468 → #2467. Base is
`feature/workflow-capacity-ground-truth`.

This is **slice B1 of Phase B, not all of Phase B.** Sizing escalation
sent separately; the census is below.

## Why this is a slice

Measured census of code lines referencing a lifecycle column literal
(comments excluded):

| Unit | Files | Sites |
|---|---|---:|
| U4 | `self-healing.ts` | 203 |
| U5 | `executor.ts` 171, `scheduler.ts` 55, `replan-target.ts` 20,
`merger-ai.ts` 5, `hold-release.ts` 4, `mesh-lease-manager.ts` 4,
`task-agent-sync.ts` 3 | 262 |
| U6 | `moves.ts` 34, `default-workflow-hooks.ts` 13, `board-config.ts`
9, `blocker-fanout.ts` 6, `task-priority.ts` 5,
`dependency-blocked-todo-report.ts` 2, `stale-paused-todo.ts` 1 | 70 |
| | **Total** | **535** |

The plan's "~207" counts the guard category only. Under the phase's
non-negotiable rule — a test that **fails before** conversion, per guard
— that is ~200 red-green cycles. Doing it as one sweep would reproduce
exactly the failure this phase exists to prevent: converted guards
nobody proved still fire.

`moves.ts` and `default-workflow-hooks.ts` stay **parked** per the
dispatch constraint (move-path convergence and the pool-id sentinel are
on an operator decision).

## Guards converted (4), each red-green

Every case below was written **first** and observed failing against the
literal implementation.

| Module | Guard | Before → After |
|---|---|---|
| `stale-paused-todo.ts` | stall detection | `column !== "todo"` →
resolved **hold** column |
| `blocker-fanout.ts` | active | `ACTIVE_COLUMNS.has(col)` →
`!terminalColumns.has(col)` |
| `blocker-fanout.ts` | hold-wait metric | `col === "todo"` → resolved
**hold** column |
| `task-priority.ts` | unblock active | `UNBLOCK_ACTIVE_COLUMNS`
**deleted**, folded into the terminal set |

Three of the seven new cases are **regression floors** that pass before
and after. One of them earned its keep immediately: it failed on my own
fixture (`activeCount` vs the public `totalCount`), catching a bad test
rather than bad code — which is the point of asserting the default path
alongside the renamed one.

### The `task-priority` finding

`UNBLOCK_ACTIVE_COLUMNS` and `DONE_COLUMNS` encoded **one concept
twice**, two lines apart, and disagreed for any custom column:
dependency counting treated a `drafting` card as unmet (correct) while
the active check treated it as inactive (wrong), zeroing the blocker's
unblock weight. The enumeration wasn't just legacy-shaped — it
contradicted its own neighbour.

## ⚠️ Behavior change, not a pure refactor

Inverting active from enumeration to exclusion means **a card in a
column that is neither terminal nor in the legacy enum now counts as
active where it previously did not.** That is the plan's stated intent,
but it is a real change for any project already using a custom column —
**Coding (Ideas)' `ideas` column is the in-tree case.** Fan-out counts
and unblock weights for such cards will rise.

## Verification

- Four affected suites green (45 tests), each conversion observed
red→green.
- `pnpm lint`, `tsc --noEmit` (core) green.

**Not verified / not done, stated plainly:**

- **Call sites are not wired.** These modules now *accept* resolved
roles; every parameter still defaults to the legacy set, so at the call
sites the vocabulary is unchanged. A caller that cannot resolve a
workflow keeps literal behavior. Threading `resolveLifecycleColumns`
through `reads.ts` and `self-healing.ts` is follow-on work — until then
the guards are *convertible*, not *converted end-to-end*.
- `dependency-blocked-todo-report.ts` and `board-config.ts` are
untouched in this slice.
- 19 core-suite failures exist on this branch; all confirmed
**pre-existing** by stashing and re-running on a clean tree
(`duplicate-guard`, `log-severity-spam-contract`, `settings-parity`,
`task-delete-caller-attribution`, `settings-defaults`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---

**Supersedes #2470**, which GitHub force-closed when its base branch was
deleted by the merge of #2469 and refuses to reopen. Same head branch,
same commits (rebased onto `main`), now based on `main` directly. The
two P1 review threads on #2470 were resolved there — one of them with a
correction noting the threading half landed in code that was
subsequently deleted as a dead feature in #2477.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Dependency and blocker reports now correctly recognize custom hold,
active, and terminal workflow columns.
* Blockers in renamed terminal columns are no longer incorrectly
reported as active.
* Stale paused-task badges and self-healing now work with
workflow-specific hold columns.
* Mixed boards with different workflow column names are handled
consistently.
* Existing default workflow behavior remains compatible, including
fallback handling when workflow details cannot be resolved.

* **Enhancements**
* Reporting and task-priority calculations now support configurable
single or multiple hold and terminal columns.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:19:32 -07:00
gsxdsm
553cc3b517 Phase A3: the in-txn capacity invariant does NOT hold — root-caused, ratcheted, not fixed (#2469)
**Stacked on #2468**, which is stacked on #2467. Base is
`feature/workflow-move-path-convergence`.

**Answer: the A2 observation was real. The documented invariant is
broken.** Not a false alarm, not a harness artifact.

`workflow-capacity.ts` states enforcement "runs INSIDE
`moveTaskInternal`'s transaction and is **NEVER bypassable** (not a
guard — runs regardless of bypassGuards/recoveryRehome/moveSource)". It
does not hold for default-workflow tasks, for two independent reasons.

## R1 — Pool-id sentinel mismatch (the defect)

| Site | Sentinel for "no workflow selection" |
|---|---|
| `moves.ts:319` (in-txn check, **asks**) | `?? "builtin:coding"` |
| `countActiveInCapacitySlotAsyncImpl` (**answers**) | `??
DEFAULT_WORKFLOW_POOL_ID` → `"__default-workflow__"` |
| `hold-release.ts:116` (sweep, second enforcement point) | `??
DEFAULT_WORKFLOW_POOL_ID` ✅ |

The check asks for occupants of a pool that no occupant is ever bucketed
into, so the count comes back `0` and the limit can never bind. The
sweep is correct, so the two enforcement points **disagree about pool
identity** — precisely what the module docstring says is impossible
("the two enforcement points can never disagree on what a limit *is*,
only on the live count").

Note the shape of the bug: it is not a missing check. The check runs,
queries correctly, and returns a confidently wrong answer.

## R2 — The `useWorkflow` gate

The whole block sits inside `if (useWorkflow && workflowIr && fromColumn
!== toColumn)` (`moves.ts:921`), and `useWorkflow` reads the raw
`experimentalFeatures.workflowColumns` key nothing in production sets.
**On the live path the check cannot run at all**, so R1 is latent today
and becomes reachable the moment A2 converges onto the flag-ON side.

## How this was established, not guessed

Three of my assumptions failed earlier in this program, so this one is
pinned by a **discriminating experiment** rather than a code reading:

| Case | Path | Selection | Result |
|---|---|---|---|
| DEFECT (R2) | inline (live) | none | accepted — check cannot run |
| DEFECT (R1) | hooks | none | accepted — sentinels disagree |
| **DISCRIMINATOR** | hooks | explicit `builtin:coding` | **refused,
`capacity-exhausted`** |

The third case changes nothing but sentinel agreement. That rules out
"capacity is simply not wired" and isolates the cause to the mismatch.

Both-path forcing reuses A2's `assertPathActive` probe. Without it this
suite would silently run one path twice and report a tautology — the
failure mode that produced sixteen false passes across A2's two harness
bugs.

## The deliverable: an invariant ratchet

The three cases above assert today's wrong behavior, so on their own
they would let the defect live forever. A fourth case states the
invariant **as written** and is marked `it.fails`:

- **today** — the body fails, so `it.fails` passes; CI stays green while
honestly recording the breach;
- **when fixed** — the body passes, `it.fails` *fails*, forcing whoever
lands the fix to flip it and the two `DEFECT:` expectations.

That is "a test that fails if the invariant is broken" in the only shape
that does not park a permanently-red test in CI.

## Blast radius (step 4) — why I did not fix it

The fix is one line: make `moves.ts:319` use `DEFAULT_WORKFLOW_POOL_ID`,
matching the sweep. The consequences are not one line.

**Today: zero.** `useWorkflow` is false everywhere, so the corrected
check still cannot run on the live path. The sentinel fix is safe to
land in isolation.

**At A2 convergence: every default-workflow task move into `in-progress`
becomes capacity-checked against `maxConcurrent` (default 2), for the
first time.** Affected movers:

- **The graph column boundary** catches `capacity-exhausted` and *parks
the run*. Runs that previously proceeded would begin suspending — this
is a scheduling behavior change across every project, not an error path.
- **`executor.ts`, `project-engine.ts`, `pr-comment-handler.ts`** each
move tasks into `in-progress` and would begin seeing a rejection they
have never seen.
- **Operator drags and the promote route** would start refusing beyond
`maxConcurrent`.

Scheduler-side admission (`maxConcurrent`) is a separate, still-live
control, so this is not "capacity is unenforced today" — it is "the
store-level check the graph and promote paths are written against
returns 0 and never binds".

**Recommendation:** land the sentinel fix on its own (provably inert
today), with the ratchet flipped in the same commit, *before* A2
convergence — so convergence does not simultaneously switch paths and
switch on a previously-dead enforcement.

## Verification

4 cases green against real PostgreSQL (3 passed + 1 expected-fail), path
flip proven live on every case. `pnpm lint` and `tsc --noEmit` green.

## Follow-up: both "not verified" items are now answered

Recorded here rather than left as open questions, since this is where
anyone investigating capacity will look.

**1. Does the sync/SQLite counter carry the same mismatch? — YES,
identically, but it is unreachable.**

`countActiveInCapacitySlotSyncImpl` (`project-store-ops.ts:767`) buckets
rows the same way as the async one:

```ts
const effectiveWorkflowId = row.wid ?? TaskStore.DEFAULT_WORKFLOW_POOL_ID;
```

So it disagrees with `moves.ts:319` in exactly the same way. **However**
its only caller is the public `TaskStore.countActiveInCapacitySlotSync`
wrapper, which has no in-repo caller at all — it is dead API surface.
The mismatch is real but currently unreachable, which makes it a
landmine for whoever wires it up rather than an active defect. Fixing
the sentinel should fix both call sites together.

**2. Are custom workflows with an explicit numeric `limit` affected? —
NO, and this is already proven by the discriminator above.**

The mismatch fires **only when the selection row is absent** — that is
what the `??` fallback is for. A custom workflow necessarily *has* a
selection row; that is what makes it custom. The DISCRIMINATOR case adds
an explicit selection and the rejection appears, which is exactly the
custom-workflow shape. So the defect is scoped to **no-selection
(default-workflow) tasks only**.

The explicit-numeric-`limit` question turns out to be orthogonal:
`resolveColumnBudgetKey` returning `col:${columnId}` decides *which
columns share a budget*, not which pool id is passed to the counter. It
does not interact with the sentinel at all.

Net effect on blast radius: **narrower than first stated.** Only
default-workflow (no-selection) tasks slip the limit today.
Custom-workflow tasks are already enforced — meaning the fix does not
switch enforcement on for them, it only closes the gap for the default
workflow.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added coverage for workflow column capacity enforcement during
transactions.
* Documented scenarios where capacity limits are bypassed, including
tasks without workflow selection.
* Verified that explicit workflow selection correctly rejects moves when
capacity is exhausted.
* Added a tracked failing test for the expected invariant once
enforcement is corrected.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 13:59:56 -07:00
gsxdsm
b941d3cba5 Phase A2 (steps 1-2): differential characterization of the two move paths (#2468)
**Stacked on #2467** (base is `feature/workflow-owned-lifecycle`, not
`main`).

Phase A2 steps 1 and 2. **Step 3 — make one path authoritative and
delete the other — is NOT done.** It is blocked on a measured
divergence, escalated to the operator. This PR is the evidence that
decision needs.

## The setup

`moves.ts` branches on `useWorkflow =
isWorkflowColumnsCompatibilityFlagEnabled(settings)`, which reads the
raw `experimentalFeatures.workflowColumns` key. Nothing in production
writes it, so the **inline branch is LIVE** and
**`default-workflow-hooks.ts` is DEAD**.

Because only one implementation runs, equivalence cannot be observed by
running the suite normally — the dead path is never entered. Every case
here forces both paths explicitly through one shared fixture and
compares a 19-field observation, not "it moved".

## Step 1-2 result: side-effect equivalence is PROVEN

Eight behaviors, field-by-field identical across both paths:

| Behavior | Verdict |
|---|---|
| `in-progress → todo` user reopen field clears | equivalent |
| Engine-source reopen does not set `userPaused` | equivalent |
| `preserveStatus` keeps status/error | equivalent |
| `preservePause` keeps an operator park (FN-7851) | equivalent |
| Timing / `cumulativeActiveMs` across exit and re-entry | equivalent |
| `preserveResumeState` step progress | equivalent |
| `preserveWorktree` | equivalent |
| Default worktree clear on reopen | equivalent |

### Why this is a proof and not a green suite

Two independent guards, both of which caught a real silent failure in
this PR's own development:

- **The forcing mechanism is self-checked.** `assertPathActive` probes
an undeclared target column — whose rejection message differs per path —
before every case. The first version of this suite wrote the flag with
`updateSettings` instead of `updateGlobalSettings`
(`experimentalFeatures` is global-scoped, which is exactly why
`moves.ts` reads it through `getSettingsFast()`), and reported **nine
passing "equivalence" cases while running the inline path twice**. The
check then caught a second failure: `updateGlobalSettings` *merges*, so
resetting with `{}` left a previous `true` in place and leaked the hooks
path into seven cases that believed they were on inline.
- **The suite is mutation-tested.** Deleting `task.blockedBy =
undefined` from `applyResetOnEntryEffects` fails the reopen case, naming
the field. Restored before commit.

Timestamps are compared by presence rather than value — the two runs
happen at different wall-clock instants by construction — but a path
that forgets to stamp `executionCompletedAt`, or wrongly clears
`firstExecutionAt`, still fails.

## ⚠️ Read this before writing any both-paths test

**A differential harness that cannot prove which path it is on will
report a tautology, confidently, and in green.** This suite hit that
twice in one afternoon:

1. **Global-scoped key written to project scope.** The first version set
the flag with `updateSettings`. `experimentalFeatures` is
**global**-scoped — which is exactly why `moves.ts` reads it through
`getSettingsFast()` (merged global + project). The write was silently
accepted and never reached `useWorkflow`. Result: **nine passing
"equivalence" cases while running the inline path twice.**
2. **Merge-on-write leaking a stale `true`.** `updateGlobalSettings`
*merges*, so resetting with `{ experimentalFeatures: {} }` left the
previous `workflowColumns: true` in place. Result: the hooks path leaked
into **seven cases that believed they were on inline.**

Neither failure produced a red test. Both were caught only by
`assertPathActive` — a per-case probe that moves a task to an undeclared
column and asserts on the rejection *message*, which differs per path
(`Valid targets: …` inline vs `Unknown column for this workflow` on
hooks).

This is the same failure class as a spy passing on a refused payload
(see #2467): **the observation confirms the assumption instead of the
behavior.** The rule that generalizes:

> When a test forces a code path, assert that the path is active using a
signal only that path can produce — before every case, not once in
setup. A forcing mechanism that can fail silently makes every assertion
downstream worthless.

Any future work touching both move paths needs this probe or it will get
a confident wrong answer.

## Why step 3 is blocked

### Divergence: rejection type and message

| | Inline (live) | Hooks (dead) |
|---|---|---|
| Validates against | legacy `VALID_TRANSITIONS` | the task's own
workflow |
| Throws | bare `Error` | `TransitionRejectionError` with a
machine-readable `rejection` |
| Message | `Valid targets: …` | `Unknown column for this workflow` |

Both reject, so neither is "broken" — but they are not interchangeable.
Making either authoritative changes what every catch site observes,
including the flag-OFF characterization suite that pins the bare-Error
contract and the callers that branch on `rejection.code`.

### Unproven, recorded as an honest negative: in-transaction capacity

The capacity block sits inside `if (useWorkflow && workflowIr &&
fromColumn !== toColumn)`, so it **cannot** run on the live path. The
natural inference is "convergence turns store-level capacity rejection
on for every project at once" — a serious blast radius, since
`capacity-exhausted` is what the graph column boundary parks on and what
the promote route surfaces to operators.

**That inference did not survive measurement.** With `maxConcurrent: 1`
and an already-occupied wip column, the second move was **accepted on
both paths**. Something further in — `resolveColumnCapacity`'s limit
resolution, or what `countActiveInCapacitySlotAsync` counts as an
occupant — keeps the check from firing even flag-on. This suite does not
establish which.

So the capacity blast radius is **unquantified, not absent**. The test
pins today's observed behavior so the investigation starts from a fact
rather than from the code reading; if a future change makes it reject,
that failure is the signal to reopen the question.

## Verification

- 10/10 green against real PostgreSQL, with the path flip proven live by
`assertPathActive` on every case.
- Mutation-tested (see above).
- `pnpm lint` and `tsc --noEmit` (core) green.

**Not verified:** whether the capacity gate would activate under some
other configuration; the plugin column-gate and post-commit plugin-hook
divergences (also inside the `useWorkflow` gate) are identified
structurally but not characterized here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Workflow lifecycle updates now provide more consistent task state
notifications and workflow activity handling.
* Task lists display resolved workflow column names and lifecycle
details more reliably.
* **Bug Fixes**
* Workflow-based task promotion and board views no longer depend on an
obsolete feature setting.
  * Improved recovery for tasks left in transitional states.
* Preserved task status, pause, progress, timing, and worktree behavior
across workflow transitions.
* **Reliability**
* Added stronger validation and durable handling for workflow events and
follow-up processing.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 13:49:55 -07:00
gsxdsm
4158cf1ab7 Phase A: workflow-owned lifecycle foundation (U1, U2, U3) (#2467)
Phase A (Foundation) of
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`.
Three units, one commit each. No operator-visible behavior change.

## U1 — Lifecycle-column resolution seam

`resolveLifecycleColumns(ir)` returns `{ intake, hold, wip, review,
complete, archived }` — the first column carrying each trait,
`undefined` for a role no column carries.
`resolveTaskLifecycleColumns(store, taskId, cache?)` is the store-aware
form; the cache is caller-owned so a sweep reads one IR per workflow
rather than one per card.

A v1/column-less IR resolves to `undefined` for the **whole struct**
rather than a struct of undefined roles. A caller must be able to
distinguish "this workflow declares no hold column" (a real shape to
honor) from "no column vocabulary at all" (skip and log) — only the
second licenses conservative fallback.

Nothing consumes the seam yet; Phases B–D convert the ~207 hardcoded
column literals onto it.

## U2 — Delete the pre-cutover parity machinery (delete-only)

**`workflow-columns-settings.ts`** — `isWorkflowColumnsEnabled` had the
body `return true`. Six live call sites branched on it, so every
flag-OFF arm was dead code that read as a supported configuration.
Deleted; surviving side inlined at self-healing's transitionPending
sweep, the scheduler's per-column capacity diagnostic, merge-trait's
policy resolver, the board-workflows payload, two task-workflow routes,
and the CLI TUI's column enrichment.

**`workflow-parity.ts`** — asserted the default workflow's adjacency
*equals* the legacy `VALID_TRANSITIONS`. U11 deliberately breaks that
equality by merging Todo into Planning, so this is not a stale assertion
to update; it is a contract against the target state. Its emitter
(`workflow-parity-observer.ts`) is already a tombstone, so
`getWorkflowParitySummary` and `computeWorkflowColumnsGraduationReport`
aggregated run-audit rows nothing writes and had no caller outside
`TaskStore`. Both store methods go with it.

`flagEnabled` stays on the board-workflows **wire** as a constant `true`
— shipped dashboard clients still branch on it, and changing the
response shape is not a deletion. U10 retires the field once no client
reads it.

The `legacy-tombstones` ratchet is extended to both files plus seven
symbols, each with the reason it is gone.

### ⚠️ Finding: the third listed deletion was NOT dead

The plan also lists "the flag-off inline move path" in
`task-store/moves.ts`. It is **not** deleted, per U2's execution note
("any behavior change found while removing a branch means the branch was
not dead").

That path is gated on `isWorkflowColumnsCompatibilityFlagEnabled`
(`store.ts:38`) — a **different** function from the always-true public
helper. It reads the raw `experimentalFeatures.workflowColumns` setting,
which nothing in production sets (`settings-schema.ts:396` — "no default
flags are emitted"; zero non-test writers; the operator's own
`~/.fusion/settings.json` has no such key). So `useWorkflow` is false
for effectively every real project: the flag-OFF inline side effects are
the **live** default move path and the flag-ON `default-workflow-hooks`
path is the dead one. The code says so itself at `moves.ts:638`.

Deleting that branch would swap every project onto an untravelled code
path — a behavior change, not a deletion.

**Carry this into Phases B and C, stated plainly so the plan's error is
not repeated:**

> **The inline move path in `moves.ts` is LIVE.
`default-workflow-hooks.ts` (the trait-hook path) is DEAD.** KTD-6
asserted the inverse. Until the convergence unit lands, **nothing may
assume trait hooks run** — a guard, sweep, or subscriber written against
`applyDefaultWorkflowMoveEffects` would never fire in production and
would still pass its tests.

Convergence is **not** attempted here. It is its own unit (Phase A2)
with a proper equivalence proof, per operator decision.

### U3's emit point is on the LIVE path — the seam is not born dead

Worth stating explicitly because it is the failure mode that would make
every later subscriber silently never fire: the `TaskTransitioned` emit
is **not** inside the `if (useWorkflow)` branch. That block closes at
`moves.ts:1212`; the emit sits at `:1214`, beside the existing
`store.emit("task:moved", …)`, on the unconditional post-commit path. It
therefore fires on **both** the live inline path and the dead hooks
path, and the convergence unit inherits the obligation to keep it firing
on whichever path survives — same events, same order, same payloads.

The graph-side emitters (`NodeEntered`, `RunSuspended`) carry the same
risk from a different direction: the bus refuses an invalid payload
*silently* by design, so an emitter regression would stop the event with
no test failure. They are asserted end-to-end through the real bus —
"did a subscriber actually receive it", not "was emit called" — because
a spy passes on a refused payload. The `moveTaskInternalImpl` emit does
**not** yet have that end-to-end assertion against a real store move;
that proof belongs to the convergence unit, which has to build the
both-paths fixture anyway.

## U3 — Post-commit event seam with a transactional outbox

**The bus is not a queue, not a transaction participant, and not a
delivery guarantee.** Durable follow-on work uses the transactional
outbox — a `workflow_work_items` row written *inside* the transition
transaction (the shape `createCompletionHandoffWorkflowWork` already
uses). "Emit after commit, let a subscriber enqueue the work" has a
crash window where a process dies between commit and subscriber, leaving
no event *and* no work-item row, so required work is skipped permanently
with nothing to recover from. Post-commit subscribers therefore carry
only losable reactions.

Emission is consequently lossy and isolated by design: a throwing or
rejecting subscriber is caught and logged, cannot roll back the
transition, and cannot stop the others. Deliveries append to one serial
chain, so two transitions on a task deliver in commit order.

The ids/outcomes-only rule is **mechanised, not documented** —
run-audit's equivalent lives only in prose and has been violated
repeatedly. A payload carrying an object body or a prose string is
refused at the emit boundary and never reaches a subscriber or log sink.
It degrades rather than throws: the emitter is post-commit, so a shape
bug must not become a lifecycle failure.

Emit points: `TaskTransitioned` from the single post-commit point in
`moveTaskInternalImpl`; `NodeEntered` and `RunSuspended` from the graph
column boundary, the latter *after* the durable continuation is
persisted so an observed suspension implies a resumable run.

`registerWorkflowEventSubscribers` (engine) is empty on purpose —
U7/U8/U10 move real reactions onto it, each with the characterization
test proving the reaction was non-authoritative first.

## Verification

- `pnpm test:gate` — green (2/10, 16/299, 1/71).
- `pnpm lint`, `pnpm build`, `tsc --noEmit` on core and engine — green.
- U1: 20 tests in `workflow-lifecycle-traits.test.ts`, including the
fully-renamed-workflow case (fails if the resolver falls back to a
literal) and a shared-cache read-count assertion.
- U2: `legacy-tombstones.test.ts` green with the extended ratchet;
`board-workflows`, `merge-trait`, `workflow-graph-executor-parity`, and
move-hook suites green with no expectation edits.
- U3: 20 bus-invariant unit tests (isolation, ordering, the allowed-key
and required-key halves of the ids-only rule, lossiness) plus 3
end-to-end emitter-delivery tests; 5 outbox tests against a **real
PostgreSQL** work-item table (crash survival, rollback, at-least-once
redelivery on lease expiry, idempotent handler → one effect,
dropped-subscriber vs. durable work). A hand-written fake of the lease
predicate would only prove the fake redelivers.

**Not verified:** the `moveTaskInternalImpl` emit is confirmed on the
unconditional post-commit path by structure and by the surrounding
tests, but is *not* yet asserted end-to-end against a real store move on
both flag settings — that is Phase A2's fixture. The engine subscriber
registry ships empty by design, so no production subscriber exercises
the bus end-to-end yet. `settings-defaults.test.ts` has one pre-existing
failure on `main` (a logger-prefix mismatch in the
`mergeIntegrationWorktree=cwd-main` warning) — confirmed present on a
clean tree, unrelated to this branch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Workflow lifecycle columns are now derived from workflow definitions,
supporting renamed and custom workflows.
* Added post-commit lifecycle events for task transitions, node entry,
and run suspend/resume with validated payloads.
* Follow-on processing for lifecycle emissions is now more robust
(rollback-safe, at-least-once delivery, idempotent handling).
* **Bug Fixes**
* Workflow board responses, task enrichment, and promotion no longer
depend on workflow-columns feature-flag gating.
  * Subscriber failures no longer impact committed workflow transitions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 13:30:13 -07:00
gsxdsm
8b039a543e fix(desktop): advance Pi runtime pin to 0.82.1 for packaging PR lane (#2465)
## Summary
- Advance the matched Pi runtime pin (`pi-ai`, `pi-coding-agent`,
`pi-agent-core`, `pi-tui`) from **0.82.0 → 0.82.1** so
electron-builder's production-dependency walk accepts `pi-agent-core`'s
`pi-ai@^0.82.1` requirement.
- Fixes the Desktop packaging PR-lane failure:
`Production dependency @earendil-works/pi-ai not found for package
@earendil-works/pi-agent-core` (required `^0.82.1`).
- Keep the workspace override guard; update pin-policy fixtures and CLI
package-config expectations.
- Tighten the advisory packaging step-order test so it asserts against
the real `electron-builder --dir` step (not a missing release-only step
name that previously passed via `indexOf === -1`).
- Run `pnpm dedupe` so the packaging lane's lockfile dedupe
early-warning is clean.

## Context
#2439 pinned the full Pi closure at 0.82.0 and made recent main-based
packaging runs green. This advances to the current upstream patch so
deploy + electron-builder stay aligned with `pi-agent-core@0.82.1`'s
declared dependency range.

## Test plan
- [x] `node scripts/check-pi-versions-pinned.mjs`
- [x] `node --test scripts/__tests__/check-pi-versions-pinned.test.mjs`
- [x] `pnpm --filter @runfusion/fusion exec vitest run
src/__tests__/package-config.test.ts`
- [x] `pnpm --filter @fusion/desktop exec vitest run
src/__tests__/release-workflow.test.ts`
- [x] `pnpm dedupe --check`
- [ ] GitHub: Desktop packaging (should run full packaging walk —
lockfile/package.json touched)
- [ ] GitHub: PR Checks (Lint, Typecheck, Build, Gate)
2026-07-26 23:47:49 -07:00
gsxdsm
99c9f14ee0 feat: run Plan Review in the planning lane with a Plan Review badge (#2462)
## What

Plan Review, planning, and the replan loop move from the implementation
column into the **planning lane** (`todo`), so a task under
specification never holds a WIP slot. The card crosses into
`in-progress` exactly once, at `parse`, released by the scheduler.

Operators also finally see a **Plan Review** badge while the gate runs —
it was previously invisible on the default workflow.

## The part that made it possible

Moving the node is ten lines. It was attempted three times and reverted
each time, because a graph run with no durable continuation replayed
from `start` and dragged an in-progress card *backward* out of the WIP
column, firing `abort-on-exit` and stranding it in a pre-WIP column with
no releaser.

So this PR adds the graph **entry contract** —
`resolveColumnResumeNode`:

| Card is in | Resumes at |
|---|---|
| `triage` | `start` |
| `todo` | `plan` |
| `in-progress` | `parse` — never re-plans, never moves backward |
| `in-review` | first review node — gates are not skipped |

`ir.columns` is ordered and that order is the lifecycle order; rework
and failure edges are excluded so the entry point is always the main
path. The proof it's the right fix: **`executor-task-done-invariant`
passes unmodified** after failing every previous attempt.

## Also in here

- **Release gate narrowed twice.** `isUnplannedForExecution` applies its
pre-release plan-review gate only when the node's column equals the
card's column *and* the group is enabled for the task. The enablement
check fixes a real deadlock — a task with Plan Review toggled off was
held forever waiting for evidence nothing would ever write.
- **Badge cleanup.** Gate badge reads "Plan Review" instead of the
ambiguous "Reviewing" and no longer hides behind a lane restriction; the
status badge stops duplicating it; `planning` renders as "Planning"
instead of the raw engine token.
- **Coding (Ideas)** renames its planner column to "Planning" (id `todo`
unchanged) and loses its private planning-node re-home — the graph it
clones is already plan-in-place.
- **New sweep** `reconcileUndeclaredTaskColumns` re-homes a row whose
column its workflow no longer declares. Written for a follow-up, kept
because it makes any column edit survivable.

## Test changes

Scheduler and release fixtures now model a card whose Plan Review passed
— the state every real card is in when the capacity sweep sees it. A
held unreviewed card is the gate working, and that path stays owned by
`pre-release-plan-review.test.ts`.

New `workflow-graph-entry-contract.test.ts` covers the invariant at
every lifecycle position, plus the gap-column and remediation-node
cases.

## Verification

Gate 299 + 70 + 10, dashboard badge suites 672, engine
workflow/entry/executor suites 147, core 122. Lint and typecheck clean.
Full engine suite sits at the pre-existing baseline (notifier /
plugin-runner / notification-service, untouched by this).

## Follow-up

Removing the Todo column entirely is a separate ~207-site
lifecycle-vocabulary refactor — planned in
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`
(companion docs PR).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Plan Review now runs in the Planning lane before implementation
begins.
* Cards resume from their current workflow column without replaying
earlier steps.
* Added automatic recovery for cards stranded in outdated workflow
columns.
* **Improvements**
  * Renamed the Coding (Ideas) planner column to “Planning.”
* Refined Plan Review gating to respect enabled settings and the card’s
current column.
* Updated planning and Plan Review badges for clearer, consistent labels
across cards and lists.
* **Bug Fixes**
* Improved workflow transitions and release behavior around planning,
review, and execution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 23:42:46 -07:00
gsxdsm
5ae6332563 refactor: collapse dead SQLite dual-path code; keep migration-only readers (#2454)
# Remove dead SQLite dual-path code; keep migration-only readers

## Summary
PostgreSQL cutover left hundreds of production dual-path branches
(`backendMode ? PG : SQLite/store.db`) whose SQLite arms only hit
throwing `Database`/`ArchiveDatabase`/`CentralDatabase` stubs. This
change mechanically collapses those unreachable arms so production
authority is AsyncDataLayer/PostgreSQL only, while preserving the six
authorized read-only migration/recovery `DatabaseSync` seams.

## Dual-path mass removed
| Metric | Before | After |
|---|---|---|
| `if (…backendMode)` (non-test) | ~328 | ~70 |
| `store.db` / `this.db` refs in core (non-test) | ~570+ | ~375 (mostly
pure legacy MissionStore/eval/insight SQLite classes + thin getters) |
| Net diff | — | **~6.7k lines removed** across 41 files |

Remaining `backendMode` checks are intentional (incomplete-PG sync
safe-defaults, settings-sync disabled-on-PG, symbol-lock PG-only gates,
“requires PostgreSQL” config versioning throws), not live SQLite
authority.

## Subsystems cleaned
- **Core TaskStore / task-store/***: collapsed if/else and early-return
dual-path across reads, moves, lifecycle, mutations, workflow, archive,
branch/PR, artifacts, comments, audit, project ops, etc. `initImpl` is
PostgreSQL-only (SQLite startup tail deleted).
- **Satellite stores**: automation, agent, routine, plugin, secrets,
approval-request, central-core dual-path arms collapsed.
- **Plugins**: reports async methods, compound-engineering pipeline +
session stores, CLI Printing Press store — SQLite fallbacks removed; PG
required.
- **Engine**: no functional dual-path change beyond whitespace
(settings-sync / peer-exchange PG-disabled behavior kept).

## Six migration-only readers retained (allowlist unchanged)
1. `packages/core/src/postgres/sqlite-migrator.ts`
2. `packages/core/src/project-identity.ts`
3. `packages/core/src/sqlite-validation.ts`
4. `packages/core/src/postgres/startup-factory.ts`
5. `packages/cli/src/commands/db.ts`
6. `scripts/lib/start-local-project.mjs`

Plus low-level `sqlite-adapter` and migrator/startup-import tests.
Inventory ratchet still requires exactly these six `new DatabaseSync(`
production sites, all `readOnly: true`.

## Not treated as SQLite
- `.fusion/project.json`, `task.json`, `agent-log.jsonl` file storage
- AsyncDataLayer / Drizzle PG paths
- Incomplete-PG sync safe-default stubs (still return empty/false/null
under backend without consulting SQLite)

## Verification
- `sqlite-production-reader-inventory.test.ts` — 15/15 pass
- `incomplete-pg-ports.pg.test.ts` — 6/6 pass
- Targeted PG tests (create-task, move, handoff, runtime-persistence,
agent, mission, insight, central-core) — green
- `tsc --noEmit` for `@fusion/core`, `@fusion/engine`,
`@fusion/dashboard` — green
- `scripts/check-no-getdatabase.mjs` — clean

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Improved end-to-end consistency by making PostgreSQL/async persistence
the standard across core task/workflow, automation, agents, plugins,
routines, secrets, approvals, central operations, and session storage.
* Unified scheduling, settings, configuration revision writes,
run/workflow selection, queues/leases/transitions, and audit/lifecycle
updates around consistent async transaction behavior.
* **Bug Fixes**
* Fixed edge cases for archived/deleted reads, unarchive/recovery flows,
not-found handling, and task/artifact/document/log/comment operations,
including more reliable emissions and hydration across search/list and
lifecycle operations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-26 23:28:42 -07:00
gsxdsm
0e3d2a2265 refactor: delete meta-task auto-archive and automated recovery follow-ups (#2461)
Deletes two pieces of automated "meta" machinery that filed and
garbage-collected cards restating state already on the task that failed.
Net **-1015 lines**.

## Why

**Automated recovery follow-ups.** `createAutomatedFollowup` and its
dedup engine (289 lines of signature matching, 1h recurrence
rate-limiting, 24h supersedes windows) existed to file recovery cards
for verification-cap and merge-conflict give-ups. In both cases the
parent is *already* parked `failed` with a descriptive `error` and a log
entry carrying the failing command, branch, and output — the card was a
second copy of that.

**Meta-task auto-archive.** The sweeps that garbage-collected those
cards were worse than redundant: the regex classifier matched ordinary
feature work, and its positional fallback bound cards to unrelated
tasks, so **live work could be archived**.

They are removed together, because the auto-archive sweeps only existed
to clean up after the follow-up engine.

## What changed

### Deleted
- `packages/engine/src/verification-followup-dedup.ts` in full —
`createAutomatedFollowup`, `decideAutomatedFollowup`,
`AutomatedFollowupKind`, `computeVerificationFailureSignature`,
`extractFailingTestFiles`.
- `findActiveRecoveryFollowUp` — dead code, defined and never called
(`tsc` independently flagged it `6133 declared but its value is never
read`).
- The meta-task auto-archive sweeps `autoArchiveResolvedMetaTasks` /
`autoArchiveStalledMetaTasks` and helpers `classifyMetaTask` /
`resolveMetaTargetTaskId` / `computeMetaChainDepth` / `archiveMetaTask`
/ `evaluateMetaAutoArchiveGuards`, plus settings
`metaTaskStallAutoCloseMs` and `metaTaskActiveExecutionGraceMs`.
- Run-audit types `task:auto-archived-meta-resolved`,
`task:auto-archived-meta-stalled`,
`task:auto-archive-meta-resolved-skipped`,
`task:auto-archive-meta-stalled-skipped`,
`verification:followup-created`, `verification:followup-deduped`.

The two signature helpers were **deleted rather than relocated** — once
the three call sites went they were provably unreachable:
`buildVerificationFailureSignature` had exactly one caller, and it was
the only caller of `extractFailingTestFiles`.

### Call sites 1 and 2 — park kept, card dropped
Verification-cap and merge-conflict give-ups keep their park, audit
event, operator comment, and log entry. Site 1's `error` string was
reworded off `"See follow-up task for investigation."` (no follow-up
will exist) to carry the guidance itself. `autoResolveDisabled` was
**kept** — it still drives the outer park guard and the `reason` string;
only the inner branch that guarded card creation is gone.

### Call site 3 — autostash orphan, replaced not deleted
This one is a genuine data-loss guard, so it keeps a durable trail. A
`live`-classified orphan is a merger stash holding **real uncommitted
work**, and unlike sites 1–2 there is no parked parent — the parent may
already be `done` and merged, so nothing else on the board would ever
mention the stash.

The card is replaced by a `logEntry` **and** an `addTaskComment` on the
parent, preserving every fact the old description carried: the sha,
`record.label` (the handle `git stash` recovery needs),
`record.detectedByTaskId`, and `sourcePhase`. New truthful run-audit
event `task:autostash-orphan-live-detected` replaces the borrowed
`verification:followup-*` name, with ids/outcomes-only metadata per
AGENTS.md.

### Kept unchanged: the two real product features
Eval follow-ups (`eval-followups.ts`) and PR-comment follow-ups
(`pr-comment-handler.ts`) only borrowed the shared engine for its dedup
pass. Both keep their exact behavior, column, priority, `sourceType`,
and log lines, with dedup inlined as a `listTasks` scan on
`suggestionId` / `prNumber` respectively. Both fail open (create) if the
listing throws, matching the old engine.

## Test changes — read this one

Two tests asserted the *deleted* engine's rate-limited `"[verification
recurrence]"` logEntry. Those assertions were removed, **not loosened**:
both tests still assert no duplicate card is created, and the eval test
still asserts the existing id is reported back. No coverage of surviving
behavior was weakened. The three `meta-*` test files were deleted along
with the sweeps they covered.

## Verification

```
$ pnpm test:gate
 Test Files  2 passed (2)     Tests   10 passed (10)    # core
 Test Files  16 passed (16)   Tests  299 passed (299)   # engine-core
 Test Files  1 passed (1)     Tests   70 passed (70)    # ci-shape
GATE_EXIT=0

$ pnpm --filter @fusion/engine --filter @fusion/core exec tsc --noEmit -p tsconfig.json
TSC_EXIT=0   (no output)
```

Plus a file-scoped run over the touched surfaces (`eval-followups`,
`pr-comment-handler`, `merger-autostash-orphan-surface`,
`merger-autostash-cleanup`, `run-audit`, `run-audit-secret-taxonomy`,
`project-engine`, `project-engine-manager`): **213/213 passed**.

A repo-wide grep confirms no surviving references to any deleted symbol,
module, or audit event.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Failed tasks now retain recovery and verification details directly on
the original task instead of generating separate follow-up cards.
* Live autostash issues now preserve stash information in task comments
and activity logs.
* Existing evaluation and pull-request follow-ups continue to be reused
when appropriate.

* **Changes**
  * Removed automatic archival of meta-tasks.
  * Removed obsolete meta-task timing settings.

* **Documentation**
* Updated architecture and settings documentation to reflect these
workflow changes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 22:38:58 -07:00
gsxdsm
3f33cb000f feat: per-origin workflow selection + feedback-derived refinement titles
Two task origins had no workflow picker in front of the operator and always
inherited the project default: `fn task create` (CLI + the `fn_task_create`
agent tool) and refinement tasks. Add a Project General setting for each, where
blank/unset means "Selected workflow" (the operator's current Board lane,
falling back to the project default) and a concrete id pins that origin.

Because the Board lane lives in browser localStorage, non-browser callers could
not resolve "Selected workflow" at all. `boardSelectedWorkflowId` mirrors the
lane into project settings so they can. Note this makes the mirrored lane
project-scoped: two operators on one project share it, last switch wins. The
Board never reads it back, so the only effect is which workflow a newly created
task inherits.

Resolution is `TaskStore.resolveOriginWorkflowOverrideId(origin)`: pinned
setting -> mirrored lane -> `undefined` to inherit each caller's existing
default-workflow path unchanged. A deleted or fragment id degrades to inherit
rather than throwing, so a stale settings value can never break task creation.
An explicit `workflow_id` argument to `fn_task_create` still wins.

Separately, a refinement is now titled by the operator's own feedback via the
shared `deriveFallbackTaskTitle`, not `Refinement: <parent title>`. Ten
refinements of one task previously rendered ten identical titles, so the board
could not tell them apart while the text saying what each one asked for sat in
the description. Provenance moves to a `Refines <id>` card chip alongside the
existing detail-view parent link and dependency edge.

Verified: merge gate (299 tests), lint, full build, and typecheck for core, CLI,
and dashboard all pass. New coverage: origin resolution across both origins and
the full precedence ladder, the two settings pickers, the board-lane mirror,
refinement titling (including sibling distinctness), and the card chip.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 20:33:25 -07:00
gsxdsm
256c64a7bd chore(release): v0.74.0-beta.5
Version bump via changesets.
2026-07-26 18:11:47 -07:00
gsxdsm
beebd270bd fix: make Queued to plan / Ready badges agree with the planning lane
TaskCard inferred "unplanned" from steps.length === 0 while triage's
todo-discovery and the scheduler's dispatch filter both decide from
PROMPT.md seed-ness, so the badges disagreed with the engine in both
directions: a real spec that parsed to zero steps read as "Queued to
plan" while the scheduler already treated it as a WIP-slot candidate, and
a re-seeded card still carrying old steps read as "Ready" while triage
was about to plan it. Either way the badge sent operators to the wrong
cap.

Adds the shared isTaskAwaitingPlanning predicate (replan park, missing
spec, seed-vs-real content) used by both triage's discovery and a new
best-effort `awaitingPlanning` enrichment on GET /api/tasks. TaskCard
derives both badges from that one value — strict complements — and keeps
the step count only as a fallback for SSE payloads and older servers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 17:48:46 -07:00
gsxdsm
0022621d22 chore(release): v0.74.0-beta.4
Version bump via changesets.
2026-07-26 17:00:59 -07:00
gsxdsm
2bb8537352 FN-8616: make agent tool-output limits configurable
Expose the shared agent tool-output budget as a scoped operator setting with an explicit no-limit option.

- Resolve global and project output caps with a safe finite default and zero sentinel.
- Propagate configured budgets through PI and plugin runtime tool wrappers.
- Add settings controls, localized labels, documentation, and regression coverage.

Files changed:
 .changeset/fn-8616-tool-output-budget-setting.md   |  7 ++++
 docs/agents.md                                     |  4 +-
 docs/settings-reference.md                         |  1 +
 .../core/src/__tests__/tool-output-budget.test.ts  | 23 ++++++++---
 packages/core/src/index.gate.ts                    |  2 +
 packages/core/src/index.ts                         |  2 +
 packages/core/src/settings-schema.ts               | 12 ++++++
 packages/core/src/tool-output-budget.ts            | 31 +++++++++++++--
 packages/core/src/types/settings-scope.ts          |  8 ++++
 .../app/components/settings/save-split.ts          |  1 +
 .../sections/GlobalGeneralSection.search.ts        | 20 ++++++++++
 .../settings/sections/GlobalGeneralSection.tsx     | 26 ++++++++++++
 ...lobalGeneralSection.tool-output-budget.test.tsx | 46 ++++++++++++++++++++++
 .../settings-default-descriptions.test.tsx         |  1 +
 .../src/__tests__/agent-session-helpers.test.ts    | 20 ++++++++++
 .../src/__tests__/runtime-resolution.test.ts       | 15 +++++++
 .../__tests__/tool-output-budget-wrapper.test.ts   | 45 ++++++++++++++++-----
 packages/engine/src/agent-runtime.ts               |  2 +
 packages/engine/src/agent-session-helpers.ts       | 18 +++++++--
 packages/engine/src/pi.ts                          | 29 ++++++++++----
 packages/engine/src/runtime-resolution.ts          | 10 ++++-
 packages/i18n/locales/en/app.json                  |  4 ++
 packages/i18n/locales/es/app.json                  |  6 ++-
 packages/i18n/locales/fr/app.json                  |  6 ++-
 packages/i18n/locales/ko/app.json                  |  6 ++-
 packages/i18n/locales/zh-CN/app.json               |  6 ++-
 packages/i18n/locales/zh-TW/app.json               |  6 ++-
 packages/i18n/src/resources.d.ts                   |  4 ++
 28 files changed, 323 insertions(+), 38 deletions(-)

Fusion-Task-Id: FN-8616

Fusion-Task-Lineage: 3ca99a61-d6ae-48ff-98d2-f14a153aa2b7

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 16:16:27 -07:00
gsxdsm
a9b30013bb fix(core): approval dedupe lookup matches on PostgreSQL instead of minting duplicates
`findLatestByDedupeKey` read `targetContext` through the string-only `fromJson`.
In backend (PostgreSQL) mode that column is jsonb and Drizzle returns it ALREADY
PARSED, so the dedupe scan never matched: every gate retry minted a duplicate
approval request, and an approved grant could never be redeemed. The live
database shows the signature plainly — 17 approved requests, 0 completed.

Normalize both shapes in one place (`normalizeTargetContext`), applied at
`rowToRequest` and both dedupe scan sites, so a row resolves whether it arrives
as a JSON string (SQLite) or a parsed object (Postgres).

The regression test asserts shape-independence rather than the single reported
case: the same stored key must resolve in BOTH shapes, and must not match a
different key or an absent context in either. Mutation-checked — reverting the
scan sites fails exactly the parsed-object case.

Cherry-picked ahead of #2457, which carries the wider approval/permission
hardening pass, because this one is an active production defect on its own.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 15:38:05 -07:00
gsxdsm
07c8c95b10 FN-8614: cap agent tool output
Bound every engine-injected tool result to preserve agent context capacity.

- Add shared 16,000-character total text budgets with deterministic truncation markers and validated overrides.
- Apply outermost output clamps to Pi and non-Pi plugin tool paths, with semantic caps for high-volume reads.
- Cover budget behavior and document the operator-facing configuration contract.

Files changed:
 .changeset/fn-8614-tool-output-budget.md           |  7 ++
 docs/agents.md                                     |  8 ++
 .../core/src/__tests__/tool-output-budget.test.ts  | 58 +++++++++++++
 packages/core/src/index.gate.ts                    |  7 ++
 packages/core/src/index.ts                         |  7 ++
 packages/core/src/tool-output-budget.ts            | 97 ++++++++++++++++++++++
 .../src/__tests__/agent-artifact-tools.test.ts     | 10 +++
 .../src/__tests__/agent-document-tools.test.ts     | 10 +++
 .../__tests__/agent-task-logs-read-tools.test.ts   |  8 ++
 .../__tests__/tool-output-budget-wrapper.test.ts   | 67 +++++++++++++++
 packages/engine/src/agent-session-helpers.ts       |  7 +-
 packages/engine/src/agent-tools.ts                 | 43 ++++++++--
 packages/engine/src/pi.ts                          | 54 +++++++++++-
 13 files changed, 374 insertions(+), 9 deletions(-)

Fusion-Task-Id: FN-8614

Fusion-Task-Lineage: b6a76ccd-d7b4-4b43-af7e-cfd16ffb7fc8

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 14:42:06 -07:00
gsxdsm
93a403af67 fix(dashboard): import delete-attribution constants via browser-safe subpath
The client bundle aliases `@fusion/core` to the leaf `core/src/types.ts` to
keep Node-only dependencies out of the browser, so a package-root import of
`FUSION_CLIENT_HEADER`/`FUSION_DASHBOARD_UI_CLIENT` typechecked but failed
`vite build`:

  "FUSION_CLIENT_HEADER" is not exported by "../core/src/types.ts"

Follow the documented pattern instead of widening the root alias: declare a
`./task-delete-attribution` subpath export, add the matching Vite alias ahead
of the broader `@fusion/core` key (Vite matches in order), register the module
in the browser-safe-core allowlist, and import the subpath from the client.
`task-delete-attribution.ts` has no imports at all, so it is a safe leaf.

`app/utils/detectContentLanguage.ts` already warned about exactly this trap;
the miss was mine for verifying with typecheck, lint and test:gate but not
`pnpm build`, which is one of the four checks CI blocks on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 13:22:18 -07:00
gsxdsm
15a2fb18cc Merge branch 'fix/incomplete-pg-ports'
Wire incomplete PostgreSQL ports for archive, reconcile, health, settings
cache, agent cache, and async prompt overrides.
2026-07-26 13:16:13 -07:00
gsxdsm
ab87d0d803 fix(api): return 404 for missing tasks, and make task deletions attributable
Three related fixes, all originating from a `[api:error] Request failed`
log line showing a 500 on `GET /api/tasks/FN-8610/runtime-fallback`.

1. Missing/deleted tasks now return 404 instead of 500.
   `getTaskImpl` signalled a miss with a bare `Error`, and route catches
   only mapped errno `ENOENT` to 404 — a leftover from the file-backed
   storage era. In Postgres mode nothing sets an errno code, so every
   unknown/missing/soft-deleted/wrong-project read returned 500. Adds a
   typed `TaskNotFoundError` (message byte-identical) plus a shared
   `task-lookup-error` mapper applied across the task, session-diff,
   git/GitHub, workflow and file-workspace route registrars. The same
   bare throw existed on both archive-lifecycle delete paths, so
   `DELETE /tasks/:id` was affected too.

2. 5xx logs now carry the origin stack.
   `rethrowAsApiError` constructed a fresh `ApiError` from the message
   and discarded the original, so the `FNXC:ApiErrorDiagnostics`
   contract logged the rethrow site rather than the throw site — the
   reported log entry had no stack at all. Threads `cause` through the
   error factories and walks the chain (bounded, cycle-guarded).

3. Task deletions are attributable, and non-operator deletes notify.
   `task:deleted` audit rows recorded `agentId: "system"` for every HTTP
   delete, making an operator click indistinguishable from a script or
   an agent; the calling agent's task id was accepted by the store and
   then never persisted. Adds a `callerKind` union recorded in audit
   metadata, tags every delete call site, and stamps a self-reported
   `x-fusion-client` header from the dashboard client. When the caller
   is `agent-tool` or `api-unattributed`, a best-effort notice is sent
   to the operator mailbox; operator and engine deletes stay silent.

`x-fusion-client` is attribution, not authentication — anything can send
it. No delete-blocking, gating or permission logic is added here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 13:15:36 -07:00
gsxdsm
2b55077546 fix: wire incomplete PostgreSQL ports for archive, reconcile, health
Replace empty backendMode stubs with real AsyncDataLayer paths: archive ID
reservation and isTaskArchivedAsync, orphaned task.json re-import, health
snapshots via checkPostgresHealth, settings/agent memory caches for sync
readers, async builtin prompt overrides, and self-healing audit/health
callers that previously used dead sync SQLite fallbacks.
2026-07-26 13:10:20 -07:00
gsxdsm
c5a38d8884 test: cover orphan reconcile and database-health PG stubs
Pin reconcileOrphanedTaskDirsImpl empty result and getDatabaseHealthImpl
always-healthy sentinel under backendMode so inventory category (e) stays
aligned with production self-healing and health call sites.
2026-07-26 12:58:48 -07:00
gsxdsm
d8123dc340 test: expand SQLite incomplete-PG-port inventory coverage
Drive real sync-reader stubs that empty-return under backendMode
(merge request, workflow selection/overrides/settings, run audit,
legacy step snapshot, settings/health) so category (e) of the
migration inventory stays pinned to shipped behavior.
2026-07-26 12:53:44 -07:00
gsxdsm
1290948530 test: ratchet authorized production SQLite DatabaseSync readers
Inventory analysis found exactly six read-only legacy openers; pin them in a
structural scan and assert incomplete archive guards stay SQLite-free in
backend mode so new production SQLite construction fails CI.
2026-07-26 12:45:58 -07:00
gsxdsm
3b83282273 feat(engine): attribute review-gate leases to a node so dead local leases reclaim fast
Groundwork for FN-8603's remaining ~14-minute wait. Liveness for a pending
review gate is judged purely by a 15-minute staleness floor because a lease
records WHO took it (`leaseOwner` = run id) but not WHERE, and under multi-node
every engine sees every other engine's leases. A fresh-but-unknown lease might
be running on a peer, so the floor was the only safe test -- and a lease left by
this node's own crashed process is indistinguishable from it.

Adds `WorkflowStepResult.leaseNodeId` plus an optional `LocalNodeLeaseIdentity`
argument to `classifyReviewLease`. One narrow new case: a lease stamped with the
caller's OWN node id whose `startedAt` predates the caller's process boot is
provably dead -- the process that could have owned it is gone -- so it
classifies as `reclaim` immediately rather than aging out. Deliberately narrow,
because widening it is a double-dispatch risk: absent (legacy) or peer node ids
keep the floor, and a lease taken by this process after boot is still adopted.

InProcessRuntime.start() resolves the local node id from CentralCore (fail-soft;
on error it stays undefined and floor-only semantics apply) and passes it to
SelfHealingManager. The graph executor stamps the field when deps.localNodeId is
set.

NOT YET WIRED, so this is inert in production and behavior is unchanged end to
end: `localNodeId` is not threaded from WorkflowGraphTaskRunner /
WorkflowTaskRuntime down into the executor deps, so no lease actually carries a
`leaseNodeId` yet. The reader is ready; the writer needs that pass-through
(WorkflowGraphTaskRunnerDeps gains the field, the runner forwards it, and the
runtime supplies this.localNodeId). Stopping here rather than half-threading it.

Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green
(299 + 70), core workflow-step-results suite green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 12:24:29 -07:00
gsxdsm
00011b0113 fix(engine): recover restart-orphaned review steps in one cycle, raise fix budget
FN-8603 sat in-review for ~36 minutes after an engine restart killed its Code
Review session 34 seconds in. It did recover on its own; the cost was latency,
not a terminal park.

Sweep ordering. reconcile-orphaned-pending-step-results PRODUCES the failed
results that recover-failed-pre-merge-steps CONSUMES, but in the periodic
maintenance list it ran ~15 entries after it. A step orphaned in cycle N was
therefore rewritten to failed only after recovery had already scanned, so
nothing re-ran it until cycle N+1. Moved it immediately before its consumer and
removed the now-duplicated later entry. Startup recovery already ordered the two
correctly.

Post-review fix budget. Default raised 3 -> 10 per operator request. Three
passes is below the observed convergence length for the gates this fallback
actually governs -- Browser Verification and custom optional gates -- since Plan
Review and Code Review already resolve to "unbounded" when unset, and exhausting
the budget parks the card for a human. The declaration default and five inline
`settings.maxPostReviewFixes ?? 3` call sites in executor.ts/self-healing.ts had
drifted into separate literals, so raising one alone would have left every
unset-settings path on the old value; they now share the exported
DEFAULT_MAX_POST_REVIEW_FIXES.

Not done, and why. Re-dispatching a restart-orphaned lease immediately at
startup is the change that would close the remaining ~14-minute wait, but it is
unsound as specified: liveness is judged by a 15-minute lease-staleness floor
because leases carry no node attribution, so treating a pre-boot lease as dead
would let one node orphan another node's genuinely running review. Needs a node
id on the lease record first. Left the floor intact.

Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green,
self-healing orphaned-pending-step-results and optional-step-revision suites
green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 12:08:03 -07:00
gsxdsm
cca13737b6 FN-8603: reduce steady-state diagnostic log noise
Route routine core, engine, and dashboard diagnostics through debug-gated shared loggers.

- Demote steady-state diagnostic sites while preserving warnings and errors for actionable failures.
- Add cross-package severity contracts and manifest coverage for demoted log sites.
- Document logging severity guidance and add a patch changeset.

Files changed:
 .changeset/fn-8603-log-severity.md                 |  7 ++
 docs/diagnostics.md                                | 20 ++++--
 .../__tests__/log-severity-spam-contract.test.ts   | 71 ++++++++++++++++++
 packages/core/src/activity-analytics.ts            |  5 +-
 packages/core/src/ai-summarize.ts                  | 61 +++++++---------
 packages/core/src/async-mission-store.ts           |  5 +-
 packages/core/src/async-secrets-store.ts           |  7 +-
 packages/core/src/central-core.ts                  | 17 ++---
 packages/core/src/docker-provisioning.ts           | 13 ++--
 packages/core/src/index.ts                         |  1 +
 packages/core/src/master-key.ts                    |  9 ++-
 packages/core/src/memory-compaction.ts             | 29 ++++----
 packages/core/src/memory-insights.ts               |  7 +-
 packages/core/src/migration-orchestrator.ts        |  7 +-
 packages/core/src/mission-store.ts                 |  5 +-
 packages/core/src/node-discovery.ts                |  7 +-
 packages/core/src/notification/dispatcher.ts       |  9 ++-
 .../core/src/plugins/bundled-plugin-install.ts     | 11 +--
 packages/core/src/reflection-store.ts              |  5 +-
 packages/core/src/secrets-store.ts                 |  7 +-
 packages/core/src/task-store/agent-logs.ts         | 21 +++---
 packages/core/src/task-store/async-events.ts       |  5 +-
 packages/core/src/task-store/async-maintenance.ts  |  7 +-
 packages/core/src/task-store/comments-ops.ts       |  7 +-
 packages/core/src/task-store/task-mutation-ops.ts  | 11 +--
 packages/core/src/task-store/workflow-integrity.ts |  9 ++-
 packages/core/src/types/merge-policy.ts            |  5 +-
 packages/core/src/usage-events.ts                  |  5 +-
 .../__tests__/log-severity-spam-contract.test.ts   | 48 +++++++++++++
 packages/dashboard/src/ai-refine.ts                |  5 +-
 packages/dashboard/src/ai-session-diagnostics.ts   | 10 +--
 packages/dashboard/src/chat.ts                     |  8 ++-
 packages/dashboard/src/devserver-manager.ts        |  9 ++-
 packages/dashboard/src/file-service.ts             |  5 +-
 packages/dashboard/src/github-tracking-comments.ts |  7 +-
 .../dashboard/src/github-tracking-reconciler.ts    |  5 +-
 packages/dashboard/src/github-tracking-state.ts    |  5 +-
 packages/dashboard/src/gitlab-lifecycle.ts         |  5 +-
 packages/dashboard/src/insights-routes.ts          |  9 ++-
 packages/dashboard/src/issue-image-attachments.ts  |  5 +-
 packages/dashboard/src/knowledge-index.ts          |  5 +-
 packages/dashboard/src/plugin-routes.ts            |  7 +-
 packages/dashboard/src/routes/board-workflows.ts   |  5 +-
 packages/dashboard/src/routes/context.ts           |  5 +-
 .../dashboard/src/routes/register-auth-routes.ts   | 13 ++--
 .../routes/register-docker-provisioning-routes.ts  |  7 +-
 .../dashboard/src/routes/register-git-github.ts    | 21 +++---
 packages/dashboard/src/routes/register-gitlab.ts   |  7 +-
 .../src/routes/register-session-diff-routes.ts     |  9 ++-
 .../src/routes/register-settings-memory-routes.ts  |  7 +-
 .../src/routes/register-setup-activity-routes.ts   |  7 +-
 .../dashboard/src/routes/register-signal-routes.ts |  5 +-
 .../src/routes/register-task-workflow-routes.ts    | 11 +--
 packages/dashboard/src/runtime-logger.ts           | 11 +--
 packages/dashboard/src/server.ts                   |  7 +-
 packages/dashboard/src/sse.ts                      |  8 ++-
 packages/dashboard/src/terminal-service.ts         | 34 ++++-----
 packages/dashboard/src/view-chunk-manifest.ts      |  5 +-
 .../engine/src/__tests__/log-severity-manifest.ts  | 83 ++++++++++++++++++++++
 .../__tests__/log-severity-spam-contract.test.ts   | 40 ++++++++++-
 .../src/__tests__/logger-debug-gating.test.ts      |  7 +-
 packages/engine/src/goal-anchoring-audit.ts        |  5 +-
 packages/engine/src/plugin-runner.ts               | 44 ++++++------
 packages/engine/src/pty-native.ts                  |  9 ++-
 .../engine/src/runtimes/child-process-worker.ts    |  4 +-
 packages/engine/src/self-healing.ts                | 12 ++--
 packages/engine/src/worktree-hooks.ts              | 10 ++-
 67 files changed, 632 insertions(+), 250 deletions(-)

Fusion-Task-Id: FN-8603

Fusion-Task-Lineage: 53901db6-1af2-4bd7-b5ea-49507e048ef2

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 12:01:19 -07:00
gsxdsm
71279ed042 fix(FN-8600): recover a duplicate verdict the planner reported in its reply
The prompt fix stops planners writing the verdict in prose, but it relies on
every model reading one sentence correctly. This closes the hole underneath it.

When the finalize read finds no spec at all, the planner's streamed reply is
searched for a line that is exactly `DUPLICATE: FN-NNNN`. If found, the engine
writes the canonical marker file and continues — so marker parsing, keep/delete
resolution, and the sourceMetadata.nearDuplicateOf that renders the operator's
decision all run on the unchanged file contract rather than a second code path
that could drift from it.

Deliberately narrow. The marker must occupy a whole line, only the first counts,
and recovery is gated on the plan being genuinely absent — a planner that wrote
a real spec is never overridden by something it said in passing. The text tail
is bounded because the verdict lands in the closing summary, and it tees off
onText rather than reading AgentLogger, whose buffer is flushed on a timer.

Verified both directions: the tests fail without the recovery block, and the
"wrote a real spec while mentioning a marker" case keeps its spec.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 10:23:00 -07:00
gsxdsm
9bad0e1233 fix(engine): demote high-frequency TUI log spam to debug
Route process spawn/exit, verification success paths, MCP connect, skill info listings, createFnAgent/session bookkeeping, and executor dispatch chatter through FUSION_DEBUG so the operator log pane keeps real lifecycle outcomes.
2026-07-26 09:50:43 -07:00
gsxdsm
86c892b67c fix(FN-8600): make the duplicate-report instruction writable, not self-contradictory
The planning prompt said "do not write PROMPT.md" and, in the same breath,
"write DUPLICATE: {id} to the output file" — where the output file IS
PROMPT.md. A planner that took the first clause literally wrote no file and
reported the duplicate in prose.

The engine only ever reads the verdict from PROMPT.md's contents, so that
duplicate was invisible: the task failed deterministic validation as
"PROMPT.md file not found or empty", retried, terminalized to failed, emitted
a task-wedge mail, was recovered to todo by self-healing, and re-planned —
three full Opus planning cycles on FN-8600 before it was caught, with no
operator decision ever surfaced because sourceMetadata.nearDuplicateOf is only
set on the branch that parses the file.

Both prompt sites now say to write PROMPT.md with the marker as its entire
contents, and say why prose alone is not recorded.

Note the engine ordering is already correct — tryFinalizeExplicitDuplicateMarker
runs before validateGeneratedPrompt, and a worktree-local spec is recovered
first. Nothing to reorder; the file simply never existed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 09:42:24 -07:00
gsxdsm
ae512aec2b FN-8601: enforce foreach merge proof
Require complete foreach execution evidence before workflow merge review.

- Add reusable foreach instance coverage proof evaluation.
- Block checklist projection and merge admission on incomplete or failed node results.
- Cover core proof logic and PostgreSQL merge-boundary behavior.
- Add a patch changeset for the merge safeguard.

Files changed:
 .changeset/fn-8601-foreach-merge-proof.md          |   7 ++
 .../src/__tests__/workflow-merge-proof.test.ts     |  43 ++++++++
 packages/core/src/index.gate.ts                    |   2 +
 packages/core/src/index.ts                         |   2 +
 packages/core/src/workflow-merge-proof.ts          |  74 +++++++++++++
 ...xecutor-merge-boundary-foreach-proof.pg.test.ts | 111 +++++++++++++++++++
 packages/engine/src/executor.ts                    | 117 +++++++++++++--------
 7 files changed, 314 insertions(+), 42 deletions(-)

Fusion-Task-Id: FN-8601

Fusion-Task-Lineage: 40578171-0b13-4538-8f38-3948ed1e92c0

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 09:40:37 -07:00
gsxdsm
d4aa79b66c FN-8598: preserve legacy task cost badges
Restore cost badges for tasks with valid legacy token totals.

- Preserve usage records when optional timestamps and cache-write totals are absent
- Use task creation time to satisfy legacy usage timestamp requirements
- Cover card badge rendering, unpriced mixed usage, and mobile visibility
- Add a patch changeset for the restored badge behavior

Files changed:
 .changeset/fn-8598-cost-badge-fix.md               |   7 +
 .../task-token-usage-serialization.test.ts         |  45 +++++++
 packages/core/src/task-store/serialization.ts      |  16 ++-
 .../__tests__/TaskCard.cost-badge.test.tsx         | 146 +++++++++++++++++++++
 .../app/utils/__tests__/taskTokenCost.test.ts      |  11 ++
 5 files changed, 221 insertions(+), 4 deletions(-)

Fusion-Task-Id: FN-8598

Fusion-Task-Lineage: 83fb4051-8e0f-4ee9-9f00-0e4d5cb8661e

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 09:06:41 -07:00
gsxdsm
795a38c018 fix(engine): quiet graph review-entry audits and label engine aborts truthfully
Recognise workflow-graph moves into in-review so gate entry no longer emits handoff-invariant violations, and split pause-abort provenance so engine teardowns are engine-abort instead of hard-cancel.
2026-07-26 08:56:58 -07:00
gsxdsm
af897d9e3c FN-8596: isolate cross-root plugin MCP discovery
Prevent cross-root MCP discovery from unloading active plugin runtimes.

- Isolate discovery loader lifecycle and runtime-state persistence.
- Preserve shared plugin owners when non-owner loader participants stop.
- Cover core, dashboard, and engine cross-root discovery behavior.

Files changed:
 .changeset/fn-8596-plugin-discovery-isolation.md   |   7 ++
 .../plugin-loader-lifecycle-scope.test.ts          |  12 +++
 .../plugin-mcp-servers-discovery-isolation.test.ts | 115 +++++++++++++++++++++
 packages/core/src/plugin-loader.ts                 |  34 +++++-
 packages/core/src/plugin-mcp-servers.ts            |   8 +-
 .../context-plugin-mcp-discovery-isolation.test.ts |  48 +++++++++
 packages/dashboard/src/routes/context.ts           |  39 ++++++-
 ...-runtime-plugin-mcp-discovery-isolation.test.ts |  69 +++++++++++++
 packages/engine/src/runtimes/in-process-runtime.ts |  47 +++++++--
 9 files changed, 364 insertions(+), 15 deletions(-)

Fusion-Task-Id: FN-8596

Fusion-Task-Lineage: 231e53b6-a9a3-4a65-9732-3dabe44da198

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 08:35:04 -07:00
gsxdsm
26dcccb7c3 fix(workflow): harden review-gate lifecycle interactions in In review
Follow-ups to running the pre-merge review gates in `in-review`. Each was
verified against the code before being fixed; one reported issue was
refuted and is noted below.

1. Symbol locks (packages/core/src/task-store/moves.ts)
   FN-8306 made the lifecycle transition the symbol-lock RELEASE authority
   but wrote no counterpart. That was harmless while a task only left WIP
   at handoff/terminal; the gate crossing now releases the task's declared
   symbols and the remediation node re-enters `in-progress` to edit the
   same files in the same live worktree with its locks gone. Neither
   acquire site (scheduler dispatch, claimDueWorkflowWorkItem) is on the
   graph re-entry path. Adds a symmetric re-acquire on `!wip -> wip`.
   Best-effort by design: a contended symbol logs and proceeds, which is
   exactly the pre-fix posture, rather than parking the remediation behind
   another holder and re-creating the stranding this change set removed.

2. Premature merge (packages/engine/src/self-healing.ts)
   `recoverMergeableReviewTasks` was the only in-review sweep with no
   liveness gate. The graph commits the column crossing at node entry and
   writes the gate's pending lease two DB round trips later, and
   `getTaskMergeBlocker` has no notion of "enabled but resultless", so in
   that window the sweep could enqueue a merge with Code Review never run.
   Filters `executingIds`, matching recoverGhostReviewTasks.

3. Orphan sweep (packages/engine/src/self-healing.ts)
   The reported restart hazard is REFUTED: nothing re-attaches an in-review
   graph run, so those leases are genuinely dead and marking them failed is
   correct FN-8492 behavior. But the sweep also runs from periodic
   maintenance in the same live process, where a tick between the lease
   write and session registration could fail a gate that just started.
   Honors a within-floor `classifyReviewLease`, matching the semantics Plan
   Review already had. Cleanup of dead leases is delayed by the staleness
   floor, not defeated. The audit event gains `needsOperatorBypass` for
   `autoMerge:false` rows, which self-healing deliberately skips and only
   fn_task_bypass_review can clear — previously indistinguishable from an
   auto-recoverable rewrite.

4. Stall detection (packages/engine/src/planner-overseer.ts)
   The `reviewer` and `merger` stages had no time-based check at all and
   returned `progressing` unconditionally, so a hung gate produced no
   signal however long it sat. Adds gate-anchored detection on both (a
   plain in-review card with no reviewState resolves to `merger`, not
   `reviewer`), keyed on the pending lease's own `startedAt` rather than
   `columnMovedAt` so it cannot fire during a legitimate human merge-wait.

`cumulativeActiveMs` is documented, not changed: it now excludes gate
runtime, but adding the `timing` trait to `in-review` would count arbitrary
human merge-wait as active work — a worse distortion than the omission.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 02:28:04 -07:00
gsxdsm
47d030215c feat(workflow): run pre-merge review gates in the In review column
Code Review and Browser Verification now run with the card in `in-review`
instead of `in-progress`, so the board shows the card under review with the
running step as a badge (matching the Coding (Ideas) preset). Their paired
remediation nodes stay in `in-progress`, so a changes-requested verdict
visibly sends the card back to implementation.

The column move IS the badge switch: the dashboard badge was already
lane-gated on `column === "in-review"`. Applied to the shared stepwise
coding IR, so it is inherited by builtin:coding (the default),
builtin:stepwise-coding, builtin:brainstorming and builtin:coding-ideas;
builtin:legacy-coding keeps its historical placement.

Two consequences handled:

- Capacity: `in-review` has no `wip` trait, so the slot is released during
  review and the remediation crossing back into `in-progress` can hit the
  non-bypassable in-transaction capacity check. The column boundary now
  PARKS the run on a `capacity-exhausted` rejection instead of failing it,
  preserving the failed gate result and worktree so the next graph run
  retries once a slot frees. Non-capacity rejections still propagate.

- Reopen clears: `applyReopenFieldClears` wiped `workflowStepResults` on
  every in-review -> in-progress move, which the remediation crossing now
  performs routinely. That destroyed the remediation input, made
  `routeRetryableRemediationGraphFailureToPreMergeFix` and
  `recoverFailedPreMergeWorkflowStep` silently no-op, and — worse — made
  both `getTaskMergeBlocker` branches vacuously false, so a card could
  return to `in-review` and be mergeable with its gate never re-run. Now
  exempted for graph-owned in-review -> in-progress crossings only;
  operator reopens, merge bounces and every -> todo/triage rebound still
  clear, so the executor's documented bounce invariant is unchanged.

Adds regression coverage for both (there was previously none for the
reopen clear in either direction), and annotates the unreachable legacy
scheduler dispatch block rather than mirroring the fix into dead code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-26 02:04:40 -07:00
gsxdsm
fd073e287f FN-8592: self-heal stranded hold continuations
Restore graph-owned plan-review continuations for eligible hold-column cards stranded after planning cancellation.

- Detect real-spec hold cards with no active workflow continuation and re-seed Plan Review safely.
- Serialize workflow continuation seeding, review-result writes, and lease claims to prevent duplicate recovery.
- Add recovery diagnostics, release warnings, regression coverage, and a patch changeset.

Files changed:
 .changeset/fn-8592-stranded-hold-continuation.md   |   7 +
 AGENTS.md                                          |   1 +
 docs/architecture.md                               |   4 +
 .../workflow-task-serialization-protocol.test.ts   | 119 +++++++++++++
 .../workflow-work-items-conditional-seed.test.ts   | 191 +++++++++++++++++++++
 packages/core/src/store.ts                         |   5 +-
 .../src/task-store/async-workflow-workitems.ts     | 123 +++++++++----
 packages/core/src/task-store/project-store-ops.ts  |  14 ++
 .../src/task-store/workflow-task-create-ops.ts     |  16 +-
 .../src/task-store/workflow-workitems-ops-2.ts     |  91 ++++++----
 .../src/__tests__/pre-release-plan-review.test.ts  |  17 ++
 ...self-healing-stranded-hold-continuation.test.ts | 171 ++++++++++++++++++
 packages/engine/src/hold-release.ts                |  57 +++++-
 packages/engine/src/plan-review-continuation.ts    |  94 ++++++++++
 packages/engine/src/runtimes/in-process-runtime.ts |  30 +---
 packages/engine/src/self-healing.ts                | 100 ++++++++++-
 16 files changed, 945 insertions(+), 95 deletions(-)

Fusion-Task-Id: FN-8592

Fusion-Task-Lineage: fe7ffd34-96e4-4418-a879-7418e6293d30

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-26 00:46:07 -07:00
gsxdsm
106c61e6ee fix(agent-tools): close the fn_delegate_task Deny bypass and the store's window clamp
Follow-up to 13a2b2a9d, from a multi-agent review of that commit. Three of its
claims did not hold.

1. fn_delegate_task bypassed the gate entirely (P0). It reaches the same
   createAgentTask primitive, was registered unconditionally in both session
   lanes, and validated only that the TARGET agent is non-ephemeral — never the
   caller. Under Deny an ephemeral worker could enumerate agents and delegate
   unlimited tasks. It is now withheld under Deny, and also under
   upon_validation: delegation has no proposal channel, so leaving it available
   would launder a create past the operator review that policy requires.

2. The widened dedupe window was capped at 5 minutes. The store query in
   branch-and-pr-entities.ts carried its own independent `?? 60_000` /
   `min(300_000, …)` pair, so widening only duplicate-guard.ts under-delivered
   and made the new ceiling unreachable. Both sites now share
   FINGERPRINT_WINDOW_DEFAULT_MS / FINGERPRINT_WINDOW_MAX_MS.

3. The pi-extension gate does not fire at all. pi's ExtensionContext carries no
   agentId — the read is a speculative cast and only tests supply one, so every
   real call short-circuits as a human caller. The fail-closed direction is kept
   for the day an identity signal exists, but the limitation is now documented
   instead of implied to be enforcement.

Also: the session prompt now states when creation is disabled and names
fn_task_log as the fallback (the base prompt still taught fn_task_create, which
is the same instruction/capability mismatch that fed the retry storm);
suppression emits an `agent:task-create-withheld` run-audit event; and the two
source-text ratchet tests are replaced with behavioral assertions on the tool
list the executor actually hands the model — verified to fail when the guard is
broken, which the string assertions did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:44:56 -07:00
gsxdsm
13a2b2a9da fix(agent-tools): hide fn_task_create under Deny and widen the dedupe window
Operator report: with project policy "Ephemeral agent follow-up tasks = Deny",
an executing agent filed ten follow-up tasks — five parallel fn_task_create
calls it reported as timed out, then five sequential retries.

Two defects:

1. Deny was advisory. fn_task_create was registered for every session and only
   refused inside execute(), so the model still saw the tool, planned around it,
   and retried it. The pi extension's isEphemeralCallerAgent also failed OPEN
   whenever the caller id did not resolve to an agent row — which is the normal
   shape of an ephemeral task-worker — so on that lane Deny was a no-op.

2. The deterministic content-fingerprint duplicate window was 60s, which only
   covered concurrent in-flight creates. A retry two minutes later saw nothing
   and filed a second task.

Fixes: isAgentTaskCreateToolAvailable() withholds the tool from ephemeral
sessions under Deny in both engine lanes (outer execution session, per-step
workflow session); isEphemeralCallerAgent fails closed on an unresolvable
caller id; the fingerprint window goes 60s -> 10m (clamp ceiling 5m -> 1h).
upon_validation keeps the tool, and permanent-agent and human/chat callers are
unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 23:03:46 -07:00
gsxdsm
99b80ad748 feat(dashboard): add opt-in auto-update and harden restart supervision
Add the `autoUpdateAndRestart` global setting (default off, Settings ->
General next to Release channel). When enabled, the dashboard host installs
available updates on the selected channel by itself and requests the
supervised in-place restart. Supervised hosts only: without a parent to
respawn, installing would leave a running process whose code no longer
matches its own install.

Fix two ways the restart affordance could silently do nothing:

- The supervisor now stamps FUSION_SUPERVISOR_PID and supervision is only
  counted when that pid is the real parent. FUSION_RESTART_SUPERVISED is
  inherited by every process Fusion spawns, so `fn dashboard` launched from
  an agent terminal skipped its own supervisor while still advertising
  restart support -- a restart request then killed it for good.
- Settings and the update banner probe /system/info on mount and treat
  capability as advisory: the button always issues the request and shows the
  server's actual refusal instead of sitting disabled after a failed probe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 22:52:53 -07:00
gsxdsm
bf317f6340 chore(release): v0.74.0-beta.3
Version bump via changesets.
2026-07-25 21:06:33 -07:00
gsxdsm
1efff4e83c test(core): speed up schema-applier PG tests by dropping psql subprocess spawns
Route admin CREATE/DROP DATABASE through a short-lived postgres.js maintenance
connection instead of spawning psql via execSync per call, and remove the
redundant DROP-before-CREATE (db names are pid+random, never pre-exist).
Cuts ~2 of 3 subprocess forks per test across ~55 tests; the slowest core
test file drops from ~90s under full-suite contention (32.6s->27s standalone)
with all 75 tests still green.

Fusion-Task-Id: FN-SLOW-TEST
2026-07-25 17:20:18 -07:00
gsxdsm
2560944663 chore(release): v0.74.0-beta.2
Version bump via changesets.
2026-07-25 16:16:07 -07:00
gsxdsm
a0496c175c chore(release): v0.74.0-beta.1
Version bump via changesets.
2026-07-25 10:08:15 -07:00
gsxdsm
10df734bd1 fix(core): seed the prompt for quick-add Start creates; instrument hold-release
Quick-add "Start" collapses create+promote into one request: it submits the
workflow id AND the post-intake `todo` column together, so the card lands in
`todo` having never sat in the workflow's manual intake column. The intake test
in task-creation.ts only matched `triage` or the resolved intake column, so the
card got generateSpecifiedPrompt — whose hard-coded boilerplate steps
("Implement the required changes") no planner ever wrote.

That stranded the card permanently: triage's todo-discovery admits a card only
when its PROMPT.md reads as a seed, so the placeholder spec was classified
"already planned" and never planned, while nothing could execute it either
(steps: []). It sat in Todo forever with no log line in any lane. Observed on
FN-8587.

Creates into `todo` on a manual-intake workflow (resolved intake is not the
legacy `triage`) now get the bootstrap seed. The pinned contract for a plain
direct create into todo on the default workflow — which intentionally keeps
generateSpecifiedPrompt — is untouched, and both create sites are fixed in step.

Also instrument the hold/release sweep, which had reasons but no timings:
per-task held duration reported on release, a per-sweep summary breaking out the
prefetch cost (a sequential await per non-archived task, so it scales with board
size rather than with held cards), and a warn when a sweep exceeds 2s — so a
"ready card doesn't move" delay can be attributed between poll cadence, sweep
cost, and a card genuinely queued on capacity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 09:52:18 -07:00
gsxdsm
b5578448fa fix(workflow): start planning immediately when a task is started
Pressing Start on a Coding (Ideas) card only writes a column move — there is no
dispatch call in that path — so planning did not begin until the triage
processor's next timer tick, up to pollIntervalMs (15s default) later. The
"Started planning" toast was optimistic and the card just sat in Todo.

- Wake planning discovery on the store's task:updated/task:created event when a
  task lands in todo/triage. Binding the wake to the store event rather than the
  Start button covers every move surface (board drag, context menu, task detail,
  List view, CLI, agent tools, POST /tasks/:id/move) by construction. The wake is
  advisory: it only advances WHEN the poll runs, so every pause, seed-prompt,
  dependency, and concurrency gate still applies.
- Admit a todo task whose PROMPT.md is missing instead of dropping it through a
  silent `catch {}`. The scheduler KEEPS a candidate whose prompt it cannot read,
  so such a card was invisible to planning while still visible to dispatch, with
  no log line in either lane. Unreadable (non-ENOENT) prompts now log.
- Route the scheduler's dispatch filter through the shared isUnplannedSeedPrompt
  predicate. Its open-coded strict bootstrap compare disagreed with triage on the
  refinement-seed shape, leaving hold-release as the only thing between an
  executor and a prompt containing just the operator's feedback text. The
  predicate also normalizes line endings/trailing whitespace, so a CRLF or
  trailing-newline round-trip no longer reclassifies an unplanned card as planned.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 09:14:11 -07:00
gsxdsm
927efb1477 fix(workflow): allow Coding (Ideas) cards to move back from Todo to Ideas
A legacy source column (todo/in-progress/...) validated moves only against the
closed VALID_TRANSITIONS map, which cannot know about a workflow-declared
column, so Todo -> Ideas was rejected even though the board drag pre-check and
context menu both offered it. Legacy sources now union VALID_TRANSITIONS with
the task's workflow-resolved adjacency, resolved lazily only when the legacy
table alone would reject. builtin:coding adjacency is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-25 08:45:52 -07:00
gsxdsm
9a2aea6120 FN-8571: add Parakeet STT transcription backend
Add a configurable sherpa-onnx Parakeet v3 speech-to-text service and dashboard API.

- Add model download, validation, caching, and lifecycle management for the bundled STT runtime.
- Register multipart-safe voice transcription routes with request-size handling and service-backed responses.
- Expose STT settings, documentation, release metadata, and focused integration coverage.

Files changed:
 .changeset/fn-8571-voice-stt-backend.md            |   7 +
 docs/settings-reference.md                         |  12 +
 .../core/src/__tests__/settings-parity.test.ts     |  10 +
 packages/core/src/index.ts                         |   2 +-
 packages/core/src/settings-schema.ts               |   2 +
 packages/core/src/types.ts                         |   2 +
 packages/core/src/types/settings-scope.ts          |  19 ++
 packages/dashboard/package.json                    |   3 +
 .../voice-body-parser-integration.test.ts          |  67 ++++++
 packages/dashboard/src/routes.ts                   |   2 +
 packages/dashboard/src/routes/README.md            |  74 +++---
 .../routes/__tests__/register-voice-routes.test.ts | 159 ++++++++++++
 .../src/routes/create-api-routes-mount-sequence.ts |   2 +-
 .../dashboard/src/routes/register-voice-routes.ts  |  92 +++++++
 packages/dashboard/src/server.ts                   |  21 +-
 .../src/stt/__tests__/model-manager.test.ts        | 171 +++++++++++++
 .../dashboard/src/stt/__tests__/voice-stt.test.ts  |  48 ++++
 packages/dashboard/src/stt/model-manager.ts        | 268 +++++++++++++++++++++
 packages/dashboard/src/stt/parakeet-service.ts     |  77 ++++++
 packages/dashboard/src/stt/types.ts                |  14 ++
 pnpm-lock.yaml                                     |  65 +++++
 21 files changed, 1086 insertions(+), 31 deletions(-)

Fusion-Task-Id: FN-8571

Fusion-Task-Lineage: b11e0cea-9c9c-41d4-a141-c47e5ed56dba

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-25 02:17:31 -07:00
gsxdsm
0056d75314 FN-8569: surface unrecoverable report health
Classify parked direct reports as operator-actionable even when their stored state appears live.

- Add a reusable Reports Health classifier that prioritizes pause markers.
- Clear stale pause markers during live-state resumes without removing diagnostic errors.
- Cover desynchronized report states and document the health invariant.

Files changed:
 .../fn-8569-reports-health-error-unrecoverable.md  |  7 +++
 docs/architecture.md                               |  1 +
 .../agent-store-pause-marker-clear.test.ts         | 65 +++++++++++++++++++
 packages/core/src/agent-store.ts                   | 12 ++++
 .../src/__tests__/heartbeat-executor.test.ts       | 27 ++++++--
 .../engine/src/__tests__/reports-health.test.ts    | 73 ++++++++++++++++++++++
 packages/engine/src/agent-heartbeat.ts             | 36 ++++++-----
 packages/engine/src/index.ts                       |  6 ++
 packages/engine/src/reports-health.ts              | 70 +++++++++++++++++++++
 9 files changed, 278 insertions(+), 19 deletions(-)

Fusion-Task-Id: FN-8569

Fusion-Task-Lineage: 37c798a0-1f2b-4221-a0c6-ccbff8d72696

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-07-24 23:57:49 -07:00