Test-only, no production change. Independent of my other open PRs.
## The number nobody was counting
This program has two censuses, and **both count converted things**: the
unproven-sites ledger (callers of the lifecycle-role resolvers) and
`raw-workflow-columns-flag-census` (reads of the `workflowColumns`
flag).
Neither counts what is still keyed to a legacy column id — **which is
where every defect this program has found actually lived**:
| defect | the literal |
|---|---|
| pool-id sentinel (capacity gate never bound) | `?? "builtin:coding"`
vs the counter's sentinel |
| agent-link leak (slot consumed forever) | terminal column matched
against a fixed id set |
| stale-paused badge silent on renamed boards | `task.column !== "todo"`
|
| merge chokepoint threw on a finished card | the `done`/`archived` pair
|
| recovered card stranded harder | `?? "todo"` |
Every one was found **by hand, one at a time, by whoever happened to
look.**
## Measured
**438 lifecycle decisions keyed to a legacy column name** (417
comparisons + 31 `??` column fallbacks, minus 8 agent-id false positives
and 2 lines carrying both shapes), across 85+ production files — 94 in
`self-healing.ts`, 70 in `executor.ts`, 26 in the dashboard
task-workflow routes.
That is the real size of the remaining surface. It dwarfs the 15-site
resolver census I've spent this unit closing, which is worth knowing
before anyone calls the vocabulary work finished.
## A hit is not a bug
Many are correct — documented legacy fallbacks, the legacy-adoption
path, code genuinely about the built-in workflow. The census claims only
that each site decides by **name** rather than by **role**, and
therefore needs a human judgment. Reporting 417 as a bug count would be
exactly the overclaiming this program keeps correcting.
## A ceiling, not an equality — deliberate
The sibling flag census fails in both directions. That number moves only
when two units touch it. **This** one moves whenever any of a dozen
concurrent conversion slices lands, and an exact-equality assertion
would go red on work heading the *right* way.
A test that's red for good reasons gets suppressed, and a suppressed
ratchet is worse than none — the failure mode AGENTS.md's quarantine
rule exists to prevent. So the count may fall freely and may never rise;
when it falls, the failure message says to lower the pin.
## Verified in both directions
- green at 417
- adding **one** literal to `replan-target.ts` → `census ROSE to 418
(ceiling 417)`
- the regex is unit-tested to count a **decision**, not a mention: a
column id in a fixture, a log line, or a `moveTask` argument is not
counted — inflating the number into noise is how a census stops being
acted on
- unreadable sources **fail closed** rather than silently shrinking the
count
## Follow-up (a8c150b12): the census was blind to three of the five
defects it cites
I ran the census against its own header. It lists five motivating
defects; the comparison-only regex counted **two**. The pool-id
sentinel, the rebound strand and the terminal fallback are all `??`
**defaults** — invisible to a `.column === "x"` pattern.
A census that cannot see three of the five bugs it names as its reason
to exist is worse than none: it reports a number that *feels* like
coverage. That is precisely the overclaim this unit keeps catching in
other people's work — caught here in mine, and only because the header
wrote the examples down somewhere they could be tested against.
It now counts two shapes — deciding **by** a name (`===`/`!==`) and
**defaulting** to one (`??`) — and pins the five motivating examples as
a test case, so the pattern cannot narrow back without failing.
**Measured: 417 comparisons + 31 fallbacks, of which 2 lines carry both
shapes → 446 lines.** Ceiling raised 417 → 446 to cover the missing
shape, not to excuse new debt.
`?? "builtin:coding"` stays deliberately uncounted: it defaults a
*workflow* id rather than a column and is legitimately correct at most
sites. It already has a stronger guard —
`scripts/check-capacity-pool-id.mjs` bans it only where the value
reaches a capacity counter, which is the only place it's wrong.
Verified both directions: green at 446; adding one fallback of the
newly-counted shape → `census ROSE to 447 (ceiling 446)`.
## Follow-up 2 (98f4264fd): 8 false positives removed — 446 → 438
Then I checked the census against real source instead of trusting the
pattern. Its top-scoring fallback file was `triage.ts` with 8 hits — and
**every one is `agentId: task.assignedAgentId ?? "triage"`**, an *agent*
id, not a column. `"triage"` is both a column id and the synthetic agent
id triage stamps on its audit rows.
Eight of ~34 fallbacks is a quarter of that shape: enough to make the
number **wrong** rather than merely imprecise. A census with known false
positives is one people learn to discount — the same end state as not
having one, which is exactly what its own header warns about.
Excluded, and the exclusion is **pinned as a test case** so it can't
creep back: the three agent-id spellings must match the raw shape *and*
be filtered, while a genuine column fallback that also mentions triage
(`first("intake") ?? "triage"`) must still count.
**Residual imprecision is stated rather than tuned away.** A couple of
counted lines are display defaults (a column rendered in CLI output).
They stay: the census claims each site *needs a human judgment*, and a
display default passes that judgment in seconds. Chasing them costs more
than the precision buys and makes the pattern too clever to trust.
Agent-ids were excluded because they're a quarter of the shape — not
because any false positive is intolerable.
Ceiling 446 → **438**. Verified both directions: green at 438; one new
fallback → `census ROSE to 439`.
## Verification
- census 3/3; engine `tsc --noEmit` clean; `pnpm test:gate` green (414 +
10 + 71)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**U9, PR9.** Config only — one line added to the `engine-core`
allow-list, plus its justification.
The merge half of U9's safeguards now fires in blocking CI (#2526).
**This is the review half, none of which did.**
## What's admitted
`workflow-step-verdict-parsing.test.ts` holds
`proseSignalsClearApproval`'s leniency guard: **a prose REJECTION must
never be promoted to APPROVE.** Removing the
REVISE/RETHINK/negated-approval disqualifiers fails **11** of its cases.
This is a **fail-open** defect on the path to an irreversible merge — a
review saying *"looks good, but this must be fixed before merging"*
would read as an approval. That belongs in the gate, not in a
non-blocking run hours after the merge.
Measured across 3 runs:
| | Files | Tests | Wall |
|---|---|---|---|
| before | 19 | 414 | 6.16 / 6.25 / 6.21s |
| after | 20 | 482 | 6.28 / 6.55 / 6.33s |
**+~0.2s** against a ~60s ceiling.
**Gate fires — verified, not assumed:** removing the disqualifiers →
`pnpm test:gate` exits 1 (11 failed / 471 passed); restored → exits 0.
## What is deliberately NOT admitted, and why
`reviewer.test.ts` holds the sibling family — *"a provider outage is not
a review verdict"*. I verified by mutation that it genuinely guards
this: removing the escalation branch fails **5** tests covering
"escalates a rate limit as `ReviewerProviderError` instead of an
`UNAVAILABLE` verdict", "does not burn the reviewer fallback retry
budget on a provider outage", and "escalates as transient once the
network retry budget is exhausted". That budget exists to bound *bad
reviews*; spending it on an outage fails tasks that have nothing wrong
with them.
It is green in `engine-default` but **fails 72 cases under
`engine-core`**, because that project resolves `@fusion/core` through
the **reduced** `index.gate.ts` barrel/bundle and the suite reaches
exports it does not carry (`__vite_ssr_import_0__.has…` TypeError).
Admitting it would mean widening the gate barrel — which trades away the
bundle's entire reason for existing (FN-7669 measured the barrel import
phase as the gate's dominant wall-time cost).
**I tried it, measured the 72 failures, and backed it out** rather than
either shipping a red gate or — the tempting version — loosening the
test until it passed under the reduced barrel. The reason is recorded in
the config next to the allow-list so the next person doesn't rediscover
it. Widening the barrel for this suite is a real option, but it is a
gate-performance decision with its own measurement, not a side effect of
a test-coverage PR.
## Review-lane characterization status
By-name coverage search performed first in every case, per the lesson
from #2520:
| Invariant | Verdict |
|---|---|
| FN-8492 orphaned pending results rewritten, never deleted | covered
(NEW=2) |
| FN-7720 bypass writes `skipped` | covered (NEW=1) |
| FN-7720 bypass never fabricates a verdict | **was vacuous** — fixed in
#2541 |
| Provider outage escalates, never becomes a verdict | covered (NEW=5),
outside the gate — see above |
| Prose rejection never promoted to APPROVE | covered (NEW=11) — **now
gated** |
| testMode never issues real AI calls | **was permanently red** — fixed
in #2547 |
Still uncharacterized, stated rather than implied: branch-group member
integration and promotion sequencing (the FN-5819 scoped exception).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**U9, PR5.** Closes the gap #2520 measured. Tests + gate config only; no
production behavior change.
## The gap
#2520 found that safeguards **1 (user pause)** and **4 (capacity
single-flight)** had **zero test coverage**. Deleting either guard
produced no new failure anywhere in the merge, project-engine,
self-healing, or concurrency suites. Both guards work correctly today —
nothing would have noticed if they stopped. U9 moves merge behind graph
nodes, so this is exactly the state not to convert on top of.
## Two tests
- **`merge admission excludes a user-paused card`** — safeguard 1, the
pause invariant re-ratified in #2486. Without the `paused || userPaused`
filter, the admission provider offers a user-paused card to the merge
pump.
- **`drainMergeQueue is single-flight`** — safeguard 4. Asserted via
`reconcileStaleMergeActive`, the first statement *inside* the guard, so
the probe isolates the guard rather than dispatching a real merge.
(Driving a real drain crashed the vitest worker; probing the guard
directly is both safer and more precise.)
**Both are two-sided** — they assert the guard blocks *and* permits. A
one-sided test would still pass against a guard that rejects everything,
which is a real failure mode for a filter.
## Proven by mutation delta
Baseline fail-set vs mutated fail-set on the identical selection, NEW
failures only:
| Mutation | NEW failures |
|---|---|
| remove the pause filter | **1** — the pause test, and only it |
| remove the single-flight guard | **1** — the single-flight test, and
only it |
| filter rejects *everything* | **1** — proves not one-sided |
| drain *always* refuses | **1** — proves not one-sided |
## Gate admission
`project-engine.test.ts` joins the `engine-core` allow-list. **One file
proves five safeguards** — user pause, `autoMerge:false`, capacity
single-flight, the pre-enqueue merge-proof consult, and at-most-once
enqueue.
Before this, **none of the six safeguards was defended by blocking CI**.
A regression surfaced only in non-blocking full-suite, after the merge.
Measured, not assumed:
| | Files | Tests | Wall (3 runs) |
|---|---|---|---|
| before | 17 | 309 | 5.19 / 5.51 / 5.19s |
| after | 18 | 412 | 6.19 / 6.24 / 6.21s |
**+~1.0s against a ~60s ceiling.**
**Verified the gate fires**, rather than assuming the allow-list edit
took — the failure mode greptile caught in #2494:
- remove safeguard 1 → `pnpm test:gate` **exits 1** (1 failed / 411
passed)
- remove safeguard 4 → **exits 1** likewise
- restored → **exits 0**
Deterministic: store, runtime, merger and notifier all mocked; no real
git, no network, no real timers in these two cases.
## Reversible calls I made rather than asking
- **Added to `project-engine.test.ts` rather than a new file.** A
dedicated file would need ~200 lines of duplicated `vi.mock`
scaffolding; reusing the existing harness also means one gate admission
covers five safeguards instead of two.
- **Did not wait for U8.** These guard code that exists today and the
conversion needs them in place first.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**U9, PR1 of several.** Test-only, no production code touched. This is
the characterization baseline the plan's Execution note asks for before
the merge lane converts.
## The finding
`builtin-coding-workflow-ir.ts` declares merge-region policy that **no
engine code reads**:
| IR declaration | Consumed by |
|---|---|
| `merge-retry` → `{ policy: "merge", maxAttempts: 3 }` | nothing —
`retry-backoff` handler is `async () => ({ outcome: "success" })`
(`workflow-node-handlers.ts:728`) |
| `merge-manual-hold` → `{ release: "manual" }` | nothing — returns a
constant `manual-required` |
| `branch-group-*` → `{ maxReworkCycles: 3 }` | nothing — returns a
constant `success` |
Live merge policy authority is elsewhere, on two separate axes:
- **conflict** retries — `settings.maxAutoMergeRetries` (default 3),
already covered by `auto-merge-retry-cap-settings.test.ts`
- **transient** retries —
`ProjectEngine.MAX_AUTO_MERGE_TRANSIENT_RETRIES = 5`
(`project-engine.ts:545`)
So the IR is a **third, dead authority**. These are different axes, not
a same-axis contradiction — but a reader looking at the IR would
reasonably take the declared numbers as live, and nothing currently says
otherwise. U9's acceptance criterion is "merge policy changes via IR
config alone, with no code change"; that fails today and this pins why.
## Why characterization rather than a fix
Making these handlers config-driven is a **merge behavior change**, and
the `requestMerge` primitive it routes through lives at
`executor.ts:7383` — inside U8's blast radius. U9 is sequenced behind U8
precisely so the merge lane converts onto an executor that is already
substrate. Landing the behavior change now would change merge semantics
on an executor about to be reshaped. It lands inside U9 proper.
When U9 wires a node kind onto its IR config, the matching case here
goes **red** and the U9 commit must move that kind out of
`CONFIG_BLIND_MERGE_REGION_KINDS`. That is the ratchet working.
## Proof it fails when reverted
A test that passes with the change reverted is not a test. The "change"
here is the test itself, so the honest analogue is mutating the
characterized production behavior. Three independent mutations, each
reverted after measuring:
| Mutation | Result |
|---|---|
| `retry-backoff` honours `config.maxAttempts` (what U9 will do) | **2
failed** / 6 passed |
| `manual-merge-hold` honours `config.release === "external-event"` |
**2 failed** / 6 passed |
| IR declaration drift: `maxAttempts: 3` → `7` | **1 failed** / 7 passed
|
Measured: 8 tests, 4.16s. `pnpm lint` clean. Tree restored to clean
after each mutation.
The assertions are behavioral, not string matches: each handler is
invoked with two contradictory configs (opposite budgets, opposite
release modes, disjoint surfaces) and asserted to return deep-equal
results.
## Six safeguards
This PR changes no production behavior, so no safeguard is altered by
it. The full six-row table with test attribution is the required
artifact for the **conversion** PR, not this one. Baseline located so
far, to be completed and verified by mutation before any conversion
lands:
| # | Safeguard | Consulted at (today) | Test attribution |
|---|---|---|---|
| 1 | user pause | `project-engine.ts:645` (`task.paused \|\|
task.userPaused`) | not yet verified |
| 2 | `autoMerge:false` | `allowsAutoMergeProcessing` —
`project-engine.ts:2797`, `merger.ts:7178` | not yet verified |
| 3 | dependency gating | not yet located | not yet verified |
| 4 | capacity | not yet located |
`workflow-column-boundary-capacity.test.ts` (unverified) |
| 5 | merge-proof | `getTaskMergeBlocker` — `project-engine.ts:2609` |
`merger-file-scope-invariant.test.ts`,
`merger-diff-volume-gate.slow.test.ts` (unverified) |
| 6 | at-most-once merge | `activeMergeTaskId` single-flight —
`project-engine.ts:693`/`:2729` | not yet verified |
Rows 3, 4 and all attributions are honestly incomplete rather than
asserted — I will not present a table I have not earned.
## Also found, for the coordinator
- **Slice statuses are stale.** S02/S03/S04 in
`docs/plans/workflow-owned-merge-stack/` are all marked
`draft-stack-handoff` but S04 has **landed** (the merge-region IR nodes
above), S03's `claimDueWorkflowWorkItem` is implemented and wired via
`workflow-work-processor.ts`, and S02's
`projectMergeRequestToWorkflowWorkItem` is implemented with **zero
production callers**. S06/S07/S08 are genuinely not started. Doc
correction coming as its own small PR.
- **S1 prerequisite verified present, not assumed** — all four store
methods live in `store.ts`, migration `0031` in tree. No S1-completion
gap.
- **Second control plane into the merge lane:** `self-healing.ts:3198`
and `:7200` call `enqueueMerge` directly, bypassing the graph. That
needs to become a recovery-fact/wake (the stack's R6) during U9.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Added a new test suite to cover U9 merge-region behavior across
supported workflow node types.
* Verified merge-region results are consistent across built-in,
contradictory, and missing configuration inputs.
* Documented current behavior for retry backoff (always succeeds) and
manual merge hold (fails as manual-required).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Summary
Restores green merge-gate and package-default suites after repeated
`origin/main` merges brought workflow-graph ownership cutover drift into
CI.
- Align engine/dashboard/core tests with post-cutover contracts
(`moveTaskIf`/`deleteTaskIf`, graph handoff, worktree-pool reclaim via
`removeWorktree` + `RemovalReason`, multi-step RESUMING parse,
soft-pause merge requester, graph-terminal failure surfaces).
- Small product fixes needed for real regressions uncovered by the
suite: soft-delete refuse before graph routing, skip DUPLICATE
step-heading withhold when an explicit marker is present, PG schema
applier guards, and related bookkeeping (research promote tool inventory
/ migration seed, stop shell `psql` in PG admin DDL).
- Quarantine/ledger hygiene only where required by standing rules; no
timeout/worker appeasement.
## Verification
- `pnpm test:gate` ×2 green
- `@fusion/engine` full package suite green (~9083 tests)
- Targeted core/dashboard clusters green (schema applier, agent-runs UI,
settings descriptions, mobile close)
## Test plan
- [x] `pnpm test:gate` (twice)
- [x] `pnpm --filter @fusion/engine test`
- [ ] CI full suite / PR checks on this branch
- [ ] Confirm no unrelated product behavior changes beyond the listed
regression fixes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added support for `roadmap-item` native structure kinds, including
native structure embeds and metadata validation.
* Added Stable and Beta release channel options in General settings.
* Added per-action reporting target configuration with clearer “unset”
guidance.
* **Bug Fixes**
* Improved heartbeat/prompt behavior when patrol is disabled.
* Prevented deleted tasks from continuing through execution.
* Made recovery for explicit duplicate redirects more permissive.
* Hardened database migration and test database cleanup to reduce flaky
failures.
* **Documentation**
* Updated settings text for release channels, reporting targets, and
inheritance/unset behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- Default `createAgentTask` in dashboard `@fusion/engine` mock so
planning/subtask create routes return 201 (FN-8277).
- Mock `findRecentTasksBySourceParentTaskId` on github/planning route
stores.
- Quarantine `merge-reuse-task-worktree.slow.test.ts` (engine-slow load
flake, run 29663725381).
## Evidence
- Prior full green: Full Suite run **29663526777** on #2325.
- Tip red class: routes-github/planning 500 + engine-slow lease
residual.
## Test plan
- [x] routes subtask create-tasks / shared branch groups tests green
locally
- [ ] Full Suite all 4 shards + engine-slow green on main tip after
merge
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved task and subtask creation test coverage to correctly handle
parent-scoped duplicate checks.
* Updated test behavior to return reliable task creation results.
* **Tests**
* Quarantined a flaky integration test from the slow test suite to
improve test run reliability.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
After #2289, Full Suite shard 4 still failed on the **OMP** twin of the
Grok process-lifecycle stress test (`import("../index.js")` × 15 under
shard transform load → 5s timeout).
Apply the same fix class as grok-runtime:
- Symbol.for exit reaper on `process-manager`
- Stress test reimports that module
- 15s timeout for cold transform
## Test plan
- [x] Local OMP process-lifecycle green
- [ ] PR gate
- [ ] Post-merge Full Suite
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Improved cleanup of OMP ACP processes when the application exits.
- Prevented duplicate exit handlers and excess listener growth during
runtime reloads.
- Preserved reliable process lifecycle behavior under repeated module
loading.
- **Tests**
- Added lifecycle coverage for repeated process-manager reloads.
- Optimized the stress test to complete more efficiently while retaining
cleanup assertions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Fixes shard 4 full-suite failures: chat_sessions schema baseline gap +
two remaining PG auth bugs missed by PR #2086.
**Scope: shard 4 only.** Shards 1/2 (engine timeouts) and shard 3
(compound-engineering CI-only failure) are separate issues not addressed
here.
## Changes
### Schema baseline gap — `chat_sessions` missing columns (42703 error)
- **`0000_initial.sql`**: Added `validator_thinking_level` and
`planning_thinking_level` columns to `CREATE TABLE
project.chat_sessions`. These exist in the Drizzle schema
(`project.ts:1492-1493`) but were missing from the SQL baseline, causing
`column does not exist` on all chat_sessions inserts in fresh test
databases.
- **`postgres-health.ts`**: Added both columns to
`EXPECTED_PROJECT_COLUMNS` self-heal list so existing databases also get
them via ALTER TABLE.
**Fixes**: `chat-store-content-search-edit.pg.test.ts` (5 tests),
`satellite-db-injected-stores.test.ts` (2 tests)
### Remaining auth bugs (password auth failed for user "runner")
- **`allocator-cross-project.test.ts`**: Still had `process.env.USER` in
inline adminExec — missed by PR #2086's batch fix. Replaced with
`PG_TEST_URL_BASE` connection string.
- **`connection.test.ts`**: Used `FUSION_PG_TEST_URL` (not set on CI)
with a bare default URL lacking credentials. `postgres.js` fell back to
OS user `runner`. Changed to derive from `FUSION_PG_TEST_URL_BASE` which
includes credentials.
**Fixes**: `allocator-cross-project.test.ts` (2 tests),
`connection.test.ts` (3 tests)
## Verification
| Check | Result |
|---|---|
| Merge gate (`pnpm test:gate`) | ✅ 294 + 114 + 63 = 471 passed |
| chat-store-content-search-edit | ✅ 5 passed |
| satellite-db-injected-stores | ✅ 10 passed |
| allocator-cross-project | ✅ 2 passed |
| connection | ✅ 13 passed |
| Lint | ✅ exit 0 |
| Typecheck | ✅ clean |
## Not in scope
- **Shards 1/2**: Engine test suite timeouts with
`getAsyncLayer`/`updateSettings` mock warnings. Pre-existing.
- **Shard 3**: `compound-engineering stage-skill-loading.test.ts` — 14
tests fail on CI (`TypeError: Cannot read properties of undefined
(reading 'close')`), pass locally. Likely CI-specific teardown issue.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added separate `validator_thinking_level` and
`planning_thinking_level` fields to chat session data, including
database schema and health-check recognition.
* **Bug Fixes**
* Improved PostgreSQL test connectivity by using configured connection
URL settings instead of hardcoded local defaults.
* Made Postgres-related test teardown null-safe to avoid failures when
setup doesn’t complete.
* **Tests**
* Updated automated test quarantine/exclusions for known failing engine
and reliability-interaction cases.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
# Migrate storage from SQLite to PostgreSQL — full dashboard cutover
Migrates Fusion's storage layer to the embedded PostgreSQL
`AsyncDataLayer` (the default backend) and **completes the
satellite-store + feature cutover** so every dashboard and Command
Center surface works in PG mode.
## Status — every surface works in embedded-PG mode
Verified live against a running embedded-Postgres dashboard (all
**200**, zero 5xx) and gate-tested (**23 files / 99 tests** on embedded
PG, plus engine-core 294 and ci-shape 63 in the blocking merge gate;
core/engine/cli/dashboard typecheck clean).
| Area | Surfaces | State |
|---|---|---|
| Satellite stores | workflows, todos, insights, research, missions,
goals, mailbox | ✅ |
| Views | artifacts, documents, evals | ✅ |
| Command Center | activity, productivity, team, tokens, tools,
**workflows**, **github**, **signals**, **plugin-activations**, **live**
(all 10) | ✅ |
| Run execution | insight generation, research run execution | ✅
(store-path; AI step needs a provider) |
| Live updates | SSE push for mission/research/insight events | ✅ |
| Workflow editing | create / update / delete / select (+ id counter) |
✅ |
| Engine | mission autopilot, incident-signal ingestion, regression
storm-guard, agent wake-on-message | ✅ |
| Core | tasks, agents, secrets, automations, memory, chat, usage, PRs,
git | ✅ |
## Approach
Each satellite store gets an `Async<Store>` wrapper exposing the sync
store's method names over the existing `async-*-store.ts` helpers;
`get<Store>Store()` returns a `Sync | Async` union; consumers `await`
(harmless on sync), and engine/CLI paths that can't convert use
`instanceof Sync` graceful fallback. Analytics aggregators branch on
`"ping" in dbOrLayer` to run schema-qualified raw SQL over `project.*`
(snake_case) in PG. Executors/orchestrators/autopilot are
await-converted to drive the union store; the async store wrappers
extend `EventEmitter` so SSE live-push fires in both backends.
Not-yet-ported capabilities degrade gracefully (never 500) and are
individually called out in commits.
## Sync with main
The branch is kept continuously merged with `main` (currently through
FN-7845, 2026-07-12); the earlier "final rebase deferred" note no longer
applies. Use **Create a merge commit** (or squash) to land it — GitHub's
rebase-merge cannot replay a merge-maintained branch.
## Residual Review Findings
Multi-agent code review of the PostgreSQL satellite-store ports (U1–U5)
applied 3 safe fixes (see `fix(review): apply autofix feedback`). The
following are **real but gated** — recorded here as follow-up work
rather than auto-applied. All are SQLite→PostgreSQL
**concurrency/atomicity regressions**: the sync stores were immune only
by SQLite's single-writer, single-threaded-handler execution; the async
ports open multi-await read-modify-write windows. **Reachability is low
today** because the execution engines that generate concurrent same-run
mutations (insight run executor, research orchestrator/dispatcher) are
`instanceof`-gated to sync mode in PG. No process-crash class survived
(all engine fallbacks correctly guard the sync store).
- **[P1] Research `appendResearchEvent` dual-write is non-atomic**
(`packages/core/src/async-research-store.ts`, corroborated: adversarial
+ reliability). The `research_run_events` insert (own transaction) and
the `run.events` jsonb update are separate writes — a crash between
them, or two concurrent appends, splits the table count from the jsonb
array. **Fix:** perform the seq-insert and the jsonb update in one
`layer.transactionImmediate`.
- **[P1] Research run terminal-reversion via stale full-row persist**
(`async-research-store.ts` `persistResearchRun`/`updateResearchStatus`).
Concurrent `PATCH /runs/:id/status` + `POST /runs/:id/events` can revert
a terminal run to `running` by overwriting the whole row, bypassing the
transition guard. **Fix:** scoped column `UPDATE`s with a `WHERE status
…` guard, or optimistic version column.
- **[P2] `updateResearchRun`/`updateInsightRun` read-then-write TOCTOU**
— concurrent PATCHes last-writer-wins on the lifecycle merge. **Fix:**
`SELECT … FOR UPDATE` / enclosing transaction.
- **[P2] `upsertRun`/`createRunOrThrowConflict` check-then-create race**
(`async-insight-store.ts`) — two callers can each create an "active"
run. **Fix:** partial unique index on `(projectId, trigger) WHERE status
IN ('pending','running')`.
- **[P3] `createResearchRetryRun` return-value divergence** — sync
returns the pre-update `queued` snapshot; async returns the reloaded
`retry_waiting` run (persisted state is identical). Pick one side for
cross-backend parity.
- **[P2/perf] Mission `getMissionWithHierarchy`/`getMissionHealth` N+1
fan-out** — O(milestones×slices) sequential round-trips hold one pool
slot per request; can starve the pool for large hierarchies. **Fix:**
batched/joined reads.
- **Testing gaps:** no PG-mode concurrency tests (interleaved
status/event mutations), no sync↔async parity assertion for the
lifecycle-error codes, and no mission status/health rollup parity test
vs the sync `MissionStore`.
~~Out of scope (deferred): AI run *execution* (insight/research) +
mission autopilot + live SSE mission events remain sync-gated/degraded
in PG mode.~~ **Since ported** — insight/research run execution, mission
autopilot, and SSE live push all run on the async layer now, which also
makes the concurrency findings above genuinely reachable; they remain
open follow-ups.
---
## Update — 2026-07-12: production-readiness hardening & live acceptance
Everything below landed on this branch since the description above was
written:
**Production blockers from review — fixed**
- `recoverStaleTransitionPending` ported to the async layer (backend
moves write + clear the crash-safe marker; startup/maintenance sweeps no
longer throw).
- Lost-update class fixed: `atomicWriteTaskJson`/`WithAudit` write
changed columns only (full-row upserts silently resurrected stale fields
across concurrent store instances — the "task stuck unplanned forever"
bug).
- First-boot **auto-migration**: booting the PG backend over a project
with a legacy `fusion.db` migrates it automatically (loud failure,
SQLite kept as backup), and the dashboard shows a one-time **"your data
was migrated" banner** with the backup paths and a Need-help Discord
link.
- `pg_dump`/`pg_restore` discovered from common install locations for
embedded-mode backups.
- The PG suite is part of the blocking merge gate (`test:pg-gate`).
**Multi-project isolation (PR #2007, merged into this branch)**
- `project_id` partition key on tasks / archived tasks / config,
`taskProjectScope` threaded through every scan/claim/count, per-project
config rows, layer bound to the project at startup.
- Review P1 follow-up: the shared cold-storage `archive.archived_tasks`
table is also partitioned and all archived-board reads/counts/searches
are scoped.
- Schema drift self-heal generalized to schema-qualified columns so
existing databases upgrade in place.
**Other changes**
- Node settings sync **removed** in PG mode (409
`settings-sync-disabled-postgres`) — nodes share state by connecting to
the same database; auth sync kept (per-machine file).
- Perf (review findings): `listTasks` pushes column filter + ORDER BY +
LIMIT/OFFSET into SQL; `getConversation` capped to the most recent 200
messages.
- Fixed a false "operator action required" pause-abort log fired on
every successfully auto-merged task.
**Live acceptance — PASSED (2026-07-12)**
A sandboxed instance (isolated HOME, embedded PG, real Opus executor)
ran a task through the complete cycle: create → triage (AI spec) →
execute → in-review → AI squash-merge landed on the project's `main` →
done. A write+read sweep of every data surface (settings, comments,
documents, attachments + artifact bridge + artifact edit, chat with real
generation, goals, missions, agent mail, secrets, workflows, memory, CC
analytics) was green on embedded PG.
**Known remaining work**
- The per-project `config` PK re-key has no upgrade path for
pre-isolation embedded-PG databases (needs a real `DROP
CONSTRAINT`/re-key migration; fresh databases are fine).
- `pg_dump`/`pg_restore` binaries are not yet bundled in release
artifacts (PATH/common-location discovery only).
- The satellite-store concurrency findings listed above.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Phil Larson <hello@phillarson.xyz>
Co-authored-by: fusion-merge <fusion-merge@local>
Narrative: FN-7673 re-attempted the engine-core gate bundle lever with a single combined-entry engine-graph design (all 14 mock-safe roots redirected through a resolveId plugin to one synthetic packages/engine/.gate-bundle/engine.mjs) after FN-7670's 14-separate-root attempt was inconclusive. This update records the negative A/B result and closes the lever.
- Documented that the combined-entry design achieved its structural goal (149 first-party inputs -> 1 output file) and full 335/335 coverage parity
- Recorded a true interleaved A/B (5 warm + 1 cold pair) showing the combined-entry bundle is consistently slower than the @fusion/core-only baseline (warm median +29.1%, import-phase aggregate +74.0%)
- Captured the working theory: funnelling 14 relative-import sites through a resolveId-plugin redirect to one large synthetic export-* file adds more transform/resolution overhead than it saves, unlike @fusion/core's plain resolve.alias
- Noted the experiment was NOT landed; wiring (engine-graph scans, combined-entry builder, resolveId plugin) was fully reverted
- Marked this lever (bundling the @fusion/engine relative-import graph for the engine-core gate, in either 14-file or single-combined-entry shape) as CLOSED absent new evidence
Files changed:
packages/engine/vitest.config.ts | 32 +++++++++++++++++++++++++++++---
1 file changed, 29 insertions(+), 3 deletions(-)
Fusion-Task-Id: FN-7673
Fusion-Task-Lineage: 46951e5f-e7dc-4f7c-9601-0cfa0b082d70
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Removes a dead test-file reference from the engine-core vitest gate include list, with a code comment documenting why.
- Remove the nonexistent `src/__tests__/merger-post-merge.test.ts` entry from packages/engine/vitest.config.ts's engine-core include list (retired by FN-7039; graph is now sole post-merge owner)
- Add FNXC comment noting the entry matched zero files and that graph post-merge coverage lives in workflow-graph-post-merge.test.ts (engine-default)
Files changed:
packages/engine/vitest.config.ts | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-7671
Fusion-Task-Lineage: 73447412-7b8a-4578-a2b8-07f83e381548
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Prototyped extending the @fusion/core pre-bundle alias lever to @fusion/engine's relative-import production graph reached by the 18 gate files, but an A/B showed no clear win over the @fusion/core-only bundle, so the change was not landed and only the rationale is recorded.
- Added an FNXC:EngineTests comment block in packages/engine/vitest.config.ts documenting the FN-7670 prototype (171 first-party files → 35 output files via esbuild multi-entry splitting)
- Recorded the negative A/B result: byte-size growth of 14 separate large root bundles offset per-file-dispatch savings, with no clear win beyond host run-to-run noise
- Left the vitest alias wiring unchanged at the @fusion/core-only bundle state, pointing future attempts to FN-7670's task docs for full analysis and to consider a single combined engine-graph entry instead of 14 separate root entries
Files changed:
packages/engine/vitest.config.ts | 19 +++++++++++++++++++
1 file changed, 19 insertions(+)
Fusion-Task-Id: FN-7670
Fusion-Task-Lineage: efd27f94-a6c4-49c7-a78e-50213fd42a24
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Prototype and land a rebuilt-every-run esbuild bundle of the @fusion/core gate-safe barrel closure, collapsing the engine-core gate's per-fork Vite SSR import-phase cost (18 forks x ~430-file closure re-resolved from scratch) into a single file load per fork.
- Add scripts/build-engine-core-gate-bundle.mjs: esbuild-bundles packages/core/src/index.gate.ts (220 first-party files, packages:"external" so third-party/node: imports stay external, treeShaking:false to preserve side effects) into packages/core/.gate-bundle/core.mjs + core.meta.json
- Wire the builder into packages/engine/vitest.config.ts's engine-core project globalSetup (alongside the existing vitest-teardown hook) so the bundle is rebuilt fresh before every gate invocation, and repoint the @fusion/core resolve.alias at the bundled output instead of index.gate.ts source
- Place the bundle output at packages/core/.gate-bundle/ as a sibling of packages/core/node_modules/ (not nested inside it) to avoid Vite SSR's external-dep heuristic, which would otherwise silently defeat vi.mock interception for imports nested in the bundle
- Gitignore packages/core/.gate-bundle/ and add a matching ESLint ignore entry so the generated bundle text is never linted or committed
- Add esbuild ^0.25.12 as a root devDependency (pnpm-lock.yaml updated accordingly)
- Document the pre-bundling rationale, placement constraints, and measured A/B wall-time results in docs/testing.md
Verified: pnpm test:gate passes (335/335 engine-core tests, 63/63 CLI ci-shape tests), engine package typecheck clean, eslint clean on touched files.
Files changed:
.gitignore | 11 ++
docs/testing.md | 3 +
eslint.config.mjs | 10 ++
package.json | 1 +
packages/engine/vitest.config.ts | 50 ++++++++-
pnpm-lock.yaml | 3 +
scripts/build-engine-core-gate-bundle.mjs | 174 ++++++++++++++++++++++++++++++
7 files changed, 247 insertions(+), 5 deletions(-)
Fusion-Task-Id: FN-7669
Fusion-Task-Lineage: 62b06b2a-4ac6-45ae-ac79-9771132bc303
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Introduces a project-scoped @fusion/core barrel used only by the engine-core
gate project, so new feature modules added to the full barrel don't silently
inflate the gate's transform/import cost.
- Add packages/core/src/index.gate.ts, a copy of the full @fusion/core barrel
minus export statements for modules added since the last re-audit baseline
(i.e. it still re-exports everything the full barrel does except newly
added, gate-irrelevant feature modules).
- Update packages/engine/vitest.config.ts to add a project-scoped
resolve.alias mapping @fusion/core -> packages/core/src/index.gate.ts for
the engine-core project only; engine-default/engine-reliability/engine-slow
and @fusion/engine continue to resolve the full barrel.
- Document the gate-safe barrel and its audit procedure in docs/testing.md.
Files changed:
docs/testing.md | 3 +
packages/core/src/index.gate.ts | 2102 ++++++++++++++++++++++++++++++++++++++
packages/engine/vitest.config.ts | 17 +
3 files changed, 2122 insertions(+)
Fusion-Task-Id: FN-7667
Fusion-Task-Lineage: 054ec89a-d973-44dd-b9ac-ad266f553f01
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Quarantine test files consistently failing on the non-blocking full-suite
CI on main, per the AGENTS.md deletion-ratchet policy:
Engine-default (shard 1-2): ce-workflow-step-conventions,
executor-column-agent-principal, restart.integration,
scheduler-node-unreachable-audit, scheduler-overlap-starvation,
scheduler-ephemeral-toggle, user-configured-command-no-execsync
Engine-reliability (shard 1): lease-recovery-central-claim,
owning-node-unavailable-interactions, todo-inprogress-flapping
CLI (shard 3): extension.test.ts
Each has a matching entry in scripts/lib/test-quarantine.json with the
failing CI run link and quarantinedAt date. Tests will be deleted
after 14 days unless rescued with a root-cause fix.
Quarantine three test files consistently failing on the non-blocking
full-suite CI on main, per the AGENTS.md deletion-ratchet policy:
- engine self-healing-fn-5488-fast-path-regressions.test.ts (shard 1)
- engine in-review-merge-stall-deadlock-recovery.test.ts (shard 2)
- dashboard DevServerView.mobile.test.tsx (shard 4)
Each has a matching entry in scripts/lib/test-quarantine.json with
the failing CI run link and quarantinedAt date. Tests will be deleted
after 14 days unless rescued with a root-cause fix.
Quarantine unrelated engine test flakes so the gate avoids known unstable fixtures.
- Add the soft-delete blocker residue reliability test to the quarantine ledger and reliability interaction exclusions.
- Add the bubblewrap sandbox backend test to the quarantine ledger and engine-core exclusions.
Files changed:
packages/engine/vitest.config.ts | 6 +++++-
scripts/lib/test-quarantine.json | 10 ++++++++++
2 files changed, 15 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-6319
Fusion-Task-Lineage: 9a47ae20-76b8-47d8-8858-f4fe51ea807e
Keep task detail PR and review affordances synchronized with the current project auto-merge setting.
- Thread the live auto-merge value into task detail modals and list split-pane detail views.
- Prefer the live auto-merge setting over the stale fetched modal snapshot while preserving task-level overrides.
- Cover PR and review tab behavior for auto-merge on/off and document the dashboard behavior.
- Evict flaky engine gate entries and add a patch changeset for the published CLI package.
Files changed:
.changeset/fn-6247-automerge-off-modal-stale.md | 5 +
docs/dashboard-guide.md | 1 +
packages/dashboard/app/App.tsx | 3 +-
packages/dashboard/app/components/AppModals.tsx | 2 +
packages/dashboard/app/components/ListView.tsx | 3 +
.../dashboard/app/components/TaskDetailModal.tsx | 4 +-
.../__tests__/TaskDetailModal.create-pr.test.tsx | 171 ++++++++++++++++++++-
packages/engine/vitest.config.ts | 2 -
8 files changed, 182 insertions(+), 9 deletions(-)
Fusion-Task-Id: FN-6247
Fusion-Task-Lineage: 1321c03a-216d-4b15-bf4f-95621d68c9ae