- workflow-graph-executor: wrap each post-merge walk() in try/catch so a malformed
post-merge IR / traversal error is logged and skipped, never flipping an already-
merged task to failed (non-blocking post-merge contract) [T9, real bug].
- Refresh stale FNXC comments now that graphNativePostMerge is default-ON and the
legacy merger post-merge path was removed (experimental-features, workflow-graph-
executor, workflow-graph-post-merge.test) [T6/T7/T8].
- Normalize FNXC timestamps to yyyy-MM-dd-hh:mm (TaskCard.test, taskProgress.test) [T2/T3].
- Changeset: category fix → feature to match the minor bump [T0].
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Final cutover — nothing reads workflow_steps at runtime, so migration 131 drops it.
- Removed the merger post-merge execution path entirely (runPostMergeWorkflowSteps,
hasEnabledPostMergeWorkflowSteps, executePostMerge{Prompt,Script}Step, post-merge
worktree helpers + call site). Graph owns post-merge.
- Executor recovery no longer reads getWorkflowStep().gateMode; gate-ness comes from the
recorded WorkflowStepResult.status.
- Removed store CRUD (create/update/deleteWorkflowStep), materializeWorkflowSteps, and
migrateLegacyWorkflowSteps; selectTaskWorkflow now seeds default-on optional-group node
ids (consistent with create-time). KEPT the plugin step-template palette (getWorkflowStep
plugin-only resolver / listWorkflowSteps plugin-only) — never touches the table.
Removed the dashboard migrate-legacy-steps route + editor migration UI.
- SCHEMA_VERSION 130→131; migration 131 DROP TABLE IF EXISTS workflow_steps; SCHEMA_SQL
table def removed; historical migrations 77/105/109/130 guarded with tableExists().
Proof nothing stranded: the graph executes IR nodes resolved from workflowId (never
stepIds/compiled rows) — materializeWorkflowSteps writes were vestigial. Full @fusion/core
suite (6290), reliability backstop (154), boot smoke, and a seed-at-130 drop test all pass
with the table gone.
Plan U7c.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Makes graph-native post-merge the default and cuts the merger's legacy post-merge
path over so post-merge runs exactly once (via the graph), with migration data prep.
- experimentalFeatures.graphNativePostMerge → default ON; merger
hasEnabledPostMergeWorkflowSteps/runPostMergeWorkflowSteps become inert when on
(no double-run; proven by a new no-double-run test). Merger code kept (U7c removes it).
- Migration 130 (SCHEMA_VERSION 129→130) rewrites each task's enabledWorkflowSteps
entries that are legacy built-in pre-merge workflow_steps row ids → the optional-group
node id (browser-verification/code-review); dedupes; idempotent; identity-stable;
leaves node-ids/compiled/custom entries untouched. Table KEPT (U7c drops it).
- Investigation (real DBs): NO custom/plugin post-merge steps exist; the only post-merge
step is compound-engineering's 'document' graph node — so the merger no-op strands
nothing. Custom-step re-pointing was verified unnecessary and skipped.
Safe-to-drop in U7c still blocked by live readers: merger post-merge fns, store CRUD,
migrateLegacyWorkflowSteps/readConfig materialization (executor recovery reader is
already null-safe→advisory).
Plan U7b.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds graph-native post-merge step execution behind experimentalFeatures.graphNativePostMerge
(default OFF — byte-identical behavior until enabled). After a successful merge-attempt
(the merge seam awaits the merge Promise), the graph runs post-merge optional-group nodes
and records phase:"post-merge" results, non-blocking. The merge-region traversal hop is
inert when the flag is off, and empty for builtin:coding even flag-on (its merge exits only
reach merge-region nodes or end), so the parity oracle holds. Optional-group recording now
derives phase + log prefix from config.phase (defaults pre-merge). Adds postMergeOptionalGroupNode
factory for migration/custom workflows. Legacy merger post-merge path untouched (U7b cutover).
Plan U7a.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
From the U1-U6 code review (correctness/adversarial/reliability/maintainability):
- Delete dead code left by the runWorkflowSteps removal: parkTaskAfterWorkflowStepPause,
handleWorkflowRevisionRequest (+ createWorkflowRevisionFollowUpTask,
injectWorkflowRevisionInstructions), handleWorkflowStepFailure, the dead
partitionWorkflowRevisionFeedback export + its test, and 2 orphaned jsdocs;
reword 2 stale comments (executor.ts FN-6722, self-healing.ts jsdoc).
- Record step output/notes on graph workflow-step results: runGraphCustomNode now
emits contextPatch:{output,notes} so the Workflow tab shows real review feedback
and [pre-merge] revision logs carry detail (was always the fallback before).
- Sync the workflow-step seam contract: core (workflow-compiler SEAM_NAMES/order,
workflow-ir column map, builtin-workflow-prompts) now rejects the workflow-step
seam to match the engine's resolveSeamName, preventing a latent run-time crash on
a persisted/cloned def that core would otherwise parse.
Residual (tracked in the PR): re-introduce the FN-4343 per-step scope gate on the
graph path; the parked-failed recovery log wording; malformed-advisory->passed edge
case; recording for non-optional-group/split-branch step realizations.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
U6: delete the built-in WORKFLOW_STEP_TEMPLATES catalog + its materializer
(getBuiltInWorkflowTemplate/ensureWorkflowStepForTemplate/toBuiltInWorkflowStep);
inline the browser-verification + code-review name/prompt/toolMode/gateMode into
their optional-group IR builders (node bytes unchanged); simplify
resolveEnabledWorkflowSteps to an identity-stable pass-through (no materialization,
so the optionalGroupIdSet collision guard is no longer needed). Plugin-contributed
step templates are kept as the editor palette.
U5: remove the legacy /api/workflow-steps REST surface (GET/POST/PATCH/DELETE +
/refine + /workflow-step-templates/:id/create), the dead client fns, and the
Settings management UI; GET /api/workflow-step-templates now serves plugin
templates only. The create-time optional-step toggles remain.
Scope: the workflow_steps store CRUD + table are intentionally KEPT — still consumed
by the engine (merger/recovery) and needed by U7's migration; their removal + the
table drop land in U7.
Plan U5 + U6.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Problem
Operators intermittently saw **"all my global settings reset"** —
including the global concurrency cap surfacing as `gate=semaphore` with
a low value in scheduler queue logs.
## Root cause
`CentralCore` is supposed to live at `~/.fusion/fusion-central.db`, but
`resolveGlobalDir(dir)` returns an explicit dir **verbatim**, and
several production call sites passed the **project** `.fusion/` dir:
- `store.getSecretsStore()` → `new CentralCore(store.getFusionDir())`
- dashboard secrets/proxy/node/secrets-sync/settings-sync routes → `new
CentralCore(store.getFusionDir())`
Each spawned a **stray per-project central DB**
(`<project>/.fusion/fusion-central.db`) seeded with **default** global
state (`globalMaxConcurrent=4`, empty secrets, default
`centralSettings`) that shadowed the real global DB whenever a
read/write hit one of those paths. Confirmed on disk: 14+ stray DBs at
default `4` vs the real `~/.fusion` at the operator's actual value.
## Fix
- Add `TaskStore.getGlobalSettingsDir()` returning the **resolved global
dir** (`string`); route the secrets store + all the affected dashboard
routes through it instead of `getFusionDir()`.
- `getSecretsStore()` also passes the global dir to `MasterKeyManager`
(co-locates the master key; makes the path test-exercisable).
- Add a `resolveGlobalDir()` **guard** that throws on a project-local
`.fusion/` dir (basename `.fusion` with a `.git` parent), with an
explicit `FUSION_ALLOW_PROJECT_LOCAL_GLOBAL_DIR` opt-out for
legitimately version-controlled custom global dirs. Inert under VITEST.
## Tests
- `global-settings-guard.test.ts` — guard rejects project/worktree
`.fusion` dirs, allows home + custom non-repo dirs.
- `store-secrets-store-global-dir.test.ts` — **symptom-based**:
`getSecretsStore()` creates the central DB in the global dir and
**never** spawns a stray project-local `fusion-central.db`.
- Added `getGlobalSettingsDir()` to route-test mock stores (the new
getter is now called by the routes).
## Operator note
Existing stray project-local `fusion-central.db` files (all
default/empty in practice) should be removed; on the affected machine
they were quarantined to `.legacy-central-db-backup-*` folders.
## Verification
- `pnpm --filter @fusion/core run typecheck` / `@fusion/dashboard`
typecheck — clean
- core guard + regression tests pass; previously-affected route suites
(proxy, nodes-sync, browse-directory, secrets-sync) pass (313/313)
- `pnpm check:changesets` — clean; eslint on changed files — clean
Companion to #1786 (global concurrency slider UI), which is independent.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- stage-review-badge-begin -->
---
<a href="https://stagereview.app/Runfusion/Fusion/pull/1787">
<picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://stagereview.app/assets/gh-open-in-stage-dark.svg">
<img src="https://stagereview.app/assets/gh-open-in-stage-light.svg"
alt="Open in Stage">
</picture>
</a>
<!-- stage-review-badge-end -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed intermittent resets of global settings (including the global
concurrency cap).
* Ensured global settings/secrets are read from and written to the
correct global storage location rather than a project-local one.
* Added a safeguard to prevent using a project-local settings directory
when running inside a repository.
* **Tests**
* Added regression coverage for global directory resolution/guard
behavior.
* Updated route and secrets sync tests to validate the global directory
selection used for proxying and sync.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
- Avoid the legacy-migration side effect on the hot path: run the cheap,
read-only checks (basename === ".fusion" && parent has .git) FIRST, and
only call resolveGlobalDirForHome() (which can perform a one-time rename)
when a dir actually looks project-local.
- Normalize paths before the home-dir comparison (realpathSync when present,
else resolve) so a trailing slash, doubled separator, or symlinked home
doesn't make the legitimate home global dir trip the guard.
- Harden the browse-directory regression test: the mock returns a global dir
DISTINCT from getFusionDir() and asserts CentralCore is constructed with the
global dir, so a revert to getFusionDir() now fails the test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Removes real wall-clock waits and per-test rebuilds from the slowest
test files, replacing them with deterministic seams. **No assertions
weakened, no timeouts widened, no retries added** — this is anti-pattern
removal per FN-5048, verified by re-running each file.
## Changes
| File | What | Result |
|---|---|---|
| `dashboard/.../insights-routes.test.ts` | Boot server+store **once**
in `beforeAll` (was `createServer` + `TaskStore.init` per test ×24);
reset insight tables per test for isolation; drive sweeper via fake
timers | test-exec **~3.7s → ~0.8s** |
| `core/.../db.test.ts` | Fixed 150ms write-lock hold → manual stdin
signal-release (keeps the real OS-lock contention under test); fixed a
real EPIPE on redundant release | 152 pass, non-flaky / 8 runs; −300ms
dead wait |
| `core/.../mission-store.test.ts` | 4 real `setTimeout` sleeps (only
there to force distinct timestamps) → `vi.setSystemTime` controlled
clock | anti-pattern removed |
| `core/.../agent-store.test.ts` | 1 real ordering-sleep → injected
`renewedAt` clock; **assertions strengthened** to pin exact timestamp
values | anti-pattern removed |
| `engine/.../in-process-runtime.test.ts` | Fake the one real 25ms
sleep; drop its inflated 30s per-test timeout | anti-pattern removed |
## Honest accounting
- The **real wins** are `insights-routes` (per-test server boot
eliminated, ~75% execution-time cut) and `db` (dead lock-hold removed).
- The **timestamp-sleep removals** (mission-store, agent-store,
in-process-runtime) are small absolute wins — the headline per-file
durations (16–25s) were **full-suite shard contention, not in-file dead
time** (each runs in 3–10s isolated). But they eliminate the FN-5048
real-wait anti-pattern, so a hub edit no longer drags real sleeps into
every `--changed` selection.
- **`workflow-routes.test.ts` was evaluated for splitting and
deliberately NOT split.** A measured A/B showed the 4-way split
*regressed* wall-clock (6s → 11s): the file is import/transform-bound
(per-file esbuild + `@fusion/core`/express import ≈ 5s > the ~4.3s test
runtime), and per-test store migration was already amortized by
`installInMemoryDbSnapshot`. Splitting only multiplies the dominant
fixed cost. Left intact.
## Verification
- `core` 612/612, `dashboard` 24/24, `engine` 78/78 (file-scoped).
- `tsc --noEmit` clean on all 3 packages; eslint clean.
Follow-up (not in this PR): `scripts/test-timings.json` is stale (its
former #1 file no longer exists) — refresh via `pnpm test:velocity --
--measure --write-report` so the watchdog budgets and velocity report
reflect reality.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- stage-review-badge-begin -->
---
<a href="https://stagereview.app/Runfusion/Fusion/pull/1784">
<picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://stagereview.app/assets/gh-open-in-stage-dark.svg">
<img src="https://stagereview.app/assets/gh-open-in-stage-light.svg"
alt="Open in Stage">
</picture>
</a>
<!-- stage-review-badge-end -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Made several test suites more deterministic by replacing real-time
delays with controlled timers and fixed timestamps.
* Improved lock and task checkout tests to use manual release signals,
reducing timing-related flakiness.
* Streamlined route test setup/teardown for faster, more reliable runs.
* Added safer cleanup around timer-based tests to avoid intermittent
failures.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Why
Diagnosing "tasks take too long" showed the dominant end-to-end
wall-clock is **waiting**, not work: tasks sit in `todo` (queue) and
`in-review` (merge wait) far longer than the agent actually runs. Today
only `cumulativeActiveMs` exists, which measures **in-progress time
only** — every other stage's dwell had to be reconstructed by hand from
agent logs.
## What
Adds `columnDwellMs?: Record<string, number>` to `Task` — a per-column
accumulator (column name → cumulative ms), recorded at the **same store
column-transition seam** as `cumulativeActiveMs` (`moveTaskInternal`).
On each move it adds `columnMovedAt(new) − columnMovedAt(prev)` to the
bucket for the column being left:
- clamped `>= 0` (clock skew safe);
- unparseable/missing prior timestamp and zero-dwell moves are skipped
(no spurious buckets);
- second visits **add** to the existing bucket (multi-visit churn is
captured);
- flag-independent — keys off the generic `columnMovedAt` delta, so it
runs for both the workflow-hook and legacy-inline move paths.
This makes per-stage dwell directly queryable, the same way
`productivity-analytics.ts` already consumes `cumulativeActiveMs`.
## Persistence
JSON-text task column, following the v129 `workspaceWorktrees` precedent
exactly: `SCHEMA_SQL` column + `SCHEMA_VERSION` 129→130 + a versioned
`addColumnIfMissing` migration. Additive and behavior-preserving —
pre-existing rows start NULL and accumulate from their next transition.
Survives archive/restore (added to the archive-entry mapping).
## Tests
`src/__tests__/store-execution-timing.test.ts` — new regression asserts
dwell across
`todo→in-progress→in-review→done→todo→in-progress→in-review` accumulates
the right per-column ms (second visits add) and survives a `getTask` DB
round-trip.
```
pnpm --filter @fusion/core exec vitest run src/__tests__/store-execution-timing.test.ts ... --reporter=dot
→ store-execution-timing 5/5, schema suites (goals/secrets) green, 17/17 total
```
Migration chain verified end-to-end (secrets-schema test climbs v11/v82
→ v130). No hardcoded literal version assertions in the suite; schema
tests assert against the `SCHEMA_VERSION` constant.
`@fusion/core` is private — no changeset.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- stage-review-badge-begin -->
---
<a href="https://stagereview.app/Runfusion/Fusion/pull/1781">
<picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://stagereview.app/assets/gh-open-in-stage-dark.svg">
<img src="https://stagereview.app/assets/gh-open-in-stage-light.svg"
alt="Open in Stage">
</picture>
</a>
<!-- stage-review-badge-end -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added per-column dwell timing to tasks, showing how long work spent in
each column across multiple visits.
* Preserved this timing data when tasks are archived and restored.
* **Bug Fixes**
* Task timing now updates correctly during column moves, including
repeated returns to the same column.
* Existing data can be upgraded to the new timing format without
breaking stored tasks.
* **Tests**
* Added coverage for multi-step task movement and data reloading to
verify timing totals stay accurate.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Addresses code-review findings on the global-settings reset fix:
- getGlobalSettingsDir() now returns a resolved `string` (was `string |
undefined`), so the project-local `.fusion` guard fires at the getter
call site instead of leaking CentralCore's undefined-default semantics.
- getSecretsStore() passes the resolved global dir to MasterKeyManager so
the master key co-locates with the global central DB and the path is
exercisable under tests (a bare new MasterKeyManager() throws in VITEST).
- resolveGlobalDir() guard gains an explicit FUSION_ALLOW_PROJECT_LOCAL_GLOBAL_DIR
opt-out so a legitimately version-controlled custom global dir (dotfiles
repo with a .git parent) is not hard-rejected.
- Add a symptom-based regression test (store-secrets-store-global-dir) proving
the secrets central DB lands in the global dir and never spawns a stray
project-local fusion-central.db.
- Add getGlobalSettingsDir() to route-test mock stores (CentralCore is mocked,
so it mirrors getFusionDir()) and FNXC comments to the remaining route sites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace real wall-clock waits and per-test rebuilds in the slowest test files
with deterministic seams. No assertions weakened, no timeouts widened, no
retries added — anti-pattern removal only.
- insights-routes.test.ts: boot the server + store ONCE in beforeAll (was a full
createServer + TaskStore.init per test x24), reset insight tables per test for
isolation, drive the sweeper via fake timers. Test-execution time ~3.7s -> ~0.8s.
- db.test.ts: convert the fixed 150ms write-lock hold to manual stdin signal-
release; keeps the real OS-lock contention under test, removes 2x150ms dead
wait. Fixed a real EPIPE on redundant release. 152 pass, non-flaky over 8 runs.
- mission-store.test.ts / agent-store.test.ts: replace real setTimeout sleeps
used only to force distinct timestamps with a controlled clock (vi.setSystemTime
/ injected renewedAt). agent-store assertions strengthened to pin exact values.
- in-process-runtime.test.ts: fake the one real 25ms sleep, drop its inflated
30s per-test timeout.
Honest note: the timestamp-sleep removals are small absolute wins (the headline
per-file durations were full-suite shard contention, not in-file dead time) but
eliminate the FN-5048 real-wait anti-pattern. workflow-routes.test.ts was
evaluated for splitting and deliberately NOT split — measured A/B showed the
split regressed wall-clock (the file is import/transform-bound, already amortized
by installInMemoryDbSnapshot), so splitting only multiplies fixed import cost.
Verified: core 612/612, dashboard 24/24, engine 78/78; typecheck + eslint clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Production code constructed `new CentralCore(store.getFusionDir())`, pointing
the central/global DB at the project's `.fusion/` instead of `~/.fusion/`.
`resolveGlobalDir()` returns an explicit dir verbatim, so this spawned stray
per-project `fusion-central.db` files seeded with default global settings
(globalMaxConcurrent=4, empty secrets) that shadowed the real global DB
whenever a read/write hit one of those paths — surfacing as intermittent
"all my global settings reset".
- Add TaskStore.getGlobalSettingsDir() (resolved global dir; undefined→~/.fusion)
- Route the secrets store + secrets/proxy/node/secrets-sync/settings-sync
dashboard routes through it instead of getFusionDir()
- Add a resolveGlobalDir() guard that throws on a project-local `.fusion/`
dir (basename `.fusion` with a `.git` parent); inert under VITEST
- Regression tests in global-settings-guard.test.ts
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
optionalGroupIdSet falls back to builtin:coding (mirroring the executor's
unselected-task resolution) so a toggled built-in group id like
browser-verification is no longer downgraded to a legacy WS-xxx step row the
graph executor never matches. Create-time optional-step controls resolve
builtin:coding when no project default workflow is set so the toggles render.
First unit of the graph-native workflow-step refactor (see
docs/plans/2026-06-25-001-refactor-workflow-steps-graph-native-plan.md).
Fusion-Task-Id: FN-7039
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds `columnDwellMs?: Record<string, number>` to Task — a per-column
accumulator (column name -> cumulative ms) recorded at the same store
column-transition seam as `cumulativeActiveMs`. On every move it adds
`columnMovedAt(new) - columnMovedAt(prev)` to the bucket for the column
being left, clamped >= 0; unparseable/missing prior timestamps and 0-dwell
moves are skipped, and second visits add to the existing bucket.
Motivation: `cumulativeActiveMs` only measures in-progress time. Diagnosis
of slow tasks showed the dominant wall-clock is *waiting* (queue time in
todo, review wait in in-review), which previously had to be reconstructed
from agent logs. This makes per-stage dwell directly queryable, like
productivity-analytics already consumes cumulativeActiveMs.
Persisted as a JSON-text task column following the v129 workspaceWorktrees
precedent: SCHEMA_SQL column + SCHEMA_VERSION 129->130 + versioned
addColumnIfMissing migration. Additive and behavior-preserving; pre-existing
rows start NULL and accumulate from their next transition. Survives
archive/restore. @fusion/core is private — no changeset.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Refinement: Code Review is now a DEFAULT-ON but toggleable `optional-group` in the
built-in coding and stepwise coding workflows (defaultOn:true), not a standard
always-on node. It is part of the existing pre-merge flow (execute →
[browser-verification optional] → code-review → review) and runs for every coding
task by default, yet an operator can toggle it off per task by removing `code-review`
from enabledWorkflowSteps; disabled → byte-inert pass-through. Advisory gateMode keeps
it non-blocking (operators can promote to a gate); toolMode readonly.
- Restore the optional-group builder (builtin-code-review-node.ts → -group.ts) with
config.defaultOn:true; stable group id `code-review`, inner id `code-review-step`.
- Wire the default-on optional-group into both built-in coding IRs.
- Fix store default-workflow seeding: interpreter-deferred built-ins (which carry
optional-group nodes) previously bailed to `undefined` in
materializeDefaultWorkflowSteps, dropping default-on group seeding under a
project-default workflow. Now they seed resolveDefaultOnOptionalGroupIds, mirroring
the explicit-workflow path, so defaultOn:true actually takes effect (the executor
enables a group strictly via enabledWorkflowSteps.includes(node.id)).
- Update tests + changeset; full @fusion/core suite green (356 files).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design correction: Code Review is now a STANDARD, default-ON step in the
built-in coding and stepwise coding workflows — not a default-off optional-group
toggle. It is a regular advisory `prompt` node on the pre-merge success path
(execute → [browser-verification optional] → code-review → review), so it runs
for every coding task with no enabledWorkflowSteps gating. Advisory gateMode means
it does not change merge outcomes; operators can promote it to a blocking gate.
- Replace the optional-group module with a standard prompt-node builder
(builtin-code-review-group.ts → builtin-code-review-node.ts).
- Keep the `code-review` WORKFLOW_STEP_TEMPLATE in the catalog (editor palette).
- Edges unchanged: code-review → review on success, code-review → end on failure
(mirrors the existing review node, no dead-end).
- Update tests + changeset for the standard always-on (no-toggle) semantics.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add a configurable "Code Review" diff-review step to the built-in coding and
stepwise coding workflows as a default-OFF optional-group prompt gate. It reuses
the existing workflow-step machinery and the shared trailing-verdict convention
(REVISE blocks, APPROVE/APPROVE_WITH_NOTES pass) — no engine verification code.
- New `code-review` WORKFLOW_STEP_TEMPLATE (toolMode readonly, gateMode advisory,
phase pre-merge) focused on the correctness value tests miss: logic bugs, edge
cases, intent-vs-implementation drift, regressions, error handling, contracts.
- New builtin-code-review-group.ts mirroring builtin-browser-verification-group.ts
(stable group id `code-review`, distinct inner node id `code-review-step`).
- Wired into builtin-coding-workflow-ir.ts and builtin-stepwise-coding-workflow-ir.ts
on the pre-merge path next to browser-verification, default OFF / opt-in.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
db.init() replays SCHEMA_SQL + ~129 migrations on every fresh in-memory
DB (~40ms each), which is minutes of pure setup across thousands of
DB-backed tests. Add a test-only migrated-schema snapshot: migrate ONE
in-memory DB per test file, serialize it, and deserialize a fresh copy
per test instead of re-migrating. Each test still gets a brand-new,
fully-isolated in-memory DB; only the migration cost is amortized.
- sqlite-adapter: expose serialize()/deserialize() (node:sqlite + bun)
- db.ts: setInMemoryTemplateSnapshot() hook (test-only, null in prod) +
serializeSnapshot(); constructor deserializes the snapshot for
in-memory DBs so init() short-circuits migrate()+compat at v129
- store-test-helpers: install/clearInMemoryDbSnapshot harness
- dashboard: db-snapshot-helper mirror (core __tests__ is cross-package)
- convert agent-store, mission-store, workflow-routes suites
Measured (raw db.init(): 43ms -> 5ms, 8x):
- agent-store 13.12s -> 3.32s
- mission-store 17.62s -> 5.69s (min of 3)
- workflow-routes tests 4.38s -> 2.79s (min of 5; not init-dominated)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Diff-proportional verification (deriveFileScopedPnpmTestCommand) + scope-aware
verification timeout, so merge/step checks finish in seconds. Propagated to this
worktree directly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Document why the goals schema-version test follows the exported schema constant.
- Add an FNXC note explaining that fresh database version assertions should track SCHEMA_VERSION.
- Preserve the dynamic schema-version expectation so migration bumps do not leave stale literals behind.
Files changed:
packages/core/src/__tests__/goals-schema.test.ts | 4 ++++
1 file changed, 4 insertions(+)
Fusion-Task-Id: FN-6709
Fusion-Task-Lineage: 05abc460-0816-4f31-a83c-fb0d3ae427e1
TaskStore shutdown now tears down its cached plugin store to avoid leaking plugin database handles.
- Dispose the cached PluginStore during TaskStore.close(), clear the cache, and remove listeners before closing.
- Keep close idempotent when no plugin store was created or after a prior teardown.
- Add regression coverage for direct TaskStore.close() and disk-backed harness reopen lifecycle.
Files changed:
.../src/__tests__/store-plugin-store-close.test.ts | 63 ++++++++++++++++++++++
packages/core/src/store.ts | 15 ++++++
2 files changed, 78 insertions(+)
Fusion-Task-Id: FN-7005
Fusion-Task-Lineage: fe2242d8-7004-4d29-904d-58ee77861d99
Stabilize @fusion/core tests by resetting stale plugin-store database handles and aligning fixture assertions.
- Add a PluginStore close method that disposes local and central database connections.
- Reset the lazy plugin store before the shared test harness clears global settings storage.
- Cover plugin-store reopening from the built-in workflow harness.
- Assert the runtime FN task-prefix fallback without requiring serialized defaults.
Files changed:
packages/core/src/__tests__/builtin-workflows.test.ts | 6 ++++++
packages/core/src/__tests__/store-test-helpers.ts | 8 ++++++++
packages/core/src/__tests__/test-project.test.ts | 8 +++++++-
packages/core/src/plugin-store.ts | 12 ++++++++++++
4 files changed, 33 insertions(+), 1 deletion(-)
Fusion-Task-Id: FN-7003
Fusion-Task-Lineage: ec4d6e45-61ea-4e6a-b296-3af3acf7ba4b
Fix 6 failing tests caused by intentional source changes that landed
without updating dependent test assertions:
- Core test-project: taskPrefix default changed from "FN" to undefined
(commit 800f845e1, derived from project name at runtime)
- Dashboard ScriptsModal.css: replace banned --text-primary with --text
- CLI package-config: update expected pi dep version ^0.79.1 -> ^0.79.9
- CLI skill-sync: document 4 new engine tools in engine-tools.md
- CLI version: update expected release:version script to include
run-ci-distill.mjs
- CLI bundled-plugin-freshness: rebuild stale dist directories
## Summary
Running more than one fusion process on a host (multiple dashboards/CLIs
across worktrees, all attaching `~/.fusion/fusion-central.db`) could
crash a `node` process at random — instantly, with no JS stack and
nothing in the logs. This happened 3 times in 3 days on one machine.
After this change those processes coexist without crashing.
The crash was an OS-level `SIGBUS` (`EXC_BAD_ACCESS`, `FS pagein error`
/ kernel `cluster_pagein past EOF`) inside SQLite's `walIndexReadHdr`.
In WAL mode every connection coordinates through a memory-mapped `-shm`
wal-index; on macOS/APFS, when one process resizes/rebuilds that file
during a checkpoint while another has it mmap'd, the reader faults on
the now-out-of-bounds page. A hardware memory fault can't be caught by
`node:sqlite` or JS, so the whole process dies.
The fix switches the central DB to `journal_mode = DELETE` (rollback
journal), which uses no `-shm` memory map and coordinates cross-process
access via POSIX byte-range locks instead — removing the faulting
surface entirely while keeping multi-process access. The existing
`busy_timeout` absorbs the writer serialization that DELETE mode trades
for WAL's reader/writer concurrency. Per-project DBs (`db.ts`) are
intentionally left on WAL: they're single-process-per-project and don't
hit this cross-process fault. SQLite migrates the existing WAL database
on first open (checkpoints `-wal` into the main file and removes
`-wal`/`-shm`), so there is no data loss.
## Test plan
- New regression tests in `central-db.test.ts` assert the central DB
reports `journal_mode = delete` (not `wal`) and that **no `-shm`
wal-index file is ever created** even after write traffic — i.e. the
exact faulted surface is gone.
- All 6 central-DB suites pass (221 tests); `@fusion/core` typechecks
clean.
---
[](https://github.com/EveryInc/compound-engineering-plugin)

<!-- stage-review-badge-begin -->
---
<a href="https://stagereview.app/Runfusion/Fusion/pull/1752">
<picture>
<source media="(prefers-color-scheme: dark)"
srcset="https://stagereview.app/assets/gh-open-in-stage-dark.svg">
<img src="https://stagereview.app/assets/gh-open-in-stage-light.svg"
alt="Open in Stage">
</picture>
</a>
<!-- stage-review-badge-end -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved stability when multiple dashboards or CLIs run on the same
machine.
* Switched the local database to a safer journaling mode to reduce rare
crash issues on macOS/APFS.
* Prevented creation of extra database side files during normal
operation, while keeping data durability and lock-based coordination in
place.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
- Verify the WAL->DELETE journal-mode switch instead of discarding exec()'s
result. During a rolling upgrade a lingering WAL holder blocks the exclusive
lock the switch needs, so SQLite either throws SQLITE_BUSY or no-ops and
returns "wal". Capture both outcomes and warn loudly so the residual -shm
SIGBUS surface is observable, rather than silently swallowed.
- Do not rethrow: the condition is transient and self-healing (the next start
after the last WAL holder exits migrates cleanly); hard-failing would make the
central DB unopenable during the very upgrade window it describes.
- Add a migration-path regression test (a WAL holder blocking the switch) that
the prior fresh-DB-only tests did not cover.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>