05151a25dbcfe7b6d89f10f9912e2fb7682cd57b
28 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
05151a25db |
feat: faster dashboard and serve startup (#2132)
## Summary Speeds up **time-to-HTTP-ready** for `fn dashboard` and `fn serve` after the PostgreSQL cutover without reintroducing the historical 3s cwd-engine race that degraded webhooks. - **Dashboard store share (serve parity):** inject the factory-booted `TaskStore` as `externalTaskStore` so cwd `ensureEngine` does not open a second pool; share only when store root matches project working directory (multi-project safe). - **Serve multi-project:** stop awaiting `startAll()` before listen; await only the primary engine; background the rest + reconciliation. - **Defer non-route-critical engine work:** ordered OAuth (refresh → monitor), automation schedule syncs, and auto-merge **enqueue** after the engine handle is returnable. - **Critical-path merge status clear:** still clear stale `merging`/`merging-pr` before ready so manual merge is not blocked after crash. - **Serve `--paused`:** apply `enginePaused` before `ensureEngine`/`startAll` (dashboard ordering). - **Stop safety:** generation counter so deferred tails cannot resume after `stop()` clears `shuttingDown`. - **Phase timing:** shared `phaseTime` helper, factory substep logs, serve time-to-listen. Plan: `docs/plans/2026-07-14-001-feat-faster-startup-plan.md` ## Test plan - [x] `packages/engine` — `project-engine-manager.test.ts` (path-matched external store) - [x] `packages/engine` — `project-engine-deferred-startup.test.ts` (status clear, OAuth order, stop generation) - [x] `packages/cli` — `startup-phase.test.ts` - [x] `packages/cli` — `serve.test.ts` (60 tests, including `--paused`) - [ ] Local: warm `fn dashboard` / `fn serve` and compare `startup phase *` / `time-to-listen` logs - [ ] `pnpm smoke:boot` (real serve `/api/health` on ephemeral port) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Performance** * Improved dashboard and serve startup times, including faster time-to-listen and time-to-ready. * Moved non-essential background initialization off the critical startup path. * Parallelized dashboard service initialization where possible. * **Reliability** * Improved multi-project startup handling and project selection. * Prevented cross-project task-store sharing. * Added safer shutdown behavior for partially completed startup. * **Diagnostics** * Added startup phase timing logs to help identify performance bottlenecks. * **Tests** * Expanded coverage for deferred startup, shutdown, project isolation, and startup timing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
78ef3075f6 |
fix(core): prevent plugin migration startup crash
Run retained SQLite plugin recovery through the privileged startup connection before handing stores to the restricted PostgreSQL runtime role. |
||
|
|
a242f1b449 |
fix(FN-7952): migrate bundled plugins to PostgreSQL (#2111)
## Summary Bundled plugins now persist shared runtime state in project-scoped PostgreSQL tables instead of maintaining independent SQLite authority. Reports, CLI Printing Press, Compound Engineering, Roadmap, Even Realities, and WhatsApp all follow the same ownership and startup contract as Fusion core. ## Design decisions - Plugin schema hooks run through the host’s PostgreSQL owner and enforce project isolation. - The SDK exposes the host contract needed by bundled plugins without importing engine internals. - Legacy Roadmap ownership fixtures use the supported empty-owner sentinel, preserving current composite primary/foreign keys while exercising backfill behavior. - The lockfile travels with the Even Realities PostgreSQL dependency so packaged installs remain reproducible. ## Validation - All six affected plugin builds pass. - Affected plugin suites pass: 773 tests across Printing Press, Compound Engineering, Even Realities, Reports, Roadmap, and WhatsApp. - `pnpm test:gate` passes all 478 gate tests. - This PR changes 40 files. ## Stack - Depends on #2110 → #2109 → #2108. - The documentation/release PR completes the stack. Related: #2105 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Breaking Changes** * PostgreSQL is now required for runtime storage; SQLite files are used only as one-time migration inputs. * The legacy `FUSION_NO_EMBEDDED_PG` fallback has been removed. * **New Features** * Added project-isolated PostgreSQL storage for plugins, reports, tasks, notifications, and other plugin data. * Added agent tools for reports and CLI service drafts. * Added PostgreSQL schema initialization support for plugin authors. * **Bug Fixes** * Improved migration and recovery of legacy plugin state. * Prevented cross-project data access and strengthened transactional schema updates. * **Documentation** * Updated storage, migration, deployment, plugin authoring, CLI, and dashboard guidance for PostgreSQL. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
6c008418fe |
fix(core): boot embedded Postgres under non-admin user on elevated Windows (#2117)
## Summary Windows embedded Postgres verification (CI `windows-latest` and elevated desktop) fails because PostgreSQL refuses to run under an administrative token: > Execution of PostgreSQL by a user with administrative permissions is not permitted. GitHub Actions runners execute as `runneradmin` elevated, so the existing `test:embedded-postgres` smoke (and any elevated Local-mode desktop launch) cannot start the server. ### Fix - When `isWindowsElevatedAdmin()` is true, **initdb / clients stay as the launcher**, but the **postgres server** is started as a dedicated non-admin local user (`fusion-pg`) via PowerShell `Start-Process -Credential`. - Readiness waits on the postgres log line `database system is ready to accept connections` with a lightweight poll (no per-iteration `tasklist`). - Real-process vitest cases use a **180s** timeout on Windows (package default is 15s, which killed healthy boots mid-start). - Builds on top of the packaged-desktop asar materialization work already on main (#2106). ## Test plan - [x] `pnpm --filter @fusion/core test:embedded-postgres` on macOS (33/33) - [ ] `desktop-windows.yml` on `feature/win-pg-verify`: - [ ] Smoke embedded Postgres on Windows - [ ] Build + package Windows EXE - [ ] Verify app.asar assets - [ ] Optional: download portable EXE and manual Local mode smoke on a Windows host <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved embedded PostgreSQL startup on Windows when Fusion runs with elevated administrator privileges. - When elevated, the embedded database now boots under a dedicated non-administrator local account, with more reliable readiness detection, logging, and shutdown cleanup. - Enhanced database provisioning and now prefers `127.0.0.1` for Windows connection addressing. - **Tests** - Added coverage for Windows elevation detection without starting embedded PostgreSQL. - Increased platform-dependent timeouts for embedded real-process tests to avoid premature failures on Windows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
2e4fcfcaea |
fix(FN-7952): establish PostgreSQL core authority (#2108)
## Summary Fusion’s core runtime now treats PostgreSQL as the authoritative metadata store without leaving current CLI, dashboard, desktop, or engine composition roots uncompilable between stack layers. This is the 99-file foundation for the larger cutover: subsequent PRs migrate the remaining consumers, plugins, and operator surfaces. ## Design decisions - Runtime store construction fails closed when an asynchronous PostgreSQL layer is unavailable; SQLite remains readable only at explicit migration and identity-recovery boundaries. - Project ownership is enforced across active, archived, workflow, mission, analytics, and plugin-schema data. - The small set of cross-package files in this layer are compatibility-critical call sites required for a green intermediate commit, not the complete consumer migration. - Schema migration 0008 remains assigned to session-advisor state from current `main`; mission lineage idempotency advances to 0009 so neither invariant can be skipped. ## Validation - All affected package typechecks pass: Core, Engine, Dashboard, CLI, and Desktop. - `pnpm test:gate` passes: 478 tests across the engine gate, PostgreSQL core gate, and CLI workflow shape. - The PR changes exactly 99 files. ## Stack This is the base PR. Engine/dashboard, CLI/desktop/ops, plugins, and docs/release follow as stacked PRs, each below 100 changed files. Related: #2105 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PostgreSQL is now the standard runtime backend, with embedded PostgreSQL enabled by default. * Added project-scoped storage for tasks, archives, chat sessions, missions, knowledge pages, and operational data. * Improved archived-task search, filtering, pagination, and restoration. * Added safer plugin schema initialization with validation and project isolation. * Added PostgreSQL-backed workflow, mission, validator, and dashboard capabilities. * **Bug Fixes** * Improved startup timeout cancellation and resource cleanup. * Prevented cross-project data access and phantom reservation cleanup errors. * Ensured archived tasks remain read-only and asynchronous writes complete reliably. * Retired SQLite opt-out settings with clear startup errors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
f4e78abeb7 |
fix: resolve CREATE ROLE fusion_runtime race condition in migration 0006 (#2104)
## Summary Fixes `CREATE ROLE fusion_runtime` race condition in migration `0006_project_ownership.sql` that causes 30 compound-engineering test failures on CI. ## Root Cause Concurrent test databases on the same PostgreSQL service container race on `CREATE ROLE fusion_runtime`: the `IF NOT EXISTS` check is not atomic (roles are cluster-wide, not per-database). Between the check and the `CREATE ROLE`, another session can create the role, causing error `23505` (unique_violation). ## Fix Replace the non-atomic `IF NOT EXISTS` guard with a `BEGIN...EXCEPTION WHEN duplicate_object OR unique_violation THEN NULL; END;` block that safely handles the race. ## Verification | Check | Result | |---|---| | compound-engineering (pipeline-store + orchestrator + session-routes) | ✅ 41 passed | | Engine shard 1/2 | ✅ 3826 passed, 0 failed | | Merge gate | ✅ 471 passed | | Lint | ✅ exit 0 | |
||
|
|
7568705244 |
fix(desktop): boot embedded Postgres in packaged app and ship omp dist (#2106)
## Summary Packaged Fusion desktop Local mode failed after the SQLite→Postgres cutover: 1. **Embedded Postgres** could not start from `app.asar` — platform packages resolve `initdb`/`postgres` via `import.meta.url` into the asar virtual path, and `spawn` fails with `ENOTDIR`. 2. **After Postgres was fixed**, Local mode still fell back to the mode chooser because `@fusion-plugin-examples/omp-runtime` was never built into `dist/` (dashboard imports it from `runtime-provider-probes.ts`). This PR makes packaged Local mode boot embedded Postgres reliably and keep the dashboard shell up. ### Changes - **CJS bootstrap** (`main-bootstrap.cjs`) as Electron `main`: patches `child_process.spawn` / `fs.promises.stat|chmod` before the ESM main loads so asar binary paths rewrite to real files. - **Materialize** the full native PG install (`bin` + `lib` + `share`) under `~/.fusion/embedded-postgres/runtime-bin/<plat-arch>/`. - **electron-builder**: full `asarUnpack` of embedded-postgres packages; allowlist PG deps and `@fusion-plugin-examples/**/*` (+ plugin-sdk / ACP SDK). - **Build** `fusion-plugin-omp-runtime` with the other dashboard-static runtime plugins; export `DASHBOARD_RUNTIME_PLUGIN_PACKAGES` for tests. - Unit coverage for asar path rewrite, packaging allowlists, and omp build inclusion. ## Test plan - [x] `pnpm --filter @fusion/core test:embedded-postgres` (23/23) - [x] Desktop packaging unit tests (`build-bundling`, `electron-builder-config`) - [x] Packaged macOS `Fusion.app` Local mode: - [x] `embedded postgres: ready on port … (database "fusion")` - [x] `desktopMode` stays `"local"` (no chooser fallback) - [x] `GET /api/health` → `status: ok`, `database.healthy: true`, `engine.available: true` - [x] Linux embedded binary lifecycle smoke (Docker aarch64, `@embedded-postgres/linux-arm64`) — initdb/start/persist/restart - [ ] CI release desktop jobs (macOS/Linux) when this lands - [ ] Windows packaged desktop Local + PG (separate agent / host) ## Verification notes | Platform | Embedded Postgres | Packaged Local shell | |----------|-------------------|----------------------| | macOS | Working | Working after this PR | | Linux | Native binary smoke pass | Full AppImage not built on this host | | Windows | Out of scope here | Separate verification | <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved embedded PostgreSQL reliability in Electron-packaged apps by rewriting bundled `app.asar` binary paths to their unpacked/materialized locations. * Ensured embedded PostgreSQL runtime binaries resolve correctly across platforms/architectures, with best-effort executable permissions and macOS dylib link normalization. * **Packaging** * Updated the desktop Electron entry to use a bootstrap module for embedded PostgreSQL binary resolution. * Expanded Electron Builder inclusion and asar-unpack rules for embedded-postgres and related packages, plus required runtime plugin/sdk assets. * **Tests** * Updated and added checks to match the new packaging and plugin/runtime expectations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
4f037679ad |
feat: planner overseer session advisor (OMP advisor parity) (#2082)
## Summary Adds a **session advisor** to the planner overseer so Fusion can review live executor transcripts the way [oh-my-pi’s advisor](https://github.com/can1357/oh-my-pi/tree/main/packages/coding-agent/src/advisor) does — without replacing the existing lifecycle supervisor (stage watch, retry, merge confirmation, human-control withhold). ### What ships - **Emission guard** (`OverseerEmissionGuard`) — content-free phrase filter, session dedupe with severity-rank escalation, one accept per advisor update - **Session delta runtime** — queues agent-log deltas, drains through an advisor agent, drops backlog after 3 failures - **Session advisor service** — model gate, level matrix (`observe` / `steer` / `autonomous`), human-control re-check at inject, `[session-advisor]` steering comments - **OVERSEER.md / WATCHDOG.md** discovery for project review priorities - **AgentLogger `onEntriesFlushed`** + poll-backed agent-log cursor for durable deltas - Workflow settings: `plannerOverseerAdvisorProvider` + `plannerOverseerAdvisorModelId` (both required; empty = soft-disabled for cost safety) - Docs + changeset ### What does not ship (deferred) - Multi-advisor YAML roster, mutating advisor tools, reviewer/merger shadowing, true tool-abort interrupt ### Plan `docs/plans/2026-07-13-001-feat-overseer-advisor-parity-plan.md` ## Enablement 1. Set workflow **Session advisor model provider** + **Session advisor model id** 2. Oversight level `observe` (log only), `steer`, or `autonomous` (inject) 3. Optional: add `OVERSEER.md` or `WATCHDOG.md` in the project ## Test plan - [x] `pnpm --filter @fusion/core exec vitest run src/__tests__/overseer-emission-guard.test.ts` - [x] `pnpm --filter @fusion/engine exec vitest run` overseer-* unit tests (21 tests) - [x] Related planner-overseer / intervention regression tests - [x] `@fusion/engine` + `@fusion/core` typecheck - [ ] Manual: configure advisor model, run an executor task, confirm `[session-advisor]` inject + timeline metadata when concern is raised ## Residual Review Findings None from autofix pass (log-cursor ordering fix already committed). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an off-by-default “session advisor” that can review live execution activity and provide severity-based guidance. * Added project and per-task controls to enable it, including a default enable switch and Quick Add / Task Detail toggles. * Enhanced advisor prompting by discovering and incorporating `OVERSEER.md`/`WATCHDOG.md` review files. * **Documentation** * Added architecture and settings documentation for the new session-advisor parity behavior. * **Bug Fixes** * Improved fail-soft handling so advisor behavior won’t disrupt execution. * Fixed concurrent PostgreSQL migration startup failures. * **Tests** * Added coverage for advice parsing, emission guarding, runtime behavior, and watchdog discovery. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
cdf67c1d98 |
fix(dashboard): stop Planning Mode retry loop, make AI sessions multi-tab (#2101)
## Problem Reported: planning gets stuck in a cycle of retrying and regenerating after a response was already supplied. After the user answers a planning question, `submitResponse` pushed the answer to history but left `session.currentQuestion` pointing at the just-answered question for the whole next generation. The planning SSE route's catch-up path re-emits `currentQuestion` to every fresh connection — and each FN-7946 auto-retry (#2073) opens a fresh connection. So after any generation error: 1. Auto-retry connects a fresh stream → the server re-emits the **already-answered** question. 2. The client treats any question event as progress: it **resets the 3-attempt auto-retry budget** and re-shows the answered question. 3. The retry regenerates; if it errors again the cycle repeats with a fresh budget — an unbounded retry/regenerate loop. Re-answering the stale question also 409-collided with the in-flight generation, feeding the same loop. ## Fix Invariant: `currentQuestion` is only set while the session is genuinely awaiting user input. - `submitResponse` clears it the moment an answer is accepted (normal turns and the deepening checkpoint), while preserving the legacy 200 respond contract on generation failure (the modal ignores the body and lets the SSE error drive recovery). - `retrySession` scrubs stale questions persisted by pre-fix builds before regenerating. - `buildSessionFromRow` only restores a question when the persisted row is `awaiting_input`. - `didSubmitSameAnswer` now compares against the last history entry so the duplicate-submit 409 message survives. - Agent onboarding gets the same fix (its SSE route also re-emits `currentQuestion` on connect); retry now asks the next question instead of re-asking the answered one. Surface enumeration: mission and milestone interviews keep questions the same way but their SSE routes never re-emit on connect, and the auto-retry budget machinery is Planning-Mode-only — planning + onboarding were the two affected surfaces. ## Symptom Verification - **Original symptom:** after answering a question, Planning Mode loops between "Retrying…" and regenerating, re-showing the already-answered question, with the auto-retry budget never exhausting. - **Exact reproduction:** answer a question, have the next generation fail (stuck watchdog/provider error), let the client auto-retry open a fresh SSE connection. - **Assertion it is gone:** new regression suite `planning-answered-question-reemit.test.ts` asserts `currentQuestion` is cleared mid-generation, on generation failure, on retry, and on restore from non-`awaiting_input` rows — so the SSE catch-up path has nothing stale to re-emit. All 5 tests fail against pre-fix code and pass with the fix; an onboarding regression test covers the sibling surface. ## Verification - New regression tests: 5/5 fail on pre-fix code, pass with the fix (plus 1 onboarding test). - Existing suites: 137 planning server tests pass (3 failures in `routes-planning.test.ts` fail identically without this change — pre-existing on the branch); all 69 `PlanningModeModal.planning-flow` client tests pass; `tsc --noEmit` clean; `pnpm check:changesets` passes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Made Planning Mode (and related planning controls) lock-free and multi-tab—no more take-over/active-in-another-tab lock overlays. * **Bug Fixes** * Fixed Planning Mode retry/generation flows where already-answered questions could reappear. * Ensured answered questions clear immediately and aren’t re-emitted during session recovery/SSE catch-up. * Improved session restoration and preserved legacy recovery behavior when generation fails after an answer. * **Tests** * Added regression coverage for the answered-question invariant and updated existing tests to reflect lock-free behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ## Follow-up: Planning Mode is now multi-tab via DB state (lock-free) Second commit removes all cross-tab coordination from planning — the persisted session row is the single source of truth and multiple tabs can read and interact with the same session: - **Server:** `/planning/*` routes no longer run `checkSessionLock` or parse `tabId`; a stale `tabId` from an older client is ignored instead of 409'd. Subtask/mission interview routes keep their existing lock behavior. - **Client:** `PlanningModeModal` drops `useSessionLock`, the `useAiSessionSync` BroadcastChannel broadcasts, `sessionTabId`/`lockSessionId` state, and the "Take Control" overlay. Tabs stay current via the per-session SSE stream plus the global `ai_session:updated` events `useBackgroundSessions` already consumes; concurrent writes resolve via the server's generation-in-progress guard (409). - **API client:** planning functions lose their `tabId` params. - **Fix uncovered by the refactor:** the 8s stuck-poll now resolves the session id inside each tick — the removed lock state was what previously re-armed the poll after Start Planning resolved the session id. - Also fixes a pre-existing PG-cutover break in `planning-generation-cancellation.test.ts` (`getSession` is async). Verification: 144 client planning tests and 137 server planning tests pass (the 3 remaining `routes-planning.test.ts` failures are pre-existing on the branch and fail identically without these changes); `tsc --noEmit` and eslint clean on changed files; `pnpm check:changesets` passes. Lock-conflict route tests were rewritten to assert lock-free semantics, plus a new modal test proving a session stays fully interactive with no lock acquisition even when another tab is active. --- ## Follow-up 2: the per-tab session lock is gone entirely Third commit extends the multi-tab model from planning to **every** AI interview surface (planning, subtask breakdown, mission interview, milestone/slice interview) and deletes the lock machinery root and branch. **Server** - Deleted the `/ai-sessions/:id/lock`, `/lock/force`, and `/lock/beacon` routes. - Dropped `checkSessionLock` from every planning/subtask/mission/milestone route (both copies — `routes.ts` and `mission-routes.ts`). A `tabId` from an older client is ignored, never 409'd; all `tabId` body parsing is gone. - Dropped `acquireLock` / `releaseLock` / `forceAcquireLock` / `getLockHolder` / `releaseStaleLocks` from `AiSessionStore`, plus the `@fusion/core` async helpers (`acquireAiSessionLock` et al) and core's re-exports. - Removed `lockedByTab`/`lockedAt` from `AiSessionRow`/`AiSessionSummary`, the upsert SQL, and all four session producers. **Client** - Deleted `useSessionLock` and the now-orphaned `getSessionTabId` util. - Removed the Take Control overlay, the "active in another tab" banners, and `BackgroundTasksIndicator`'s active-elsewhere gate (the confirm prompt and lock badge — sessions now just open). - Reduced `useAiSessionSync` to what its own comments already called it — a low-latency *status* supplement to SSE: no `activeTabMap`, `broadcastLock/Unlock/Heartbeat`, `owningTabId`, `tab:*` messages, or stale-heartbeat sweep. - Dropped `tabId` from every session API client function; removed the lock CSS. **Deliberately kept: the two DB columns.** `ai_sessions.locked_by_tab` / `locked_at` remain as dead, always-NULL columns with a deprecation note. Dropping them is an irreversible migration, and released binaries still name those columns explicitly in their upsert — an older install pointed at the same database would fail every session write. They can be dropped once no such binary can reach it. No code reads or writes them. **Verification**: 397 client tests and 137 server planning tests pass (the same 3 `routes-planning.test.ts` failures are pre-existing — verified identical on a clean stash); `tsc --noEmit` clean for `@fusion/core` and `@fusion/dashboard`; eslint clean on all changed files; the 30 PG `schema-applier` tests pass (they exercise the retained columns); `pnpm check:changesets` passes. The lock-conflict route tests and both modal lock tests were rewritten to assert the inverse: routes and modals stay fully interactive while another tab "holds" a lock, and the lock API is never called. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
be55d0a987 |
fix(cli): reuse project stores for skill discovery (#2102)
## Summary - reuse the dashboard command's backend-aware per-project `TaskStore` cache during project-scoped plugin skill discovery - obtain plugin state through `TaskStore.getPluginStore()` instead of constructing bare SQLite-default `PluginStore` / `TaskStore` instances - keep cached project stores alive for the dashboard process while still stopping request-scoped plugin loaders - add a regression covering the real Skills adapter callback and refresh the dashboard test fixture with `getAsyncLayer()` ## Root cause `GET /api/skills/discovered` resolved the project correctly, then `getProjectScopedPluginSkills()` constructed new stores without an `AsyncDataLayer`. After `VAL-REMOVAL-005`, that enters the physically removed synchronous SQLite runtime and returns HTTP 500 even when PostgreSQL health, projects, tasks, and both project engines are healthy. The existing route tests mocked the Skills adapter callback, so they did not exercise this CLI wiring. ## Verification - targeted dashboard regression: 1 passed, 91 skipped - `pnpm lint` - `pnpm --filter @runfusion/fusion typecheck` - `pnpm --filter @runfusion/fusion build` - `pnpm check:changesets --strict` - `git diff --check` Live Atlas validation against the migrated embedded PostgreSQL runtime: - `/api/skills/discovered?projectId=proj_84f4645c2da64288`: HTTP 200, 36 skills - `/api/skills/discovered?projectId=proj_7538a9dd46c24c5f`: HTTP 200, 36 skills - local dashboard and Tailscale dashboard: HTTP 200 - controlled SIGTERM: launchd restarted the dashboard and both Skills routes remained healthy <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed dashboard project-scoped plugin-skill discovery in PostgreSQL mode with safer store reuse/teardown and request-scoped plugin-loader lifecycle. - Improved dashboard cleanup to avoid duplicate concurrent store closes and ensured proper shutdown behavior per root type. - Made `fusion_runtime` role creation race-safe during concurrent PostgreSQL migrations. - **New Features** - Added `persistRuntimeState` option to control whether plugin runtime state changes are persisted. - **Tests** - Expanded dashboard and core hot-reload tests to verify scoped, non-persistent runtime behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
678265a526 |
fix(cli): show live SQLite migration progress
Report source scans, per-table copy milestones, checksum phases, verification outcomes, and unambiguous failure or finalization status during first-boot and manual migrations. |
||
|
|
0312d2e140 | fix(core): preserve late SQLite columns during cutover | ||
|
|
7677ab07dc |
fix: add chat_sessions columns to schema baseline + fix remaining PG auth bugs (shard 4) (#2096)
## Summary Fixes shard 4 full-suite failures: chat_sessions schema baseline gap + two remaining PG auth bugs missed by PR #2086. **Scope: shard 4 only.** Shards 1/2 (engine timeouts) and shard 3 (compound-engineering CI-only failure) are separate issues not addressed here. ## Changes ### Schema baseline gap — `chat_sessions` missing columns (42703 error) - **`0000_initial.sql`**: Added `validator_thinking_level` and `planning_thinking_level` columns to `CREATE TABLE project.chat_sessions`. These exist in the Drizzle schema (`project.ts:1492-1493`) but were missing from the SQL baseline, causing `column does not exist` on all chat_sessions inserts in fresh test databases. - **`postgres-health.ts`**: Added both columns to `EXPECTED_PROJECT_COLUMNS` self-heal list so existing databases also get them via ALTER TABLE. **Fixes**: `chat-store-content-search-edit.pg.test.ts` (5 tests), `satellite-db-injected-stores.test.ts` (2 tests) ### Remaining auth bugs (password auth failed for user "runner") - **`allocator-cross-project.test.ts`**: Still had `process.env.USER` in inline adminExec — missed by PR #2086's batch fix. Replaced with `PG_TEST_URL_BASE` connection string. - **`connection.test.ts`**: Used `FUSION_PG_TEST_URL` (not set on CI) with a bare default URL lacking credentials. `postgres.js` fell back to OS user `runner`. Changed to derive from `FUSION_PG_TEST_URL_BASE` which includes credentials. **Fixes**: `allocator-cross-project.test.ts` (2 tests), `connection.test.ts` (3 tests) ## Verification | Check | Result | |---|---| | Merge gate (`pnpm test:gate`) | ✅ 294 + 114 + 63 = 471 passed | | chat-store-content-search-edit | ✅ 5 passed | | satellite-db-injected-stores | ✅ 10 passed | | allocator-cross-project | ✅ 2 passed | | connection | ✅ 13 passed | | Lint | ✅ exit 0 | | Typecheck | ✅ clean | ## Not in scope - **Shards 1/2**: Engine test suite timeouts with `getAsyncLayer`/`updateSettings` mock warnings. Pre-existing. - **Shard 3**: `compound-engineering stage-skill-loading.test.ts` — 14 tests fail on CI (`TypeError: Cannot read properties of undefined (reading 'close')`), pass locally. Likely CI-specific teardown issue. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added separate `validator_thinking_level` and `planning_thinking_level` fields to chat session data, including database schema and health-check recognition. * **Bug Fixes** * Improved PostgreSQL test connectivity by using configured connection URL settings instead of hardcoded local defaults. * Made Postgres-related test teardown null-safe to avoid failures when setup doesn’t complete. * **Tests** * Updated automated test quarantine/exclusions for known failing engine and reliability-interaction cases. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
379d450c38 |
fix(core): preserve required empty JSON during migration (#2099)
## Summary - preserve empty and whitespace-only legacy SQLite text as JSON string scalars when the PostgreSQL target is required `jsonb` without a default - keep nullable/defaulted JSON behavior unchanged - canonicalize converted JSON before source/target checksum comparison - cover empty, whitespace, malformed, and scalar workflow IR values ## Test plan - `FUSION_PG_TEST_URL_BASE=postgresql://127.0.0.1:55432 nix shell nixpkgs#postgresql_15 -c bash -c 'corepack pnpm --filter @fusion/core exec vitest run src/__tests__/postgres/sqlite-migrator.test.ts -t "preserves empty, whitespace, malformed, and scalar values" --reporter=dot'`\n- `corepack pnpm --filter @fusion/core typecheck`\n- `corepack pnpm check:changesets --strict`\n- `corepack pnpm --filter @runfusion/fusion build` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved SQLite-to-PostgreSQL migrations for required `jsonb` fields. - Preserves empty, whitespace-only, malformed, and scalar JSON values instead of replacing them with defaults or `NULL`. - Maintains existing `nullable` and default-value behavior. - Improved migration verification for converted `jsonb` data. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
945d629e3b |
fix(core): make SQLite cutover lossless and project-local
Preserve legacy-only tables, recover partial migration ownership, and enforce project-local keys, relationships, agents, merge queues, task IDs, archives, and monitor state with PostgreSQL RLS. Report successful cutovers once in the dashboard and system inbox with retained SQLite paths and Discord support details. |
||
|
|
7c8a84fb2f |
fix(core): converge multi-project SQLite cutover
Migrate central SQLite state once per cluster, isolate project metadata, and preserve file-local revision identities while verifying accumulated shared tables. |
||
|
|
12a4fbe9bb | fix(core): complete legacy SQLite cutover | ||
|
|
99870ba329 | fix(core): recover partial PostgreSQL migrations | ||
|
|
c25f8b796d |
Harden PostgreSQL migration foundation (#2088)
## Summary - make SQLite-to-PostgreSQL cutover retryable, fail-closed, versioned, and transactionally serialized - isolate migration sessions from runtime traffic and apply schema upgrades through `0002` - enforce tenant ownership across automations, analytics, activity, usage, agent runs, evals, and todos - replace expired SQLite-only coverage with PostgreSQL parity and concurrency coverage This is PR 1 of 2. The stacked follow-up restores PostgreSQL parity for CLI, engine, dashboard, and bundled integrations. ## Verification - `pnpm check:changesets --strict` - `pnpm --filter @fusion/core typecheck` - migration schema, connection, and SQLite cutover suite: 57 tests passed - `pnpm test:gate`: 463 tests passed ## Post-Deploy Monitoring & Validation - take a restorable PostgreSQL backup before deploy - confirm `fusion_schema_migrations` contains `0002` - confirm each expected project has a complete `fusion_sqlite_migrations` row - verify no null or empty tenant ownership in automations, activity logs, agent runs, and usage events - monitor for ownership inference failures, cutover verification failures, and migration session errors - restore the backup for data rollback; do not downgrade the tenant-isolation schema in place <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PostgreSQL-backed analytics and live dashboard metrics are now project-scoped (activity, tools, monitor, signals, and live snapshots). * Evaluation runs and scheduled eval batches received lifecycle improvements (ordering, updates, and execution flow). * Todo list changes now emit events; WhatsApp persistence and project-scoped roadmap data are supported. * **Bug Fixes** * SQLite-to-PostgreSQL cutovers now fail safely with stronger verification, serialized cutover handling, and safer project ownership. * PostgreSQL backend writes and reads are now strictly project-isolated and fail closed when project context is missing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
bc348345a4 |
fix(engine): break Plan Review REVISE replan loop (feedback + bounded cap) (#2078)
## Problem A task whose Plan Review step returns verdict `REVISE` can loop forever: plan → plan-review REVISE → `needs-replan` → re-plan → near-identical plan → REVISE → repeat. The triage **pre-execution** Plan Review gate (`runPlanReviewBeforeExecution`) sets `status: "needs-replan"` on REVISE with **no cap and no escape to `awaiting-approval`** — unlike the executor graph path, which already has `PLAN_REVIEW_REPLAN_HARD_CAP`. Under `planApprovalMode: require-all` there is also no human exit, because the task never reaches `awaiting-approval`. Separately, replan feedback (`triage.ts`) was derived only from `task.log` comment actions + the latest user comment; it never consulted the plan-review verdict stored in `task.workflowStepResults`. ## Fix 1. **Thread plan-review feedback into replan** — when re-planning with no comment-derived feedback, seed `buildSpecificationPrompt` from the most recent `plan-review` REVISE `output` in `workflowStepResults` (existing user/AI-comment precedence preserved). 2. **Bounded cap** — new `planReviewReplanCount` counter (`types.ts`, `store.ts` column + updateTask, `db.ts` migration 146, `manual-retry-reset.ts`). After `PLAN_REVIEW_GATE_REPLAN_CAP = 3` consecutive REVISE replans the task escalates to `awaiting-approval` (`awaitingApprovalReason: "plan-review-replan-cap"`) instead of replanning. Counter resets on APPROVE. ## Tests Adds `triage-replan-feedback-from-plan-review.test.ts` and `triage-plan-review-replan-cap.test.ts`. Merge gate green locally (`verify:fast`, `test:gate` 337+63, `lint`); changeset included. Made with Claude (see `Co-Authored-By` trailer). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Prevented Plan Review “REVISE” from looping indefinitely by enforcing a bounded replan cap. * After repeated Plan Review replans, tasks now escalate to an approval-hold state with a dedicated reason. * Improved replan feedback by seeding from the latest Plan Review output when no explicit feedback is available; the counter clears when Plan Review approves. * Manual retries now reset the Plan Review replan cap counter. * **Documentation** * Added release notes describing the Plan Review replan safeguards. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
b5c76af700 |
fix(core): preserve jsonb defaults during PostgreSQL migration (#2080)
## Summary - preserve target defaults when legacy SQLite rows contain `NULL` or empty strings for `NOT NULL` jsonb columns - derive the fallback from PostgreSQL column metadata instead of hard-coding table or column names - keep migration checksum conversion aligned with inserted values - add regression coverage for legacy null JSON fields ## Test plan - `corepack pnpm@10.33.0 --filter @fusion/core typecheck` - `FUSION_PG_TEST_SKIP=1 corepack pnpm@10.33.0 --filter @fusion/core exec vitest run src/__tests__/postgres/sqlite-migrator.test.ts` - `corepack pnpm@10.33.0 --filter @fusion/core build` The PostgreSQL-backed integration suite requires `psql`, which is unavailable in this environment; CI should exercise the added migration case against PostgreSQL. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved SQLite-to-PostgreSQL migration for legacy rows containing `NULL` or empty JSON values. * For eligible `NOT NULL` `jsonb` columns, the migrator now preserves/apply compatible PostgreSQL column defaults instead of writing SQL `NULL`. * Migration verification now aligns with the final values inserted into PostgreSQL to prevent checksum mismatches. * **Tests** * Added an end-to-end legacy migration case to confirm `jsonb` fields materialize as empty defaults (e.g., `[]`) rather than staying `NULL`. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
8e4514e585 |
fix: key workflow settings by the central project id and stamp all partitioned tables on both migration paths
Closes the remaining PG-cutover partitioning gaps: - getWorkflowSettingsProjectId resolves the bound AsyncDataLayer's central- registry id first. In backend mode the SQLite stub's getProjectIdentity() throws, so the old fallback ALWAYS keyed workflow_settings / workflow_prompt_overrides by the rootDir path string — a namespace nothing else reads, making workflow settings appear reset after cutover. - Stamping is extracted into core stampMigratedProjectRows (tasks/archived NULL->id, config ''->id, workflow_settings + workflow_prompt_overrides rootDir-key->id, all guarded against clobbering per-project rows), shared by startup-factory Step 5.5 and 'fn db migrate', which now resolves the registered project by path after the copy and warns when unregistered. - The task-id allocator and merge_queue are verified safe WITHOUT project partitioning: task ids are a global PK, the per-prefix sequence scans are intentionally global (only the per-project config floor can raise them), so two projects sharing a prefix cannot mint duplicate ids. FNXC comments lock the invariant; a cross-project PG regression test proves it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3ccc9f96e6 |
fix: bind rootDir boots to the central project registry and re-key the migrated config row
Implements the central-project-identity architecture: cwd/rootDir is ONLY a
lookup key into central.projects; project identity (the partition key for
every task/config read and write) comes from the registry.
- createTaskStoreForBackend resolves the registered project id by path for
rootDir-only boots and binds the AsyncDataLayer to it. Previously
'fn dashboard' / 'fn serve' / desktop booted their main store UNBOUND, so
unscoped API requests wrote NULL-project_id rows the projectId-bound engine
could never see, and unbound config reads (id = 1) were indeterminate once
multiple per-project rows existed. The engine already worked registry-first
(resolveLocalProjectWorkingDirectory); this brings the store boots in line.
- Step 5.5 auto-migration now also re-keys the migrated legacy config row
('' -> project id, guarded against clobbering an existing per-project row).
configScope() has no bound->'' fallback, so the migrated project settings,
workflowSteps, taskPrefix, and nextId counters were silently invisible to
bound readers right after a successful migration.
- Unregistered paths resolve to undefined and boot unbound, preserving legacy
single-project behavior with unfiltered readers.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0f3a3d3f49 |
fix: stamp migrated task rows with the central-registry project id on rootDir-only boots
The SQLite -> PostgreSQL auto-migration leaves project_id NULL and Step 5.5 only stamped rows when options.projectId was bound — but 'fn dashboard' in the project directory (the main cutover path) boots with rootDir only, so every migrated row stayed NULL, project-bound readers (engine InProcessRuntime, dashboard project-store-resolver) filtered them all out, and the board showed no tasks right after a successful migration. The stamping id is now resolved from the freshly-migrated central registry by matching the registered project path to rootDir; projects never registered centrally keep NULL rows, matching their unbound readers. Integration test covers the rootDir-only stamp. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
7aa969892a |
fix: count actually-inserted rows in the SQLite -> PostgreSQL migrator via RETURNING
insertBatch read the driver wrapper's count (result.count ?? result.rowCount ?? rows.length), which reported 0 through drizzle's execute even when every row landed — migration reports showed 'inserted 0' for fully-migrated tables and the startup banner's migratedRows total was wrong. ON CONFLICT DO NOTHING RETURNING 1 yields exactly one row per row actually inserted, making the count driver-agnostic and correctly excluding conflict-skipped rows. Idempotency test now asserts first-run insertedRows == sourceRows and re-run insertedRows == 0. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fbcd00204a |
fix: snake_case legacy SQLite table names in the PG migrator and keep PG-mode boots from touching SQLite
Two post-cutover fixes: 1. The SQLite -> PostgreSQL migrator matched table names verbatim while only column names were snake_cased, so all 22 legacy camelCase tables (activityLog, runAuditEvents, mergeQueue, taskClaims, projectNodePathMappings, ...) resolved zero PostgreSQL columns and were silently skipped as 'no PostgreSQL counterpart'. First observed as 'Project/node path mapping not found' on engine start because central.project_node_path_mappings was never populated. TablePlan now carries a snake_cased pgTable used for every PostgreSQL-side operation; regression test migrates a camelCase activityLog into project.activity_log. 2. The first-boot auto-migration guard opened .fusion/fusion.db with a read-write DatabaseSync on every boot (isValidSqliteDatabaseFile), which performs WAL recovery + checkpoint — writing the legacy file on each PG boot. The PG emptiness count now runs before the SQLite probe, so steady-state PG boots never open the legacy SQLite files at all. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
eb5c81cc59 |
fix: widen ce_sessions.last_activity_at to bigint so PG first-boot migration survives epoch-ms values
project.ce_sessions.last_activity_at stores Date.now() epoch milliseconds but was declared integer in both the Drizzle shape and the CE plugin schema-hook DDL, overflowing PG int4 during the SQLite -> PostgreSQL first-boot auto-migration and blocking startup at task-store init. Now bigint in both sites, with an idempotent ALTER for datadirs that already materialized the integer column, plus a schema-wide invariant test that no numeric *_at/*_time/*_timestamp column is 32-bit integer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
c15c78feeb |
feat: migrate storage from SQLite to PostgreSQL (#1793)
# Migrate storage from SQLite to PostgreSQL — full dashboard cutover Migrates Fusion's storage layer to the embedded PostgreSQL `AsyncDataLayer` (the default backend) and **completes the satellite-store + feature cutover** so every dashboard and Command Center surface works in PG mode. ## Status — every surface works in embedded-PG mode Verified live against a running embedded-Postgres dashboard (all **200**, zero 5xx) and gate-tested (**23 files / 99 tests** on embedded PG, plus engine-core 294 and ci-shape 63 in the blocking merge gate; core/engine/cli/dashboard typecheck clean). | Area | Surfaces | State | |---|---|---| | Satellite stores | workflows, todos, insights, research, missions, goals, mailbox | ✅ | | Views | artifacts, documents, evals | ✅ | | Command Center | activity, productivity, team, tokens, tools, **workflows**, **github**, **signals**, **plugin-activations**, **live** (all 10) | ✅ | | Run execution | insight generation, research run execution | ✅ (store-path; AI step needs a provider) | | Live updates | SSE push for mission/research/insight events | ✅ | | Workflow editing | create / update / delete / select (+ id counter) | ✅ | | Engine | mission autopilot, incident-signal ingestion, regression storm-guard, agent wake-on-message | ✅ | | Core | tasks, agents, secrets, automations, memory, chat, usage, PRs, git | ✅ | ## Approach Each satellite store gets an `Async<Store>` wrapper exposing the sync store's method names over the existing `async-*-store.ts` helpers; `get<Store>Store()` returns a `Sync | Async` union; consumers `await` (harmless on sync), and engine/CLI paths that can't convert use `instanceof Sync` graceful fallback. Analytics aggregators branch on `"ping" in dbOrLayer` to run schema-qualified raw SQL over `project.*` (snake_case) in PG. Executors/orchestrators/autopilot are await-converted to drive the union store; the async store wrappers extend `EventEmitter` so SSE live-push fires in both backends. Not-yet-ported capabilities degrade gracefully (never 500) and are individually called out in commits. ## Sync with main The branch is kept continuously merged with `main` (currently through FN-7845, 2026-07-12); the earlier "final rebase deferred" note no longer applies. Use **Create a merge commit** (or squash) to land it — GitHub's rebase-merge cannot replay a merge-maintained branch. ## Residual Review Findings Multi-agent code review of the PostgreSQL satellite-store ports (U1–U5) applied 3 safe fixes (see `fix(review): apply autofix feedback`). The following are **real but gated** — recorded here as follow-up work rather than auto-applied. All are SQLite→PostgreSQL **concurrency/atomicity regressions**: the sync stores were immune only by SQLite's single-writer, single-threaded-handler execution; the async ports open multi-await read-modify-write windows. **Reachability is low today** because the execution engines that generate concurrent same-run mutations (insight run executor, research orchestrator/dispatcher) are `instanceof`-gated to sync mode in PG. No process-crash class survived (all engine fallbacks correctly guard the sync store). - **[P1] Research `appendResearchEvent` dual-write is non-atomic** (`packages/core/src/async-research-store.ts`, corroborated: adversarial + reliability). The `research_run_events` insert (own transaction) and the `run.events` jsonb update are separate writes — a crash between them, or two concurrent appends, splits the table count from the jsonb array. **Fix:** perform the seq-insert and the jsonb update in one `layer.transactionImmediate`. - **[P1] Research run terminal-reversion via stale full-row persist** (`async-research-store.ts` `persistResearchRun`/`updateResearchStatus`). Concurrent `PATCH /runs/:id/status` + `POST /runs/:id/events` can revert a terminal run to `running` by overwriting the whole row, bypassing the transition guard. **Fix:** scoped column `UPDATE`s with a `WHERE status …` guard, or optimistic version column. - **[P2] `updateResearchRun`/`updateInsightRun` read-then-write TOCTOU** — concurrent PATCHes last-writer-wins on the lifecycle merge. **Fix:** `SELECT … FOR UPDATE` / enclosing transaction. - **[P2] `upsertRun`/`createRunOrThrowConflict` check-then-create race** (`async-insight-store.ts`) — two callers can each create an "active" run. **Fix:** partial unique index on `(projectId, trigger) WHERE status IN ('pending','running')`. - **[P3] `createResearchRetryRun` return-value divergence** — sync returns the pre-update `queued` snapshot; async returns the reloaded `retry_waiting` run (persisted state is identical). Pick one side for cross-backend parity. - **[P2/perf] Mission `getMissionWithHierarchy`/`getMissionHealth` N+1 fan-out** — O(milestones×slices) sequential round-trips hold one pool slot per request; can starve the pool for large hierarchies. **Fix:** batched/joined reads. - **Testing gaps:** no PG-mode concurrency tests (interleaved status/event mutations), no sync↔async parity assertion for the lifecycle-error codes, and no mission status/health rollup parity test vs the sync `MissionStore`. ~~Out of scope (deferred): AI run *execution* (insight/research) + mission autopilot + live SSE mission events remain sync-gated/degraded in PG mode.~~ **Since ported** — insight/research run execution, mission autopilot, and SSE live push all run on the async layer now, which also makes the concurrency findings above genuinely reachable; they remain open follow-ups. --- ## Update — 2026-07-12: production-readiness hardening & live acceptance Everything below landed on this branch since the description above was written: **Production blockers from review — fixed** - `recoverStaleTransitionPending` ported to the async layer (backend moves write + clear the crash-safe marker; startup/maintenance sweeps no longer throw). - Lost-update class fixed: `atomicWriteTaskJson`/`WithAudit` write changed columns only (full-row upserts silently resurrected stale fields across concurrent store instances — the "task stuck unplanned forever" bug). - First-boot **auto-migration**: booting the PG backend over a project with a legacy `fusion.db` migrates it automatically (loud failure, SQLite kept as backup), and the dashboard shows a one-time **"your data was migrated" banner** with the backup paths and a Need-help Discord link. - `pg_dump`/`pg_restore` discovered from common install locations for embedded-mode backups. - The PG suite is part of the blocking merge gate (`test:pg-gate`). **Multi-project isolation (PR #2007, merged into this branch)** - `project_id` partition key on tasks / archived tasks / config, `taskProjectScope` threaded through every scan/claim/count, per-project config rows, layer bound to the project at startup. - Review P1 follow-up: the shared cold-storage `archive.archived_tasks` table is also partitioned and all archived-board reads/counts/searches are scoped. - Schema drift self-heal generalized to schema-qualified columns so existing databases upgrade in place. **Other changes** - Node settings sync **removed** in PG mode (409 `settings-sync-disabled-postgres`) — nodes share state by connecting to the same database; auth sync kept (per-machine file). - Perf (review findings): `listTasks` pushes column filter + ORDER BY + LIMIT/OFFSET into SQL; `getConversation` capped to the most recent 200 messages. - Fixed a false "operator action required" pause-abort log fired on every successfully auto-merged task. **Live acceptance — PASSED (2026-07-12)** A sandboxed instance (isolated HOME, embedded PG, real Opus executor) ran a task through the complete cycle: create → triage (AI spec) → execute → in-review → AI squash-merge landed on the project's `main` → done. A write+read sweep of every data surface (settings, comments, documents, attachments + artifact bridge + artifact edit, chat with real generation, goals, missions, agent mail, secrets, workflows, memory, CC analytics) was green on embedded PG. **Known remaining work** - The per-project `config` PK re-key has no upgrade path for pre-isolation embedded-PG databases (needs a real `DROP CONSTRAINT`/re-key migration; fresh databases are fine). - `pg_dump`/`pg_restore` binaries are not yet bundled in release artifacts (PATH/common-location discovery only). - The satellite-store concurrency findings listed above. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Phil Larson <hello@phillarson.xyz> Co-authored-by: fusion-merge <fusion-merge@local> |