# Migrate storage from SQLite to PostgreSQL — full dashboard cutover Migrates Fusion's storage layer to the embedded PostgreSQL `AsyncDataLayer` (the default backend) and **completes the satellite-store + feature cutover** so every dashboard and Command Center surface works in PG mode. ## Status — every surface works in embedded-PG mode Verified live against a running embedded-Postgres dashboard (all **200**, zero 5xx) and gate-tested (**23 files / 99 tests** on embedded PG, plus engine-core 294 and ci-shape 63 in the blocking merge gate; core/engine/cli/dashboard typecheck clean). | Area | Surfaces | State | |---|---|---| | Satellite stores | workflows, todos, insights, research, missions, goals, mailbox | ✅ | | Views | artifacts, documents, evals | ✅ | | Command Center | activity, productivity, team, tokens, tools, **workflows**, **github**, **signals**, **plugin-activations**, **live** (all 10) | ✅ | | Run execution | insight generation, research run execution | ✅ (store-path; AI step needs a provider) | | Live updates | SSE push for mission/research/insight events | ✅ | | Workflow editing | create / update / delete / select (+ id counter) | ✅ | | Engine | mission autopilot, incident-signal ingestion, regression storm-guard, agent wake-on-message | ✅ | | Core | tasks, agents, secrets, automations, memory, chat, usage, PRs, git | ✅ | ## Approach Each satellite store gets an `Async<Store>` wrapper exposing the sync store's method names over the existing `async-*-store.ts` helpers; `get<Store>Store()` returns a `Sync | Async` union; consumers `await` (harmless on sync), and engine/CLI paths that can't convert use `instanceof Sync` graceful fallback. Analytics aggregators branch on `"ping" in dbOrLayer` to run schema-qualified raw SQL over `project.*` (snake_case) in PG. Executors/orchestrators/autopilot are await-converted to drive the union store; the async store wrappers extend `EventEmitter` so SSE live-push fires in both backends. Not-yet-ported capabilities degrade gracefully (never 500) and are individually called out in commits. ## Sync with main The branch is kept continuously merged with `main` (currently through FN-7845, 2026-07-12); the earlier "final rebase deferred" note no longer applies. Use **Create a merge commit** (or squash) to land it — GitHub's rebase-merge cannot replay a merge-maintained branch. ## Residual Review Findings Multi-agent code review of the PostgreSQL satellite-store ports (U1–U5) applied 3 safe fixes (see `fix(review): apply autofix feedback`). The following are **real but gated** — recorded here as follow-up work rather than auto-applied. All are SQLite→PostgreSQL **concurrency/atomicity regressions**: the sync stores were immune only by SQLite's single-writer, single-threaded-handler execution; the async ports open multi-await read-modify-write windows. **Reachability is low today** because the execution engines that generate concurrent same-run mutations (insight run executor, research orchestrator/dispatcher) are `instanceof`-gated to sync mode in PG. No process-crash class survived (all engine fallbacks correctly guard the sync store). - **[P1] Research `appendResearchEvent` dual-write is non-atomic** (`packages/core/src/async-research-store.ts`, corroborated: adversarial + reliability). The `research_run_events` insert (own transaction) and the `run.events` jsonb update are separate writes — a crash between them, or two concurrent appends, splits the table count from the jsonb array. **Fix:** perform the seq-insert and the jsonb update in one `layer.transactionImmediate`. - **[P1] Research run terminal-reversion via stale full-row persist** (`async-research-store.ts` `persistResearchRun`/`updateResearchStatus`). Concurrent `PATCH /runs/:id/status` + `POST /runs/:id/events` can revert a terminal run to `running` by overwriting the whole row, bypassing the transition guard. **Fix:** scoped column `UPDATE`s with a `WHERE status …` guard, or optimistic version column. - **[P2] `updateResearchRun`/`updateInsightRun` read-then-write TOCTOU** — concurrent PATCHes last-writer-wins on the lifecycle merge. **Fix:** `SELECT … FOR UPDATE` / enclosing transaction. - **[P2] `upsertRun`/`createRunOrThrowConflict` check-then-create race** (`async-insight-store.ts`) — two callers can each create an "active" run. **Fix:** partial unique index on `(projectId, trigger) WHERE status IN ('pending','running')`. - **[P3] `createResearchRetryRun` return-value divergence** — sync returns the pre-update `queued` snapshot; async returns the reloaded `retry_waiting` run (persisted state is identical). Pick one side for cross-backend parity. - **[P2/perf] Mission `getMissionWithHierarchy`/`getMissionHealth` N+1 fan-out** — O(milestones×slices) sequential round-trips hold one pool slot per request; can starve the pool for large hierarchies. **Fix:** batched/joined reads. - **Testing gaps:** no PG-mode concurrency tests (interleaved status/event mutations), no sync↔async parity assertion for the lifecycle-error codes, and no mission status/health rollup parity test vs the sync `MissionStore`. ~~Out of scope (deferred): AI run *execution* (insight/research) + mission autopilot + live SSE mission events remain sync-gated/degraded in PG mode.~~ **Since ported** — insight/research run execution, mission autopilot, and SSE live push all run on the async layer now, which also makes the concurrency findings above genuinely reachable; they remain open follow-ups. --- ## Update — 2026-07-12: production-readiness hardening & live acceptance Everything below landed on this branch since the description above was written: **Production blockers from review — fixed** - `recoverStaleTransitionPending` ported to the async layer (backend moves write + clear the crash-safe marker; startup/maintenance sweeps no longer throw). - Lost-update class fixed: `atomicWriteTaskJson`/`WithAudit` write changed columns only (full-row upserts silently resurrected stale fields across concurrent store instances — the "task stuck unplanned forever" bug). - First-boot **auto-migration**: booting the PG backend over a project with a legacy `fusion.db` migrates it automatically (loud failure, SQLite kept as backup), and the dashboard shows a one-time **"your data was migrated" banner** with the backup paths and a Need-help Discord link. - `pg_dump`/`pg_restore` discovered from common install locations for embedded-mode backups. - The PG suite is part of the blocking merge gate (`test:pg-gate`). **Multi-project isolation (PR #2007, merged into this branch)** - `project_id` partition key on tasks / archived tasks / config, `taskProjectScope` threaded through every scan/claim/count, per-project config rows, layer bound to the project at startup. - Review P1 follow-up: the shared cold-storage `archive.archived_tasks` table is also partitioned and all archived-board reads/counts/searches are scoped. - Schema drift self-heal generalized to schema-qualified columns so existing databases upgrade in place. **Other changes** - Node settings sync **removed** in PG mode (409 `settings-sync-disabled-postgres`) — nodes share state by connecting to the same database; auth sync kept (per-machine file). - Perf (review findings): `listTasks` pushes column filter + ORDER BY + LIMIT/OFFSET into SQL; `getConversation` capped to the most recent 200 messages. - Fixed a false "operator action required" pause-abort log fired on every successfully auto-merged task. **Live acceptance — PASSED (2026-07-12)** A sandboxed instance (isolated HOME, embedded PG, real Opus executor) ran a task through the complete cycle: create → triage (AI spec) → execute → in-review → AI squash-merge landed on the project's `main` → done. A write+read sweep of every data surface (settings, comments, documents, attachments + artifact bridge + artifact edit, chat with real generation, goals, missions, agent mail, secrets, workflows, memory, CC analytics) was green on embedded PG. **Known remaining work** - The per-project `config` PK re-key has no upgrade path for pre-isolation embedded-PG databases (needs a real `DROP CONSTRAINT`/re-key migration; fresh databases are fine). - `pg_dump`/`pg_restore` binaries are not yet bundled in release artifacts (PATH/common-location discovery only). - The satellite-store concurrency findings listed above. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Phil Larson <hello@phillarson.xyz> Co-authored-by: fusion-merge <fusion-merge@local>
178 lines
6.1 KiB
JavaScript
178 lines
6.1 KiB
JavaScript
#!/usr/bin/env node
|
|
/*
|
|
FNXC:PostgresCutover 2026-07-05-13:00:
|
|
Ported from direct node:sqlite access on .fusion/fusion.db to the PostgreSQL
|
|
backend (scripts/lib/backend-db.mjs). The FN-3899 recovery logic
|
|
(planRecoverBlockedBy) is pure and operates on plain row arrays so tests need
|
|
no database; only the thin apply step touches PostgreSQL. In PG the `log`
|
|
column is jsonb (already an array), unlike the SQLite JSON-string column.
|
|
*/
|
|
import fs from "node:fs";
|
|
import path from "node:path";
|
|
import process from "node:process";
|
|
import { execSync } from "node:child_process";
|
|
import { openBackend, rowsOf } from "./lib/backend-db.mjs";
|
|
|
|
function parseArgs(argv) {
|
|
const flags = new Set(argv.slice(2));
|
|
return {
|
|
dryRun: !flags.has("--apply"),
|
|
apply: flags.has("--apply"),
|
|
};
|
|
}
|
|
|
|
export function parseFileScopeFromPromptText(promptText) {
|
|
const headerMatch = promptText.match(/^##\s+File Scope\s*$/m);
|
|
if (!headerMatch || headerMatch.index === undefined) return [];
|
|
const start = headerMatch.index + headerMatch[0].length;
|
|
const rest = promptText.slice(start);
|
|
const nextHeader = rest.search(/^##\s+/m);
|
|
const section = nextHeader >= 0 ? rest.slice(0, nextHeader) : rest;
|
|
const paths = [];
|
|
const regex = /`([^`]+)`/g;
|
|
let match;
|
|
while ((match = regex.exec(section)) !== null) {
|
|
const value = match[1].trim();
|
|
if (value) paths.push(value);
|
|
}
|
|
return [...new Set(paths)];
|
|
}
|
|
|
|
export function pathsOverlap(a, b) {
|
|
for (const pa of a) {
|
|
const prefixA = pa.endsWith("/*") ? pa.slice(0, -1) : null;
|
|
for (const pb of b) {
|
|
const prefixB = pb.endsWith("/*") ? pb.slice(0, -1) : null;
|
|
const cleanA = prefixA ? pa.slice(0, -2) : pa;
|
|
const cleanB = prefixB ? pb.slice(0, -2) : pb;
|
|
if (cleanA === cleanB) return true;
|
|
if (prefixA && pb.startsWith(prefixA)) return true;
|
|
if (prefixB && pa.startsWith(prefixB)) return true;
|
|
if (prefixA && prefixB && (prefixA.startsWith(prefixB) || prefixB.startsWith(prefixA))) return true;
|
|
if (pa === pb) return true;
|
|
}
|
|
}
|
|
return false;
|
|
}
|
|
|
|
function loadScope(tasksDir, taskId) {
|
|
const promptPath = path.join(tasksDir, taskId, "PROMPT.md");
|
|
if (!fs.existsSync(promptPath)) return [];
|
|
return parseFileScopeFromPromptText(fs.readFileSync(promptPath, "utf8"));
|
|
}
|
|
|
|
function isTerminalColumn(column) {
|
|
return column === "done" || column === "archived";
|
|
}
|
|
|
|
/**
|
|
* Pure FN-3899 planning: rows are { id, column, blockedBy, worktree, paused }.
|
|
* Returns findings; entries with newBlocker === null are the repairs.
|
|
*/
|
|
export function planRecoverBlockedBy({ rows, tasksDir }) {
|
|
const byId = new Map(rows.map((row) => [row.id, row]));
|
|
|
|
const activeScopes = new Map();
|
|
for (const row of rows) {
|
|
const isActive = row.column === "in-progress" || (row.column === "in-review" && row.worktree && !row.paused);
|
|
if (!isActive) continue;
|
|
const scope = loadScope(tasksDir, row.id);
|
|
if (scope.length > 0) activeScopes.set(row.id, scope);
|
|
}
|
|
|
|
const findings = [];
|
|
for (const row of rows) {
|
|
if (row.column !== "todo" || !row.blockedBy) continue;
|
|
|
|
const blocker = byId.get(row.blockedBy);
|
|
const taskScope = loadScope(tasksDir, row.id);
|
|
let reason = null;
|
|
|
|
if (!blocker) {
|
|
reason = "blocker-missing";
|
|
} else if (isTerminalColumn(blocker.column)) {
|
|
reason = `blocker-terminal:${blocker.column}`;
|
|
} else if (blocker.column === "in-review" && !blocker.worktree) {
|
|
reason = "blocker-in-review-without-worktree";
|
|
} else {
|
|
const blockerScope = activeScopes.get(blocker.id) ?? [];
|
|
if (taskScope.length === 0 || blockerScope.length === 0 || !pathsOverlap(taskScope, blockerScope)) {
|
|
reason = "scope-no-overlap";
|
|
}
|
|
}
|
|
|
|
if (!reason) {
|
|
findings.push({ taskId: row.id, oldBlocker: row.blockedBy, newBlocker: row.blockedBy, reason: "unchanged" });
|
|
continue;
|
|
}
|
|
|
|
findings.push({ taskId: row.id, oldBlocker: row.blockedBy, newBlocker: null, reason });
|
|
}
|
|
|
|
return findings;
|
|
}
|
|
|
|
export async function recoverBlockedBy({ backend, tasksDir, dryRun = true }) {
|
|
const { asyncLayer, sql } = backend;
|
|
const rows = rowsOf(
|
|
await asyncLayer.db.execute(sql`
|
|
SELECT id, "column", blocked_by AS "blockedBy", worktree, paused, log
|
|
FROM project."tasks"
|
|
WHERE deleted_at IS NULL
|
|
`),
|
|
);
|
|
|
|
const findings = planRecoverBlockedBy({ rows, tasksDir });
|
|
if (dryRun) return findings;
|
|
|
|
const byId = new Map(rows.map((row) => [row.id, row]));
|
|
const now = new Date().toISOString();
|
|
for (const finding of findings) {
|
|
if (finding.newBlocker === finding.oldBlocker) continue;
|
|
const row = byId.get(finding.taskId);
|
|
const log = Array.isArray(row?.log) ? [...row.log] : [];
|
|
log.push({
|
|
at: now,
|
|
message: "Recovered: cleared stale blockedBy via FN-3899 recovery",
|
|
outcome: `Recovered: cleared stale blockedBy via FN-3899 recovery (reason: ${finding.reason})`,
|
|
});
|
|
await asyncLayer.db.execute(sql`
|
|
UPDATE project."tasks"
|
|
SET blocked_by = NULL, log = ${JSON.stringify(log)}::jsonb, updated_at = ${now}
|
|
WHERE id = ${finding.taskId}
|
|
`);
|
|
}
|
|
|
|
return findings;
|
|
}
|
|
|
|
function resolveProjectRoot() {
|
|
const commonDir = execSync("git rev-parse --git-common-dir", { encoding: "utf8" }).trim();
|
|
return path.resolve(commonDir, "..");
|
|
}
|
|
|
|
function printFindings(findings, dryRun) {
|
|
const changed = findings.filter((row) => row.oldBlocker !== row.newBlocker);
|
|
console.log(dryRun ? "Mode: DRY RUN" : "Mode: APPLY");
|
|
console.log("taskId\toldBlocker\tnewBlocker\treason");
|
|
for (const row of findings) {
|
|
if (row.oldBlocker === row.newBlocker) continue;
|
|
console.log(`${row.taskId}\t${row.oldBlocker}\t${row.newBlocker ?? "NULL"}\t${row.reason}`);
|
|
}
|
|
console.log(`Repairs: ${changed.length}`);
|
|
}
|
|
|
|
if (import.meta.url === `file://${process.argv[1]}`) {
|
|
const { dryRun } = parseArgs(process.argv);
|
|
const projectRoot = resolveProjectRoot();
|
|
const tasksDir = path.join(projectRoot, ".fusion", "tasks");
|
|
|
|
const backend = await openBackend(projectRoot);
|
|
try {
|
|
const findings = await recoverBlockedBy({ backend, tasksDir, dryRun });
|
|
printFindings(findings, dryRun);
|
|
} finally {
|
|
await backend.shutdown().catch(() => {});
|
|
}
|
|
}
|