Files
fusion/scripts/audit-branch-cross-contamination.mjs
gsxdsm c15c78feeb feat: migrate storage from SQLite to PostgreSQL (#1793)
# Migrate storage from SQLite to PostgreSQL — full dashboard cutover

Migrates Fusion's storage layer to the embedded PostgreSQL
`AsyncDataLayer` (the default backend) and **completes the
satellite-store + feature cutover** so every dashboard and Command
Center surface works in PG mode.

## Status — every surface works in embedded-PG mode

Verified live against a running embedded-Postgres dashboard (all
**200**, zero 5xx) and gate-tested (**23 files / 99 tests** on embedded
PG, plus engine-core 294 and ci-shape 63 in the blocking merge gate;
core/engine/cli/dashboard typecheck clean).

| Area | Surfaces | State |
|---|---|---|
| Satellite stores | workflows, todos, insights, research, missions,
goals, mailbox | ✅ |
| Views | artifacts, documents, evals | ✅ |
| Command Center | activity, productivity, team, tokens, tools,
**workflows**, **github**, **signals**, **plugin-activations**, **live**
(all 10) | ✅ |
| Run execution | insight generation, research run execution | ✅
(store-path; AI step needs a provider) |
| Live updates | SSE push for mission/research/insight events | ✅ |
| Workflow editing | create / update / delete / select (+ id counter) |
✅ |
| Engine | mission autopilot, incident-signal ingestion, regression
storm-guard, agent wake-on-message | ✅ |
| Core | tasks, agents, secrets, automations, memory, chat, usage, PRs,
git | ✅ |

## Approach

Each satellite store gets an `Async<Store>` wrapper exposing the sync
store's method names over the existing `async-*-store.ts` helpers;
`get<Store>Store()` returns a `Sync | Async` union; consumers `await`
(harmless on sync), and engine/CLI paths that can't convert use
`instanceof Sync` graceful fallback. Analytics aggregators branch on
`"ping" in dbOrLayer` to run schema-qualified raw SQL over `project.*`
(snake_case) in PG. Executors/orchestrators/autopilot are
await-converted to drive the union store; the async store wrappers
extend `EventEmitter` so SSE live-push fires in both backends.

Not-yet-ported capabilities degrade gracefully (never 500) and are
individually called out in commits.

## Sync with main

The branch is kept continuously merged with `main` (currently through
FN-7845, 2026-07-12); the earlier "final rebase deferred" note no longer
applies. Use **Create a merge commit** (or squash) to land it — GitHub's
rebase-merge cannot replay a merge-maintained branch.

## Residual Review Findings

Multi-agent code review of the PostgreSQL satellite-store ports (U1–U5)
applied 3 safe fixes (see `fix(review): apply autofix feedback`). The
following are **real but gated** — recorded here as follow-up work
rather than auto-applied. All are SQLite→PostgreSQL
**concurrency/atomicity regressions**: the sync stores were immune only
by SQLite's single-writer, single-threaded-handler execution; the async
ports open multi-await read-modify-write windows. **Reachability is low
today** because the execution engines that generate concurrent same-run
mutations (insight run executor, research orchestrator/dispatcher) are
`instanceof`-gated to sync mode in PG. No process-crash class survived
(all engine fallbacks correctly guard the sync store).

- **[P1] Research `appendResearchEvent` dual-write is non-atomic**
(`packages/core/src/async-research-store.ts`, corroborated: adversarial
+ reliability). The `research_run_events` insert (own transaction) and
the `run.events` jsonb update are separate writes — a crash between
them, or two concurrent appends, splits the table count from the jsonb
array. **Fix:** perform the seq-insert and the jsonb update in one
`layer.transactionImmediate`.
- **[P1] Research run terminal-reversion via stale full-row persist**
(`async-research-store.ts` `persistResearchRun`/`updateResearchStatus`).
Concurrent `PATCH /runs/:id/status` + `POST /runs/:id/events` can revert
a terminal run to `running` by overwriting the whole row, bypassing the
transition guard. **Fix:** scoped column `UPDATE`s with a `WHERE status
…` guard, or optimistic version column.
- **[P2] `updateResearchRun`/`updateInsightRun` read-then-write TOCTOU**
— concurrent PATCHes last-writer-wins on the lifecycle merge. **Fix:**
`SELECT … FOR UPDATE` / enclosing transaction.
- **[P2] `upsertRun`/`createRunOrThrowConflict` check-then-create race**
(`async-insight-store.ts`) — two callers can each create an "active"
run. **Fix:** partial unique index on `(projectId, trigger) WHERE status
IN ('pending','running')`.
- **[P3] `createResearchRetryRun` return-value divergence** — sync
returns the pre-update `queued` snapshot; async returns the reloaded
`retry_waiting` run (persisted state is identical). Pick one side for
cross-backend parity.
- **[P2/perf] Mission `getMissionWithHierarchy`/`getMissionHealth` N+1
fan-out** — O(milestones×slices) sequential round-trips hold one pool
slot per request; can starve the pool for large hierarchies. **Fix:**
batched/joined reads.
- **Testing gaps:** no PG-mode concurrency tests (interleaved
status/event mutations), no sync↔async parity assertion for the
lifecycle-error codes, and no mission status/health rollup parity test
vs the sync `MissionStore`.

~~Out of scope (deferred): AI run *execution* (insight/research) +
mission autopilot + live SSE mission events remain sync-gated/degraded
in PG mode.~~ **Since ported** — insight/research run execution, mission
autopilot, and SSE live push all run on the async layer now, which also
makes the concurrency findings above genuinely reachable; they remain
open follow-ups.







---

## Update — 2026-07-12: production-readiness hardening & live acceptance

Everything below landed on this branch since the description above was
written:

**Production blockers from review — fixed**
- `recoverStaleTransitionPending` ported to the async layer (backend
moves write + clear the crash-safe marker; startup/maintenance sweeps no
longer throw).
- Lost-update class fixed: `atomicWriteTaskJson`/`WithAudit` write
changed columns only (full-row upserts silently resurrected stale fields
across concurrent store instances — the "task stuck unplanned forever"
bug).
- First-boot **auto-migration**: booting the PG backend over a project
with a legacy `fusion.db` migrates it automatically (loud failure,
SQLite kept as backup), and the dashboard shows a one-time **"your data
was migrated" banner** with the backup paths and a Need-help Discord
link.
- `pg_dump`/`pg_restore` discovered from common install locations for
embedded-mode backups.
- The PG suite is part of the blocking merge gate (`test:pg-gate`).

**Multi-project isolation (PR #2007, merged into this branch)**
- `project_id` partition key on tasks / archived tasks / config,
`taskProjectScope` threaded through every scan/claim/count, per-project
config rows, layer bound to the project at startup.
- Review P1 follow-up: the shared cold-storage `archive.archived_tasks`
table is also partitioned and all archived-board reads/counts/searches
are scoped.
- Schema drift self-heal generalized to schema-qualified columns so
existing databases upgrade in place.

**Other changes**
- Node settings sync **removed** in PG mode (409
`settings-sync-disabled-postgres`) — nodes share state by connecting to
the same database; auth sync kept (per-machine file).
- Perf (review findings): `listTasks` pushes column filter + ORDER BY +
LIMIT/OFFSET into SQL; `getConversation` capped to the most recent 200
messages.
- Fixed a false "operator action required" pause-abort log fired on
every successfully auto-merged task.

**Live acceptance — PASSED (2026-07-12)**
A sandboxed instance (isolated HOME, embedded PG, real Opus executor)
ran a task through the complete cycle: create → triage (AI spec) →
execute → in-review → AI squash-merge landed on the project's `main` →
done. A write+read sweep of every data surface (settings, comments,
documents, attachments + artifact bridge + artifact edit, chat with real
generation, goals, missions, agent mail, secrets, workflows, memory, CC
analytics) was green on embedded PG.

**Known remaining work**
- The per-project `config` PK re-key has no upgrade path for
pre-isolation embedded-PG databases (needs a real `DROP
CONSTRAINT`/re-key migration; fresh databases are fine).
- `pg_dump`/`pg_restore` binaries are not yet bundled in release
artifacts (PATH/common-location discovery only).
- The satellite-store concurrency findings listed above.

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Phil Larson <hello@phillarson.xyz>
Co-authored-by: fusion-merge <fusion-merge@local>
2026-07-13 19:07:58 -07:00

225 lines
7.5 KiB
JavaScript

#!/usr/bin/env node
/*
FNXC:PostgresCutover 2026-07-05-13:00:
Ported from the sqlite3 CLI on .fusion/fusion.db to the PostgreSQL backend
(scripts/lib/backend-db.mjs). The git contamination analysis
(analyzeBranchCrossContamination) is pure given task rows, so tests inject
rows directly; only row loading touches PostgreSQL.
*/
import { execFileSync } from "node:child_process";
import fs from "node:fs";
import path from "node:path";
import { openBackend, rowsOf } from "./lib/backend-db.mjs";
function runGit(projectRoot, args, { allowFailure = false } = {}) {
try {
return execFileSync("git", args, { cwd: projectRoot, encoding: "utf8", stdio: ["ignore", "pipe", "pipe"] }).trim();
} catch (error) {
if (allowFailure) return null;
throw error;
}
}
function parseArgs(argv) {
const options = {
projectRoot: process.cwd(),
outPath: null,
};
for (let i = 0; i < argv.length; i += 1) {
const arg = argv[i];
if (arg === "--project-root") {
options.projectRoot = path.resolve(argv[i + 1] ?? process.cwd());
i += 1;
} else if (arg.startsWith("--out=")) {
options.outPath = path.resolve(arg.slice("--out=".length));
} else if (arg === "--out") {
options.outPath = path.resolve(argv[i + 1] ?? "audit-branch-cross-contamination.json");
i += 1;
}
}
return options;
}
function parseTaskIdFromSubject(subject) {
const match = String(subject).match(/^[a-z]+\((FN-\d+)\):/i);
return match ? match[1].toUpperCase() : null;
}
function parseTaskIdFromBody(body) {
const match = String(body).match(/(?:^|\n)Fusion-Task-Id:\s*(FN-\d+)\s*(?:\n|$)/i);
return match ? match[1].toUpperCase() : null;
}
function expectedBranch(taskId, branch) {
if (branch && String(branch).trim()) return String(branch).trim();
return `fusion/${String(taskId).toLowerCase()}`;
}
function branchExists(projectRoot, branchName) {
return runGit(projectRoot, ["rev-parse", "--verify", `refs/heads/${branchName}`], { allowFailure: true }) !== null;
}
function resolveMainRef(projectRoot) {
if (runGit(projectRoot, ["rev-parse", "--verify", "origin/main"], { allowFailure: true })) {
return "origin/main";
}
return "main";
}
function resolveBaseCommit(projectRoot, branchName, taskBaseCommitSha) {
if (taskBaseCommitSha && String(taskBaseCommitSha).trim()) {
return { baseCommitSha: String(taskBaseCommitSha).trim(), source: "task.baseCommitSha" };
}
const mainRef = resolveMainRef(projectRoot);
const fallback = runGit(projectRoot, ["merge-base", mainRef, branchName], { allowFailure: true });
if (fallback) {
return { baseCommitSha: fallback.trim(), source: `merge-base(${mainRef},${branchName})` };
}
return { baseCommitSha: null, source: "unresolved" };
}
function parseCommits(raw) {
if (!raw) return [];
const lines = raw.split("\n").map((line) => line.trim()).filter(Boolean);
return lines.map((line) => {
const [sha, subject, body] = line.split("\u001f");
const trailerTaskId = parseTaskIdFromBody(body ?? "");
const subjectTaskId = parseTaskIdFromSubject(subject ?? "");
return {
sha,
subject,
trailerTaskId,
subjectTaskId,
attributedTaskId: trailerTaskId ?? subjectTaskId,
};
});
}
/** Pure analysis over injected task rows ({ id, title, branch, baseCommitSha, columnName }). */
export function analyzeBranchCrossContamination({ projectRoot = process.cwd(), taskRows }) {
const report = {
generatedAt: new Date().toISOString(),
projectRoot,
scannedTaskCount: taskRows.length,
scannedColumns: ["triage", "todo", "in-progress", "in-review"],
taintedTaskCount: 0,
missingBranchCount: 0,
tasks: [],
};
for (const task of taskRows) {
const taskId = String(task.id).toUpperCase();
const branchName = expectedBranch(taskId, task.branch);
const baseResolution = resolveBaseCommit(projectRoot, branchName, task.baseCommitSha);
const baseCommitSha = baseResolution.baseCommitSha;
if (!branchExists(projectRoot, branchName)) {
report.missingBranchCount += 1;
process.stderr.write(`[skip] ${taskId}: branch not found locally (${branchName})\n`);
report.tasks.push({
taskId,
title: task.title,
branchName,
baseCommitSha,
column: task.columnName,
skipped: true,
reason: "branch-missing-local",
});
continue;
}
if (!baseCommitSha) {
report.tasks.push({
taskId,
title: task.title,
branchName,
baseCommitSha: null,
baseResolutionSource: baseResolution.source,
column: task.columnName,
skipped: true,
reason: "missing-baseCommitSha",
});
continue;
}
const rawLog = runGit(projectRoot, ["log", `${baseCommitSha}..${branchName}`, "--format=%H%x1f%s%x1f%b"]);
const commits = parseCommits(rawLog);
const taintedCommits = commits.filter((commit) => commit.attributedTaskId && commit.attributedTaskId !== taskId);
const ownCommits = commits.filter((commit) => commit.attributedTaskId === taskId);
const isTainted = taintedCommits.length > 0;
if (isTainted) report.taintedTaskCount += 1;
report.tasks.push({
taskId,
title: task.title,
branchName,
baseCommitSha,
baseResolutionSource: baseResolution.source,
column: task.columnName,
totalCommits: commits.length,
taskAttributedCommitCount: ownCommits.length,
tainted: isTainted,
taintedCommits: taintedCommits.map((commit) => ({
sha: commit.sha,
subject: commit.subject,
foreignTaskId: commit.attributedTaskId,
})),
recommendation: !isTainted ? "clean" : ownCommits.length > 0 ? "refile" : "force-reset",
commits,
});
}
return report;
}
export async function auditBranchCrossContamination({ projectRoot = process.cwd() } = {}) {
const backend = await openBackend(projectRoot);
let taskRows;
try {
const { asyncLayer, sql } = backend;
taskRows = rowsOf(await asyncLayer.db.execute(sql`
SELECT id, title, branch, base_commit_sha AS "baseCommitSha", "column" AS "columnName"
FROM project."tasks"
WHERE deleted_at IS NULL AND "column" IN ('triage','todo','in-progress','in-review')
ORDER BY id
`));
} finally {
await backend.shutdown().catch(() => {});
}
return analyzeBranchCrossContamination({ projectRoot, taskRows });
}
function renderSummary(report) {
const lines = [
`Branch cross-contamination audit`,
`Scanned tasks: ${report.scannedTaskCount}`,
`Tainted branches: ${report.taintedTaskCount}`,
`Missing local branches: ${report.missingBranchCount}`,
"",
];
for (const task of report.tasks) {
if (task.skipped) continue;
if (!task.tainted) continue;
lines.push(`${task.taskId} | branch=${task.branchName} | base=${task.baseCommitSha} | commits=${task.totalCommits} | recommendation=${task.recommendation}`);
for (const commit of task.taintedCommits) {
lines.push(` - ${commit.sha.slice(0, 12)} ${commit.subject} [foreign=${commit.foreignTaskId ?? "unknown"}]`);
}
}
return lines.join("\n");
}
if (import.meta.url === `file://${process.argv[1]}`) {
const options = parseArgs(process.argv.slice(2));
const report = await auditBranchCrossContamination({ projectRoot: options.projectRoot });
const json = JSON.stringify(report, null, 2);
process.stdout.write(`${json}\n`);
process.stderr.write(`${renderSummary(report)}\n`);
if (options.outPath) {
fs.mkdirSync(path.dirname(options.outPath), { recursive: true });
fs.writeFileSync(options.outPath, json);
process.stderr.write(`\nWrote report: ${options.outPath}\n`);
}
}