38d128ca29287ee6b8abfd18c316cc81bd0ff526
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
25b3c06d2d |
fix(plugins): compound-engineering pipelines stalled forever on a renamed board (#3022)
Closes #3020 — which I filed **instead of** fixing, on a rationale that turned out to be wrong. I said the plugin had no scaffolding for faking `CePipelineStore` + `taskStore` together. It does: `_harness.ts` already builds a real `PluginContext` over a live PostgreSQL layer. The gap was **two missing readers on its task-store stub**, not missing infrastructure. I checked the harness only after filing. ## The defect `TERMINAL_COLUMNS` is `{in-review, done}`, and the reconciler advances a pipeline only when **every** current-stage board task is in that set. On a board whose review and completion lanes are renamed that's false for every task, permanently: - the pipeline never advances a stage - it never creates its outbound task - it sits `running` indefinitely Nothing errors, so it reads as work that hasn't finished. Unlike the display defects in this family (#3014, #3017), the CE flow actually **stops**. ## Shape The decision is extracted to an exported `isStageTerminalColumn` because it *is* the whole decision. Left private it could only be reached through a pipeline-state + links + board-tasks fixture, and the half that needed proving is that a renamed board resolves to its own lanes through this store. It uses `resolveReviewColumns` rather than re-deriving the union — that helper is the documented review **set** (`mergeOrchestration ∪ mergeBlocker ∪ humanReview`), so a board splitting those across a merge lane and a human lane is covered without this site drifting from it. ## Two things my first attempt got wrong **The fixture spelled traits in camelCase** — `{ trait: "humanReview" }`. Trait **ids** are kebab-case (`human-review`, `merge-blocker`, `wip`); the camelCase names are the resolved **flags**. Those columns therefore resolved to *no roles at all*, silently, because an unknown trait isn't an error. `complete` is spelled identically in both vocabularies, which is exactly what made the first run look like *"complete works, review is broken"* rather than *"the fixture is wrong"* — I nearly went debugging the production union. **The harness extension is additive** and inert until a test seeds it, so all 24 existing plugin suites see the previous shape. ## Measured | check | result | |---|---| | new suite | **4/4** | | reverting to the literal-only gate | fails **exactly 2** — the renamed-terminal case, and a board declaring a NON-terminal column named `done` — while the legacy control and the WIP/intake negative still pass | | plugin suite | **24 files, 184 tests green** | | `tsc` + all five gates | clean | That second row is the one that matters: the `done`-without-`complete` board is the only shape where a real resolution and a legacy fallback disagree, so it's what separates the fix from a lucky agreement. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5ae6332563 |
refactor: collapse dead SQLite dual-path code; keep migration-only readers (#2454)
# Remove dead SQLite dual-path code; keep migration-only readers ## Summary PostgreSQL cutover left hundreds of production dual-path branches (`backendMode ? PG : SQLite/store.db`) whose SQLite arms only hit throwing `Database`/`ArchiveDatabase`/`CentralDatabase` stubs. This change mechanically collapses those unreachable arms so production authority is AsyncDataLayer/PostgreSQL only, while preserving the six authorized read-only migration/recovery `DatabaseSync` seams. ## Dual-path mass removed | Metric | Before | After | |---|---|---| | `if (…backendMode)` (non-test) | ~328 | ~70 | | `store.db` / `this.db` refs in core (non-test) | ~570+ | ~375 (mostly pure legacy MissionStore/eval/insight SQLite classes + thin getters) | | Net diff | — | **~6.7k lines removed** across 41 files | Remaining `backendMode` checks are intentional (incomplete-PG sync safe-defaults, settings-sync disabled-on-PG, symbol-lock PG-only gates, “requires PostgreSQL” config versioning throws), not live SQLite authority. ## Subsystems cleaned - **Core TaskStore / task-store/***: collapsed if/else and early-return dual-path across reads, moves, lifecycle, mutations, workflow, archive, branch/PR, artifacts, comments, audit, project ops, etc. `initImpl` is PostgreSQL-only (SQLite startup tail deleted). - **Satellite stores**: automation, agent, routine, plugin, secrets, approval-request, central-core dual-path arms collapsed. - **Plugins**: reports async methods, compound-engineering pipeline + session stores, CLI Printing Press store — SQLite fallbacks removed; PG required. - **Engine**: no functional dual-path change beyond whitespace (settings-sync / peer-exchange PG-disabled behavior kept). ## Six migration-only readers retained (allowlist unchanged) 1. `packages/core/src/postgres/sqlite-migrator.ts` 2. `packages/core/src/project-identity.ts` 3. `packages/core/src/sqlite-validation.ts` 4. `packages/core/src/postgres/startup-factory.ts` 5. `packages/cli/src/commands/db.ts` 6. `scripts/lib/start-local-project.mjs` Plus low-level `sqlite-adapter` and migrator/startup-import tests. Inventory ratchet still requires exactly these six `new DatabaseSync(` production sites, all `readOnly: true`. ## Not treated as SQLite - `.fusion/project.json`, `task.json`, `agent-log.jsonl` file storage - AsyncDataLayer / Drizzle PG paths - Incomplete-PG sync safe-default stubs (still return empty/false/null under backend without consulting SQLite) ## Verification - `sqlite-production-reader-inventory.test.ts` — 15/15 pass - `incomplete-pg-ports.pg.test.ts` — 6/6 pass - Targeted PG tests (create-task, move, handoff, runtime-persistence, agent, mission, insight, central-core) — green - `tsc --noEmit` for `@fusion/core`, `@fusion/engine`, `@fusion/dashboard` — green - `scripts/check-no-getdatabase.mjs` — clean <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Improved end-to-end consistency by making PostgreSQL/async persistence the standard across core task/workflow, automation, agents, plugins, routines, secrets, approvals, central operations, and session storage. * Unified scheduling, settings, configuration revision writes, run/workflow selection, queues/leases/transitions, and audit/lifecycle updates around consistent async transaction behavior. * **Bug Fixes** * Fixed edge cases for archived/deleted reads, unarchive/recovery flows, not-found handling, and task/artifact/document/log/comment operations, including more reliable emissions and hydration across search/list and lifecycle operations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
a242f1b449 |
fix(FN-7952): migrate bundled plugins to PostgreSQL (#2111)
## Summary Bundled plugins now persist shared runtime state in project-scoped PostgreSQL tables instead of maintaining independent SQLite authority. Reports, CLI Printing Press, Compound Engineering, Roadmap, Even Realities, and WhatsApp all follow the same ownership and startup contract as Fusion core. ## Design decisions - Plugin schema hooks run through the host’s PostgreSQL owner and enforce project isolation. - The SDK exposes the host contract needed by bundled plugins without importing engine internals. - Legacy Roadmap ownership fixtures use the supported empty-owner sentinel, preserving current composite primary/foreign keys while exercising backfill behavior. - The lockfile travels with the Even Realities PostgreSQL dependency so packaged installs remain reproducible. ## Validation - All six affected plugin builds pass. - Affected plugin suites pass: 773 tests across Printing Press, Compound Engineering, Even Realities, Reports, Roadmap, and WhatsApp. - `pnpm test:gate` passes all 478 gate tests. - This PR changes 40 files. ## Stack - Depends on #2110 → #2109 → #2108. - The documentation/release PR completes the stack. Related: #2105 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Breaking Changes** * PostgreSQL is now required for runtime storage; SQLite files are used only as one-time migration inputs. * The legacy `FUSION_NO_EMBEDDED_PG` fallback has been removed. * **New Features** * Added project-isolated PostgreSQL storage for plugins, reports, tasks, notifications, and other plugin data. * Added agent tools for reports and CLI service drafts. * Added PostgreSQL schema initialization support for plugin authors. * **Bug Fixes** * Improved migration and recovery of legacy plugin state. * Prevented cross-project data access and strengthened transactional schema updates. * **Documentation** * Updated storage, migration, deployment, plugin authoring, CLI, and dashboard guidance for PostgreSQL. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
c15c78feeb |
feat: migrate storage from SQLite to PostgreSQL (#1793)
# Migrate storage from SQLite to PostgreSQL — full dashboard cutover Migrates Fusion's storage layer to the embedded PostgreSQL `AsyncDataLayer` (the default backend) and **completes the satellite-store + feature cutover** so every dashboard and Command Center surface works in PG mode. ## Status — every surface works in embedded-PG mode Verified live against a running embedded-Postgres dashboard (all **200**, zero 5xx) and gate-tested (**23 files / 99 tests** on embedded PG, plus engine-core 294 and ci-shape 63 in the blocking merge gate; core/engine/cli/dashboard typecheck clean). | Area | Surfaces | State | |---|---|---| | Satellite stores | workflows, todos, insights, research, missions, goals, mailbox | ✅ | | Views | artifacts, documents, evals | ✅ | | Command Center | activity, productivity, team, tokens, tools, **workflows**, **github**, **signals**, **plugin-activations**, **live** (all 10) | ✅ | | Run execution | insight generation, research run execution | ✅ (store-path; AI step needs a provider) | | Live updates | SSE push for mission/research/insight events | ✅ | | Workflow editing | create / update / delete / select (+ id counter) | ✅ | | Engine | mission autopilot, incident-signal ingestion, regression storm-guard, agent wake-on-message | ✅ | | Core | tasks, agents, secrets, automations, memory, chat, usage, PRs, git | ✅ | ## Approach Each satellite store gets an `Async<Store>` wrapper exposing the sync store's method names over the existing `async-*-store.ts` helpers; `get<Store>Store()` returns a `Sync | Async` union; consumers `await` (harmless on sync), and engine/CLI paths that can't convert use `instanceof Sync` graceful fallback. Analytics aggregators branch on `"ping" in dbOrLayer` to run schema-qualified raw SQL over `project.*` (snake_case) in PG. Executors/orchestrators/autopilot are await-converted to drive the union store; the async store wrappers extend `EventEmitter` so SSE live-push fires in both backends. Not-yet-ported capabilities degrade gracefully (never 500) and are individually called out in commits. ## Sync with main The branch is kept continuously merged with `main` (currently through FN-7845, 2026-07-12); the earlier "final rebase deferred" note no longer applies. Use **Create a merge commit** (or squash) to land it — GitHub's rebase-merge cannot replay a merge-maintained branch. ## Residual Review Findings Multi-agent code review of the PostgreSQL satellite-store ports (U1–U5) applied 3 safe fixes (see `fix(review): apply autofix feedback`). The following are **real but gated** — recorded here as follow-up work rather than auto-applied. All are SQLite→PostgreSQL **concurrency/atomicity regressions**: the sync stores were immune only by SQLite's single-writer, single-threaded-handler execution; the async ports open multi-await read-modify-write windows. **Reachability is low today** because the execution engines that generate concurrent same-run mutations (insight run executor, research orchestrator/dispatcher) are `instanceof`-gated to sync mode in PG. No process-crash class survived (all engine fallbacks correctly guard the sync store). - **[P1] Research `appendResearchEvent` dual-write is non-atomic** (`packages/core/src/async-research-store.ts`, corroborated: adversarial + reliability). The `research_run_events` insert (own transaction) and the `run.events` jsonb update are separate writes — a crash between them, or two concurrent appends, splits the table count from the jsonb array. **Fix:** perform the seq-insert and the jsonb update in one `layer.transactionImmediate`. - **[P1] Research run terminal-reversion via stale full-row persist** (`async-research-store.ts` `persistResearchRun`/`updateResearchStatus`). Concurrent `PATCH /runs/:id/status` + `POST /runs/:id/events` can revert a terminal run to `running` by overwriting the whole row, bypassing the transition guard. **Fix:** scoped column `UPDATE`s with a `WHERE status …` guard, or optimistic version column. - **[P2] `updateResearchRun`/`updateInsightRun` read-then-write TOCTOU** — concurrent PATCHes last-writer-wins on the lifecycle merge. **Fix:** `SELECT … FOR UPDATE` / enclosing transaction. - **[P2] `upsertRun`/`createRunOrThrowConflict` check-then-create race** (`async-insight-store.ts`) — two callers can each create an "active" run. **Fix:** partial unique index on `(projectId, trigger) WHERE status IN ('pending','running')`. - **[P3] `createResearchRetryRun` return-value divergence** — sync returns the pre-update `queued` snapshot; async returns the reloaded `retry_waiting` run (persisted state is identical). Pick one side for cross-backend parity. - **[P2/perf] Mission `getMissionWithHierarchy`/`getMissionHealth` N+1 fan-out** — O(milestones×slices) sequential round-trips hold one pool slot per request; can starve the pool for large hierarchies. **Fix:** batched/joined reads. - **Testing gaps:** no PG-mode concurrency tests (interleaved status/event mutations), no sync↔async parity assertion for the lifecycle-error codes, and no mission status/health rollup parity test vs the sync `MissionStore`. ~~Out of scope (deferred): AI run *execution* (insight/research) + mission autopilot + live SSE mission events remain sync-gated/degraded in PG mode.~~ **Since ported** — insight/research run execution, mission autopilot, and SSE live push all run on the async layer now, which also makes the concurrency findings above genuinely reachable; they remain open follow-ups. --- ## Update — 2026-07-12: production-readiness hardening & live acceptance Everything below landed on this branch since the description above was written: **Production blockers from review — fixed** - `recoverStaleTransitionPending` ported to the async layer (backend moves write + clear the crash-safe marker; startup/maintenance sweeps no longer throw). - Lost-update class fixed: `atomicWriteTaskJson`/`WithAudit` write changed columns only (full-row upserts silently resurrected stale fields across concurrent store instances — the "task stuck unplanned forever" bug). - First-boot **auto-migration**: booting the PG backend over a project with a legacy `fusion.db` migrates it automatically (loud failure, SQLite kept as backup), and the dashboard shows a one-time **"your data was migrated" banner** with the backup paths and a Need-help Discord link. - `pg_dump`/`pg_restore` discovered from common install locations for embedded-mode backups. - The PG suite is part of the blocking merge gate (`test:pg-gate`). **Multi-project isolation (PR #2007, merged into this branch)** - `project_id` partition key on tasks / archived tasks / config, `taskProjectScope` threaded through every scan/claim/count, per-project config rows, layer bound to the project at startup. - Review P1 follow-up: the shared cold-storage `archive.archived_tasks` table is also partitioned and all archived-board reads/counts/searches are scoped. - Schema drift self-heal generalized to schema-qualified columns so existing databases upgrade in place. **Other changes** - Node settings sync **removed** in PG mode (409 `settings-sync-disabled-postgres`) — nodes share state by connecting to the same database; auth sync kept (per-machine file). - Perf (review findings): `listTasks` pushes column filter + ORDER BY + LIMIT/OFFSET into SQL; `getConversation` capped to the most recent 200 messages. - Fixed a false "operator action required" pause-abort log fired on every successfully auto-merged task. **Live acceptance — PASSED (2026-07-12)** A sandboxed instance (isolated HOME, embedded PG, real Opus executor) ran a task through the complete cycle: create → triage (AI spec) → execute → in-review → AI squash-merge landed on the project's `main` → done. A write+read sweep of every data surface (settings, comments, documents, attachments + artifact bridge + artifact edit, chat with real generation, goals, missions, agent mail, secrets, workflows, memory, CC analytics) was green on embedded PG. **Known remaining work** - The per-project `config` PK re-key has no upgrade path for pre-isolation embedded-PG databases (needs a real `DROP CONSTRAINT`/re-key migration; fresh databases are fine). - `pg_dump`/`pg_restore` binaries are not yet bundled in release artifacts (PATH/common-location discovery only). - The satellite-store concurrency findings listed above. --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Phil Larson <hello@phillarson.xyz> Co-authored-by: fusion-merge <fusion-merge@local> |
||
|
|
f66bfae6eb | feat(compound-engineering): clarify the pipeline flow | ||
|
|
b69dd8eb14 |
FN-7103: alias brainstorm artifacts to unified plans
Compound Engineering now treats brainstorm output as the requirements-only form of unified plan artifacts. - Point the brainstorm stage at docs/plans while keeping the existing stage and skill identifiers for compatibility. - Surface CE plan frontmatter metadata during artifact discovery and reading. - Harden CE artifact discovery against symlink escapes for conventional artifact paths. - Update docs, tests, and changeset coverage for the unified brainstorm/plan artifact flow. Files changed: .changeset/fn-7103-ce-plan-alias.md | 7 ++ docs/plugins/compound-engineering.md | 21 +++- .../fusion-plugin-compound-engineering/README.md | 36 ++++-- .../orchestrator-interrupt-resume.test.ts | 8 +- .../src/__tests__/session-store.test.ts | 38 ++++++ .../src/__tests__/stage-registry.test.ts | 18 +++ .../src/__tests__/sync.test.ts | 36 +++++- .../src/artifacts/__tests__/discovery.test.ts | 133 +++++++++++++++++++- .../src/artifacts/discovery.ts | 136 +++++++++++++++++++-- .../src/session/stage-registry.ts | 10 +- .../src/sync/reconciler.ts | 2 +- 11 files changed, 413 insertions(+), 32 deletions(-) Fusion-Task-Id: FN-7103 Fusion-Task-Lineage: 284a5c5c-40f7-44b2-8e3e-1de671dde36a Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
08312bb59d |
fix(compound-engineering): correctness fixes from code review
- reconciler: a deleted current-stage board task no longer wedges the pipeline in 'running' forever (terminality computed over existing tasks only; all-deleted is a no-op, not a wedge) - session-store: a human-slow awaiting_input session is no longer misclassified stale (interval rubric applies only to in-flight active/launching turns) - stage-registry: pipeline progression uses an explicit order ordinal instead of registry insertion order, so out-of-order registration can't corrupt advancement - orchestrator.answer(): validate questionId before mutating state, so a stale id can't destroy the persisted currentQuestion recovery anchor - orchestrator.resume(): rehydrate a live interactive session by replaying persisted history (side effects suppressed) so a resumed session is actually answerable instead of dead-ending; honest interrupted+error fallback when no factory is available 95 tests (6 new regression tests, each confirmed failing pre-fix). |
||
|
|
061297b432 |
refactor(compound-engineering): consolidate duplication from simplify pass
Quality cleanup across the 9-unit build (behavior-preserving, 89 tests green): - extract createCeTaskWithLink so the work bridge and reconciler share one provenance+link contract (prevents drift) - discovery list scan probes readability via accessSync instead of reading and discarding full file bytes - makeError helper replaces ~7 duplicated CeArtifactError literals - shared asString route helper; drop dead pipelineIds set; collapse a double pipeline-state write and a redundant link re-query in advance - resolveStageSkillCwd no longer takes params it ignores |
||
|
|
4dd7509f53 |
feat(compound-engineering): CE-pipeline state model + bidirectional sync (U8)
Add a ce_pipeline_state machine (currentStage/status) kept distinct from board-task ownership (task column) — separate tables, no shared column, per FN-5719. onTaskMoved/onTaskCompleted hooks do only an indexed lookup + enqueue and return well under the 5s budget; advancement happens in an on-demand reconciler sweep that re-derives correct pipeline state from board truth, so a dropped hook event still converges (no tight poll loop). Outbound CE-flow changes create the next-stage board task. Conflict policy: board authoritative for task state, CE flow authoritative for artifact/pipeline content (inbound read and outbound write target different rows, so they cannot contend). Host-scheduler note: no host timer wired; sweeps run on hook-drain and on the dashboard/route refresh surface (tracked for the host event-publish follow-up). |
||
|
|
b71e531735 |
feat(compound-engineering): work->board bridge with traceable tasks (U7)
When the work stage completes with a derived task list, create Fusion tasks via ctx.taskStore.createTask tagged CE-originated (sourceType workflow_step + sourceMetadata marker) and record an authoritative ce_pipeline_links row per task (back-reference lives in the link table, not task-row JSON, per FN-5719). Tasks then run the normal lifecycle untouched. Zero derived tasks is a clean no-op. The link store is intentionally minimal for U8 to extend with the bidirectional pipeline-state machine. |