a3dda8eaffc8cc33f21b6c1927fbbd3089978975
67 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a3dda8eaff |
fix(core): prevent concurrent startup database failures (#2330)
## Summary Concurrent PostgreSQL project initialization no longer causes transient dashboard failures, including repeated `GET /api/remote/status` 500 responses. The failure was a database deadlock between project-row identity promotion and schema/plugin DDL, which previously acquired overlapping locks in inconsistent orders. This establishes one advisory-lock order across SQLite cutover, project identity promotion, and schema mutations. Focused regression coverage proves schema DDL waits behind an active migration transaction and that identity stamping acquires the migration lock before reading project-owned tables. ## Validation - 25 focused unit tests passed. - 3 focused real-PostgreSQL regression tests passed. - `@fusion/core` typecheck passed. - Strict changeset validation passed. - Fast workspace verification passed, including the CLI build and boot health check. --- [](https://github.com/EveryInc/compound-engineering-plugin) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Prevented transient dashboard failures caused by PostgreSQL startup and migration deadlocks. * Improved serialization when multiple projects initialize or update database schemas concurrently. * Ensured migration state updates and schema changes occur in a consistent order. * **Tests** * Added coverage for migration lock ordering, concurrent schema operations, and recovery after lock contention. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
c95e08ea21 |
FN-8297: add research finding promotion to mission features
Bridge completed research findings into durable mission roadmap features. - Persist stable finding and citation provenance on mission features - Add idempotent promotion APIs, dashboard controls, and agent tooling - Document the promotion flow and cover finding identity and feature synchronization Files changed: .changeset/fn-8297-research-mission-bridge.md | 7 ++++ docs/missions.md | 4 +++ docs/research.md | 4 +++ .../__tests__/research-finding-identity.test.ts | 15 +++++++++ packages/core/src/async-mission-store-queries.ts | 26 +++++++++++++++ packages/core/src/async-mission-store.ts | 16 +++++++++ packages/core/src/index.gate.ts | 3 ++ packages/core/src/index.ts | 3 ++ packages/core/src/mission-types.ts | 13 ++++++++ .../core/src/postgres/migrations/0000_initial.sql | 4 +++ .../0023_research_feature_provenance.sql | 8 +++++ packages/core/src/postgres/schema-applier.ts | 13 +++++++- packages/core/src/postgres/schema/project.ts | 5 +++ packages/core/src/research-feature-promotion.ts | 38 ++++++++++++++++++++++ packages/core/src/research-types.ts | 21 ++++++++++++ packages/core/src/types.ts | 1 + packages/dashboard/app/api/legacy.ts | 21 ++++++++++++ packages/dashboard/app/components/ResearchView.tsx | 19 ++++++----- packages/dashboard/app/hooks/useResearch.ts | 3 ++ packages/dashboard/src/chat.ts | 4 +-- packages/dashboard/src/research-routes.ts | 37 +++++++++++++++++---- .../src/__tests__/agent-mission-tools.test.ts | 14 +++++++- .../src/__tests__/mission-feature-sync.test.ts | 20 ++++++++++++ packages/engine/src/agent-tools.ts | 26 +++++++++++++++ packages/engine/src/mission-feature-sync.ts | 1 + 25 files changed, 308 insertions(+), 18 deletions(-) Fusion-Task-Id: FN-8297 Fusion-Task-Lineage: ffef4d26-1372-466a-8d3d-2b9b5cc9a53b Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
7b1a89d1bc |
fix(core): join running embedded Postgres correctly
Read the live port from PostgreSQL's actual postmaster.pid field so extension TaskStore boot reuses the existing server instead of wedging on a colliding start. |
||
|
|
d51ce46db5 |
FN-8295: add persisted ideation mission handoffs
Persist bounded ideation sessions through agent tools, the Command Center, and atomic Mission convergence. - Store ideation sessions and divergent candidates in PostgreSQL with async APIs and migration support. - Expose gated ideation tools, chat routes, and agent lifecycle integration. - Add the Ideation panel, documentation, release metadata, and regression coverage. Files changed: .changeset/fn-8295-ideation-diverge-converge.md | 7 ++ docs/ideation/persisted-diverge-converge.md | 22 ++++ docs/missions.md | 4 + .../__tests__/postgres/ideation-store.pg.test.ts | 57 +++++++++ packages/core/src/async-ideation-store-queries.ts | 117 +++++++++++++++++ packages/core/src/async-ideation-store.ts | 138 +++++++++++++++++++++ packages/core/src/async-mission-store.ts | 27 ++-- packages/core/src/ideation-types.ts | 69 +++++++++++ packages/core/src/index.ts | 3 + .../core/src/postgres/migrations/0022_ideation.sql | 67 ++++++++++ packages/core/src/postgres/schema-applier.ts | 18 ++- packages/core/src/postgres/schema/project.ts | 49 +++++++- packages/core/src/store.ts | 8 +- packages/core/src/task-store/remaining-ops-8.ts | 14 +++ .../components/command-center/CommandCenter.tsx | 7 +- .../components/command-center/IdeationPanel.css | 18 +++ .../components/command-center/IdeationPanel.tsx | 58 +++++++++ .../__tests__/CommandCenter.test.tsx | 6 +- .../dashboard/src/__tests__/chat-manager.test.ts | 1 + packages/dashboard/src/__tests__/chat.test.ts | 1 + .../__tests__/ideation-tool-route-parity.test.ts | 29 +++++ packages/dashboard/src/chat.ts | 4 + packages/dashboard/src/ideation-routes.ts | 50 ++++++++ .../src/routes/register-integrated-routers.ts | 2 + .../src/__tests__/agent-ideation-tools.test.ts | 40 ++++++ .../src/__tests__/gating-classifications.test.ts | 16 +++ .../src/__tests__/permanent-agent-gating.test.ts | 2 + packages/engine/src/agent-heartbeat.ts | 4 +- packages/engine/src/agent-tools.ts | 67 ++++++++++ packages/engine/src/executor.ts | 2 + packages/engine/src/gating-classifications.ts | 8 ++ packages/engine/src/index.ts | 1 + packages/engine/src/triage.ts | 2 + 33 files changed, 897 insertions(+), 21 deletions(-) Fusion-Task-Id: FN-8295 Fusion-Task-Lineage: 1b8b0752-22bd-4b2f-aebd-4305c63abcf9 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
bcb1256d6b |
FN-8282: add configuration revision history and rollback
Record configuration revisions so prior settings can be restored safely. - Add revision storage interfaces, PostgreSQL schema, migration, and pruning. - Capture changes to global settings, routines, automations, and async task settings. - Expose revision listing and restore operations with coverage and storage documentation. Files changed: .changeset/fn-8282-config-versioning.md | 7 + docs/storage.md | 6 + .../__tests__/configuration-revision-store.test.ts | 26 +++ .../src/__tests__/postgres/schema-applier.test.ts | 15 +- .../core/src/async-configuration-revision-store.ts | 192 +++++++++++++++++++++ packages/core/src/automation-store.ts | 88 +++++++++- packages/core/src/configuration-revision-store.ts | 33 ++++ packages/core/src/global-settings.ts | 172 +++++++++++++++++- packages/core/src/index.gate.ts | 3 + packages/core/src/index.ts | 3 + .../migrations/0021_configuration_revisions.sql | 40 +++++ packages/core/src/postgres/schema-applier.ts | 19 +- packages/core/src/postgres/schema/project.ts | 27 +++ packages/core/src/routine-store.ts | 103 ++++++++++- packages/core/src/store.ts | 21 ++- packages/core/src/task-store/async-settings.ts | 14 +- packages/core/src/task-store/remaining-ops-2.ts | 86 ++++++++- packages/core/src/task-store/settings-ops.ts | 104 +++++++---- packages/core/src/types.ts | 36 ++++ packages/engine/src/agent-tools.ts | 7 +- 20 files changed, 928 insertions(+), 74 deletions(-) Fusion-Task-Id: FN-8282 Fusion-Task-Lineage: fd681798-10b0-47ce-b96b-6f33eecdf70a Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
2c17fa70ab | feat(postgres): make embedded connection cap configurable | ||
|
|
7c23771433 |
FN-8265: add task follow-up proposal creation
Enable configured ephemeral workers to propose and create follow-up tasks from mailbox messages. - Add persisted task-proposal claim state, migrations, and async messaging APIs. - Register task-proposal creation routes, SSE events, agent tool support, and CLI integration. - Add mailbox creation controls, settings, documentation, localization, and regression coverage. Files changed: .changeset/fn-8265-task-follow-up-policy.md | 7 ++ docs/dashboard-guide.md | 2 +- docs/settings-reference.md | 12 ++- packages/cli/src/extension.ts | 34 ++++--- .../src/__tests__/postgres/sqlite-migrator.test.ts | 14 ++- .../postgres/task-proposal-claim.pg.test.ts | 51 +++++++++++ packages/core/src/async-message-store.ts | 39 ++++++++ packages/core/src/index.gate.ts | 4 +- packages/core/src/index.ts | 4 +- packages/core/src/message-store.ts | 46 ++++++++++ .../core/src/postgres/migrations/0000_initial.sql | 2 + .../migrations/0020_task_proposal_claim.sql | 4 + packages/core/src/postgres/schema-applier.ts | 13 ++- packages/core/src/postgres/schema/project.ts | 3 + packages/core/src/settings-schema.ts | 19 +++- packages/core/src/task-store/async-persistence.ts | 2 +- packages/core/src/task-store/persistence.ts | 4 +- packages/core/src/task-store/serialization.ts | 1 + packages/core/src/task-store/task-creation.ts | 51 +++++++++++ packages/core/src/task-store/task-row-mappers.ts | 2 +- packages/core/src/types.ts | 56 +++++++++++- packages/dashboard/app/api/legacy.ts | 5 + packages/dashboard/app/components/MailboxModal.tsx | 4 + .../app/components/MailboxTaskProposal.css | 3 + .../app/components/MailboxTaskProposal.tsx | 33 +++++++ packages/dashboard/app/components/MailboxView.tsx | 4 + .../__tests__/MailboxTaskProposal.test.tsx | 43 +++++++++ .../app/components/settings/section-keys.ts | 2 +- .../settings/sections/GeneralSection.search.ts | 13 ++- .../settings/sections/GeneralSection.tsx | 22 +++-- .../settings-default-descriptions.test.tsx | 4 +- .../routes/__tests__/task-proposal-routes.test.ts | 99 ++++++++++++++++++++ .../src/routes/register-messaging-scripts.ts | 101 +++++++++++++++++++++ packages/dashboard/src/sse.ts | 7 ++ packages/engine/src/agent-tools.ts | 30 ++++-- packages/engine/src/executor.ts | 9 +- packages/engine/src/step-session-executor.ts | 8 +- packages/i18n/locales/en/app.json | 5 + 38 files changed, 696 insertions(+), 66 deletions(-) Fusion-Task-Id: FN-8265 Fusion-Task-Lineage: 4e864a2f-3485-4a54-8be7-1699b5479a94 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
68ce5c31da |
fix: harden onboarding/git/desktop batch per multi-agent review findings
- uninstaller: taskkill only the first, digits-only postmaster.pid line (the for /f loop ran taskkill on the port/epoch lines — potential unrelated-process kill) - git-missing dialogs use new ConfirmOptions.alwaysAsk so global skip-confirmations cannot silently pick an unseen choice - Windows quit prompt: embedded-local runtimes only, skipped during OS session end (sync dialog blocked Windows shutdown) - 'leave it running' detaches the embedded lifecycle (disarms its process shutdown hook) so Electron exit cannot kill the postmaster the operator chose to keep (new detachKeepingEmbedded) - wizard: ref-based double-submit guard around the async git preflight - clone route: ENOENT invalidate-and-retry matching runGitCommand - openExternalUrl: drop the async window.open fallback (always popup-blocked); log bridge failures instead - DirectoryPicker: close the panel when listing the created folder fails so Select cannot re-commit the parent - git status probe bounded to two spawns (PATH + first candidate) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
95e011f890 |
fix(core): auto-repair empty non-UTF-8 embedded Postgres clusters on boot (#2286)
Users whose embedded cluster was initdb'd with an OS-locale encoding by a pre-fix version now self-heal with zero manual steps: on the encoding-conversion schema failure the startup factory proves the cluster is non-UTF-8 AND empty (the baseline transaction never applied, so no schema or migrated data can exist) and that this process owns the postmaster, then deletes the data dir and reboots once with the UTF-8 initdb defaults. Joined instances and unproven states keep the manual re-init hint; one retry ever, so no loops. Verified on the elevated windows-latest runner: CI seeds a real WIN1252 cluster via initdb and proves a stock 'fn serve' auto-recovers it to a healthy /api/health (run 29633351848, all jobs green). Also caps the desktop-windows embedded-PG smoke at 30 min and adds a skip input. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d4ee80a818 |
fix(core): Windows embedded Postgres — no local user account, UTF-8 clusters, diagnosable boot errors
Squash of feature/win-elevated-no-user, verified end-to-end on the elevated windows-latest runner (restricted-token double boot + full 'fn serve' /api/health smoke, both green). - Elevated Windows boots embedded PostgreSQL via pg_ctl's built-in restricted-token re-exec instead of creating a 'fusion-pg' local user (operator requirement: Fusion must never create accounts). Removes the credential launcher, icacls grants, and cmd/PowerShell wrapper — and with them the 'directory name is invalid' and wrapper-log EBUSY field failures. Leftover fusion-pg accounts are deleted on start. - Embedded clusters are always initdb'd --encoding=UTF8 --locale=C (GitHub issue #2286: OS-locale WIN1252/WIN1254 clusters could not store the UTF-8 schema and crash-looped the dashboard). Existing non-UTF-8 clusters get an actionable re-init hint at boot. - Schema-backend boot failures now surface the full error cause chain (DrizzleQueryError hid the real PostgresError behind the SQL text). - Elevated stop() waits until the port closes and postmaster.pid is gone before resolving. - CI: branch verification workflow (restricted-token proof + elevated boot smoke + account-absence assertions); boot-smoke stderr tail widened for diagnosability. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fd87c3f23f |
fix(core): pin non-admin postgres launcher working directory on elevated Windows
Start-Process -Credential (CreateProcessWithLogonW) validates the working directory as the TARGET user. The launcher inherited the desktop app's cwd (admin profile / install dir), which the dedicated fusion-pg user cannot read, so elevated desktop boots died with "The directory name is invalid" before postgres ever started. launch.ps1 now pins -WorkingDirectory to the .pgrunner run dir inside the data dir the user was just granted full control on. CI never caught it because runner cwds are world-traversable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0b6c4cd4ca |
feat(dashboard,desktop): show live database-migration progress during boot
The one-time SQLite→PostgreSQL migration runs inside createTaskStoreForBackend before any HTTP server listens, so browsers saw "connection refused" and open tabs failed silently for minutes. Now: - CLI: a temporary holding server binds the dashboard port for the boot window, serving an auto-reloading "Database migration in progress" page and an /api/health payload with status "migrating" + structured progress; the port is handed off (awaited) to the real app.listen(). - Dashboard SPA: already-open tabs render the new MigrationInProgressBanner from the 15s health poll when status is "migrating". - Desktop: LocalRuntimeManager publishes migration progress on DesktopRuntimeStatus via the new core onMigrationProgress option; DesktopLaunchGate shows the live label and extends its 30s startup timeout while progress advances (2min stall cap), in both boot and first-run flows. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
48b0d04322 |
fix(core): sanitize NUL (u0000) characters in SQLite-to-PostgreSQL migration
Legacy SQLite databases can hold U+0000 in TEXT cells and inside stored JSON, which PostgreSQL rejects in text and jsonb columns and which aborted the first-boot auto-migration. Strip NUL from plain text cells, JSON string values and object keys, malformed-JSON scalars, and opaque legacy-preservation cells; content-checksum verification compares the sanitized source against the sanitized target so migrations still verify. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8517a5d3ba |
fix(core): make embedded Postgres shared_memory_type default platform-aware
shared_memory_type=mmap (defaulted 2026-07-16 for SysV shm exhaustion) is invalid on Windows — PostgreSQL only accepts "windows" there and dies with FATAL invalid value for parameter before opening the port. Every Windows embedded start broke, failing the Windows release smoke in both the v0.70.0 and v0.70.1 tag runs. Default flags now come from defaultEmbeddedPostgresFlagsFor(platform): empty on win32 (no override needed; SysV exhaustion cannot occur there), mmap elsewhere. Regression test asserts the per-platform flag invariant. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6b893f78ec |
fix(cli): make the standalone fn binary boot PostgreSQL in both modes
The bun-compiled exe has been unbootable since the PG cutover: bun
standalone binaries do no node_modules resolution, so the deliberately
out-of-graph require("embedded-postgres") failed from /$bunfs, and
readFile'd migration .sql files were never embedded, so even external
DATABASE_URL mode died at schema init.
- schema-applier: resolveMigrationsDir() — FUSION_MIGRATIONS_DIR env >
module-relative dist/migrations (npm/desktop, unchanged) >
execPath-relative migrations/ (standalone exe), probe-based.
- embedded-lifecycle: require("embedded-postgres") first (npm/desktop
untouched), falling back to a self-contained staged bundle at
<execDir>/runtime/<platform>/embedded-postgres/dist/index.cjs
(FUSION_EMBEDDED_PG_RUNTIME_DIR override) with the native
initdb/pg_ctl/postgres payload beside it.
- build.ts: stage dist/migrations plus the per-target embedded-postgres
bundle + native payload (warn when a cross-target payload is absent on
the host, mirroring desktop's verifyEmbeddedPostgresPayloads).
- release.yml: package fn-cli-<os>-<arch>.tar.gz (binary + migrations +
runtime + client) with sha256 per leg; prune staged payload files from
the release-collection globs; bare fn-cli-* binaries still uploaded.
E2E-verified on the compiled binary: embedded mode initdb→/api/health
200 database healthy; DATABASE_URL mode applied migrations 0000–0019
(109 tables). Core typecheck clean; schema-applier 58/58 and
embedded-lifecycle 44/44 tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d8735b3dbe |
FN-8249: persist GitHub translations and show status
Persist legacy GitHub translation cache entries and expose background translation progress to import operators. - Backfill historic unscoped translation cache partitions during schema migration. - Display accessible translating and failure status in the GitHub issues import list. - Add migration, service, and UI coverage plus operator documentation. Files changed: .../fn-8249-github-import-translation-status.md | 7 ++ docs/dashboard-guide.md | 2 +- .../postgres/import-translation-cache.pg.test.ts | 13 +++- .../src/__tests__/postgres/schema-applier.test.ts | 48 ++++++++++++ ...translation_cache_legacy_partition_backfill.sql | 31 ++++++++ packages/core/src/postgres/schema-applier.ts | 29 ++++++- .../dashboard/app/components/GitHubImportModal.css | 42 ++++++++++ .../dashboard/app/components/GitHubImportModal.tsx | 29 +++++++ .../__tests__/GitHubImportModal.test.tsx | 90 ++++++++++++++++++++++ .../src/__tests__/import-translate-service.test.ts | 62 +++++++++++++++ 10 files changed, 346 insertions(+), 7 deletions(-) Fusion-Task-Id: FN-8249 Fusion-Task-Lineage: 79c26d85-50f1-4d01-9341-a17db5e57f8f Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
d1e9b563f7 |
fix(FN-8141): add forward migration for tasks.bulk_completion_refusal_at
PR #2260 added project.tasks.bulk_completion_refusal_at to the Drizzle model and the 0000 baseline but shipped no forward migration. Databases created before #2260 already carry the 0000 marker, so the applier skips the baseline and they never gained the column — every such cluster crashed on the first TaskStore SELECT ("column bulk_completion_refusal_at does not exist"), taking down dashboard/app boot. Adds forward migration 0018 (wired via BULK_COMPLETION_REFUSAL_AT_VERSION; SCHEMA_BASELINE_VERSION -> "0018") so existing clusters heal on next startup. Prevention: - Per-column upgrade regression test reproducing the exact existing-DB failure. - Migration-wiring-integrity guard (no PostgreSQL): SCHEMA_BASELINE_VERSION must equal the highest migration file, and every .sql must be registered in the applier so none silently never runs. - Repairs 6 pre-existing schema-applier tests left stale by the 0017 addition (baseline-marker identity + version-list enumerations). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e72629c251 |
FN-8126: add per-task merger model controls
Enable Quick Add and task editing to select merger models and thinking levels. - Persist merger model and thinking overrides through task APIs, storage, and PostgreSQL migrations. - Add merger-lane selection controls to Quick Add and model settings interfaces. - Apply task merger settings to merger and PR fallback sessions, with regression coverage. - Document the merger lane and include a release changeset. Files changed: .changeset/fn-8126-quick-add-merger-lane.md | 7 ++ docs/dashboard-guide.md | 2 + docs/settings-reference.md | 7 +- .../core/src/__tests__/model-resolution.test.ts | 9 ++ packages/core/src/index.gate.ts | 1 + packages/core/src/index.ts | 1 + packages/core/src/model-resolution.ts | 20 ++++ .../core/src/postgres/migrations/0000_initial.sql | 3 + .../migrations/0017_task_merger_model_lane.sql | 4 + packages/core/src/postgres/schema-applier.ts | 14 ++- packages/core/src/postgres/schema/project.ts | 3 + packages/core/src/store.ts | 2 +- .../core/src/task-store/archive-lifecycle-2.ts | 6 ++ packages/core/src/task-store/persistence.ts | 6 ++ packages/core/src/task-store/remaining-ops-2.ts | 4 +- packages/core/src/task-store/remaining-ops-6.ts | 2 +- packages/core/src/task-store/serialization.ts | 6 ++ packages/core/src/task-store/task-creation.ts | 6 ++ packages/core/src/task-store/task-row-mappers.ts | 4 +- packages/core/src/task-store/task-update.ts | 6 ++ packages/core/src/types.ts | 18 ++++ packages/dashboard/app/api/tasks.ts | 13 +++ .../dashboard/app/components/InlineCreateCard.tsx | 33 ++++++- .../app/components/ModelSelectionModal.tsx | 29 ++++++ .../dashboard/app/components/ModelSelectorTab.tsx | 101 +++++++++++++++++++-- .../dashboard/app/components/QuickEntryBox.tsx | 44 +++++++-- .../__tests__/ModelSelectionModal.test.tsx | 20 ++++ .../components/__tests__/ModelSelectorTab.test.tsx | 37 +++++++- .../src/routes/register-task-workflow-routes.ts | 27 +++++- .../src/__tests__/agent-session-helpers.test.ts | 8 ++ packages/engine/src/agent-session-helpers.ts | 16 +++- packages/engine/src/merger-ai.ts | 14 +-- packages/engine/src/merger.ts | 35 ++++--- packages/engine/src/pr-response-run-ops.ts | 7 +- 34 files changed, 451 insertions(+), 64 deletions(-) Fusion-Task-Id: FN-8126 Fusion-Task-Lineage: 3fc81801-6d77-4e11-9cf0-3af37313930e Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
a136535f15 |
fix(engine): taint steps skipped after a bulk-completion refusal so they cannot auto-promote (#2260)
## What & why
**FN-8141 laundered a failed task into `done` with zero net changes and
no sign-off.** After the executor's
`bulk-step-completion-without-review` refusal fired (steps had no
APPROVE verdicts), the agent used the sanctioned skip affordance
(`fn_task_update status="skipped"`) on the remaining unreviewed steps.
Because every completion check counts `skipped` as complete, the task
then satisfied the exact condition the refusal was protecting, and
downstream **automatic** promotion (implicit `fn_task_done`,
self-healing `recoverStrandedCompletedTodoTasks`) moved it to in-review
— where the AI merger found an empty diff and finalized it as a no-op
`done`.
This PR restores the invariant: **steps skipped while a
bulk-step-completion refusal marker is active on the task are "tainted"
and cannot carry the task to review through any automatic path.** The
taint clears on an honest exit — an accepted `fn_task_done` (explicit or
non-tainted implicit) or an operator manual retry — so the legitimate
`PREMISE STALE` skip-then-done flow is unaffected.
## Design
- **Persisted marker**: new nullable `Task.bulkCompletionRefusalAt` (ISO
timestamp), stamped when the `bulk-step-completion-without-review`
refusal fires (explicit `fn_task_done` handler + implicit
`handleImplicitTaskDoneRefusal`). Survives requeue so a refusal on
attempt N taints attempt N+1's promotion. Full store plumbing (types,
descriptors, serialization, SQLite/PG schema + health self-heal).
- **Pure evaluator** `evaluateSkipBypassTaint(task)` in `@fusion/core`
(next to `evaluateNoCommitsNoOpFinalize`): `blocked` iff the marker is
set AND ≥1 step is `skipped`. Single rule every AUTO-promotion check
calls.
- **Clearing**: accepted explicit `fn_task_done`, accepted
implicit/retry completion (the success-reset `updateTask`s), and
`buildManualRetryResetPatch` (operator retry). A fresh lifecycle that
genuinely re-does the work leaves zero skipped steps, so it is never
blocked even if a marker lingers.
## Surface enumeration (every consumer of "all steps done/skipped" that
gates AUTO-promotion)
- **executor.ts**: `getCompletedTaskFinalizationDecision` (gated on the
`isTaskWorkComplete` branch only, never on an accepted `taskDone`);
`recoverCompletedTask` (shared chokepoint for unpause resume,
completed-task watchdog, orphan resume);
`evaluateImplicitCompletionRefusal` (both implicit-completion loops);
`isTaskAlreadyCompleteForNonContinuableSession`; graph merge-boundary
`getWorkflowMergeImplementationProofFailure`.
- **self-healing.ts**: `recoverCompletedTasks` (stuck in-progress) and
`recoverStrandedCompletedTodoTasks` (the exact FN-8141 promoter).
- **Verified-safe, left as-is**: per-step graph node projections
(executor ~6274/6298) and progress-render checks — they don't gate
whole-task auto-promotion.
## Test evidence
Scoped runs (all green):
```
CORE: pnpm --filter @fusion/core exec vitest run \
src/__tests__/skip-bypass-taint-guard.test.ts \
src/__tests__/skip-bypass-taint-persistence.test.ts \
src/__tests__/manual-retry-reset.test.ts
→ 17 passed
ENGINE: pnpm --filter @fusion/engine exec vitest run \
src/__tests__/executor-skip-bypass-taint.test.ts \
src/__tests__/self-healing.test.ts
→ 401 passed
```
Coverage: pure-evaluator (skip-before-refusal counts, skip-after-refusal
doesn't, taint-clearing, empty-marker/empty-steps edges); store
round-trip of the marker (set→read→clear); executor white-box (implicit
completion refused when tainted, allowed when clean or fully re-done,
graph merge-boundary reports missing proof, and the **explicit
`fn_task_done` PREMISE-STALE honest exit stays accepted**); self-healing
(FN-8141 sequence does not promote from either recovery path; a clean
legitimately-skipped task still promotes); manual-retry clears the
marker.
## Note on `pnpm verify:fast`
`verify:fast` currently fails at the workspace-artifact bootstrap on
**pre-existing** pi-SDK type errors in
`packages/engine/src/{auth-storage,pi,provider-registration}.ts` — the
FN-8145 upstream migration breakage (pi 0.80.x removed
`AuthStorage`/`ModelRegistry.create`). **None of those files are in this
diff.** `@fusion/core` builds clean (`packages/core build: Done`), and
`@fusion/engine` `tsc` reports **no errors in the files this PR
touches** (`executor.ts`, `self-healing.ts`); the only engine build
errors are the FN-8145 files. This base failure is the same condition
FN-8141 describes and is out of scope for this task.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus <noreply@anthropic.com>
|
||
|
|
c449379d00 |
FN-8172: persist import translations across restarts
Persist GitHub and GitLab import translation caches across application restarts. - Normalize cache ownership scope for reads, writes, pruning, and PostgreSQL RLS - Add forward migration 0016 to repair existing cache table partitioning - Cover durable cache reuse and migration behavior with PostgreSQL and service tests Files changed: .changeset/github-translation-cache-persistence.md | 7 ++ docs/settings-reference.md | 2 +- .../postgres/import-translation-cache.pg.test.ts | 105 +++++++++++++++++++++ .../src/__tests__/postgres/schema-applier.test.ts | 84 +++++++++++++++-- .../migrations/0010_import_translation_cache.sql | 6 +- .../0016_import_translation_cache_scope_fix.sql | 50 ++++++++++ packages/core/src/postgres/schema-applier.ts | 29 +++++- packages/core/src/postgres/schema/project.ts | 8 +- packages/core/src/task-store/remaining-ops-8.ts | 21 +++-- .../src/__tests__/import-translate-service.test.ts | 32 ++++++- packages/dashboard/src/import-translate-service.ts | Bin 10392 -> 11109 bytes 11 files changed, 323 insertions(+), 21 deletions(-) Fusion-Task-Id: FN-8172 Fusion-Task-Lineage: b3d18f58-dc0b-47e4-a1bf-acb8b2360869 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
c6be0b158b |
FN-8129: centralize database backup settings
Move database backup policy and scheduling to shared global configuration. - Split project memory backups from cluster-wide database backup settings. - Migrate legacy backup values and routines safely into central global storage. - Schedule and dispatch one shared PostgreSQL backup routine across project engines. Files changed: .changeset/fn-8129-backup-settings-scope-split.md | 7 + docs/dashboard-guide.md | 2 + docs/settings-reference.md | 10 +- packages/cli/src/commands/backup.ts | 3 +- .../__tests__/backup-settings-migration.test.ts | 50 ++++++ .../src/__tests__/backup-settings-scope.test.ts | 27 +++ packages/core/src/backup-settings-migration.ts | 188 +++++++++++++++++++++ packages/core/src/backup.ts | 77 +++++---- packages/core/src/global-routine-store.ts | 104 ++++++++++++ packages/core/src/index.gate.ts | 6 +- packages/core/src/index.ts | 6 +- .../core/src/postgres/migrations/0000_initial.sql | 19 +++ .../postgres/migrations/0015_global_routines.sql | 19 +++ packages/core/src/postgres/schema-applier.ts | 19 ++- packages/core/src/postgres/schema/central.ts | 21 ++- packages/core/src/postgres/startup-factory.ts | 11 ++ packages/core/src/settings-schema.ts | 14 +- packages/core/src/types.ts | 31 +++- .../dashboard/app/components/SettingsModal.tsx | 10 +- .../settings/__tests__/section-keys.test.ts | 1 + .../app/components/settings/save-split.ts | 2 + .../search/__tests__/settings-search-index.test.ts | 1 + .../settings/search/entries.ts | 2 + .../app/components/settings/section-keys.ts | 4 - .../settings/sections/BackupsSection.search.ts | 40 ----- .../settings/sections/BackupsSection.tsx | 112 +----------- .../sections/DatabaseBackupsSection.search.ts | 51 ++++++ .../settings/sections/DatabaseBackupsSection.tsx | 142 ++++++++++++++++ .../settings-default-descriptions.test.tsx | 1 + packages/dashboard/src/routes.ts | 12 +- .../src/routes/register-settings-memory-routes.ts | 41 ++--- .../engine/src/__tests__/routine-scheduler.test.ts | 55 +++++- packages/engine/src/cron-runner.ts | 4 +- packages/engine/src/routine-runner.ts | 67 +++++--- packages/engine/src/routine-scheduler.ts | 35 +++- 35 files changed, 929 insertions(+), 265 deletions(-) Fusion-Task-Id: FN-8129 Fusion-Task-Lineage: af17f39a-7f1c-40ff-8a4a-cd63895cd532 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
478f226a54 |
test: green full-suite CI after main drift (#2229)
## Summary Restores green **Full Suite (non-blocking)** runs on `main`. Recent main merges left i18n key parity, schema baseline bookkeeping (0011→0012), heartbeat tool inventory (FN-8058 `fn_task_logs_read`), and merger whitespace-classification mocks (execFile `git diff -p -w :2: :3:`) out of date, so all four test shards failed. ## Root causes observed on main - **Shard 4 / `@fusion/i18n`**: missing `skipConfirmationDialogs*` + `reviewBudgetExhausted` in non-en locales; orphan `awaitingApprovalPlanReviewReplanCap` - **Shard 3 / `@fusion/core`**: `SCHEMA_BASELINE_VERSION` advanced to `0012` while tests still equated it with `OWNER_PROJECT_ID_SPLIT_VERSION` (`0011`) and omitted `0012` from applied-migration lists - **Shards 1–2 / `@fusion/engine`**: tool count/snapshot drift for `fn_task_logs_read`; merger tests still mocked `git diff-tree` for trivial classification after the execFile `:2:`/`:3:` cutover; mock provider `updateTask` arity drift ## Changes - Locale catalogs: add missing keys, drop orphan key - Schema applier tests: immutable 0011 identity + baseline 0012 lists - Heartbeat + gating snapshots: include `fn_task_logs_read` - Merger unit mocks: recognize `git diff -p -w :2:path :3:path` - Mock provider: accept optional third `updateTask` arg ## Test plan - [x] `pnpm --filter @fusion/i18n exec vitest run` — 23/23 - [x] `pnpm --filter @fusion/core exec vitest run src/__tests__/postgres/schema-applier.test.ts` (immutable + automation upgrade) — pass - [x] `pnpm --filter @fusion/core exec vitest run` project-identity + satellite-fusiondir — pass - [x] Engine suites from failed CI shards (file-scoped, hermes/openclaw/paperclip/grok, reliability post-finalize/mission, heartbeat, gating, merger recovery/prompt, mock-provider, etc.) — pass - [ ] Full Suite workflow green on merge to main <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Improved project data isolation across backend operations. - Added safer optional toast handling when UI components render outside the full application shell. - Added support for reading task logs during agent heartbeat sessions. - **Bug Fixes** - Prevented runtime probes from hanging and avoided scanning large binary files. - Improved path handling for workspaces with missing descendants. - Corrected task retry state resets and GitHub import/issue-close behavior. - **Style** - Improved chat, terminal, and settings spacing. - Added clearer accessibility labeling for the auto-merge control. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
d870878a23 |
FN-7998: add executor alternate model escalation
Add opt-in executor escalation after same-model tool-failure retries are exhausted. - Persist escalation settings and one-shot task state across SQLite and PostgreSQL stores. - Retry once on a configured alternate model or scheduler node and audit escalation outcomes. - Expose escalation controls, documentation, translations, migration, and regression coverage. Files changed: .changeset/fn-7998-executor-escalation.md | 7 ++ AGENTS.md | 1 + docs/settings-reference.md | 13 ++- .../core/src/__tests__/settings-defaults.test.ts | 23 ++++- packages/core/src/in-review-stall.ts | 29 ++++++ packages/core/src/index.gate.ts | 3 +- packages/core/src/index.ts | 3 +- packages/core/src/manual-retry-reset.ts | 1 + .../0014_executor_escalation_attempt.sql | 2 + packages/core/src/postgres/schema-applier.ts | 17 ++++ packages/core/src/postgres/schema/project.ts | 1 + packages/core/src/settings-schema.ts | 4 + packages/core/src/store.ts | 2 +- packages/core/src/task-store/persistence.ts | 2 + packages/core/src/task-store/remaining-ops-2.ts | 2 +- packages/core/src/task-store/remaining-ops-3.ts | 2 +- packages/core/src/task-store/remaining-ops-6.ts | 2 +- packages/core/src/task-store/serialization.ts | 1 + packages/core/src/task-store/task-update.ts | 2 + packages/core/src/types.ts | 13 +++ .../dashboard/app/components/SettingsModal.tsx | 12 +++ .../app/components/settings/section-keys.ts | 4 + .../settings/sections/SchedulingSection.search.ts | 36 +++++++ .../settings/sections/SchedulingSection.tsx | 6 ++ .../settings-default-descriptions.test.tsx | 4 + .../__tests__/executor-tool-failure-retry.test.ts | 91 +++++++++++++++++- packages/engine/src/executor.ts | 104 +++++++++++++++++++-- packages/i18n/locales/en/app.json | 8 ++ 28 files changed, 376 insertions(+), 19 deletions(-) Fusion-Task-Id: FN-7998 Fusion-Task-Lineage: bbce767d-c61a-4667-be62-abc0cc54d8be Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
60b6e3e048 |
FN-7996: add configurable executor tool-failure retries
Add bounded, durable same-model retry handling for qualifying consecutive executor tool errors. - Persist retry claims, cursors, and audit markers with PostgreSQL migrations. - Expose project retry count, backoff, and failure threshold settings in the dashboard. - Cover retry, exhaustion, reset, and stale-run safety behavior with tests. Files changed: .changeset/fn-7996-executor-tool-failure-retry.md | 7 + AGENTS.md | 1 + docs/architecture.md | 1 + docs/settings-reference.md | 10 ++ .../executor-tool-failure-retry-claim.test.ts | 17 +++ .../core/src/__tests__/manual-retry-reset.test.ts | 3 + .../core/src/__tests__/settings-defaults.test.ts | 15 +- packages/core/src/in-review-stall.ts | 20 +++ packages/core/src/index.gate.ts | 6 + packages/core/src/index.ts | 6 + packages/core/src/manual-retry-reset.ts | 3 + .../0013_executor_tool_failure_retry.sql | 4 + packages/core/src/postgres/schema-applier.ts | 17 +++ packages/core/src/postgres/schema/project.ts | 3 + packages/core/src/settings-schema.ts | 3 + packages/core/src/store.ts | 10 +- packages/core/src/task-store/persistence.ts | 7 + packages/core/src/task-store/remaining-ops-2.ts | 2 +- packages/core/src/task-store/remaining-ops-3.ts | 2 +- packages/core/src/task-store/remaining-ops-6.ts | 65 ++++++++- packages/core/src/task-store/serialization.ts | 3 + packages/core/src/task-store/task-update.ts | 6 + packages/core/src/types.ts | 16 +++ .../dashboard/app/components/SettingsModal.tsx | 15 ++ .../app/components/settings/section-keys.ts | 3 + .../settings/sections/SchedulingSection.search.ts | 27 ++++ .../settings/sections/SchedulingSection.tsx | 4 + .../settings-default-descriptions.test.tsx | 3 + .../__tests__/executor-tool-failure-retry.test.ts | 160 +++++++++++++++++++++ packages/engine/src/executor.ts | 87 ++++++++++- packages/i18n/locales/en/app.json | 6 + 31 files changed, 523 insertions(+), 9 deletions(-) Fusion-Task-Id: FN-7996 Fusion-Task-Lineage: d1682ef8-534c-410e-b74c-1f2cf176eac2 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
d4914eb8b3 |
FN-8127: fix embedded PostgreSQL backups
Enable backup managers to resolve active embedded PostgreSQL runtime URLs safely. - Track embedded backend URLs with generation-aware lifecycle leases. - Keep backup resolution current through owner shutdown and joiner release. - Document PostgreSQL client-tool requirements and add regression coverage. Files changed: .changeset/fn-8127-embedded-backup.md | 7 ++ docs/settings-reference.md | 3 + packages/core/src/__tests__/backup.test.ts | 115 ++++++++++++++++++++ packages/core/src/backup.ts | 17 ++- packages/core/src/index.gate.ts | 9 ++ packages/core/src/index.ts | 9 ++ .../core/src/postgres/active-backend-registry.ts | 119 +++++++++++++++++++++ packages/core/src/postgres/embedded-lifecycle.ts | 5 + packages/core/src/postgres/index.ts | 9 ++ packages/core/src/postgres/startup-factory.ts | 118 +++++++++++++++++--- 10 files changed, 389 insertions(+), 22 deletions(-) Fusion-Task-Id: FN-8127 Fusion-Task-Lineage: 6125be5c-d1d5-4228-a6b1-290311de70d9 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
de1638e262 |
FN-8090: use mmap shared memory for embedded PostgreSQL
Enable constrained-host embedded PostgreSQL startup without SysV shared-memory exhaustion. - Default embedded lifecycle flags to mmap-backed shared memory while preserving caller overrides - Cover normal and elevated Windows launch paths with deterministic flag propagation tests - Document the 64MB /dev/shm support floor and add a patch changeset Files changed: .changeset/fn-8090-embedded-pg-shm.md | 7 ++ docs/postgres-migration-review-2026-07-14.md | 4 + docs/storage.md | 5 ++ .../__tests__/postgres/embedded-lifecycle.test.ts | 88 ++++++++++++++++++++++ .../postgres/embedded-windows-admin.test.ts | 17 +++++ packages/core/src/postgres/embedded-lifecycle.ts | 50 +++++++++++- .../core/src/postgres/embedded-windows-admin.ts | 2 +- 7 files changed, 168 insertions(+), 5 deletions(-) Fusion-Task-Id: FN-8090 Fusion-Task-Lineage: ac175843-69ba-4c9c-9692-aff095fc351f Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
e87b51bd07 |
FN-8054: add pinned chat conversations
Add durable, scoped pinning for Direct chat conversations. - Add pinned session persistence, migration coverage, and archive-safe row locking. - Enforce a three-conversation per-project pin limit through the chat API. - Add desktop and mobile pin controls, sorting, indicators, and regression tests. Files changed: .changeset/fn-8054-pin-conversations.md | 7 ++ docs/dashboard-guide.md | 2 + .../postgres/satellite-db-injected-stores.test.ts | 13 ++++ packages/core/src/async-chat-store.ts | 27 ++++++++ packages/core/src/chat-store.ts | 63 ++++++++++++++++-- packages/core/src/chat-types.ts | 9 +++ .../core/src/postgres/migrations/0000_initial.sql | 1 + .../postgres/migrations/0012_chat_session_pins.sql | 8 +++ packages/core/src/postgres/postgres-health.ts | 3 + packages/core/src/postgres/schema-applier.ts | 30 ++++++++- packages/core/src/postgres/schema/project.ts | 3 + packages/dashboard/app/api/legacy.ts | 1 + packages/dashboard/app/components/ChatView.css | 32 +++++++++- packages/dashboard/app/components/ChatView.tsx | 74 ++++++++++++++++++++-- .../dashboard/app/hooks/__tests__/useChat.test.ts | 21 ++++++ packages/dashboard/app/hooks/useChat.ts | 72 +++++++++++++++++---- .../dashboard/src/routes/register-chat-routes.ts | 25 +++++++- 17 files changed, 366 insertions(+), 25 deletions(-) Fusion-Task-Id: FN-8054 Fusion-Task-Lineage: 088cb01c-582b-4f56-a222-214da90ff356 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
375368e147 |
FN-8051: ensure PostgreSQL schemas initialize before plugin hooks
Ensure required PostgreSQL namespaces exist before plugin initialization on every boot. - Create project, central, and archive schemas under the schema advisory lock before hooks run - Cover marker-present databases with a plugin-hook schema availability regression test - Add a patch changeset for the reliability fix Files changed: .changeset/fn-8051-schema-init.md | 7 ++++ .../src/__tests__/postgres/schema-applier.test.ts | 43 ++++++++++++++++++++++ packages/core/src/postgres/schema-applier.ts | 12 ++++++ 3 files changed, 62 insertions(+) Fusion-Task-Id: FN-8051 Fusion-Task-Lineage: a3b20683-a742-4a8c-9cfc-fbf316c5649b Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
261901343e |
fix(core): split the domain project field from the RLS partition column (#2165)
## Problem Migration 0006 made `project_id` the RLS isolation partition on every `project`-schema table — stamped by a BEFORE INSERT trigger from the `fusion.project_id` session GUC, with every PK/unique/FK rewritten to composite `(project_id, …)`. Eleven tables **also** carried a caller-supplied domain `projectId` on their TS types and wrote that domain value into the same physical column. When the domain value differs from the session GUC, the parent row lands in the domain partition while child rows (`research_run_events`, `experiment_session_records`, `eval_task_results`, …) land in the session partition — and the composite FK fails with SQLSTATE 23503. Appending an event to a project-owned research run could not persist. ## Fix **Decision (operator): separate domain column; `project_id` stays the partition.** - **Migration `0011_owner_project_id.sql`** adds a nullable `owner_project_id` domain column to the 11 conflated tables (`research_runs`, `experiment_sessions`, `todo_lists`, `eval_runs`, `chat_sessions`, `chat_rooms`, `ai_sessions`, `chat_token_usage`, `project_insights`, `project_insight_runs`, `cli_sessions`), backfills it from `project_id` (identical in production, so exact; the `__legacy_unscoped__` sentinel backfills to NULL), and indexes it. Idempotent, `to_regclass`-guarded per the 0007 pattern. - **Stores** (`async-research-store`, `async-experiment-session-store`, `async-todo-store`, `async-chat-store`, `async-ai-session-store`, `async-eval-store`, `async-insight-store`, `cli-session-store`, …) stop writing `project_id` entirely — the trigger/GUC owns the partition — and map their domain `projectId` field to `owner_project_id` for both reads and filters. TS types unchanged. - **Applier** registers `OWNER_PROJECT_ID_SPLIT_VERSION = "0011"` and advances `SCHEMA_BASELINE_VERSION`. ## Verification (re-run independently of the implementing agent) - Core `tsc --noEmit`: exit 0 · `pnpm lint`: exit 0 · `pnpm check:changesets`: exit 0 · `pnpm test:gate`: 185/185 - Full postgres suite: **5 failed / 807 passed** vs a **7 / 804** baseline — the two conflation round-trips (`satellite-db-injected-stores` ResearchStore + ExperimentSessionStore) go green, zero new failures. The remaining 5 are pre-existing unbound-harness `__meta`/identity failures, unrelated to this change. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Corrected project-scoped persistence and queries across AI sessions, chats (rooms + token usage), evaluations/experiments, insights, research, and todos by separating domain ownership from RLS partitioning. * Prevented foreign-key and row-level security violations when storing or retrieving project-scoped data, including legacy records. * **Database / New Features** * Added migration 0011 introducing `owner_project_id` and backfilling existing rows to preserve ownership while improving isolation. * **Tests** * Updated migration-parity coverage to include the new baseline step. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b51de02a54 |
Revert "fix(core): resolve unbound project ids to a real partition or no filter"
This reverts commit
|
||
|
|
a048a619fc |
fix(core): resolve unbound project ids to a real partition or no filter
Six of the eight postgres-suite failures shared one root cause: writes
normalize project_id, reads did not. The fusion_assign_project_id trigger
(migration 0006) rewrites a blank project_id to the session's fusion.project_id
or '__legacy_unscoped__', but helpers reached as `layer.projectId ?? ""` then
filtered on the literal '' -- a value the database never stores. Every unbound
read missed rows it had just written.
AsyncDataLayer.projectId is optional by design (undefined = project-agnostic),
so `?? ""` is the bug: it turns "no scope" into a scope that matches nothing.
The resolution differs by what the rows are, and conflating them corrupts data:
- Data and analytics reads (usage events, agent runs, research runs) take
projectScopeFor(): a bound id filters, an unbound one reads across projects.
This matches the contract taskProjectScope already documents ("when undefined
the scope filter is a no-op").
- __meta migration guards (project-identity stamps, agent-store markers) take
projectPartitionId(): an unbound id resolves to the shared sentinel
partition. projectScopeFor would be wrong here -- dropping the predicate lets
an unbound getMetaValue return whichever project's marker it finds first, so
on the shared cluster project A's "migration complete" marker would tell
project B to skip a migration it never ran. upsertMetaValue already documented
this: "the empty binding remains the explicit project-agnostic compatibility
partition". Writing the sentinel explicitly also keeps the partition
deterministic -- a blank write from a session carrying fusion.project_id would
otherwise land in that project's stamp.
Names the sentinel (LEGACY_UNSCOPED_PROJECT_ID) instead of open-coding it, and
puts both helpers next to taskProjectScope so the convention has one home.
Fixes taskstore-remaining (24/24), project-identity (6/6), and
satellite-fusiondir-stores (16/16).
The remaining two failures are a different bug and are NOT addressed here: the
child tables research_run_events and experiment_session_records never declared
project_id in schema-as-code, though migration 0006 added the column and
rewrote their FKs to composite (project_id, parent_id). Drizzle therefore cannot
write the parent's partition, the trigger stamps '__legacy_unscoped__', and the
FK fails against a project-owned parent. That needs a schema-as-code change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
a588c38784 |
fix(core): read usage events across projects when the layer is unbound
An unbound (project-agnostic) data layer read zero usage events it had just
written. AsyncDataLayer.projectId is optional by design -- undefined means a
project-agnostic layer for single-project / global / analytics reads -- but
helpers taking `projectId: string` are called as `layer.projectId ?? ""`, which
turns "no scope" into a literal '' scope.
'' never matches: the fusion_assign_project_id BEFORE INSERT trigger (migration
0006) rewrites a written '' to the session's fusion.project_id or
'__legacy_unscoped__', so a read filtering on '' looks for a value the database
never stores. Writes normalize, reads did not. Proven by probe: the row is
present with project_id '__legacy_unscoped__', emitUsageEvent returns true, and
queryUsageEvents returns [] even with no other filters.
Treat blank as unbound and drop the scope predicate, matching the contract
taskProjectScope already documents ("when undefined the scope filter is a
no-op"). Restricting an unbound reader to '__legacy_unscoped__' rows instead
would make an unscoped analytics read silently partial.
Adds projectScopeFor() next to taskProjectScope so the convention has one home
rather than a third open-coded variant.
Note the write path is already live: remaining-ops-7.ts emits with
`layer.projectId ?? ""` under backendMode, so unscoped events are accumulating
under the sentinel today. The async reader has no production caller yet, which
is why nothing user-facing broke.
Fixes taskstore-remaining.test.ts (24/24). The remaining failures in that suite
share this root cause but not this resolution -- see the follow-up.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d1bda3683c |
fix(core): reap the losing wrapper and stop self-joining on a startup race
Two related leaks on the embedded Postgres startup-race join. The flagged one: the catch dropped `nonAdminHandle` to null without stopping it, so a wrapper that onLaunched had already published leaked. The obvious fix -- call handle.stop() first -- is worse than the leak. stop() runs killAll(), which resolves its target by reading line 1 of the data dir's postmaster.pid. On this path that file belongs to the process that WON the race, so stop() would taskkill the instance we are joining. pg.stop() is the same trap via pg_ctl -D on the shared dir, which is why settleCancelledStart (it calls both) cannot be reused here. Added NonAdminServerHandle.stopWrapperOnly(), which kills only our wrapper pid and its children, and called it before the handle is dropped. A racing winner is another process's child, so /t cannot reach it. The one found while making that safe: the catch joined on ANY start failure. A start that took the lock and then failed later (readiness timeout, non-admin poll error) reads back its OWN postmaster.pid, so isAlreadyRunning hands back our own port and we "join" ourselves with ownsProcess=false -- nothing ever stops it, orphaning a live postmaster for the life of the host. The join now fires only on a lock-collision error, which is the one failure proving our postgres refused to start and someone else owns the dir. Every other failure returns to the existing cancellation/cleanup paths, which stop what they started. That is also what makes the wrapper-only kill provably safe: on this path our postgres never took the lock. Tests: a non-lock failure must propagate even with a postmaster.pid present (fails without the fix -- the old catch swallowed it and joined), and a lock collision must still join. Both always-on with a mocked ctor. Pre-existing and unrelated: taskstore-remaining.test.ts fails identically on a clean tree with these changes stashed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
130c70286b |
fix(core): create the database when joining a racing embedded Postgres
A lifecycle that joins an already-running instance returned a connection URL before the owner had created the database. The owner calls ensureDatabase() only after its own start() resolves, but the signals a joiner detects the instance by -- the runningInstances entry and, decisively, postmaster.pid, which postgres itself writes -- both appear earlier. A joiner landing in that window handed back a URL to a database that did not exist and failed at the caller's first connect. Reordering the owner's publish does not fix it: isAlreadyRunning falls back to the pid file, whose timing postgres owns, so the joiner must verify. Both join paths (preflight and the startup-race catch) now create the database if absent. Creating from the joiner is safe rather than a second writer -- CREATE DATABASE is atomic and both sides tolerate the duplicate, so whoever loses treats the winner's database as its own success. Verification takes the joined instance's port explicitly. getPort() resolves to `options.port ?? resolvedPort`, which on a join with an explicitly configured port is this instance's requested port, not the one being joined. It is best-effort by contract: isAlreadyRunning joins optimistically without probing (a stale pid file from a crash still resolves to a port), so a probe failure logs and returns the URL exactly as before, letting the connection layer report an unreachable cluster. A hard throw would turn every stale-pid start into a startup failure. Duplicate tolerance covers both codes a real cluster produces: 42P04 duplicate_database when the winner committed before our catalog probe, and 23505 unique_violation on pg_database_datname_index when the two CREATEs collide inside the catalog insert. The concurrent-ensureDatabase test caught the 23505 arm -- tolerating only 42P04 left the tighter half of the race throwing. Tests: a real-process test proving a joiner creates the database the owner has not (drop-the-database reproduces the window), a real-process concurrent ensureDatabase race, and an always-on test pinning the best-effort contract for an unreachable join. All three fail without the fix. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
e9537c9e85 |
docs(core): correct the ensureDatabase comment on the postgres join path
The preflight join carried "// Ensure the database exists on the running instance" above a line that only builds a URL. No ensureDatabase() call has ever followed it, so the comment described behavior the code does not have. Replace it with why the call is absent: a joiner has no cluster of its own to ensure, the owning process creates the database after its own start(), and ensureDatabase() would throw here anyway because it requires `this.running` -- which the join path leaves false by design so stop() never reaps an instance we did not start. Also records the ordering assumption the path rests on: the owner publishes runningInstances / writes postmaster.pid before its ensureDatabase() resolves, so a joiner winning that window fails at the connection layer rather than silently using a missing database. Comment-only; no behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8023aa2d08 |
fix(core): do not rescue a cancelled embedded Postgres start into a success
The startup-race join added in
|
||
|
|
e33039ad0f |
fix(core): join competing postmaster when embedded Postgres startup races
Starting a second Fusion process could fail with `lock file "postmaster.pid" already exists`. The singleton preflight check and `pg.start()` are not atomic, so another process can create the lock in between — the loser surfaced the collision to the TUI as an error instead of simply joining the live instance. `EmbeddedPostgresLifecycle.start()` now wraps the start path in a try/catch. On failure it re-reads `postmaster.pid` via `isAlreadyRunning()`; when a live instance is found it connects to that port with `ownsProcess=false` (so this process never stops a server it did not start) and logs the race. Failures with no live instance rethrow unchanged, so genuine startup errors are unaffected. Regression test lives outside the real-process `embeddedDescribe` block — it uses a mocked ctor, and nesting it there would skip it under FUSION_EMBEDDED_TEST_SKIP=1 (the gate/CI default), leaving the fix unprotected. Verified: 35/35 embedded-lifecycle tests pass, core typecheck clean, lint clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5445693e51 |
fix(FN-8009): quiet embedded backend TUI logs
Suppress routine embedded-backend resolution messages while retaining redacted external-backend diagnostics. |
||
|
|
0863c0fb58 |
feat(dashboard): auto-translate foreign-language GitHub issues on import (#2141)
## Why The Import Tasks panel routinely lists issues in languages the operator cannot read. Translation already shipped in #2128, but deliberately **opt-in and preview-only** — its header comment read *"Translation is opt-in (never automatic) so import provenance stays faithful until the operator asks."* This reverses that decision **behind a default-off setting**, so operators who never opt in keep byte-faithful import provenance. The superseded comment is kept and annotated rather than deleted, so the reason the rule changed stays in the code. ### The structural gap #2128 left `POST /github/issues/import` accepts only `{owner, repo, issueNumber}` and **re-fetches the issue server-side**. A translation held in React state could never reach the created task, and the in-memory cache died with the modal. That is why the cache here is server-side rather than in the hook — it's what makes "imported issues carry the translated version" actually true. ## What operators get Auto-translate is **off by default**. When enabled: - The **50 most recent OPEN** foreign-language issues translate on panel load — **list titles**, not just the preview, so the list reads in your language before you click anything. - Translations show **by default**, with a toggle back to the original (hover a translated list title to see the original). - Translations **persist until the issue closes**, so re-opening the panel neither waits nor re-bills. - **Both single and batch import** carry the translation, so the created task reads like the preview you approved. - A **target language** setting (unset = follow the dashboard language) and a dedicated **model lane**, so you can pin a cheap/fast model without dragging the summarization lane onto it. ## Notable decisions | Decision | Why | |---|---| | Detect **before** the model | An issue already in the target language is never sent. Without this, an English repo with the setting on would bill every issue to return its input unchanged. | | Detection moved to `@fusion/core` | The panel and the server must not disagree about which issues are foreign; two copies of a heuristic drift. | | Own rate-limit budget | Translation shared a 10/hour budget with refine/goal-draft. Fanning out per-issue would fail partway **and** starve refine for the hour. | | Cache keyed on a **source hash** | An edited issue misses the cache and re-translates instead of serving stale prose. | | Import is **cache-read only** | A miss imports the original. Import must never block on, or fail because of, translation. | | `project_id` leads the cache PK + full RLS contract | All projects share one flat `project` schema. `verification_cache`'s PK predates that discipline; this table does not copy that mistake. | ## Verification - ✅ `pnpm lint`, `@fusion/core` + `@fusion/dashboard` typecheck - ✅ `pnpm verify:fast` — build + scoped typecheck + real boot smoke (`/api/health`) - ✅ `pnpm test:gate` — 479 tests - ✅ 19 new tests covering the billing invariants (off/closed/same-language ⇒ **no model call**), cache hit/miss-on-edit, the 50 cap, and per-item fail-soft - ✅ `schema-applier` real-Postgres suite (46 tests) exercises migration `0010` and its isolation invariant **Pre-existing failures NOT touched** (confirmed red on `HEAD` before this branch): `AppearanceSection`'s task-popup test, and two PG-cutover keys (`sqliteMigrationNotice`, `postgresMigrationInboxMessageSentAt`) missing description mappings. I left the latter rather than guess an allowlist entry that could mask a real coverage gap. ## Reviewer notes - Short Latin-script prose (a one-line Spanish title) rates only *medium* confidence and won't auto-translate — the existing heuristic is deliberately conservative so English issues are never billed. CJK detects regardless of length. The threshold is the knob if you'd rather bias toward translating. - The RLS/isolation contract in migration `0010` is the part most worth a careful look. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
05151a25db |
feat: faster dashboard and serve startup (#2132)
## Summary Speeds up **time-to-HTTP-ready** for `fn dashboard` and `fn serve` after the PostgreSQL cutover without reintroducing the historical 3s cwd-engine race that degraded webhooks. - **Dashboard store share (serve parity):** inject the factory-booted `TaskStore` as `externalTaskStore` so cwd `ensureEngine` does not open a second pool; share only when store root matches project working directory (multi-project safe). - **Serve multi-project:** stop awaiting `startAll()` before listen; await only the primary engine; background the rest + reconciliation. - **Defer non-route-critical engine work:** ordered OAuth (refresh → monitor), automation schedule syncs, and auto-merge **enqueue** after the engine handle is returnable. - **Critical-path merge status clear:** still clear stale `merging`/`merging-pr` before ready so manual merge is not blocked after crash. - **Serve `--paused`:** apply `enginePaused` before `ensureEngine`/`startAll` (dashboard ordering). - **Stop safety:** generation counter so deferred tails cannot resume after `stop()` clears `shuttingDown`. - **Phase timing:** shared `phaseTime` helper, factory substep logs, serve time-to-listen. Plan: `docs/plans/2026-07-14-001-feat-faster-startup-plan.md` ## Test plan - [x] `packages/engine` — `project-engine-manager.test.ts` (path-matched external store) - [x] `packages/engine` — `project-engine-deferred-startup.test.ts` (status clear, OAuth order, stop generation) - [x] `packages/cli` — `startup-phase.test.ts` - [x] `packages/cli` — `serve.test.ts` (60 tests, including `--paused`) - [ ] Local: warm `fn dashboard` / `fn serve` and compare `startup phase *` / `time-to-listen` logs - [ ] `pnpm smoke:boot` (real serve `/api/health` on ephemeral port) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Performance** * Improved dashboard and serve startup times, including faster time-to-listen and time-to-ready. * Moved non-essential background initialization off the critical startup path. * Parallelized dashboard service initialization where possible. * **Reliability** * Improved multi-project startup handling and project selection. * Prevented cross-project task-store sharing. * Added safer shutdown behavior for partially completed startup. * **Diagnostics** * Added startup phase timing logs to help identify performance bottlenecks. * **Tests** * Expanded coverage for deferred startup, shutdown, project isolation, and startup timing. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
78ef3075f6 |
fix(core): prevent plugin migration startup crash
Run retained SQLite plugin recovery through the privileged startup connection before handing stores to the restricted PostgreSQL runtime role. |
||
|
|
a242f1b449 |
fix(FN-7952): migrate bundled plugins to PostgreSQL (#2111)
## Summary Bundled plugins now persist shared runtime state in project-scoped PostgreSQL tables instead of maintaining independent SQLite authority. Reports, CLI Printing Press, Compound Engineering, Roadmap, Even Realities, and WhatsApp all follow the same ownership and startup contract as Fusion core. ## Design decisions - Plugin schema hooks run through the host’s PostgreSQL owner and enforce project isolation. - The SDK exposes the host contract needed by bundled plugins without importing engine internals. - Legacy Roadmap ownership fixtures use the supported empty-owner sentinel, preserving current composite primary/foreign keys while exercising backfill behavior. - The lockfile travels with the Even Realities PostgreSQL dependency so packaged installs remain reproducible. ## Validation - All six affected plugin builds pass. - Affected plugin suites pass: 773 tests across Printing Press, Compound Engineering, Even Realities, Reports, Roadmap, and WhatsApp. - `pnpm test:gate` passes all 478 gate tests. - This PR changes 40 files. ## Stack - Depends on #2110 → #2109 → #2108. - The documentation/release PR completes the stack. Related: #2105 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Breaking Changes** * PostgreSQL is now required for runtime storage; SQLite files are used only as one-time migration inputs. * The legacy `FUSION_NO_EMBEDDED_PG` fallback has been removed. * **New Features** * Added project-isolated PostgreSQL storage for plugins, reports, tasks, notifications, and other plugin data. * Added agent tools for reports and CLI service drafts. * Added PostgreSQL schema initialization support for plugin authors. * **Bug Fixes** * Improved migration and recovery of legacy plugin state. * Prevented cross-project data access and strengthened transactional schema updates. * **Documentation** * Updated storage, migration, deployment, plugin authoring, CLI, and dashboard guidance for PostgreSQL. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
6c008418fe |
fix(core): boot embedded Postgres under non-admin user on elevated Windows (#2117)
## Summary Windows embedded Postgres verification (CI `windows-latest` and elevated desktop) fails because PostgreSQL refuses to run under an administrative token: > Execution of PostgreSQL by a user with administrative permissions is not permitted. GitHub Actions runners execute as `runneradmin` elevated, so the existing `test:embedded-postgres` smoke (and any elevated Local-mode desktop launch) cannot start the server. ### Fix - When `isWindowsElevatedAdmin()` is true, **initdb / clients stay as the launcher**, but the **postgres server** is started as a dedicated non-admin local user (`fusion-pg`) via PowerShell `Start-Process -Credential`. - Readiness waits on the postgres log line `database system is ready to accept connections` with a lightweight poll (no per-iteration `tasklist`). - Real-process vitest cases use a **180s** timeout on Windows (package default is 15s, which killed healthy boots mid-start). - Builds on top of the packaged-desktop asar materialization work already on main (#2106). ## Test plan - [x] `pnpm --filter @fusion/core test:embedded-postgres` on macOS (33/33) - [ ] `desktop-windows.yml` on `feature/win-pg-verify`: - [ ] Smoke embedded Postgres on Windows - [ ] Build + package Windows EXE - [ ] Verify app.asar assets - [ ] Optional: download portable EXE and manual Local mode smoke on a Windows host <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved embedded PostgreSQL startup on Windows when Fusion runs with elevated administrator privileges. - When elevated, the embedded database now boots under a dedicated non-administrator local account, with more reliable readiness detection, logging, and shutdown cleanup. - Enhanced database provisioning and now prefers `127.0.0.1` for Windows connection addressing. - **Tests** - Added coverage for Windows elevation detection without starting embedded PostgreSQL. - Increased platform-dependent timeouts for embedded real-process tests to avoid premature failures on Windows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
2e4fcfcaea |
fix(FN-7952): establish PostgreSQL core authority (#2108)
## Summary Fusion’s core runtime now treats PostgreSQL as the authoritative metadata store without leaving current CLI, dashboard, desktop, or engine composition roots uncompilable between stack layers. This is the 99-file foundation for the larger cutover: subsequent PRs migrate the remaining consumers, plugins, and operator surfaces. ## Design decisions - Runtime store construction fails closed when an asynchronous PostgreSQL layer is unavailable; SQLite remains readable only at explicit migration and identity-recovery boundaries. - Project ownership is enforced across active, archived, workflow, mission, analytics, and plugin-schema data. - The small set of cross-package files in this layer are compatibility-critical call sites required for a green intermediate commit, not the complete consumer migration. - Schema migration 0008 remains assigned to session-advisor state from current `main`; mission lineage idempotency advances to 0009 so neither invariant can be skipped. ## Validation - All affected package typechecks pass: Core, Engine, Dashboard, CLI, and Desktop. - `pnpm test:gate` passes: 478 tests across the engine gate, PostgreSQL core gate, and CLI workflow shape. - The PR changes exactly 99 files. ## Stack This is the base PR. Engine/dashboard, CLI/desktop/ops, plugins, and docs/release follow as stacked PRs, each below 100 changed files. Related: #2105 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * PostgreSQL is now the standard runtime backend, with embedded PostgreSQL enabled by default. * Added project-scoped storage for tasks, archives, chat sessions, missions, knowledge pages, and operational data. * Improved archived-task search, filtering, pagination, and restoration. * Added safer plugin schema initialization with validation and project isolation. * Added PostgreSQL-backed workflow, mission, validator, and dashboard capabilities. * **Bug Fixes** * Improved startup timeout cancellation and resource cleanup. * Prevented cross-project data access and phantom reservation cleanup errors. * Ensured archived tasks remain read-only and asynchronous writes complete reliably. * Retired SQLite opt-out settings with clear startup errors. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
f4e78abeb7 |
fix: resolve CREATE ROLE fusion_runtime race condition in migration 0006 (#2104)
## Summary Fixes `CREATE ROLE fusion_runtime` race condition in migration `0006_project_ownership.sql` that causes 30 compound-engineering test failures on CI. ## Root Cause Concurrent test databases on the same PostgreSQL service container race on `CREATE ROLE fusion_runtime`: the `IF NOT EXISTS` check is not atomic (roles are cluster-wide, not per-database). Between the check and the `CREATE ROLE`, another session can create the role, causing error `23505` (unique_violation). ## Fix Replace the non-atomic `IF NOT EXISTS` guard with a `BEGIN...EXCEPTION WHEN duplicate_object OR unique_violation THEN NULL; END;` block that safely handles the race. ## Verification | Check | Result | |---|---| | compound-engineering (pipeline-store + orchestrator + session-routes) | ✅ 41 passed | | Engine shard 1/2 | ✅ 3826 passed, 0 failed | | Merge gate | ✅ 471 passed | | Lint | ✅ exit 0 | |
||
|
|
7568705244 |
fix(desktop): boot embedded Postgres in packaged app and ship omp dist (#2106)
## Summary Packaged Fusion desktop Local mode failed after the SQLite→Postgres cutover: 1. **Embedded Postgres** could not start from `app.asar` — platform packages resolve `initdb`/`postgres` via `import.meta.url` into the asar virtual path, and `spawn` fails with `ENOTDIR`. 2. **After Postgres was fixed**, Local mode still fell back to the mode chooser because `@fusion-plugin-examples/omp-runtime` was never built into `dist/` (dashboard imports it from `runtime-provider-probes.ts`). This PR makes packaged Local mode boot embedded Postgres reliably and keep the dashboard shell up. ### Changes - **CJS bootstrap** (`main-bootstrap.cjs`) as Electron `main`: patches `child_process.spawn` / `fs.promises.stat|chmod` before the ESM main loads so asar binary paths rewrite to real files. - **Materialize** the full native PG install (`bin` + `lib` + `share`) under `~/.fusion/embedded-postgres/runtime-bin/<plat-arch>/`. - **electron-builder**: full `asarUnpack` of embedded-postgres packages; allowlist PG deps and `@fusion-plugin-examples/**/*` (+ plugin-sdk / ACP SDK). - **Build** `fusion-plugin-omp-runtime` with the other dashboard-static runtime plugins; export `DASHBOARD_RUNTIME_PLUGIN_PACKAGES` for tests. - Unit coverage for asar path rewrite, packaging allowlists, and omp build inclusion. ## Test plan - [x] `pnpm --filter @fusion/core test:embedded-postgres` (23/23) - [x] Desktop packaging unit tests (`build-bundling`, `electron-builder-config`) - [x] Packaged macOS `Fusion.app` Local mode: - [x] `embedded postgres: ready on port … (database "fusion")` - [x] `desktopMode` stays `"local"` (no chooser fallback) - [x] `GET /api/health` → `status: ok`, `database.healthy: true`, `engine.available: true` - [x] Linux embedded binary lifecycle smoke (Docker aarch64, `@embedded-postgres/linux-arm64`) — initdb/start/persist/restart - [ ] CI release desktop jobs (macOS/Linux) when this lands - [ ] Windows packaged desktop Local + PG (separate agent / host) ## Verification notes | Platform | Embedded Postgres | Packaged Local shell | |----------|-------------------|----------------------| | macOS | Working | Working after this PR | | Linux | Native binary smoke pass | Full AppImage not built on this host | | Windows | Out of scope here | Separate verification | <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved embedded PostgreSQL reliability in Electron-packaged apps by rewriting bundled `app.asar` binary paths to their unpacked/materialized locations. * Ensured embedded PostgreSQL runtime binaries resolve correctly across platforms/architectures, with best-effort executable permissions and macOS dylib link normalization. * **Packaging** * Updated the desktop Electron entry to use a bootstrap module for embedded PostgreSQL binary resolution. * Expanded Electron Builder inclusion and asar-unpack rules for embedded-postgres and related packages, plus required runtime plugin/sdk assets. * **Tests** * Updated and added checks to match the new packaging and plugin/runtime expectations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
4f037679ad |
feat: planner overseer session advisor (OMP advisor parity) (#2082)
## Summary Adds a **session advisor** to the planner overseer so Fusion can review live executor transcripts the way [oh-my-pi’s advisor](https://github.com/can1357/oh-my-pi/tree/main/packages/coding-agent/src/advisor) does — without replacing the existing lifecycle supervisor (stage watch, retry, merge confirmation, human-control withhold). ### What ships - **Emission guard** (`OverseerEmissionGuard`) — content-free phrase filter, session dedupe with severity-rank escalation, one accept per advisor update - **Session delta runtime** — queues agent-log deltas, drains through an advisor agent, drops backlog after 3 failures - **Session advisor service** — model gate, level matrix (`observe` / `steer` / `autonomous`), human-control re-check at inject, `[session-advisor]` steering comments - **OVERSEER.md / WATCHDOG.md** discovery for project review priorities - **AgentLogger `onEntriesFlushed`** + poll-backed agent-log cursor for durable deltas - Workflow settings: `plannerOverseerAdvisorProvider` + `plannerOverseerAdvisorModelId` (both required; empty = soft-disabled for cost safety) - Docs + changeset ### What does not ship (deferred) - Multi-advisor YAML roster, mutating advisor tools, reviewer/merger shadowing, true tool-abort interrupt ### Plan `docs/plans/2026-07-13-001-feat-overseer-advisor-parity-plan.md` ## Enablement 1. Set workflow **Session advisor model provider** + **Session advisor model id** 2. Oversight level `observe` (log only), `steer`, or `autonomous` (inject) 3. Optional: add `OVERSEER.md` or `WATCHDOG.md` in the project ## Test plan - [x] `pnpm --filter @fusion/core exec vitest run src/__tests__/overseer-emission-guard.test.ts` - [x] `pnpm --filter @fusion/engine exec vitest run` overseer-* unit tests (21 tests) - [x] Related planner-overseer / intervention regression tests - [x] `@fusion/engine` + `@fusion/core` typecheck - [ ] Manual: configure advisor model, run an executor task, confirm `[session-advisor]` inject + timeline metadata when concern is raised ## Residual Review Findings None from autofix pass (log-cursor ordering fix already committed). <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an off-by-default “session advisor” that can review live execution activity and provide severity-based guidance. * Added project and per-task controls to enable it, including a default enable switch and Quick Add / Task Detail toggles. * Enhanced advisor prompting by discovering and incorporating `OVERSEER.md`/`WATCHDOG.md` review files. * **Documentation** * Added architecture and settings documentation for the new session-advisor parity behavior. * **Bug Fixes** * Improved fail-soft handling so advisor behavior won’t disrupt execution. * Fixed concurrent PostgreSQL migration startup failures. * **Tests** * Added coverage for advice parsing, emission guarding, runtime behavior, and watchdog discovery. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
cdf67c1d98 |
fix(dashboard): stop Planning Mode retry loop, make AI sessions multi-tab (#2101)
## Problem Reported: planning gets stuck in a cycle of retrying and regenerating after a response was already supplied. After the user answers a planning question, `submitResponse` pushed the answer to history but left `session.currentQuestion` pointing at the just-answered question for the whole next generation. The planning SSE route's catch-up path re-emits `currentQuestion` to every fresh connection — and each FN-7946 auto-retry (#2073) opens a fresh connection. So after any generation error: 1. Auto-retry connects a fresh stream → the server re-emits the **already-answered** question. 2. The client treats any question event as progress: it **resets the 3-attempt auto-retry budget** and re-shows the answered question. 3. The retry regenerates; if it errors again the cycle repeats with a fresh budget — an unbounded retry/regenerate loop. Re-answering the stale question also 409-collided with the in-flight generation, feeding the same loop. ## Fix Invariant: `currentQuestion` is only set while the session is genuinely awaiting user input. - `submitResponse` clears it the moment an answer is accepted (normal turns and the deepening checkpoint), while preserving the legacy 200 respond contract on generation failure (the modal ignores the body and lets the SSE error drive recovery). - `retrySession` scrubs stale questions persisted by pre-fix builds before regenerating. - `buildSessionFromRow` only restores a question when the persisted row is `awaiting_input`. - `didSubmitSameAnswer` now compares against the last history entry so the duplicate-submit 409 message survives. - Agent onboarding gets the same fix (its SSE route also re-emits `currentQuestion` on connect); retry now asks the next question instead of re-asking the answered one. Surface enumeration: mission and milestone interviews keep questions the same way but their SSE routes never re-emit on connect, and the auto-retry budget machinery is Planning-Mode-only — planning + onboarding were the two affected surfaces. ## Symptom Verification - **Original symptom:** after answering a question, Planning Mode loops between "Retrying…" and regenerating, re-showing the already-answered question, with the auto-retry budget never exhausting. - **Exact reproduction:** answer a question, have the next generation fail (stuck watchdog/provider error), let the client auto-retry open a fresh SSE connection. - **Assertion it is gone:** new regression suite `planning-answered-question-reemit.test.ts` asserts `currentQuestion` is cleared mid-generation, on generation failure, on retry, and on restore from non-`awaiting_input` rows — so the SSE catch-up path has nothing stale to re-emit. All 5 tests fail against pre-fix code and pass with the fix; an onboarding regression test covers the sibling surface. ## Verification - New regression tests: 5/5 fail on pre-fix code, pass with the fix (plus 1 onboarding test). - Existing suites: 137 planning server tests pass (3 failures in `routes-planning.test.ts` fail identically without this change — pre-existing on the branch); all 69 `PlanningModeModal.planning-flow` client tests pass; `tsc --noEmit` clean; `pnpm check:changesets` passes. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Made Planning Mode (and related planning controls) lock-free and multi-tab—no more take-over/active-in-another-tab lock overlays. * **Bug Fixes** * Fixed Planning Mode retry/generation flows where already-answered questions could reappear. * Ensured answered questions clear immediately and aren’t re-emitted during session recovery/SSE catch-up. * Improved session restoration and preserved legacy recovery behavior when generation fails after an answer. * **Tests** * Added regression coverage for the answered-question invariant and updated existing tests to reflect lock-free behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --- ## Follow-up: Planning Mode is now multi-tab via DB state (lock-free) Second commit removes all cross-tab coordination from planning — the persisted session row is the single source of truth and multiple tabs can read and interact with the same session: - **Server:** `/planning/*` routes no longer run `checkSessionLock` or parse `tabId`; a stale `tabId` from an older client is ignored instead of 409'd. Subtask/mission interview routes keep their existing lock behavior. - **Client:** `PlanningModeModal` drops `useSessionLock`, the `useAiSessionSync` BroadcastChannel broadcasts, `sessionTabId`/`lockSessionId` state, and the "Take Control" overlay. Tabs stay current via the per-session SSE stream plus the global `ai_session:updated` events `useBackgroundSessions` already consumes; concurrent writes resolve via the server's generation-in-progress guard (409). - **API client:** planning functions lose their `tabId` params. - **Fix uncovered by the refactor:** the 8s stuck-poll now resolves the session id inside each tick — the removed lock state was what previously re-armed the poll after Start Planning resolved the session id. - Also fixes a pre-existing PG-cutover break in `planning-generation-cancellation.test.ts` (`getSession` is async). Verification: 144 client planning tests and 137 server planning tests pass (the 3 remaining `routes-planning.test.ts` failures are pre-existing on the branch and fail identically without these changes); `tsc --noEmit` and eslint clean on changed files; `pnpm check:changesets` passes. Lock-conflict route tests were rewritten to assert lock-free semantics, plus a new modal test proving a session stays fully interactive with no lock acquisition even when another tab is active. --- ## Follow-up 2: the per-tab session lock is gone entirely Third commit extends the multi-tab model from planning to **every** AI interview surface (planning, subtask breakdown, mission interview, milestone/slice interview) and deletes the lock machinery root and branch. **Server** - Deleted the `/ai-sessions/:id/lock`, `/lock/force`, and `/lock/beacon` routes. - Dropped `checkSessionLock` from every planning/subtask/mission/milestone route (both copies — `routes.ts` and `mission-routes.ts`). A `tabId` from an older client is ignored, never 409'd; all `tabId` body parsing is gone. - Dropped `acquireLock` / `releaseLock` / `forceAcquireLock` / `getLockHolder` / `releaseStaleLocks` from `AiSessionStore`, plus the `@fusion/core` async helpers (`acquireAiSessionLock` et al) and core's re-exports. - Removed `lockedByTab`/`lockedAt` from `AiSessionRow`/`AiSessionSummary`, the upsert SQL, and all four session producers. **Client** - Deleted `useSessionLock` and the now-orphaned `getSessionTabId` util. - Removed the Take Control overlay, the "active in another tab" banners, and `BackgroundTasksIndicator`'s active-elsewhere gate (the confirm prompt and lock badge — sessions now just open). - Reduced `useAiSessionSync` to what its own comments already called it — a low-latency *status* supplement to SSE: no `activeTabMap`, `broadcastLock/Unlock/Heartbeat`, `owningTabId`, `tab:*` messages, or stale-heartbeat sweep. - Dropped `tabId` from every session API client function; removed the lock CSS. **Deliberately kept: the two DB columns.** `ai_sessions.locked_by_tab` / `locked_at` remain as dead, always-NULL columns with a deprecation note. Dropping them is an irreversible migration, and released binaries still name those columns explicitly in their upsert — an older install pointed at the same database would fail every session write. They can be dropped once no such binary can reach it. No code reads or writes them. **Verification**: 397 client tests and 137 server planning tests pass (the same 3 `routes-planning.test.ts` failures are pre-existing — verified identical on a clean stash); `tsc --noEmit` clean for `@fusion/core` and `@fusion/dashboard`; eslint clean on all changed files; the 30 PG `schema-applier` tests pass (they exercise the retained columns); `pnpm check:changesets` passes. The lock-conflict route tests and both modal lock tests were rewritten to assert the inverse: routes and modals stay fully interactive while another tab "holds" a lock, and the lock API is never called. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
be55d0a987 |
fix(cli): reuse project stores for skill discovery (#2102)
## Summary - reuse the dashboard command's backend-aware per-project `TaskStore` cache during project-scoped plugin skill discovery - obtain plugin state through `TaskStore.getPluginStore()` instead of constructing bare SQLite-default `PluginStore` / `TaskStore` instances - keep cached project stores alive for the dashboard process while still stopping request-scoped plugin loaders - add a regression covering the real Skills adapter callback and refresh the dashboard test fixture with `getAsyncLayer()` ## Root cause `GET /api/skills/discovered` resolved the project correctly, then `getProjectScopedPluginSkills()` constructed new stores without an `AsyncDataLayer`. After `VAL-REMOVAL-005`, that enters the physically removed synchronous SQLite runtime and returns HTTP 500 even when PostgreSQL health, projects, tasks, and both project engines are healthy. The existing route tests mocked the Skills adapter callback, so they did not exercise this CLI wiring. ## Verification - targeted dashboard regression: 1 passed, 91 skipped - `pnpm lint` - `pnpm --filter @runfusion/fusion typecheck` - `pnpm --filter @runfusion/fusion build` - `pnpm check:changesets --strict` - `git diff --check` Live Atlas validation against the migrated embedded PostgreSQL runtime: - `/api/skills/discovered?projectId=proj_84f4645c2da64288`: HTTP 200, 36 skills - `/api/skills/discovered?projectId=proj_7538a9dd46c24c5f`: HTTP 200, 36 skills - local dashboard and Tailscale dashboard: HTTP 200 - controlled SIGTERM: launchd restarted the dashboard and both Skills routes remained healthy <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Fixed dashboard project-scoped plugin-skill discovery in PostgreSQL mode with safer store reuse/teardown and request-scoped plugin-loader lifecycle. - Improved dashboard cleanup to avoid duplicate concurrent store closes and ensured proper shutdown behavior per root type. - Made `fusion_runtime` role creation race-safe during concurrent PostgreSQL migrations. - **New Features** - Added `persistRuntimeState` option to control whether plugin runtime state changes are persisted. - **Tests** - Expanded dashboard and core hot-reload tests to verify scoped, non-persistent runtime behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
678265a526 |
fix(cli): show live SQLite migration progress
Report source scans, per-table copy milestones, checksum phases, verification outcomes, and unambiguous failure or finalization status during first-boot and manual migrations. |