## Summary
Restores green merge-gate and package-default suites after repeated
`origin/main` merges brought workflow-graph ownership cutover drift into
CI.
- Align engine/dashboard/core tests with post-cutover contracts
(`moveTaskIf`/`deleteTaskIf`, graph handoff, worktree-pool reclaim via
`removeWorktree` + `RemovalReason`, multi-step RESUMING parse,
soft-pause merge requester, graph-terminal failure surfaces).
- Small product fixes needed for real regressions uncovered by the
suite: soft-delete refuse before graph routing, skip DUPLICATE
step-heading withhold when an explicit marker is present, PG schema
applier guards, and related bookkeeping (research promote tool inventory
/ migration seed, stop shell `psql` in PG admin DDL).
- Quarantine/ledger hygiene only where required by standing rules; no
timeout/worker appeasement.
## Verification
- `pnpm test:gate` ×2 green
- `@fusion/engine` full package suite green (~9083 tests)
- Targeted core/dashboard clusters green (schema applier, agent-runs UI,
settings descriptions, mobile close)
## Test plan
- [x] `pnpm test:gate` (twice)
- [x] `pnpm --filter @fusion/engine test`
- [ ] CI full suite / PR checks on this branch
- [ ] Confirm no unrelated product behavior changes beyond the listed
regression fixes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added support for `roadmap-item` native structure kinds, including
native structure embeds and metadata validation.
* Added Stable and Beta release channel options in General settings.
* Added per-action reporting target configuration with clearer “unset”
guidance.
* **Bug Fixes**
* Improved heartbeat/prompt behavior when patrol is disabled.
* Prevented deleted tasks from continuing through execution.
* Made recovery for explicit duplicate redirects more permissive.
* Hardened database migration and test database cleanup to reduce flaky
failures.
* **Documentation**
* Updated settings text for release channels, reporting targets, and
inheritance/unset behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Store-open adoption runs as fusion_runtime, which lacked grants on
public.fusion_schema_migrations, so the drained-marker write failed every
boot. Migration 0032 grants SELECT plus a SECURITY DEFINER helper limited
to the exact marker, and store-open calls that helper instead of raw INSERT.
## Summary
The Coding (Ideas) workflow now behaves like the board it presents:
Ideas stays inert, Todo owns planning and plan review, In progress owns
implementation, and In review owns code review and merge. The restored
preset is intentionally limited to that five-stage path, while the
existing Coding workflow remains unchanged.
Workflow execution now suspends at Todo→In progress instead of running
the implementation node early. A durable, single-owner continuation
records the exact resume node and survives process restarts; the
scheduler remains the only component allowed to admit the task into WIP.
Disabled optional review groups traverse the same boundary without
invoking a reviewer, avoiding the prior stuck-task behavior.
Workflow validation also rejects capacity holds with no reachable WIP
destination, so deterministic lifecycle deadlocks fail at authoring time
rather than after a task is running.
Session-settled decisions carried from planning: columns are execution
invariants, scheduler-owned WIP admission is preserved, the existing
Coding (Ideas) preset is restored and simplified, and invalid release
topology is rejected (user-approved).
## Validation
- `pnpm lint`
- `pnpm verify:fast`
- `pnpm test:gate` (296 engine, 128 PostgreSQL core, and 63 CI-shape
tests)
- Focused workflow lifecycle tests (106 assertions)
- PostgreSQL regression coverage proves atomic continuation replacement
and database rejection of a second active owner
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added durable, resumable workflow execution across capacity boundaries
(including explicit suspend/resume at the correct node).
* Introduced Todo “plan review” workflow continuations and automated
planning/capacity draining.
* Restored Coding (Ideas) as a selectable built-in and updated its lane
placement; improved optional-step group enablement support.
* **Bug Fixes**
* User moves back to Todo now cancels active workflow continuations.
* Rejected workflow boundary transitions now surface as errors (instead
of silently continuing).
* Workflows with undriveable capacity-hold configurations are now
rejected.
* **Tests / Data**
* Expanded coverage for workflow suspension, continuations, and
continuation replacement; updated database schema to persist
continuation metadata and enforce single active continuation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Chat messages, chat room messages, and agent/user mailbox sends could
crash mid-conversation when the persisted content or metadata contained
a raw U+0000 (NUL) byte — e.g. Windows CLI diagnostic/tool output piped
directly into a message body. PostgreSQL text/jsonb columns reject NUL
outright (`unsupported Unicode escape sequence` / `\u0000 cannot be
converted to text`), which surfaced as an uncaught `PostgresError` that
aborted the write and killed the conversation turn.
A NUL-byte sanitizer already existed for the one-time SQLite →
PostgreSQL first-boot migration (`sqlite-migrator.ts`'s
`stripNulChars`/`deepStripNulChars`), but it was never wired into the
**live** write paths — only into that one-shot migration.
## What changed
- Extracted `stripNulChars`/`deepStripNulChars` into a shared
`packages/core/src/postgres/nul-sanitize.ts` module
(`sqlite-migrator.ts` now imports from it instead of defining its own
copy).
- Wired sanitization into the three live write paths that persist
free-form content/metadata:
- `async-chat-store.ts`: `addChatMessage`, `addChatRoomMessage`
- `async-message-store.ts`: `sendMessage`
- Each of these functions now also **returns the sanitized value** —
previously they returned the original, unsanitized input object even
though the sanitized value is what was actually persisted to the
database, which was a latent inconsistency I found while adding test
coverage.
## Bonus fix: embedded-Postgres startup race
While rebuilding and testing this locally via `pnpm smoke:boot`, I hit a
separate, pre-existing, reproducible race: a process joining an existing
embedded-Postgres data dir (via `postmaster.pid`, per the existing
`FNXC:PostgresStartupRace 2026-07-15-20:45` comment in
`embedded-lifecycle.ts`) can race the true owner's TCP listener bind and
get `ECONNREFUSED` on its very first connection attempt.
`bootSchemaBackendOnce` turned this into a hard `startup-factory: failed
to initialize PostgreSQL schema backend` failure with no retry.
I verified this is **not** caused by my NUL-sanitize change — it
reproduces identically on unmodified `main` (confirmed via `git stash`).
Added `JoinedInstanceUnreachableError` and one retry (mirroring the
existing `NonUtf8EmbeddedClusterError` one-retry pattern already in the
same file) instead of failing the whole boot outright.
## Tests
- New unit tests for the shared sanitizer:
`packages/core/src/__tests__/nul-sanitize.test.ts` (10 tests, including
a regression test reproducing the exact production failure signature).
- New PostgreSQL integration test coverage in the existing `.pg.test.ts`
suites, reproducing the exact production failure payload for both
`addChatMessage` and `sendMessage` and asserting both the in-memory
return value and the re-read-from-database value are NUL-free.
- Verified end-to-end against a real, disposable PostgreSQL 16 instance
(outside the vitest harness, since this dev machine lacked a local
`psql`/`pg_dump` client at the time) using a standalone script that
calls the actual patched functions with the production crash payload —
all checks passed before and after the return-value fix was added.
- `pnpm --filter @fusion/core typecheck` clean.
## Changeset
Included (`patch`, category `fix`).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Prevented crashes and PostgreSQL insertion failures when chat or
mailbox content/JSON metadata contains raw NUL (`U+0000`) bytes.
* NUL characters are now stripped from message text and deeply from
nested metadata (including JSON object keys) before writes, and
sanitized values are reflected in returned messages.
* Improved embedded PostgreSQL startup reliability by retrying once on
transient joined-instance connection-refused failures.
* **Tests**
* Added unit and PostgreSQL regression coverage for NUL sanitization
across message/chat paths and for the embedded startup retry scenario.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Grant the restricted runtime role read-only access to its own SQLite cutover marker. Repair existing databases with migration 0030 and apply the same row-scoped policy when first-boot migration creates the ledger.
## Summary
PostgreSQL runtime roles without `CREATE` permission on `public` no
longer trigger schema writes during migration-marker health reads, so
`permission denied for schema public` is not mislabeled as database
corruption. Once connectivity and task-ID integrity pass, an unavailable
migration marker is treated as advisory instead of making the whole
database unhealthy. Dashboard and notification guidance now describes a
PostgreSQL health failure accurately and renders actionable log and
recovery links in every supported locale.
## Validation
- 54 targeted tests passed across core, dashboard, engine, and i18n.
- Typechecks passed for all four affected packages.
- Scoped ESLint, strict changeset validation, and diff checks passed.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* PostgreSQL health failures are now reported as degraded health checks
rather than database corruption.
* Migration-status lookup failures no longer incorrectly mark an
otherwise healthy database as unhealthy.
* Migration-state checks are now read-only and avoid creating or
modifying database structures.
* **UI & Localization**
* Updated database health banner messaging and recovery guidance across
supported languages.
* The banner now appears for broader PostgreSQL health failures and
links to storage documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Part **1 of 3** of the IR-driven lifecycle cutover (split from #2335 to
fit review-tool file limits; plan:
docs/plans/2026-07-18-001-refactor-ir-driven-lifecycle-cutover-plan.md,
included here).
**Scope (48 files, packages/core + docs/plans):** shared transition
policy + validator (KTD-5), IR validation hardening incl. the benchmark
capability floor, CAS review leases (KTD-4), pooled WIP capacity budgets
(KTD-9), lifecycle-trait helpers, durable IR pin/drift detection
(KTD-3), review-level creation-time preset, legacy adoption module +
census + migration 0026 + stale-binary guard (KTD-8), core-side builtin
workflow fixes (single default-IR authority, no-merge complete-column
support).
Note: `workflow-cutover.ts` (interpreter parity scaffolding) stays alive
in this PR — its last consumer dies in part 2/3, which retires it.
**Merge order:** this PR → #TBD-2 (engine) → #2335 (dashboard/top).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added workflow-trait-driven task lifecycle transitions, including WIP
capacity pooling and workflow-aware recovery (IR pinning + drift
detection).
* Added legacy adoption/backfill for pre-cutover task states, with
unmappable rows safely parked.
* Added create-time `reviewLevel` presets to automatically configure
enabled workflow steps.
* **Bug Fixes**
* Fixed workflow moves when no workflow selection exists.
* Improved merge-blocker validation to be keyed to the workflow’s actual
review-lane identity, preventing invalid moves and misclassified
terminal states.
* **Tests**
* Added end-to-end and unit/integration coverage for workflow
validation, legacy adoption, migrations/schema guards, leases, review
presets, and transition rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
SQLite INTEGER is effectively int64, but the PostgreSQL baseline mapped
unbounded token/usage counters on `project.tasks` and
`project.chat_token_usage` to `integer` (int4). Real data contains
values > 2,147,483,647, causing the SQLite-to-PostgreSQL migration to
fail with `value ... is out of range for type integer`.
Changes:
- Change baseline DDL to `bigint` for the affected columns.
- Update Drizzle schema to `bigint({ mode: "number" })` to preserve JS
`number` semantics.
- Add forward migration `0024_bigint_counters.sql` for existing
clusters.
- Bump `SCHEMA_BASELINE_VERSION` to `0024`.
Fixes the int4 overflow observed during migration of large token/usage
counters.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Expanded token-usage and activity/lease counters to 64-bit integers to
prevent overflow on large workloads.
* Improved distributed task ID state/reservations to be isolated per
project and to merge/update conflicting entries more reliably.
* **Chores**
* Added an idempotent PostgreSQL migration for bigint counter support
and advanced schema baseline tracking.
* Updated dashboard build support by adding `html2canvas` type
definitions and the production dependency.
* **Tests**
* Updated schema-applier migration checks to include the new
bigint-counters baseline identity.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
Prevent automated tests from inheriting production PostgreSQL URLs and route global test-mode startups to a dedicated external or embedded test database.
## Summary
Concurrent PostgreSQL project initialization no longer causes transient
dashboard failures, including repeated `GET /api/remote/status` 500
responses. The failure was a database deadlock between project-row
identity promotion and schema/plugin DDL, which previously acquired
overlapping locks in inconsistent orders.
This establishes one advisory-lock order across SQLite cutover, project
identity promotion, and schema mutations. Focused regression coverage
proves schema DDL waits behind an active migration transaction and that
identity stamping acquires the migration lock before reading
project-owned tables.
## Validation
- 25 focused unit tests passed.
- 3 focused real-PostgreSQL regression tests passed.
- `@fusion/core` typecheck passed.
- Strict changeset validation passed.
- Fast workspace verification passed, including the CLI build and boot
health check.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Prevented transient dashboard failures caused by PostgreSQL startup
and migration deadlocks.
* Improved serialization when multiple projects initialize or update
database schemas concurrently.
* Ensured migration state updates and schema changes occur in a
consistent order.
* **Tests**
* Added coverage for migration lock ordering, concurrent schema
operations, and recovery after lock contention.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Read the live port from PostgreSQL's actual postmaster.pid field so extension TaskStore boot reuses the existing server instead of wedging on a colliding start.
- uninstaller: taskkill only the first, digits-only postmaster.pid line
(the for /f loop ran taskkill on the port/epoch lines — potential
unrelated-process kill)
- git-missing dialogs use new ConfirmOptions.alwaysAsk so global
skip-confirmations cannot silently pick an unseen choice
- Windows quit prompt: embedded-local runtimes only, skipped during OS
session end (sync dialog blocked Windows shutdown)
- 'leave it running' detaches the embedded lifecycle (disarms its
process shutdown hook) so Electron exit cannot kill the postmaster
the operator chose to keep (new detachKeepingEmbedded)
- wizard: ref-based double-submit guard around the async git preflight
- clone route: ENOENT invalidate-and-retry matching runGitCommand
- openExternalUrl: drop the async window.open fallback (always
popup-blocked); log bridge failures instead
- DirectoryPicker: close the panel when listing the created folder
fails so Select cannot re-commit the parent
- git status probe bounded to two spawns (PATH + first candidate)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Users whose embedded cluster was initdb'd with an OS-locale encoding by
a pre-fix version now self-heal with zero manual steps: on the
encoding-conversion schema failure the startup factory proves the
cluster is non-UTF-8 AND empty (the baseline transaction never applied,
so no schema or migrated data can exist) and that this process owns the
postmaster, then deletes the data dir and reboots once with the UTF-8
initdb defaults. Joined instances and unproven states keep the manual
re-init hint; one retry ever, so no loops.
Verified on the elevated windows-latest runner: CI seeds a real WIN1252
cluster via initdb and proves a stock 'fn serve' auto-recovers it to a
healthy /api/health (run 29633351848, all jobs green). Also caps the
desktop-windows embedded-PG smoke at 30 min and adds a skip input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Squash of feature/win-elevated-no-user, verified end-to-end on the
elevated windows-latest runner (restricted-token double boot + full
'fn serve' /api/health smoke, both green).
- Elevated Windows boots embedded PostgreSQL via pg_ctl's built-in
restricted-token re-exec instead of creating a 'fusion-pg' local user
(operator requirement: Fusion must never create accounts). Removes
the credential launcher, icacls grants, and cmd/PowerShell wrapper —
and with them the 'directory name is invalid' and wrapper-log EBUSY
field failures. Leftover fusion-pg accounts are deleted on start.
- Embedded clusters are always initdb'd --encoding=UTF8 --locale=C
(GitHub issue #2286: OS-locale WIN1252/WIN1254 clusters could not
store the UTF-8 schema and crash-looped the dashboard). Existing
non-UTF-8 clusters get an actionable re-init hint at boot.
- Schema-backend boot failures now surface the full error cause chain
(DrizzleQueryError hid the real PostgresError behind the SQL text).
- Elevated stop() waits until the port closes and postmaster.pid is
gone before resolving.
- CI: branch verification workflow (restricted-token proof + elevated
boot smoke + account-absence assertions); boot-smoke stderr tail
widened for diagnosability.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Start-Process -Credential (CreateProcessWithLogonW) validates the working
directory as the TARGET user. The launcher inherited the desktop app's cwd
(admin profile / install dir), which the dedicated fusion-pg user cannot
read, so elevated desktop boots died with "The directory name is invalid"
before postgres ever started. launch.ps1 now pins -WorkingDirectory to the
.pgrunner run dir inside the data dir the user was just granted full
control on. CI never caught it because runner cwds are world-traversable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The one-time SQLite→PostgreSQL migration runs inside createTaskStoreForBackend
before any HTTP server listens, so browsers saw "connection refused" and open
tabs failed silently for minutes. Now:
- CLI: a temporary holding server binds the dashboard port for the boot window,
serving an auto-reloading "Database migration in progress" page and an
/api/health payload with status "migrating" + structured progress; the port
is handed off (awaited) to the real app.listen().
- Dashboard SPA: already-open tabs render the new MigrationInProgressBanner
from the 15s health poll when status is "migrating".
- Desktop: LocalRuntimeManager publishes migration progress on
DesktopRuntimeStatus via the new core onMigrationProgress option;
DesktopLaunchGate shows the live label and extends its 30s startup timeout
while progress advances (2min stall cap), in both boot and first-run flows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Legacy SQLite databases can hold U+0000 in TEXT cells and inside stored
JSON, which PostgreSQL rejects in text and jsonb columns and which
aborted the first-boot auto-migration. Strip NUL from plain text cells,
JSON string values and object keys, malformed-JSON scalars, and opaque
legacy-preservation cells; content-checksum verification compares the
sanitized source against the sanitized target so migrations still verify.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
shared_memory_type=mmap (defaulted 2026-07-16 for SysV shm exhaustion)
is invalid on Windows — PostgreSQL only accepts "windows" there and
dies with FATAL invalid value for parameter before opening the port.
Every Windows embedded start broke, failing the Windows release smoke
in both the v0.70.0 and v0.70.1 tag runs. Default flags now come from
defaultEmbeddedPostgresFlagsFor(platform): empty on win32 (no override
needed; SysV exhaustion cannot occur there), mmap elsewhere. Regression
test asserts the per-platform flag invariant.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The bun-compiled exe has been unbootable since the PG cutover: bun
standalone binaries do no node_modules resolution, so the deliberately
out-of-graph require("embedded-postgres") failed from /$bunfs, and
readFile'd migration .sql files were never embedded, so even external
DATABASE_URL mode died at schema init.
- schema-applier: resolveMigrationsDir() — FUSION_MIGRATIONS_DIR env >
module-relative dist/migrations (npm/desktop, unchanged) >
execPath-relative migrations/ (standalone exe), probe-based.
- embedded-lifecycle: require("embedded-postgres") first (npm/desktop
untouched), falling back to a self-contained staged bundle at
<execDir>/runtime/<platform>/embedded-postgres/dist/index.cjs
(FUSION_EMBEDDED_PG_RUNTIME_DIR override) with the native
initdb/pg_ctl/postgres payload beside it.
- build.ts: stage dist/migrations plus the per-target embedded-postgres
bundle + native payload (warn when a cross-target payload is absent on
the host, mirroring desktop's verifyEmbeddedPostgresPayloads).
- release.yml: package fn-cli-<os>-<arch>.tar.gz (binary + migrations +
runtime + client) with sha256 per leg; prune staged payload files from
the release-collection globs; bare fn-cli-* binaries still uploaded.
E2E-verified on the compiled binary: embedded mode initdb→/api/health
200 database healthy; DATABASE_URL mode applied migrations 0000–0019
(109 tables). Core typecheck clean; schema-applier 58/58 and
embedded-lifecycle 44/44 tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PR #2260 added project.tasks.bulk_completion_refusal_at to the Drizzle model
and the 0000 baseline but shipped no forward migration. Databases created
before #2260 already carry the 0000 marker, so the applier skips the baseline
and they never gained the column — every such cluster crashed on the first
TaskStore SELECT ("column bulk_completion_refusal_at does not exist"), taking
down dashboard/app boot.
Adds forward migration 0018 (wired via BULK_COMPLETION_REFUSAL_AT_VERSION;
SCHEMA_BASELINE_VERSION -> "0018") so existing clusters heal on next startup.
Prevention:
- Per-column upgrade regression test reproducing the exact existing-DB failure.
- Migration-wiring-integrity guard (no PostgreSQL): SCHEMA_BASELINE_VERSION must
equal the highest migration file, and every .sql must be registered in the
applier so none silently never runs.
- Repairs 6 pre-existing schema-applier tests left stale by the 0017 addition
(baseline-marker identity + version-list enumerations).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
## What & why
**FN-8141 laundered a failed task into `done` with zero net changes and
no sign-off.** After the executor's
`bulk-step-completion-without-review` refusal fired (steps had no
APPROVE verdicts), the agent used the sanctioned skip affordance
(`fn_task_update status="skipped"`) on the remaining unreviewed steps.
Because every completion check counts `skipped` as complete, the task
then satisfied the exact condition the refusal was protecting, and
downstream **automatic** promotion (implicit `fn_task_done`,
self-healing `recoverStrandedCompletedTodoTasks`) moved it to in-review
— where the AI merger found an empty diff and finalized it as a no-op
`done`.
This PR restores the invariant: **steps skipped while a
bulk-step-completion refusal marker is active on the task are "tainted"
and cannot carry the task to review through any automatic path.** The
taint clears on an honest exit — an accepted `fn_task_done` (explicit or
non-tainted implicit) or an operator manual retry — so the legitimate
`PREMISE STALE` skip-then-done flow is unaffected.
## Design
- **Persisted marker**: new nullable `Task.bulkCompletionRefusalAt` (ISO
timestamp), stamped when the `bulk-step-completion-without-review`
refusal fires (explicit `fn_task_done` handler + implicit
`handleImplicitTaskDoneRefusal`). Survives requeue so a refusal on
attempt N taints attempt N+1's promotion. Full store plumbing (types,
descriptors, serialization, SQLite/PG schema + health self-heal).
- **Pure evaluator** `evaluateSkipBypassTaint(task)` in `@fusion/core`
(next to `evaluateNoCommitsNoOpFinalize`): `blocked` iff the marker is
set AND ≥1 step is `skipped`. Single rule every AUTO-promotion check
calls.
- **Clearing**: accepted explicit `fn_task_done`, accepted
implicit/retry completion (the success-reset `updateTask`s), and
`buildManualRetryResetPatch` (operator retry). A fresh lifecycle that
genuinely re-does the work leaves zero skipped steps, so it is never
blocked even if a marker lingers.
## Surface enumeration (every consumer of "all steps done/skipped" that
gates AUTO-promotion)
- **executor.ts**: `getCompletedTaskFinalizationDecision` (gated on the
`isTaskWorkComplete` branch only, never on an accepted `taskDone`);
`recoverCompletedTask` (shared chokepoint for unpause resume,
completed-task watchdog, orphan resume);
`evaluateImplicitCompletionRefusal` (both implicit-completion loops);
`isTaskAlreadyCompleteForNonContinuableSession`; graph merge-boundary
`getWorkflowMergeImplementationProofFailure`.
- **self-healing.ts**: `recoverCompletedTasks` (stuck in-progress) and
`recoverStrandedCompletedTodoTasks` (the exact FN-8141 promoter).
- **Verified-safe, left as-is**: per-step graph node projections
(executor ~6274/6298) and progress-render checks — they don't gate
whole-task auto-promotion.
## Test evidence
Scoped runs (all green):
```
CORE: pnpm --filter @fusion/core exec vitest run \
src/__tests__/skip-bypass-taint-guard.test.ts \
src/__tests__/skip-bypass-taint-persistence.test.ts \
src/__tests__/manual-retry-reset.test.ts
→ 17 passed
ENGINE: pnpm --filter @fusion/engine exec vitest run \
src/__tests__/executor-skip-bypass-taint.test.ts \
src/__tests__/self-healing.test.ts
→ 401 passed
```
Coverage: pure-evaluator (skip-before-refusal counts, skip-after-refusal
doesn't, taint-clearing, empty-marker/empty-steps edges); store
round-trip of the marker (set→read→clear); executor white-box (implicit
completion refused when tainted, allowed when clean or fully re-done,
graph merge-boundary reports missing proof, and the **explicit
`fn_task_done` PREMISE-STALE honest exit stays accepted**); self-healing
(FN-8141 sequence does not promote from either recovery path; a clean
legitimately-skipped task still promotes); manual-retry clears the
marker.
## Note on `pnpm verify:fast`
`verify:fast` currently fails at the workspace-artifact bootstrap on
**pre-existing** pi-SDK type errors in
`packages/engine/src/{auth-storage,pi,provider-registration}.ts` — the
FN-8145 upstream migration breakage (pi 0.80.x removed
`AuthStorage`/`ModelRegistry.create`). **None of those files are in this
diff.** `@fusion/core` builds clean (`packages/core build: Done`), and
`@fusion/engine` `tsc` reports **no errors in the files this PR
touches** (`executor.ts`, `self-healing.ts`); the only engine build
errors are the FN-8145 files. This base failure is the same condition
FN-8141 describes and is out of scope for this task.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus <noreply@anthropic.com>
## Problem
Migration 0006 made `project_id` the RLS isolation partition on every
`project`-schema table — stamped by a BEFORE INSERT trigger from the
`fusion.project_id` session GUC, with every PK/unique/FK rewritten to
composite `(project_id, …)`. Eleven tables **also** carried a
caller-supplied domain `projectId` on their TS types and wrote that
domain value into the same physical column.
When the domain value differs from the session GUC, the parent row lands
in the domain partition while child rows (`research_run_events`,
`experiment_session_records`, `eval_task_results`, …) land in the
session partition — and the composite FK fails with SQLSTATE 23503.
Appending an event to a project-owned research run could not persist.
## Fix
**Decision (operator): separate domain column; `project_id` stays the
partition.**
- **Migration `0011_owner_project_id.sql`** adds a nullable
`owner_project_id` domain column to the 11 conflated tables
(`research_runs`, `experiment_sessions`, `todo_lists`, `eval_runs`,
`chat_sessions`, `chat_rooms`, `ai_sessions`, `chat_token_usage`,
`project_insights`, `project_insight_runs`, `cli_sessions`), backfills
it from `project_id` (identical in production, so exact; the
`__legacy_unscoped__` sentinel backfills to NULL), and indexes it.
Idempotent, `to_regclass`-guarded per the 0007 pattern.
- **Stores** (`async-research-store`, `async-experiment-session-store`,
`async-todo-store`, `async-chat-store`, `async-ai-session-store`,
`async-eval-store`, `async-insight-store`, `cli-session-store`, …) stop
writing `project_id` entirely — the trigger/GUC owns the partition — and
map their domain `projectId` field to `owner_project_id` for both reads
and filters. TS types unchanged.
- **Applier** registers `OWNER_PROJECT_ID_SPLIT_VERSION = "0011"` and
advances `SCHEMA_BASELINE_VERSION`.
## Verification (re-run independently of the implementing agent)
- Core `tsc --noEmit`: exit 0 · `pnpm lint`: exit 0 · `pnpm
check:changesets`: exit 0 · `pnpm test:gate`: 185/185
- Full postgres suite: **5 failed / 807 passed** vs a **7 / 804**
baseline — the two conflation round-trips
(`satellite-db-injected-stores` ResearchStore + ExperimentSessionStore)
go green, zero new failures. The remaining 5 are pre-existing
unbound-harness `__meta`/identity failures, unrelated to this change.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Corrected project-scoped persistence and queries across AI sessions,
chats (rooms + token usage), evaluations/experiments, insights,
research, and todos by separating domain ownership from RLS
partitioning.
* Prevented foreign-key and row-level security violations when storing
or retrieving project-scoped data, including legacy records.
* **Database / New Features**
* Added migration 0011 introducing `owner_project_id` and backfilling
existing rows to preserve ownership while improving isolation.
* **Tests**
* Updated migration-parity coverage to include the new baseline step.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Six of the eight postgres-suite failures shared one root cause: writes
normalize project_id, reads did not. The fusion_assign_project_id trigger
(migration 0006) rewrites a blank project_id to the session's fusion.project_id
or '__legacy_unscoped__', but helpers reached as `layer.projectId ?? ""` then
filtered on the literal '' -- a value the database never stores. Every unbound
read missed rows it had just written.
AsyncDataLayer.projectId is optional by design (undefined = project-agnostic),
so `?? ""` is the bug: it turns "no scope" into a scope that matches nothing.
The resolution differs by what the rows are, and conflating them corrupts data:
- Data and analytics reads (usage events, agent runs, research runs) take
projectScopeFor(): a bound id filters, an unbound one reads across projects.
This matches the contract taskProjectScope already documents ("when undefined
the scope filter is a no-op").
- __meta migration guards (project-identity stamps, agent-store markers) take
projectPartitionId(): an unbound id resolves to the shared sentinel
partition. projectScopeFor would be wrong here -- dropping the predicate lets
an unbound getMetaValue return whichever project's marker it finds first, so
on the shared cluster project A's "migration complete" marker would tell
project B to skip a migration it never ran. upsertMetaValue already documented
this: "the empty binding remains the explicit project-agnostic compatibility
partition". Writing the sentinel explicitly also keeps the partition
deterministic -- a blank write from a session carrying fusion.project_id would
otherwise land in that project's stamp.
Names the sentinel (LEGACY_UNSCOPED_PROJECT_ID) instead of open-coding it, and
puts both helpers next to taskProjectScope so the convention has one home.
Fixes taskstore-remaining (24/24), project-identity (6/6), and
satellite-fusiondir-stores (16/16).
The remaining two failures are a different bug and are NOT addressed here: the
child tables research_run_events and experiment_session_records never declared
project_id in schema-as-code, though migration 0006 added the column and
rewrote their FKs to composite (project_id, parent_id). Drizzle therefore cannot
write the parent's partition, the trigger stamps '__legacy_unscoped__', and the
FK fails against a project-owned parent. That needs a schema-as-code change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An unbound (project-agnostic) data layer read zero usage events it had just
written. AsyncDataLayer.projectId is optional by design -- undefined means a
project-agnostic layer for single-project / global / analytics reads -- but
helpers taking `projectId: string` are called as `layer.projectId ?? ""`, which
turns "no scope" into a literal '' scope.
'' never matches: the fusion_assign_project_id BEFORE INSERT trigger (migration
0006) rewrites a written '' to the session's fusion.project_id or
'__legacy_unscoped__', so a read filtering on '' looks for a value the database
never stores. Writes normalize, reads did not. Proven by probe: the row is
present with project_id '__legacy_unscoped__', emitUsageEvent returns true, and
queryUsageEvents returns [] even with no other filters.
Treat blank as unbound and drop the scope predicate, matching the contract
taskProjectScope already documents ("when undefined the scope filter is a
no-op"). Restricting an unbound reader to '__legacy_unscoped__' rows instead
would make an unscoped analytics read silently partial.
Adds projectScopeFor() next to taskProjectScope so the convention has one home
rather than a third open-coded variant.
Note the write path is already live: remaining-ops-7.ts emits with
`layer.projectId ?? ""` under backendMode, so unscoped events are accumulating
under the sentinel today. The async reader has no production caller yet, which
is why nothing user-facing broke.
Fixes taskstore-remaining.test.ts (24/24). The remaining failures in that suite
share this root cause but not this resolution -- see the follow-up.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two related leaks on the embedded Postgres startup-race join.
The flagged one: the catch dropped `nonAdminHandle` to null without stopping
it, so a wrapper that onLaunched had already published leaked. The obvious fix
-- call handle.stop() first -- is worse than the leak. stop() runs killAll(),
which resolves its target by reading line 1 of the data dir's postmaster.pid.
On this path that file belongs to the process that WON the race, so stop()
would taskkill the instance we are joining. pg.stop() is the same trap via
pg_ctl -D on the shared dir, which is why settleCancelledStart (it calls both)
cannot be reused here. Added NonAdminServerHandle.stopWrapperOnly(), which
kills only our wrapper pid and its children, and called it before the handle is
dropped. A racing winner is another process's child, so /t cannot reach it.
The one found while making that safe: the catch joined on ANY start failure. A
start that took the lock and then failed later (readiness timeout, non-admin
poll error) reads back its OWN postmaster.pid, so isAlreadyRunning hands back
our own port and we "join" ourselves with ownsProcess=false -- nothing ever
stops it, orphaning a live postmaster for the life of the host. The join now
fires only on a lock-collision error, which is the one failure proving our
postgres refused to start and someone else owns the dir. Every other failure
returns to the existing cancellation/cleanup paths, which stop what they
started. That is also what makes the wrapper-only kill provably safe: on this
path our postgres never took the lock.
Tests: a non-lock failure must propagate even with a postmaster.pid present
(fails without the fix -- the old catch swallowed it and joined), and a lock
collision must still join. Both always-on with a mocked ctor.
Pre-existing and unrelated: taskstore-remaining.test.ts fails identically on a
clean tree with these changes stashed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>