After a hard host crash (SIGKILL, power loss), postmaster.pid survives with no
postmaster behind it. The optimistic join handed every subsequent boot a URL to
the dead port, so the dashboard could never start again without a manual pid
delete. Probe the recorded pid with signal 0: provably dead (ESRCH) rebuts the
live-lock presumption and the boot takes an owned start — PostgreSQL itself
re-validates and reclaims the stale lock file, so a recycled live pid keeps the
old join-then-fail behavior and a genuinely live postmaster still surfaces the
lock collision we already join on. EPERM counts as alive (fail-closed).
Verified end to end: real cluster started, postmaster SIGKILLed leaving the pid
file + interrupted WAL, fresh lifecycle detected the stale lock, ran an owned
start, and crash recovery preserved the marker row.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
beta.4 follow-up from the issue thread: on an interrupted (not-cleanly-shut-down)
cluster, the elevated Windows launcher declared readiness on a bare TCP accept
while crash recovery still rejected every connection with 57P03, so
ensureDatabase failed and the start cleanup fast-shutdown the recovering
postmaster ~0.2s after launch; the retry then joined the instance it had just
told to stop, got ECONNREFUSED, and parked the dashboard in a dead shell. The
30s ".pgrunner sharing violation" stall was recovery's SyncDataDirectory fsync
walk hitting Fusion's own pgctl log inside the data dir.
- Move the pgctl runner dir to a sibling .pgrunner-<dataDirName> outside the
data dir (and sweep the legacy in-dataDir .pgrunner), so recovery's fsync
walk can never contend with the postmaster's inherited log handle.
- Ignore 57P03 recovery rejections in the elevated readiness fatal scan.
- Owned starts wait for the cluster to genuinely accept connections (retrying
57P03/socket errors, bounded by the start timeout) before ensureDatabase —
never stop a postmaster that is still in recovery.
- Join-path database verify retries the 57P03 recovery signal for up to 15s;
socket errors keep the instant optimistic-join contract for stale pids.
- startup-factory's joined-instance-unreachable retry backs off across ~15s
instead of a single 500ms attempt.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## Summary
Two-part fix for #2411 — Windows embedded PostgreSQL backends dying with
exception `0xC0000142` and taking the whole dashboard down.
### 1. Crash hardening + recovery (FN-8522)
- Child-only native `PATH` hardening so forked backends can always
resolve their runtime DLLs.
- Non-blocking `.pgrunner` log monitoring (shared read), eliminating the
self-inflicted ~30s `sharing violation` retry window at boot.
- Detection of the ordered 0xC0000142 shutdown sequence with a single
automatic restart of owned clusters on their resolved port, plus
operator diagnostics.
### 2. Platform-aware `max_connections` default (follow-up from
[operator
report](https://github.com/Runfusion/Fusion/issues/2411#issuecomment-5054900702))
On Windows every PostgreSQL connection is a separate process; the
embedded cluster's unconfigured `max_connections=500` cap lets backend
spawn bursts exhaust the non-interactive desktop heap, which kills
forked backends with exactly `0xC0000142`. The reporter confirmed
stability after lowering the cap.
- `embeddedPostgresMaxConnections` is now schema-unset so the server can
distinguish "operator never set it" from an explicit choice
(`getSettings()` merges schema defaults, which previously pinned 500
unconditionally and made the runtime fallback dead code).
- New `resolveEmbeddedMaxConnections()` resolves the unset default
platform-aware: **150 on win32, 500 elsewhere**. Explicit settings are
honored on every platform, clamped to [32, 2000] as before.
- Settings UI renders the cap empty ("auto") with platform-aware help
copy across all six locales.
- Fixed a latent reset bug this exposed: global "Reset this menu" wrote
`undefined` for undefined-default keys, which JSON serialization drops —
the stored value silently survived reset. Now uses null-as-delete.
## Testing
- New unit tests for `resolveEmbeddedMaxConnections` (platform defaults,
clamping, non-integer handling).
- Updated settings-defaults, default-descriptions, and SettingsModal
tests; embedded lifecycle + recovery coverage from FN-8522.
- `@fusion/core` builds clean; changesets included for both parts.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Summary
Restores green merge-gate and package-default suites after repeated
`origin/main` merges brought workflow-graph ownership cutover drift into
CI.
- Align engine/dashboard/core tests with post-cutover contracts
(`moveTaskIf`/`deleteTaskIf`, graph handoff, worktree-pool reclaim via
`removeWorktree` + `RemovalReason`, multi-step RESUMING parse,
soft-pause merge requester, graph-terminal failure surfaces).
- Small product fixes needed for real regressions uncovered by the
suite: soft-delete refuse before graph routing, skip DUPLICATE
step-heading withhold when an explicit marker is present, PG schema
applier guards, and related bookkeeping (research promote tool inventory
/ migration seed, stop shell `psql` in PG admin DDL).
- Quarantine/ledger hygiene only where required by standing rules; no
timeout/worker appeasement.
## Verification
- `pnpm test:gate` ×2 green
- `@fusion/engine` full package suite green (~9083 tests)
- Targeted core/dashboard clusters green (schema applier, agent-runs UI,
settings descriptions, mobile close)
## Test plan
- [x] `pnpm test:gate` (twice)
- [x] `pnpm --filter @fusion/engine test`
- [ ] CI full suite / PR checks on this branch
- [ ] Confirm no unrelated product behavior changes beyond the listed
regression fixes
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added support for `roadmap-item` native structure kinds, including
native structure embeds and metadata validation.
* Added Stable and Beta release channel options in General settings.
* Added per-action reporting target configuration with clearer “unset”
guidance.
* **Bug Fixes**
* Improved heartbeat/prompt behavior when patrol is disabled.
* Prevented deleted tasks from continuing through execution.
* Made recovery for explicit duplicate redirects more permissive.
* Hardened database migration and test database cleanup to reduce flaky
failures.
* **Documentation**
* Updated settings text for release channels, reporting targets, and
inheritance/unset behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Store-open adoption runs as fusion_runtime, which lacked grants on
public.fusion_schema_migrations, so the drained-marker write failed every
boot. Migration 0032 grants SELECT plus a SECURITY DEFINER helper limited
to the exact marker, and store-open calls that helper instead of raw INSERT.
## Summary
The Coding (Ideas) workflow now behaves like the board it presents:
Ideas stays inert, Todo owns planning and plan review, In progress owns
implementation, and In review owns code review and merge. The restored
preset is intentionally limited to that five-stage path, while the
existing Coding workflow remains unchanged.
Workflow execution now suspends at Todo→In progress instead of running
the implementation node early. A durable, single-owner continuation
records the exact resume node and survives process restarts; the
scheduler remains the only component allowed to admit the task into WIP.
Disabled optional review groups traverse the same boundary without
invoking a reviewer, avoiding the prior stuck-task behavior.
Workflow validation also rejects capacity holds with no reachable WIP
destination, so deterministic lifecycle deadlocks fail at authoring time
rather than after a task is running.
Session-settled decisions carried from planning: columns are execution
invariants, scheduler-owned WIP admission is preserved, the existing
Coding (Ideas) preset is restored and simplified, and invalid release
topology is rejected (user-approved).
## Validation
- `pnpm lint`
- `pnpm verify:fast`
- `pnpm test:gate` (296 engine, 128 PostgreSQL core, and 63 CI-shape
tests)
- Focused workflow lifecycle tests (106 assertions)
- PostgreSQL regression coverage proves atomic continuation replacement
and database rejection of a second active owner
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added durable, resumable workflow execution across capacity boundaries
(including explicit suspend/resume at the correct node).
* Introduced Todo “plan review” workflow continuations and automated
planning/capacity draining.
* Restored Coding (Ideas) as a selectable built-in and updated its lane
placement; improved optional-step group enablement support.
* **Bug Fixes**
* User moves back to Todo now cancels active workflow continuations.
* Rejected workflow boundary transitions now surface as errors (instead
of silently continuing).
* Workflows with undriveable capacity-hold configurations are now
rejected.
* **Tests / Data**
* Expanded coverage for workflow suspension, continuations, and
continuation replacement; updated database schema to persist
continuation metadata and enforce single active continuation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Chat messages, chat room messages, and agent/user mailbox sends could
crash mid-conversation when the persisted content or metadata contained
a raw U+0000 (NUL) byte — e.g. Windows CLI diagnostic/tool output piped
directly into a message body. PostgreSQL text/jsonb columns reject NUL
outright (`unsupported Unicode escape sequence` / `\u0000 cannot be
converted to text`), which surfaced as an uncaught `PostgresError` that
aborted the write and killed the conversation turn.
A NUL-byte sanitizer already existed for the one-time SQLite →
PostgreSQL first-boot migration (`sqlite-migrator.ts`'s
`stripNulChars`/`deepStripNulChars`), but it was never wired into the
**live** write paths — only into that one-shot migration.
## What changed
- Extracted `stripNulChars`/`deepStripNulChars` into a shared
`packages/core/src/postgres/nul-sanitize.ts` module
(`sqlite-migrator.ts` now imports from it instead of defining its own
copy).
- Wired sanitization into the three live write paths that persist
free-form content/metadata:
- `async-chat-store.ts`: `addChatMessage`, `addChatRoomMessage`
- `async-message-store.ts`: `sendMessage`
- Each of these functions now also **returns the sanitized value** —
previously they returned the original, unsanitized input object even
though the sanitized value is what was actually persisted to the
database, which was a latent inconsistency I found while adding test
coverage.
## Bonus fix: embedded-Postgres startup race
While rebuilding and testing this locally via `pnpm smoke:boot`, I hit a
separate, pre-existing, reproducible race: a process joining an existing
embedded-Postgres data dir (via `postmaster.pid`, per the existing
`FNXC:PostgresStartupRace 2026-07-15-20:45` comment in
`embedded-lifecycle.ts`) can race the true owner's TCP listener bind and
get `ECONNREFUSED` on its very first connection attempt.
`bootSchemaBackendOnce` turned this into a hard `startup-factory: failed
to initialize PostgreSQL schema backend` failure with no retry.
I verified this is **not** caused by my NUL-sanitize change — it
reproduces identically on unmodified `main` (confirmed via `git stash`).
Added `JoinedInstanceUnreachableError` and one retry (mirroring the
existing `NonUtf8EmbeddedClusterError` one-retry pattern already in the
same file) instead of failing the whole boot outright.
## Tests
- New unit tests for the shared sanitizer:
`packages/core/src/__tests__/nul-sanitize.test.ts` (10 tests, including
a regression test reproducing the exact production failure signature).
- New PostgreSQL integration test coverage in the existing `.pg.test.ts`
suites, reproducing the exact production failure payload for both
`addChatMessage` and `sendMessage` and asserting both the in-memory
return value and the re-read-from-database value are NUL-free.
- Verified end-to-end against a real, disposable PostgreSQL 16 instance
(outside the vitest harness, since this dev machine lacked a local
`psql`/`pg_dump` client at the time) using a standalone script that
calls the actual patched functions with the production crash payload —
all checks passed before and after the return-value fix was added.
- `pnpm --filter @fusion/core typecheck` clean.
## Changeset
Included (`patch`, category `fix`).
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Prevented crashes and PostgreSQL insertion failures when chat or
mailbox content/JSON metadata contains raw NUL (`U+0000`) bytes.
* NUL characters are now stripped from message text and deeply from
nested metadata (including JSON object keys) before writes, and
sanitized values are reflected in returned messages.
* Improved embedded PostgreSQL startup reliability by retrying once on
transient joined-instance connection-refused failures.
* **Tests**
* Added unit and PostgreSQL regression coverage for NUL sanitization
across message/chat paths and for the embedded startup retry scenario.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Grant the restricted runtime role read-only access to its own SQLite cutover marker. Repair existing databases with migration 0030 and apply the same row-scoped policy when first-boot migration creates the ledger.
## Summary
PostgreSQL runtime roles without `CREATE` permission on `public` no
longer trigger schema writes during migration-marker health reads, so
`permission denied for schema public` is not mislabeled as database
corruption. Once connectivity and task-ID integrity pass, an unavailable
migration marker is treated as advisory instead of making the whole
database unhealthy. Dashboard and notification guidance now describes a
PostgreSQL health failure accurately and renders actionable log and
recovery links in every supported locale.
## Validation
- 54 targeted tests passed across core, dashboard, engine, and i18n.
- Typechecks passed for all four affected packages.
- Scoped ESLint, strict changeset validation, and diff checks passed.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* PostgreSQL health failures are now reported as degraded health checks
rather than database corruption.
* Migration-status lookup failures no longer incorrectly mark an
otherwise healthy database as unhealthy.
* Migration-state checks are now read-only and avoid creating or
modifying database structures.
* **UI & Localization**
* Updated database health banner messaging and recovery guidance across
supported languages.
* The banner now appears for broader PostgreSQL health failures and
links to storage documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Part **1 of 3** of the IR-driven lifecycle cutover (split from #2335 to
fit review-tool file limits; plan:
docs/plans/2026-07-18-001-refactor-ir-driven-lifecycle-cutover-plan.md,
included here).
**Scope (48 files, packages/core + docs/plans):** shared transition
policy + validator (KTD-5), IR validation hardening incl. the benchmark
capability floor, CAS review leases (KTD-4), pooled WIP capacity budgets
(KTD-9), lifecycle-trait helpers, durable IR pin/drift detection
(KTD-3), review-level creation-time preset, legacy adoption module +
census + migration 0026 + stale-binary guard (KTD-8), core-side builtin
workflow fixes (single default-IR authority, no-merge complete-column
support).
Note: `workflow-cutover.ts` (interpreter parity scaffolding) stays alive
in this PR — its last consumer dies in part 2/3, which retires it.
**Merge order:** this PR → #TBD-2 (engine) → #2335 (dashboard/top).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added workflow-trait-driven task lifecycle transitions, including WIP
capacity pooling and workflow-aware recovery (IR pinning + drift
detection).
* Added legacy adoption/backfill for pre-cutover task states, with
unmappable rows safely parked.
* Added create-time `reviewLevel` presets to automatically configure
enabled workflow steps.
* **Bug Fixes**
* Fixed workflow moves when no workflow selection exists.
* Improved merge-blocker validation to be keyed to the workflow’s actual
review-lane identity, preventing invalid moves and misclassified
terminal states.
* **Tests**
* Added end-to-end and unit/integration coverage for workflow
validation, legacy adoption, migrations/schema guards, leases, review
presets, and transition rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
SQLite INTEGER is effectively int64, but the PostgreSQL baseline mapped
unbounded token/usage counters on `project.tasks` and
`project.chat_token_usage` to `integer` (int4). Real data contains
values > 2,147,483,647, causing the SQLite-to-PostgreSQL migration to
fail with `value ... is out of range for type integer`.
Changes:
- Change baseline DDL to `bigint` for the affected columns.
- Update Drizzle schema to `bigint({ mode: "number" })` to preserve JS
`number` semantics.
- Add forward migration `0024_bigint_counters.sql` for existing
clusters.
- Bump `SCHEMA_BASELINE_VERSION` to `0024`.
Fixes the int4 overflow observed during migration of large token/usage
counters.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Expanded token-usage and activity/lease counters to 64-bit integers to
prevent overflow on large workloads.
* Improved distributed task ID state/reservations to be isolated per
project and to merge/update conflicting entries more reliably.
* **Chores**
* Added an idempotent PostgreSQL migration for bigint counter support
and advanced schema baseline tracking.
* Updated dashboard build support by adding `html2canvas` type
definitions and the production dependency.
* **Tests**
* Updated schema-applier migration checks to include the new
bigint-counters baseline identity.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
Prevent automated tests from inheriting production PostgreSQL URLs and route global test-mode startups to a dedicated external or embedded test database.
## Summary
Concurrent PostgreSQL project initialization no longer causes transient
dashboard failures, including repeated `GET /api/remote/status` 500
responses. The failure was a database deadlock between project-row
identity promotion and schema/plugin DDL, which previously acquired
overlapping locks in inconsistent orders.
This establishes one advisory-lock order across SQLite cutover, project
identity promotion, and schema mutations. Focused regression coverage
proves schema DDL waits behind an active migration transaction and that
identity stamping acquires the migration lock before reading
project-owned tables.
## Validation
- 25 focused unit tests passed.
- 3 focused real-PostgreSQL regression tests passed.
- `@fusion/core` typecheck passed.
- Strict changeset validation passed.
- Fast workspace verification passed, including the CLI build and boot
health check.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Prevented transient dashboard failures caused by PostgreSQL startup
and migration deadlocks.
* Improved serialization when multiple projects initialize or update
database schemas concurrently.
* Ensured migration state updates and schema changes occur in a
consistent order.
* **Tests**
* Added coverage for migration lock ordering, concurrent schema
operations, and recovery after lock contention.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Read the live port from PostgreSQL's actual postmaster.pid field so extension TaskStore boot reuses the existing server instead of wedging on a colliding start.
- uninstaller: taskkill only the first, digits-only postmaster.pid line
(the for /f loop ran taskkill on the port/epoch lines — potential
unrelated-process kill)
- git-missing dialogs use new ConfirmOptions.alwaysAsk so global
skip-confirmations cannot silently pick an unseen choice
- Windows quit prompt: embedded-local runtimes only, skipped during OS
session end (sync dialog blocked Windows shutdown)
- 'leave it running' detaches the embedded lifecycle (disarms its
process shutdown hook) so Electron exit cannot kill the postmaster
the operator chose to keep (new detachKeepingEmbedded)
- wizard: ref-based double-submit guard around the async git preflight
- clone route: ENOENT invalidate-and-retry matching runGitCommand
- openExternalUrl: drop the async window.open fallback (always
popup-blocked); log bridge failures instead
- DirectoryPicker: close the panel when listing the created folder
fails so Select cannot re-commit the parent
- git status probe bounded to two spawns (PATH + first candidate)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Users whose embedded cluster was initdb'd with an OS-locale encoding by
a pre-fix version now self-heal with zero manual steps: on the
encoding-conversion schema failure the startup factory proves the
cluster is non-UTF-8 AND empty (the baseline transaction never applied,
so no schema or migrated data can exist) and that this process owns the
postmaster, then deletes the data dir and reboots once with the UTF-8
initdb defaults. Joined instances and unproven states keep the manual
re-init hint; one retry ever, so no loops.
Verified on the elevated windows-latest runner: CI seeds a real WIN1252
cluster via initdb and proves a stock 'fn serve' auto-recovers it to a
healthy /api/health (run 29633351848, all jobs green). Also caps the
desktop-windows embedded-PG smoke at 30 min and adds a skip input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Squash of feature/win-elevated-no-user, verified end-to-end on the
elevated windows-latest runner (restricted-token double boot + full
'fn serve' /api/health smoke, both green).
- Elevated Windows boots embedded PostgreSQL via pg_ctl's built-in
restricted-token re-exec instead of creating a 'fusion-pg' local user
(operator requirement: Fusion must never create accounts). Removes
the credential launcher, icacls grants, and cmd/PowerShell wrapper —
and with them the 'directory name is invalid' and wrapper-log EBUSY
field failures. Leftover fusion-pg accounts are deleted on start.
- Embedded clusters are always initdb'd --encoding=UTF8 --locale=C
(GitHub issue #2286: OS-locale WIN1252/WIN1254 clusters could not
store the UTF-8 schema and crash-looped the dashboard). Existing
non-UTF-8 clusters get an actionable re-init hint at boot.
- Schema-backend boot failures now surface the full error cause chain
(DrizzleQueryError hid the real PostgresError behind the SQL text).
- Elevated stop() waits until the port closes and postmaster.pid is
gone before resolving.
- CI: branch verification workflow (restricted-token proof + elevated
boot smoke + account-absence assertions); boot-smoke stderr tail
widened for diagnosability.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Start-Process -Credential (CreateProcessWithLogonW) validates the working
directory as the TARGET user. The launcher inherited the desktop app's cwd
(admin profile / install dir), which the dedicated fusion-pg user cannot
read, so elevated desktop boots died with "The directory name is invalid"
before postgres ever started. launch.ps1 now pins -WorkingDirectory to the
.pgrunner run dir inside the data dir the user was just granted full
control on. CI never caught it because runner cwds are world-traversable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The one-time SQLite→PostgreSQL migration runs inside createTaskStoreForBackend
before any HTTP server listens, so browsers saw "connection refused" and open
tabs failed silently for minutes. Now:
- CLI: a temporary holding server binds the dashboard port for the boot window,
serving an auto-reloading "Database migration in progress" page and an
/api/health payload with status "migrating" + structured progress; the port
is handed off (awaited) to the real app.listen().
- Dashboard SPA: already-open tabs render the new MigrationInProgressBanner
from the 15s health poll when status is "migrating".
- Desktop: LocalRuntimeManager publishes migration progress on
DesktopRuntimeStatus via the new core onMigrationProgress option;
DesktopLaunchGate shows the live label and extends its 30s startup timeout
while progress advances (2min stall cap), in both boot and first-run flows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Legacy SQLite databases can hold U+0000 in TEXT cells and inside stored
JSON, which PostgreSQL rejects in text and jsonb columns and which
aborted the first-boot auto-migration. Strip NUL from plain text cells,
JSON string values and object keys, malformed-JSON scalars, and opaque
legacy-preservation cells; content-checksum verification compares the
sanitized source against the sanitized target so migrations still verify.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
shared_memory_type=mmap (defaulted 2026-07-16 for SysV shm exhaustion)
is invalid on Windows — PostgreSQL only accepts "windows" there and
dies with FATAL invalid value for parameter before opening the port.
Every Windows embedded start broke, failing the Windows release smoke
in both the v0.70.0 and v0.70.1 tag runs. Default flags now come from
defaultEmbeddedPostgresFlagsFor(platform): empty on win32 (no override
needed; SysV exhaustion cannot occur there), mmap elsewhere. Regression
test asserts the per-platform flag invariant.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The bun-compiled exe has been unbootable since the PG cutover: bun
standalone binaries do no node_modules resolution, so the deliberately
out-of-graph require("embedded-postgres") failed from /$bunfs, and
readFile'd migration .sql files were never embedded, so even external
DATABASE_URL mode died at schema init.
- schema-applier: resolveMigrationsDir() — FUSION_MIGRATIONS_DIR env >
module-relative dist/migrations (npm/desktop, unchanged) >
execPath-relative migrations/ (standalone exe), probe-based.
- embedded-lifecycle: require("embedded-postgres") first (npm/desktop
untouched), falling back to a self-contained staged bundle at
<execDir>/runtime/<platform>/embedded-postgres/dist/index.cjs
(FUSION_EMBEDDED_PG_RUNTIME_DIR override) with the native
initdb/pg_ctl/postgres payload beside it.
- build.ts: stage dist/migrations plus the per-target embedded-postgres
bundle + native payload (warn when a cross-target payload is absent on
the host, mirroring desktop's verifyEmbeddedPostgresPayloads).
- release.yml: package fn-cli-<os>-<arch>.tar.gz (binary + migrations +
runtime + client) with sha256 per leg; prune staged payload files from
the release-collection globs; bare fn-cli-* binaries still uploaded.
E2E-verified on the compiled binary: embedded mode initdb→/api/health
200 database healthy; DATABASE_URL mode applied migrations 0000–0019
(109 tables). Core typecheck clean; schema-applier 58/58 and
embedded-lifecycle 44/44 tests pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PR #2260 added project.tasks.bulk_completion_refusal_at to the Drizzle model
and the 0000 baseline but shipped no forward migration. Databases created
before #2260 already carry the 0000 marker, so the applier skips the baseline
and they never gained the column — every such cluster crashed on the first
TaskStore SELECT ("column bulk_completion_refusal_at does not exist"), taking
down dashboard/app boot.
Adds forward migration 0018 (wired via BULK_COMPLETION_REFUSAL_AT_VERSION;
SCHEMA_BASELINE_VERSION -> "0018") so existing clusters heal on next startup.
Prevention:
- Per-column upgrade regression test reproducing the exact existing-DB failure.
- Migration-wiring-integrity guard (no PostgreSQL): SCHEMA_BASELINE_VERSION must
equal the highest migration file, and every .sql must be registered in the
applier so none silently never runs.
- Repairs 6 pre-existing schema-applier tests left stale by the 0017 addition
(baseline-marker identity + version-list enumerations).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
## What & why
**FN-8141 laundered a failed task into `done` with zero net changes and
no sign-off.** After the executor's
`bulk-step-completion-without-review` refusal fired (steps had no
APPROVE verdicts), the agent used the sanctioned skip affordance
(`fn_task_update status="skipped"`) on the remaining unreviewed steps.
Because every completion check counts `skipped` as complete, the task
then satisfied the exact condition the refusal was protecting, and
downstream **automatic** promotion (implicit `fn_task_done`,
self-healing `recoverStrandedCompletedTodoTasks`) moved it to in-review
— where the AI merger found an empty diff and finalized it as a no-op
`done`.
This PR restores the invariant: **steps skipped while a
bulk-step-completion refusal marker is active on the task are "tainted"
and cannot carry the task to review through any automatic path.** The
taint clears on an honest exit — an accepted `fn_task_done` (explicit or
non-tainted implicit) or an operator manual retry — so the legitimate
`PREMISE STALE` skip-then-done flow is unaffected.
## Design
- **Persisted marker**: new nullable `Task.bulkCompletionRefusalAt` (ISO
timestamp), stamped when the `bulk-step-completion-without-review`
refusal fires (explicit `fn_task_done` handler + implicit
`handleImplicitTaskDoneRefusal`). Survives requeue so a refusal on
attempt N taints attempt N+1's promotion. Full store plumbing (types,
descriptors, serialization, SQLite/PG schema + health self-heal).
- **Pure evaluator** `evaluateSkipBypassTaint(task)` in `@fusion/core`
(next to `evaluateNoCommitsNoOpFinalize`): `blocked` iff the marker is
set AND ≥1 step is `skipped`. Single rule every AUTO-promotion check
calls.
- **Clearing**: accepted explicit `fn_task_done`, accepted
implicit/retry completion (the success-reset `updateTask`s), and
`buildManualRetryResetPatch` (operator retry). A fresh lifecycle that
genuinely re-does the work leaves zero skipped steps, so it is never
blocked even if a marker lingers.
## Surface enumeration (every consumer of "all steps done/skipped" that
gates AUTO-promotion)
- **executor.ts**: `getCompletedTaskFinalizationDecision` (gated on the
`isTaskWorkComplete` branch only, never on an accepted `taskDone`);
`recoverCompletedTask` (shared chokepoint for unpause resume,
completed-task watchdog, orphan resume);
`evaluateImplicitCompletionRefusal` (both implicit-completion loops);
`isTaskAlreadyCompleteForNonContinuableSession`; graph merge-boundary
`getWorkflowMergeImplementationProofFailure`.
- **self-healing.ts**: `recoverCompletedTasks` (stuck in-progress) and
`recoverStrandedCompletedTodoTasks` (the exact FN-8141 promoter).
- **Verified-safe, left as-is**: per-step graph node projections
(executor ~6274/6298) and progress-render checks — they don't gate
whole-task auto-promotion.
## Test evidence
Scoped runs (all green):
```
CORE: pnpm --filter @fusion/core exec vitest run \
src/__tests__/skip-bypass-taint-guard.test.ts \
src/__tests__/skip-bypass-taint-persistence.test.ts \
src/__tests__/manual-retry-reset.test.ts
→ 17 passed
ENGINE: pnpm --filter @fusion/engine exec vitest run \
src/__tests__/executor-skip-bypass-taint.test.ts \
src/__tests__/self-healing.test.ts
→ 401 passed
```
Coverage: pure-evaluator (skip-before-refusal counts, skip-after-refusal
doesn't, taint-clearing, empty-marker/empty-steps edges); store
round-trip of the marker (set→read→clear); executor white-box (implicit
completion refused when tainted, allowed when clean or fully re-done,
graph merge-boundary reports missing proof, and the **explicit
`fn_task_done` PREMISE-STALE honest exit stays accepted**); self-healing
(FN-8141 sequence does not promote from either recovery path; a clean
legitimately-skipped task still promotes); manual-retry clears the
marker.
## Note on `pnpm verify:fast`
`verify:fast` currently fails at the workspace-artifact bootstrap on
**pre-existing** pi-SDK type errors in
`packages/engine/src/{auth-storage,pi,provider-registration}.ts` — the
FN-8145 upstream migration breakage (pi 0.80.x removed
`AuthStorage`/`ModelRegistry.create`). **None of those files are in this
diff.** `@fusion/core` builds clean (`packages/core build: Done`), and
`@fusion/engine` `tsc` reports **no errors in the files this PR
touches** (`executor.ts`, `self-healing.ts`); the only engine build
errors are the FN-8145 files. This base failure is the same condition
FN-8141 describes and is out of scope for this task.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus <noreply@anthropic.com>