Found by a subsystem audit of FN-8764's work-item and role-routing design,
prompted by three production deadlocks already fixed in it.
1. executor.ts — closing out the continuation could skip handleGraphFailure.
The two transitions that close a run's continuation sat outside the
interpreter try/catch with no handler, unlike their siblings in the same
function. The row is usually ALREADY terminal by then: the run's first fence
write retires the continuation it resumed on, which is what makes the
handover atomic. So `succeeded -> failed` hit the store's terminal guard and
threw, escaping executeWorkflowGraph and skipping handleGraphFailure — a
failed run's card was left sitting in its wip column, unparked, with no error
recorded. Closing the continuation is bookkeeping and must never pre-empt the
lifecycle action.
2. executor.ts — capacity attemptId dropped its run-id fallback.
`resolvedRunId` is optional by construction (a definition load failure leaves
it undefined) and this interpolated it raw, producing the literal attempt id
`undefined:<nodeInstance>` shared by every task in the project that hit that
failure. The lease is keyed on (projectId, attemptId) and returns "acquired"
for a pre-existing row regardless of agent, so colliding tasks bypass both the
project and per-agent caps and one task's release deletes another's live
lease. The two durable writes on either side already used the fallback.
3. workflow-task-runtime.ts — failWorkItem dropped a promise bare.
The write is deliberately fire-and-forget, but an unhandled rejection (most
likely the terminal guard when a peer already closed the row) crossed into
process-level unhandled-rejection territory while the caller had already
returned "failed" as if it were persisted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tasks were parking with:
Workflow principal fence write failed at node 'step-execute' — Workflow work
item <id> is terminal (cancelled) and cannot be requeued as running
A work item is keyed on (project_id, run_id, task_id, node_id, kind), and a
node's fence run id is DERIVED rather than per-attempt
(`<taskId>:<workflowId>:<nodeInstanceId>`), so every attempt at a node instance
targets the same row. Fence rows are terminalized at the end of every run, so
the SECOND visit to any node hit upsertWorkflowWorkItem's terminal guard and
failed the fence closed. Retries, rework cycles (maxReworkCycles) and operator
re-dispatch all make a second visit ordinary, so this caught any task that did
not clear a node on its first pass — every step-execute row in the affected
database was already terminal (succeeded, failed, or cancelled).
The guard is correct; the identity was wrong — a new attempt is a new work item.
replaceActiveTaskWorkflowContinuation is the sanctioned single writer for a task
continuation and its contract is already "retire the predecessor, install the
successor", so it now drops a TERMINAL row occupying the exact target key inside
the same locked transaction before upserting. Bounded and non-destructive: at
most one row per (task, node instance), the dropped row is finished work, and
its transitions are already in run-audit. Terminal rows for other nodes are
untouched and upsertWorkflowWorkItem still refuses to resurrect a terminal row
for every other caller.
Existing stuck rows heal on their next dispatch; no migration or cleanup needed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A principal fence written for a node inside a foreach template stores the
TEMPLATE node id (step-execute) with the materialized instance in
nodeInstanceId (steps#0:step-execute). The template node lives under the
foreach's config.template and is never in ir.nodes, so handing it to the
interpreter as a start node resolved to nothing and threw WorkflowIrError.
executeWorkflowGraph's catch turned that into a terminal graph failure, so a
healthy card was parked on every dispatch:
[workflow-graph] FN-8825 could not resolve workflow — parking task instead of
legacy fallback: interpreter-error: Workflow IR missing start node
Latent since FN-8764 introduced these fences, and reachable only once a
step-execute fence could become the task's sole active continuation — which the
atomic-handover change in dd40691ca2 made routine.
The executor now passes a continuation node id as the resume point only when the
task's resolved IR actually contains it. Otherwise it falls back to the graph
entry contract: with no explicit start node the run re-enters at the card's own
column, so an in-progress card re-enters at parse, finds the foreach already
expanded, and hands control back to steps. The instance resumes from its own row
in workflow_run_step_instances, so nothing is replayed. Already-persisted
template-node continuations therefore heal on their next dispatch with no
migration.
Also splits the error message. One string covered a genuinely malformed IR and a
caller asking to resume at an unknown node, and reporting the second as "missing
start node" sends the reader to inspect a workflow definition that is fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
FN-8764 made AgentStore.init() unconditionally provision the four durable
built-in workflow-owner agents, and that provisioning needs a bound
asyncLayer.projectId — backendProjectId rejects the empty/unbound partition so a
shared cluster cannot mix ownership. Every AgentStore-backed PG test therefore
threw in init(); without this change all 10 cases in agent-instructions.pg.test
fail at agent-store.ts:526.
createTaskStoreForTest / createSharedPgTaskStoreTestHarness gain an OPT-IN
projectId. Undefined keeps the historical project-agnostic harness (RLS bypass,
empty-string partition) that the rest of the core suite relies on. When set, the
connection GUC `fusion.project_id`, the AsyncDataLayer, and the seeded config row
all share one partition — so agents (explicit project_id) and their config
revisions (GUC-default project_id) land together and the
(project_id, agent_id) FK on agent_config_revisions holds.
Also folds in two already-merged consequences: serve.test expects the
consumerId: "engine" that serve.ts:326 already passes (FN-8685), and the
auto-generated Fusion skill docs pick up fn_workflow_step_resume, the roles /
max_workflow_sessions agent fields, and the deprecated singular role.
Uncommitted in the working tree; reviewed, verified against the real database,
and committed as-is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Code review of ef8828f14 found the continuation handover it introduced was a
hand-rolled, non-atomic replacement for a primitive this repo already has, with
six P1 defects — two of which recreated the very deadlock it was written to fix.
The invariant: a task may hold ONE active kind="task" work item
(idx_workflow_work_items_one_active_task_continuation), and that partial unique
index is NOT what a plain upsert's ON CONFLICT targets. So a predecessor the run
has already left makes the write RAISE.
Every continuation write in the executor and triage now goes through
replaceActiveTaskWorkflowContinuation, which retires non-matching active rows
and installs the successor in ONE transaction under the task advisory lock:
- Sibling foreach instances share the template nodeId and differ only by runId,
so the old node-identity guard released nothing and instance #1 re-deadlocked.
- Reacting to a FAILED write could not tell an index conflict from a transient
database error, so it destroyed legitimate held continuations.
- Read-then-write across separate transactions let a concurrent engine lose a
live claim; the lock now serializes it.
- A failed retry left the task with zero active rows and no error, because the
hold then transitioned an already-terminal row and the throw was swallowed.
- The same unguarded write existed on the executor's hold path and at both of
triage's planning-continuation writes; a throw there degraded a recoverable
availability hold into a terminal graph failure.
Coverage moves from a fake store to the real index: the new PG suite proves the
bare upsert raises and that replace handles a different node, a sibling foreach
instance, a held predecessor, and re-entry, plus a drift guard tying the SQL
predicate to ACTIVE_WORKFLOW_WORK_ITEM_STATES. The hand-rolled handover is
tombstoned so it cannot return as a "conflict fix", and both new run-audit
events are documented in the AGENTS.md inventory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every task sat in progress with no session, no log, and no error after the
FN-8764 role-agent rollout. Two independent deadlocks, both invisible:
1. The in-process runtime built its AgentStore but never passed it into
TaskExecutorOptions, so the executor's fail-closed role-routing gate refused
every classified node (execute/step-execute/review/merge).
2. A resumed run keeps the continuation work item it woke on active until the
interpreter returns, so the next node's principal-fence upsert violated
idx_workflow_work_items_one_active_task_continuation — a different index than
its ON CONFLICT target — and raised. The run re-suspended on every dispatch;
only an operator bouncing the card to the hold column cleared it.
Both refusals were swallowed as recoverable "principal holds" that write no log,
audit row, or task error, which is why a fully deadlocked board looked idle.
- Wire agentStore into the executor; assert the shared instance at every runtime
seam in the PG composition test.
- Supersede an active work item for a node the run has already left, then retry
the fence write once; never touch a claim on the node currently executing.
- Record task:workflow-run-suspended and task:workflow-continuation-superseded;
log principal holds, routing-unavailable faults, and fence-write errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
provisionBuiltinWorkflowRoleAgents took a project-scoped pg_advisory_xact_lock
inside transactionImmediate, then ran provision() against the pool. The lock
holder therefore needed a SECOND pooled connection to finish while concurrent
callers occupied the remaining slots blocking on that same advisory lock. With
DEFAULT_POOL_MAX=3 this self-deadlocked: the holder could never complete, the
waiters could never take the lock, and every subsequent query -- that is, every
DB-backed API route -- queued forever behind an exhausted pool. Observed as a
dashboard that booted ("Ready in 6.5s") and then answered no /api request while
the event loop sat idle in kevent; pg_stat_activity showed one session idle in
transaction holding the lock and two active sessions waiting on it.
Thread an optional QueryHandle through listAgents, findAgentByName, createAgent,
and writeAgent so provisioning runs on tx and the lock and its work share one
connection.
Regression test bounds the pool to a single connection, which makes any second
checkout unsatisfiable and fails deterministically rather than racing. Verified
by reverting the one-line fix: the suite hangs past 300s instead of passing in
under 4s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both plugins declared dashboardViews only in their src/index.ts module.
PluginLoader.getCurrentManifestDashboardViews treats a successfully-read
manifest as authoritative and returns an empty list when the key is absent,
and getPluginDashboardViews only falls back to the module definition when the
manifest read fails. The module-level entries were therefore discarded and
neither view appeared on any nav surface -- header overflow, desktop sidebar,
or mobile More sheet all consume the same array from usePluginDashboardViews.
Mirror the manifest shape used by fusion-plugin-compound-engineering.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## Summary
Adds an **operator-only** escape hatch for a card stranded `in-review`
(or `in-progress`) with a workflow step permanently stuck in `pending`
status — the leading real-world cause being a dispatched prompt node
(e.g. `code-review`) whose verdict callback was never received (see
#1946). Transitions the stuck `pending` pre-merge step to `status:
"failed"` with resume audit metadata, so the existing
`fn_task_bypass_review` escape hatch can then clear the merge blocker.
## What changed
- **`WorkflowStepResult`** gains resume audit fields: `resumedBy`,
`resumedAt`, `resumeReason`, `resumedFromStatus`. They are pure audit
trail and **do not** participate in merge-blocking
(`getTaskMergeBlocker`).
- **`findPendingPreMergeStep`** (new helper, exported from
`@fusion/core`) summarizes the stuck-pending pre-merge state for
operator tooling. Ignores post-merge steps; returns the newest pending
pre-merge result.
- **`TaskStore.resumeWorkflowStep(id, { stepId, reason, actor })`** —
the store primitive (eligibility-gated: task must be
`in-review`/`in-progress`, not paused; step must exist and be `pending`;
a mandatory non-blank `reason` and `stepId` are required). Runs under
`withTaskLock`, writes the resume as a terminal `failed` result, appends
a task-log breadcrumb, and emits the new `task:resume-step` run-audit
event.
- **`fn_workflow_step_resume`** — new CLI/pi-extension tool registered
**only** on the operator surface (deliberately **not** wired into
executor/reviewer/triage agent tool lists). Accepts `{ id, stepId,
reason }`; the actor defaults to `cli-operator`.
- **Run-audit**: new `task:resume-step` `DatabaseMutationType` member.
## Why
A prompt-node verdict callback can be lost (dispatched prompt never
receives a verdict), leaving the step `pending` forever. Previously the
only recourse was `fn_task_bypass_review`, which requires a terminal
*failed* pre-merge step to clear the blocker — a permanently `pending`
step could not be bypassed. This PR bridges that gap: resume (pending →
failed) then bypass (failed merge-blocker cleared).
## Verification
- **Typecheck**: `@fusion/core`, `@fusion/engine`, `@runfusion/fusion`
all clean.
- **`task-merge-bypass.test.ts`**: 15/15 pass (incl. 5 new
`findPendingPreMergeStep` cases).
- **`store-resume-step.test.ts`** (new, PG-backed): 9/9 pass —
eligibility gating, resume rewrite + audit fields, run-audit event,
non-pending/non-found/blank-argument rejection, in-progress column
support, property preservation.
- **`extension.test.ts`**: 75/75 pass (expected-tool registration
includes the new tool).
## Files
- `packages/core/src/types/workflow/workflow-steps.ts`
- `packages/core/src/merge/task-merge.ts`
- `packages/core/src/store.ts`
- `packages/core/src/index.ts`
- `packages/core/src/__tests__/store-resume-step.test.ts` (new)
- `packages/core/src/__tests__/task-merge-bypass.test.ts`
- `packages/engine/src/util/run-audit.ts`
- `packages/cli/src/extension.ts`
- `packages/cli/src/__tests__/extension.test.ts`
- `.changeset/stas-032-resume-workflow-step.md` (minor, feature)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an operator-only workflow recovery tool for permanently pending
pre-merge steps.
* Operators can mark eligible pending steps as failed by providing a
required audit reason.
* Recovery actions record operator details, timestamps, reasons, prior
status, task logs, and audit events.
* **Bug Fixes**
* Improved selection of the latest pending pre-merge workflow step while
excluding post-merge steps.
* Added validation to prevent recovery of paused, invalid, or
out-of-scope workflow steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: schindler <schindler@users.noreply.github.com>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Warm the extension-host task stores **up front** at dashboard startup
instead of letting the first `fn_task_*` call lazily boot a second
PostgreSQL pool per project.
## What changed
`packages/cli/src/commands/dashboard.ts`:
- After the dashboard boots, iterate every registered project (from
`centralCoreForEngine.listProjects()`) and call
`setHostTaskStore(p.path, engine.getTaskStore())` for each non-cwd
project that already has a running `ProjectEngine`.
- Reuses each engine's **existing** `TaskStore` directly — no new
backend connection, no schema advisory-lock contention, no extra
connection-pool exhaustion.
- `cwd` is skipped because its store is already injected at startup.
- Per-project failures are non-fatal (warn) and a failed project listing
logs a single warn — dashboard startup never blocks on this.
- `.changeset/extension-host-store-warmup.md` (patch, fix).
## Why
Left on its own, the first extension tool call (`fn_task_update`,
`fn_task_archive`, `fn_agent_show`, …) for a non-cwd project falls
through to `createTaskStoreForBackend`, which boots a **second**
PostgreSQL connection pool on demand. On busy hosts that lazy boot can
time out, or the call stalls behind pool/startup contention — the
classic "first `fn_task_*` call is slow or errors" experience.
Pre-populating from the already-running engines removes that lazy
worst-case path entirely.
## Verification
- `pnpm verify:fast` — PASS (13 steps, 115s): CLI `tsup` build green,
scoped typecheck/build green, boot smoke green (`fn --help` + real
`serve` with `GET /api/health` 200).
- Cherry-picked cleanly onto current `origin/main` (`5532019fd`); branch
is up-to-date with `origin/main` at PR time.
## Files
- `packages/cli/src/commands/dashboard.ts` (+30)
- `.changeset/extension-host-store-warmup.md` (new)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved dashboard startup reliability by reusing existing project
task connections.
* Prevented extension task tools from creating duplicate connection
pools.
* Added non-blocking warnings when individual project initialization or
discovery fails.
* Dashboard startup now reports how many project task stores were
successfully prepared.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Restores the non-blocking full suite on `main` after consistent shard
failures (latest red: [run
30982276306](https://github.com/Runfusion/Fusion/actions/runs/30982276306);
all four shards failed on `@fusion/core`, `@fusion/engine`, and
`@fusion/plugin-sdk`).
### Fixes
- **Path / import drift** after code-organization peels: update
static-guard and integration tests to new module locations (`central/`,
`board/`, `execution/`, `merge/`, `worktree/`, `plugins/`, `types/*`
barrels, etc.).
- **Inventory re-pins**:
- SQLite production `DatabaseSync` allowlist
(`central/project-identity.ts`, `db/sqlite-validation.ts`)
- Engine blocking-shellout allowlist regenerated from live source (33
audited sites)
- Core log-severity manifest paths for peeled modules
- **Partial protocol assert update** for `isPlanReviewSatisfied` (file
also quarantined until full rescue)
### Quarantine (deletion ratchet)
Remaining behavioral reds quarantined on sight — no
timeout/retry/assertion appeasement:
- **14 core** files (incomplete unit fakes for `layer.db.select`,
ledger/census drift, 15s wedge timeout, serialization protocol drift)
- **13 engine** files (mock-hoist errors, fake-store/census/behavior
drift under suite)
Paired updates: `scripts/lib/test-quarantine.json` + package vitest
excludes. Deletion clock starts `2026-08-05`.
### Local verification
- Path-fixed core scanners: 173 passed
- Path-fixed engine scanners: 58 passed
- `@fusion/plugin-sdk` full: 16 passed
- PG smokes: mission-autopilot, research-execution, satellite,
transition-pending, workflow-sync
## Test plan
- [ ] CI PR checks green (lint/typecheck/build/gate)
- [ ] Full suite on merge to main: all 4 shards green or only
intentional non-blocking signal
- [ ] Confirm quarantined files appear in ledger + vitest excludes and
are not executed
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated test coverage to reflect reorganized source locations and
module paths.
* Refreshed static checks, allowlists, and source-based assertions
without changing tested behavior.
* **Chores**
* Quarantined failing core and engine test suites with documented
tracking details.
* Updated test configuration and quarantine records to improve suite
stability and reporting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->