Every task sat in progress with no session, no log, and no error after the
FN-8764 role-agent rollout. Two independent deadlocks, both invisible:
1. The in-process runtime built its AgentStore but never passed it into
TaskExecutorOptions, so the executor's fail-closed role-routing gate refused
every classified node (execute/step-execute/review/merge).
2. A resumed run keeps the continuation work item it woke on active until the
interpreter returns, so the next node's principal-fence upsert violated
idx_workflow_work_items_one_active_task_continuation — a different index than
its ON CONFLICT target — and raised. The run re-suspended on every dispatch;
only an operator bouncing the card to the hold column cleared it.
Both refusals were swallowed as recoverable "principal holds" that write no log,
audit row, or task error, which is why a fully deadlocked board looked idle.
- Wire agentStore into the executor; assert the shared instance at every runtime
seam in the PG composition test.
- Supersede an active work item for a node the run has already left, then retry
the fence write once; never touch a claim on the node currently executing.
- Record task:workflow-run-suspended and task:workflow-continuation-superseded;
log principal holds, routing-unavailable faults, and fence-write errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
provisionBuiltinWorkflowRoleAgents took a project-scoped pg_advisory_xact_lock
inside transactionImmediate, then ran provision() against the pool. The lock
holder therefore needed a SECOND pooled connection to finish while concurrent
callers occupied the remaining slots blocking on that same advisory lock. With
DEFAULT_POOL_MAX=3 this self-deadlocked: the holder could never complete, the
waiters could never take the lock, and every subsequent query -- that is, every
DB-backed API route -- queued forever behind an exhausted pool. Observed as a
dashboard that booted ("Ready in 6.5s") and then answered no /api request while
the event loop sat idle in kevent; pg_stat_activity showed one session idle in
transaction holding the lock and two active sessions waiting on it.
Thread an optional QueryHandle through listAgents, findAgentByName, createAgent,
and writeAgent so provisioning runs on tx and the lock and its work share one
connection.
Regression test bounds the pool to a single connection, which makes any second
checkout unsatisfiable and fails deterministically rather than racing. Verified
by reverting the one-line fix: the suite hangs past 300s instead of passing in
under 4s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both plugins declared dashboardViews only in their src/index.ts module.
PluginLoader.getCurrentManifestDashboardViews treats a successfully-read
manifest as authoritative and returns an empty list when the key is absent,
and getPluginDashboardViews only falls back to the module definition when the
manifest read fails. The module-level entries were therefore discarded and
neither view appeared on any nav surface -- header overflow, desktop sidebar,
or mobile More sheet all consume the same array from usePluginDashboardViews.
Mirror the manifest shape used by fusion-plugin-compound-engineering.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## Summary
Adds an **operator-only** escape hatch for a card stranded `in-review`
(or `in-progress`) with a workflow step permanently stuck in `pending`
status — the leading real-world cause being a dispatched prompt node
(e.g. `code-review`) whose verdict callback was never received (see
#1946). Transitions the stuck `pending` pre-merge step to `status:
"failed"` with resume audit metadata, so the existing
`fn_task_bypass_review` escape hatch can then clear the merge blocker.
## What changed
- **`WorkflowStepResult`** gains resume audit fields: `resumedBy`,
`resumedAt`, `resumeReason`, `resumedFromStatus`. They are pure audit
trail and **do not** participate in merge-blocking
(`getTaskMergeBlocker`).
- **`findPendingPreMergeStep`** (new helper, exported from
`@fusion/core`) summarizes the stuck-pending pre-merge state for
operator tooling. Ignores post-merge steps; returns the newest pending
pre-merge result.
- **`TaskStore.resumeWorkflowStep(id, { stepId, reason, actor })`** —
the store primitive (eligibility-gated: task must be
`in-review`/`in-progress`, not paused; step must exist and be `pending`;
a mandatory non-blank `reason` and `stepId` are required). Runs under
`withTaskLock`, writes the resume as a terminal `failed` result, appends
a task-log breadcrumb, and emits the new `task:resume-step` run-audit
event.
- **`fn_workflow_step_resume`** — new CLI/pi-extension tool registered
**only** on the operator surface (deliberately **not** wired into
executor/reviewer/triage agent tool lists). Accepts `{ id, stepId,
reason }`; the actor defaults to `cli-operator`.
- **Run-audit**: new `task:resume-step` `DatabaseMutationType` member.
## Why
A prompt-node verdict callback can be lost (dispatched prompt never
receives a verdict), leaving the step `pending` forever. Previously the
only recourse was `fn_task_bypass_review`, which requires a terminal
*failed* pre-merge step to clear the blocker — a permanently `pending`
step could not be bypassed. This PR bridges that gap: resume (pending →
failed) then bypass (failed merge-blocker cleared).
## Verification
- **Typecheck**: `@fusion/core`, `@fusion/engine`, `@runfusion/fusion`
all clean.
- **`task-merge-bypass.test.ts`**: 15/15 pass (incl. 5 new
`findPendingPreMergeStep` cases).
- **`store-resume-step.test.ts`** (new, PG-backed): 9/9 pass —
eligibility gating, resume rewrite + audit fields, run-audit event,
non-pending/non-found/blank-argument rejection, in-progress column
support, property preservation.
- **`extension.test.ts`**: 75/75 pass (expected-tool registration
includes the new tool).
## Files
- `packages/core/src/types/workflow/workflow-steps.ts`
- `packages/core/src/merge/task-merge.ts`
- `packages/core/src/store.ts`
- `packages/core/src/index.ts`
- `packages/core/src/__tests__/store-resume-step.test.ts` (new)
- `packages/core/src/__tests__/task-merge-bypass.test.ts`
- `packages/engine/src/util/run-audit.ts`
- `packages/cli/src/extension.ts`
- `packages/cli/src/__tests__/extension.test.ts`
- `.changeset/stas-032-resume-workflow-step.md` (minor, feature)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an operator-only workflow recovery tool for permanently pending
pre-merge steps.
* Operators can mark eligible pending steps as failed by providing a
required audit reason.
* Recovery actions record operator details, timestamps, reasons, prior
status, task logs, and audit events.
* **Bug Fixes**
* Improved selection of the latest pending pre-merge workflow step while
excluding post-merge steps.
* Added validation to prevent recovery of paused, invalid, or
out-of-scope workflow steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: schindler <schindler@users.noreply.github.com>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Warm the extension-host task stores **up front** at dashboard startup
instead of letting the first `fn_task_*` call lazily boot a second
PostgreSQL pool per project.
## What changed
`packages/cli/src/commands/dashboard.ts`:
- After the dashboard boots, iterate every registered project (from
`centralCoreForEngine.listProjects()`) and call
`setHostTaskStore(p.path, engine.getTaskStore())` for each non-cwd
project that already has a running `ProjectEngine`.
- Reuses each engine's **existing** `TaskStore` directly — no new
backend connection, no schema advisory-lock contention, no extra
connection-pool exhaustion.
- `cwd` is skipped because its store is already injected at startup.
- Per-project failures are non-fatal (warn) and a failed project listing
logs a single warn — dashboard startup never blocks on this.
- `.changeset/extension-host-store-warmup.md` (patch, fix).
## Why
Left on its own, the first extension tool call (`fn_task_update`,
`fn_task_archive`, `fn_agent_show`, …) for a non-cwd project falls
through to `createTaskStoreForBackend`, which boots a **second**
PostgreSQL connection pool on demand. On busy hosts that lazy boot can
time out, or the call stalls behind pool/startup contention — the
classic "first `fn_task_*` call is slow or errors" experience.
Pre-populating from the already-running engines removes that lazy
worst-case path entirely.
## Verification
- `pnpm verify:fast` — PASS (13 steps, 115s): CLI `tsup` build green,
scoped typecheck/build green, boot smoke green (`fn --help` + real
`serve` with `GET /api/health` 200).
- Cherry-picked cleanly onto current `origin/main` (`5532019fd`); branch
is up-to-date with `origin/main` at PR time.
## Files
- `packages/cli/src/commands/dashboard.ts` (+30)
- `.changeset/extension-host-store-warmup.md` (new)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved dashboard startup reliability by reusing existing project
task connections.
* Prevented extension task tools from creating duplicate connection
pools.
* Added non-blocking warnings when individual project initialization or
discovery fails.
* Dashboard startup now reports how many project task stores were
successfully prepared.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Restores the non-blocking full suite on `main` after consistent shard
failures (latest red: [run
30982276306](https://github.com/Runfusion/Fusion/actions/runs/30982276306);
all four shards failed on `@fusion/core`, `@fusion/engine`, and
`@fusion/plugin-sdk`).
### Fixes
- **Path / import drift** after code-organization peels: update
static-guard and integration tests to new module locations (`central/`,
`board/`, `execution/`, `merge/`, `worktree/`, `plugins/`, `types/*`
barrels, etc.).
- **Inventory re-pins**:
- SQLite production `DatabaseSync` allowlist
(`central/project-identity.ts`, `db/sqlite-validation.ts`)
- Engine blocking-shellout allowlist regenerated from live source (33
audited sites)
- Core log-severity manifest paths for peeled modules
- **Partial protocol assert update** for `isPlanReviewSatisfied` (file
also quarantined until full rescue)
### Quarantine (deletion ratchet)
Remaining behavioral reds quarantined on sight — no
timeout/retry/assertion appeasement:
- **14 core** files (incomplete unit fakes for `layer.db.select`,
ledger/census drift, 15s wedge timeout, serialization protocol drift)
- **13 engine** files (mock-hoist errors, fake-store/census/behavior
drift under suite)
Paired updates: `scripts/lib/test-quarantine.json` + package vitest
excludes. Deletion clock starts `2026-08-05`.
### Local verification
- Path-fixed core scanners: 173 passed
- Path-fixed engine scanners: 58 passed
- `@fusion/plugin-sdk` full: 16 passed
- PG smokes: mission-autopilot, research-execution, satellite,
transition-pending, workflow-sync
## Test plan
- [ ] CI PR checks green (lint/typecheck/build/gate)
- [ ] Full suite on merge to main: all 4 shards green or only
intentional non-blocking signal
- [ ] Confirm quarantined files appear in ledger + vitest excludes and
are not executed
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated test coverage to reflect reorganized source locations and
module paths.
* Refreshed static checks, allowlists, and source-based assertions
without changing tested behavior.
* **Chores**
* Quarantined failing core and engine test suites with documented
tracking details.
* Updated test configuration and quarantine records to improve suite
stability and reporting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- prepares emitted dashboard CSS and binds the fixture server before
starting Chrome
- prevents cold client builds from consuming the supervised browser
lifetime
- adds a regression that holds browser launch until fixture preparation
resolves
## Test plan
- `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest
run --project dashboard-app-quality-foundation-ui --pool=vmThreads
--maxWorkers=1 --silent=passed-only --reporter=dot
app/__tests__/browser-layout-smoke-fixture.test.ts`
- `pnpm --filter @fusion/dashboard typecheck`
- `pnpm --filter @fusion/dashboard build`
- `FUSION_BROWSER_SMOKE_REQUIRE=1 pnpm --filter @fusion/dashboard
test:browser-smoke`
- `pnpm exec eslint packages/dashboard/scripts/browser-layout-smoke.mjs
packages/dashboard/app/__tests__/browser-layout-smoke-fixture.test.ts`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved browser smoke test setup to ensure fixture preparation
completes before the browser launches.
* Added cleanup handling when browser startup fails, preventing leftover
test resources.
* Preserved the primary browser launch error when cleanup also fails,
while recording the cleanup issue.
* **Tests**
* Added coverage for fixture startup order, launch failures, and cleanup
behavior.
* Preserved existing HTML fixture validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->