Ensure group-merge coordinator fakes initialize production merge-lane state.
- Add a reusable fixture for ProjectEngine merge-lane fields.
- Use the fixture in the group-merge routing integration test.
- Include capacity and PR-retry state required by the production drain.
Files changed:
.../_project-engine-merge-lane-fixture.ts | 62 ++++++++++++++++++++++
.../src/__tests__/group-merge-coordinator.test.ts | 15 +++---
2 files changed, 69 insertions(+), 8 deletions(-)
Fusion-Task-Id: FN-8871
Fusion-Task-Lineage: 9bde0439-fa88-481b-a7f3-9feb7e663883
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Operator directed deletion of tests that test pre-refactor behavior no
longer in the codebase (removed APIs, mock shape drift, stale assertions
from the 2026-08-05 full-suite quarantine wave, run 30982276306).
All 27 entries were permanently red — not flaky — testing APIs removed
during the PG cutover and workflow peel refactors (getBuiltinWorkflow,
resolveWorkflowIrForTaskWithProvenance, layer.db.select mock shapes,
vi.mock hoist errors, stale serialization/count literals).
Kept 3 actionable entries that catch real issues:
- register-model-routes-kimi-k3-supplemental (real CI flake, rescue feature ready)
- project-engine.test.ts (catches real 60s→120s assertion drift)
- PlanningModeModal.planning-flow (second-sighting real race)
Vitest config exclusions and quarantine ledger updated in lockstep.
Align the delegation role test with the pluralized runtime error message.
- Update the reviewer delegation rejection expectation to use the roles label.
Files changed:
packages/engine/src/__tests__/agent-tools-delegation.test.ts | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Fusion-Task-Id: FN-8844
Fusion-Task-Lineage: d9a34bda-71de-4b47-b93d-a27dfeb018d1
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Found by a subsystem audit of FN-8764's work-item and role-routing design,
prompted by three production deadlocks already fixed in it.
1. executor.ts — closing out the continuation could skip handleGraphFailure.
The two transitions that close a run's continuation sat outside the
interpreter try/catch with no handler, unlike their siblings in the same
function. The row is usually ALREADY terminal by then: the run's first fence
write retires the continuation it resumed on, which is what makes the
handover atomic. So `succeeded -> failed` hit the store's terminal guard and
threw, escaping executeWorkflowGraph and skipping handleGraphFailure — a
failed run's card was left sitting in its wip column, unparked, with no error
recorded. Closing the continuation is bookkeeping and must never pre-empt the
lifecycle action.
2. executor.ts — capacity attemptId dropped its run-id fallback.
`resolvedRunId` is optional by construction (a definition load failure leaves
it undefined) and this interpolated it raw, producing the literal attempt id
`undefined:<nodeInstance>` shared by every task in the project that hit that
failure. The lease is keyed on (projectId, attemptId) and returns "acquired"
for a pre-existing row regardless of agent, so colliding tasks bypass both the
project and per-agent caps and one task's release deletes another's live
lease. The two durable writes on either side already used the fallback.
3. workflow-task-runtime.ts — failWorkItem dropped a promise bare.
The write is deliberately fire-and-forget, but an unhandled rejection (most
likely the terminal guard when a peer already closed the row) crossed into
process-level unhandled-rejection territory while the caller had already
returned "failed" as if it were persisted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A principal fence written for a node inside a foreach template stores the
TEMPLATE node id (step-execute) with the materialized instance in
nodeInstanceId (steps#0:step-execute). The template node lives under the
foreach's config.template and is never in ir.nodes, so handing it to the
interpreter as a start node resolved to nothing and threw WorkflowIrError.
executeWorkflowGraph's catch turned that into a terminal graph failure, so a
healthy card was parked on every dispatch:
[workflow-graph] FN-8825 could not resolve workflow — parking task instead of
legacy fallback: interpreter-error: Workflow IR missing start node
Latent since FN-8764 introduced these fences, and reachable only once a
step-execute fence could become the task's sole active continuation — which the
atomic-handover change in dd40691ca2 made routine.
The executor now passes a continuation node id as the resume point only when the
task's resolved IR actually contains it. Otherwise it falls back to the graph
entry contract: with no explicit start node the run re-enters at the card's own
column, so an in-progress card re-enters at parse, finds the foreach already
expanded, and hands control back to steps. The instance resumes from its own row
in workflow_run_step_instances, so nothing is replayed. Already-persisted
template-node continuations therefore heal on their next dispatch with no
migration.
Also splits the error message. One string covered a genuinely malformed IR and a
caller asking to resume at an unknown node, and reporting the second as "missing
start node" sends the reader to inspect a workflow definition that is fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Code review of ef8828f14 found the continuation handover it introduced was a
hand-rolled, non-atomic replacement for a primitive this repo already has, with
six P1 defects — two of which recreated the very deadlock it was written to fix.
The invariant: a task may hold ONE active kind="task" work item
(idx_workflow_work_items_one_active_task_continuation), and that partial unique
index is NOT what a plain upsert's ON CONFLICT targets. So a predecessor the run
has already left makes the write RAISE.
Every continuation write in the executor and triage now goes through
replaceActiveTaskWorkflowContinuation, which retires non-matching active rows
and installs the successor in ONE transaction under the task advisory lock:
- Sibling foreach instances share the template nodeId and differ only by runId,
so the old node-identity guard released nothing and instance #1 re-deadlocked.
- Reacting to a FAILED write could not tell an index conflict from a transient
database error, so it destroyed legitimate held continuations.
- Read-then-write across separate transactions let a concurrent engine lose a
live claim; the lock now serializes it.
- A failed retry left the task with zero active rows and no error, because the
hold then transitioned an already-terminal row and the throw was swallowed.
- The same unguarded write existed on the executor's hold path and at both of
triage's planning-continuation writes; a throw there degraded a recoverable
availability hold into a terminal graph failure.
Coverage moves from a fake store to the real index: the new PG suite proves the
bare upsert raises and that replace handles a different node, a sibling foreach
instance, a held predecessor, and re-entry, plus a drift guard tying the SQL
predicate to ACTIVE_WORKFLOW_WORK_ITEM_STATES. The hand-rolled handover is
tombstoned so it cannot return as a "conflict fix", and both new run-audit
events are documented in the AGENTS.md inventory.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every task sat in progress with no session, no log, and no error after the
FN-8764 role-agent rollout. Two independent deadlocks, both invisible:
1. The in-process runtime built its AgentStore but never passed it into
TaskExecutorOptions, so the executor's fail-closed role-routing gate refused
every classified node (execute/step-execute/review/merge).
2. A resumed run keeps the continuation work item it woke on active until the
interpreter returns, so the next node's principal-fence upsert violated
idx_workflow_work_items_one_active_task_continuation — a different index than
its ON CONFLICT target — and raised. The run re-suspended on every dispatch;
only an operator bouncing the card to the hold column cleared it.
Both refusals were swallowed as recoverable "principal holds" that write no log,
audit row, or task error, which is why a fully deadlocked board looked idle.
- Wire agentStore into the executor; assert the shared instance at every runtime
seam in the PG composition test.
- Supersede an active work item for a node the run has already left, then retry
the fence write once; never touch a claim on the node currently executing.
- Record task:workflow-run-suspended and task:workflow-continuation-superseded;
log principal holds, routing-unavailable faults, and fence-write errors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## Summary
Adds an **operator-only** escape hatch for a card stranded `in-review`
(or `in-progress`) with a workflow step permanently stuck in `pending`
status — the leading real-world cause being a dispatched prompt node
(e.g. `code-review`) whose verdict callback was never received (see
#1946). Transitions the stuck `pending` pre-merge step to `status:
"failed"` with resume audit metadata, so the existing
`fn_task_bypass_review` escape hatch can then clear the merge blocker.
## What changed
- **`WorkflowStepResult`** gains resume audit fields: `resumedBy`,
`resumedAt`, `resumeReason`, `resumedFromStatus`. They are pure audit
trail and **do not** participate in merge-blocking
(`getTaskMergeBlocker`).
- **`findPendingPreMergeStep`** (new helper, exported from
`@fusion/core`) summarizes the stuck-pending pre-merge state for
operator tooling. Ignores post-merge steps; returns the newest pending
pre-merge result.
- **`TaskStore.resumeWorkflowStep(id, { stepId, reason, actor })`** —
the store primitive (eligibility-gated: task must be
`in-review`/`in-progress`, not paused; step must exist and be `pending`;
a mandatory non-blank `reason` and `stepId` are required). Runs under
`withTaskLock`, writes the resume as a terminal `failed` result, appends
a task-log breadcrumb, and emits the new `task:resume-step` run-audit
event.
- **`fn_workflow_step_resume`** — new CLI/pi-extension tool registered
**only** on the operator surface (deliberately **not** wired into
executor/reviewer/triage agent tool lists). Accepts `{ id, stepId,
reason }`; the actor defaults to `cli-operator`.
- **Run-audit**: new `task:resume-step` `DatabaseMutationType` member.
## Why
A prompt-node verdict callback can be lost (dispatched prompt never
receives a verdict), leaving the step `pending` forever. Previously the
only recourse was `fn_task_bypass_review`, which requires a terminal
*failed* pre-merge step to clear the blocker — a permanently `pending`
step could not be bypassed. This PR bridges that gap: resume (pending →
failed) then bypass (failed merge-blocker cleared).
## Verification
- **Typecheck**: `@fusion/core`, `@fusion/engine`, `@runfusion/fusion`
all clean.
- **`task-merge-bypass.test.ts`**: 15/15 pass (incl. 5 new
`findPendingPreMergeStep` cases).
- **`store-resume-step.test.ts`** (new, PG-backed): 9/9 pass —
eligibility gating, resume rewrite + audit fields, run-audit event,
non-pending/non-found/blank-argument rejection, in-progress column
support, property preservation.
- **`extension.test.ts`**: 75/75 pass (expected-tool registration
includes the new tool).
## Files
- `packages/core/src/types/workflow/workflow-steps.ts`
- `packages/core/src/merge/task-merge.ts`
- `packages/core/src/store.ts`
- `packages/core/src/index.ts`
- `packages/core/src/__tests__/store-resume-step.test.ts` (new)
- `packages/core/src/__tests__/task-merge-bypass.test.ts`
- `packages/engine/src/util/run-audit.ts`
- `packages/cli/src/extension.ts`
- `packages/cli/src/__tests__/extension.test.ts`
- `.changeset/stas-032-resume-workflow-step.md` (minor, feature)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an operator-only workflow recovery tool for permanently pending
pre-merge steps.
* Operators can mark eligible pending steps as failed by providing a
required audit reason.
* Recovery actions record operator details, timestamps, reasons, prior
status, task logs, and audit events.
* **Bug Fixes**
* Improved selection of the latest pending pre-merge workflow step while
excluding post-merge steps.
* Added validation to prevent recovery of paused, invalid, or
out-of-scope workflow steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: schindler <schindler@users.noreply.github.com>
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Restores the non-blocking full suite on `main` after consistent shard
failures (latest red: [run
30982276306](https://github.com/Runfusion/Fusion/actions/runs/30982276306);
all four shards failed on `@fusion/core`, `@fusion/engine`, and
`@fusion/plugin-sdk`).
### Fixes
- **Path / import drift** after code-organization peels: update
static-guard and integration tests to new module locations (`central/`,
`board/`, `execution/`, `merge/`, `worktree/`, `plugins/`, `types/*`
barrels, etc.).
- **Inventory re-pins**:
- SQLite production `DatabaseSync` allowlist
(`central/project-identity.ts`, `db/sqlite-validation.ts`)
- Engine blocking-shellout allowlist regenerated from live source (33
audited sites)
- Core log-severity manifest paths for peeled modules
- **Partial protocol assert update** for `isPlanReviewSatisfied` (file
also quarantined until full rescue)
### Quarantine (deletion ratchet)
Remaining behavioral reds quarantined on sight — no
timeout/retry/assertion appeasement:
- **14 core** files (incomplete unit fakes for `layer.db.select`,
ledger/census drift, 15s wedge timeout, serialization protocol drift)
- **13 engine** files (mock-hoist errors, fake-store/census/behavior
drift under suite)
Paired updates: `scripts/lib/test-quarantine.json` + package vitest
excludes. Deletion clock starts `2026-08-05`.
### Local verification
- Path-fixed core scanners: 173 passed
- Path-fixed engine scanners: 58 passed
- `@fusion/plugin-sdk` full: 16 passed
- PG smokes: mission-autopilot, research-execution, satellite,
transition-pending, workflow-sync
## Test plan
- [ ] CI PR checks green (lint/typecheck/build/gate)
- [ ] Full suite on merge to main: all 4 shards green or only
intentional non-blocking signal
- [ ] Confirm quarantined files appear in ledger + vitest excludes and
are not executed
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated test coverage to reflect reorganized source locations and
module paths.
* Refreshed static checks, allowlists, and source-based assertions
without changing tested behavior.
* **Chores**
* Quarantined failing core and engine test suites with documented
tracking details.
* Updated test configuration and quarantine records to improve suite
stability and reporting.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Both gate-admitted census tests re-scanned the whole ~1200-file corpus
redundantly: legacy-column-literal-census ran census() twice in full
(~900ms each) and no-legacy-move-targets re-read the entire tree a second
time in its anti-vacuity case. Memoize the census result and share a
comment-stripped read cache so each tree is scanned once per run.
Fusion-Task-Id: none
## Summary
Add an opt-in source-development loop that restarts the dashboard and
engine when runtime TypeScript or JSON changes. Use `pnpm dev:watch`;
`pnpm dev:hmr` now combines Vite UI HMR with the same supervised
API/engine restart path.
The watcher filters tests, fixtures, generated declarations, build
output, and task state. It coalesces bursts with a two-second maximum
wait, waits for the child to acknowledge its IPC listener, and rebuilds
runtime dist artifacts before a source-triggered respawn.
## Safety model
- Close scheduler, triage, heartbeat, mission, routine, self-healing,
and merge admission before checking for active work.
- Let already-running agents reach a safe boundary; do not mutate
durable pause settings.
- Enter the existing graceful exit-code-86 shutdown and supervised
respawn path.
- Retry failed liveness reads and declined restart requests instead of
dropping the pending change.
- Keep ordinary `pnpm dev` behavior unchanged; inherited watch state
does not break nested non-dashboard development commands.
A development restart intentionally replaces the dashboard process, so
transient dashboard connections and project dev-server children
reconnect or restart with it. Agent work is the protected boundary.
## Validation
- `pnpm lint`
- `pnpm test:gate` (753 tests passed across engine, core, PostgreSQL
gate, and CI-shape suites)
- Focused CLI watcher/restart/supervision suites: 40 tests passed
- Focused engine drain/manager suites: 52 tests passed
- `pnpm --filter @runfusion/fusion typecheck`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm verify:fast` (13 steps passed, including CLI build and real
health boot smoke)
- Manual unsupported-command probe confirms explicit `--watch` fails
clearly outside the dashboard command
## Post-Deploy Monitoring & Validation
- Watch for `[fusion:dev] source changed`, `source restart deferred`,
`active work drained`, and `restart requested` logs during the first
watched development session.
- Healthy behavior is one exit-86 respawn per edit batch, no interrupted
active agents, refreshed dist artifacts, and a healthy dashboard after
respawn.
- Investigate repeated restart loops, watcher attachment warnings,
declined restart retries, or liveness-read failures.
- Immediate mitigation is to use ordinary `pnpm dev` without `--watch`;
no production runtime behavior or durable setting needs rollback.
- Validation owner: Fusion maintainers during the first source edit
after merge.
---
[](https://github.com/EveryInc/compound-engineering-plugin)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added `pnpm dev:watch` to automatically restart development runtime
processes when source files change.
* Development restarts now wait for active work to finish, preventing
new work from starting during the transition.
* Enhanced `pnpm dev:hmr` with graceful runtime source restarts while
keeping the dashboard available.
* Rapid source changes are grouped to avoid unnecessary restarts.
* **Documentation**
* Updated development setup and contribution guides with the new watch
workflow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- update the CE workflow-step fixture runtime import to the moved
`cli-runtime/skill-resolver.js` path
- align the runtime specifier with the existing type-only import
## Test plan
- `pnpm --filter @fusion/engine exec vitest run --project=engine-default
src/__tests__/ce-workflow-step-executor.test.ts --silent=passed-only
--reporter=dot --maxWorkers=1 --no-file-parallelism` (52 passed)
- `pnpm --filter @fusion/engine typecheck`
- `pnpm check:changesets`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated test configuration to reference the correct skill resolver
location.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
Planning can no longer approve or execute against evidence from a
superseded dependency episode. Dependency mutations, approval decisions,
recovery, and execution admission now share serialized lifecycle rules,
so stale planner work cannot restore an invalid approval or release an
unplanned task.
Review also converges instead of discovering one blocker per round.
Planning performs a repository-grounded completeness pass up front; Plan
Review batches all independently discoverable blockers and carries an
episode-scoped decision ledger across revisions; code review traces
changed invariants through production consumers and tests. Repeated
feedback still advances the safety budget, while provider failures and
superseded episodes stay outside the remediation ledger.
The dashboard now exposes manual approval only for the intended
exhausted-review state, and refusal/recovery audit events make rejected
lifecycle transitions diagnosable without leaking prompt content.
## Validation
- `pnpm verify:fast` — scoped typechecks/builds, CLI build, and boot
smoke passed.
- Focused Core and Engine regression suites — 511 tests passed.
- `pnpm lint`, strict changeset validation, Core/Engine typechecks, and
package builds passed.
Fixes#3325.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Improved Plan Review approvals, rejections, and replan-cap handling
across task workflows.
* Added cumulative feedback and attempt tracking across repeated
planning reviews.
* Added safer recovery for stalled planning handoffs and interrupted
approval updates.
* **Bug Fixes**
* Prevented stale approvals and unplanned execution after dependency
changes.
* Improved concurrent approval handling, retryability, and
refusal-record deduplication.
* Refined dashboard approval indicators and responsive approval views.
* **Quality Improvements**
* Strengthened planning and code-review completeness checks and
blocking-finding coverage.
* Preserved review history while clearly marking outdated approvals.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Replace ~200ms of real-time sleep() waits with deterministic fake-timer
advances. The dispatcher clamps tickIntervalMs to a 100ms floor, so the
old 30-40ms real sleeps never fired a second interval tick — the
double-dispatch and stop-cancels-timer cases now genuinely exercise a
follow-up tick, making them both faster and stronger.
Fusion-Task-Id: FN-perf