## Symptom
A tunnel process that dies unexpectedly (crash, OOM, network hiccup, the
24h `maxLifetimeMs` supervision cutoff) leaves `TunnelProcessManager` in
a terminal `failed` state. `handleUnexpectedExit` records the error and
stops — the operator's remote-access tunnel silently goes dark until a
human notices and restarts it.
## Fix
The manager now remembers the desired tunnel (`provider` + `config`)
across the start/switch lifecycle and respawns it after an unexpected
exit with exponential backoff — 1s base doubling to a 30s cap, both
configurable via new `restartBaseDelayMs` / `restartMaxDelayMs` options,
and `autoRestart: false` to opt out. The attempt counter resets once the
tunnel reports ready.
Intentional shutdowns stay shut down: `stop()` and `switchProvider()`
clear the desired tunnel and cancel any pending restart timer.
Hardening that falls out of restart support:
- Exit/spawn-error/readiness-timeout handlers are guarded by
process-handle identity, so a stale `close` event from a superseded
child cannot clobber the state of the restarted tunnel.
- The readiness-timeout path now SIGTERMs the stalled child instead of
leaking it while marking the tunnel failed.
- A synchronous `superviseSpawn` throw now records a redacted
`start_failed` status instead of escaping unhandled.
## Verification
- New tests (fake timers, no real waits): restart after unexpected exit
for both cloudflare and tailscale providers; backoff progression and
delay cap across repeated pre-readiness failures; explicit `stop()`
cancels a pending restart.
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/tunnel-process-manager.test.ts` — 17/17 pass.
- `pnpm verify:fast` — 14 steps green (scoped typecheck/build + CLI
build + boot smoke).
- Changeset included (`patch`, category `fix`).
## Provenance
This is the tunnel half of a fork-side fix (FN-915) that has been
running in production since July; the merge-blocker-preservation half of
that same fix already landed upstream (present since v0.76.0-beta.0).
Ported onto current `main` — the original patch applied cleanly.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Tunnels now automatically restart after unexpected crashes.
* Restart attempts use increasing delays up to a configurable maximum.
* Automatic recovery is enabled by default and can be configured.
* **Bug Fixes**
* Prevented stale process events from triggering unwanted restarts.
* Explicit stops and provider changes now cancel pending restart
attempts.
* Failed or unready tunnel processes are terminated cleanly.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: v <v@v.speedport.ip>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
## Summary
- fail closed when a persisted external remediation route lacks a
concrete checkout path
- verify recovery, remediation, dependency-abort cleanup, and completion
validation use the live persisted task route
- use unique missing-checkout fixtures and exact observed-path
assertions
## Test plan
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/verify-worktree-invariants-missing.test.ts
src/__tests__/executor-triage-column-audit.test.ts --silent=passed-only
--reporter=dot`
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/executor-fast-mode-workflows.test.ts -t 'completed-task
recovery captures the live external checkout|pre-merge remediation'
--silent=passed-only --reporter=dot`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm exec eslint packages/engine/src/executor.ts
packages/engine/src/__tests__/verify-worktree-invariants-missing.test.ts
packages/engine/src/__tests__/executor-triage-column-audit.test.ts
packages/engine/src/__tests__/executor-fast-mode-workflows.test.ts`
- `pnpm check:fnxc-future-dates`
- `pnpm check:changesets --strict`
- `git diff --check`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved recovery for externally executed tasks by using the latest
routing information instead of outdated task data.
* External remediation now stops safely when a checkout location is
missing or invalid, preventing execution in an unintended location.
* Improved cleanup behavior to preserve operator-owned checkouts.
* Enhanced validation and error reporting for missing or invalid
checkout paths.
* **Tests**
* Expanded coverage for recovery, remediation safety, checkout
ownership, and worktree validation scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
The board stopped moving. Work items churned held -> running -> held at
~3.5/sec across every task, pinning a core and writing ~19k workflowWorkItem
audit rows/hour while nothing executed. Hold reason:
workflow-principal-role-pool-exhausted:executor.
provisionBuiltinWorkflowRoleAgents seeded the four permanent owners (triage,
executor, reviewer, merger) with runtimeConfig.enabled=false, while the
router's available() treats enabled===false as unavailable. The only permanent
principals for every built-in role were unroutable BY CONSTRUCTION — shipped
that way, so any instance without operator-created role agents deadlocks at its
first workflow node. Nothing self-recovers: a pool only changes by operator
action.
Routability of these four is an invariant, not a setting. Unlike an operator's
agent, disabling one does not opt an agent out — it removes the only thing that
can run that stage, and there is no fallback.
- seed built-ins enabled; converge existing rows on provisioning
- enforceBuiltinWorkflowRoleRoutability coerces enabled back at the durable
writeAgent seam, so no REST/UI/plugin/restore path can reintroduce the
deadlock. Other runtimeConfig keys are preserved; operator-owned agents keep
their off switch
- share the static routability predicate (isWorkflowPrincipalEligible) between
provisioning and the router so the two cannot drift apart again
Also fix the spin itself: a principal hold had no cooldown, so the scheduler
re-dispatched instantly and the run re-entered only to re-fence and re-park.
It now records a backoff ladder (15s -> 5m) checked before graph entry, and
logs once per distinct reason instead of every pass — the same self-recovering
shape as holdForSessionContention. The hold never increments `attempt`, so no
existing guard could ever fire.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## Summary
- recover task-ID-pinned worktrees when their directory remains but Git
metadata/registration is gone
- preserve the orphan directory under `.fusion/recovery/worktrees`
before recreating the worktree
- retain existing fail-closed behavior for active, foreign, repo-root,
or out-of-root paths
## Problem
A task-pinned worktree can lose `.git` metadata while leaving build
artifacts behind. Fusion classifies that path as incomplete or
unregistered, but then calls `git worktree remove --force`. Git cannot
remove a directory it no longer recognizes as a worktree, so acquisition
aborts and the scheduler can repeat the same recovery indefinitely.
The observed reproduction left `.build` and `.swiftpm` under the pinned
path after Git registration was gone.
## Fix
For an inactive path that is both:
1. inside the configured worktree root, and
2. classified as incomplete or unregistered,
hold the shared worktree-path reservation across classification,
preservation, and recreation. Recovery directories are created one
canonical, project-contained component at a time, then the orphan is
atomically moved into `.fusion/recovery/worktrees` and the pinned
worktree is recreated. Moving rather than deleting preserves any unknown
task artifacts for operator inspection. Other classifications continue
through the existing guarded Git-removal path.
## Verification
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/worktree-acquisition.test.ts --silent=passed-only
--reporter=dot` — 30 passed
- `pnpm --filter @fusion/engine typecheck`
- `pnpm --filter @fusion/engine build`
- `pnpm test:gate:static`
- `pnpm check:changesets`
- `git diff --check`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Improved recovery of task-pinned worktrees when incomplete or stale
directories occupy the expected location.
- Preserves eligible inactive or unregistered worktree contents in a
recovery area instead of deleting them.
- Prevents unsafe recovery through symbolic links and handles concurrent
recovery attempts reliably.
- Supports recovery when directories span different storage devices.
- Ensures interrupted or invalid worktree states can be recreated safely
without disrupting active sessions.
- **Documentation**
- Added a patch changeset documenting the improved worktree recovery
behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- keep persisted operator-routed external checkouts authoritative during
executor recovery, remediation, verification, and cleanup
- fail closed when a configured external route is invalid instead of
falling back to a Fusion-managed worktree
- prevent Fusion from cleaning up operator-owned external checkouts
- add dashboard and executor regression coverage for the routing
handoffs
## Test plan
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/verify-worktree-invariants-missing.test.ts`
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/external-execution-checkout.test.ts
src/__tests__/executor-triage-column-audit.test.ts`
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/executor-fast-mode-workflows.test.ts -t 'external
execution|authoritative executor route|completed-task recovery captures
the live external|pre-merge remediation reuses the live external'`
- `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest
run src/__tests__/routes-tasks-near-duplicate.test.ts -t 'PATCH
external-checkout persists one clean Git checkout for execution and
review'`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm --filter @fusion/dashboard typecheck`
- `pnpm test:gate:static`
- `pnpm check:changesets --strict`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Bug Fixes**
- Improved external checkout routing across execution, verification,
recovery, retries, and remediation.
- Operations now use the latest persisted checkout details, preventing
stale routing information from directing work to the wrong location.
- Invalid or missing checkout routes fail safely with clear verification
errors.
- External checkouts are protected from unintended managed worktree or
branch cleanup.
- **Documentation**
- Clarified external checkout routing and validation behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
A stale base is an optimization miss, not an execution failure. FN-8693's
dispatch-time refresh refused dirty checkouts and own-commit rebase conflicts
with executionSafe:false, and the refusal threw out of acquireTaskWorktree into
execute()'s generic terminal sink — parking the task `failed` and paging the
operator. Run-audit for 2026-08-01..09: 99 of 136 execution failures were these
refusals (74 dirty-worktree, 25 stale-base-conflict), and the bounded
non-parking lane built for them fired 0 times because it only ever saw refusals
published as typed graph node values and no code node enables refreshStaleBase.
82 of the 99 landed within five minutes of "Task marked done by agent": they
were code-review-remediation re-entries into execute() on the task's own warm
worktree — exactly the checkout the refresh must leave alone. Dispatch-time
rebase has no conflict resolution, so on a busy main it could only ever fail;
the merge lane already rebases with AI arbitration before landing and
deliberately leaves refreshStaleBase off.
- refreshReusedWorktreeBase: dirty tree, own-commit conflict, unresolvable base
and compensated persistence failures now return skipped/executionSafe — keep
the local base and run. Only an unproven tree (failed compensation, so a
half-rebased checkout may be on disk) still refuses.
- Check whether a mutation is needed before consulting the working tree: a
worktree already on the current base was refused just for carrying WIP.
- executor: catch WorktreeBaseRefreshError first and route it into
holdForWorktreeBaseRefresh, one shared non-parking lane the graph path now
uses too, so the two entry points cannot drift.
- run-audit: worktree:base-refresh-skipped separates a declined refresh from a
genuine block.
reset-to-base — the actual FN-8693 requirement — is preserved and tested.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
## Summary
- apply Fusion's existing stale-base reconciliation to freshly
reacquired and pooled execution worktrees
- advance retained task branches to the current local integration commit
after dependencies land
- keep planning and Worktrunk behavior unchanged while preserving
dirty/conflict fail-closed handling
## Problem
A dependent task can be planned before its dependency lands. If the
dependency merges and its branch is deleted, a later execution retry may
recreate the dependent worktree from its already-existing task branch.
That branch can still point at the pre-dependency commit.
Fusion already refreshes reused execution worktrees, but fresh
acquisition returned without calling the same reconciliation primitive.
The dependent task therefore executed without the landed dependency
output even though Fusion marked the dependency complete.
## Fix
When `refreshStaleBase` is enabled, run `refreshReusedWorktreeBase`
after a native fresh or pooled worktree is acquired and before cleanup,
init, or session execution. Track the actual backend used by injected
and fallback creators so a native fallback still refreshes while
Worktrunk-managed paths remain excluded. The existing primitive:
- resolves the current local integration branch without requiring a
remote
- resets branches with no task-owned commits
- rebases branches with task-owned commits
- blocks dirty or conflicting worktrees
- persists the integration commit as `baseCommitSha`
If refresh blocks a pooled checkout, clear the task's durable binding
before releasing the checkout for reuse.
Planning callers do not enable `refreshStaleBase`, so planning worktrees
remain unchanged.
## Verification
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/worktree-base-refresh.test.ts
src/__tests__/worktree-acquisition.test.ts --silent=passed-only
--reporter=dot` — 34 passed
- `pnpm --filter @fusion/engine typecheck`
- `pnpm --filter @fusion/engine build`
- `pnpm test:gate:static`
- `pnpm check:changesets`
- `git diff --check`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved worktree acquisition by refreshing stale branches against the
current integration branch.
* Added refresh support for recreated, pooled, and native fallback
worktrees.
* Prevented task execution when refresh fails and safely released
affected pooled worktrees.
* Avoided unnecessary refreshes for newly created Worktrunk worktrees.
* **Tests**
* Added coverage for stale-base refresh behavior across supported
acquisition scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Summary
- normalize Vitest mock-call tuples before selecting graph
implementation sessions
- restore six clean-main executor assertions that otherwise misclassify
a routed implementation call as missing
## Test plan
- `pnpm --filter @fusion/engine exec vitest run --reporter=dot
src/__tests__/executor-review-verdicts.test.ts
src/__tests__/executor-worktree-liveness.test.ts`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm test:gate:static`
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Tests**
* Updated workflow routing and worktree liveness tests to use normalized
agent configuration when validating implementation-session selection.
* Improved coverage for fresh worktrees, configured directories, and
routed implementation sessions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
- add an explicit API route that persists one clean external Git
checkout for task execution and enforced review
- fence execution to the checkout's persisted branch and fail closed
when the route becomes invalid
- allow completion invariants to validate explicitly routed checkouts
outside the project worktree directory
## Test plan
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/external-execution-checkout.test.ts
src/__tests__/review-checkout.test.ts
src/__tests__/engine-no-blocking-shellout.test.ts --silent=passed-only
--reporter=dot`
- `pnpm --filter @fusion/engine exec vitest run
src/__tests__/executor-fast-mode-workflows.test.ts -t 'prepares a
persisted external execution checkout' --silent=passed-only
--reporter=dot`
- `FUSION_DASHBOARD_DEEP=1 pnpm --filter @fusion/dashboard exec vitest
run src/__tests__/routes-tasks-near-duplicate.test.ts -t 'PATCH
external-checkout' --project dashboard-api --silent=passed-only
--reporter=dot`
- `pnpm --filter @fusion/engine typecheck`
- `pnpm --filter @fusion/dashboard exec tsc --noEmit`
- `pnpm verify:fast`
- `pnpm test:gate` (605 non-PostgreSQL tests pass; local PostgreSQL
suites cannot authenticate because the configured client returns an
empty password)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added support for routing task execution and review through
operator-selected external Git checkouts.
* External checkouts are validated for valid Git repositories, attached
branches, clean status, and branch consistency.
* Tasks can clear previously configured external checkout routing.
* Valid routed checkouts are used directly without creating a separate
worktree.
* **Bug Fixes**
* Invalid, incomplete, dirty, or mismatched checkout configurations now
fail early with clear validation errors.
* Missing tasks return the appropriate not-found response.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com>
## Summary
Resolves merge conflicts from #3392 against current `main`.
`main` already landed the core #3386 fix via FN-8909 (`includeArchived:
false` live-row enumeration + per-task isolation). This PR rebases the
remaining #3392 refinements on top of that:
- Best-effort, secret-free audit emission (per-task and no-action)
- Distinguish `no-eligible-orphan` vs `all-attempts-failed` /
`no-finalization` no-action outcomes
- Redacted `errorType=` warn logs so poisoned-row failures cannot abort
the sweep or leak error prose
Supersedes #3392 (fork head is not writable from this environment
despite `maintainer_can_modify`).
## Test plan
- [x] `pnpm --filter @fusion/engine exec vitest run
src/__tests__/self-healing.test.ts --project engine-default --run -t
"finalizeOrphanedPlanningSegments"` — 8 passed
- [ ] CI PR checks green
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved recovery of orphaned planning segments when individual
finalization attempts fail.
* Recovery now continues successfully even if audit recording encounters
an error.
* Added clearer recovery outcomes, distinguishing cases where no
segments qualify, all attempts fail, or only some segments are
finalized.
* Warning messages now provide structured error details without exposing
sensitive information.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Co-authored-by: Codex <codex@openai.com>
Ensure group-merge coordinator tests exercise the production store seams without warning suppression.
- Provide scoped settings and branch-group task fake methods used by promotion evaluation
- Await promotion evaluation and fail the fixture on missing-store-method warnings
- Assert branch-group task lookup is reached during the merge drain
Files changed:
packages/engine/src/__tests__/group-merge-coordinator.test.ts | 59 ++++++++++++++++++----
1 file changed, 49 insertions(+), 10 deletions(-)
Fusion-Task-Id: FN-8895
Fusion-Task-Lineage: 99748c07-2083-41f9-b776-da644bf1857a
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Ensure group-merge coordinator fakes initialize production merge-lane state.
- Add a reusable fixture for ProjectEngine merge-lane fields.
- Use the fixture in the group-merge routing integration test.
- Include capacity and PR-retry state required by the production drain.
Files changed:
.../_project-engine-merge-lane-fixture.ts | 62 ++++++++++++++++++++++
.../src/__tests__/group-merge-coordinator.test.ts | 15 +++---
2 files changed, 69 insertions(+), 8 deletions(-)
Fusion-Task-Id: FN-8871
Fusion-Task-Lineage: 9bde0439-fa88-481b-a7f3-9feb7e663883
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
Operator directed deletion of tests that test pre-refactor behavior no
longer in the codebase (removed APIs, mock shape drift, stale assertions
from the 2026-08-05 full-suite quarantine wave, run 30982276306).
All 27 entries were permanently red — not flaky — testing APIs removed
during the PG cutover and workflow peel refactors (getBuiltinWorkflow,
resolveWorkflowIrForTaskWithProvenance, layer.db.select mock shapes,
vi.mock hoist errors, stale serialization/count literals).
Kept 3 actionable entries that catch real issues:
- register-model-routes-kimi-k3-supplemental (real CI flake, rescue feature ready)
- project-engine.test.ts (catches real 60s→120s assertion drift)
- PlanningModeModal.planning-flow (second-sighting real race)
Vitest config exclusions and quarantine ledger updated in lockstep.
Align the delegation role test with the pluralized runtime error message.
- Update the reviewer delegation rejection expectation to use the roles label.
Files changed:
packages/engine/src/__tests__/agent-tools-delegation.test.ts | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
Fusion-Task-Id: FN-8844
Fusion-Task-Lineage: d9a34bda-71de-4b47-b93d-a27dfeb018d1
Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>