## Why FN-7996 sat in a dispatch→park loop **all day** (06:06→16:35): its `worktree` metadata pointed at recycled pool worktrees (`coral-badger`, `grand-ridge` — the latter actually belonged to FN-8069), Plan Review refused to start in the missing directory, and the task terminal-parked `failed` every cycle while the planner overseer blindly retried. Root-cause chain: 1. **`graphFailureValue()` couldn't read optional-group results.** `runOptionalGroup` publishes context under the group id (`node:plan-review:value`) and the unqualified template id, but the failed node is recorded as the materialized `plan-review::plan-review-step` — the lookup only understood `#` foreach ids. FN-7977's provider-failure hold *did* classify this failure, but its hold value was invisible to routing. 2. **No graph-failure router handled the `assertValidWorktreeSession` refusal**, so it fell to the terminal sink, which parked the task and *overwrote* `task.error` with a generic message — erasing the signature the missing-worktree self-healing sweep (in-review-only anyway) classifies on. 3. **Plan Review didn't need the worktree at all** — its spec is store-injected (FN-7561) — yet it launched its reviewer in whatever stale `task.worktree` said. ## What - `handleGraphFailure` routes unusable-worktree node failures (any node, any error key, `::`/`#` materialized ids) into the existing bounded worktree-session recovery: clear stale worktree/branch/session metadata, requeue to todo, budgeted by `worktreeSessionRetryCount`. An exhausted budget still falls through to the visible terminal park for human inspection. - `graphFailureValue` resolves `group::template` ids (group value first — it carries post-classification routing intent — then the unqualified template value). Foreach `#` behavior unchanged. - Plan Review falls back to the repo root when its recorded worktree is missing on disk; other read-only gates intentionally keep failing fast into the new recovery (silently retargeting them to root would review the wrong tree). - `recoverMissingWorktreeSessionStartFailure` returns its outcome so the graph router can distinguish requeue from escalate-exhausted; existing truthy callers unchanged. ## Symptom Verification - **Original symptom:** graph-node session-start refusal → `Workflow graph terminated with failure at node 'plan-review::plan-review-step'`, task parked failed with stale metadata intact, no recovery. - **Reproduction:** `graph-node-missing-worktree-recovery.test.ts` drives `handleGraphFailure` with the exact FN-7996 result shape (optional-group materialized id + `Refusing to start coding agent in missing worktree` node error). - **Assertion it is gone:** the task is requeued to `todo` with `worktree`/`branch`/`sessionFile` cleared and retry budget incremented — and is *not* marked `failed`; budget exhaustion still parks visibly. ## Surface Enumeration - Optional-group template nodes (Plan Review — the repro), write-capable review gates, and any custom graph node: covered by the `handleGraphFailure` router (scans exact/materialized/unqualified `:error` keys). - Execute-seam session start: already covered by the pre-existing recovery (unchanged, still passes). - In-review / merge-active columns: already covered by self-healing sweeps (unchanged). - Paused / user-paused / deleted / done tasks: explicitly left to their owning machinery (guard tests). - Budget exhaustion: falls through to the visible terminal park (test). ## Testing - `pnpm --filter @fusion/engine exec vitest run src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts` — 13 passed - Adjacent suites (`worktree-incomplete-session-start`, `executor-graph-requeue-gate`, `workflow-graph-optional-group`, `executor-paused-abort-todo-benign`) — 78 passed - `tsc --noEmit` on `@fusion/engine` — clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved recovery when workflow tasks encounter missing or recycled worktrees. * Automatically retries affected tasks with stale worktree details cleared, up to the configured retry limit. * Escalates tasks after recovery attempts are exhausted. * Improved failure routing for optional workflow groups and template instances. * Plan Review now falls back to the repository root when its recorded worktree is unavailable. * **Tests** * Added regression coverage for recovery, routing, retry limits, and repository-root fallback behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
684 B
684 B
@runfusion/fusion
| @runfusion/fusion |
|---|
| patch |
summary: Auto-recover tasks whose workflow step hits a missing or recycled worktree instead of parking them failed forever.
category: fix
dev: FN-7996 root cause set — handleGraphFailure routes assertValidWorktreeSession refusals from any graph node into the bounded worktree-session recovery (clear stale metadata, requeue todo, budgeted by worktreeSessionRetryCount); graphFailureValue now resolves optional-group group::template materialized ids so group routing values (e.g. the FN-7977 plan-review provider-failure hold) are visible; Plan Review runs from the repo root when its recorded worktree is gone (spec is store-injected).