From aa07a78f1817e57fac2cfd41df748bf8ab872d1c Mon Sep 17 00:00:00 2001 From: gsxdsm Date: Thu, 16 Jul 2026 10:28:11 -0700 Subject: [PATCH] fix: recover graph-node missing-worktree failures instead of terminal-parking (FN-7996) (#2231) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Why FN-7996 sat in a dispatch→park loop **all day** (06:06→16:35): its `worktree` metadata pointed at recycled pool worktrees (`coral-badger`, `grand-ridge` — the latter actually belonged to FN-8069), Plan Review refused to start in the missing directory, and the task terminal-parked `failed` every cycle while the planner overseer blindly retried. Root-cause chain: 1. **`graphFailureValue()` couldn't read optional-group results.** `runOptionalGroup` publishes context under the group id (`node:plan-review:value`) and the unqualified template id, but the failed node is recorded as the materialized `plan-review::plan-review-step` — the lookup only understood `#` foreach ids. FN-7977's provider-failure hold *did* classify this failure, but its hold value was invisible to routing. 2. **No graph-failure router handled the `assertValidWorktreeSession` refusal**, so it fell to the terminal sink, which parked the task and *overwrote* `task.error` with a generic message — erasing the signature the missing-worktree self-healing sweep (in-review-only anyway) classifies on. 3. **Plan Review didn't need the worktree at all** — its spec is store-injected (FN-7561) — yet it launched its reviewer in whatever stale `task.worktree` said. ## What - `handleGraphFailure` routes unusable-worktree node failures (any node, any error key, `::`/`#` materialized ids) into the existing bounded worktree-session recovery: clear stale worktree/branch/session metadata, requeue to todo, budgeted by `worktreeSessionRetryCount`. An exhausted budget still falls through to the visible terminal park for human inspection. - `graphFailureValue` resolves `group::template` ids (group value first — it carries post-classification routing intent — then the unqualified template value). Foreach `#` behavior unchanged. - Plan Review falls back to the repo root when its recorded worktree is missing on disk; other read-only gates intentionally keep failing fast into the new recovery (silently retargeting them to root would review the wrong tree). - `recoverMissingWorktreeSessionStartFailure` returns its outcome so the graph router can distinguish requeue from escalate-exhausted; existing truthy callers unchanged. ## Symptom Verification - **Original symptom:** graph-node session-start refusal → `Workflow graph terminated with failure at node 'plan-review::plan-review-step'`, task parked failed with stale metadata intact, no recovery. - **Reproduction:** `graph-node-missing-worktree-recovery.test.ts` drives `handleGraphFailure` with the exact FN-7996 result shape (optional-group materialized id + `Refusing to start coding agent in missing worktree` node error). - **Assertion it is gone:** the task is requeued to `todo` with `worktree`/`branch`/`sessionFile` cleared and retry budget incremented — and is *not* marked `failed`; budget exhaustion still parks visibly. ## Surface Enumeration - Optional-group template nodes (Plan Review — the repro), write-capable review gates, and any custom graph node: covered by the `handleGraphFailure` router (scans exact/materialized/unqualified `:error` keys). - Execute-seam session start: already covered by the pre-existing recovery (unchanged, still passes). - In-review / merge-active columns: already covered by self-healing sweeps (unchanged). - Paused / user-paused / deleted / done tasks: explicitly left to their owning machinery (guard tests). - Budget exhaustion: falls through to the visible terminal park (test). ## Testing - `pnpm --filter @fusion/engine exec vitest run src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts` — 13 passed - Adjacent suites (`worktree-incomplete-session-start`, `executor-graph-requeue-gate`, `workflow-graph-optional-group`, `executor-paused-abort-todo-benign`) — 78 passed - `tsc --noEmit` on `@fusion/engine` — clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) ## Summary by CodeRabbit * **Bug Fixes** * Improved recovery when workflow tasks encounter missing or recycled worktrees. * Automatically retries affected tasks with stale worktree details cleared, up to the configured retry limit. * Escalates tasks after recovery attempts are exhausted. * Improved failure routing for optional workflow groups and template instances. * Plan Review now falls back to the repository root when its recorded worktree is unavailable. * **Tests** * Added regression coverage for recovery, routing, retry limits, and repository-root fallback behavior. --------- Co-authored-by: Claude Fable 5 --- .../github-import-and-issue-close-fixes.md | 2 +- .../graph-node-missing-worktree-recovery.md | 7 + ...aph-node-missing-worktree-recovery.test.ts | 344 ++++++++++++++++++ packages/engine/src/executor.ts | 152 +++++++- 4 files changed, 501 insertions(+), 4 deletions(-) create mode 100644 .changeset/graph-node-missing-worktree-recovery.md create mode 100644 packages/engine/src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts diff --git a/.changeset/github-import-and-issue-close-fixes.md b/.changeset/github-import-and-issue-close-fixes.md index e05d07caf7..9f64c75c53 100644 --- a/.changeset/github-import-and-issue-close-fixes.md +++ b/.changeset/github-import-and-issue-close-fixes.md @@ -2,6 +2,6 @@ "@runfusion/fusion": minor --- -summary: GitHub issue import now pages through all open issues with Previous/Next controls, and linked issues reliably close when their task reaches Done. +summary: GitHub issue import pages through all open issues, and linked issues reliably close when tasks reach Done. category: feature dev: The import picker (GitHubImportModal) fetches up to 300 open issues in one request and pages the result client-side at 30/page with Prev/Next controls and a page indicator; a truncation notice appears past the cap. NewTaskModal's reference picker limit rose 30→100. GitHubClient.listIssues now pages the REST path (per_page loop until limit/exhaustion, PR-filtering no longer stops paging early) and lifts the gh path's 100 cap (gh --limit paginates internally); gh-CLI label filtering fetches the full cap before client-side OR filtering. Separately, the GitHub-tracking reconcile sweep now isolates its three passes in runSweep so a throw in one pass no longer silently starves the others — previously a failure in the first pass disabled the entire close-on-Done backstop, leaving linked/imported issues open; failures are now logged instead of swallowed. diff --git a/.changeset/graph-node-missing-worktree-recovery.md b/.changeset/graph-node-missing-worktree-recovery.md new file mode 100644 index 0000000000..46fbc4bcc1 --- /dev/null +++ b/.changeset/graph-node-missing-worktree-recovery.md @@ -0,0 +1,7 @@ +--- +"@runfusion/fusion": patch +--- + +summary: Auto-recover tasks whose workflow step hits a missing or recycled worktree instead of parking them failed forever. +category: fix +dev: FN-7996 root cause set — `handleGraphFailure` routes `assertValidWorktreeSession` refusals from any graph node into the bounded worktree-session recovery (clear stale metadata, requeue todo, budgeted by `worktreeSessionRetryCount`); `graphFailureValue` now resolves optional-group `group::template` materialized ids so group routing values (e.g. the FN-7977 plan-review provider-failure hold) are visible; Plan Review runs from the repo root when its recorded worktree is gone (spec is store-injected). diff --git a/packages/engine/src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts b/packages/engine/src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts new file mode 100644 index 0000000000..5cbd295a26 --- /dev/null +++ b/packages/engine/src/__tests__/reliability-interactions/graph-node-missing-worktree-recovery.test.ts @@ -0,0 +1,344 @@ +import { describe, expect, it, vi, beforeEach } from "vitest"; +import type { TaskDetail } from "@fusion/core"; +import "../executor-test-helpers.js"; +import { PLAN_REVIEW_PROVIDER_FAILURE_HOLD_VALUE } from "../../workflow-graph-executor.js"; +import { TaskExecutor } from "../../executor.js"; +import { createMockStore, mockedExecSync, mockedExistsSync, resetExecutorMocks } from "../executor-test-helpers.js"; +import { MAX_WORKTREE_SESSION_RETRIES } from "../../self-healing.js"; + +/* +FNXC:MissingWorktreeRecovery 2026-07-16-18:40: +FN-7996 regression coverage. A session-start unusable-worktree refusal thrown inside a +workflow-graph NODE (Plan Review ran with stale task.worktree metadata pointing at a recycled +worktree) fell through every graph-failure router into the terminal park, erasing the error +signature and looping dispatch→park all day. The invariant: any graph-node failure carrying the +assertValidWorktreeSession refusal routes into the bounded worktree-session recovery (clear +stale metadata, requeue todo) and only an exhausted budget may terminal-park; additionally +graphFailureValue must resolve optional-group materialized ids (`group::template`) so group +routing values (e.g. FN-7977's provider-failure hold) are never invisible. +*/ + +const MISSING_WT_ERROR = "Refusing to start coding agent in missing worktree: /tmp/stale-wt"; + +function makeTask(overrides: Partial = {}): TaskDetail { + const now = new Date().toISOString(); + return { + id: "FN-7996-T", + title: "Graph node missing worktree", + description: "Desc", + column: "in-progress", + dependencies: [], + steps: [{ name: "Preflight", status: "pending" }], + currentStep: 0, + log: [], + worktree: "/tmp/stale-wt", + branch: "fusion/fn-7996-t", + status: null, + error: null, + paused: false, + userPaused: false, + createdAt: now, + updatedAt: now, + ...overrides, + } as TaskDetail; +} + +function planReviewGraphFailure(context: Record) { + return { + disposition: "failed", + outcome: "failure" as const, + visitedNodeIds: ["start", "plan-review", "plan-review::plan-review-step"], + context, + }; +} + +function trackingStore(initial: TaskDetail) { + const store = createMockStore(); + let live = initial; + store.getTask.mockImplementation(async () => live as any); + store.updateTask.mockImplementation(async (_id: string, updates: Record) => { + live = { ...live, ...updates } as TaskDetail; + return live as any; + }); + store.moveTask.mockImplementation(async (_id: string, column: string) => { + live = { ...live, column } as TaskDetail; + }); + return { store, getLive: () => live }; +} + +describe("graphFailureValue optional-group materialized ids", () => { + beforeEach(() => { + resetExecutorMocks(); + mockedExecSync.mockReturnValue("" as any); + }); + + it("prefers the group's published value for a `group::template` failed node", () => { + const executor = new TaskExecutor(createMockStore(), "/tmp/test"); + const value = (executor as any).graphFailureValue({ + visitedNodeIds: ["plan-review", "plan-review::plan-review-step"], + context: { + "node:plan-review:value": PLAN_REVIEW_PROVIDER_FAILURE_HOLD_VALUE, + "node:plan-review-step:value": "exception", + }, + }); + expect(value).toBe(PLAN_REVIEW_PROVIDER_FAILURE_HOLD_VALUE); + }); + + it("falls back to the unqualified template value when the group has none", () => { + const executor = new TaskExecutor(createMockStore(), "/tmp/test"); + const value = (executor as any).graphFailureValue({ + visitedNodeIds: ["plan-review::plan-review-step"], + context: { "node:plan-review-step:value": "exception" }, + }); + expect(value).toBe("exception"); + }); + + it("keeps resolving foreach `#` instance ids through the container key", () => { + const executor = new TaskExecutor(createMockStore(), "/tmp/test"); + const value = (executor as any).graphFailureValue({ + visitedNodeIds: ["steps#0:step-execute"], + context: { "node:steps:value": "awaiting-user-input" }, + }); + expect(value).toBe("awaiting-user-input"); + }); +}); + +describe("graph-node unusable-worktree failure recovery (FN-7996)", () => { + beforeEach(() => { + resetExecutorMocks(); + mockedExecSync.mockReturnValue("" as any); + }); + + it("requeues to todo with cleared worktree metadata instead of terminal-parking", async () => { + const initial = makeTask(); + const { store, getLive } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + await (executor as any).handleGraphFailure(initial, planReviewGraphFailure({ + "node:plan-review-step:error": MISSING_WT_ERROR, + "node:plan-review-step:value": "exception", + })); + + const live = getLive(); + expect(live.column).toBe("todo"); + expect(live.status).toBeNull(); + expect(live.worktree).toBeNull(); + expect(live.branch).toBeNull(); + expect(live.worktreeSessionRetryCount).toBe(1); + expect(store.updateTask).not.toHaveBeenCalledWith( + initial.id, + expect.objectContaining({ status: "failed" }), + expect.anything(), + ); + expect(store.moveTask).toHaveBeenCalledWith( + initial.id, + "todo", + expect.objectContaining({ moveSource: "engine", recoveryRehome: true }), + ); + }); + + it("recovers when the refusal is only present under the materialized instance error key", async () => { + const initial = makeTask(); + const { store, getLive } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + await (executor as any).handleGraphFailure(initial, planReviewGraphFailure({ + "node:plan-review::plan-review-step:error": MISSING_WT_ERROR, + })); + + expect(getLive().column).toBe("todo"); + expect(getLive().worktree).toBeNull(); + }); + + it("terminal-parks visibly once the worktree-session retry budget is exhausted", async () => { + const initial = makeTask({ worktreeSessionRetryCount: MAX_WORKTREE_SESSION_RETRIES }); + const { store, getLive } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + await (executor as any).handleGraphFailure(initial, planReviewGraphFailure({ + "node:plan-review-step:error": MISSING_WT_ERROR, + })); + + const live = getLive(); + expect(live.column).toBe("in-progress"); + expect(live.status).toBe("failed"); + expect(String(live.error)).toContain("plan-review::plan-review-step"); + expect(store.moveTask).not.toHaveBeenCalledWith(initial.id, "todo", expect.anything()); + }); + + it("does not intercept graph failures without the worktree refusal signature", async () => { + const initial = makeTask(); + const { store } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + planReviewGraphFailure({ "node:plan-review-step:error": "model API key missing" }), + ); + + expect(handled).toBe(false); + expect(store.moveTask).not.toHaveBeenCalled(); + }); + + it("ignores stale error keys from earlier nodes when a later node failed differently", async () => { + const initial = makeTask(); + const { store } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + { + disposition: "failed", + outcome: "failure", + visitedNodeIds: ["start", "plan-review", "plan-review::plan-review-step", "merge"], + context: { + // Earlier, already-handled node error must not misroute the merge failure. + "node:plan-review-step:error": MISSING_WT_ERROR, + "node:merge:error": "merge conflict in packages/engine/src/executor.ts", + }, + }, + ); + + expect(handled).toBe(false); + expect(store.moveTask).not.toHaveBeenCalled(); + }); + + it("detects the refusal on foreach `container#N:template` materialized ids", async () => { + const initial = makeTask(); + const { store, getLive } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + { + disposition: "failed", + outcome: "failure", + visitedNodeIds: ["steps#0:step-execute"], + context: { "node:step-execute:error": MISSING_WT_ERROR }, + }, + ); + + expect(handled).toBe(true); + expect(getLive().column).toBe("todo"); + }); + + it("leaves auto-merge-off in-review tasks terminal for human merge (FN-5147)", async () => { + const initial = makeTask({ column: "in-review" as const, status: "failed" }); + const { store } = trackingStore(initial); + store.getSettings.mockResolvedValue({ autoMerge: false } as any); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + planReviewGraphFailure({ "node:plan-review-step:error": MISSING_WT_ERROR }), + ); + + expect(handled).toBe(false); + expect(store.moveTask).not.toHaveBeenCalled(); + expect(store.updateTask).not.toHaveBeenCalledWith( + initial.id, + expect.objectContaining({ worktree: null }), + expect.anything(), + ); + }); + + it("still recovers in-review tasks when auto-merge processing is allowed", async () => { + const initial = makeTask({ column: "in-review" as const, status: "failed" }); + const { store, getLive } = trackingStore(initial); + store.getSettings.mockResolvedValue({ autoMerge: true } as any); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + planReviewGraphFailure({ "node:plan-review-step:error": MISSING_WT_ERROR }), + ); + + expect(handled).toBe(true); + expect(getLive().column).toBe("todo"); + }); + + it.each([ + ["paused", { paused: true }], + ["user-paused", { userPaused: true }], + ["deleted", { deletedAt: new Date().toISOString() }], + ["done", { column: "done" as const }], + ])("leaves %s tasks to their owning machinery", async (_label, overrides) => { + const initial = makeTask(overrides as Partial); + const { store } = trackingStore(initial); + const executor = new TaskExecutor(store, "/tmp/test"); + + const handled = await (executor as any).routeUnusableWorktreeGraphFailureToRecovery( + initial, + initial, + planReviewGraphFailure({ "node:plan-review-step:error": MISSING_WT_ERROR }), + ); + + expect(handled).toBe(false); + expect(store.moveTask).not.toHaveBeenCalled(); + }); +}); + +describe("Plan Review missing-worktree repo-root fallback (FN-7996)", () => { + beforeEach(() => { + resetExecutorMocks(); + mockedExecSync.mockReturnValue("" as any); + }); + + it("runs the Plan Review reviewer from the repo root when the recorded worktree is gone", async () => { + const store = createMockStore(); + const executor = new TaskExecutor(store, "/tmp/test"); + mockedExistsSync.mockImplementation((path: unknown) => path !== "/tmp/stale-wt"); + + const captured: { worktreePath?: string } = {}; + vi.spyOn(executor as any, "executeWorkflowStep").mockImplementation(async (...args: any[]) => { + captured.worktreePath = args[2]; + return { success: true, output: "APPROVE" }; + }); + + const node = { + id: "plan-review-step", + kind: "prompt", + config: { name: "Plan Review", prompt: "Review the plan." }, + }; + const live = makeTask({ worktree: "/tmp/stale-wt" }); + store.getTask.mockResolvedValue(live as any); + const result = await (executor as any).runGraphCustomNode(node, live, {}, undefined); + + expect(result.outcome).toBe("success"); + expect(captured.worktreePath).toBe("/tmp/test"); + expect(store.logEntry).toHaveBeenCalledWith( + live.id, + expect.stringContaining("running the reviewer from the repo root"), + undefined, + undefined, + ); + }); + + it("keeps other read-only nodes on the recorded path so they fail fast into recovery", async () => { + const store = createMockStore(); + const executor = new TaskExecutor(store, "/tmp/test"); + mockedExistsSync.mockImplementation((path: unknown) => path !== "/tmp/stale-wt"); + + const captured: { worktreePath?: string } = {}; + vi.spyOn(executor as any, "executeWorkflowStep").mockImplementation(async (...args: any[]) => { + captured.worktreePath = args[2]; + return { success: true, output: "ok" }; + }); + + const node = { + id: "custom-gate", + kind: "prompt", + config: { name: "Custom Gate", prompt: "Check something.", toolMode: "readonly" }, + }; + const live = makeTask({ worktree: "/tmp/stale-wt" }); + store.getTask.mockResolvedValue(live as any); + await (executor as any).runGraphCustomNode(node, live, { reviewerInlineFixes: false }, undefined); + + expect(captured.worktreePath).toBe("/tmp/stale-wt"); + }); +}); diff --git a/packages/engine/src/executor.ts b/packages/engine/src/executor.ts index e5ffb462f7..c4049ec9cf 100644 --- a/packages/engine/src/executor.ts +++ b/packages/engine/src/executor.ts @@ -7845,7 +7845,34 @@ export class TaskExecutor { return { outcome: "failure", value: "no-worktree-for-write-node" }; } - const worktreePath = executionTarget.worktree || this.rootDir; + /* + FNXC:PlanReviewWorktree 2026-07-16-18:30: + FN-7996: Plan Review runs pre-execution and reviews the store-injected PROMPT.md (see + FNXC:PlanReviewSpecInjection) — it does not need worktree contents at all. But it inherited + whatever stale task.worktree metadata survived earlier park/requeue cycles, and a recycled or + pruned path made session start refuse ("Refusing to start coding agent in missing worktree"), + terminal-parking the task. When the recorded worktree is absent on disk, run Plan Review from + the repo root instead. Scoped strictly to Plan Review: other read-only gates review + implementation diffs, so silently retargeting them to the root would review the wrong tree — + they keep failing fast and route through the unusable-worktree graph-failure recovery. + */ + const nodeDisplayName = typeof cfg.name === "string" && cfg.name.trim() ? cfg.name.trim() : node.id; + const isPlanReviewNode = node.id === "plan-review-step" || nodeDisplayName === "Plan Review" || optionalGroupId === "plan-review"; + let worktreePath = executionTarget.worktree || this.rootDir; + if ( + isPlanReviewNode + && !writeCapable + && executionTarget.worktree + && !existsSync(executionTarget.worktree) + ) { + await this.store.logEntry( + live.id, + `Plan Review worktree ${executionTarget.worktree} is missing on disk — running the reviewer from the repo root (spec is store-injected)`, + undefined, + this.getRunContextFor(live.id), + ); + worktreePath = this.rootDir; + } let prompt = typeof cfg.prompt === "string" ? cfg.prompt : ""; let modelProvider = typeof cfg.modelProvider === "string" && cfg.modelProvider.trim() ? cfg.modelProvider : undefined; let modelId = typeof cfg.modelId === "string" && cfg.modelId.trim() ? cfg.modelId : undefined; @@ -8302,6 +8329,25 @@ export class TaskExecutor { if (!failedNode || !result.context) return undefined; const value = result.context[`node:${failedNode}:value`]; if (typeof value === "string") return value; + /* + FNXC:WorkflowLifecycle 2026-07-16-18:20: + Optional-group template failures record materialized `::` ids in + visitedNodeIds, but runOptionalGroup publishes context values under the UNQUALIFIED + template id, and the group wrapper publishes the group's FINAL routing value (e.g. + FN-7977's plan-review provider-failure hold) under the group id. FN-7996 parked + terminally because this lookup only understood `#` foreach ids, so every graph-failure + router (provider hold, awaiting states) missed group-template failures. Prefer the + group's own value (it carries post-classification routing intent), then the template's. + */ + const groupInstanceDelimiter = failedNode.indexOf("::"); + if (groupInstanceDelimiter !== -1) { + const groupNode = failedNode.slice(0, groupInstanceDelimiter); + const groupValue = result.context[`node:${groupNode}:value`]; + if (typeof groupValue === "string") return groupValue; + const templateNode = failedNode.slice(groupInstanceDelimiter + 2); + const templateValue = result.context[`node:${templateNode}:value`]; + return typeof templateValue === "string" ? templateValue : undefined; + } const foreachInstanceDelimiter = failedNode.indexOf("#"); if (foreachInstanceDelimiter === -1) return undefined; /* @@ -8317,6 +8363,88 @@ export class TaskExecutor { return value === "awaiting-user-input" || value === "awaiting-cli-approval"; } + /* + FNXC:MissingWorktreeRecovery 2026-07-16-18:25: + FN-7996: a session-start unusable-worktree refusal (assertValidWorktreeSession in pi.ts) + thrown inside ANY workflow graph node (Plan Review, code review, custom gates) surfaced as a + generic node "exception" and fell through every graph-failure router into the terminal park, + which also OVERWROTE task.error with a generic message — erasing the signature the in-review + missing-worktree self-healing sweep classifies on. The overseer then blindly re-dispatched the + same stale task.worktree all day. Extract the underlying node error from the graph context so + handleGraphFailure can route these into the same bounded recovery the execute session-start + path already uses (clear stale worktree/branch/session metadata, requeue to todo, budgeted by + worktreeSessionRetryCount). + */ + private extractUnusableWorktreeGraphFailure(result: WorkflowGraphTaskRunResult): string | null { + if (!result.context) return null; + const failedNode = result.visitedNodeIds[result.visitedNodeIds.length - 1]; + if (!failedNode) return null; + /* + FNXC:MissingWorktreeRecovery 2026-07-16-19:40: + Detection is scoped to the FAILED node's error keys only (exact id, plus the + `group::template` / `container#N:template` materialized-id derivations under which + runOptionalGroup/foreach publish template context). A catch-all scan over every + `node:*:error` entry would match a STALE error left by an earlier, already-handled node + and misroute an unrelated later failure into worktree recovery (greptile PR#2231 P1). + */ + const candidateKeys: string[] = [`node:${failedNode}:error`]; + const groupInstanceDelimiter = failedNode.indexOf("::"); + if (groupInstanceDelimiter !== -1) { + candidateKeys.push(`node:${failedNode.slice(groupInstanceDelimiter + 2)}:error`); + candidateKeys.push(`node:${failedNode.slice(0, groupInstanceDelimiter)}:error`); + } + const foreachInstanceDelimiter = failedNode.indexOf("#"); + if (foreachInstanceDelimiter !== -1) { + candidateKeys.push(`node:${failedNode.slice(0, foreachInstanceDelimiter)}:error`); + const instanceRest = failedNode.slice(foreachInstanceDelimiter + 1); + const templateDelimiter = instanceRest.indexOf(":"); + if (templateDelimiter !== -1) { + candidateKeys.push(`node:${instanceRest.slice(templateDelimiter + 1)}:error`); + } + } + for (const key of candidateKeys) { + const value = result.context[key]; + if (typeof value === "string" && isMissingWorktreeSessionStartFailure(value)) return value; + } + return null; + } + + private async routeUnusableWorktreeGraphFailureToRecovery( + task: Task, + live: TaskDetail, + result: WorkflowGraphTaskRunResult, + ): Promise { + if (live.deletedAt) return false; + if (live.paused || live.userPaused === true) return false; + if (live.column === "done" || live.column === "archived") return false; + // Pause/abort provenance owns aborted runs; a genuine abort never carries the + // session-start refusal as its terminal node error in the same walk. + if (this.pausedAborted.has(task.id)) return false; + const errorText = this.extractUnusableWorktreeGraphFailure(result); + if (!errorText) return false; + /* + FNXC:MissingWorktreeRecovery 2026-07-16-19:40: + FN-5147: with auto-merge off, `in-review` is terminal-until-human-merged — recovery must + not move those tasks backward or re-enqueue them. Mirrors the gating the in-review + self-healing sweep (recoverMissingWorktreeReviewFailures) applies before the same recovery. + */ + if (live.column === "in-review") { + const settings = await this.store.getSettings(); + if (!allowsAutoMergeProcessing(live, settings)) return false; + } + const stalePath = extractMissingWorktreePathFromSessionStartFailure(errorText) ?? live.worktree ?? ""; + const audit = createRunAuditor(this.store, { + runId: this.getRunContextFor(task.id)?.runId ?? generateSyntheticRunId("graph-worktree-recovery", task.id), + agentId: this.getRunContextFor(task.id)?.agentId ?? (task.assignedAgentId ?? "executor"), + taskId: task.id, + phase: "execute", + }); + const outcome = await this.recoverMissingWorktreeSessionStartFailure(live, stalePath, new Error(errorText), audit); + // escalate-exhausted intentionally returns false: the failure falls through to the + // visible terminal park so a human inspects the task instead of it looping silently. + return outcome === "requeue-todo"; + } + private isMergeGraphFailure(failedNode: string | undefined): boolean { /* FNXC:WorkflowLifecycle 2026-06-19-00:00: @@ -8970,6 +9098,17 @@ export class TaskExecutor { return; } const live = loadedLive; + /* + FNXC:MissingWorktreeRecovery 2026-07-16-18:25: + An unusable-worktree session-start refusal inside a graph node must route to the bounded + worktree-session recovery BEFORE any other classifier: FN-7977's provider-failure hold + would otherwise retry the same stale worktree in place, and the terminal sink would park + the task failed with the signature erased (FN-7996 looped dispatch→park all day). + */ + if (await this.routeUnusableWorktreeGraphFailureToRecovery(task, live, result)) { + await this.persistTokenUsage(task.id); + return; + } if (this.graphFailureValue(result) === PLAN_REVIEW_PROVIDER_FAILURE_HOLD_VALUE) { /* * FNXC:PlanReviewReplan 2026-07-15-16:35: @@ -16966,12 +17105,19 @@ You have access to the file system to review changes.${inlineFixBlock}${verdictB } } + /* + FNXC:MissingWorktreeRecovery 2026-07-16-18:35: + Returns the recovery outcome (not a bare boolean) so the FN-7996 graph-failure router can + distinguish "requeued for clean retry" (handled — stop failure processing) from + "escalate-exhausted" (fall through to the visible terminal park for human inspection). + Existing session-start callers treat any truthy outcome as handled, unchanged. + */ private async recoverMissingWorktreeSessionStartFailure( task: Task, worktreePath: string, error: unknown, audit: RunAuditor, - ): Promise { + ): Promise { const errorText = error instanceof Error ? error.message : String(error); const missingWorktreeFailure = isMissingWorktreeSessionStartFailure(errorText); const missingTaskJsonFailure = isTransientMissingTaskJsonError(error, task); @@ -17047,7 +17193,7 @@ You have access to the file system to review changes.${inlineFixBlock}${verdictB this.getRunContextFor(task.id), ); } - return true; + return recovery.outcome === "escalate-exhausted" ? "escalate-exhausted" : "requeue-todo"; } private async emitWorktreeReanchoredAudit(