FN-9186: Bound no-progress self-healing requeues
Bound no-progress self-healing retries so persistent failures stop consuming agent sessions. - Persist a three-attempt retry budget with exponential backoff and an idempotent terminal sentinel. - Prevent restart and wedge recovery paths from reopening exhausted failures. - Emit bounded retry and exhaustion audit events with regression coverage and operator documentation. Files changed: .changeset/fn-9186-no-progress-requeue-budget.md | 7 ++ AGENTS.md | 1 + docs/architecture.md | 2 + docs/run-audit.md | 2 + docs/self-healing-backward-move-audit.md | 2 +- .../src/__tests__/default-workflow-hooks.test.ts | 18 +++ .../non-executor-run-audit-sink-health.test.ts | 43 +++++++ .../__tests__/restart-recovery-coordinator.test.ts | 14 +++ ...self-healing-no-progress-requeue-budget.test.ts | 131 +++++++++++++++++++++ packages/engine/src/__tests__/self-healing.test.ts | 59 ++++++---- .../src/healing/no-progress-requeue-budget.ts | 7 ++ .../src/healing/restart-recovery-coordinator.ts | 10 +- .../__tests__/task-wedge-notification.test.ts | 10 +- .../src/notification/task-wedge-notification.ts | 9 ++ packages/engine/src/self-healing.ts | 95 +++++++++++++-- packages/engine/src/util/run-audit.ts | 3 + 16 files changed, 378 insertions(+), 35 deletions(-) Fusion-Task-Id: FN-9186 Fusion-Task-Lineage: 06d29a60-b6ad-4506-a3c1-f3c52411946a Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
7
.changeset/fn-9186-no-progress-requeue-budget.md
Normal file
7
.changeset/fn-9186-no-progress-requeue-budget.md
Normal file
@@ -0,0 +1,7 @@
|
||||
---
|
||||
"@runfusion/fusion": patch
|
||||
---
|
||||
|
||||
summary: Bound self-healing retries for failed no-progress tasks.
|
||||
category: fix
|
||||
dev: Uses persisted retry budget and exponential backoff before terminal operator parking.
|
||||
@@ -435,3 +435,4 @@ Note: the embedded main-content views Workflows (`_WorkflowEditorView`), Import
|
||||
<!-- FNXC:RunAudit 2026-08-20-05:49: FN-9177 requires new core best-effort emitters to use the core-owned bounded seam. -->
|
||||
- FN-9177: New core best-effort emitters must use `packages/core/src/run-audit/emit-bounded-run-audit.ts`. It deliberately mirrors the engine seam because `@fusion/core` cannot import `@fusion/engine`; transactional and deliberately awaited durability writers remain unbounded.
|
||||
- FN-9178: Awaited core audit exclusions are classified with evidence in `docs/run-audit.md`: class A candidates, class B outcome-signalled candidates, and class C forensic/durability records. Transactional writers remain permanently unbounded because they share the mutation transaction.
|
||||
- FN-9186: self-healing emits `task:no-progress-no-task-done-requeue` for each bounded, backed-off zero-progress requeue and `task:no-progress-no-task-done-requeue-exhausted` once when its `taskDoneRetryCount` budget is spent. Metadata remains ids/counts/outcomes-only (`taskId`, column, attempt, maxAttempts, optional delayMs, outcome); the `NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED:` park is protected from generic terminal-failure and restart recovery, and audit emission is bounded best-effort.
|
||||
|
||||
@@ -2460,3 +2460,5 @@ Every populated or cleared task `branch` mutation declares `branchWriteOrigin`:
|
||||
Workspace landing derives its repository obligations from confirmed repository scope and durable per-repository evidence. A repository with a recorded `landedSha` remains an obligation on a finalize-once retry even when its current task-branch diff is empty. Undefined scope, duplicate repository declarations or worktree paths, and unexplained empty obligations fail closed; only an explicit commit-free task may take the no-op path.
|
||||
|
||||
Every land and recovery door evaluates graph-owned pre-merge blockers before changing merge state, acquiring leases, or writing Git. Recovery uses the persisted transient merge counter and reports scheduling separately from observed finalization. Main-checkout committed-work detection requires task ownership after the repository baseline; historical task-ID commits, recorded landings, and foreign shared-checkout commits do not become task violations. A scope revision atomically clears both its approval fingerprints and Code Review remediation target; a current-scope approval likewise clears that target, so a successor cannot inherit a stale remediation coordinator.
|
||||
|
||||
- **Bounded no-progress recovery (FN-9186):** `recoverNoProgressNoTaskDoneFailures()` spends the durable `taskDoneRetryCount` budget (maximum three) and writes the `recoveryRetryCount`/`nextRecoveryAt` display mirror using exponential backoff before requeuing a clean zero-progress failure. `recoveryRetryCount` cannot be the budget because terminal-failure recovery clears an expired mirror after a re-failure. On exhaustion the task stays failed with `NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED:`; the specific wedge descriptor prevents generic terminal-failure recovery from reopening it, and restart recovery also preserves the park. Manual Retry clears the error and counters to grant a fresh budget.
|
||||
|
||||
@@ -48,6 +48,8 @@ Reconciliation-scoped auto-recover/reclaim events the self-healing sweep surface
|
||||
| `task:auto-recover-paused-abort-park` | Self-healing clears a benign pause-abort operator park and requeues the task. |
|
||||
| `task:auto-rebound-paused-scope-decay` | Self-healing rebounds a task whose paused scope decayed past its floor, unblocking followers. |
|
||||
| `task:auto-archive-failure-budget-exhausted` | Self-healing abandons a repeatedly failing stale-task archive and surfaces it for operator action. |
|
||||
| `task:no-progress-no-task-done-requeue` | A zero-progress no-task-done failure consumes one bounded self-healing retry and records its backoff. Metadata is task ID, column, attempt, maximum, delay, and fixed outcome only. |
|
||||
| `task:no-progress-no-task-done-requeue-exhausted` | The bounded no-progress requeue budget parks a task once. Metadata is task ID, column, attempt, maximum, and fixed outcome only; bounded best-effort emission never gates the park. |
|
||||
| `task:reclaim-phantom-executor-binding` | Self-healing proves an in-memory executor-active binding is stale and requeues the task. |
|
||||
| `task:reconcile-orphaned-pending-step-results` | Self-healing rewrites orphaned `pending` workflow-step results (no live session) to `failed`. |
|
||||
| `task:reconcile-stale-duplicate-decision` | Self-healing clears a recurring duplicate-decision pause with no canonical target. |
|
||||
|
||||
@@ -54,7 +54,7 @@ Stages that cannot satisfy all three must either (a) tighten predicate to requir
|
||||
| recoverDriftedAgentTaskLinks | 6126 | agent assigned to terminal/missing task | n/a | task-link mismatch | clear assignment | RECONCILE-ONLY | keep | n/a | n/a |
|
||||
| recoverOrphanedAgents | 6180 | dead parent/direct-report linkage | n/a | org topology checks | pause/delete/reparent decisions | RECONCILE-ONLY | keep | n/a | n/a |
|
||||
| recoverStaleHeartbeatRuns | 6372 | stale heartbeat run records | run age thresholds | pid/age mismatch | terminate run records | RECONCILE-ONLY | keep | n/a | n/a |
|
||||
| recoverNoProgressNoTaskDoneFailures | 6451 | in-progress failed no-task-done no progress | implicit (no explicit grace) | no-step-progress + no git work + not executing | clear metadata + move to todo | BACKWARD | tighten | triple proof + no-progress checks + recent liveness-audit absence | gate move; emit `task:no-progress-no-task-done-no-action` |
|
||||
| recoverNoProgressNoTaskDoneFailures | 6451 | in-progress failed no-task-done no progress | `nextRecoveryAt` exponential backoff; maximum three attempts | no-step-progress + no git work + not executing + due budget | clear metadata + move to todo, then terminally park exhausted failures | BACKWARD | tightened (FN-9186) | triple proof + no-progress checks + persisted `taskDoneRetryCount` + due backoff | gate move; emit `task:no-progress-no-task-done-no-action`, bounded requeue audit, or one exhaustion audit |
|
||||
| recoverMissingWorktreeReviewFailures | 6516 | in-review failed OR merge-active (`merging`/`merging-pr`/`merging-fix`) session-start missing/unusable worktree | classifier-based | error classifier proof only | autoRecover requeue to todo | BACKWARD | tightened | triple proof + classifier proof + `allowsAutoMergeProcessing` + workspace-task exclusion | gate requeue; emit `task:missing-worktree-review-no-action` or `task:reconcile-missing-worktree-merge-active-no-action`; successful merge-active recovery clears `worktree`/`branch`/`sessionFile`, resets the worktree-session retry budget, increments `recoveryRetryCount` as the bounded stale-metadata clear counter, emits `task:reconcile-missing-worktree-merge-active`, and requeues to todo preserving progress |
|
||||
| recoverPartialProgressNoTaskDoneFailures | 6586 | in-review failed no-task-done with partial progress | bounded by `MAX_TASK_DONE_RETRIES` | no-task-done + partial progress + retry budget | clear error + move to todo preserveProgress | BACKWARD | tighten | triple proof + retry-budget predicates | gate move; emit `task:partial-progress-no-task-done-no-action` |
|
||||
| recoverApprovedTriageTasks | 6706 | triage planning specified stale | `APPROVED_TRIAGE_RECOVERY_GRACE_MS` | planning idle + valid PROMPT.md | recoverApprovedTriageTask callback | FORWARD | keep | n/a | n/a |
|
||||
|
||||
@@ -103,6 +103,24 @@ describe("default-workflow-hooks registry wiring", () => {
|
||||
expect(ctx.task.userPaused).toBe(true);
|
||||
});
|
||||
|
||||
/*
|
||||
FNXC:SelfHealing 2026-08-21-16:06:
|
||||
FN-9186 writes the no-progress backoff immediately before its engine wip-to-todo
|
||||
rebound. Only review-origin moves clear this display mirror, so this pins the
|
||||
move-hook contract that keeps the scheduler from immediately redispatching it.
|
||||
*/
|
||||
it("preserves no-progress recovery backoff on an engine wip-to-rebound move", () => {
|
||||
registerDefaultWorkflowHooks();
|
||||
const ctx = makeCtx({ fromColumn: "in-progress", toColumn: "todo", moveSource: "engine", options: { recoveryRehome: true } });
|
||||
ctx.task.recoveryRetryCount = 1;
|
||||
ctx.task.nextRecoveryAt = "2026-08-21T16:07:00.000Z";
|
||||
|
||||
applyDefaultWorkflowMoveEffects(ctx);
|
||||
|
||||
expect(ctx.task.recoveryRetryCount).toBe(1);
|
||||
expect(ctx.task.nextRecoveryAt).toBe("2026-08-21T16:07:00.000Z");
|
||||
});
|
||||
|
||||
it("preservePause never SETS a pause on an unpaused reopen, and default reopen still clears one", () => {
|
||||
registerDefaultWorkflowHooks();
|
||||
// preservePause on an unpaused task: nothing appears.
|
||||
|
||||
@@ -134,6 +134,49 @@ describe("FN-9175 non-executor audit sink health", () => {
|
||||
manager.stop();
|
||||
});
|
||||
|
||||
it.each(hostileModes)("keeps no-progress retry and exhaustion mutations bounded with a %s audit sink", async (mode) => {
|
||||
const sink = sinkFor(mode);
|
||||
const task = {
|
||||
id: `FN-9186-${mode}`,
|
||||
column: "in-progress",
|
||||
status: "failed",
|
||||
error: "Agent finished without calling fn_task_done: sandbox unavailable",
|
||||
paused: false,
|
||||
steps: [],
|
||||
};
|
||||
const store = {
|
||||
...sink.host,
|
||||
getSettings: vi.fn().mockResolvedValue({ maintenanceIntervalMs: 0 }),
|
||||
listTasks: vi.fn(async ({ column }: { column: string }) => column === task.column ? [{ ...task }] : []),
|
||||
updateTaskAtomic: vi.fn(async (_id: string, updater: (live: typeof task) => Partial<typeof task> | null) => {
|
||||
const patch = await updater({ ...task });
|
||||
if (patch) Object.assign(task, patch);
|
||||
return { ...task };
|
||||
}),
|
||||
moveTask: vi.fn(async () => { task.column = "todo"; }),
|
||||
logEntry: vi.fn().mockResolvedValue(undefined),
|
||||
getRootDir: vi.fn(() => "/tmp/fn-9175"),
|
||||
};
|
||||
const manager = new SelfHealingManager(store as any, { rootDir: "/tmp/fn-9175", getExecutingTaskIds: () => new Set() });
|
||||
vi.spyOn(manager as any, "hasRecoverableGitWork").mockResolvedValue(false);
|
||||
vi.spyOn(manager as any, "evaluateBackwardMoveTripleProof").mockResolvedValue({ ok: true });
|
||||
await expect(settleBounded(sink, () => manager.recoverNoProgressNoTaskDoneFailures())).resolves.toBe(1);
|
||||
expect(store.moveTask).toHaveBeenCalledOnce();
|
||||
|
||||
Object.assign(task, {
|
||||
column: "in-progress",
|
||||
status: "failed",
|
||||
error: "Agent finished without calling fn_task_done: sandbox unavailable",
|
||||
taskDoneRetryCount: 3,
|
||||
recoveryRetryCount: 3,
|
||||
nextRecoveryAt: undefined,
|
||||
});
|
||||
await expect(settleBounded(sink, () => manager.recoverNoProgressNoTaskDoneFailures())).resolves.toBe(0);
|
||||
expect(task.error).toMatch(/^NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED:/);
|
||||
expect(store.moveTask).toHaveBeenCalledOnce();
|
||||
manager.stop();
|
||||
});
|
||||
|
||||
it.each(hostileModes)("reclaims a wedged active merge with a %s audit sink", async (mode) => {
|
||||
const sink = sinkFor(mode);
|
||||
const task = { id: "FN-WEDGE", title: "wedged", description: "", column: "in-review", status: "reviewing", paused: false, dependencies: [], steps: [], updatedAt: "2026-01-01T00:00:00.000Z" };
|
||||
|
||||
@@ -10,6 +10,7 @@ import {
|
||||
isRecoverableMissingWorktreeReviewFailureNoProgress,
|
||||
isRecoverableMissingWorktreeReviewFailureWithProgress,
|
||||
} from "../healing/restart-recovery-coordinator.js";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "../healing/no-progress-requeue-budget.js";
|
||||
|
||||
function createTask(overrides: Partial<Task>): Task {
|
||||
return {
|
||||
@@ -184,6 +185,19 @@ describe("RestartRecoveryCoordinator", () => {
|
||||
expect(executor.resumeOrphaned).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
|
||||
it("does not reopen an exhausted no-progress park", async () => {
|
||||
const store = {
|
||||
listTasks: vi.fn().mockResolvedValue([
|
||||
createTask({ id: "FN-parked", status: "failed", error: `${NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX} 3/3 attempts spent. Agent finished without calling fn_task_done`, steps: [] }),
|
||||
]),
|
||||
updateTask: vi.fn(), logEntry: vi.fn(), moveTask: vi.fn(),
|
||||
} as unknown as TaskStore;
|
||||
const executor = { resumeOrphaned: vi.fn().mockResolvedValue(undefined) } as any;
|
||||
await new RestartRecoveryCoordinator(store, executor).recoverInterruptedRuns();
|
||||
expect(store.updateTask).not.toHaveBeenCalled();
|
||||
expect(store.moveTask).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("does not requeue when step progress exists", async () => {
|
||||
const store = {
|
||||
listTasks: vi.fn().mockResolvedValue([
|
||||
|
||||
@@ -0,0 +1,131 @@
|
||||
import { EventEmitter } from "node:events";
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { buildManualRetryResetPatch, type Task, type TaskStore } from "@fusion/core";
|
||||
import {
|
||||
MAX_TASK_DONE_RETRIES,
|
||||
SelfHealingManager,
|
||||
} from "../self-healing.js";
|
||||
import { classifyTerminalFailureAutoRecoveryForTask } from "../notification/task-wedge-notification.js";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "../healing/no-progress-requeue-budget.js";
|
||||
|
||||
function candidate(overrides: Partial<Task> = {}): Task {
|
||||
return {
|
||||
id: "FN-9186",
|
||||
column: "in-progress",
|
||||
status: "failed",
|
||||
error: "Agent finished without calling fn_task_done: sandbox unavailable",
|
||||
paused: false,
|
||||
steps: [],
|
||||
...overrides,
|
||||
} as Task;
|
||||
}
|
||||
|
||||
/** Production-shaped mutable store: atomic updates serialize competing maintenance owners. */
|
||||
function createStore(initial: Task) {
|
||||
let task = { ...initial };
|
||||
let tail = Promise.resolve();
|
||||
const emitter = new EventEmitter();
|
||||
const moveTask = vi.fn(async (_id: string, column: string) => {
|
||||
task = { ...task, column } as Task;
|
||||
return task;
|
||||
});
|
||||
const store = Object.assign(emitter, {
|
||||
getSettings: vi.fn().mockResolvedValue({ maintenanceIntervalMs: 60_000, autoRecovery: { mode: "on" } }),
|
||||
listTasks: vi.fn(async ({ column }: { column?: string } = {}) => !column || column === task.column ? [{ ...task }] : []),
|
||||
listWorkflowDefinitions: vi.fn().mockResolvedValue([]),
|
||||
getTask: vi.fn(async () => ({ ...task })),
|
||||
updateTaskAtomic: vi.fn(async (_id: string, updater: (live: Task) => Partial<Task> | null) => {
|
||||
const previous = tail;
|
||||
let release!: () => void;
|
||||
tail = new Promise<void>((resolve) => { release = resolve; });
|
||||
await previous;
|
||||
try {
|
||||
const patch = await updater({ ...task });
|
||||
if (patch) task = { ...task, ...patch } as Task;
|
||||
return { ...task };
|
||||
} finally {
|
||||
release();
|
||||
}
|
||||
}),
|
||||
moveTask,
|
||||
updateTask: vi.fn(),
|
||||
logEntry: vi.fn().mockResolvedValue(undefined),
|
||||
recordRunAuditEvent: vi.fn().mockResolvedValue(undefined),
|
||||
claimTerminalFailureAutoRecoveryAttempt: vi.fn(),
|
||||
applyTerminalFailureAutoRecoveryRetry: vi.fn(),
|
||||
getRootDir: vi.fn().mockReturnValue("/tmp/test-project"),
|
||||
}) as unknown as TaskStore & EventEmitter;
|
||||
return {
|
||||
store,
|
||||
moveTask,
|
||||
read: () => ({ ...task }),
|
||||
failAgain: () => { task = candidate({ taskDoneRetryCount: task.taskDoneRetryCount }); },
|
||||
applyManualRetryAndFailAgain: () => {
|
||||
task = candidate({ ...task, ...buildManualRetryResetPatch(), error: candidate().error });
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
describe("no-progress no-task_done recovery budget", () => {
|
||||
it("serializes competing sweeps so the persisted budget permits only three moves", async () => {
|
||||
const fixture = createStore(candidate());
|
||||
const first = new SelfHealingManager(fixture.store, { rootDir: "/tmp/test-project", getExecutingTaskIds: () => new Set() });
|
||||
const second = new SelfHealingManager(fixture.store, { rootDir: "/tmp/test-project", getExecutingTaskIds: () => new Set() });
|
||||
vi.spyOn(first as any, "hasRecoverableGitWork").mockResolvedValue(false);
|
||||
vi.spyOn(second as any, "hasRecoverableGitWork").mockResolvedValue(false);
|
||||
vi.spyOn(first as any, "evaluateBackwardMoveTripleProof").mockResolvedValue({ ok: true });
|
||||
vi.spyOn(second as any, "evaluateBackwardMoveTripleProof").mockResolvedValue({ ok: true });
|
||||
|
||||
for (let attempt = 0; attempt < 10; attempt += 1) {
|
||||
await Promise.all([first.recoverNoProgressNoTaskDoneFailures(), second.recoverNoProgressNoTaskDoneFailures()]);
|
||||
const current = fixture.read();
|
||||
if (current.taskDoneRetryCount && current.taskDoneRetryCount <= MAX_TASK_DONE_RETRIES && current.column === "todo") {
|
||||
fixture.failAgain();
|
||||
}
|
||||
}
|
||||
|
||||
expect(fixture.moveTask).toHaveBeenCalledTimes(MAX_TASK_DONE_RETRIES);
|
||||
expect(fixture.read().column).toBe("in-progress");
|
||||
expect(fixture.read().status).toBe("failed");
|
||||
expect(fixture.read().error?.startsWith(NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX)).toBe(true);
|
||||
expect(fixture.store.logEntry).toHaveBeenCalledTimes(MAX_TASK_DONE_RETRIES + 1);
|
||||
first.stop();
|
||||
second.stop();
|
||||
});
|
||||
|
||||
it("keeps the exhausted park out of terminal-failure recovery while ordinary failures retain that owner", async () => {
|
||||
const fixture = createStore(candidate({
|
||||
taskDoneRetryCount: MAX_TASK_DONE_RETRIES,
|
||||
error: `${NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX} 3/3 Agent finished without calling fn_task_done: sandbox unavailable`,
|
||||
}));
|
||||
const manager = new SelfHealingManager(fixture.store, { rootDir: "/tmp/test-project", getExecutingTaskIds: () => new Set() });
|
||||
|
||||
expect(classifyTerminalFailureAutoRecoveryForTask(fixture.read(), { autoRecoveryEnabled: true }))
|
||||
.toEqual({ action: "skip", reason: "not-generic-terminal-failure" });
|
||||
await expect(manager.autoRecoverTerminalFailures()).resolves.toBe(0);
|
||||
expect(fixture.store.claimTerminalFailureAutoRecoveryAttempt).not.toHaveBeenCalled();
|
||||
expect(fixture.store.applyTerminalFailureAutoRecoveryRetry).not.toHaveBeenCalled();
|
||||
expect(fixture.moveTask).not.toHaveBeenCalled();
|
||||
expect(fixture.read().error?.startsWith(NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX)).toBe(true);
|
||||
expect(fixture.read().column).toBe("in-progress");
|
||||
|
||||
const ordinary = createStore(candidate({ id: "FN-9186-ordinary", column: "todo", error: "opaque terminal failure" }));
|
||||
(ordinary.store.claimTerminalFailureAutoRecoveryAttempt as ReturnType<typeof vi.fn>)
|
||||
.mockResolvedValue({ outcome: "claimed", attempt: 1, applyToken: "test-token" });
|
||||
(ordinary.store.applyTerminalFailureAutoRecoveryRetry as ReturnType<typeof vi.fn>)
|
||||
.mockResolvedValue({ outcome: "applied" });
|
||||
const ordinaryManager = new SelfHealingManager(ordinary.store, { rootDir: "/tmp/test-project", getExecutingTaskIds: () => new Set() });
|
||||
await expect(ordinaryManager.autoRecoverTerminalFailures()).resolves.toBe(1);
|
||||
expect(ordinary.store.claimTerminalFailureAutoRecoveryAttempt).toHaveBeenCalledOnce();
|
||||
expect(ordinary.store.applyTerminalFailureAutoRecoveryRetry).toHaveBeenCalledOnce();
|
||||
ordinaryManager.stop();
|
||||
|
||||
fixture.applyManualRetryAndFailAgain();
|
||||
vi.spyOn(manager as any, "hasRecoverableGitWork").mockResolvedValue(false);
|
||||
vi.spyOn(manager as any, "evaluateBackwardMoveTripleProof").mockResolvedValue({ ok: true });
|
||||
await expect(manager.recoverNoProgressNoTaskDoneFailures()).resolves.toBe(1);
|
||||
expect(fixture.moveTask).toHaveBeenCalledTimes(1);
|
||||
expect(fixture.read().taskDoneRetryCount).toBe(1);
|
||||
manager.stop();
|
||||
});
|
||||
});
|
||||
@@ -141,7 +141,7 @@ vi.mock("../merger.js", () => ({
|
||||
classifyOwnedLandedEvidence: vi.fn(),
|
||||
}));
|
||||
|
||||
import { SelfHealingManager, isBranchAheadOfBase, MAX_AUTO_MERGE_RETRIES } from "../self-healing.js";
|
||||
import { SelfHealingManager, isBranchAheadOfBase, MAX_AUTO_MERGE_RETRIES, MAX_TASK_DONE_RETRIES } from "../self-healing.js";
|
||||
import { HEARTBEAT_ERROR_RECOVERY_METADATA_KEY, HEARTBEAT_ERROR_RETRY_EXHAUSTED_PAUSE_REASON, HEARTBEAT_ERROR_UNRECOVERABLE_PAUSE_REASON, readHeartbeatErrorRetryCount } from "../agent-heartbeat.js";
|
||||
import { PlanningLifecycleLockTransportError, TaskDeletedError, TaskNotFoundError, type TaskStore, type Settings, type Task, type AgentStore, type Agent, type NotificationProvider } from "@fusion/core";
|
||||
import { EventEmitter } from "node:events";
|
||||
@@ -2279,27 +2279,22 @@ describe("SelfHealingManager", () => {
|
||||
});
|
||||
vi.spyOn(managerWithRecovery as any, "hasRecoverableGitWork").mockReturnValue(false);
|
||||
|
||||
(store.listTasks as ReturnType<typeof vi.fn>).mockResolvedValue([
|
||||
{
|
||||
id: "FN-1473",
|
||||
column: "in-progress",
|
||||
status: "failed",
|
||||
error: "Agent finished without calling fn_task_done (after retry)",
|
||||
paused: false,
|
||||
steps: [],
|
||||
},
|
||||
]);
|
||||
const candidate = {
|
||||
id: "FN-1473",
|
||||
column: "in-progress",
|
||||
status: "failed",
|
||||
error: "Agent finished without calling fn_task_done (after retry)",
|
||||
paused: false,
|
||||
steps: [],
|
||||
};
|
||||
(store.listTasks as ReturnType<typeof vi.fn>).mockResolvedValue([candidate]);
|
||||
store.updateTaskAtomic = vi.fn(async (_id, updater) => ({ ...candidate, ...(await updater(candidate as Task)) })) as any;
|
||||
|
||||
const result = await managerWithRecovery.recoverNoProgressNoTaskDoneFailures();
|
||||
|
||||
expect(result).toBe(1);
|
||||
expect(store.listTasks).toHaveBeenCalledWith({ column: "in-progress", slim: true });
|
||||
expect(store.updateTask).toHaveBeenCalledWith("FN-1473", {
|
||||
status: "stuck-killed",
|
||||
worktree: null,
|
||||
branch: null,
|
||||
branchWriteOrigin: "engine",
|
||||
});
|
||||
expect(store.updateTaskAtomic).toHaveBeenCalledWith("FN-1473", expect.any(Function));
|
||||
expect(store.logEntry).toHaveBeenCalledWith(
|
||||
"FN-1473",
|
||||
expect.stringContaining("no-progress no-task_done failure"),
|
||||
@@ -2309,6 +2304,31 @@ describe("SelfHealingManager", () => {
|
||||
managerWithRecovery.stop();
|
||||
});
|
||||
|
||||
it("parks an exhausted no-progress budget once without moving the task", async () => {
|
||||
const managerWithRecovery = new SelfHealingManager(store, {
|
||||
rootDir: "/tmp/test-project",
|
||||
getExecutingTaskIds: () => new Set<string>(),
|
||||
});
|
||||
vi.spyOn(managerWithRecovery as any, "hasRecoverableGitWork").mockReturnValue(false);
|
||||
const candidate = {
|
||||
id: "FN-9186", column: "in-progress", status: "failed",
|
||||
error: "Agent finished without calling fn_task_done", taskDoneRetryCount: MAX_TASK_DONE_RETRIES,
|
||||
paused: false, steps: [],
|
||||
};
|
||||
(store.listTasks as ReturnType<typeof vi.fn>).mockResolvedValue([candidate]);
|
||||
store.updateTaskAtomic = vi.fn(async (_id, updater) => ({ ...candidate, ...(await updater(candidate as Task)) })) as any;
|
||||
|
||||
expect(await managerWithRecovery.recoverNoProgressNoTaskDoneFailures()).toBe(0);
|
||||
const exhaustedPatch = await (store.updateTaskAtomic as any).mock.calls[0][1](candidate);
|
||||
expect(exhaustedPatch).toEqual(expect.objectContaining({
|
||||
error: expect.stringMatching(/^NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED:/),
|
||||
recoveryRetryCount: null,
|
||||
nextRecoveryAt: null,
|
||||
}));
|
||||
expect(store.moveTask).not.toHaveBeenCalled();
|
||||
managerWithRecovery.stop();
|
||||
});
|
||||
|
||||
it("skips no-task_done failures with step progress", async () => {
|
||||
const managerWithRecovery = new SelfHealingManager(store, {
|
||||
rootDir: "/tmp/test-project",
|
||||
@@ -3781,7 +3801,6 @@ describe("SelfHealingManager", () => {
|
||||
error: null,
|
||||
worktreeSessionRetryCount: 1,
|
||||
worktree: liveWorktree,
|
||||
branch: "fusion/fn-3900",
|
||||
sessionFile: null,
|
||||
});
|
||||
/*
|
||||
@@ -3900,7 +3919,7 @@ describe("SelfHealingManager", () => {
|
||||
error: null,
|
||||
worktreeSessionRetryCount: 1,
|
||||
worktree: null,
|
||||
branch: expectedBranch,
|
||||
...(expectedBranch === branch ? {} : { branch: expectedBranch, branchWriteOrigin: "engine" }),
|
||||
sessionFile: null,
|
||||
});
|
||||
managerWithRecovery.stop();
|
||||
@@ -4021,6 +4040,7 @@ describe("SelfHealingManager", () => {
|
||||
worktreeSessionRetryCount: 1,
|
||||
worktree: null,
|
||||
branch: null,
|
||||
branchWriteOrigin: "engine",
|
||||
sessionFile: null,
|
||||
});
|
||||
expect(store.logEntry).toHaveBeenCalledWith(
|
||||
@@ -4229,7 +4249,6 @@ describe("SelfHealingManager", () => {
|
||||
expect(result).toBe(1);
|
||||
expect(store.updateTask).toHaveBeenCalledWith("FN-7802-WORKSPACE", expect.objectContaining({
|
||||
worktree: null,
|
||||
branch: null,
|
||||
sessionFile: null,
|
||||
}));
|
||||
expect(store.moveTask).toHaveBeenCalledWith("FN-7802-WORKSPACE", "todo", { preserveProgress: true, moveSource: "engine", recoveryRehome: true });
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
/**
|
||||
* FNXC:SelfHealing 2026-08-21-15:44:
|
||||
* Issue #3496 showed that no-progress task failures can be environmental and
|
||||
* repeat indefinitely. All recovery owners share this sentinel so an exhausted
|
||||
* park cannot be reopened by a differently scoped recovery path.
|
||||
*/
|
||||
export const NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX = "NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED:";
|
||||
@@ -3,6 +3,7 @@ import type { TaskExecutor } from "../executor.js";
|
||||
import { createLogger } from "../logger.js";
|
||||
import { resolveProjectColumnsForRoles, resolveReboundTargetForTask } from "@fusion/core";
|
||||
import { setImmediate as setImmediateCb } from "node:timers";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "./no-progress-requeue-budget.js";
|
||||
|
||||
/*
|
||||
FNXC:WorkflowResolvedColumns 2026-07-31-15:20 (fleet — the one arm left after main made the others required):
|
||||
@@ -222,7 +223,14 @@ export class RestartRecoveryCoordinator {
|
||||
}
|
||||
|
||||
private mustSafeRetry(task: Task): boolean {
|
||||
return isNoTaskDoneFailure(task) && !hasStepProgress(task);
|
||||
/*
|
||||
FNXC:SelfHealing 2026-08-21-15:44:
|
||||
Restart recovery is intentionally budget-free because it runs once per start,
|
||||
but must not erase #3496's terminal park and recreate the loop it did not own.
|
||||
*/
|
||||
return isNoTaskDoneFailure(task)
|
||||
&& !task.error?.startsWith(NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX)
|
||||
&& !hasStepProgress(task);
|
||||
}
|
||||
|
||||
private async safeRequeue(task: Task): Promise<void> {
|
||||
|
||||
@@ -2,7 +2,8 @@ import { describe, expect, it, vi } from "vitest";
|
||||
import { TaskNotFoundError, WEDGE_RENOTIFY_COOLDOWN_MS, type NotificationProvider, type Settings, type Task } from "@fusion/core";
|
||||
import { NotificationService } from "../notification-service.js";
|
||||
import { MAX_AUTO_MERGE_TRANSIENT_RETRIES } from "../../errors/transient-merge-error-classifier.js";
|
||||
import { describeSelfHealingNoActionWedge, describeTaskRecoveryOwner, describeTaskWedge, isTaskProgressing } from "../task-wedge-notification.js";
|
||||
import { classifyTerminalFailureAutoRecoveryForTask, describeSelfHealingNoActionWedge, describeTaskRecoveryOwner, describeTaskWedge, isTaskProgressing } from "../task-wedge-notification.js";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "../../healing/no-progress-requeue-budget.js";
|
||||
|
||||
type Listener = (task: Task) => void;
|
||||
|
||||
@@ -779,6 +780,13 @@ describe("task wedge notifications", () => {
|
||||
expect(describeTaskWedge(task({ error: "internal stack trace or opaque failure" }))).toMatchObject({ reasonKey: "terminal-failed" });
|
||||
});
|
||||
|
||||
it("keeps an exhausted no-progress requeue park specific despite preserved error text", () => {
|
||||
const { task } = fixture();
|
||||
const parked = task({ error: `${NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX} 3/3 attempts spent. BLOCKED: check:changeset-format tool failure` });
|
||||
expect(describeTaskWedge(parked)).toMatchObject({ reasonKey: "no-progress-requeue-budget-exhausted" });
|
||||
expect(classifyTerminalFailureAutoRecoveryForTask(parked, { autoRecoveryEnabled: true })).toEqual({ action: "skip", reason: "not-generic-terminal-failure" });
|
||||
});
|
||||
|
||||
it("changes the active reason into a new episode without raw error keys", async () => {
|
||||
const { store, service, sendMessageOnce, task } = fixture();
|
||||
await service.start();
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
import { classifyTerminalFailureAutoRecovery, type TerminalFailureAutoRecoveryDecision, type Task } from "@fusion/core";
|
||||
import { hasTransientMergeRecoveryOwner } from "../errors/transient-merge-error-classifier.js";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "../healing/no-progress-requeue-budget.js";
|
||||
|
||||
/** A bounded, operator-safe description of a task that cannot make progress. */
|
||||
export interface TaskWedgeDescriptor {
|
||||
@@ -202,6 +203,14 @@ export function describeTaskWedge(task: Task): TaskWedgeDescriptor | null {
|
||||
if (error.startsWith("EXECUTION_DISPATCH_LOOP_EXHAUSTED")) {
|
||||
return { reasonKey: "execution-dispatch-loop-exhausted", reason: "Execution re-queued without progress until its retry budget was exhausted.", action: "Retry, decompose, or rescope the task." };
|
||||
}
|
||||
/*
|
||||
FNXC:TaskWedgeNotifications 2026-08-21-15:44:
|
||||
#3496's sentinel must precede preserved failure text matchers. Otherwise the
|
||||
terminal owner treats this park as generic, clears error, and restarts the loop.
|
||||
*/
|
||||
if (error.startsWith(NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX)) {
|
||||
return { reasonKey: "no-progress-requeue-budget-exhausted", reason: "Self-healing exhausted its no-progress requeue budget.", action: "Repair the environment or task, then retry the task." };
|
||||
}
|
||||
if (error.includes("tool failure") || error.includes("Tool failure")) {
|
||||
return { reasonKey: "tool-failure-retry-exhausted", reason: "Execution tool-failure retries were exhausted.", action: "Inspect the failing tool and retry the task." };
|
||||
}
|
||||
|
||||
@@ -136,6 +136,7 @@ import {
|
||||
TERMINAL_FAILURE_CLAIM_APPLY_GRACE_MS,
|
||||
} from "@fusion/core";
|
||||
import { BASE_DELAY_MS, computeRecoveryDecision, formatDelay, MAX_DELAY_MS, MAX_RECOVERY_RETRIES } from "./healing/recovery-policy.js";
|
||||
import { NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX } from "./healing/no-progress-requeue-budget.js";
|
||||
|
||||
export {
|
||||
COMPLETED_BLOCKED_PAUSE_REASON,
|
||||
@@ -585,9 +586,15 @@ const ORPHANED_WITH_WORKTREE_GRACE_MS = 300_000;
|
||||
/**
|
||||
* Maximum times a task can be auto-requeued after the agent exits without
|
||||
* calling `fn_task_done`. Bounded so a persistently-broken task cannot loop
|
||||
* forever; when exhausted the task stays in `in-review` for human inspection.
|
||||
* forever; when exhausted the task stays failed in its wip lane for human inspection.
|
||||
*/
|
||||
const MAX_TASK_DONE_RETRIES = 3;
|
||||
/**
|
||||
* FNXC:SelfHealing 2026-08-21-15:44:
|
||||
* Issue #3496 requires a hard cap on no-progress automatic requeues. This
|
||||
* durable budget uses taskDoneRetryCount, not recoveryRetryCount, because the
|
||||
* terminal-failure owner clears the latter display mirror after each failure.
|
||||
*/
|
||||
export const MAX_TASK_DONE_RETRIES = 3;
|
||||
const RECONCILE_SCOPE_OVERRIDE_MERGE_ACTIVE_STATUS_SET = new Set<string>(MERGE_ACTIVE_MISSING_WORKTREE_STATUSES);
|
||||
/**
|
||||
* FNXC:WorkflowLifecycle 2026-06-20-00:00: single source of truth for the
|
||||
@@ -15316,7 +15323,8 @@ const movedTask = await this.store.moveTask(task.id, completeLane);
|
||||
!task.paused &&
|
||||
!executingIds.has(task.id) &&
|
||||
!isTaskWorkComplete(task) &&
|
||||
!hasStepProgress(task),
|
||||
!hasStepProgress(task) &&
|
||||
isRecoveryRetryDue(task, Date.now()),
|
||||
);
|
||||
|
||||
if (candidates.length === 0) return 0;
|
||||
@@ -15342,17 +15350,80 @@ const movedTask = await this.store.moveTask(task.id, completeLane);
|
||||
continue;
|
||||
}
|
||||
|
||||
await this.store.updateTask(task.id, {
|
||||
status: "stuck-killed",
|
||||
worktree: null,
|
||||
branch: null, branchWriteOrigin: "engine" as const,
|
||||
const auditor = createRunAuditor(this.store, {
|
||||
runId: generateSyntheticRunId("no-progress-no-task-done", task.id),
|
||||
agentId: "self-healing",
|
||||
taskId: task.id,
|
||||
taskLineageId: task.lineageId,
|
||||
phase: "no-progress-no-task-done",
|
||||
});
|
||||
await this.store.logEntry(
|
||||
task.id,
|
||||
"Auto-recovered no-progress no-task_done failure — clean worktree, moved back to todo",
|
||||
);
|
||||
// #1411: backward recovery — skip order-derived adjacency.
|
||||
const now = Date.now();
|
||||
let transition:
|
||||
| { kind: "retry"; prior: number; delayMs: number }
|
||||
| { kind: "exhausted"; prior: number }
|
||||
| undefined;
|
||||
/*
|
||||
FNXC:SelfHealing 2026-08-21-15:44:
|
||||
taskDoneRetryCount survives the terminal-failure mirror clear, unlike
|
||||
recoveryRetryCount. Claim the failed row under its store lock before the
|
||||
backward move: concurrent startup/maintenance sweeps otherwise read the
|
||||
same count and spend #3496's three-attempt cap more than once.
|
||||
*/
|
||||
await this.store.updateTaskAtomic(task.id, (live) => {
|
||||
if (
|
||||
live.status !== "failed" ||
|
||||
!isNoTaskDoneFailure(live) ||
|
||||
live.paused ||
|
||||
isTaskWorkComplete(live) ||
|
||||
hasStepProgress(live) ||
|
||||
!isRecoveryRetryDue(live, now)
|
||||
) return null;
|
||||
const prior = live.taskDoneRetryCount ?? 0;
|
||||
/*
|
||||
FNXC:SelfHealing 2026-08-21-15:44:
|
||||
The sentinel is an idempotence fence. It prevents later sweeps from
|
||||
duplicating the terminal log/audit escalation after #3496 is exhausted.
|
||||
*/
|
||||
if (prior >= MAX_TASK_DONE_RETRIES) {
|
||||
if (live.error?.startsWith(NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX)) return null;
|
||||
transition = { kind: "exhausted", prior };
|
||||
return {
|
||||
error: `${NO_PROGRESS_REQUEUE_BUDGET_EXHAUSTED_PREFIX} ${prior}/${MAX_TASK_DONE_RETRIES} attempts spent. ${live.error ?? ""}`,
|
||||
recoveryRetryCount: null,
|
||||
nextRecoveryAt: null,
|
||||
};
|
||||
}
|
||||
const delayMs = computeRecoveryDecision({ recoveryRetryCount: prior }).delayMs;
|
||||
transition = { kind: "retry", prior, delayMs };
|
||||
return {
|
||||
status: "stuck-killed",
|
||||
worktree: null,
|
||||
branch: null,
|
||||
branchWriteOrigin: "engine" as const,
|
||||
taskDoneRetryCount: prior + 1,
|
||||
recoveryRetryCount: prior + 1,
|
||||
nextRecoveryAt: new Date(now + delayMs).toISOString(),
|
||||
};
|
||||
});
|
||||
if (!transition) continue;
|
||||
if (transition.kind === "exhausted") {
|
||||
await this.store.logEntry(task.id, `No-progress no-task_done recovery exhausted after ${transition.prior}/${MAX_TASK_DONE_RETRIES} attempts; task remains failed for operator action`);
|
||||
await auditor.database({
|
||||
type: "task:no-progress-no-task-done-requeue-exhausted",
|
||||
target: task.id,
|
||||
metadata: { taskId: task.id, column: task.column, attempt: transition.prior, maxAttempts: MAX_TASK_DONE_RETRIES, outcome: "exhausted" },
|
||||
});
|
||||
continue;
|
||||
}
|
||||
|
||||
await this.store.logEntry(task.id, `Auto-recovered no-progress no-task_done failure — retry ${transition.prior + 1}/${MAX_TASK_DONE_RETRIES} in ${formatDelay(transition.delayMs)}, moved back to todo`);
|
||||
// #1411: the locked status claim fences duplicate backward moves before this public move acquires its own lock.
|
||||
await this.store.moveTask(task.id, await resolveReboundTargetForTask(this.store, task.id), { moveSource: "engine", recoveryRehome: true });
|
||||
await auditor.database({
|
||||
type: "task:no-progress-no-task-done-requeue",
|
||||
target: task.id,
|
||||
metadata: { taskId: task.id, column: task.column, attempt: transition.prior + 1, maxAttempts: MAX_TASK_DONE_RETRIES, delayMs: transition.delayMs, outcome: "requeued" },
|
||||
});
|
||||
recovered++;
|
||||
} catch (err: unknown) { const errorMessage = err instanceof Error ? err.message : String(err);
|
||||
log.error(`Failed to recover no-progress no-task_done failure ${task.id}: ${errorMessage}`);
|
||||
|
||||
@@ -533,6 +533,9 @@ export type DatabaseMutationType =
|
||||
*/
|
||||
| "task:auto-recover-terminal-failure"
|
||||
| "task:auto-recover-terminal-failure-exhausted"
|
||||
/** Metadata: { taskId, column, attempt, maxAttempts, delayMs?, outcome } — ids/counts/outcomes only. */
|
||||
| "task:no-progress-no-task-done-requeue"
|
||||
| "task:no-progress-no-task-done-requeue-exhausted"
|
||||
| "task:auto-recover-finalize-already-on-main"
|
||||
/** Metadata: { taskId, previousColumn, targetColumn, commitSha, status, blockedBy, overlapBlockedBy, reason } */
|
||||
| "task:auto-merge-finalize-column-mismatch-reconciled"
|
||||
|
||||
Reference in New Issue
Block a user