FN-7672: recover durable agents stuck in error state despite active manager
Root-causes 4 correlated CTO-report agent failures where durable non-ephemeral agents got stuck in `error` state indefinitely because the heartbeat scheduler stops ticking error-state agents entirely, and self-healing's recovery sweep previously only considered them when their manager row was missing. - SelfHealingManager: scope the `managerMissing` gate to the "running" orphan-detection path only, so "error"-state durable agents with a present/active manager now fall through to the existing transient/operator-actionable/active-execution/cooldown/retry-budget recovery guards instead of being skipped outright - Add FNXC:AgentHeartbeat comment documenting the FN-7672 incident and rationale for the scoping change - Extend self-healing.test.ts with coverage for manager-present durable agents in error state - Add changeset (patch) describing the fix for release notes - Update docs/agents.md accordingly Files changed: .changeset/fn-7672-durable-agent-recovery.md | 7 ++ docs/agents.md | 2 + packages/engine/src/__tests__/self-healing.test.ts | 129 ++++++++++++++++++++- packages/engine/src/self-healing.ts | 24 +++- 4 files changed, 160 insertions(+), 2 deletions(-) Fusion-Task-Id: FN-7672 Fusion-Task-Lineage: 6676dc9e-66e7-4f70-804a-cccf77e8d337 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
7
.changeset/fn-7672-durable-agent-recovery.md
Normal file
7
.changeset/fn-7672-durable-agent-recovery.md
Normal file
@@ -0,0 +1,7 @@
|
|||||||
|
---
|
||||||
|
"@runfusion/fusion": patch
|
||||||
|
---
|
||||||
|
|
||||||
|
summary: Auto-recover durable agents stuck in transient error state even when their manager is active.
|
||||||
|
category: fix
|
||||||
|
dev: SelfHealingManager durable-error recovery no longer requires a missing manager; manager-present durable non-ephemeral agents with a transient lastError and no active run are recovered under the existing cooldown/backoff/retry-budget guards (FN-7672).
|
||||||
@@ -364,6 +364,8 @@ On restart attempts, the runtime triggers the normal heartbeat pipeline with `so
|
|||||||
|
|
||||||
Self-healing intentionally leaves agents in `error` (no auto-restart) when blockers are operator-actionable or non-transient, when cooldown has not elapsed, when active execution is present, or when retry budget is exhausted.
|
Self-healing intentionally leaves agents in `error` (no auto-restart) when blockers are operator-actionable or non-transient, when cooldown has not elapsed, when active execution is present, or when retry budget is exhausted.
|
||||||
|
|
||||||
|
**Manager presence does not gate this sweep (FN-7672):** eligibility for durable `state="error"` recovery does *not* depend on whether the agent's `reportsTo` manager is present/active. `HeartbeatTriggerScheduler` clears timers entirely once an agent enters `state="error"`, so this recovery sweep is the *only* path back to a healthy heartbeat for a durable agent stuck in `error` — a present manager does not make the agent any less stuck. (A separate, unrelated `managerMissing` check still gates recovery of orphaned `state="running"` agents — a different failure mode where a live process's manager row was deleted.) FN-7672 root-caused a correlated 4-agent error cluster reporting to one active manager (a transient upstream auth/session blip) that could never have self-healed under the old manager-missing-only gate, even once the underlying cause resolved.
|
||||||
|
|
||||||
- **Timer trigger:** run completes and the durable agent returns to `state="active"` (recoverable soft-fail).
|
- **Timer trigger:** run completes and the durable agent returns to `state="active"` (recoverable soft-fail).
|
||||||
- **Assignment / on-demand trigger:** run completes with `resultJson.actionRequired = true`, then the durable agent is paused with `pauseReason="heartbeat-model-unavailable"` and `lastError` set to actionable credential guidance (including the missing provider name when detectable).
|
- **Assignment / on-demand trigger:** run completes with `resultJson.actionRequired = true`, then the durable agent is paused with `pauseReason="heartbeat-model-unavailable"` and `lastError` set to actionable credential guidance (including the missing provider name when detectable).
|
||||||
|
|
||||||
|
|||||||
@@ -920,7 +920,7 @@ describe("SelfHealingManager", () => {
|
|||||||
expect(result).toBe(0);
|
expect(result).toBe(0);
|
||||||
});
|
});
|
||||||
|
|
||||||
it("skips agents with valid manager", async () => {
|
it("skips a manager-present error-state agent whose lastError is not transient (default permanent classification)", async () => {
|
||||||
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
const now = Date.now();
|
const now = Date.now();
|
||||||
const agentStore = createMockAgentStore([
|
const agentStore = createMockAgentStore([
|
||||||
@@ -931,11 +931,138 @@ describe("SelfHealingManager", () => {
|
|||||||
|
|
||||||
const result = await managerWithAgents.recoverOrphanedAgents();
|
const result = await managerWithAgents.recoverOrphanedAgents();
|
||||||
|
|
||||||
|
// No lastError at all classifies as "permanent" (default), so this agent
|
||||||
|
// is correctly skipped — but via the transient-classification guard, not
|
||||||
|
// because its manager is present. See the "manager-present" tests below
|
||||||
|
// for FN-7672's actual invariant: manager presence alone must no longer
|
||||||
|
// exclude a durable error-state agent from the recovery sweep.
|
||||||
expect(result).toBe(0);
|
expect(result).toBe(0);
|
||||||
expect(agentStore.updateAgent).not.toHaveBeenCalled();
|
expect(agentStore.updateAgent).not.toHaveBeenCalled();
|
||||||
managerWithAgents.stop();
|
managerWithAgents.stop();
|
||||||
});
|
});
|
||||||
|
|
||||||
|
/*
|
||||||
|
* FNXC:AgentHeartbeat 2026-07-08-12:20:
|
||||||
|
* FN-7672 regression: 4 of the CTO's 6 durable direct reports went
|
||||||
|
* simultaneously `error` (shared upstream auth/session blip) while
|
||||||
|
* reporting to an ACTIVE/PRESENT CTO. Because HeartbeatTriggerScheduler
|
||||||
|
* clears timers on `state === "error"`, the only way back to a healthy
|
||||||
|
* heartbeat is this recovery sweep — and it previously required
|
||||||
|
* `managerMissing`, so a manager-present durable agent could never
|
||||||
|
* self-heal even with a genuinely transient cause. These tests assert
|
||||||
|
* the manager-present path is now considered (subject to all existing
|
||||||
|
* guards, unweakened).
|
||||||
|
*/
|
||||||
|
it("recovers a transient error-state agent even when its manager is present and active", async () => {
|
||||||
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
|
const now = Date.now();
|
||||||
|
const agentStore = createMockAgentStore([
|
||||||
|
{ id: "manager-1", state: "active", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{
|
||||||
|
id: "report-1",
|
||||||
|
state: "error",
|
||||||
|
reportsTo: "manager-1",
|
||||||
|
lastError: "socket hang up",
|
||||||
|
metadata: {},
|
||||||
|
updatedAt: new Date(now - 120_000).toISOString(),
|
||||||
|
} as Agent,
|
||||||
|
]);
|
||||||
|
const restartDurableAgentHeartbeat = vi.fn().mockResolvedValue(true);
|
||||||
|
const managerWithAgents = new SelfHealingManager(store, {
|
||||||
|
rootDir: "/tmp/test-project",
|
||||||
|
agentStore,
|
||||||
|
restartDurableAgentHeartbeat,
|
||||||
|
});
|
||||||
|
|
||||||
|
const result = await managerWithAgents.recoverOrphanedAgents();
|
||||||
|
|
||||||
|
expect(result).toBe(1);
|
||||||
|
expect(agentStore.updateAgentState).toHaveBeenCalledWith("report-1", "active");
|
||||||
|
expect(agentStore.updateAgent).toHaveBeenLastCalledWith("report-1", { lastError: undefined });
|
||||||
|
expect(restartDurableAgentHeartbeat).toHaveBeenCalledWith("report-1", { reason: "transient-error", attempt: 1 });
|
||||||
|
managerWithAgents.stop();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("does NOT auto-recover a manager-present agent whose error is operator-actionable (FN-7672 auth-credential cluster shape)", async () => {
|
||||||
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
|
const now = Date.now();
|
||||||
|
const agentStore = createMockAgentStore([
|
||||||
|
{ id: "manager-1", state: "active", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{
|
||||||
|
id: "report-auth",
|
||||||
|
state: "error",
|
||||||
|
reportsTo: "manager-1",
|
||||||
|
lastError:
|
||||||
|
'Error: 401 {"type":"error","error":{"type":"authentication_error","message":"Invalid authentication credentials"},"request_id":"req_011CcpL6f3iXHxeHfMUjg9o8"}',
|
||||||
|
updatedAt: new Date(now - 120_000).toISOString(),
|
||||||
|
} as Agent,
|
||||||
|
]);
|
||||||
|
const managerWithAgents = new SelfHealingManager(store, { rootDir: "/tmp/test-project", agentStore });
|
||||||
|
|
||||||
|
const result = await managerWithAgents.recoverOrphanedAgents();
|
||||||
|
|
||||||
|
expect(result).toBe(0);
|
||||||
|
expect(agentStore.updateAgentState).not.toHaveBeenCalled();
|
||||||
|
managerWithAgents.stop();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("recovers only the eligible manager-present agent among a mixed cluster without touching healthy siblings", async () => {
|
||||||
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
|
const now = Date.now();
|
||||||
|
const agentStore = createMockAgentStore([
|
||||||
|
{ id: "cto", state: "active", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{ id: "sibling-healthy-1", state: "active", reportsTo: "cto", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{ id: "sibling-healthy-2", state: "active", reportsTo: "cto", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{
|
||||||
|
id: "report-transient",
|
||||||
|
state: "error",
|
||||||
|
reportsTo: "cto",
|
||||||
|
lastError: "socket hang up",
|
||||||
|
metadata: {},
|
||||||
|
updatedAt: new Date(now - 120_000).toISOString(),
|
||||||
|
} as Agent,
|
||||||
|
{
|
||||||
|
id: "report-auth-1",
|
||||||
|
state: "error",
|
||||||
|
reportsTo: "cto",
|
||||||
|
lastError: "Invalid authentication credentials",
|
||||||
|
updatedAt: new Date(now - 120_000).toISOString(),
|
||||||
|
} as Agent,
|
||||||
|
]);
|
||||||
|
const restartDurableAgentHeartbeat = vi.fn().mockResolvedValue(true);
|
||||||
|
const managerWithAgents = new SelfHealingManager(store, {
|
||||||
|
rootDir: "/tmp/test-project",
|
||||||
|
agentStore,
|
||||||
|
restartDurableAgentHeartbeat,
|
||||||
|
});
|
||||||
|
|
||||||
|
const result = await managerWithAgents.recoverOrphanedAgents();
|
||||||
|
|
||||||
|
expect(result).toBe(1);
|
||||||
|
expect(agentStore.updateAgentState).toHaveBeenCalledTimes(1);
|
||||||
|
expect(agentStore.updateAgentState).toHaveBeenCalledWith("report-transient", "active");
|
||||||
|
expect(agentStore.updateAgentState).not.toHaveBeenCalledWith("sibling-healthy-1", expect.anything());
|
||||||
|
expect(agentStore.updateAgentState).not.toHaveBeenCalledWith("sibling-healthy-2", expect.anything());
|
||||||
|
expect(agentStore.updateAgentState).not.toHaveBeenCalledWith("report-auth-1", expect.anything());
|
||||||
|
managerWithAgents.stop();
|
||||||
|
});
|
||||||
|
|
||||||
|
it("still gates a manager-missing RUNNING orphan on managerMissing (unchanged behavior)", async () => {
|
||||||
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
|
const now = Date.now();
|
||||||
|
const agentStore = createMockAgentStore([
|
||||||
|
{ id: "manager-1", state: "active", updatedAt: new Date(now).toISOString() } as Agent,
|
||||||
|
{ id: "running-1", state: "running", reportsTo: "manager-1", updatedAt: new Date(now - 120_000).toISOString() } as Agent,
|
||||||
|
]);
|
||||||
|
const managerWithAgents = new SelfHealingManager(store, { rootDir: "/tmp/test-project", agentStore });
|
||||||
|
|
||||||
|
const result = await managerWithAgents.recoverOrphanedAgents();
|
||||||
|
|
||||||
|
expect(result).toBe(0);
|
||||||
|
expect(agentStore.updateAgentState).not.toHaveBeenCalled();
|
||||||
|
managerWithAgents.stop();
|
||||||
|
});
|
||||||
|
|
||||||
it("recovers orphaned agent in transient error state", async () => {
|
it("recovers orphaned agent in transient error state", async () => {
|
||||||
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
vi.mocked(store.getSettings).mockResolvedValue({ taskStuckTimeoutMs: 60_000 } as unknown as Settings);
|
||||||
const now = Date.now();
|
const now = Date.now();
|
||||||
|
|||||||
@@ -10143,8 +10143,30 @@ export class SelfHealingManager {
|
|||||||
if (agent.state !== "running" && agent.state !== "error") {
|
if (agent.state !== "running" && agent.state !== "error") {
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
|
/*
|
||||||
|
* FNXC:AgentHeartbeat 2026-07-08-12:20:
|
||||||
|
* FN-7672: 4 of the CTO's 6 durable direct reports went simultaneously
|
||||||
|
* `error` (correlated auth/session blip) and stayed stuck for hours —
|
||||||
|
* HeartbeatTriggerScheduler clears timers entirely on `state === "error"`
|
||||||
|
* (isTickableState excludes it), so a durable error-state agent can ONLY
|
||||||
|
* come back via this recovery sweep; there is no natural self-heal via
|
||||||
|
* the normal tick loop. The `managerMissing` gate below previously
|
||||||
|
* applied uniformly to BOTH orphaned "running" agents (a genuinely
|
||||||
|
* different failure mode — a live process whose manager row vanished)
|
||||||
|
* AND "error" agents, which meant a durable agent in `error` with a
|
||||||
|
* present/active manager was structurally never even considered for
|
||||||
|
* recovery, regardless of how transient its `lastError` was. That is
|
||||||
|
* the systemic gap: manager presence has no bearing on whether a
|
||||||
|
* durable agent's own error is transient and safe to retry. Restrict
|
||||||
|
* `managerMissing` to the "running" orphan-detection path (unchanged
|
||||||
|
* behavior) and let "error" state proceed to the existing transient /
|
||||||
|
* operator-actionable / active-execution / cooldown / retry-budget
|
||||||
|
* guards below, which already exist specifically to prevent restart
|
||||||
|
* loops on genuinely broken (non-transient/operator-actionable)
|
||||||
|
* credentials — those guards are NOT weakened here.
|
||||||
|
*/
|
||||||
const managerMissing = !agent.reportsTo || !allAgentIds.has(agent.reportsTo);
|
const managerMissing = !agent.reportsTo || !allAgentIds.has(agent.reportsTo);
|
||||||
if (!managerMissing) {
|
if (agent.state === "running" && !managerMissing) {
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
const updatedAt = Date.parse(agent.updatedAt ?? "");
|
const updatedAt = Date.parse(agent.updatedAt ?? "");
|
||||||
|
|||||||
Reference in New Issue
Block a user