FN-7998: add executor alternate model escalation

Add opt-in executor escalation after same-model tool-failure retries are exhausted.

- Persist escalation settings and one-shot task state across SQLite and PostgreSQL stores.
- Retry once on a configured alternate model or scheduler node and audit escalation outcomes.
- Expose escalation controls, documentation, translations, migration, and regression coverage.

Files changed:
 .changeset/fn-7998-executor-escalation.md          |   7 ++
 AGENTS.md                                          |   1 +
 docs/settings-reference.md                         |  13 ++-
 .../core/src/__tests__/settings-defaults.test.ts   |  23 ++++-
 packages/core/src/in-review-stall.ts               |  29 ++++++
 packages/core/src/index.gate.ts                    |   3 +-
 packages/core/src/index.ts                         |   3 +-
 packages/core/src/manual-retry-reset.ts            |   1 +
 .../0014_executor_escalation_attempt.sql           |   2 +
 packages/core/src/postgres/schema-applier.ts       |  17 ++++
 packages/core/src/postgres/schema/project.ts       |   1 +
 packages/core/src/settings-schema.ts               |   4 +
 packages/core/src/store.ts                         |   2 +-
 packages/core/src/task-store/persistence.ts        |   2 +
 packages/core/src/task-store/remaining-ops-2.ts    |   2 +-
 packages/core/src/task-store/remaining-ops-3.ts    |   2 +-
 packages/core/src/task-store/remaining-ops-6.ts    |   2 +-
 packages/core/src/task-store/serialization.ts      |   1 +
 packages/core/src/task-store/task-update.ts        |   2 +
 packages/core/src/types.ts                         |  13 +++
 .../dashboard/app/components/SettingsModal.tsx     |  12 +++
 .../app/components/settings/section-keys.ts        |   4 +
 .../settings/sections/SchedulingSection.search.ts  |  36 +++++++
 .../settings/sections/SchedulingSection.tsx        |   6 ++
 .../settings-default-descriptions.test.tsx         |   4 +
 .../__tests__/executor-tool-failure-retry.test.ts  |  91 +++++++++++++++++-
 packages/engine/src/executor.ts                    | 104 +++++++++++++++++++--
 packages/i18n/locales/en/app.json                  |   8 ++
 28 files changed, 376 insertions(+), 19 deletions(-)

Fusion-Task-Id: FN-7998

Fusion-Task-Lineage: bbce767d-c61a-4667-be62-abc0cc54d8be

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
gsxdsm
2026-07-16 14:31:59 -07:00
parent 4df5c62854
commit d870878a23
28 changed files with 376 additions and 19 deletions

View File

@@ -0,0 +1,7 @@
---
"@runfusion/fusion": minor
---
summary: Optionally escalate an executor run to a stronger model or configured node after same-model retries are exhausted.
category: feature
dev: New opt-in project settings provide one model/node escalation attempt with durable task state and audit events.

View File

@@ -276,6 +276,7 @@ Scoped exception (FN-5819): shared-branch-group members (`branchContext.assignme
- FN-7514: the planner overseer's per-task oversight loop (`PlannerRecoveryController.tick`) emits `overseer:oversight-withheld-human-control` when the pure `evaluateOverseerHumanControl` guard withholds ALL oversight action (no steering, retry, targeted-fix, or pending confirmation) for a task that is user-paused (`task.userPaused===true`, or `task.paused===true` with no `pausedReason`) or ineligible for auto-merge processing per `allowsAutoMergeProcessing` (`autoMerge:false`/PR-based human-review terminal contract). The guard runs BEFORE FN-7513's confirmation classification, so a withheld task never records a pending confirmation. Metadata: `{ taskId, reason: "user-paused" | "auto-merge-off-human-review", stage, oversightLevel }`; deduped per (taskId, withheld reason) so it is not re-emitted every poll while the reason is unchanged.
- FN-7720: `TaskStore.bypassFailedPreMergeReviewStep` emits `task:bypass-review` when a privileged operator bypasses the latest failed pre-merge review step of an `in-review` task; metadata includes `workflowStepId`, `workflowStepName`, `bypassedFromStatus`, `bypassedFromVerdict`, and the mandatory `reason`. The bypass rewrites the step's `status` to `"skipped"` with `bypassedBy`/`bypassedAt`/`bypassReason`/`bypassedFromStatus` fields; it never fabricates a reviewer `verdict` and clears only the failed-pre-merge-step `getTaskMergeBlocker` reason. Reachable via `fn_task_bypass_review` (CLI/pi-extension operator tool surface only — not executor/reviewer/triage) and `POST /tasks/:id/bypass-review`.
- FN-7996: executor emits `task:execution-tool-failure-retry` for a claimed same-model consecutive-tool-failure retry and `task:execution-tool-failure-retry-exhausted` when the matching run budget is spent. Metadata is ids/counts/outcomes-only; the exhausted event is emitted once through a project-scoped compare-and-set while terminal parking remains idempotent.
- FN-7998: executor emits `task:execution-escalation-retry` when its opt-in, single alternate model/node attempt is persisted after FN-7996 exhaustion, and `task:execution-escalation-exhausted` when that attempt also reaches the terminal park. Metadata remains ids/counts/outcomes-only (`taskId`, graph node id, target booleans, and prior retry count); no model identifiers or prose are persisted in run-audit.
- FN-8004: `agent:heartbeat-move-skipped-soft-delete` records a heartbeat move that races a soft-deleted task without parking the durable agent. Metadata remains ids/timestamps/source only (`agentId`, optional `taskId`/`deletedAt`, `moveAttemptedAt`, optional `source`); it never stores error prose.

View File

@@ -1755,4 +1755,15 @@ Standardize executor/validator pairs; auto-selectable by task size (Small → Bu
| `executorToolFailureRetryBackoffMs` | integer, `2000` | Unref'd delay before the rerun. |
| `executorToolFailureThreshold` | integer, `3` | Consecutive terminal tool failures required to qualify. |
Values are project-scoped and finite values are floored; count/backoff must be at least `0`, and threshold at least `1`, otherwise their defaults apply. The executor evaluates this bounded policy before its terminal graph-failure park: it counts `tool_error` completion entries, resets only on `tool_result`, and ignores `tool` invocation markers. The detector is scoped to the current executor-run agent-log cursor. Its project-scoped atomic claim prevents concurrent retries and classifies cursor mismatch before an exhausted cap so stale handlers do not park newer work. The exhausted audit is compare-and-set deduplicated while the terminal park remains idempotent. FN-7998 may extend this same-model foundation with escalation.
Values are project-scoped and finite values are floored; count/backoff must be at least `0`, and threshold at least `1`, otherwise their defaults apply. The executor evaluates this bounded policy before its terminal graph-failure park: it counts `tool_error` completion entries, resets only on `tool_result`, and ignores `tool` invocation markers. The detector is scoped to the current executor-run agent-log cursor. Its project-scoped atomic claim prevents concurrent retries and classifies cursor mismatch before an exhausted cap so stale handlers do not park newer work. The exhausted audit is compare-and-set deduplicated while the terminal park remains idempotent.
### Executor escalation after tool-failure retry exhaustion
| Setting | Type/default | Behavior |
| --- | --- | --- |
| `executorModelEscalationEnabled` | boolean, `false` | Opt in to one alternate attempt after same-model retries exhaust. |
| `executorEscalationProvider` | string, unset | Provider for an alternate model; requires `executorEscalationModelId`. |
| `executorEscalationModelId` | string, unset | Alternate model ID; requires `executorEscalationProvider`. |
| `executorEscalationNodeId` | string, unset | Optional configured node target. |
Escalation is enabled only when the toggle is true and either a complete provider/model pair or a node ID is configured. It is single-shot: after FN-7996 exhausts same-model retries, Fusion persists the override and tries once before the existing terminal park. The alternate model enters the [model-selection hierarchy](#model-selection-hierarchy) as a task-level override; a node target enters `resolveEffectiveNode` as a task-level routing override and is requeued so scheduler routing is recalculated. This remains opt-in by default to avoid unexpected model cost or execution behavior. Column-agent overrides still govern their sessions and can supersede a task-level model target.

View File

@@ -1,5 +1,5 @@
import { afterEach, describe, expect, it, vi } from "vitest";
import { CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD, DEFAULT_CONSECUTIVE_TOOL_FAILURE_RETRY_BACKOFF_MS, DEFAULT_MAX_CONSECUTIVE_TOOL_FAILURE_RETRIES, DEFAULT_MAX_AUTO_MERGE_RETRIES, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries } from "../in-review-stall.js";
import { CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD, DEFAULT_CONSECUTIVE_TOOL_FAILURE_RETRY_BACKOFF_MS, DEFAULT_MAX_CONSECUTIVE_TOOL_FAILURE_RETRIES, DEFAULT_MAX_AUTO_MERGE_RETRIES, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveExecutorEscalationTarget, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries } from "../in-review-stall.js";
import { isExperimentalFeatureEnabled } from "../experimental-features.js";
import { DEFAULT_GLOBAL_SETTINGS, DEFAULT_PROJECT_SETTINGS, GLOBAL_SETTINGS_KEYS, PROJECT_SETTINGS_KEYS, isGlobalOnlySettingsKey } from "../settings-schema.js";
import { isWorkflowColumnsEnabled } from "../workflow-columns-settings.js";
@@ -64,6 +64,27 @@ describe("settings defaults invariants", () => {
expect(resolveMaxAutoMergeRetries({ maxAutoMergeRetries: Number.NaN })).toBe(3);
});
it("defaults executor escalation off and resolves only complete opt-in targets", () => {
expect(DEFAULT_PROJECT_SETTINGS.executorModelEscalationEnabled).toBe(false);
expect(PROJECT_SETTINGS_KEYS).toEqual(expect.arrayContaining([
"executorModelEscalationEnabled",
"executorEscalationProvider",
"executorEscalationModelId",
"executorEscalationNodeId",
]));
expect(resolveExecutorEscalationTarget({
executorModelEscalationEnabled: false,
executorEscalationProvider: "anthropic",
executorEscalationModelId: "claude",
executorEscalationNodeId: "node-1",
}).enabled).toBe(false);
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true })).toEqual({ enabled: false });
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic" })).toEqual({ enabled: false });
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude" })).toEqual({ enabled: true, provider: "anthropic", modelId: "claude" });
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationNodeId: "node-1" })).toEqual({ enabled: true, nodeId: "node-1" });
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude", executorEscalationNodeId: "node-1" })).toEqual({ enabled: true, provider: "anthropic", modelId: "claude", nodeId: "node-1" });
});
it("resolves worktrunk as disabled when both scopes are unset or empty", () => {
expect(resolveWorktrunkSettings(undefined, undefined).enabled).toBe(false);
expect(resolveWorktrunkSettings({}, {}).enabled).toBe(false);

View File

@@ -73,6 +73,35 @@ export function resolveConsecutiveToolFailureThreshold(settings?: { executorTool
return Number.isFinite(numeric) && Math.floor(numeric) >= 1 ? Math.floor(numeric) : CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD;
}
export interface ExecutorEscalationTarget {
enabled: boolean;
provider?: string;
modelId?: string;
nodeId?: string;
}
/**
* FNXC:ExecutorEscalation 2026-07-16-21:00:
* Escalation remains opt-in and only accepts a complete model pair or a node target. Keeping this resolver in core makes engine and settings surfaces agree that incomplete targets silently preserve FN-7996 terminal behavior.
*/
export function resolveExecutorEscalationTarget(settings?: {
executorModelEscalationEnabled?: unknown;
executorEscalationProvider?: unknown;
executorEscalationModelId?: unknown;
executorEscalationNodeId?: unknown;
} | null): ExecutorEscalationTarget {
const provider = typeof settings?.executorEscalationProvider === "string" ? settings.executorEscalationProvider.trim() : "";
const modelId = typeof settings?.executorEscalationModelId === "string" ? settings.executorEscalationModelId.trim() : "";
const nodeId = typeof settings?.executorEscalationNodeId === "string" ? settings.executorEscalationNodeId.trim() : "";
const hasModelTarget = Boolean(provider && modelId);
const hasNodeTarget = Boolean(nodeId);
return {
enabled: settings?.executorModelEscalationEnabled === true && (hasModelTarget || hasNodeTarget),
...(hasModelTarget ? { provider, modelId } : {}),
...(hasNodeTarget ? { nodeId } : {}),
};
}
export const IN_REVIEW_STALL_LOG_PREFIX = "In-review stall surfaced [";
export const IN_REVIEW_STALL_DEADLOCK_LOG_PREFIX = "In-review stall auto-disposed [";
export const IN_REVIEW_STALL_TERMINAL_LOG_PREFIX = "In-review stall terminal disposed [";

View File

@@ -993,8 +993,9 @@ export {
resolveMaxConsecutiveToolFailureRetries,
resolveConsecutiveToolFailureRetryBackoffMs,
resolveConsecutiveToolFailureThreshold,
resolveExecutorEscalationTarget,
} from "./in-review-stall.js";
export type { InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
export type { ExecutorEscalationTarget, InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
export {
getStalePausedReviewSignal,
DEFAULT_STALE_PAUSED_REVIEW_THRESHOLD_MS,

View File

@@ -1027,8 +1027,9 @@ export {
resolveMaxConsecutiveToolFailureRetries,
resolveConsecutiveToolFailureRetryBackoffMs,
resolveConsecutiveToolFailureThreshold,
resolveExecutorEscalationTarget,
} from "./in-review-stall.js";
export type { InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
export type { ExecutorEscalationTarget, InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
export {
getStalePausedReviewSignal,
DEFAULT_STALE_PAUSED_REVIEW_THRESHOLD_MS,

View File

@@ -43,6 +43,7 @@ export function buildAutoPauseClearPatch(
export function buildManualRetryResetPatch(options?: { resetMergeRetries?: boolean }): Partial<Task> {
const patch: Partial<Task> = {
nextRecoveryAt: null as unknown as Task["nextRecoveryAt"],
executorEscalationAttempted: false,
toolFailureDetectorLogCursor: null,
toolFailureRetryExhaustedAuditEmitted: false,
};

View File

@@ -0,0 +1,2 @@
-- FNXC:ExecutorEscalation 2026-07-16-21:00: Persist the single-shot escalation latch so executor restarts cannot repeat a costly alternate model/node attempt.
ALTER TABLE project.tasks ADD COLUMN IF NOT EXISTS executor_escalation_attempted integer DEFAULT 0;

View File

@@ -74,6 +74,8 @@ already applied the baseline before Direct conversations can be pinned.
export const CHAT_SESSION_PINS_VERSION = "0012";
/** FNXC:ExecutorToolFailureRetry 2026-07-16-12:00: upgrades existing PostgreSQL task rows before retry-state reads. */
export const EXECUTOR_TOOL_FAILURE_RETRY_VERSION = "0013";
/** FNXC:ExecutorEscalation 2026-07-16-21:00: Existing clusters need the durable single-shot latch before executor reads it during post-FN-7996 escalation. */
export const EXECUTOR_ESCALATION_ATTEMPT_VERSION = "0014";
/** Bookkeeping table for the fresh Drizzle migration history. */
export const MIGRATION_BOOKKEEPING_TABLE = "fusion_schema_migrations";
@@ -145,6 +147,11 @@ const EXECUTOR_TOOL_FAILURE_RETRY_MIGRATION_PATH = join(
"migrations",
"0013_executor_tool_failure_retry.sql",
);
const EXECUTOR_ESCALATION_ATTEMPT_MIGRATION_PATH = join(
__dirname,
"migrations",
"0014_executor_escalation_attempt.sql",
);
/**
* Ensure the migration bookkeeping table exists. Lives in the public schema so
@@ -227,6 +234,7 @@ export async function applySchemaBaseline(
const ownerProjectIdSplitAlreadyApplied = applied.includes(OWNER_PROJECT_ID_SPLIT_VERSION);
const chatSessionPinsAlreadyApplied = applied.includes(CHAT_SESSION_PINS_VERSION);
const executorToolFailureRetryAlreadyApplied = applied.includes(EXECUTOR_TOOL_FAILURE_RETRY_VERSION);
const executorEscalationAttemptAlreadyApplied = applied.includes(EXECUTOR_ESCALATION_ATTEMPT_VERSION);
let schemaChanged = false;
if (!baselineAlreadyApplied) {
@@ -502,6 +510,15 @@ export async function applySchemaBaseline(
schemaChanged = true;
}
if (!executorEscalationAttemptAlreadyApplied) {
const executorEscalationAttemptSql = await readFile(EXECUTOR_ESCALATION_ATTEMPT_MIGRATION_PATH, "utf8");
await tx.execute(sql.raw(executorEscalationAttemptSql));
await tx.execute(
sql`INSERT INTO public.${sql.identifier(MIGRATION_BOOKKEEPING_TABLE)} (version) VALUES (${EXECUTOR_ESCALATION_ATTEMPT_VERSION}) ON CONFLICT (version) DO NOTHING`,
);
schemaChanged = true;
}
return { applied: schemaChanged, pluginHooksRun: pluginHooks.length };
});
}

View File

@@ -101,6 +101,7 @@ export const tasks = projectSchema.table("tasks", {
resumeLimboCount: integer("resume_limbo_count").default(0),
graphResumeRetryCount: integer("graph_resume_retry_count").default(0),
consecutiveToolFailureRetryCount: integer("consecutive_tool_failure_retry_count").default(0),
executorEscalationAttempted: integer("executor_escalation_attempted").default(0),
toolFailureDetectorLogCursor: integer("tool_failure_detector_log_cursor"),
toolFailureRetryExhaustedAuditEmitted: integer("tool_failure_retry_exhausted_audit_emitted").default(0),
resumeLimboTipSha: text("resume_limbo_tip_sha"),

View File

@@ -496,6 +496,10 @@ export const DEFAULT_PROJECT_SETTINGS = {
executorToolFailureRetryCount: 2,
executorToolFailureRetryBackoffMs: 2000,
executorToolFailureThreshold: 3,
executorModelEscalationEnabled: false,
executorEscalationProvider: undefined,
executorEscalationModelId: undefined,
executorEscalationNodeId: undefined,
/**
* FNXC:Merge 2026-06-26-00:00:
* New and unconfigured projects default AI merge to sync a dirty checked-out integration branch, restoring the legacy stash → fast-forward → restore landing behavior. Explicit persisted merger.allowDirtyLocalCheckoutSync values still win, and no existing-project migration stamps this default into storage.

View File

@@ -1166,7 +1166,7 @@ export class TaskStore extends EventEmitter<TaskStoreEvents> {
}
async updateTask(
id: string,
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("./types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("./types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("./types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("./types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("./types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("./types.js").TaskReview | null; reviewState?: import("./types.js").TaskReviewState | null; workflowStepResults?: import("./types.js").WorkflowStepResult[] | null; mergeDetails?: import("./types.js").MergeDetails | null; sourceIssue?: import("./types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("./types.js").TaskGithubTracking | null; tokenUsage?: import("./types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("./types.js").WorkflowTransitionNotificationMarker | undefined; plannerOversightLevel?: string | null; sessionAdvisorEnabled?: boolean | null; approvedPlanFingerprint?: string | null }, runContext?: RunMutationContext,
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("./types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("./types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("./types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("./types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("./types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; executorEscalationAttempted?: boolean | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("./types.js").TaskReview | null; reviewState?: import("./types.js").TaskReviewState | null; workflowStepResults?: import("./types.js").WorkflowStepResult[] | null; mergeDetails?: import("./types.js").MergeDetails | null; sourceIssue?: import("./types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("./types.js").TaskGithubTracking | null; tokenUsage?: import("./types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("./types.js").WorkflowTransitionNotificationMarker | undefined; plannerOversightLevel?: string | null; sessionAdvisorEnabled?: boolean | null; approvedPlanFingerprint?: string | null }, runContext?: RunMutationContext,
): Promise<Task> {
return updateTaskImpl(this, id, updates, runContext);
}

View File

@@ -48,6 +48,7 @@ export interface TaskRow {
resumeLimboCount: number | null;
graphResumeRetryCount: number | null;
consecutiveToolFailureRetryCount: number | null;
executorEscalationAttempted: number | null;
toolFailureDetectorLogCursor: number | null;
toolFailureRetryExhaustedAuditEmitted: number | null;
resumeLimboTipSha: string | null;
@@ -232,6 +233,7 @@ export const TASK_COLUMN_DESCRIPTORS: TaskColumnDescriptor[] = [
defineTaskColumn("resumeLimboCount", (task) => task.resumeLimboCount ?? 0),
defineTaskColumn("graphResumeRetryCount", (task) => task.graphResumeRetryCount === undefined ? 0 : task.graphResumeRetryCount),
defineTaskColumn("consecutiveToolFailureRetryCount", (task) => task.consecutiveToolFailureRetryCount ?? 0),
defineTaskColumn("executorEscalationAttempted", (task) => task.executorEscalationAttempted ? 1 : 0),
defineTaskColumn("toolFailureDetectorLogCursor", (task) => task.toolFailureDetectorLogCursor ?? null),
// FNXC:ExecutorToolFailureRetry 2026-07-16-13:30: PostgreSQL stores this CAS marker as integer (0/1), matching its schema and legacy SQLite flag representation.
defineTaskColumn("toolFailureRetryExhaustedAuditEmitted", (task) => task.toolFailureRetryExhaustedAuditEmitted ? 1 : 0),

View File

@@ -44,7 +44,7 @@ export function getTaskSelectClauseWithActivityLogLimitImpl(store: TaskStore, li
"modelPresetId", "modelProvider", "modelId",
"validatorModelProvider", "validatorModelId",
"planningModelProvider", "planningModelId",
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "executorEscalationAttempted", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
"error", "summary", "thinkingLevel", "validatorThinkingLevel", "planningThinkingLevel", "executionMode",
"tokenUsageInputTokens", "tokenUsageOutputTokens", "tokenUsageCachedTokens", "tokenUsageCacheWriteTokens", "tokenUsageTotalTokens", "tokenUsageFirstUsedAt", "tokenUsageLastUsedAt", "tokenUsageModelProvider", "tokenUsageModelId", "tokenUsagePerModel", "tokenBudgetSoftAlertedAt", "tokenBudgetHardAlertedAt", "tokenBudgetOverride",
"createdAt", "updatedAt", "columnMovedAt", "firstExecutionAt", "cumulativeActiveMs", "executionStartedAt", "executionCompletedAt",

View File

@@ -34,7 +34,7 @@ export function getTaskSelectClauseImpl2(store: TaskStore, slim: boolean, tableA
"modelPresetId", "modelProvider", "modelId",
"validatorModelProvider", "validatorModelId",
"planningModelProvider", "planningModelId",
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "executorEscalationAttempted", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
"error", "summary", "thinkingLevel", "validatorThinkingLevel", "planningThinkingLevel", "executionMode",
"tokenUsageInputTokens", "tokenUsageOutputTokens", "tokenUsageCachedTokens", "tokenUsageCacheWriteTokens", "tokenUsageTotalTokens", "tokenUsageFirstUsedAt", "tokenUsageLastUsedAt", "tokenUsageModelProvider", "tokenUsageModelId", "tokenUsagePerModel", "tokenBudgetSoftAlertedAt", "tokenBudgetHardAlertedAt", "tokenBudgetOverride",
"createdAt", "updatedAt", "columnMovedAt", "firstExecutionAt", "cumulativeActiveMs", "executionStartedAt", "executionCompletedAt",

View File

@@ -624,7 +624,7 @@ export async function resetPromptCheckboxesImpl(store: TaskStore, dir: string):
export async function updateTaskImpl(store: TaskStore,
id: string,
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("../types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("../types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("../types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("../types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("../types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("../types.js").TaskReview | null; reviewState?: import("../types.js").TaskReviewState | null; workflowStepResults?: import("../types.js").WorkflowStepResult[] | null; mergeDetails?: import("../types.js").MergeDetails | null; sourceIssue?: import("../types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("../types.js").TaskGithubTracking | null; tokenUsage?: import("../types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("../types.js").WorkflowTransitionNotificationMarker | undefined; sessionAdvisorEnabled?: boolean | null }, runContext?: RunMutationContext,
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("../types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("../types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("../types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("../types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("../types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; executorEscalationAttempted?: boolean | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("../types.js").TaskReview | null; reviewState?: import("../types.js").TaskReviewState | null; workflowStepResults?: import("../types.js").WorkflowStepResult[] | null; mergeDetails?: import("../types.js").MergeDetails | null; sourceIssue?: import("../types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("../types.js").TaskGithubTracking | null; tokenUsage?: import("../types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("../types.js").WorkflowTransitionNotificationMarker | undefined; sessionAdvisorEnabled?: boolean | null }, runContext?: RunMutationContext,
): Promise<Task> {
/*
FNXC:StateMachine 2026-07-07-12:00:

View File

@@ -102,6 +102,7 @@ export function rowToTask(row: TaskRow): Task {
resumeLimboCount: row.resumeLimboCount ?? undefined,
graphResumeRetryCount: row.graphResumeRetryCount ?? undefined,
consecutiveToolFailureRetryCount: row.consecutiveToolFailureRetryCount ?? undefined,
executorEscalationAttempted: row.executorEscalationAttempted ? true : undefined,
toolFailureDetectorLogCursor: row.toolFailureDetectorLogCursor ?? undefined,
toolFailureRetryExhaustedAuditEmitted: row.toolFailureRetryExhaustedAuditEmitted ? true : undefined,
resumeLimboTipSha: row.resumeLimboTipSha || undefined,

View File

@@ -367,6 +367,8 @@ export async function updateTaskUnlockedImpl(store: TaskStore, id: string, updat
}
if (updates.consecutiveToolFailureRetryCount === null) task.consecutiveToolFailureRetryCount = null;
else if (updates.consecutiveToolFailureRetryCount !== undefined) task.consecutiveToolFailureRetryCount = updates.consecutiveToolFailureRetryCount;
if (updates.executorEscalationAttempted === null) task.executorEscalationAttempted = null;
else if (updates.executorEscalationAttempted !== undefined) task.executorEscalationAttempted = updates.executorEscalationAttempted;
if (updates.toolFailureDetectorLogCursor === null) task.toolFailureDetectorLogCursor = null;
else if (updates.toolFailureDetectorLogCursor !== undefined) task.toolFailureDetectorLogCursor = updates.toolFailureDetectorLogCursor;
if (updates.toolFailureRetryExhaustedAuditEmitted === null) task.toolFailureRetryExhaustedAuditEmitted = null;

View File

@@ -1618,6 +1618,11 @@ export interface Task {
* FN-7996 persists the bounded same-model retry budget for consecutive terminal tool errors. The executor atomically claims it per run cursor so concurrent failures cannot exceed the configured cap.
*/
consecutiveToolFailureRetryCount?: number | null;
/**
* FNXC:ExecutorEscalation 2026-07-16-21:00:
* Records consumption of the one opt-in alternate model/node attempt after FN-7996 exhausts same-model retries. Reset with the retry window so unrelated failure surfaces receive their own bounded escalation.
*/
executorEscalationAttempted?: boolean | null;
/** Agent-log boundary captured at executor-run start; only later terminal outcomes qualify. */
toolFailureDetectorLogCursor?: number | null;
/** Durable compare-and-set marker which permits one exhaustion audit per retry window. */
@@ -3052,6 +3057,14 @@ export interface ProjectSettings {
executorToolFailureRetryCount?: number;
executorToolFailureRetryBackoffMs?: number;
executorToolFailureThreshold?: number;
/**
* FNXC:ExecutorEscalation 2026-07-16-21:00:
* Opt-in single-shot escalation runs only after FN-7996 exhausts same-model retries. The task model target enters the model-selection hierarchy as an override and the node target enters routing as a task override; default off prevents surprise cost or behavior changes.
*/
executorModelEscalationEnabled?: boolean;
executorEscalationProvider?: string;
executorEscalationModelId?: string;
executorEscalationNodeId?: string;
/**
* FNXC:VerificationConcurrency 2026-07-15-03:35:
* Max concurrent verification subprocesses (fn_run_verification / merge testCommand builds) across all tasks in this process. Caps stacked monorepo typecheck/build pegging CPU when many tasks are in-progress. Default 1. Raise only on high-core hosts.

View File

@@ -1140,6 +1140,10 @@ export function SettingsModal({
executorToolFailureRetryCount: 2,
executorToolFailureRetryBackoffMs: 2000,
executorToolFailureThreshold: 3,
executorModelEscalationEnabled: false,
executorEscalationProvider: "",
executorEscalationModelId: "",
executorEscalationNodeId: "",
mergeIntegrationWorktree: "reuse-task-worktree",
mergeAdvanceAutoSync: "stash-and-ff",
merger: { mode: "ai", maxReviewPasses: 3, allowDirtyLocalCheckoutSync: true },
@@ -1724,6 +1728,10 @@ export function SettingsModal({
executorToolFailureRetryCount: resolveNonNegativeExecutorToolFailureSetting(s.executorToolFailureRetryCount, 2),
executorToolFailureRetryBackoffMs: resolveNonNegativeExecutorToolFailureSetting(s.executorToolFailureRetryBackoffMs, 2000),
executorToolFailureThreshold: Math.max(1, Math.floor(Number(s.executorToolFailureThreshold ?? 3) || 3)),
executorModelEscalationEnabled: s.executorModelEscalationEnabled === true,
executorEscalationProvider: s.executorEscalationProvider ?? "",
executorEscalationModelId: s.executorEscalationModelId ?? "",
executorEscalationNodeId: s.executorEscalationNodeId ?? "",
worktreeCopyFiles: Array.isArray(s.worktreeCopyFiles) ? s.worktreeCopyFiles : [],
};
setForm(normalizedSettings);
@@ -3340,6 +3348,10 @@ export function SettingsModal({
executorToolFailureRetryCount: resolveNonNegativeExecutorToolFailureSetting(form.executorToolFailureRetryCount, 2),
executorToolFailureRetryBackoffMs: resolveNonNegativeExecutorToolFailureSetting(form.executorToolFailureRetryBackoffMs, 2000),
executorToolFailureThreshold: Math.max(1, Math.floor(Number(form.executorToolFailureThreshold ?? 3) || 3)),
executorModelEscalationEnabled: form.executorModelEscalationEnabled === true,
executorEscalationProvider: form.executorEscalationProvider?.trim() || undefined,
executorEscalationModelId: form.executorEscalationModelId?.trim() || undefined,
executorEscalationNodeId: form.executorEscalationNodeId?.trim() || undefined,
taskPrefix: form.taskPrefix?.trim() || undefined,
githubTrackingDefaultRepo: form.githubTrackingDefaultRepo?.trim() || undefined,
/*

View File

@@ -118,6 +118,10 @@ const PROJECT_SECTION_KEYS: Record<string, readonly string[]> = {
"executorToolFailureRetryCount",
"executorToolFailureRetryBackoffMs",
"executorToolFailureThreshold",
"executorModelEscalationEnabled",
"executorEscalationProvider",
"executorEscalationModelId",
"executorEscalationNodeId",
"groupOverlappingFiles",
"heartbeatScopeDiscipline",
"ignoreHiddenOverlapPaths",

View File

@@ -66,6 +66,42 @@ export const schedulingSearchEntries: SettingsSearchEntry[] = [
helpFallback: "Terminal tool errors required before retrying. Default: 3.",
keywords: ["executor", "tool error", "threshold", "auto retry"],
},
{
sectionId: "scheduling",
key: "executorModelEscalationEnabled",
labelKey: "settings.scheduling.executorModelEscalationEnabled",
labelFallback: "Escalate after tool-failure retries",
helpKey: "settings.scheduling.executorModelEscalationEnabledHelp",
helpFallback: "After same-model retries are exhausted, try one configured alternate model or node. Disabled by default.",
keywords: ["executor", "model", "node", "escalation", "tool error"],
},
{
sectionId: "scheduling",
key: "executorEscalationProvider",
labelKey: "settings.scheduling.executorEscalationProvider",
labelFallback: "Escalation provider",
helpKey: "settings.scheduling.executorEscalationProviderHelp",
helpFallback: "Provider for the alternate model. Requires an alternate model ID.",
keywords: ["executor", "model", "provider", "escalation"],
},
{
sectionId: "scheduling",
key: "executorEscalationModelId",
labelKey: "settings.scheduling.executorEscalationModelId",
labelFallback: "Escalation model ID",
helpKey: "settings.scheduling.executorEscalationModelIdHelp",
helpFallback: "Alternate model ID. Requires an escalation provider.",
keywords: ["executor", "model", "escalation"],
},
{
sectionId: "scheduling",
key: "executorEscalationNodeId",
labelKey: "settings.scheduling.executorEscalationNodeId",
labelFallback: "Escalation node ID",
helpKey: "settings.scheduling.executorEscalationNodeIdHelp",
helpFallback: "Optional configured node; a node target re-enters scheduler routing.",
keywords: ["executor", "node", "routing", "escalation"],
},
{
sectionId: "scheduling",
key: "pollIntervalMs",

View File

@@ -3,6 +3,7 @@ import { MovedSettingsStub } from "./MovedSettingsStub";
import { SettingsToggleRow } from "../SettingsToggleRow";
import { SettingsSelectRow } from "../SettingsSelectRow";
import { SettingsNumberRow } from "../SettingsNumberRow";
import { SettingsTextRow } from "../SettingsTextRow";
import { SettingsHelpTip } from "../SettingsHelpTip";
import type { SettingsFormState, SetSettingsForm } from "./context";
const MS_PER_DAY = 24 * 60 * 60 * 1000;
@@ -37,6 +38,11 @@ export function SchedulingSection({ form, setForm, concurrencyLoading = false, o
<SettingsNumberRow descriptor={{ key: "executorToolFailureRetryCount", label: t("settings.scheduling.executorToolFailureRetryCount", "Executor tool-failure retries"), help: t("settings.scheduling.executorToolFailureRetryCountHelp", "Same-model retries after consecutive tool-call failures. Set 0 to disable. Default: 2."), scope: "project", min: 0, step: 1 }} value={form.executorToolFailureRetryCount ?? 2} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureRetryCount: Math.max(0, Math.floor(v ?? 2)) } as SettingsFormState))} />
<SettingsNumberRow descriptor={{ key: "executorToolFailureRetryBackoffMs", label: t("settings.scheduling.executorToolFailureRetryBackoffMs", "Tool-failure retry backoff (ms)"), help: t("settings.scheduling.executorToolFailureRetryBackoffMsHelp", "Unref'd wait before retrying. Default: 2000."), scope: "project", min: 0, step: 1 }} value={form.executorToolFailureRetryBackoffMs ?? 2000} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureRetryBackoffMs: Math.max(0, Math.floor(v ?? 2000)) } as SettingsFormState))} />
<SettingsNumberRow descriptor={{ key: "executorToolFailureThreshold", label: t("settings.scheduling.executorToolFailureThreshold", "Consecutive tool failures"), help: t("settings.scheduling.executorToolFailureThresholdHelp", "Terminal tool errors required before retrying. Default: 3."), scope: "project", min: 1, step: 1 }} value={form.executorToolFailureThreshold ?? 3} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureThreshold: Math.max(1, Math.floor(v ?? 3)) } as SettingsFormState))} />
{/* FNXC:ExecutorEscalation 2026-07-16-21:00: Keep alternate model/node escalation opt-in and adjacent to its FN-7996 retry policy; a complete model pair or node id is required before the executor consumes its one extra attempt. */}
<SettingsToggleRow descriptor={{ key: "executorModelEscalationEnabled", label: t("settings.scheduling.executorModelEscalationEnabled", "Escalate after tool-failure retries"), help: t("settings.scheduling.executorModelEscalationEnabledHelp", "After same-model retries are exhausted, try one configured alternate model or node. Disabled by default."), scope: "project" }} value={form.executorModelEscalationEnabled === true} onChange={(value) => setForm((f) => ({ ...f, executorModelEscalationEnabled: value === true } as SettingsFormState))} />
<SettingsTextRow descriptor={{ key: "executorEscalationProvider", label: t("settings.scheduling.executorEscalationProvider", "Escalation provider"), help: t("settings.scheduling.executorEscalationProviderHelp", "Provider for the alternate model. Requires an alternate model ID."), scope: "project" }} value={form.executorEscalationProvider ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationProvider: value ?? "" } as SettingsFormState))} />
<SettingsTextRow descriptor={{ key: "executorEscalationModelId", label: t("settings.scheduling.executorEscalationModelId", "Escalation model ID"), help: t("settings.scheduling.executorEscalationModelIdHelp", "Alternate model ID. Requires an escalation provider."), scope: "project" }} value={form.executorEscalationModelId ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationModelId: value ?? "" } as SettingsFormState))} />
<SettingsTextRow descriptor={{ key: "executorEscalationNodeId", label: t("settings.scheduling.executorEscalationNodeId", "Escalation node ID"), help: t("settings.scheduling.executorEscalationNodeIdHelp", "Optional configured node; a node target re-enters scheduler routing."), scope: "project" }} value={form.executorEscalationNodeId ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationNodeId: value ?? "" } as SettingsFormState))} />
<SettingsNumberRow
descriptor={{
key: "maxConcurrent",

View File

@@ -201,6 +201,10 @@ const SETTING_DESCRIPTION_KEYS: Record<string, string> = {
executorToolFailureRetryCount: "scheduling.executorToolFailureRetryCountHelp",
executorToolFailureRetryBackoffMs: "scheduling.executorToolFailureRetryBackoffMsHelp",
executorToolFailureThreshold: "scheduling.executorToolFailureThresholdHelp",
executorModelEscalationEnabled: "scheduling.executorModelEscalationEnabledHelp",
executorEscalationProvider: "scheduling.executorEscalationProviderHelp",
executorEscalationModelId: "scheduling.executorEscalationModelIdHelp",
executorEscalationNodeId: "scheduling.executorEscalationNodeIdHelp",
taskStuckTimeoutMs: "scheduling.timeoutInMinutesForDetectingStuckTasksWhen",
staleHighFanoutBlockerAgeThresholdMs: "scheduling.escalateHighFanOutBlockersOnlyAfterThey",
preserveProgressOnStuckRequeue: "scheduling.whenTheStuckDetectorKillsAndReQueues",

View File

@@ -41,9 +41,9 @@ function graphFailure() {
};
}
function makeHarness(options: { retries: number; entries: Array<{ type: string }> }) {
function makeHarness(options: { retries: number; entries: Array<{ type: string }>; settings?: Record<string, unknown>; task?: Partial<TaskDetail> }) {
const store = createMockStore();
const task = makeTask();
const task = makeTask(options.task);
store.getTask.mockResolvedValue(task);
store.getSettings.mockResolvedValue({
maxConcurrent: 2,
@@ -53,10 +53,12 @@ function makeHarness(options: { retries: number; entries: Array<{ type: string }
executorToolFailureRetryCount: options.retries,
executorToolFailureRetryBackoffMs: 0,
executorToolFailureThreshold: 3,
...options.settings,
});
store.getAgentLogCount = vi.fn().mockResolvedValue(options.entries.length);
store.getAgentLogs = vi.fn().mockResolvedValue(options.entries);
store.claimNextToolFailureRetry = vi.fn().mockResolvedValue({ outcome: "claimed", attempt: 1 });
store.updateTask.mockImplementation(async (_id: string, patch: Partial<TaskDetail>) => Object.assign(task, patch));
store.updateTaskAtomic = vi.fn(async (_id: string, updater: (current: TaskDetail) => Partial<TaskDetail> | null) => {
const updates = updater(task);
if (updates) Object.assign(task, updates);
@@ -123,6 +125,91 @@ describe("executor consecutive tool-failure retry (FN-7996)", () => {
});
});
it("escalates once to a configured model after same-model retries exhaust", async () => {
const { executor, store, task } = makeHarness({
retries: 2,
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
});
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
const execute = vi.spyOn(executor as any, "execute").mockResolvedValue(undefined);
await (executor as any).handleGraphFailure(task, graphFailure());
await vi.advanceTimersByTimeAsync(0);
expect(task).toMatchObject({ modelProvider: "anthropic", modelId: "claude-sonnet", executorEscalationAttempted: true, status: null, error: null });
expect(execute).toHaveBeenCalledWith(task);
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-retry", metadata: expect.objectContaining({ taskId: task.id, hasModelTarget: true, hasNodeTarget: false }) }));
expect(store.updateTaskAtomic).toHaveBeenCalledWith(task.id, expect.any(Function), undefined);
});
it("requeues a node escalation for scheduler effective-node resolution", async () => {
const { executor, store, task } = makeHarness({
retries: 2,
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
settings: { executorModelEscalationEnabled: true, executorEscalationNodeId: "cursor-node" },
});
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
const execute = vi.spyOn(executor as any, "execute").mockResolvedValue(undefined);
await (executor as any).handleGraphFailure(task, graphFailure());
expect(task).toMatchObject({ nodeId: "cursor-node", column: "todo", executorEscalationAttempted: true, status: null, error: null });
expect(execute).not.toHaveBeenCalled();
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-retry", metadata: expect.objectContaining({ hasNodeTarget: true }) }));
});
it("parks the single escalated attempt and records escalation exhaustion", async () => {
const { executor, store, task } = makeHarness({
retries: 2,
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
task: { executorEscalationAttempted: true },
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
});
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
await (executor as any).handleGraphFailure(task, graphFailure());
expect(task.status).toBe("failed");
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-exhausted" }));
});
it("does not let a concurrent exhausted handler park the escalation it lost", async () => {
const { executor, store, task } = makeHarness({
retries: 2,
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
});
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
task.executorEscalationAttempted = true;
// The atomic escalation claim invalidates its exhausted cursor before scheduling.
task.toolFailureDetectorLogCursor = null;
await (executor as any).handleGraphFailure(task, graphFailure());
expect(task).toMatchObject({ status: null, error: null, executorEscalationAttempted: true });
expect(store.updateTask).not.toHaveBeenCalledWith(task.id, expect.objectContaining({ status: "failed" }), expect.anything());
expect(store.recordRunAuditEvent).not.toHaveBeenCalledWith(expect.objectContaining({
mutationType: "task:execution-escalation-exhausted",
}));
});
it("audits a terminal escalated failure even after escalation is disabled", async () => {
const { executor, store, task } = makeHarness({
retries: 0,
entries: [],
task: { executorEscalationAttempted: true, modelProvider: "anthropic", modelId: "claude-sonnet" },
});
await (executor as any).handleGraphFailure(task, graphFailure());
expect(task.status).toBe("failed");
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({
mutationType: "task:execution-escalation-exhausted",
metadata: expect.objectContaining({ hadModelTarget: true, hadNodeTarget: false }),
}));
});
it("does not let an exhausted stale handler park a newer cursor-owned run", async () => {
const { executor, store, task } = makeHarness({
retries: 2,

View File

@@ -13,7 +13,7 @@ import { existsSync, lstatSync, realpathSync } from "node:fs";
import { readFile, rm, writeFile } from "node:fs/promises";
import type { TaskStore, Task, TaskDetail, TaskTokenUsage, StepStatus, Settings, WorkflowStep, MissionStore, AsyncMissionStore, Slice, AgentState, AgentCapability, RunMutationContext, AgentHeartbeatConfig, Agent, AgentMemoryInclusionMode, ProjectSettings, MergeResult, WorkflowIrNode, WorkflowIrNodeKind, WorkflowStepResult as CoreWorkflowStepResult, ThinkingLevel } from "@fusion/core";
import { getUnmetSchedulingDependencies } from "./scheduler.js";
import { RetryStormError, serializeRetryStormError, isExperimentalFeatureEnabled, resolveWorkflowIrForTask, resolveColumnAgentBinding, resolveEffectiveAgent, instanceNodeId, getWorkflowExtensionRegistry, getBuiltinWorkflow, parseNoOpCompletionMarker, allowsAutoMergeProcessing, resolveEffectiveAutoMerge, isLiveSharedBranchGroupMemberIntegration, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveOptionalStepRevisionBudget, resolveOptionalReviewRevisionBudget, COMPLETION_SUMMARY_NODE_ID, upsertWorkflowStepResult, AWAITING_APPROVAL_PAUSE_REASON, THINKING_LEVELS, AgentStore, resolveExecutorFallbackModel } from "@fusion/core";
import { RetryStormError, serializeRetryStormError, isExperimentalFeatureEnabled, resolveWorkflowIrForTask, resolveColumnAgentBinding, resolveEffectiveAgent, instanceNodeId, getWorkflowExtensionRegistry, getBuiltinWorkflow, parseNoOpCompletionMarker, allowsAutoMergeProcessing, resolveEffectiveAutoMerge, isLiveSharedBranchGroupMemberIntegration, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveExecutorEscalationTarget, resolveOptionalStepRevisionBudget, resolveOptionalReviewRevisionBudget, COMPLETION_SUMMARY_NODE_ID, upsertWorkflowStepResult, AWAITING_APPROVAL_PAUSE_REASON, THINKING_LEVELS, AgentStore, resolveExecutorFallbackModel } from "@fusion/core";
import { finalizeProvenAutoMergeTask } from "./auto-merge-finalization.js";
import { mergeEffectiveSettings } from "./effective-settings.js";
import { moveTaskToReplanColumn, resolveReplanTargetColumn } from "./replan-target.js";
@@ -5419,7 +5419,7 @@ export class TaskExecutor {
await this.finalizeMergeConfirmedWorkflowGraphTask(task.id, "graph-completed");
}
if ((live.graphResumeRetryCount ?? 0) !== 0 || (live.consecutiveToolFailureRetryCount ?? 0) !== 0) {
await this.store.updateTask(task.id, { graphResumeRetryCount: 0, consecutiveToolFailureRetryCount: 0, toolFailureDetectorLogCursor: null, toolFailureRetryExhaustedAuditEmitted: false }, this.getRunContextFor(task.id));
await this.store.updateTask(task.id, { graphResumeRetryCount: 0, consecutiveToolFailureRetryCount: 0, executorEscalationAttempted: false, toolFailureDetectorLogCursor: null, toolFailureRetryExhaustedAuditEmitted: false }, this.getRunContextFor(task.id));
}
}
return true;
@@ -9690,12 +9690,13 @@ export class TaskExecutor {
return;
}
const message = `Workflow graph terminated with failure at node '${failedNode ?? "unknown"}'`;
const maxToolFailureRetries = resolveMaxConsecutiveToolFailureRetries(await this.store.getSettings());
const settings = await this.store.getSettings();
const maxToolFailureRetries = resolveMaxConsecutiveToolFailureRetries(settings);
const isExecuteFailure = failedNode === "execute" || failedNode?.endsWith(":step-execute") === true || failedNode === "step-execute";
if (maxToolFailureRetries > 0 && isExecuteFailure && !live.paused && !live.userPaused && !live.deletedAt && live.column === "in-progress") {
// Prefer the execution-local boundary; recovery paths refetch durable state rather than use the stale failure snapshot.
const cursor = this.graphToolFailureRunCursors.get(task.id) ?? (await this.store.getTask(task.id))?.toolFailureDetectorLogCursor;
const threshold = resolveConsecutiveToolFailureThreshold(await this.store.getSettings());
const threshold = resolveConsecutiveToolFailureThreshold(settings);
if (await this.hasTrailingConsecutiveToolFailures(task.id, cursor, threshold)) {
const claim = await this.store.claimNextToolFailureRetry(task.id, cursor!, maxToolFailureRetries);
if (claim.outcome === "claimed") {
@@ -9703,11 +9704,56 @@ export class TaskExecutor {
await this.store.logEntry(task.id, `Consecutive tool-call failures — auto-retrying same model (${claim.attempt}/${maxToolFailureRetries}) instead of parking`, undefined, this.getRunContextFor(task.id));
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("tool-failure-retry", task.id), domain: "database", mutationType: "task:execution-tool-failure-retry", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", attempt: claim.attempt, maxAttempts: maxToolFailureRetries, consecutiveToolFailures: threshold, mode: "same-model" } });
const schedule = () => { void (async () => { const resume = await this.store.getTask(task.id); if (resume && !resume.deletedAt && !resume.paused && !resume.userPaused && resume.column === "in-progress") await this.execute(resume); })().catch((error) => executorLog.error(`${task.id}: tool-failure retry failed`, error)); };
const delay = resolveConsecutiveToolFailureRetryBackoffMs(await this.store.getSettings());
const delay = resolveConsecutiveToolFailureRetryBackoffMs(settings);
setTimeout(schedule, delay).unref?.();
return;
}
if (claim.outcome === "already-claimed-for-run") { await this.store.getTask(task.id); return; }
/*
FNXC:ExecutorEscalation 2026-07-16-21:00:
FN-7998 inserts exactly one opt-in recovery between FN-7996 exhaustion and the unchanged terminal park. Refetch before writing so a pause, deletion, or later run cannot inherit a costly model/node override from this stale graph result.
*/
const escalationTarget = resolveExecutorEscalationTarget(settings);
const hasModelTarget = escalationTarget.provider !== undefined && escalationTarget.modelId !== undefined;
const hasNodeTarget = escalationTarget.nodeId !== undefined;
let claimedEscalation = false;
let priorEscalationRetryCount = 0;
/*
FNXC:ExecutorEscalation 2026-07-16-22:30:
The one-shot latch is claimed under the TaskStore lock. Concurrent exhausted
graph handlers for the same detector cursor must not both schedule an alternate
run; a loser leaves the winner's in-progress row untouched.
*/
await this.store.updateTaskAtomic(task.id, (current) => {
const ownsFailureRun = current.toolFailureDetectorLogCursor === cursor
&& current.column === "in-progress"
&& !current.paused
&& !current.userPaused
&& !current.deletedAt;
if (!ownsFailureRun || current.executorEscalationAttempted === true || !escalationTarget.enabled) return null;
claimedEscalation = true;
priorEscalationRetryCount = current.consecutiveToolFailureRetryCount ?? 0;
return {
...(hasModelTarget ? { modelProvider: escalationTarget.provider, modelId: escalationTarget.modelId } : {}),
...(hasNodeTarget ? { nodeId: escalationTarget.nodeId, column: "todo" as const } : {}),
executorEscalationAttempted: true,
/* FNXC:ExecutorEscalation 2026-07-16-22:40: Invalidate the exhausted run cursor before releasing the claim so concurrent stale handlers cannot park or audit the alternate execution; the alternate captures its own cursor at startup. */
toolFailureDetectorLogCursor: null,
status: null,
error: null,
};
}, this.getRunContextFor(task.id));
if (claimedEscalation) {
await this.store.logEntry(task.id, "Same-model retries exhausted — escalating to alternate model/node (one attempt) instead of parking", undefined, this.getRunContextFor(task.id));
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-retry", task.id), domain: "database", mutationType: "task:execution-escalation-retry", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hasModelTarget, hasNodeTarget, priorConsecutiveToolFailureRetryCount: priorEscalationRetryCount } });
if (!hasNodeTarget) {
const scheduleEscalation = () => { void (async () => { const resumeTask = await this.store.getTask(task.id); if (resumeTask && !resumeTask.deletedAt && !resumeTask.paused && !resumeTask.userPaused && resumeTask.column === "in-progress") await this.execute(resumeTask); })().catch((error) => executorLog.error(`${task.id}: escalation retry failed`, error)); };
const handle = setTimeout(scheduleEscalation, resolveConsecutiveToolFailureRetryBackoffMs(settings));
handle.unref?.();
}
return;
}
/*
FNXC:ExecutorToolFailureRetry 2026-07-16-20:45:
Exhaustion belongs to the graph run that supplied `cursor`, not a later run
@@ -9717,6 +9763,9 @@ export class TaskExecutor {
old terminal handler from parking a newer in-progress executor run.
*/
let cursorOwnedTerminalPark = false;
let escalationAttemptFailed = false;
let escalationHadModelTarget = false;
let escalationHadNodeTarget = false;
await this.store.updateTaskAtomic(task.id, (current) => {
if (
current.toolFailureDetectorLogCursor !== cursor
@@ -9724,28 +9773,65 @@ export class TaskExecutor {
|| current.paused
|| current.userPaused
|| current.deletedAt
|| current.status !== null
) {
return null;
}
cursorOwnedTerminalPark = true;
escalationAttemptFailed = current.executorEscalationAttempted === true;
escalationHadModelTarget = current.modelProvider != null && current.modelId != null;
escalationHadNodeTarget = current.nodeId != null;
return { error: message, status: "failed" };
}, this.getRunContextFor(task.id));
if (!cursorOwnedTerminalPark) return;
if (await this.store.markToolFailureRetryExhaustedAudit(task.id)) {
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("tool-failure-retry-exhausted", task.id), domain: "database", mutationType: "task:execution-tool-failure-retry-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", attempts: maxToolFailureRetries, limit: maxToolFailureRetries, outcome: "terminal-park" } });
}
if (escalationAttemptFailed) {
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-exhausted", task.id), domain: "database", mutationType: "task:execution-escalation-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hadModelTarget: escalationHadModelTarget, hadNodeTarget: escalationHadNodeTarget } });
}
executorLog.warn(`${task.id}: ${message}`);
await this.store.logEntry(task.id, message, undefined, this.getRunContextFor(task.id));
await this.persistTokenUsage(task.id);
return;
}
}
if (live.executorEscalationAttempted === true) {
const failureCursor = task.toolFailureDetectorLogCursor;
let escalationTerminalParked = false;
let escalationHadModelTarget = false;
let escalationHadNodeTarget = false;
/*
FNXC:ExecutorEscalation 2026-07-16-22:35:
Once the durable escalation latch is set, every terminal failure of that
alternate run emits the exhaustion audit even if an operator disables the
setting mid-run. Cursor ownership prevents an old concurrent handler from
parking the newly scheduled alternate execution.
*/
await this.store.updateTaskAtomic(task.id, (current) => {
if (
current.toolFailureDetectorLogCursor !== failureCursor
|| current.column !== "in-progress"
|| current.paused
|| current.userPaused
|| current.deletedAt
|| current.status !== null
) return null;
escalationTerminalParked = true;
escalationHadModelTarget = current.modelProvider != null && current.modelId != null;
escalationHadNodeTarget = current.nodeId != null;
return { error: message, status: "failed" };
}, this.getRunContextFor(task.id));
if (!escalationTerminalParked) return;
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-exhausted", task.id), domain: "database", mutationType: "task:execution-escalation-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hadModelTarget: escalationHadModelTarget, hadNodeTarget: escalationHadNodeTarget } });
} else {
// status "failed" doubles as the self-healing exemption: review-task
// revival sweeps skip tasks carrying a non-null status, preventing the
// FN-5704-style loop of re-running the graph from scratch.
await this.store.updateTask(task.id, { error: message, status: "failed" }, this.getRunContextFor(task.id));
}
executorLog.warn(`${task.id}: ${message}`);
await this.store.logEntry(task.id, message, undefined, this.getRunContextFor(task.id));
// status "failed" doubles as the self-healing exemption: review-task
// revival sweeps skip tasks carrying a non-null status, preventing the
// FN-5704-style loop of re-running the graph from scratch.
await this.store.updateTask(task.id, { error: message, status: "failed" }, this.getRunContextFor(task.id));
await this.persistTokenUsage(task.id);
} catch (err) {
executorLog.error(

View File

@@ -6634,6 +6634,14 @@
"executorToolFailureRetryBackoffMsHelp": "Unref'd wait before retrying. Default: 2000.",
"executorToolFailureThreshold": "Consecutive tool failures",
"executorToolFailureThresholdHelp": "Terminal tool errors required before retrying. Default: 3.",
"executorModelEscalationEnabled": "Escalate after tool-failure retries",
"executorModelEscalationEnabledHelp": "After same-model retries are exhausted, try one configured alternate model or node. Default: disabled.",
"executorEscalationProvider": "Escalation provider",
"executorEscalationProviderHelp": "Provider for the alternate model. No default — unset; requires an alternate model ID.",
"executorEscalationModelId": "Escalation model ID",
"executorEscalationModelIdHelp": "Alternate model ID. No default — unset; requires an escalation provider.",
"executorEscalationNodeId": "Escalation node ID",
"executorEscalationNodeIdHelp": "Optional configured node; a node target re-enters scheduler routing. No default — unset.",
"fullAgentLog": "Full agent log",
"globalMaxConcurrent": "Global Max Concurrent",
"heartbeatScopeDiscipline": "Heartbeat Scope Discipline",