FN-7998: add executor alternate model escalation
Add opt-in executor escalation after same-model tool-failure retries are exhausted. - Persist escalation settings and one-shot task state across SQLite and PostgreSQL stores. - Retry once on a configured alternate model or scheduler node and audit escalation outcomes. - Expose escalation controls, documentation, translations, migration, and regression coverage. Files changed: .changeset/fn-7998-executor-escalation.md | 7 ++ AGENTS.md | 1 + docs/settings-reference.md | 13 ++- .../core/src/__tests__/settings-defaults.test.ts | 23 ++++- packages/core/src/in-review-stall.ts | 29 ++++++ packages/core/src/index.gate.ts | 3 +- packages/core/src/index.ts | 3 +- packages/core/src/manual-retry-reset.ts | 1 + .../0014_executor_escalation_attempt.sql | 2 + packages/core/src/postgres/schema-applier.ts | 17 ++++ packages/core/src/postgres/schema/project.ts | 1 + packages/core/src/settings-schema.ts | 4 + packages/core/src/store.ts | 2 +- packages/core/src/task-store/persistence.ts | 2 + packages/core/src/task-store/remaining-ops-2.ts | 2 +- packages/core/src/task-store/remaining-ops-3.ts | 2 +- packages/core/src/task-store/remaining-ops-6.ts | 2 +- packages/core/src/task-store/serialization.ts | 1 + packages/core/src/task-store/task-update.ts | 2 + packages/core/src/types.ts | 13 +++ .../dashboard/app/components/SettingsModal.tsx | 12 +++ .../app/components/settings/section-keys.ts | 4 + .../settings/sections/SchedulingSection.search.ts | 36 +++++++ .../settings/sections/SchedulingSection.tsx | 6 ++ .../settings-default-descriptions.test.tsx | 4 + .../__tests__/executor-tool-failure-retry.test.ts | 91 +++++++++++++++++- packages/engine/src/executor.ts | 104 +++++++++++++++++++-- packages/i18n/locales/en/app.json | 8 ++ 28 files changed, 376 insertions(+), 19 deletions(-) Fusion-Task-Id: FN-7998 Fusion-Task-Lineage: bbce767d-c61a-4667-be62-abc0cc54d8be Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
7
.changeset/fn-7998-executor-escalation.md
Normal file
7
.changeset/fn-7998-executor-escalation.md
Normal file
@@ -0,0 +1,7 @@
|
||||
---
|
||||
"@runfusion/fusion": minor
|
||||
---
|
||||
|
||||
summary: Optionally escalate an executor run to a stronger model or configured node after same-model retries are exhausted.
|
||||
category: feature
|
||||
dev: New opt-in project settings provide one model/node escalation attempt with durable task state and audit events.
|
||||
@@ -276,6 +276,7 @@ Scoped exception (FN-5819): shared-branch-group members (`branchContext.assignme
|
||||
- FN-7514: the planner overseer's per-task oversight loop (`PlannerRecoveryController.tick`) emits `overseer:oversight-withheld-human-control` when the pure `evaluateOverseerHumanControl` guard withholds ALL oversight action (no steering, retry, targeted-fix, or pending confirmation) for a task that is user-paused (`task.userPaused===true`, or `task.paused===true` with no `pausedReason`) or ineligible for auto-merge processing per `allowsAutoMergeProcessing` (`autoMerge:false`/PR-based human-review terminal contract). The guard runs BEFORE FN-7513's confirmation classification, so a withheld task never records a pending confirmation. Metadata: `{ taskId, reason: "user-paused" | "auto-merge-off-human-review", stage, oversightLevel }`; deduped per (taskId, withheld reason) so it is not re-emitted every poll while the reason is unchanged.
|
||||
- FN-7720: `TaskStore.bypassFailedPreMergeReviewStep` emits `task:bypass-review` when a privileged operator bypasses the latest failed pre-merge review step of an `in-review` task; metadata includes `workflowStepId`, `workflowStepName`, `bypassedFromStatus`, `bypassedFromVerdict`, and the mandatory `reason`. The bypass rewrites the step's `status` to `"skipped"` with `bypassedBy`/`bypassedAt`/`bypassReason`/`bypassedFromStatus` fields; it never fabricates a reviewer `verdict` and clears only the failed-pre-merge-step `getTaskMergeBlocker` reason. Reachable via `fn_task_bypass_review` (CLI/pi-extension operator tool surface only — not executor/reviewer/triage) and `POST /tasks/:id/bypass-review`.
|
||||
- FN-7996: executor emits `task:execution-tool-failure-retry` for a claimed same-model consecutive-tool-failure retry and `task:execution-tool-failure-retry-exhausted` when the matching run budget is spent. Metadata is ids/counts/outcomes-only; the exhausted event is emitted once through a project-scoped compare-and-set while terminal parking remains idempotent.
|
||||
- FN-7998: executor emits `task:execution-escalation-retry` when its opt-in, single alternate model/node attempt is persisted after FN-7996 exhaustion, and `task:execution-escalation-exhausted` when that attempt also reaches the terminal park. Metadata remains ids/counts/outcomes-only (`taskId`, graph node id, target booleans, and prior retry count); no model identifiers or prose are persisted in run-audit.
|
||||
- FN-8004: `agent:heartbeat-move-skipped-soft-delete` records a heartbeat move that races a soft-deleted task without parking the durable agent. Metadata remains ids/timestamps/source only (`agentId`, optional `taskId`/`deletedAt`, `moveAttemptedAt`, optional `source`); it never stores error prose.
|
||||
|
||||
|
||||
|
||||
@@ -1755,4 +1755,15 @@ Standardize executor/validator pairs; auto-selectable by task size (Small → Bu
|
||||
| `executorToolFailureRetryBackoffMs` | integer, `2000` | Unref'd delay before the rerun. |
|
||||
| `executorToolFailureThreshold` | integer, `3` | Consecutive terminal tool failures required to qualify. |
|
||||
|
||||
Values are project-scoped and finite values are floored; count/backoff must be at least `0`, and threshold at least `1`, otherwise their defaults apply. The executor evaluates this bounded policy before its terminal graph-failure park: it counts `tool_error` completion entries, resets only on `tool_result`, and ignores `tool` invocation markers. The detector is scoped to the current executor-run agent-log cursor. Its project-scoped atomic claim prevents concurrent retries and classifies cursor mismatch before an exhausted cap so stale handlers do not park newer work. The exhausted audit is compare-and-set deduplicated while the terminal park remains idempotent. FN-7998 may extend this same-model foundation with escalation.
|
||||
Values are project-scoped and finite values are floored; count/backoff must be at least `0`, and threshold at least `1`, otherwise their defaults apply. The executor evaluates this bounded policy before its terminal graph-failure park: it counts `tool_error` completion entries, resets only on `tool_result`, and ignores `tool` invocation markers. The detector is scoped to the current executor-run agent-log cursor. Its project-scoped atomic claim prevents concurrent retries and classifies cursor mismatch before an exhausted cap so stale handlers do not park newer work. The exhausted audit is compare-and-set deduplicated while the terminal park remains idempotent.
|
||||
|
||||
### Executor escalation after tool-failure retry exhaustion
|
||||
|
||||
| Setting | Type/default | Behavior |
|
||||
| --- | --- | --- |
|
||||
| `executorModelEscalationEnabled` | boolean, `false` | Opt in to one alternate attempt after same-model retries exhaust. |
|
||||
| `executorEscalationProvider` | string, unset | Provider for an alternate model; requires `executorEscalationModelId`. |
|
||||
| `executorEscalationModelId` | string, unset | Alternate model ID; requires `executorEscalationProvider`. |
|
||||
| `executorEscalationNodeId` | string, unset | Optional configured node target. |
|
||||
|
||||
Escalation is enabled only when the toggle is true and either a complete provider/model pair or a node ID is configured. It is single-shot: after FN-7996 exhausts same-model retries, Fusion persists the override and tries once before the existing terminal park. The alternate model enters the [model-selection hierarchy](#model-selection-hierarchy) as a task-level override; a node target enters `resolveEffectiveNode` as a task-level routing override and is requeued so scheduler routing is recalculated. This remains opt-in by default to avoid unexpected model cost or execution behavior. Column-agent overrides still govern their sessions and can supersede a task-level model target.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import { afterEach, describe, expect, it, vi } from "vitest";
|
||||
import { CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD, DEFAULT_CONSECUTIVE_TOOL_FAILURE_RETRY_BACKOFF_MS, DEFAULT_MAX_CONSECUTIVE_TOOL_FAILURE_RETRIES, DEFAULT_MAX_AUTO_MERGE_RETRIES, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries } from "../in-review-stall.js";
|
||||
import { CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD, DEFAULT_CONSECUTIVE_TOOL_FAILURE_RETRY_BACKOFF_MS, DEFAULT_MAX_CONSECUTIVE_TOOL_FAILURE_RETRIES, DEFAULT_MAX_AUTO_MERGE_RETRIES, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveExecutorEscalationTarget, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries } from "../in-review-stall.js";
|
||||
import { isExperimentalFeatureEnabled } from "../experimental-features.js";
|
||||
import { DEFAULT_GLOBAL_SETTINGS, DEFAULT_PROJECT_SETTINGS, GLOBAL_SETTINGS_KEYS, PROJECT_SETTINGS_KEYS, isGlobalOnlySettingsKey } from "../settings-schema.js";
|
||||
import { isWorkflowColumnsEnabled } from "../workflow-columns-settings.js";
|
||||
@@ -64,6 +64,27 @@ describe("settings defaults invariants", () => {
|
||||
expect(resolveMaxAutoMergeRetries({ maxAutoMergeRetries: Number.NaN })).toBe(3);
|
||||
});
|
||||
|
||||
it("defaults executor escalation off and resolves only complete opt-in targets", () => {
|
||||
expect(DEFAULT_PROJECT_SETTINGS.executorModelEscalationEnabled).toBe(false);
|
||||
expect(PROJECT_SETTINGS_KEYS).toEqual(expect.arrayContaining([
|
||||
"executorModelEscalationEnabled",
|
||||
"executorEscalationProvider",
|
||||
"executorEscalationModelId",
|
||||
"executorEscalationNodeId",
|
||||
]));
|
||||
expect(resolveExecutorEscalationTarget({
|
||||
executorModelEscalationEnabled: false,
|
||||
executorEscalationProvider: "anthropic",
|
||||
executorEscalationModelId: "claude",
|
||||
executorEscalationNodeId: "node-1",
|
||||
}).enabled).toBe(false);
|
||||
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true })).toEqual({ enabled: false });
|
||||
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic" })).toEqual({ enabled: false });
|
||||
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude" })).toEqual({ enabled: true, provider: "anthropic", modelId: "claude" });
|
||||
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationNodeId: "node-1" })).toEqual({ enabled: true, nodeId: "node-1" });
|
||||
expect(resolveExecutorEscalationTarget({ executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude", executorEscalationNodeId: "node-1" })).toEqual({ enabled: true, provider: "anthropic", modelId: "claude", nodeId: "node-1" });
|
||||
});
|
||||
|
||||
it("resolves worktrunk as disabled when both scopes are unset or empty", () => {
|
||||
expect(resolveWorktrunkSettings(undefined, undefined).enabled).toBe(false);
|
||||
expect(resolveWorktrunkSettings({}, {}).enabled).toBe(false);
|
||||
|
||||
@@ -73,6 +73,35 @@ export function resolveConsecutiveToolFailureThreshold(settings?: { executorTool
|
||||
return Number.isFinite(numeric) && Math.floor(numeric) >= 1 ? Math.floor(numeric) : CONSECUTIVE_TOOL_FAILURE_RETRY_THRESHOLD;
|
||||
}
|
||||
|
||||
export interface ExecutorEscalationTarget {
|
||||
enabled: boolean;
|
||||
provider?: string;
|
||||
modelId?: string;
|
||||
nodeId?: string;
|
||||
}
|
||||
|
||||
/**
|
||||
* FNXC:ExecutorEscalation 2026-07-16-21:00:
|
||||
* Escalation remains opt-in and only accepts a complete model pair or a node target. Keeping this resolver in core makes engine and settings surfaces agree that incomplete targets silently preserve FN-7996 terminal behavior.
|
||||
*/
|
||||
export function resolveExecutorEscalationTarget(settings?: {
|
||||
executorModelEscalationEnabled?: unknown;
|
||||
executorEscalationProvider?: unknown;
|
||||
executorEscalationModelId?: unknown;
|
||||
executorEscalationNodeId?: unknown;
|
||||
} | null): ExecutorEscalationTarget {
|
||||
const provider = typeof settings?.executorEscalationProvider === "string" ? settings.executorEscalationProvider.trim() : "";
|
||||
const modelId = typeof settings?.executorEscalationModelId === "string" ? settings.executorEscalationModelId.trim() : "";
|
||||
const nodeId = typeof settings?.executorEscalationNodeId === "string" ? settings.executorEscalationNodeId.trim() : "";
|
||||
const hasModelTarget = Boolean(provider && modelId);
|
||||
const hasNodeTarget = Boolean(nodeId);
|
||||
return {
|
||||
enabled: settings?.executorModelEscalationEnabled === true && (hasModelTarget || hasNodeTarget),
|
||||
...(hasModelTarget ? { provider, modelId } : {}),
|
||||
...(hasNodeTarget ? { nodeId } : {}),
|
||||
};
|
||||
}
|
||||
|
||||
export const IN_REVIEW_STALL_LOG_PREFIX = "In-review stall surfaced [";
|
||||
export const IN_REVIEW_STALL_DEADLOCK_LOG_PREFIX = "In-review stall auto-disposed [";
|
||||
export const IN_REVIEW_STALL_TERMINAL_LOG_PREFIX = "In-review stall terminal disposed [";
|
||||
|
||||
@@ -993,8 +993,9 @@ export {
|
||||
resolveMaxConsecutiveToolFailureRetries,
|
||||
resolveConsecutiveToolFailureRetryBackoffMs,
|
||||
resolveConsecutiveToolFailureThreshold,
|
||||
resolveExecutorEscalationTarget,
|
||||
} from "./in-review-stall.js";
|
||||
export type { InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
|
||||
export type { ExecutorEscalationTarget, InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
|
||||
export {
|
||||
getStalePausedReviewSignal,
|
||||
DEFAULT_STALE_PAUSED_REVIEW_THRESHOLD_MS,
|
||||
|
||||
@@ -1027,8 +1027,9 @@ export {
|
||||
resolveMaxConsecutiveToolFailureRetries,
|
||||
resolveConsecutiveToolFailureRetryBackoffMs,
|
||||
resolveConsecutiveToolFailureThreshold,
|
||||
resolveExecutorEscalationTarget,
|
||||
} from "./in-review-stall.js";
|
||||
export type { InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
|
||||
export type { ExecutorEscalationTarget, InReviewStallSignal, InReviewStallCode, ProviderErrorClassification } from "./in-review-stall.js";
|
||||
export {
|
||||
getStalePausedReviewSignal,
|
||||
DEFAULT_STALE_PAUSED_REVIEW_THRESHOLD_MS,
|
||||
|
||||
@@ -43,6 +43,7 @@ export function buildAutoPauseClearPatch(
|
||||
export function buildManualRetryResetPatch(options?: { resetMergeRetries?: boolean }): Partial<Task> {
|
||||
const patch: Partial<Task> = {
|
||||
nextRecoveryAt: null as unknown as Task["nextRecoveryAt"],
|
||||
executorEscalationAttempted: false,
|
||||
toolFailureDetectorLogCursor: null,
|
||||
toolFailureRetryExhaustedAuditEmitted: false,
|
||||
};
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
-- FNXC:ExecutorEscalation 2026-07-16-21:00: Persist the single-shot escalation latch so executor restarts cannot repeat a costly alternate model/node attempt.
|
||||
ALTER TABLE project.tasks ADD COLUMN IF NOT EXISTS executor_escalation_attempted integer DEFAULT 0;
|
||||
@@ -74,6 +74,8 @@ already applied the baseline before Direct conversations can be pinned.
|
||||
export const CHAT_SESSION_PINS_VERSION = "0012";
|
||||
/** FNXC:ExecutorToolFailureRetry 2026-07-16-12:00: upgrades existing PostgreSQL task rows before retry-state reads. */
|
||||
export const EXECUTOR_TOOL_FAILURE_RETRY_VERSION = "0013";
|
||||
/** FNXC:ExecutorEscalation 2026-07-16-21:00: Existing clusters need the durable single-shot latch before executor reads it during post-FN-7996 escalation. */
|
||||
export const EXECUTOR_ESCALATION_ATTEMPT_VERSION = "0014";
|
||||
|
||||
/** Bookkeeping table for the fresh Drizzle migration history. */
|
||||
export const MIGRATION_BOOKKEEPING_TABLE = "fusion_schema_migrations";
|
||||
@@ -145,6 +147,11 @@ const EXECUTOR_TOOL_FAILURE_RETRY_MIGRATION_PATH = join(
|
||||
"migrations",
|
||||
"0013_executor_tool_failure_retry.sql",
|
||||
);
|
||||
const EXECUTOR_ESCALATION_ATTEMPT_MIGRATION_PATH = join(
|
||||
__dirname,
|
||||
"migrations",
|
||||
"0014_executor_escalation_attempt.sql",
|
||||
);
|
||||
|
||||
/**
|
||||
* Ensure the migration bookkeeping table exists. Lives in the public schema so
|
||||
@@ -227,6 +234,7 @@ export async function applySchemaBaseline(
|
||||
const ownerProjectIdSplitAlreadyApplied = applied.includes(OWNER_PROJECT_ID_SPLIT_VERSION);
|
||||
const chatSessionPinsAlreadyApplied = applied.includes(CHAT_SESSION_PINS_VERSION);
|
||||
const executorToolFailureRetryAlreadyApplied = applied.includes(EXECUTOR_TOOL_FAILURE_RETRY_VERSION);
|
||||
const executorEscalationAttemptAlreadyApplied = applied.includes(EXECUTOR_ESCALATION_ATTEMPT_VERSION);
|
||||
let schemaChanged = false;
|
||||
|
||||
if (!baselineAlreadyApplied) {
|
||||
@@ -502,6 +510,15 @@ export async function applySchemaBaseline(
|
||||
schemaChanged = true;
|
||||
}
|
||||
|
||||
if (!executorEscalationAttemptAlreadyApplied) {
|
||||
const executorEscalationAttemptSql = await readFile(EXECUTOR_ESCALATION_ATTEMPT_MIGRATION_PATH, "utf8");
|
||||
await tx.execute(sql.raw(executorEscalationAttemptSql));
|
||||
await tx.execute(
|
||||
sql`INSERT INTO public.${sql.identifier(MIGRATION_BOOKKEEPING_TABLE)} (version) VALUES (${EXECUTOR_ESCALATION_ATTEMPT_VERSION}) ON CONFLICT (version) DO NOTHING`,
|
||||
);
|
||||
schemaChanged = true;
|
||||
}
|
||||
|
||||
return { applied: schemaChanged, pluginHooksRun: pluginHooks.length };
|
||||
});
|
||||
}
|
||||
|
||||
@@ -101,6 +101,7 @@ export const tasks = projectSchema.table("tasks", {
|
||||
resumeLimboCount: integer("resume_limbo_count").default(0),
|
||||
graphResumeRetryCount: integer("graph_resume_retry_count").default(0),
|
||||
consecutiveToolFailureRetryCount: integer("consecutive_tool_failure_retry_count").default(0),
|
||||
executorEscalationAttempted: integer("executor_escalation_attempted").default(0),
|
||||
toolFailureDetectorLogCursor: integer("tool_failure_detector_log_cursor"),
|
||||
toolFailureRetryExhaustedAuditEmitted: integer("tool_failure_retry_exhausted_audit_emitted").default(0),
|
||||
resumeLimboTipSha: text("resume_limbo_tip_sha"),
|
||||
|
||||
@@ -496,6 +496,10 @@ export const DEFAULT_PROJECT_SETTINGS = {
|
||||
executorToolFailureRetryCount: 2,
|
||||
executorToolFailureRetryBackoffMs: 2000,
|
||||
executorToolFailureThreshold: 3,
|
||||
executorModelEscalationEnabled: false,
|
||||
executorEscalationProvider: undefined,
|
||||
executorEscalationModelId: undefined,
|
||||
executorEscalationNodeId: undefined,
|
||||
/**
|
||||
* FNXC:Merge 2026-06-26-00:00:
|
||||
* New and unconfigured projects default AI merge to sync a dirty checked-out integration branch, restoring the legacy stash → fast-forward → restore landing behavior. Explicit persisted merger.allowDirtyLocalCheckoutSync values still win, and no existing-project migration stamps this default into storage.
|
||||
|
||||
@@ -1166,7 +1166,7 @@ export class TaskStore extends EventEmitter<TaskStoreEvents> {
|
||||
}
|
||||
async updateTask(
|
||||
id: string,
|
||||
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("./types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("./types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("./types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("./types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("./types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("./types.js").TaskReview | null; reviewState?: import("./types.js").TaskReviewState | null; workflowStepResults?: import("./types.js").WorkflowStepResult[] | null; mergeDetails?: import("./types.js").MergeDetails | null; sourceIssue?: import("./types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("./types.js").TaskGithubTracking | null; tokenUsage?: import("./types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("./types.js").WorkflowTransitionNotificationMarker | undefined; plannerOversightLevel?: string | null; sessionAdvisorEnabled?: boolean | null; approvedPlanFingerprint?: string | null }, runContext?: RunMutationContext,
|
||||
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("./types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("./types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("./types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("./types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("./types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; executorEscalationAttempted?: boolean | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("./types.js").TaskReview | null; reviewState?: import("./types.js").TaskReviewState | null; workflowStepResults?: import("./types.js").WorkflowStepResult[] | null; mergeDetails?: import("./types.js").MergeDetails | null; sourceIssue?: import("./types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("./types.js").TaskGithubTracking | null; tokenUsage?: import("./types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("./types.js").WorkflowTransitionNotificationMarker | undefined; plannerOversightLevel?: string | null; sessionAdvisorEnabled?: boolean | null; approvedPlanFingerprint?: string | null }, runContext?: RunMutationContext,
|
||||
): Promise<Task> {
|
||||
return updateTaskImpl(this, id, updates, runContext);
|
||||
}
|
||||
|
||||
@@ -48,6 +48,7 @@ export interface TaskRow {
|
||||
resumeLimboCount: number | null;
|
||||
graphResumeRetryCount: number | null;
|
||||
consecutiveToolFailureRetryCount: number | null;
|
||||
executorEscalationAttempted: number | null;
|
||||
toolFailureDetectorLogCursor: number | null;
|
||||
toolFailureRetryExhaustedAuditEmitted: number | null;
|
||||
resumeLimboTipSha: string | null;
|
||||
@@ -232,6 +233,7 @@ export const TASK_COLUMN_DESCRIPTORS: TaskColumnDescriptor[] = [
|
||||
defineTaskColumn("resumeLimboCount", (task) => task.resumeLimboCount ?? 0),
|
||||
defineTaskColumn("graphResumeRetryCount", (task) => task.graphResumeRetryCount === undefined ? 0 : task.graphResumeRetryCount),
|
||||
defineTaskColumn("consecutiveToolFailureRetryCount", (task) => task.consecutiveToolFailureRetryCount ?? 0),
|
||||
defineTaskColumn("executorEscalationAttempted", (task) => task.executorEscalationAttempted ? 1 : 0),
|
||||
defineTaskColumn("toolFailureDetectorLogCursor", (task) => task.toolFailureDetectorLogCursor ?? null),
|
||||
// FNXC:ExecutorToolFailureRetry 2026-07-16-13:30: PostgreSQL stores this CAS marker as integer (0/1), matching its schema and legacy SQLite flag representation.
|
||||
defineTaskColumn("toolFailureRetryExhaustedAuditEmitted", (task) => task.toolFailureRetryExhaustedAuditEmitted ? 1 : 0),
|
||||
|
||||
@@ -44,7 +44,7 @@ export function getTaskSelectClauseWithActivityLogLimitImpl(store: TaskStore, li
|
||||
"modelPresetId", "modelProvider", "modelId",
|
||||
"validatorModelProvider", "validatorModelId",
|
||||
"planningModelProvider", "planningModelId",
|
||||
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
|
||||
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "executorEscalationAttempted", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
|
||||
"error", "summary", "thinkingLevel", "validatorThinkingLevel", "planningThinkingLevel", "executionMode",
|
||||
"tokenUsageInputTokens", "tokenUsageOutputTokens", "tokenUsageCachedTokens", "tokenUsageCacheWriteTokens", "tokenUsageTotalTokens", "tokenUsageFirstUsedAt", "tokenUsageLastUsedAt", "tokenUsageModelProvider", "tokenUsageModelId", "tokenUsagePerModel", "tokenBudgetSoftAlertedAt", "tokenBudgetHardAlertedAt", "tokenBudgetOverride",
|
||||
"createdAt", "updatedAt", "columnMovedAt", "firstExecutionAt", "cumulativeActiveMs", "executionStartedAt", "executionCompletedAt",
|
||||
|
||||
@@ -34,7 +34,7 @@ export function getTaskSelectClauseImpl2(store: TaskStore, slim: boolean, tableA
|
||||
"modelPresetId", "modelProvider", "modelId",
|
||||
"validatorModelProvider", "validatorModelId",
|
||||
"planningModelProvider", "planningModelId",
|
||||
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
|
||||
"mergeRetries", "workflowStepRetries", "stuckKillCount", "resumeLimboCount", "executeRequeueLoopCount", "graphResumeRetryCount", "consecutiveToolFailureRetryCount", "executorEscalationAttempted", "toolFailureDetectorLogCursor", "toolFailureRetryExhaustedAuditEmitted", "resumeLimboTipSha", "resumeLimboStepSignature", "executeRequeueLoopSignature", "postReviewFixCount", "planReviewReplanCount", "recoveryRetryCount", "taskDoneRetryCount", "worktreeSessionRetryCount", "completionHandoffLimboRecoveryCount", "verificationFailureCount", "mergeConflictBounceCount", "mergeAuditBounceCount", "mergeTransientRetryCount", "branchConflictRecoveryCount", "reviewerContextRetryCount", "reviewerFallbackRetryCount", "nextRecoveryAt",
|
||||
"error", "summary", "thinkingLevel", "validatorThinkingLevel", "planningThinkingLevel", "executionMode",
|
||||
"tokenUsageInputTokens", "tokenUsageOutputTokens", "tokenUsageCachedTokens", "tokenUsageCacheWriteTokens", "tokenUsageTotalTokens", "tokenUsageFirstUsedAt", "tokenUsageLastUsedAt", "tokenUsageModelProvider", "tokenUsageModelId", "tokenUsagePerModel", "tokenBudgetSoftAlertedAt", "tokenBudgetHardAlertedAt", "tokenBudgetOverride",
|
||||
"createdAt", "updatedAt", "columnMovedAt", "firstExecutionAt", "cumulativeActiveMs", "executionStartedAt", "executionCompletedAt",
|
||||
|
||||
@@ -624,7 +624,7 @@ export async function resetPromptCheckboxesImpl(store: TaskStore, dir: string):
|
||||
|
||||
export async function updateTaskImpl(store: TaskStore,
|
||||
id: string,
|
||||
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("../types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("../types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("../types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("../types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("../types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("../types.js").TaskReview | null; reviewState?: import("../types.js").TaskReviewState | null; workflowStepResults?: import("../types.js").WorkflowStepResult[] | null; mergeDetails?: import("../types.js").MergeDetails | null; sourceIssue?: import("../types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("../types.js").TaskGithubTracking | null; tokenUsage?: import("../types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("../types.js").WorkflowTransitionNotificationMarker | undefined; sessionAdvisorEnabled?: boolean | null }, runContext?: RunMutationContext,
|
||||
updates: { title?: string; description?: string; priority?: TaskPriority | null; prompt?: string; worktree?: string | null; workspaceWorktrees?: import("../types.js").Task["workspaceWorktrees"]; status?: string | null; dependencies?: string[]; steps?: import("../types.js").TaskStep[]; customFields?: Record<string, unknown>; currentStep?: number; blockedBy?: string | null; overlapBlockedBy?: string | null; assignedAgentId?: string | null; pausedByAgentId?: string | null; pausedReason?: string | null; tokenBudgetSoftAlertedAt?: string | null; worktrunkFallbackAlertedAt?: string | null; worktrunkFailure?: import("../types.js").Task["worktrunkFailure"] | null; tokenBudgetHardAlertedAt?: string | null; tokenBudgetOverride?: import("../types.js").TaskTokenBudgetOverride | null; dispatchStormCount?: number | null; lastDispatchAt?: string | null; assigneeUserId?: string | null; scopeOverride?: boolean | null; scopeOverrideReason?: string | null; scopeAutoWiden?: string[] | null; nodeId?: string | null; effectiveNodeId?: string | null; effectiveNodeSource?: string | null; checkedOutBy?: string | null; checkedOutAt?: string | null; checkoutNodeId?: string | null; checkoutRunId?: string | null; checkoutLeaseRenewedAt?: string | null; checkoutLeaseEpoch?: number | null; paused?: boolean; baseBranch?: string | null; autoMerge?: boolean | null; branch?: string | null; executionStartBranch?: string | null; baseCommitSha?: string | null; size?: "S" | "M" | "L"; reviewLevel?: number; executionMode?: import("../types.js").ExecutionMode | null; mergeRetries?: number; workflowStepRetries?: number; stuckKillCount?: number | null; resumeLimboCount?: number | null; executeRequeueLoopCount?: number | null; graphResumeRetryCount?: number | null; consecutiveToolFailureRetryCount?: number | null; executorEscalationAttempted?: boolean | null; toolFailureDetectorLogCursor?: number | null; toolFailureRetryExhaustedAuditEmitted?: boolean | null; resumeLimboTipSha?: string | null; resumeLimboStepSignature?: string | null; executeRequeueLoopSignature?: string | null; postReviewFixCount?: number | null; planReviewReplanCount?: number | null; recoveryRetryCount?: number | null; taskDoneRetryCount?: number | null; worktreeSessionRetryCount?: number | null; completionHandoffLimboRecoveryCount?: number | null; verificationFailureCount?: number | null; mergeConflictBounceCount?: number | null; mergeAuditBounceCount?: number | null; mergeTransientRetryCount?: number | null; branchConflictRecoveryCount?: number | null; reviewerContextRetryCount?: number | null; reviewerFallbackRetryCount?: number | null; nextRecoveryAt?: string | null; enabledWorkflowSteps?: string[]; noCommitsExpected?: boolean | null; modelProvider?: string | null; modelId?: string | null; validatorModelProvider?: string | null; validatorModelId?: string | null; planningModelProvider?: string | null; planningModelId?: string | null; thinkingLevel?: string | null; validatorThinkingLevel?: string | null; planningThinkingLevel?: string | null; error?: string | null; summary?: string | null; sessionFile?: string | null; firstExecutionAt?: string | null; cumulativeActiveMs?: number | null; executionStartedAt?: string | null; executionCompletedAt?: string | null; review?: import("../types.js").TaskReview | null; reviewState?: import("../types.js").TaskReviewState | null; workflowStepResults?: import("../types.js").WorkflowStepResult[] | null; mergeDetails?: import("../types.js").MergeDetails | null; sourceIssue?: import("../types.js").TaskSourceIssue | null; sourceMetadataPatch?: Record<string, unknown> | null; githubTracking?: import("../types.js").TaskGithubTracking | null; tokenUsage?: import("../types.js").TaskTokenUsage | null; modifiedFiles?: string[] | null; missionId?: string | null; sliceId?: string | null; workflowTransitionNotification?: import("../types.js").WorkflowTransitionNotificationMarker | undefined; sessionAdvisorEnabled?: boolean | null }, runContext?: RunMutationContext,
|
||||
): Promise<Task> {
|
||||
/*
|
||||
FNXC:StateMachine 2026-07-07-12:00:
|
||||
|
||||
@@ -102,6 +102,7 @@ export function rowToTask(row: TaskRow): Task {
|
||||
resumeLimboCount: row.resumeLimboCount ?? undefined,
|
||||
graphResumeRetryCount: row.graphResumeRetryCount ?? undefined,
|
||||
consecutiveToolFailureRetryCount: row.consecutiveToolFailureRetryCount ?? undefined,
|
||||
executorEscalationAttempted: row.executorEscalationAttempted ? true : undefined,
|
||||
toolFailureDetectorLogCursor: row.toolFailureDetectorLogCursor ?? undefined,
|
||||
toolFailureRetryExhaustedAuditEmitted: row.toolFailureRetryExhaustedAuditEmitted ? true : undefined,
|
||||
resumeLimboTipSha: row.resumeLimboTipSha || undefined,
|
||||
|
||||
@@ -367,6 +367,8 @@ export async function updateTaskUnlockedImpl(store: TaskStore, id: string, updat
|
||||
}
|
||||
if (updates.consecutiveToolFailureRetryCount === null) task.consecutiveToolFailureRetryCount = null;
|
||||
else if (updates.consecutiveToolFailureRetryCount !== undefined) task.consecutiveToolFailureRetryCount = updates.consecutiveToolFailureRetryCount;
|
||||
if (updates.executorEscalationAttempted === null) task.executorEscalationAttempted = null;
|
||||
else if (updates.executorEscalationAttempted !== undefined) task.executorEscalationAttempted = updates.executorEscalationAttempted;
|
||||
if (updates.toolFailureDetectorLogCursor === null) task.toolFailureDetectorLogCursor = null;
|
||||
else if (updates.toolFailureDetectorLogCursor !== undefined) task.toolFailureDetectorLogCursor = updates.toolFailureDetectorLogCursor;
|
||||
if (updates.toolFailureRetryExhaustedAuditEmitted === null) task.toolFailureRetryExhaustedAuditEmitted = null;
|
||||
|
||||
@@ -1618,6 +1618,11 @@ export interface Task {
|
||||
* FN-7996 persists the bounded same-model retry budget for consecutive terminal tool errors. The executor atomically claims it per run cursor so concurrent failures cannot exceed the configured cap.
|
||||
*/
|
||||
consecutiveToolFailureRetryCount?: number | null;
|
||||
/**
|
||||
* FNXC:ExecutorEscalation 2026-07-16-21:00:
|
||||
* Records consumption of the one opt-in alternate model/node attempt after FN-7996 exhausts same-model retries. Reset with the retry window so unrelated failure surfaces receive their own bounded escalation.
|
||||
*/
|
||||
executorEscalationAttempted?: boolean | null;
|
||||
/** Agent-log boundary captured at executor-run start; only later terminal outcomes qualify. */
|
||||
toolFailureDetectorLogCursor?: number | null;
|
||||
/** Durable compare-and-set marker which permits one exhaustion audit per retry window. */
|
||||
@@ -3052,6 +3057,14 @@ export interface ProjectSettings {
|
||||
executorToolFailureRetryCount?: number;
|
||||
executorToolFailureRetryBackoffMs?: number;
|
||||
executorToolFailureThreshold?: number;
|
||||
/**
|
||||
* FNXC:ExecutorEscalation 2026-07-16-21:00:
|
||||
* Opt-in single-shot escalation runs only after FN-7996 exhausts same-model retries. The task model target enters the model-selection hierarchy as an override and the node target enters routing as a task override; default off prevents surprise cost or behavior changes.
|
||||
*/
|
||||
executorModelEscalationEnabled?: boolean;
|
||||
executorEscalationProvider?: string;
|
||||
executorEscalationModelId?: string;
|
||||
executorEscalationNodeId?: string;
|
||||
/**
|
||||
* FNXC:VerificationConcurrency 2026-07-15-03:35:
|
||||
* Max concurrent verification subprocesses (fn_run_verification / merge testCommand builds) across all tasks in this process. Caps stacked monorepo typecheck/build pegging CPU when many tasks are in-progress. Default 1. Raise only on high-core hosts.
|
||||
|
||||
@@ -1140,6 +1140,10 @@ export function SettingsModal({
|
||||
executorToolFailureRetryCount: 2,
|
||||
executorToolFailureRetryBackoffMs: 2000,
|
||||
executorToolFailureThreshold: 3,
|
||||
executorModelEscalationEnabled: false,
|
||||
executorEscalationProvider: "",
|
||||
executorEscalationModelId: "",
|
||||
executorEscalationNodeId: "",
|
||||
mergeIntegrationWorktree: "reuse-task-worktree",
|
||||
mergeAdvanceAutoSync: "stash-and-ff",
|
||||
merger: { mode: "ai", maxReviewPasses: 3, allowDirtyLocalCheckoutSync: true },
|
||||
@@ -1724,6 +1728,10 @@ export function SettingsModal({
|
||||
executorToolFailureRetryCount: resolveNonNegativeExecutorToolFailureSetting(s.executorToolFailureRetryCount, 2),
|
||||
executorToolFailureRetryBackoffMs: resolveNonNegativeExecutorToolFailureSetting(s.executorToolFailureRetryBackoffMs, 2000),
|
||||
executorToolFailureThreshold: Math.max(1, Math.floor(Number(s.executorToolFailureThreshold ?? 3) || 3)),
|
||||
executorModelEscalationEnabled: s.executorModelEscalationEnabled === true,
|
||||
executorEscalationProvider: s.executorEscalationProvider ?? "",
|
||||
executorEscalationModelId: s.executorEscalationModelId ?? "",
|
||||
executorEscalationNodeId: s.executorEscalationNodeId ?? "",
|
||||
worktreeCopyFiles: Array.isArray(s.worktreeCopyFiles) ? s.worktreeCopyFiles : [],
|
||||
};
|
||||
setForm(normalizedSettings);
|
||||
@@ -3340,6 +3348,10 @@ export function SettingsModal({
|
||||
executorToolFailureRetryCount: resolveNonNegativeExecutorToolFailureSetting(form.executorToolFailureRetryCount, 2),
|
||||
executorToolFailureRetryBackoffMs: resolveNonNegativeExecutorToolFailureSetting(form.executorToolFailureRetryBackoffMs, 2000),
|
||||
executorToolFailureThreshold: Math.max(1, Math.floor(Number(form.executorToolFailureThreshold ?? 3) || 3)),
|
||||
executorModelEscalationEnabled: form.executorModelEscalationEnabled === true,
|
||||
executorEscalationProvider: form.executorEscalationProvider?.trim() || undefined,
|
||||
executorEscalationModelId: form.executorEscalationModelId?.trim() || undefined,
|
||||
executorEscalationNodeId: form.executorEscalationNodeId?.trim() || undefined,
|
||||
taskPrefix: form.taskPrefix?.trim() || undefined,
|
||||
githubTrackingDefaultRepo: form.githubTrackingDefaultRepo?.trim() || undefined,
|
||||
/*
|
||||
|
||||
@@ -118,6 +118,10 @@ const PROJECT_SECTION_KEYS: Record<string, readonly string[]> = {
|
||||
"executorToolFailureRetryCount",
|
||||
"executorToolFailureRetryBackoffMs",
|
||||
"executorToolFailureThreshold",
|
||||
"executorModelEscalationEnabled",
|
||||
"executorEscalationProvider",
|
||||
"executorEscalationModelId",
|
||||
"executorEscalationNodeId",
|
||||
"groupOverlappingFiles",
|
||||
"heartbeatScopeDiscipline",
|
||||
"ignoreHiddenOverlapPaths",
|
||||
|
||||
@@ -66,6 +66,42 @@ export const schedulingSearchEntries: SettingsSearchEntry[] = [
|
||||
helpFallback: "Terminal tool errors required before retrying. Default: 3.",
|
||||
keywords: ["executor", "tool error", "threshold", "auto retry"],
|
||||
},
|
||||
{
|
||||
sectionId: "scheduling",
|
||||
key: "executorModelEscalationEnabled",
|
||||
labelKey: "settings.scheduling.executorModelEscalationEnabled",
|
||||
labelFallback: "Escalate after tool-failure retries",
|
||||
helpKey: "settings.scheduling.executorModelEscalationEnabledHelp",
|
||||
helpFallback: "After same-model retries are exhausted, try one configured alternate model or node. Disabled by default.",
|
||||
keywords: ["executor", "model", "node", "escalation", "tool error"],
|
||||
},
|
||||
{
|
||||
sectionId: "scheduling",
|
||||
key: "executorEscalationProvider",
|
||||
labelKey: "settings.scheduling.executorEscalationProvider",
|
||||
labelFallback: "Escalation provider",
|
||||
helpKey: "settings.scheduling.executorEscalationProviderHelp",
|
||||
helpFallback: "Provider for the alternate model. Requires an alternate model ID.",
|
||||
keywords: ["executor", "model", "provider", "escalation"],
|
||||
},
|
||||
{
|
||||
sectionId: "scheduling",
|
||||
key: "executorEscalationModelId",
|
||||
labelKey: "settings.scheduling.executorEscalationModelId",
|
||||
labelFallback: "Escalation model ID",
|
||||
helpKey: "settings.scheduling.executorEscalationModelIdHelp",
|
||||
helpFallback: "Alternate model ID. Requires an escalation provider.",
|
||||
keywords: ["executor", "model", "escalation"],
|
||||
},
|
||||
{
|
||||
sectionId: "scheduling",
|
||||
key: "executorEscalationNodeId",
|
||||
labelKey: "settings.scheduling.executorEscalationNodeId",
|
||||
labelFallback: "Escalation node ID",
|
||||
helpKey: "settings.scheduling.executorEscalationNodeIdHelp",
|
||||
helpFallback: "Optional configured node; a node target re-enters scheduler routing.",
|
||||
keywords: ["executor", "node", "routing", "escalation"],
|
||||
},
|
||||
{
|
||||
sectionId: "scheduling",
|
||||
key: "pollIntervalMs",
|
||||
|
||||
@@ -3,6 +3,7 @@ import { MovedSettingsStub } from "./MovedSettingsStub";
|
||||
import { SettingsToggleRow } from "../SettingsToggleRow";
|
||||
import { SettingsSelectRow } from "../SettingsSelectRow";
|
||||
import { SettingsNumberRow } from "../SettingsNumberRow";
|
||||
import { SettingsTextRow } from "../SettingsTextRow";
|
||||
import { SettingsHelpTip } from "../SettingsHelpTip";
|
||||
import type { SettingsFormState, SetSettingsForm } from "./context";
|
||||
const MS_PER_DAY = 24 * 60 * 60 * 1000;
|
||||
@@ -37,6 +38,11 @@ export function SchedulingSection({ form, setForm, concurrencyLoading = false, o
|
||||
<SettingsNumberRow descriptor={{ key: "executorToolFailureRetryCount", label: t("settings.scheduling.executorToolFailureRetryCount", "Executor tool-failure retries"), help: t("settings.scheduling.executorToolFailureRetryCountHelp", "Same-model retries after consecutive tool-call failures. Set 0 to disable. Default: 2."), scope: "project", min: 0, step: 1 }} value={form.executorToolFailureRetryCount ?? 2} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureRetryCount: Math.max(0, Math.floor(v ?? 2)) } as SettingsFormState))} />
|
||||
<SettingsNumberRow descriptor={{ key: "executorToolFailureRetryBackoffMs", label: t("settings.scheduling.executorToolFailureRetryBackoffMs", "Tool-failure retry backoff (ms)"), help: t("settings.scheduling.executorToolFailureRetryBackoffMsHelp", "Unref'd wait before retrying. Default: 2000."), scope: "project", min: 0, step: 1 }} value={form.executorToolFailureRetryBackoffMs ?? 2000} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureRetryBackoffMs: Math.max(0, Math.floor(v ?? 2000)) } as SettingsFormState))} />
|
||||
<SettingsNumberRow descriptor={{ key: "executorToolFailureThreshold", label: t("settings.scheduling.executorToolFailureThreshold", "Consecutive tool failures"), help: t("settings.scheduling.executorToolFailureThresholdHelp", "Terminal tool errors required before retrying. Default: 3."), scope: "project", min: 1, step: 1 }} value={form.executorToolFailureThreshold ?? 3} onChange={(v) => setForm((f) => ({ ...f, executorToolFailureThreshold: Math.max(1, Math.floor(v ?? 3)) } as SettingsFormState))} />
|
||||
{/* FNXC:ExecutorEscalation 2026-07-16-21:00: Keep alternate model/node escalation opt-in and adjacent to its FN-7996 retry policy; a complete model pair or node id is required before the executor consumes its one extra attempt. */}
|
||||
<SettingsToggleRow descriptor={{ key: "executorModelEscalationEnabled", label: t("settings.scheduling.executorModelEscalationEnabled", "Escalate after tool-failure retries"), help: t("settings.scheduling.executorModelEscalationEnabledHelp", "After same-model retries are exhausted, try one configured alternate model or node. Disabled by default."), scope: "project" }} value={form.executorModelEscalationEnabled === true} onChange={(value) => setForm((f) => ({ ...f, executorModelEscalationEnabled: value === true } as SettingsFormState))} />
|
||||
<SettingsTextRow descriptor={{ key: "executorEscalationProvider", label: t("settings.scheduling.executorEscalationProvider", "Escalation provider"), help: t("settings.scheduling.executorEscalationProviderHelp", "Provider for the alternate model. Requires an alternate model ID."), scope: "project" }} value={form.executorEscalationProvider ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationProvider: value ?? "" } as SettingsFormState))} />
|
||||
<SettingsTextRow descriptor={{ key: "executorEscalationModelId", label: t("settings.scheduling.executorEscalationModelId", "Escalation model ID"), help: t("settings.scheduling.executorEscalationModelIdHelp", "Alternate model ID. Requires an escalation provider."), scope: "project" }} value={form.executorEscalationModelId ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationModelId: value ?? "" } as SettingsFormState))} />
|
||||
<SettingsTextRow descriptor={{ key: "executorEscalationNodeId", label: t("settings.scheduling.executorEscalationNodeId", "Escalation node ID"), help: t("settings.scheduling.executorEscalationNodeIdHelp", "Optional configured node; a node target re-enters scheduler routing."), scope: "project" }} value={form.executorEscalationNodeId ?? ""} onChange={(value) => setForm((f) => ({ ...f, executorEscalationNodeId: value ?? "" } as SettingsFormState))} />
|
||||
<SettingsNumberRow
|
||||
descriptor={{
|
||||
key: "maxConcurrent",
|
||||
|
||||
@@ -201,6 +201,10 @@ const SETTING_DESCRIPTION_KEYS: Record<string, string> = {
|
||||
executorToolFailureRetryCount: "scheduling.executorToolFailureRetryCountHelp",
|
||||
executorToolFailureRetryBackoffMs: "scheduling.executorToolFailureRetryBackoffMsHelp",
|
||||
executorToolFailureThreshold: "scheduling.executorToolFailureThresholdHelp",
|
||||
executorModelEscalationEnabled: "scheduling.executorModelEscalationEnabledHelp",
|
||||
executorEscalationProvider: "scheduling.executorEscalationProviderHelp",
|
||||
executorEscalationModelId: "scheduling.executorEscalationModelIdHelp",
|
||||
executorEscalationNodeId: "scheduling.executorEscalationNodeIdHelp",
|
||||
taskStuckTimeoutMs: "scheduling.timeoutInMinutesForDetectingStuckTasksWhen",
|
||||
staleHighFanoutBlockerAgeThresholdMs: "scheduling.escalateHighFanOutBlockersOnlyAfterThey",
|
||||
preserveProgressOnStuckRequeue: "scheduling.whenTheStuckDetectorKillsAndReQueues",
|
||||
|
||||
@@ -41,9 +41,9 @@ function graphFailure() {
|
||||
};
|
||||
}
|
||||
|
||||
function makeHarness(options: { retries: number; entries: Array<{ type: string }> }) {
|
||||
function makeHarness(options: { retries: number; entries: Array<{ type: string }>; settings?: Record<string, unknown>; task?: Partial<TaskDetail> }) {
|
||||
const store = createMockStore();
|
||||
const task = makeTask();
|
||||
const task = makeTask(options.task);
|
||||
store.getTask.mockResolvedValue(task);
|
||||
store.getSettings.mockResolvedValue({
|
||||
maxConcurrent: 2,
|
||||
@@ -53,10 +53,12 @@ function makeHarness(options: { retries: number; entries: Array<{ type: string }
|
||||
executorToolFailureRetryCount: options.retries,
|
||||
executorToolFailureRetryBackoffMs: 0,
|
||||
executorToolFailureThreshold: 3,
|
||||
...options.settings,
|
||||
});
|
||||
store.getAgentLogCount = vi.fn().mockResolvedValue(options.entries.length);
|
||||
store.getAgentLogs = vi.fn().mockResolvedValue(options.entries);
|
||||
store.claimNextToolFailureRetry = vi.fn().mockResolvedValue({ outcome: "claimed", attempt: 1 });
|
||||
store.updateTask.mockImplementation(async (_id: string, patch: Partial<TaskDetail>) => Object.assign(task, patch));
|
||||
store.updateTaskAtomic = vi.fn(async (_id: string, updater: (current: TaskDetail) => Partial<TaskDetail> | null) => {
|
||||
const updates = updater(task);
|
||||
if (updates) Object.assign(task, updates);
|
||||
@@ -123,6 +125,91 @@ describe("executor consecutive tool-failure retry (FN-7996)", () => {
|
||||
});
|
||||
});
|
||||
|
||||
it("escalates once to a configured model after same-model retries exhaust", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 2,
|
||||
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
|
||||
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
|
||||
});
|
||||
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
|
||||
const execute = vi.spyOn(executor as any, "execute").mockResolvedValue(undefined);
|
||||
|
||||
await (executor as any).handleGraphFailure(task, graphFailure());
|
||||
await vi.advanceTimersByTimeAsync(0);
|
||||
|
||||
expect(task).toMatchObject({ modelProvider: "anthropic", modelId: "claude-sonnet", executorEscalationAttempted: true, status: null, error: null });
|
||||
expect(execute).toHaveBeenCalledWith(task);
|
||||
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-retry", metadata: expect.objectContaining({ taskId: task.id, hasModelTarget: true, hasNodeTarget: false }) }));
|
||||
expect(store.updateTaskAtomic).toHaveBeenCalledWith(task.id, expect.any(Function), undefined);
|
||||
});
|
||||
|
||||
it("requeues a node escalation for scheduler effective-node resolution", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 2,
|
||||
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
|
||||
settings: { executorModelEscalationEnabled: true, executorEscalationNodeId: "cursor-node" },
|
||||
});
|
||||
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
|
||||
const execute = vi.spyOn(executor as any, "execute").mockResolvedValue(undefined);
|
||||
|
||||
await (executor as any).handleGraphFailure(task, graphFailure());
|
||||
|
||||
expect(task).toMatchObject({ nodeId: "cursor-node", column: "todo", executorEscalationAttempted: true, status: null, error: null });
|
||||
expect(execute).not.toHaveBeenCalled();
|
||||
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-retry", metadata: expect.objectContaining({ hasNodeTarget: true }) }));
|
||||
});
|
||||
|
||||
it("parks the single escalated attempt and records escalation exhaustion", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 2,
|
||||
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
|
||||
task: { executorEscalationAttempted: true },
|
||||
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
|
||||
});
|
||||
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
|
||||
|
||||
await (executor as any).handleGraphFailure(task, graphFailure());
|
||||
|
||||
expect(task.status).toBe("failed");
|
||||
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({ mutationType: "task:execution-escalation-exhausted" }));
|
||||
});
|
||||
|
||||
it("does not let a concurrent exhausted handler park the escalation it lost", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 2,
|
||||
entries: [{ type: "tool_error" }, { type: "tool_error" }, { type: "tool_error" }],
|
||||
settings: { executorModelEscalationEnabled: true, executorEscalationProvider: "anthropic", executorEscalationModelId: "claude-sonnet" },
|
||||
});
|
||||
store.claimNextToolFailureRetry.mockResolvedValue({ outcome: "exhausted" });
|
||||
task.executorEscalationAttempted = true;
|
||||
// The atomic escalation claim invalidates its exhausted cursor before scheduling.
|
||||
task.toolFailureDetectorLogCursor = null;
|
||||
|
||||
await (executor as any).handleGraphFailure(task, graphFailure());
|
||||
|
||||
expect(task).toMatchObject({ status: null, error: null, executorEscalationAttempted: true });
|
||||
expect(store.updateTask).not.toHaveBeenCalledWith(task.id, expect.objectContaining({ status: "failed" }), expect.anything());
|
||||
expect(store.recordRunAuditEvent).not.toHaveBeenCalledWith(expect.objectContaining({
|
||||
mutationType: "task:execution-escalation-exhausted",
|
||||
}));
|
||||
});
|
||||
|
||||
it("audits a terminal escalated failure even after escalation is disabled", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 0,
|
||||
entries: [],
|
||||
task: { executorEscalationAttempted: true, modelProvider: "anthropic", modelId: "claude-sonnet" },
|
||||
});
|
||||
|
||||
await (executor as any).handleGraphFailure(task, graphFailure());
|
||||
|
||||
expect(task.status).toBe("failed");
|
||||
expect(store.recordRunAuditEvent).toHaveBeenCalledWith(expect.objectContaining({
|
||||
mutationType: "task:execution-escalation-exhausted",
|
||||
metadata: expect.objectContaining({ hadModelTarget: true, hadNodeTarget: false }),
|
||||
}));
|
||||
});
|
||||
|
||||
it("does not let an exhausted stale handler park a newer cursor-owned run", async () => {
|
||||
const { executor, store, task } = makeHarness({
|
||||
retries: 2,
|
||||
|
||||
@@ -13,7 +13,7 @@ import { existsSync, lstatSync, realpathSync } from "node:fs";
|
||||
import { readFile, rm, writeFile } from "node:fs/promises";
|
||||
import type { TaskStore, Task, TaskDetail, TaskTokenUsage, StepStatus, Settings, WorkflowStep, MissionStore, AsyncMissionStore, Slice, AgentState, AgentCapability, RunMutationContext, AgentHeartbeatConfig, Agent, AgentMemoryInclusionMode, ProjectSettings, MergeResult, WorkflowIrNode, WorkflowIrNodeKind, WorkflowStepResult as CoreWorkflowStepResult, ThinkingLevel } from "@fusion/core";
|
||||
import { getUnmetSchedulingDependencies } from "./scheduler.js";
|
||||
import { RetryStormError, serializeRetryStormError, isExperimentalFeatureEnabled, resolveWorkflowIrForTask, resolveColumnAgentBinding, resolveEffectiveAgent, instanceNodeId, getWorkflowExtensionRegistry, getBuiltinWorkflow, parseNoOpCompletionMarker, allowsAutoMergeProcessing, resolveEffectiveAutoMerge, isLiveSharedBranchGroupMemberIntegration, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveOptionalStepRevisionBudget, resolveOptionalReviewRevisionBudget, COMPLETION_SUMMARY_NODE_ID, upsertWorkflowStepResult, AWAITING_APPROVAL_PAUSE_REASON, THINKING_LEVELS, AgentStore, resolveExecutorFallbackModel } from "@fusion/core";
|
||||
import { RetryStormError, serializeRetryStormError, isExperimentalFeatureEnabled, resolveWorkflowIrForTask, resolveColumnAgentBinding, resolveEffectiveAgent, instanceNodeId, getWorkflowExtensionRegistry, getBuiltinWorkflow, parseNoOpCompletionMarker, allowsAutoMergeProcessing, resolveEffectiveAutoMerge, isLiveSharedBranchGroupMemberIntegration, resolveMaxAutoMergeRetries, resolveMaxConsecutiveToolFailureRetries, resolveConsecutiveToolFailureRetryBackoffMs, resolveConsecutiveToolFailureThreshold, resolveExecutorEscalationTarget, resolveOptionalStepRevisionBudget, resolveOptionalReviewRevisionBudget, COMPLETION_SUMMARY_NODE_ID, upsertWorkflowStepResult, AWAITING_APPROVAL_PAUSE_REASON, THINKING_LEVELS, AgentStore, resolveExecutorFallbackModel } from "@fusion/core";
|
||||
import { finalizeProvenAutoMergeTask } from "./auto-merge-finalization.js";
|
||||
import { mergeEffectiveSettings } from "./effective-settings.js";
|
||||
import { moveTaskToReplanColumn, resolveReplanTargetColumn } from "./replan-target.js";
|
||||
@@ -5419,7 +5419,7 @@ export class TaskExecutor {
|
||||
await this.finalizeMergeConfirmedWorkflowGraphTask(task.id, "graph-completed");
|
||||
}
|
||||
if ((live.graphResumeRetryCount ?? 0) !== 0 || (live.consecutiveToolFailureRetryCount ?? 0) !== 0) {
|
||||
await this.store.updateTask(task.id, { graphResumeRetryCount: 0, consecutiveToolFailureRetryCount: 0, toolFailureDetectorLogCursor: null, toolFailureRetryExhaustedAuditEmitted: false }, this.getRunContextFor(task.id));
|
||||
await this.store.updateTask(task.id, { graphResumeRetryCount: 0, consecutiveToolFailureRetryCount: 0, executorEscalationAttempted: false, toolFailureDetectorLogCursor: null, toolFailureRetryExhaustedAuditEmitted: false }, this.getRunContextFor(task.id));
|
||||
}
|
||||
}
|
||||
return true;
|
||||
@@ -9690,12 +9690,13 @@ export class TaskExecutor {
|
||||
return;
|
||||
}
|
||||
const message = `Workflow graph terminated with failure at node '${failedNode ?? "unknown"}'`;
|
||||
const maxToolFailureRetries = resolveMaxConsecutiveToolFailureRetries(await this.store.getSettings());
|
||||
const settings = await this.store.getSettings();
|
||||
const maxToolFailureRetries = resolveMaxConsecutiveToolFailureRetries(settings);
|
||||
const isExecuteFailure = failedNode === "execute" || failedNode?.endsWith(":step-execute") === true || failedNode === "step-execute";
|
||||
if (maxToolFailureRetries > 0 && isExecuteFailure && !live.paused && !live.userPaused && !live.deletedAt && live.column === "in-progress") {
|
||||
// Prefer the execution-local boundary; recovery paths refetch durable state rather than use the stale failure snapshot.
|
||||
const cursor = this.graphToolFailureRunCursors.get(task.id) ?? (await this.store.getTask(task.id))?.toolFailureDetectorLogCursor;
|
||||
const threshold = resolveConsecutiveToolFailureThreshold(await this.store.getSettings());
|
||||
const threshold = resolveConsecutiveToolFailureThreshold(settings);
|
||||
if (await this.hasTrailingConsecutiveToolFailures(task.id, cursor, threshold)) {
|
||||
const claim = await this.store.claimNextToolFailureRetry(task.id, cursor!, maxToolFailureRetries);
|
||||
if (claim.outcome === "claimed") {
|
||||
@@ -9703,11 +9704,56 @@ export class TaskExecutor {
|
||||
await this.store.logEntry(task.id, `Consecutive tool-call failures — auto-retrying same model (${claim.attempt}/${maxToolFailureRetries}) instead of parking`, undefined, this.getRunContextFor(task.id));
|
||||
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("tool-failure-retry", task.id), domain: "database", mutationType: "task:execution-tool-failure-retry", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", attempt: claim.attempt, maxAttempts: maxToolFailureRetries, consecutiveToolFailures: threshold, mode: "same-model" } });
|
||||
const schedule = () => { void (async () => { const resume = await this.store.getTask(task.id); if (resume && !resume.deletedAt && !resume.paused && !resume.userPaused && resume.column === "in-progress") await this.execute(resume); })().catch((error) => executorLog.error(`${task.id}: tool-failure retry failed`, error)); };
|
||||
const delay = resolveConsecutiveToolFailureRetryBackoffMs(await this.store.getSettings());
|
||||
const delay = resolveConsecutiveToolFailureRetryBackoffMs(settings);
|
||||
setTimeout(schedule, delay).unref?.();
|
||||
return;
|
||||
}
|
||||
if (claim.outcome === "already-claimed-for-run") { await this.store.getTask(task.id); return; }
|
||||
/*
|
||||
FNXC:ExecutorEscalation 2026-07-16-21:00:
|
||||
FN-7998 inserts exactly one opt-in recovery between FN-7996 exhaustion and the unchanged terminal park. Refetch before writing so a pause, deletion, or later run cannot inherit a costly model/node override from this stale graph result.
|
||||
*/
|
||||
const escalationTarget = resolveExecutorEscalationTarget(settings);
|
||||
const hasModelTarget = escalationTarget.provider !== undefined && escalationTarget.modelId !== undefined;
|
||||
const hasNodeTarget = escalationTarget.nodeId !== undefined;
|
||||
let claimedEscalation = false;
|
||||
let priorEscalationRetryCount = 0;
|
||||
/*
|
||||
FNXC:ExecutorEscalation 2026-07-16-22:30:
|
||||
The one-shot latch is claimed under the TaskStore lock. Concurrent exhausted
|
||||
graph handlers for the same detector cursor must not both schedule an alternate
|
||||
run; a loser leaves the winner's in-progress row untouched.
|
||||
*/
|
||||
await this.store.updateTaskAtomic(task.id, (current) => {
|
||||
const ownsFailureRun = current.toolFailureDetectorLogCursor === cursor
|
||||
&& current.column === "in-progress"
|
||||
&& !current.paused
|
||||
&& !current.userPaused
|
||||
&& !current.deletedAt;
|
||||
if (!ownsFailureRun || current.executorEscalationAttempted === true || !escalationTarget.enabled) return null;
|
||||
claimedEscalation = true;
|
||||
priorEscalationRetryCount = current.consecutiveToolFailureRetryCount ?? 0;
|
||||
return {
|
||||
...(hasModelTarget ? { modelProvider: escalationTarget.provider, modelId: escalationTarget.modelId } : {}),
|
||||
...(hasNodeTarget ? { nodeId: escalationTarget.nodeId, column: "todo" as const } : {}),
|
||||
executorEscalationAttempted: true,
|
||||
/* FNXC:ExecutorEscalation 2026-07-16-22:40: Invalidate the exhausted run cursor before releasing the claim so concurrent stale handlers cannot park or audit the alternate execution; the alternate captures its own cursor at startup. */
|
||||
toolFailureDetectorLogCursor: null,
|
||||
status: null,
|
||||
error: null,
|
||||
};
|
||||
}, this.getRunContextFor(task.id));
|
||||
if (claimedEscalation) {
|
||||
await this.store.logEntry(task.id, "Same-model retries exhausted — escalating to alternate model/node (one attempt) instead of parking", undefined, this.getRunContextFor(task.id));
|
||||
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-retry", task.id), domain: "database", mutationType: "task:execution-escalation-retry", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hasModelTarget, hasNodeTarget, priorConsecutiveToolFailureRetryCount: priorEscalationRetryCount } });
|
||||
if (!hasNodeTarget) {
|
||||
const scheduleEscalation = () => { void (async () => { const resumeTask = await this.store.getTask(task.id); if (resumeTask && !resumeTask.deletedAt && !resumeTask.paused && !resumeTask.userPaused && resumeTask.column === "in-progress") await this.execute(resumeTask); })().catch((error) => executorLog.error(`${task.id}: escalation retry failed`, error)); };
|
||||
const handle = setTimeout(scheduleEscalation, resolveConsecutiveToolFailureRetryBackoffMs(settings));
|
||||
handle.unref?.();
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
/*
|
||||
FNXC:ExecutorToolFailureRetry 2026-07-16-20:45:
|
||||
Exhaustion belongs to the graph run that supplied `cursor`, not a later run
|
||||
@@ -9717,6 +9763,9 @@ export class TaskExecutor {
|
||||
old terminal handler from parking a newer in-progress executor run.
|
||||
*/
|
||||
let cursorOwnedTerminalPark = false;
|
||||
let escalationAttemptFailed = false;
|
||||
let escalationHadModelTarget = false;
|
||||
let escalationHadNodeTarget = false;
|
||||
await this.store.updateTaskAtomic(task.id, (current) => {
|
||||
if (
|
||||
current.toolFailureDetectorLogCursor !== cursor
|
||||
@@ -9724,28 +9773,65 @@ export class TaskExecutor {
|
||||
|| current.paused
|
||||
|| current.userPaused
|
||||
|| current.deletedAt
|
||||
|| current.status !== null
|
||||
) {
|
||||
return null;
|
||||
}
|
||||
cursorOwnedTerminalPark = true;
|
||||
escalationAttemptFailed = current.executorEscalationAttempted === true;
|
||||
escalationHadModelTarget = current.modelProvider != null && current.modelId != null;
|
||||
escalationHadNodeTarget = current.nodeId != null;
|
||||
return { error: message, status: "failed" };
|
||||
}, this.getRunContextFor(task.id));
|
||||
if (!cursorOwnedTerminalPark) return;
|
||||
if (await this.store.markToolFailureRetryExhaustedAudit(task.id)) {
|
||||
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("tool-failure-retry-exhausted", task.id), domain: "database", mutationType: "task:execution-tool-failure-retry-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", attempts: maxToolFailureRetries, limit: maxToolFailureRetries, outcome: "terminal-park" } });
|
||||
}
|
||||
if (escalationAttemptFailed) {
|
||||
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-exhausted", task.id), domain: "database", mutationType: "task:execution-escalation-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hadModelTarget: escalationHadModelTarget, hadNodeTarget: escalationHadNodeTarget } });
|
||||
}
|
||||
executorLog.warn(`${task.id}: ${message}`);
|
||||
await this.store.logEntry(task.id, message, undefined, this.getRunContextFor(task.id));
|
||||
await this.persistTokenUsage(task.id);
|
||||
return;
|
||||
}
|
||||
}
|
||||
if (live.executorEscalationAttempted === true) {
|
||||
const failureCursor = task.toolFailureDetectorLogCursor;
|
||||
let escalationTerminalParked = false;
|
||||
let escalationHadModelTarget = false;
|
||||
let escalationHadNodeTarget = false;
|
||||
/*
|
||||
FNXC:ExecutorEscalation 2026-07-16-22:35:
|
||||
Once the durable escalation latch is set, every terminal failure of that
|
||||
alternate run emits the exhaustion audit even if an operator disables the
|
||||
setting mid-run. Cursor ownership prevents an old concurrent handler from
|
||||
parking the newly scheduled alternate execution.
|
||||
*/
|
||||
await this.store.updateTaskAtomic(task.id, (current) => {
|
||||
if (
|
||||
current.toolFailureDetectorLogCursor !== failureCursor
|
||||
|| current.column !== "in-progress"
|
||||
|| current.paused
|
||||
|| current.userPaused
|
||||
|| current.deletedAt
|
||||
|| current.status !== null
|
||||
) return null;
|
||||
escalationTerminalParked = true;
|
||||
escalationHadModelTarget = current.modelProvider != null && current.modelId != null;
|
||||
escalationHadNodeTarget = current.nodeId != null;
|
||||
return { error: message, status: "failed" };
|
||||
}, this.getRunContextFor(task.id));
|
||||
if (!escalationTerminalParked) return;
|
||||
await this.store.recordRunAuditEvent?.({ taskId: task.id, agentId: "executor", runId: generateSyntheticRunId("escalation-exhausted", task.id), domain: "database", mutationType: "task:execution-escalation-exhausted", target: task.id, metadata: { taskId: task.id, nodeId: failedNode ?? "unknown", hadModelTarget: escalationHadModelTarget, hadNodeTarget: escalationHadNodeTarget } });
|
||||
} else {
|
||||
// status "failed" doubles as the self-healing exemption: review-task
|
||||
// revival sweeps skip tasks carrying a non-null status, preventing the
|
||||
// FN-5704-style loop of re-running the graph from scratch.
|
||||
await this.store.updateTask(task.id, { error: message, status: "failed" }, this.getRunContextFor(task.id));
|
||||
}
|
||||
executorLog.warn(`${task.id}: ${message}`);
|
||||
await this.store.logEntry(task.id, message, undefined, this.getRunContextFor(task.id));
|
||||
// status "failed" doubles as the self-healing exemption: review-task
|
||||
// revival sweeps skip tasks carrying a non-null status, preventing the
|
||||
// FN-5704-style loop of re-running the graph from scratch.
|
||||
await this.store.updateTask(task.id, { error: message, status: "failed" }, this.getRunContextFor(task.id));
|
||||
await this.persistTokenUsage(task.id);
|
||||
} catch (err) {
|
||||
executorLog.error(
|
||||
|
||||
@@ -6634,6 +6634,14 @@
|
||||
"executorToolFailureRetryBackoffMsHelp": "Unref'd wait before retrying. Default: 2000.",
|
||||
"executorToolFailureThreshold": "Consecutive tool failures",
|
||||
"executorToolFailureThresholdHelp": "Terminal tool errors required before retrying. Default: 3.",
|
||||
"executorModelEscalationEnabled": "Escalate after tool-failure retries",
|
||||
"executorModelEscalationEnabledHelp": "After same-model retries are exhausted, try one configured alternate model or node. Default: disabled.",
|
||||
"executorEscalationProvider": "Escalation provider",
|
||||
"executorEscalationProviderHelp": "Provider for the alternate model. No default — unset; requires an alternate model ID.",
|
||||
"executorEscalationModelId": "Escalation model ID",
|
||||
"executorEscalationModelIdHelp": "Alternate model ID. No default — unset; requires an escalation provider.",
|
||||
"executorEscalationNodeId": "Escalation node ID",
|
||||
"executorEscalationNodeIdHelp": "Optional configured node; a node target re-enters scheduler routing. No default — unset.",
|
||||
"fullAgentLog": "Full agent log",
|
||||
"globalMaxConcurrent": "Global Max Concurrent",
|
||||
"heartbeatScopeDiscipline": "Heartbeat Scope Discipline",
|
||||
|
||||
Reference in New Issue
Block a user