fix(agents): never let a built-in workflow role agent be unroutable
The board stopped moving. Work items churned held -> running -> held at ~3.5/sec across every task, pinning a core and writing ~19k workflowWorkItem audit rows/hour while nothing executed. Hold reason: workflow-principal-role-pool-exhausted:executor. provisionBuiltinWorkflowRoleAgents seeded the four permanent owners (triage, executor, reviewer, merger) with runtimeConfig.enabled=false, while the router's available() treats enabled===false as unavailable. The only permanent principals for every built-in role were unroutable BY CONSTRUCTION — shipped that way, so any instance without operator-created role agents deadlocks at its first workflow node. Nothing self-recovers: a pool only changes by operator action. Routability of these four is an invariant, not a setting. Unlike an operator's agent, disabling one does not opt an agent out — it removes the only thing that can run that stage, and there is no fallback. - seed built-ins enabled; converge existing rows on provisioning - enforceBuiltinWorkflowRoleRoutability coerces enabled back at the durable writeAgent seam, so no REST/UI/plugin/restore path can reintroduce the deadlock. Other runtimeConfig keys are preserved; operator-owned agents keep their off switch - share the static routability predicate (isWorkflowPrincipalEligible) between provisioning and the router so the two cannot drift apart again Also fix the spin itself: a principal hold had no cooldown, so the scheduler re-dispatched instantly and the run re-entered only to re-fence and re-park. It now records a backoff ladder (15s -> 5m) checked before graph entry, and logs once per distinct reason instead of every pass — the same self-recovering shape as holdForSessionContention. The hold never increments `attempt`, so no existing guard could ever fire. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,7 @@
|
|||||||
|
---
|
||||||
|
"@runfusion/fusion": patch
|
||||||
|
---
|
||||||
|
|
||||||
|
summary: Fix a deadlock where built-in workflow agents were unroutable, leaving every task stuck and spinning.
|
||||||
|
category: fix
|
||||||
|
dev: `provisionBuiltinWorkflowRoleAgents` seeded the four permanent owners with `runtimeConfig.enabled: false` while the router's `available()` rejects `enabled === false`, so no built-in role could ever be routed. Built-ins are now seeded enabled, existing rows converge on provisioning, and `enforceBuiltinWorkflowRoleRoutability` coerces them back at the durable `writeAgent` seam so no API/UI/plugin path can disable them. The static routability predicate (`isWorkflowPrincipalEligible`) is shared by provisioning and the router so they cannot drift. Separately, a workflow-principal hold now uses a backoff ladder (`PRINCIPAL_HOLD_BACKOFF_MS`, 15s→5m) checked before graph entry, instead of re-dispatching immediately — the old path spun ~3.5×/sec writing ~19k audit rows/hour with nothing executing.
|
||||||
@@ -0,0 +1,83 @@
|
|||||||
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15 (a built-in workflow owner is never unroutable — regression):
|
||||||
|
|
||||||
|
Reported symptom: every task stopped moving. Work items churned `held → running → held` at ~3.5×/second across
|
||||||
|
the whole board, pinning a CPU core and writing ~19k `workflowWorkItem` audit rows/hour while ZERO work
|
||||||
|
executed. The hold reason was `workflow-principal-role-pool-exhausted:executor`.
|
||||||
|
|
||||||
|
Root cause: `provisionBuiltinWorkflowRoleAgents` seeded the four permanent workflow owners (triage, executor,
|
||||||
|
reviewer, merger) with `runtimeConfig: { enabled: false }`, while the router's `available()` treats
|
||||||
|
`enabled === false` as unavailable. The only permanent principals for every built-in role were therefore
|
||||||
|
unroutable BY CONSTRUCTION — a defect every fresh instance ships with, not a local misconfiguration. Nothing
|
||||||
|
recovers on its own, because the pool can only change through operator action.
|
||||||
|
|
||||||
|
Invariant under test: a built-in workflow role owner is always routable. It is coerced back to routable at the
|
||||||
|
durable write seam, so no caller — REST, dashboard toggle, plugin, provisioning, config restore — can put the
|
||||||
|
system into the deadlocked state; and the static routability predicate is SHARED with the router so
|
||||||
|
"what provisioning produces" and "what routing accepts" cannot drift apart again.
|
||||||
|
*/
|
||||||
|
import { describe, expect, it } from "vitest";
|
||||||
|
import {
|
||||||
|
enforceBuiltinWorkflowRoleRoutability,
|
||||||
|
isBuiltinWorkflowRoleAgent,
|
||||||
|
isWorkflowPrincipalEligible,
|
||||||
|
} from "../agent-role-policy.js";
|
||||||
|
|
||||||
|
const builtIn = (runtimeConfig?: Record<string, unknown>) => ({
|
||||||
|
id: "agent-builtin",
|
||||||
|
metadata: { builtInWorkflowRole: true, workflowRole: "executor" },
|
||||||
|
runtimeConfig,
|
||||||
|
});
|
||||||
|
|
||||||
|
const operatorOwned = (runtimeConfig?: Record<string, unknown>) => ({
|
||||||
|
id: "agent-operator",
|
||||||
|
metadata: {},
|
||||||
|
runtimeConfig,
|
||||||
|
});
|
||||||
|
|
||||||
|
describe("built-in workflow role routability invariant", () => {
|
||||||
|
it("coerces a disabled built-in owner back to routable", () => {
|
||||||
|
const result = enforceBuiltinWorkflowRoleRoutability(builtIn({ enabled: false }));
|
||||||
|
expect(result.runtimeConfig).toMatchObject({ enabled: true });
|
||||||
|
expect(isWorkflowPrincipalEligible(result)).toBe(true);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("preserves every other runtimeConfig key while coercing", () => {
|
||||||
|
// The operator's heartbeat cadence and claim policy are theirs; only routability is non-negotiable.
|
||||||
|
const result = enforceBuiltinWorkflowRoleRoutability(
|
||||||
|
builtIn({ enabled: false, heartbeatIntervalMs: 3_600_000, autoClaimRelevantTasks: true }),
|
||||||
|
);
|
||||||
|
expect(result.runtimeConfig).toEqual({
|
||||||
|
enabled: true,
|
||||||
|
heartbeatIntervalMs: 3_600_000,
|
||||||
|
autoClaimRelevantTasks: true,
|
||||||
|
});
|
||||||
|
});
|
||||||
|
|
||||||
|
it("leaves an operator-owned agent free to be disabled", () => {
|
||||||
|
// The invariant protects the engine's own principals, not every agent — operators keep their off switch.
|
||||||
|
const result = enforceBuiltinWorkflowRoleRoutability(operatorOwned({ enabled: false }));
|
||||||
|
expect(result.runtimeConfig).toMatchObject({ enabled: false });
|
||||||
|
expect(isWorkflowPrincipalEligible(result)).toBe(false);
|
||||||
|
expect(isBuiltinWorkflowRoleAgent(result)).toBe(false);
|
||||||
|
});
|
||||||
|
|
||||||
|
it("is a no-op for an already-routable built-in owner (same reference, no churn)", () => {
|
||||||
|
const agent = builtIn({ enabled: true });
|
||||||
|
expect(enforceBuiltinWorkflowRoleRoutability(agent)).toBe(agent);
|
||||||
|
const unset = builtIn();
|
||||||
|
expect(enforceBuiltinWorkflowRoleRoutability(unset)).toBe(unset);
|
||||||
|
});
|
||||||
|
|
||||||
|
/*
|
||||||
|
The predicate the router consults. `enabled === false` was the exact bit that made the pool look exhausted;
|
||||||
|
paused/errored agents must stay excluded so the coercion never resurrects a genuinely broken principal.
|
||||||
|
*/
|
||||||
|
it("still excludes paused and errored agents from principal routing", () => {
|
||||||
|
expect(isWorkflowPrincipalEligible({ runtimeConfig: { enabled: true }, state: "paused" })).toBe(false);
|
||||||
|
expect(isWorkflowPrincipalEligible({ runtimeConfig: { enabled: true }, state: "error" })).toBe(false);
|
||||||
|
expect(isWorkflowPrincipalEligible({ runtimeConfig: { enabled: true }, state: "active" })).toBe(true);
|
||||||
|
// An unset runtimeConfig is routable: only an explicit `false` opts an agent out.
|
||||||
|
expect(isWorkflowPrincipalEligible({ state: "active" })).toBe(true);
|
||||||
|
});
|
||||||
|
});
|
||||||
@@ -79,6 +79,54 @@ export function canAgentReceiveImplementationTasks(agent: RoleTaggedAgent): bool
|
|||||||
return getAgentAssignmentPolicy(agent) !== "none";
|
return getAgentAssignmentPolicy(agent) !== "none";
|
||||||
}
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
* The STATIC half of workflow-principal routability, shared so provisioning and the router cannot drift.
|
||||||
|
* A disabled runtime, a paused/errored agent, or a transient per-task worker can never own a workflow stage.
|
||||||
|
* The router adds the dynamic half (session capacity); this predicate is the part provisioning must satisfy
|
||||||
|
* for an instance to be able to route a role at all.
|
||||||
|
*
|
||||||
|
* Extracted after every built-in workflow owner shipped `runtimeConfig: { enabled: false }` while the router
|
||||||
|
* treated `enabled === false` as unavailable — so the only permanent principals for triage/executor/reviewer/
|
||||||
|
* merger were unroutable by construction, and any instance without operator-created role agents held at its
|
||||||
|
* first workflow node.
|
||||||
|
*/
|
||||||
|
/** True for the four provenance-marked permanent owners that route built-in workflow stages. */
|
||||||
|
export function isBuiltinWorkflowRoleAgent(agent: { metadata?: Record<string, unknown> | null }): boolean {
|
||||||
|
return agent.metadata?.builtInWorkflowRole === true;
|
||||||
|
}
|
||||||
|
|
||||||
|
/**
|
||||||
|
* FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
* HARD INVARIANT: a built-in workflow role owner is never unroutable.
|
||||||
|
*
|
||||||
|
* These four are the engine's own principals for triage/executor/reviewer/merger. Unlike an operator's agent,
|
||||||
|
* disabling one does not "opt an agent out" — it removes the only thing that can run that workflow stage, and
|
||||||
|
* the engine has no fallback: every task holds at its first node of that role and the hold re-dispatches
|
||||||
|
* forever. That is not a configuration an operator can meaningfully choose, so `runtimeConfig.enabled` is
|
||||||
|
* COERCED back to true for them at the write seam rather than validated and rejected: the write still
|
||||||
|
* succeeds, every other runtimeConfig key the caller sent is preserved, and the system cannot be put into the
|
||||||
|
* deadlocked state by an API call, a UI toggle, a plugin, or a stale record.
|
||||||
|
*
|
||||||
|
* To take a built-in owner out of rotation, add your own agent with that role and route to it — that path
|
||||||
|
* leaves the role routable, which is the property this protects.
|
||||||
|
*/
|
||||||
|
export function enforceBuiltinWorkflowRoleRoutability<T extends {
|
||||||
|
metadata?: Record<string, unknown> | null;
|
||||||
|
runtimeConfig?: Record<string, unknown> | null;
|
||||||
|
}>(agent: T): T {
|
||||||
|
if (!isBuiltinWorkflowRoleAgent(agent)) return agent;
|
||||||
|
if (agent.runtimeConfig?.enabled !== false) return agent;
|
||||||
|
return { ...agent, runtimeConfig: { ...agent.runtimeConfig, enabled: true } };
|
||||||
|
}
|
||||||
|
|
||||||
|
export function isWorkflowPrincipalEligible(
|
||||||
|
agent: Pick<RoleTaggedAgent, "runtimeConfig"> & { state?: string; id?: string },
|
||||||
|
): boolean {
|
||||||
|
if (agent.runtimeConfig?.enabled === false) return false;
|
||||||
|
return agent.state !== "paused" && agent.state !== "error";
|
||||||
|
}
|
||||||
|
|
||||||
export function isImplementationTask(task: Pick<Task, "column">): boolean {
|
export function isImplementationTask(task: Pick<Task, "column">): boolean {
|
||||||
return IMPLEMENTATION_TASK_COLUMNS.has(task.column);
|
return IMPLEMENTATION_TASK_COLUMNS.has(task.column);
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -110,6 +110,7 @@ import {
|
|||||||
BUILTIN_WORKFLOW_ROLE_AGENT_DEFAULT_LIST,
|
BUILTIN_WORKFLOW_ROLE_AGENT_DEFAULT_LIST,
|
||||||
type BuiltinWorkflowRole,
|
type BuiltinWorkflowRole,
|
||||||
} from "./workflow-role-agent-defaults.js";
|
} from "./workflow-role-agent-defaults.js";
|
||||||
|
import { enforceBuiltinWorkflowRoleRoutability } from "./agent-role-policy.js";
|
||||||
|
|
||||||
const agentStoreLog = createLogger("agent-store");
|
const agentStoreLog = createLogger("agent-store");
|
||||||
|
|
||||||
@@ -2058,7 +2059,16 @@ export class AgentStore extends EventEmitter {
|
|||||||
roles: [definition.role],
|
roles: [definition.role],
|
||||||
title: definition.title,
|
title: definition.title,
|
||||||
metadata: { builtInWorkflowRole: true, workflowRole: definition.role },
|
metadata: { builtInWorkflowRole: true, workflowRole: definition.role },
|
||||||
runtimeConfig: { enabled: false },
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
These four ARE the permanent principals that route built-in workflow stages, and the router treats
|
||||||
|
`runtimeConfig.enabled === false` as unavailable. Provisioning them disabled therefore created a
|
||||||
|
system that could not route ANY built-in role: every task held at its first workflow node with
|
||||||
|
`workflow-principal-role-pool-exhausted:<role>`, and the held item was re-claimed with no backoff,
|
||||||
|
burning CPU and ~19k audit rows/hour while nothing executed. Seed them ENABLED so a fresh instance
|
||||||
|
can route out of the box; an operator can still disable one deliberately afterwards.
|
||||||
|
*/
|
||||||
|
runtimeConfig: { enabled: true },
|
||||||
instructionsText: definition.instructionsText,
|
instructionsText: definition.instructionsText,
|
||||||
soul: definition.soul,
|
soul: definition.soul,
|
||||||
bundleConfig: { ...BUILTIN_WORKFLOW_AGENT_BUNDLE_CONFIG, files: [...BUILTIN_WORKFLOW_AGENT_BUNDLE_CONFIG.files] },
|
bundleConfig: { ...BUILTIN_WORKFLOW_AGENT_BUNDLE_CONFIG, files: [...BUILTIN_WORKFLOW_AGENT_BUNDLE_CONFIG.files] },
|
||||||
@@ -2070,6 +2080,11 @@ export class AgentStore extends EventEmitter {
|
|||||||
|| (agent.bundleConfig !== undefined && !canonicalOrPartialBundle)
|
|| (agent.bundleConfig !== undefined && !canonicalOrPartialBundle)
|
||||||
|| (Boolean(agent.instructionsText?.trim()) && agent.instructionsText !== definition.instructionsText);
|
|| (Boolean(agent.instructionsText?.trim()) && agent.instructionsText !== definition.instructionsText);
|
||||||
const updates: Partial<Agent> = {};
|
const updates: Partial<Agent> = {};
|
||||||
|
// FNXC:WorkflowAgentRouting 2026-08-10-01:15: converge every existing built-in owner to routable.
|
||||||
|
// See enforceBuiltinWorkflowRoleRoutability — routability is an invariant for these four, not a setting.
|
||||||
|
if (agent.runtimeConfig?.enabled === false) {
|
||||||
|
updates.runtimeConfig = { ...agent.runtimeConfig, enabled: true };
|
||||||
|
}
|
||||||
if (!customInstructions && !agent.instructionsText?.trim()) updates.instructionsText = definition.instructionsText;
|
if (!customInstructions && !agent.instructionsText?.trim()) updates.instructionsText = definition.instructionsText;
|
||||||
if (!agent.soul?.trim()) updates.soul = definition.soul;
|
if (!agent.soul?.trim()) updates.soul = definition.soul;
|
||||||
// FNXC:WorkflowAgentIdentities 2026-08-08-06:38: Do not rewrite complete canonical
|
// FNXC:WorkflowAgentIdentities 2026-08-08-06:38: Do not rewrite complete canonical
|
||||||
@@ -3183,9 +3198,16 @@ export class AgentStore extends EventEmitter {
|
|||||||
}
|
}
|
||||||
|
|
||||||
private async writeAgent(agent: Agent, executor?: QueryHandle): Promise<void> {
|
private async writeAgent(agent: Agent, executor?: QueryHandle): Promise<void> {
|
||||||
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
The single durable write seam for agents, and therefore the only place the built-in-owner routability
|
||||||
|
invariant cannot be bypassed. Enforcing here — rather than in the REST handler — covers the API, the
|
||||||
|
dashboard toggle, plugins, provisioning, config-revision restores, and any future caller at once.
|
||||||
|
*/
|
||||||
|
const routable = enforceBuiltinWorkflowRoleRoutability(agent);
|
||||||
// FNXC:SqliteFinalRemoval 2026-06-25-23:40:
|
// FNXC:SqliteFinalRemoval 2026-06-25-23:40:
|
||||||
// Backend mode: delegate to async Drizzle writeAgent helper.
|
// Backend mode: delegate to async Drizzle writeAgent helper.
|
||||||
await writeAgentAsync(executor ?? this.asyncLayer!.db, agent, this.asyncLayer!.projectId);
|
await writeAgentAsync(executor ?? this.asyncLayer!.db, routable, this.asyncLayer!.projectId);
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -796,6 +796,9 @@ export {
|
|||||||
getAgentAssignmentPolicy,
|
getAgentAssignmentPolicy,
|
||||||
isAgentAutoAssignable,
|
isAgentAutoAssignable,
|
||||||
canAgentReceiveImplementationTasks,
|
canAgentReceiveImplementationTasks,
|
||||||
|
isWorkflowPrincipalEligible,
|
||||||
|
isBuiltinWorkflowRoleAgent,
|
||||||
|
enforceBuiltinWorkflowRoleRoutability,
|
||||||
evaluateImplementationTaskBind,
|
evaluateImplementationTaskBind,
|
||||||
assertImplementationTaskBindAllowed,
|
assertImplementationTaskBindAllowed,
|
||||||
AgentTaskRoutingPolicyError,
|
AgentTaskRoutingPolicyError,
|
||||||
|
|||||||
@@ -2,6 +2,7 @@ import {
|
|||||||
canAgentReceiveImplementationTasks,
|
canAgentReceiveImplementationTasks,
|
||||||
classifyWorkflowAgentNode,
|
classifyWorkflowAgentNode,
|
||||||
isEphemeralAgent,
|
isEphemeralAgent,
|
||||||
|
isWorkflowPrincipalEligible,
|
||||||
resolveColumnAgentBinding,
|
resolveColumnAgentBinding,
|
||||||
type Agent,
|
type Agent,
|
||||||
type TaskDetail,
|
type TaskDetail,
|
||||||
@@ -133,13 +134,10 @@ function available(agent: Agent | undefined, activeSessions: ReadonlyMap<string,
|
|||||||
* never satisfy a role-pool route, even when their singular compatibility role
|
* never satisfy a role-pool route, even when their singular compatibility role
|
||||||
* matches. A named transient identity is likewise unavailable and holds closed.
|
* matches. A named transient identity is likewise unavailable and holds closed.
|
||||||
*/
|
*/
|
||||||
if (
|
// FNXC:WorkflowAgentRouting 2026-08-10-01:15: the static half is shared with provisioning
|
||||||
!agent
|
// (isWorkflowPrincipalEligible) so "what provisioning must produce" and "what routing accepts"
|
||||||
|| isEphemeralAgent(agent)
|
// cannot drift apart again — that drift is what left every built-in owner unroutable.
|
||||||
|| agent.runtimeConfig?.enabled === false
|
if (!agent || isEphemeralAgent(agent) || !isWorkflowPrincipalEligible(agent)) return false;
|
||||||
|| agent.state === "paused"
|
|
||||||
|| agent.state === "error"
|
|
||||||
) return false;
|
|
||||||
const max = agent.runtimeConfig?.maxWorkflowSessions;
|
const max = agent.runtimeConfig?.maxWorkflowSessions;
|
||||||
return typeof max !== "number" || activeSessions.get(agent.id) === undefined || activeSessions.get(agent.id)! < max;
|
return typeof max !== "number" || activeSessions.get(agent.id) === undefined || activeSessions.get(agent.id)! < max;
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -625,6 +625,15 @@ ordinary re-dispatch rather than parked.
|
|||||||
const MAX_SESSION_CONTENTION_HOLD_RETRIES = 10;
|
const MAX_SESSION_CONTENTION_HOLD_RETRIES = 10;
|
||||||
const SESSION_CONTENTION_HOLD_BACKOFF_MS = process.env.VITEST || process.env.NODE_ENV === "test" ? 0 : 5_000;
|
const SESSION_CONTENTION_HOLD_BACKOFF_MS = process.env.VITEST || process.env.NODE_ENV === "test" ? 0 : 5_000;
|
||||||
const SESSION_CONTENTION_HOLD_MAX_BACKOFF_MS = 60_000;
|
const SESSION_CONTENTION_HOLD_MAX_BACKOFF_MS = 60_000;
|
||||||
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
Backoff ladder for a workflow-principal hold (unroutable role pool / unavailable named owner). Mirrors the
|
||||||
|
session-contention ladder, including the test-mode zero so suites do not wait on wall-clock. The ceiling is
|
||||||
|
higher than contention's because an unroutable pool clears on OPERATOR action (enable/add an agent), not on
|
||||||
|
another task finishing, so polling it every few seconds only burns CPU.
|
||||||
|
*/
|
||||||
|
const PRINCIPAL_HOLD_BACKOFF_MS = process.env.VITEST || process.env.NODE_ENV === "test" ? 0 : 15_000;
|
||||||
|
const PRINCIPAL_HOLD_MAX_BACKOFF_MS = 300_000;
|
||||||
/** How long to wait before recovering a completed task still stuck in in-progress. */
|
/** How long to wait before recovering a completed task still stuck in in-progress. */
|
||||||
const COMPLETED_TASK_WATCHDOG_MS = 60_000;
|
const COMPLETED_TASK_WATCHDOG_MS = 60_000;
|
||||||
/** How long to wait before retrying a workflow rerun handoff that never reached in-progress. */
|
/** How long to wait before retrying a workflow rerun handoff that never reached in-progress. */
|
||||||
@@ -7850,13 +7859,39 @@ export class TaskExecutor {
|
|||||||
* loudly for the never-clears variant — so the next occurrence is greppable.
|
* loudly for the never-clears variant — so the next occurrence is greppable.
|
||||||
*/
|
*/
|
||||||
const neverClears = principalHoldReason.startsWith("workflow-principal-routing-unavailable:");
|
const neverClears = principalHoldReason.startsWith("workflow-principal-routing-unavailable:");
|
||||||
|
/*
|
||||||
|
* FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
* A hold with no cooldown is a HOT LOOP. Nothing here stops the scheduler re-dispatching the task
|
||||||
|
* immediately, so the run re-enters, re-fences, re-suspends and re-parks the work item — observed at
|
||||||
|
* ~3.5 re-dispatches/second across every task, pinning a core and writing ~19k `workflowWorkItem`
|
||||||
|
* audit rows/hour while ZERO work executed. The hold also never increments `attempt`, so no retry
|
||||||
|
* budget is consumed and no existing guard can ever fire; it spins until an operator intervenes.
|
||||||
|
*
|
||||||
|
* Record a per-task backoff keyed on the hold REASON and let `execute()` skip dispatch while it is
|
||||||
|
* live — the same self-recovering shape as `holdForSessionContention`. A changed reason resets the
|
||||||
|
* ladder (genuinely new information), and the first occurrence still logs immediately so the hold
|
||||||
|
* stays greppable; repeats inside the window are silent so the task log is not flooded either.
|
||||||
|
*/
|
||||||
|
const priorHold = this.principalHoldBackoff.get(task.id);
|
||||||
|
const repeated = priorHold?.reason === principalHoldReason;
|
||||||
|
const attempt = repeated ? priorHold!.attempt + 1 : 1;
|
||||||
|
this.principalHoldBackoff.set(task.id, {
|
||||||
|
reason: principalHoldReason,
|
||||||
|
attempt,
|
||||||
|
until: Date.now() + Math.min(
|
||||||
|
PRINCIPAL_HOLD_MAX_BACKOFF_MS,
|
||||||
|
PRINCIPAL_HOLD_BACKOFF_MS * 2 ** (attempt - 1),
|
||||||
|
),
|
||||||
|
});
|
||||||
const holdMessage = `[workflow-graph] ${task.id} held at graph node — ${principalHoldReason}`;
|
const holdMessage = `[workflow-graph] ${task.id} held at graph node — ${principalHoldReason}`;
|
||||||
if (neverClears) {
|
if (!repeated) {
|
||||||
executorLog.error(`${holdMessage} (workflow principal routing is unavailable; this hold cannot self-clear)`);
|
if (neverClears) {
|
||||||
} else {
|
executorLog.error(`${holdMessage} (workflow principal routing is unavailable; this hold cannot self-clear)`);
|
||||||
executorLog.warn(holdMessage);
|
} else {
|
||||||
|
executorLog.warn(holdMessage);
|
||||||
|
}
|
||||||
|
await this.store.logEntry(task.id, `Workflow stage held — ${principalHoldReason}`).catch(() => undefined);
|
||||||
}
|
}
|
||||||
await this.store.logEntry(task.id, `Workflow stage held — ${principalHoldReason}`).catch(() => undefined);
|
|
||||||
/*
|
/*
|
||||||
* FNXC:WorkflowAgentRouting 2026-08-07-23:50:
|
* FNXC:WorkflowAgentRouting 2026-08-07-23:50:
|
||||||
* The task must end this run with EXACTLY ONE active continuation, and the hold
|
* The task must end this run with EXACTLY ONE active continuation, and the hold
|
||||||
@@ -7892,6 +7927,9 @@ export class TaskExecutor {
|
|||||||
}
|
}
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
|
// FNXC:WorkflowAgentRouting 2026-08-10-01:15: this run cleared the principal fence, so any prior hold
|
||||||
|
// is resolved — drop the ladder so a later hold starts from the short delay rather than a stale one.
|
||||||
|
this.clearPrincipalHoldBackoff(task.id);
|
||||||
/* Direct graph node fences are terminalized only after the interpreter
|
/* Direct graph node fences are terminalized only after the interpreter
|
||||||
* returns, preserving their historical principal through all handler and
|
* returns, preserving their historical principal through all handler and
|
||||||
* tool-gate calls while ensuring completed work cannot render as active.
|
* tool-gate calls while ensuring completed work cannot render as active.
|
||||||
@@ -11487,6 +11525,27 @@ export class TaskExecutor {
|
|||||||
*/
|
*/
|
||||||
private sessionContentionHoldAttempts = new Map<string, number>();
|
private sessionContentionHoldAttempts = new Map<string, number>();
|
||||||
|
|
||||||
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
Per-task cooldown for a workflow-principal hold, so an unroutable role pool is a cheap wait instead of a
|
||||||
|
dispatch hot loop. IN-MEMORY on purpose, matching the session-contention hold: it needs no schema change,
|
||||||
|
and a restart clearing it is correct — a restart is exactly when agent configuration may have changed.
|
||||||
|
*/
|
||||||
|
private principalHoldBackoff = new Map<string, { reason: string; attempt: number; until: number }>();
|
||||||
|
|
||||||
|
/** True while a principal hold is still cooling down, so dispatch should not re-enter the graph. */
|
||||||
|
private isPrincipalHoldCoolingDown(taskId: string): boolean {
|
||||||
|
const hold = this.principalHoldBackoff.get(taskId);
|
||||||
|
if (!hold) return false;
|
||||||
|
if (Date.now() >= hold.until) return false;
|
||||||
|
return true;
|
||||||
|
}
|
||||||
|
|
||||||
|
/** Clear the cooldown once the task dispatches for any other reason. */
|
||||||
|
private clearPrincipalHoldBackoff(taskId: string): void {
|
||||||
|
this.principalHoldBackoff.delete(taskId);
|
||||||
|
}
|
||||||
|
|
||||||
private clearSessionContentionHold(taskId: string): void {
|
private clearSessionContentionHold(taskId: string): void {
|
||||||
this.sessionContentionHoldAttempts.delete(taskId);
|
this.sessionContentionHoldAttempts.delete(taskId);
|
||||||
}
|
}
|
||||||
@@ -14016,6 +14075,18 @@ export class TaskExecutor {
|
|||||||
both pass the graphRouting.has gate, both enter executeWorkflowGraph, and one park
|
both pass the graphRouting.has gate, both enter executeWorkflowGraph, and one park
|
||||||
status=failed while the other still owned work (FN-8471 overseer thrash).
|
status=failed while the other still owned work (FN-8471 overseer thrash).
|
||||||
*/
|
*/
|
||||||
|
/*
|
||||||
|
FNXC:WorkflowAgentRouting 2026-08-10-01:15:
|
||||||
|
Honor an active principal-hold cooldown BEFORE the graph is entered. Without this the hold is recorded and
|
||||||
|
then immediately re-tested by the next dispatch, which is the hot loop itself: re-entering only to re-fence
|
||||||
|
and re-park costs a graph run, two work-item writes, and two audit rows per pass for a condition that can
|
||||||
|
only change when an operator enables or adds an agent. Skipping here is what makes the hold a real wait.
|
||||||
|
*/
|
||||||
|
if (this.isPrincipalHoldCoolingDown(task.id)) {
|
||||||
|
executorLog.debug(`execute() called for ${task.id} while a workflow-principal hold is cooling down — deferring`);
|
||||||
|
if (dropPreHeldExecutorSlot(task.id)) this.options.semaphore?.release();
|
||||||
|
return;
|
||||||
|
}
|
||||||
if (this.graphRouting.has(task.id)) {
|
if (this.graphRouting.has(task.id)) {
|
||||||
// Duplicate dispatch while the graph runner owns this task — drop it,
|
// Duplicate dispatch while the graph runner owns this task — drop it,
|
||||||
// mirroring the executingTaskLock duplicate-invocation behavior.
|
// mirroring the executingTaskLock duplicate-invocation behavior.
|
||||||
|
|||||||
Reference in New Issue
Block a user