Files
fusion/packages/engine/src/usage-limit-detector.ts
gsxdsm a68785a41d P0: two silent triage guards in the executor's ownership — one strands a card with nothing to rescue it (#2572)
P0 audit of the executor's assigned `triage` sites after the
Planning-column merge. **One of them can strand a card**, so leading
with that.

## The stall — `handleDepAbortCleanup`

`executor.ts` moved a dependency-aborted task to the **literal**
`triage`. The default coding lineage no longer declares that column.

A card that gains a dependency mid-execution has its work discarded and
is then parked in a column its own workflow does not define. Nothing in
the graph routes a card out of an undeclared column. The only rescue is
`reconcileUndeclaredTaskColumns`, which runs on the **next engine
start** — so between the abort and a restart the card is stalled with no
automatic recovery. It does not throw, so it would have surfaced as a
user report, not a red test.

Fixed to `resolveReboundColumnFor`, the helper the other ~16 executor
rebounds already use.

## The silent skip — `UsageLimitPauser.taskUsesProvider`

The planning lane was identified by the same literal. For a default card
the lane resolved to **no providers**, so when a provider hit a usage
limit during a *planning* session, the fan-out that pauses peers on that
provider skipped every default-workflow card and they kept hammering the
rate-limited provider.

Not a stall: the triggering task is still paused by the explicit
fallback below the filter. What was lost is blast-radius containment. A
planning session runs while the card is pre-implementation, and the
caller has already excluded `done`/`archived`, so that is exactly "not
the implementation column and not the review column" — which matches
`todo`, `triage`, `ideas`, and a renamed planner alike.

## Full audit table for my assigned sites

| Site | (a) Still fires for a default card? | (b) What silently stops |
(c) Action |
|---|---|---|---|
| `executor.ts:16395` `moveTask(id, "triage")` | **No** — writes an
undeclared column | Card parked where nothing routes it; rescue only at
next engine start | **Fixed** — `resolveReboundColumnFor` |
| `usage-limit-detector.ts:126` `column === "triage"` | **No** |
Usage-limit fan-out skips every default card; peers keep hitting the
limited provider | **Fixed** — pre-implementation predicate |
| `executor.ts:3409` `from === "todo" \|\| from === "triage"` | **Yes**,
via the `todo` arm | — | Unchanged; `triage` arm still live for
legacy-coding |
| `executor.ts:4951` `originColumn === "todo" \|\| === "triage"` |
**Yes**, via the `todo` arm | — | Unchanged |
| `executor.ts:4963` `originColumn === "triage"` double-hop | No, and
correctly so | Nothing — the extra hop exists only for shapes that
declare `triage` | Unchanged; still required by legacy-coding |
| `executor.ts:1110` `Type.Literal("triage")` | n/a | — | **Not a
column** — an agent ROLE in `spawnAgentParams` |

Counts for my ownership: **6 sites audited, 2 defects, 2 fixed, 3
correct as-is, 1 false positive.**

## Red-green

Reverting each fix fails its own test:

```
Tests  2 failed | 2 passed (4)
  × dependency-abort cleanup requeues to a DECLARED column
  × usage-limit fan-out … pauses a peer card sitting in the merged Planning column (id `todo`)
```

The other two are the regression floor and pass both ways by design: a
legacy workflow that **does** declare `triage` still fans out, and an
in-progress card is still **not** swept into the planning lane (the
guard must stay narrow — "any non-wip column" would have been the easy
wrong fix).

## Verification

- New audit suite + graph-boundary + step-session + ownership ledger —
**45 tests green**
- `pnpm test:gate` green (10 / 414 / 71); `pnpm lint` clean; `tsc
--noEmit` clean
- Changeset included (`patch`, `fix`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 19:03:10 -07:00

254 lines
12 KiB
TypeScript

/**
* Usage Limit Detector — classifies API errors as usage-limit-related and
* parks only the task routed through the unavailable provider.
*
* Usage-limit errors indicate provider-local conditions (rate limits, quota
* exceeded, billing issues, overloaded APIs). Continued retrying the affected
* task is wasteful, but unrelated providers must remain available. Transient
* server errors (500, timeout, connection refused) are NOT classified as usage-
* limit errors — they are temporary and may resolve on their own via per-session
* retry.
*/
import type { Task, TaskStore } from "@fusion/core";
import { resolveTaskLifecycleColumns, type WorkflowIr } from "@fusion/core";
import {
resolveExecutorSessionModel,
resolveMergerSessionModel,
resolvePlanningSessionModel,
resolveValidatorSessionModel,
} from "./agent-session-helpers.js";
import { createLogger } from "./logger.js";
const log = createLogger("usage-limit");
/**
* Patterns that indicate API usage/capacity/billing limits.
* These are checked case-insensitively against error messages.
*/
const USAGE_LIMIT_PATTERNS: RegExp[] = [
/overloaded/i,
/rate[_\s]?limit/i,
/too many requests/i,
/\b429\b/,
/\b529\b/,
/quota/i,
/billing/i,
/\bcredit/i,
/insufficient.*(quota|credit|balance|fund)/i,
];
/**
* Classify whether an error message indicates a usage-limit condition.
*
* Returns `true` for rate limits, overloaded errors, and quota/billing issues.
* Returns `false` for transient server errors (500/502/503/504, timeout,
* connection refused) that may resolve on their own.
*/
export function isUsageLimitError(errorMessage: string): boolean {
return USAGE_LIMIT_PATTERNS.some((pattern) => pattern.test(errorMessage));
}
/**
* Lightweight coordinator that agents call when they detect usage-limit errors.
* It parks only the task that reached the unavailable provider. A provider-local
* outage must never activate the project-wide emergency stop because doing so
* also kills healthy Codex/Claude/Grok work routed through other providers.
*/
/**
* Check if an agent session resolved with an error after exhausting retries.
*
* pi-coding-agent's `session.prompt()` does **not** throw when retries are
* exhausted — it resolves normally and stores the error on
* `session.state.errorMessage` (was `session.state.error` prior to
* pi-coding-agent 0.70). Call this immediately after every
* `await session.prompt(...)` to re-raise the swallowed error so existing
* `catch` blocks (with `isUsageLimitError` checks) can detect rate-limit
* conditions and trigger `UsageLimitPauser`.
*
* @param session — The agent session (or any object with `state.errorMessage?: string`)
* @throws {Error} If `session.state.errorMessage` is set and non-empty
*/
export function checkSessionError(session: { state: { errorMessage?: string; error?: string } }): void {
const state = session.state;
const error = state?.errorMessage ?? state?.error;
if (error) {
throw new Error(error);
}
}
export class UsageLimitPauser {
constructor(private store: TaskStore) {}
private normalizeProviderId(provider: string): string {
return provider.trim().toLowerCase().replace(/[^a-z0-9._-]+/g, "-").replace(/^-+|-+$/g, "");
}
/**
* Clear only parks created for a provider whose independent health probe has
* transitioned back to usable. Manual/user pauses and every other provider
* reason remain untouched.
*/
async onProviderAvailable(provider: string): Promise<number> {
const providerId = this.normalizeProviderId(provider);
if (!providerId) return 0;
const pausedReason = `provider-rate-limit:${providerId}`;
// FNXC:ArchitectureHotPath 2026-07-22-17:20: listTasks() must be explicit about payload shape (architecture-hot-paths contract). Recovery only reads scalar pause fields, so request slim rows to avoid loading heavy log/steps/comments for every task.
const tasks = await this.store.listTasks({ slim: true });
const recoverableTasks = tasks.filter((task) =>
task.paused === true
&& task.userPaused !== true
&& (task.pausedReason === pausedReason
|| (providerId === "unknown" && task.pausedReason === "provider-rate-limit")));
/*
FNXC:ProviderRateLimitRecovery 2026-07-19-20:15:
Provider recovery is a health-state transition, never a task call used as a probe. The daemon's independent authenticated usage/capacity monitor invokes this seam only after positive health, and this exact-reason filter ensures recovery cannot clear manual parks, unrelated failure reasons, or another provider's outage.
*/
await Promise.all(recoverableTasks.map(async (task) => {
await this.store.logEntry(task.id, `Provider ${providerId} is available again; resuming task`);
await this.store.pauseTask(task.id, false);
}));
if (recoverableTasks.length > 0) {
log.log(`Provider ${providerId} recovered; resumed ${recoverableTasks.length} task(s)`);
}
return recoverableTasks.length;
}
private taskUsesProvider(
task: Task,
provider: string,
settings: Awaited<ReturnType<TaskStore["getSettings"]>>,
agentType: string,
preImplementationColumns?: ReadonlySet<string>,
): boolean {
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-15:20 (P0 audit after the Planning-column merge):
The planning lane was identified by the LITERAL `triage`. The default coding lineage no
longer declares that column, so this comparison stopped matching for every default-workflow
card — silently. Nothing throws; the lane simply resolves to no providers, so when a provider
hits a usage limit during a PLANNING session the fan-out that pauses other tasks on that same
provider skips every default card, and they keep hammering the rate-limited provider. The
triggering task is still paused by the explicit fallback below, so no card is stranded — what
is lost is the blast-radius containment.
A planning session runs while the card is PRE-IMPLEMENTATION. The caller has already excluded
`done`/`archived`, so that is exactly "not the implementation column and not the review
column" — which matches `todo`, `triage`, `ideas`, and a renamed planner alike. The two
literals that remain here name the wip and review lanes and belong to the executor/scheduler
vocabulary conversion, not to this fix.
*/
const isPreImplementation = preImplementationColumns?.has(task.column) === true;
const providersByActiveLane = agentType === "triage"
? (isPreImplementation ? [
resolvePlanningSessionModel(task.planningModelProvider, task.planningModelId, settings).provider,
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
] : [])
: agentType === "executor"
? (task.column === "in-progress" ? [
resolveExecutorSessionModel(task.modelProvider, task.modelId, settings).provider,
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
] : [])
: agentType === "merger"
? (task.column === "in-review" ? [resolveMergerSessionModel(settings, undefined, task).provider] : [])
: [];
const resolvedProviders = providersByActiveLane;
return resolvedProviders.some((candidate) => candidate?.trim().toLowerCase() === provider);
}
/**
* Called by agents when a usage-limit error is detected after retries are exhausted.
* Parks the affected task while leaving every other provider lane running.
*
* @param agentType - The type of agent that hit the limit (e.g., "executor", "triage", "merger")
* @param taskId - The task that was being processed when the limit was hit
* @param errorMessage - The error message from the API
* @param provider - Best-effort provider identifier used in the pause reason
*/
async onUsageLimitHit(agentType: string, taskId: string, errorMessage: string, provider?: string): Promise<void> {
const providerId = this.normalizeProviderId(provider ?? "unknown") || "unknown";
const pausedReason = `provider-rate-limit:${providerId}`;
/*
FNXC:ProviderRateLimitIsolation 2026-07-19-19:10:
A 429 is provider-local, not a project emergency. Park only the task that exhausted retries on that provider so healthy provider lanes continue executing. Keep the provider id in structured pause provenance when the caller can identify it; never persist the full provider response as pause metadata.
*/
log.warn(`${agentType} hit usage limit${providerId ? ` for ${providerId}` : ""} on ${taskId}: ${errorMessage}`);
log.warn(`Matched pattern in error: "${errorMessage.slice(0, 200)}"`);
// Log the triggering error on the task
await this.store.logEntry(
taskId,
`Usage limit detected (${agentType}${providerId ? `/${providerId}` : ""}): ${errorMessage}`,
);
const [settings, tasks] = await Promise.all([
this.store.getSettings(),
// FNXC:ArchitectureHotPath 2026-07-22-17:20: slim payload — this scan only reads column/pause/model-provider scalars, never heavy detail fields.
this.store.listTasks({ slim: true }),
]);
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-20:50 (P0 audit, PR #2572 review — greptile):
The planning lane is resolved PER TASK from its own workflow, not inferred by excluding two
literals. "Not `in-progress` and not `in-review`" reads any custom non-terminal column — a
second processing lane, a manual hold, a bespoke review stage — as pre-implementation, so a
planning-provider limit would pause cards that are nowhere near planning. Trait-derived
intake/hold is the only answer that holds for a workflow this code has never seen.
One IR read per WORKFLOW, not per task: the cache is caller-owned (the U1 contract) and shared
across the whole fan-out, so a 400-card board spanning three workflows reads three IRs. A task
whose workflow cannot be resolved yields an empty set and is skipped rather than guessed into
the lane — conservative, because the cost of a wrong include is pausing work that was fine.
*/
const irCache = new Map<string, WorkflowIr>();
const preImplementationByTask = new Map<string, ReadonlySet<string>>();
if (agentType === "triage") {
await Promise.all(tasks.map(async (task) => {
const columns = await resolveTaskLifecycleColumns(this.store, task.id, irCache).catch(() => undefined);
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-22:10 (PR #2572 review — greptile, 2nd):
INTAKE ONLY. `hold` is not a synonym for "planning": a workflow may carry a hold trait on
a MID-PIPELINE wait — manual release, timed, dependency, external event — and a card
parked there is downstream of implementation, not queued for planning. Including hold
would pause it on a planning-provider limit, which is the same over-classification as the
literal-exclusion predicate this replaced, just further along.
The planning session is the one that runs on an intake card, so intake is the lane. When a
workflow's hold column IS its pre-implementation queue it is normally the same column as
intake (the merged Planning lane declares both traits) and is covered by that; where they
differ, the hold column is a wait and is deliberately excluded.
*/
const lanes = new Set<string>();
if (columns?.intake) lanes.add(columns.intake);
preImplementationByTask.set(task.id, lanes);
}));
}
const affectedTasks = tasks.filter((task) =>
task.column !== "done"
&& task.column !== "archived"
&& task.paused !== true
&& providerId !== "unknown"
&& this.taskUsesProvider(task, providerId, settings, agentType, preImplementationByTask.get(task.id)));
// Always include the task that produced the 429 even if its actual provider
// came from a runtime fallback not represented in persisted task settings.
if (!affectedTasks.some((task) => task.id === taskId)) {
const triggeringTask = await this.store.getTask(taskId).catch(() => null);
if (triggeringTask && triggeringTask.paused !== true) affectedTasks.push(triggeringTask);
}
await Promise.all(affectedTasks.map(async (task) => {
if (task.id !== taskId) {
await this.store.logEntry(
task.id,
`Paused because provider ${providerId} reached a usage limit on ${taskId}`,
);
}
await this.store.pauseTask(task.id, true, undefined, { pausedReason });
}));
log.warn(`Paused ${affectedTasks.length} task(s) routed through ${providerId}; other provider lanes remain active`);
}
}