Files
fusion/packages/engine/src/usage-limit-detector.ts
gsxdsm 31e49b684a TAKING default-workflow-hooks.ts + executor.ts + live-agent-count.ts + 6 dashboard files: reopen semantics by role, and the census's blind spot in both directions (13 sites) (#2628)
Batched conversion of every lifecycle-column guard I hold, plus the
three the census could not see. **Six files to zero, repo-wide 60 → 49
by a comment-stripped unanchored sweep.** Each conversion has an
isolated revert proof and a paired negative case, and the one code move
is a separate commit from the behavior changes.

## Per-file before → after

Counts from a comment-stripped, unanchored `(===|!==) ["']triage["']`
sweep over `packages/*/src` + `plugins/*/src`, excluding tests.

| file | before | after | note |
|---|---:|---:|---|
| `core/default-workflow-hooks.ts` | 4 | **0** | |
| `core/task-store/moves.ts` | 5 | **4** | only the flag-ON mirror
converted; the flag-OFF inline block is the parity reference and stays |
| `engine/executor.ts` | 3 | **0** | **absent from the 45-guard list** —
see below |
| `core/live-agent-count.ts` | 2 | **0** | duplication removed; answer
deliberately unchanged |
| `engine/replan-target.ts` | 2 | **0** | both were comment prose, not
guards |
| `core/agent-prompts.ts` | 3 | **0** | ROLE comparisons, never column
guards |
| `engine/usage-limit-detector.ts` | 2 | **0** | ROLE comparisons |
| `dashboard/app/components/DocumentsView.tsx` | 1 | **0** | real column
guard |
| `dashboard/app/components/TaskChatTab.tsx` | 2 | **0** | ROLE |
| `dashboard/app/components/AgentLogViewer.tsx` | 1 | **0** | ROLE |
| `dashboard/app/components/effective-model-resolution.ts` | 1 | **0** |
ROLE |
| `dashboard/app/hooks/useTasks.ts` | 1 | **0** | ROLE |
| `dashboard/…/command-center/MissionControlPanel.tsx` | 1 | 1 | alias
table, marked `DELIBERATE-LITERAL` with its reason |

## The census errs in BOTH directions

This is the finding I would most like carried into the remaining work.

- It **flagged 10 sites that were never column guards.** `role ===
"triage"` / `agentType === "triage"` compare an **AGENT ROLE**. The
planner *lane* is named `triage` and keeps that name — U11 removed the
*column*. Worse than noise: the obvious "finish the migration" edit is
to rename the role, and that silently empties the planner's prompt
template and mis-binds its model markers. `PLANNER_AGENT_ROLE` now names
it, so the two vocabularies are distinguishable by grep and a rename
fails loudly (revert proof: 4 tests, two of them pre-existing).
- It **missed 3 real guards in `executor.ts`**, because the pattern
matches `column`/`toColumn`/`fromColumn` and those locals are named
`from` and `originColumn`. A census keyed on variable names will keep
missing guards wherever a local was named for its role in the function.

## Two real defects, not tidying

**1. A renamed board could merge with its re-review never run.**
`default-workflow-hooks.ts` is named for the default workflow, but the
store runs it on the flag-ON path for *every* workflow — the trait
registry resolves hooks by trait id, not by workflow. Its reopen
predicates listed the default lineage's column names, so on a renamed
board **no reopen effect fired at all**. One of them clears
`workflowStepResults`, which `getTaskMergeBlocker` reads: a card bounced
out of review carried its old `passed` result back in, and that
satisfies the merge gate. Same regression the graph-owned-crossing
carve-out exists to prevent, arriving through the other door. (Two
smaller ones rode along: failure state never cleared on a renamed
reopen, and an operator dragging a card back to the queue never parked
it, so the scheduler re-dispatched what they had just pulled back.)

**I forgot the carve-out on my first pass, and that was worse than not
converting.** A role-resolved clear plus a *name*-matched exemption
means a renamed board takes the clear and never the exemption,
destroying the remediation input the graph had just written. My own
paired negative test caught it.

**2. The last-resort recovery for completed-but-stranded work did not
exist off the default lineage.** In `recoverCompletedTask`,
`promotedFromPlannerColumn` was false on a renamed board, so finished
work resting in the planning lane was never promoted — the code fell
through to `handoffTaskToReview` straight from the planning column, and
role adjacency has no planning → review edge, so the handoff was
rejected and the card stayed stuck with its work complete. I converted
the promotion **target** too: resolving the lane and then moving to a
literal `in-progress` is the half-conversion I have already been burned
by twice this program, where the guard starts admitting cards and the
move then sends them to a column the board does not declare.

## E2E evidence

`renamed-board-reopen.pg.test.ts` drives a **real PostgreSQL store** and
a real `moveTask` on a workflow whose columns carry the standard traits
under non-default names. The unit tests cannot show this: if `moves.ts`
passed `undefined`, every unit case still passes via the no-basis
fallback while the real board keeps the old behavior. **Proof it is
load-bearing: forcing `moveLifecycleColumns` to `undefined` fails 2 of
3.** The executor suite covers both the split-role and the MERGED
post-U11 shape.

## Revert proofs, isolated per site

| change reverted | result |
|---|---|
| reopen predicate → literal names | 4 of 10 fail |
| reopen field clears → literal names | 2 of 10 fail |
| `userPaused` hold lane → literal `todo` | 1 of 10 fail |
| graph carve-out → literal names | 1 of 10 fail |
| store passes `undefined` lifecycle columns | 2 of 3 fail (real PG) |
| `promotedFromPlannerColumn` → literals | 3 of 7 fail |
| two-hop condition → `=== "triage"` | 1 of 7 fails |
| promotion target → `"in-progress"` | 3 of 7 fail |
| `isPlannerColumnFor` → literals | 1 of 7 fails |
| live-agent-count: one arm dropped | 2 of 11 fail |
| DocumentsView: trait branch removed | 3 of 7 fail |
| planner role renamed to `"planner"` | 4 fail (2 pre-existing) |

Every conversion is paired with a negative case (a forward move, a
not-a-planner-lane card, a default-lineage card, a renamed column with
no traits), so neither "always fire" nor "never fire" can pass for
"resolve the role".

## Deliberately NOT converted, with reasons

- **`moves.ts` flag-OFF inline block (4).** That branch *is* the legacy
path, kept verbatim so the two can be parity-checked. Converting it
erases the reference implementation.
- **`live-agent-count.ts`'s no-flags fallback.** Reachable, and there is
nothing to resolve from — `enrich…FromFlags` exists for callers with
board flags rather than an IR, so a column missing from that map is the
renamed case. "Not intake" is as much a guess as "todo is intake", and
Running/Waiting are complements, so a card matching neither arm is
reported as neither and the footer's queued total under-reports it. The
real fix is at the caller; four new cases pin that flags override the
legacy answer **in both directions**. What did change is the
duplication: two hand-written copies of one rule now call one named
function.
- **`MissionControlPanel`'s `FUNNEL_STAGES`.** An alias table of column
*names* where `triage` sits beside `signal` and `backlog`. Command
Center aggregates across projects, so there is no single workflow to
resolve traits from — the honest conversion is a data change, not a
predicate change.
- **`DocumentsView` with no traits.** Same no-basis rule; the documents
list is full of historical columns absent from the current board. A case
asserts a renamed column with no traits still reads as "working",
documenting the gap rather than hiding it.

## Fixture findings

Each cost a red run that looked like the code under test:

- a `merge-blocker` column needs a reachable merge-class node, or
`parseWorkflowIr` rejects the workflow;
- a back-edge must be `kind: "rework"`, and a rework edge is legal only
**into** a node with `config.reworkRegion: true`;
- a workflow gets role-level transitions only when it declares wip +
review + complete + **archived** plus a planning lane — without the
archived column, adjacency falls back to order-derived neighbours and
`checking -> queued` is not a legal move at all;
- `recoverCompletedTask` only *reaches* the promotion seam when nothing
is left to gate; without passed `plan-review`/`code-review` rows it
re-enters the workflow graph and returns first, so a naive fixture
silently tests the wrong branch and every assertion reads "no moves
happened" for an unrelated reason.

## Verification

- `pnpm test:gate` **71/71**
- new suites: 10/10 reopen-semantics, 3/3 renamed-board-reopen (real
PG), 7/7 executor-planner-lanes, 7/7 documents-status-dot, 4/4
planner-role-is-not-a-column
- neighbours: 132 + 10 + 482 (gate shards), 350/351 engine
planning/replan suites, 64/64 agent-prompts, 51/51 usage-limit-detector,
11/11 live-agent-count, 11/11 dashboard hook/log suites
- the single engine failure (`executor-fast-mode-workflows.test.ts` ›
"raw fast mode still invokes non-executable review seam nodes")
**reproduces with my changes stashed** — pre-existing on `origin/main`
- typechecks clean for core, engine, and dashboard-app
(`tsconfig.app.json`; `tsconfig.json` checks nothing under `app/`);
`pnpm lint` clean

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 22:39:14 -07:00

256 lines
13 KiB
TypeScript

/**
* Usage Limit Detector — classifies API errors as usage-limit-related and
* parks only the task routed through the unavailable provider.
*
* Usage-limit errors indicate provider-local conditions (rate limits, quota
* exceeded, billing issues, overloaded APIs). Continued retrying the affected
* task is wasteful, but unrelated providers must remain available. Transient
* server errors (500, timeout, connection refused) are NOT classified as usage-
* limit errors — they are temporary and may resolve on their own via per-session
* retry.
*/
import type { Task, TaskStore } from "@fusion/core";
// FNXC:WorkflowLifecycleColumns 2026-07-30-11:00: `agentType` is an AGENT ROLE, not a column.
// The planner lane is named `triage` and keeps that name; only the COLUMN was removed by U11.
import { PLANNER_AGENT_ROLE, resolveTaskLifecycleColumns, type WorkflowIr } from "@fusion/core";
import {
resolveExecutorSessionModel,
resolveMergerSessionModel,
resolvePlanningSessionModel,
resolveValidatorSessionModel,
} from "./agent-session-helpers.js";
import { createLogger } from "./logger.js";
const log = createLogger("usage-limit");
/**
* Patterns that indicate API usage/capacity/billing limits.
* These are checked case-insensitively against error messages.
*/
const USAGE_LIMIT_PATTERNS: RegExp[] = [
/overloaded/i,
/rate[_\s]?limit/i,
/too many requests/i,
/\b429\b/,
/\b529\b/,
/quota/i,
/billing/i,
/\bcredit/i,
/insufficient.*(quota|credit|balance|fund)/i,
];
/**
* Classify whether an error message indicates a usage-limit condition.
*
* Returns `true` for rate limits, overloaded errors, and quota/billing issues.
* Returns `false` for transient server errors (500/502/503/504, timeout,
* connection refused) that may resolve on their own.
*/
export function isUsageLimitError(errorMessage: string): boolean {
return USAGE_LIMIT_PATTERNS.some((pattern) => pattern.test(errorMessage));
}
/**
* Lightweight coordinator that agents call when they detect usage-limit errors.
* It parks only the task that reached the unavailable provider. A provider-local
* outage must never activate the project-wide emergency stop because doing so
* also kills healthy Codex/Claude/Grok work routed through other providers.
*/
/**
* Check if an agent session resolved with an error after exhausting retries.
*
* pi-coding-agent's `session.prompt()` does **not** throw when retries are
* exhausted — it resolves normally and stores the error on
* `session.state.errorMessage` (was `session.state.error` prior to
* pi-coding-agent 0.70). Call this immediately after every
* `await session.prompt(...)` to re-raise the swallowed error so existing
* `catch` blocks (with `isUsageLimitError` checks) can detect rate-limit
* conditions and trigger `UsageLimitPauser`.
*
* @param session — The agent session (or any object with `state.errorMessage?: string`)
* @throws {Error} If `session.state.errorMessage` is set and non-empty
*/
export function checkSessionError(session: { state: { errorMessage?: string; error?: string } }): void {
const state = session.state;
const error = state?.errorMessage ?? state?.error;
if (error) {
throw new Error(error);
}
}
export class UsageLimitPauser {
constructor(private store: TaskStore) {}
private normalizeProviderId(provider: string): string {
return provider.trim().toLowerCase().replace(/[^a-z0-9._-]+/g, "-").replace(/^-+|-+$/g, "");
}
/**
* Clear only parks created for a provider whose independent health probe has
* transitioned back to usable. Manual/user pauses and every other provider
* reason remain untouched.
*/
async onProviderAvailable(provider: string): Promise<number> {
const providerId = this.normalizeProviderId(provider);
if (!providerId) return 0;
const pausedReason = `provider-rate-limit:${providerId}`;
// FNXC:ArchitectureHotPath 2026-07-22-17:20: listTasks() must be explicit about payload shape (architecture-hot-paths contract). Recovery only reads scalar pause fields, so request slim rows to avoid loading heavy log/steps/comments for every task.
const tasks = await this.store.listTasks({ slim: true });
const recoverableTasks = tasks.filter((task) =>
task.paused === true
&& task.userPaused !== true
&& (task.pausedReason === pausedReason
|| (providerId === "unknown" && task.pausedReason === "provider-rate-limit")));
/*
FNXC:ProviderRateLimitRecovery 2026-07-19-20:15:
Provider recovery is a health-state transition, never a task call used as a probe. The daemon's independent authenticated usage/capacity monitor invokes this seam only after positive health, and this exact-reason filter ensures recovery cannot clear manual parks, unrelated failure reasons, or another provider's outage.
*/
await Promise.all(recoverableTasks.map(async (task) => {
await this.store.logEntry(task.id, `Provider ${providerId} is available again; resuming task`);
await this.store.pauseTask(task.id, false);
}));
if (recoverableTasks.length > 0) {
log.log(`Provider ${providerId} recovered; resumed ${recoverableTasks.length} task(s)`);
}
return recoverableTasks.length;
}
private taskUsesProvider(
task: Task,
provider: string,
settings: Awaited<ReturnType<TaskStore["getSettings"]>>,
agentType: string,
preImplementationColumns?: ReadonlySet<string>,
): boolean {
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-15:20 (P0 audit after the Planning-column merge):
The planning lane was identified by the LITERAL `triage`. The default coding lineage no
longer declares that column, so this comparison stopped matching for every default-workflow
card — silently. Nothing throws; the lane simply resolves to no providers, so when a provider
hits a usage limit during a PLANNING session the fan-out that pauses other tasks on that same
provider skips every default card, and they keep hammering the rate-limited provider. The
triggering task is still paused by the explicit fallback below, so no card is stranded — what
is lost is the blast-radius containment.
A planning session runs while the card is PRE-IMPLEMENTATION. The caller has already excluded
`done`/`archived`, so that is exactly "not the implementation column and not the review
column" — which matches `todo`, `triage`, `ideas`, and a renamed planner alike. The two
literals that remain here name the wip and review lanes and belong to the executor/scheduler
vocabulary conversion, not to this fix.
*/
const isPreImplementation = preImplementationColumns?.has(task.column) === true;
const providersByActiveLane = agentType === PLANNER_AGENT_ROLE
? (isPreImplementation ? [
resolvePlanningSessionModel(task.planningModelProvider, task.planningModelId, settings).provider,
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
] : [])
: agentType === "executor"
? (task.column === "in-progress" ? [
resolveExecutorSessionModel(task.modelProvider, task.modelId, settings).provider,
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
] : [])
: agentType === "merger"
? (task.column === "in-review" ? [resolveMergerSessionModel(settings, undefined, task).provider] : [])
: [];
const resolvedProviders = providersByActiveLane;
return resolvedProviders.some((candidate) => candidate?.trim().toLowerCase() === provider);
}
/**
* Called by agents when a usage-limit error is detected after retries are exhausted.
* Parks the affected task while leaving every other provider lane running.
*
* @param agentType - The type of agent that hit the limit (e.g., "executor", "triage", "merger")
* @param taskId - The task that was being processed when the limit was hit
* @param errorMessage - The error message from the API
* @param provider - Best-effort provider identifier used in the pause reason
*/
async onUsageLimitHit(agentType: string, taskId: string, errorMessage: string, provider?: string): Promise<void> {
const providerId = this.normalizeProviderId(provider ?? "unknown") || "unknown";
const pausedReason = `provider-rate-limit:${providerId}`;
/*
FNXC:ProviderRateLimitIsolation 2026-07-19-19:10:
A 429 is provider-local, not a project emergency. Park only the task that exhausted retries on that provider so healthy provider lanes continue executing. Keep the provider id in structured pause provenance when the caller can identify it; never persist the full provider response as pause metadata.
*/
log.warn(`${agentType} hit usage limit${providerId ? ` for ${providerId}` : ""} on ${taskId}: ${errorMessage}`);
log.warn(`Matched pattern in error: "${errorMessage.slice(0, 200)}"`);
// Log the triggering error on the task
await this.store.logEntry(
taskId,
`Usage limit detected (${agentType}${providerId ? `/${providerId}` : ""}): ${errorMessage}`,
);
const [settings, tasks] = await Promise.all([
this.store.getSettings(),
// FNXC:ArchitectureHotPath 2026-07-22-17:20: slim payload — this scan only reads column/pause/model-provider scalars, never heavy detail fields.
this.store.listTasks({ slim: true }),
]);
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-20:50 (P0 audit, PR #2572 review — greptile):
The planning lane is resolved PER TASK from its own workflow, not inferred by excluding two
literals. "Not `in-progress` and not `in-review`" reads any custom non-terminal column — a
second processing lane, a manual hold, a bespoke review stage — as pre-implementation, so a
planning-provider limit would pause cards that are nowhere near planning. Trait-derived
intake/hold is the only answer that holds for a workflow this code has never seen.
One IR read per WORKFLOW, not per task: the cache is caller-owned (the U1 contract) and shared
across the whole fan-out, so a 400-card board spanning three workflows reads three IRs. A task
whose workflow cannot be resolved yields an empty set and is skipped rather than guessed into
the lane — conservative, because the cost of a wrong include is pausing work that was fine.
*/
const irCache = new Map<string, WorkflowIr>();
const preImplementationByTask = new Map<string, ReadonlySet<string>>();
if (agentType === PLANNER_AGENT_ROLE) {
await Promise.all(tasks.map(async (task) => {
const columns = await resolveTaskLifecycleColumns(this.store, task.id, irCache).catch(() => undefined);
/*
FNXC:WorkflowLifecycleColumns 2026-07-29-22:10 (PR #2572 review — greptile, 2nd):
INTAKE ONLY. `hold` is not a synonym for "planning": a workflow may carry a hold trait on
a MID-PIPELINE wait — manual release, timed, dependency, external event — and a card
parked there is downstream of implementation, not queued for planning. Including hold
would pause it on a planning-provider limit, which is the same over-classification as the
literal-exclusion predicate this replaced, just further along.
The planning session is the one that runs on an intake card, so intake is the lane. When a
workflow's hold column IS its pre-implementation queue it is normally the same column as
intake (the merged Planning lane declares both traits) and is covered by that; where they
differ, the hold column is a wait and is deliberately excluded.
*/
const lanes = new Set<string>();
if (columns?.intake) lanes.add(columns.intake);
preImplementationByTask.set(task.id, lanes);
}));
}
const affectedTasks = tasks.filter((task) =>
task.column !== "done"
&& task.column !== "archived"
&& task.paused !== true
&& providerId !== "unknown"
&& this.taskUsesProvider(task, providerId, settings, agentType, preImplementationByTask.get(task.id)));
// Always include the task that produced the 429 even if its actual provider
// came from a runtime fallback not represented in persisted task settings.
if (!affectedTasks.some((task) => task.id === taskId)) {
const triggeringTask = await this.store.getTask(taskId).catch(() => null);
if (triggeringTask && triggeringTask.paused !== true) affectedTasks.push(triggeringTask);
}
await Promise.all(affectedTasks.map(async (task) => {
if (task.id !== taskId) {
await this.store.logEntry(
task.id,
`Paused because provider ${providerId} reached a usage limit on ${taskId}`,
);
}
await this.store.pauseTask(task.id, true, undefined, { pausedReason });
}));
log.warn(`Paused ${affectedTasks.length} task(s) routed through ${providerId}; other provider lanes remain active`);
}
}