Batched conversion of every lifecycle-column guard I hold, plus the three the census could not see. **Six files to zero, repo-wide 60 → 49 by a comment-stripped unanchored sweep.** Each conversion has an isolated revert proof and a paired negative case, and the one code move is a separate commit from the behavior changes. ## Per-file before → after Counts from a comment-stripped, unanchored `(===|!==) ["']triage["']` sweep over `packages/*/src` + `plugins/*/src`, excluding tests. | file | before | after | note | |---|---:|---:|---| | `core/default-workflow-hooks.ts` | 4 | **0** | | | `core/task-store/moves.ts` | 5 | **4** | only the flag-ON mirror converted; the flag-OFF inline block is the parity reference and stays | | `engine/executor.ts` | 3 | **0** | **absent from the 45-guard list** — see below | | `core/live-agent-count.ts` | 2 | **0** | duplication removed; answer deliberately unchanged | | `engine/replan-target.ts` | 2 | **0** | both were comment prose, not guards | | `core/agent-prompts.ts` | 3 | **0** | ROLE comparisons, never column guards | | `engine/usage-limit-detector.ts` | 2 | **0** | ROLE comparisons | | `dashboard/app/components/DocumentsView.tsx` | 1 | **0** | real column guard | | `dashboard/app/components/TaskChatTab.tsx` | 2 | **0** | ROLE | | `dashboard/app/components/AgentLogViewer.tsx` | 1 | **0** | ROLE | | `dashboard/app/components/effective-model-resolution.ts` | 1 | **0** | ROLE | | `dashboard/app/hooks/useTasks.ts` | 1 | **0** | ROLE | | `dashboard/…/command-center/MissionControlPanel.tsx` | 1 | 1 | alias table, marked `DELIBERATE-LITERAL` with its reason | ## The census errs in BOTH directions This is the finding I would most like carried into the remaining work. - It **flagged 10 sites that were never column guards.** `role === "triage"` / `agentType === "triage"` compare an **AGENT ROLE**. The planner *lane* is named `triage` and keeps that name — U11 removed the *column*. Worse than noise: the obvious "finish the migration" edit is to rename the role, and that silently empties the planner's prompt template and mis-binds its model markers. `PLANNER_AGENT_ROLE` now names it, so the two vocabularies are distinguishable by grep and a rename fails loudly (revert proof: 4 tests, two of them pre-existing). - It **missed 3 real guards in `executor.ts`**, because the pattern matches `column`/`toColumn`/`fromColumn` and those locals are named `from` and `originColumn`. A census keyed on variable names will keep missing guards wherever a local was named for its role in the function. ## Two real defects, not tidying **1. A renamed board could merge with its re-review never run.** `default-workflow-hooks.ts` is named for the default workflow, but the store runs it on the flag-ON path for *every* workflow — the trait registry resolves hooks by trait id, not by workflow. Its reopen predicates listed the default lineage's column names, so on a renamed board **no reopen effect fired at all**. One of them clears `workflowStepResults`, which `getTaskMergeBlocker` reads: a card bounced out of review carried its old `passed` result back in, and that satisfies the merge gate. Same regression the graph-owned-crossing carve-out exists to prevent, arriving through the other door. (Two smaller ones rode along: failure state never cleared on a renamed reopen, and an operator dragging a card back to the queue never parked it, so the scheduler re-dispatched what they had just pulled back.) **I forgot the carve-out on my first pass, and that was worse than not converting.** A role-resolved clear plus a *name*-matched exemption means a renamed board takes the clear and never the exemption, destroying the remediation input the graph had just written. My own paired negative test caught it. **2. The last-resort recovery for completed-but-stranded work did not exist off the default lineage.** In `recoverCompletedTask`, `promotedFromPlannerColumn` was false on a renamed board, so finished work resting in the planning lane was never promoted — the code fell through to `handoffTaskToReview` straight from the planning column, and role adjacency has no planning → review edge, so the handoff was rejected and the card stayed stuck with its work complete. I converted the promotion **target** too: resolving the lane and then moving to a literal `in-progress` is the half-conversion I have already been burned by twice this program, where the guard starts admitting cards and the move then sends them to a column the board does not declare. ## E2E evidence `renamed-board-reopen.pg.test.ts` drives a **real PostgreSQL store** and a real `moveTask` on a workflow whose columns carry the standard traits under non-default names. The unit tests cannot show this: if `moves.ts` passed `undefined`, every unit case still passes via the no-basis fallback while the real board keeps the old behavior. **Proof it is load-bearing: forcing `moveLifecycleColumns` to `undefined` fails 2 of 3.** The executor suite covers both the split-role and the MERGED post-U11 shape. ## Revert proofs, isolated per site | change reverted | result | |---|---| | reopen predicate → literal names | 4 of 10 fail | | reopen field clears → literal names | 2 of 10 fail | | `userPaused` hold lane → literal `todo` | 1 of 10 fail | | graph carve-out → literal names | 1 of 10 fail | | store passes `undefined` lifecycle columns | 2 of 3 fail (real PG) | | `promotedFromPlannerColumn` → literals | 3 of 7 fail | | two-hop condition → `=== "triage"` | 1 of 7 fails | | promotion target → `"in-progress"` | 3 of 7 fail | | `isPlannerColumnFor` → literals | 1 of 7 fails | | live-agent-count: one arm dropped | 2 of 11 fail | | DocumentsView: trait branch removed | 3 of 7 fail | | planner role renamed to `"planner"` | 4 fail (2 pre-existing) | Every conversion is paired with a negative case (a forward move, a not-a-planner-lane card, a default-lineage card, a renamed column with no traits), so neither "always fire" nor "never fire" can pass for "resolve the role". ## Deliberately NOT converted, with reasons - **`moves.ts` flag-OFF inline block (4).** That branch *is* the legacy path, kept verbatim so the two can be parity-checked. Converting it erases the reference implementation. - **`live-agent-count.ts`'s no-flags fallback.** Reachable, and there is nothing to resolve from — `enrich…FromFlags` exists for callers with board flags rather than an IR, so a column missing from that map is the renamed case. "Not intake" is as much a guess as "todo is intake", and Running/Waiting are complements, so a card matching neither arm is reported as neither and the footer's queued total under-reports it. The real fix is at the caller; four new cases pin that flags override the legacy answer **in both directions**. What did change is the duplication: two hand-written copies of one rule now call one named function. - **`MissionControlPanel`'s `FUNNEL_STAGES`.** An alias table of column *names* where `triage` sits beside `signal` and `backlog`. Command Center aggregates across projects, so there is no single workflow to resolve traits from — the honest conversion is a data change, not a predicate change. - **`DocumentsView` with no traits.** Same no-basis rule; the documents list is full of historical columns absent from the current board. A case asserts a renamed column with no traits still reads as "working", documenting the gap rather than hiding it. ## Fixture findings Each cost a red run that looked like the code under test: - a `merge-blocker` column needs a reachable merge-class node, or `parseWorkflowIr` rejects the workflow; - a back-edge must be `kind: "rework"`, and a rework edge is legal only **into** a node with `config.reworkRegion: true`; - a workflow gets role-level transitions only when it declares wip + review + complete + **archived** plus a planning lane — without the archived column, adjacency falls back to order-derived neighbours and `checking -> queued` is not a legal move at all; - `recoverCompletedTask` only *reaches* the promotion seam when nothing is left to gate; without passed `plan-review`/`code-review` rows it re-enters the workflow graph and returns first, so a naive fixture silently tests the wrong branch and every assertion reads "no moves happened" for an unrelated reason. ## Verification - `pnpm test:gate` **71/71** - new suites: 10/10 reopen-semantics, 3/3 renamed-board-reopen (real PG), 7/7 executor-planner-lanes, 7/7 documents-status-dot, 4/4 planner-role-is-not-a-column - neighbours: 132 + 10 + 482 (gate shards), 350/351 engine planning/replan suites, 64/64 agent-prompts, 51/51 usage-limit-detector, 11/11 live-agent-count, 11/11 dashboard hook/log suites - the single engine failure (`executor-fast-mode-workflows.test.ts` › "raw fast mode still invokes non-executable review seam nodes") **reproduces with my changes stashed** — pre-existing on `origin/main` - typechecks clean for core, engine, and dashboard-app (`tsconfig.app.json`; `tsconfig.json` checks nothing under `app/`); `pnpm lint` clean 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
256 lines
13 KiB
TypeScript
256 lines
13 KiB
TypeScript
/**
|
|
* Usage Limit Detector — classifies API errors as usage-limit-related and
|
|
* parks only the task routed through the unavailable provider.
|
|
*
|
|
* Usage-limit errors indicate provider-local conditions (rate limits, quota
|
|
* exceeded, billing issues, overloaded APIs). Continued retrying the affected
|
|
* task is wasteful, but unrelated providers must remain available. Transient
|
|
* server errors (500, timeout, connection refused) are NOT classified as usage-
|
|
* limit errors — they are temporary and may resolve on their own via per-session
|
|
* retry.
|
|
*/
|
|
|
|
import type { Task, TaskStore } from "@fusion/core";
|
|
// FNXC:WorkflowLifecycleColumns 2026-07-30-11:00: `agentType` is an AGENT ROLE, not a column.
|
|
// The planner lane is named `triage` and keeps that name; only the COLUMN was removed by U11.
|
|
import { PLANNER_AGENT_ROLE, resolveTaskLifecycleColumns, type WorkflowIr } from "@fusion/core";
|
|
import {
|
|
resolveExecutorSessionModel,
|
|
resolveMergerSessionModel,
|
|
resolvePlanningSessionModel,
|
|
resolveValidatorSessionModel,
|
|
} from "./agent-session-helpers.js";
|
|
import { createLogger } from "./logger.js";
|
|
|
|
const log = createLogger("usage-limit");
|
|
|
|
/**
|
|
* Patterns that indicate API usage/capacity/billing limits.
|
|
* These are checked case-insensitively against error messages.
|
|
*/
|
|
const USAGE_LIMIT_PATTERNS: RegExp[] = [
|
|
/overloaded/i,
|
|
/rate[_\s]?limit/i,
|
|
/too many requests/i,
|
|
/\b429\b/,
|
|
/\b529\b/,
|
|
/quota/i,
|
|
/billing/i,
|
|
/\bcredit/i,
|
|
/insufficient.*(quota|credit|balance|fund)/i,
|
|
];
|
|
|
|
/**
|
|
* Classify whether an error message indicates a usage-limit condition.
|
|
*
|
|
* Returns `true` for rate limits, overloaded errors, and quota/billing issues.
|
|
* Returns `false` for transient server errors (500/502/503/504, timeout,
|
|
* connection refused) that may resolve on their own.
|
|
*/
|
|
export function isUsageLimitError(errorMessage: string): boolean {
|
|
return USAGE_LIMIT_PATTERNS.some((pattern) => pattern.test(errorMessage));
|
|
}
|
|
|
|
/**
|
|
* Lightweight coordinator that agents call when they detect usage-limit errors.
|
|
* It parks only the task that reached the unavailable provider. A provider-local
|
|
* outage must never activate the project-wide emergency stop because doing so
|
|
* also kills healthy Codex/Claude/Grok work routed through other providers.
|
|
*/
|
|
/**
|
|
* Check if an agent session resolved with an error after exhausting retries.
|
|
*
|
|
* pi-coding-agent's `session.prompt()` does **not** throw when retries are
|
|
* exhausted — it resolves normally and stores the error on
|
|
* `session.state.errorMessage` (was `session.state.error` prior to
|
|
* pi-coding-agent 0.70). Call this immediately after every
|
|
* `await session.prompt(...)` to re-raise the swallowed error so existing
|
|
* `catch` blocks (with `isUsageLimitError` checks) can detect rate-limit
|
|
* conditions and trigger `UsageLimitPauser`.
|
|
*
|
|
* @param session — The agent session (or any object with `state.errorMessage?: string`)
|
|
* @throws {Error} If `session.state.errorMessage` is set and non-empty
|
|
*/
|
|
export function checkSessionError(session: { state: { errorMessage?: string; error?: string } }): void {
|
|
const state = session.state;
|
|
const error = state?.errorMessage ?? state?.error;
|
|
if (error) {
|
|
throw new Error(error);
|
|
}
|
|
}
|
|
|
|
export class UsageLimitPauser {
|
|
constructor(private store: TaskStore) {}
|
|
|
|
private normalizeProviderId(provider: string): string {
|
|
return provider.trim().toLowerCase().replace(/[^a-z0-9._-]+/g, "-").replace(/^-+|-+$/g, "");
|
|
}
|
|
|
|
/**
|
|
* Clear only parks created for a provider whose independent health probe has
|
|
* transitioned back to usable. Manual/user pauses and every other provider
|
|
* reason remain untouched.
|
|
*/
|
|
async onProviderAvailable(provider: string): Promise<number> {
|
|
const providerId = this.normalizeProviderId(provider);
|
|
if (!providerId) return 0;
|
|
|
|
const pausedReason = `provider-rate-limit:${providerId}`;
|
|
// FNXC:ArchitectureHotPath 2026-07-22-17:20: listTasks() must be explicit about payload shape (architecture-hot-paths contract). Recovery only reads scalar pause fields, so request slim rows to avoid loading heavy log/steps/comments for every task.
|
|
const tasks = await this.store.listTasks({ slim: true });
|
|
const recoverableTasks = tasks.filter((task) =>
|
|
task.paused === true
|
|
&& task.userPaused !== true
|
|
&& (task.pausedReason === pausedReason
|
|
|| (providerId === "unknown" && task.pausedReason === "provider-rate-limit")));
|
|
|
|
/*
|
|
FNXC:ProviderRateLimitRecovery 2026-07-19-20:15:
|
|
Provider recovery is a health-state transition, never a task call used as a probe. The daemon's independent authenticated usage/capacity monitor invokes this seam only after positive health, and this exact-reason filter ensures recovery cannot clear manual parks, unrelated failure reasons, or another provider's outage.
|
|
*/
|
|
await Promise.all(recoverableTasks.map(async (task) => {
|
|
await this.store.logEntry(task.id, `Provider ${providerId} is available again; resuming task`);
|
|
await this.store.pauseTask(task.id, false);
|
|
}));
|
|
|
|
if (recoverableTasks.length > 0) {
|
|
log.log(`Provider ${providerId} recovered; resumed ${recoverableTasks.length} task(s)`);
|
|
}
|
|
return recoverableTasks.length;
|
|
}
|
|
|
|
private taskUsesProvider(
|
|
task: Task,
|
|
provider: string,
|
|
settings: Awaited<ReturnType<TaskStore["getSettings"]>>,
|
|
agentType: string,
|
|
preImplementationColumns?: ReadonlySet<string>,
|
|
): boolean {
|
|
/*
|
|
FNXC:WorkflowLifecycleColumns 2026-07-29-15:20 (P0 audit after the Planning-column merge):
|
|
The planning lane was identified by the LITERAL `triage`. The default coding lineage no
|
|
longer declares that column, so this comparison stopped matching for every default-workflow
|
|
card — silently. Nothing throws; the lane simply resolves to no providers, so when a provider
|
|
hits a usage limit during a PLANNING session the fan-out that pauses other tasks on that same
|
|
provider skips every default card, and they keep hammering the rate-limited provider. The
|
|
triggering task is still paused by the explicit fallback below, so no card is stranded — what
|
|
is lost is the blast-radius containment.
|
|
|
|
A planning session runs while the card is PRE-IMPLEMENTATION. The caller has already excluded
|
|
`done`/`archived`, so that is exactly "not the implementation column and not the review
|
|
column" — which matches `todo`, `triage`, `ideas`, and a renamed planner alike. The two
|
|
literals that remain here name the wip and review lanes and belong to the executor/scheduler
|
|
vocabulary conversion, not to this fix.
|
|
*/
|
|
const isPreImplementation = preImplementationColumns?.has(task.column) === true;
|
|
const providersByActiveLane = agentType === PLANNER_AGENT_ROLE
|
|
? (isPreImplementation ? [
|
|
resolvePlanningSessionModel(task.planningModelProvider, task.planningModelId, settings).provider,
|
|
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
|
|
] : [])
|
|
: agentType === "executor"
|
|
? (task.column === "in-progress" ? [
|
|
resolveExecutorSessionModel(task.modelProvider, task.modelId, settings).provider,
|
|
resolveValidatorSessionModel(task.validatorModelProvider, task.validatorModelId, settings).provider,
|
|
] : [])
|
|
: agentType === "merger"
|
|
? (task.column === "in-review" ? [resolveMergerSessionModel(settings, undefined, task).provider] : [])
|
|
: [];
|
|
const resolvedProviders = providersByActiveLane;
|
|
return resolvedProviders.some((candidate) => candidate?.trim().toLowerCase() === provider);
|
|
}
|
|
|
|
/**
|
|
* Called by agents when a usage-limit error is detected after retries are exhausted.
|
|
* Parks the affected task while leaving every other provider lane running.
|
|
*
|
|
* @param agentType - The type of agent that hit the limit (e.g., "executor", "triage", "merger")
|
|
* @param taskId - The task that was being processed when the limit was hit
|
|
* @param errorMessage - The error message from the API
|
|
* @param provider - Best-effort provider identifier used in the pause reason
|
|
*/
|
|
async onUsageLimitHit(agentType: string, taskId: string, errorMessage: string, provider?: string): Promise<void> {
|
|
const providerId = this.normalizeProviderId(provider ?? "unknown") || "unknown";
|
|
const pausedReason = `provider-rate-limit:${providerId}`;
|
|
|
|
/*
|
|
FNXC:ProviderRateLimitIsolation 2026-07-19-19:10:
|
|
A 429 is provider-local, not a project emergency. Park only the task that exhausted retries on that provider so healthy provider lanes continue executing. Keep the provider id in structured pause provenance when the caller can identify it; never persist the full provider response as pause metadata.
|
|
*/
|
|
log.warn(`${agentType} hit usage limit${providerId ? ` for ${providerId}` : ""} on ${taskId}: ${errorMessage}`);
|
|
log.warn(`Matched pattern in error: "${errorMessage.slice(0, 200)}"`);
|
|
|
|
// Log the triggering error on the task
|
|
await this.store.logEntry(
|
|
taskId,
|
|
`Usage limit detected (${agentType}${providerId ? `/${providerId}` : ""}): ${errorMessage}`,
|
|
);
|
|
|
|
const [settings, tasks] = await Promise.all([
|
|
this.store.getSettings(),
|
|
// FNXC:ArchitectureHotPath 2026-07-22-17:20: slim payload — this scan only reads column/pause/model-provider scalars, never heavy detail fields.
|
|
this.store.listTasks({ slim: true }),
|
|
]);
|
|
/*
|
|
FNXC:WorkflowLifecycleColumns 2026-07-29-20:50 (P0 audit, PR #2572 review — greptile):
|
|
The planning lane is resolved PER TASK from its own workflow, not inferred by excluding two
|
|
literals. "Not `in-progress` and not `in-review`" reads any custom non-terminal column — a
|
|
second processing lane, a manual hold, a bespoke review stage — as pre-implementation, so a
|
|
planning-provider limit would pause cards that are nowhere near planning. Trait-derived
|
|
intake/hold is the only answer that holds for a workflow this code has never seen.
|
|
|
|
One IR read per WORKFLOW, not per task: the cache is caller-owned (the U1 contract) and shared
|
|
across the whole fan-out, so a 400-card board spanning three workflows reads three IRs. A task
|
|
whose workflow cannot be resolved yields an empty set and is skipped rather than guessed into
|
|
the lane — conservative, because the cost of a wrong include is pausing work that was fine.
|
|
*/
|
|
const irCache = new Map<string, WorkflowIr>();
|
|
const preImplementationByTask = new Map<string, ReadonlySet<string>>();
|
|
if (agentType === PLANNER_AGENT_ROLE) {
|
|
await Promise.all(tasks.map(async (task) => {
|
|
const columns = await resolveTaskLifecycleColumns(this.store, task.id, irCache).catch(() => undefined);
|
|
/*
|
|
FNXC:WorkflowLifecycleColumns 2026-07-29-22:10 (PR #2572 review — greptile, 2nd):
|
|
INTAKE ONLY. `hold` is not a synonym for "planning": a workflow may carry a hold trait on
|
|
a MID-PIPELINE wait — manual release, timed, dependency, external event — and a card
|
|
parked there is downstream of implementation, not queued for planning. Including hold
|
|
would pause it on a planning-provider limit, which is the same over-classification as the
|
|
literal-exclusion predicate this replaced, just further along.
|
|
|
|
The planning session is the one that runs on an intake card, so intake is the lane. When a
|
|
workflow's hold column IS its pre-implementation queue it is normally the same column as
|
|
intake (the merged Planning lane declares both traits) and is covered by that; where they
|
|
differ, the hold column is a wait and is deliberately excluded.
|
|
*/
|
|
const lanes = new Set<string>();
|
|
if (columns?.intake) lanes.add(columns.intake);
|
|
preImplementationByTask.set(task.id, lanes);
|
|
}));
|
|
}
|
|
const affectedTasks = tasks.filter((task) =>
|
|
task.column !== "done"
|
|
&& task.column !== "archived"
|
|
&& task.paused !== true
|
|
&& providerId !== "unknown"
|
|
&& this.taskUsesProvider(task, providerId, settings, agentType, preImplementationByTask.get(task.id)));
|
|
|
|
// Always include the task that produced the 429 even if its actual provider
|
|
// came from a runtime fallback not represented in persisted task settings.
|
|
if (!affectedTasks.some((task) => task.id === taskId)) {
|
|
const triggeringTask = await this.store.getTask(taskId).catch(() => null);
|
|
if (triggeringTask && triggeringTask.paused !== true) affectedTasks.push(triggeringTask);
|
|
}
|
|
|
|
await Promise.all(affectedTasks.map(async (task) => {
|
|
if (task.id !== taskId) {
|
|
await this.store.logEntry(
|
|
task.id,
|
|
`Paused because provider ${providerId} reached a usage limit on ${taskId}`,
|
|
);
|
|
}
|
|
await this.store.pauseTask(task.id, true, undefined, { pausedReason });
|
|
}));
|
|
log.warn(`Paused ${affectedTasks.length} task(s) routed through ${providerId}; other provider lanes remain active`);
|
|
}
|
|
}
|