Files
fusion/packages/core/src/workflow-capacity.ts
gsxdsm 743df98aa4 capacity, part 1: merge pinned at 1, worktrees-off mode, and one dead knob deleted (#2502)
First slice of the capacity simplification. Operator: *"just have two
capacity — overall per project agent count and max worktrees. Remove all
other capacities and counts."* Plus two later additions: **merge is
always 1, fixed**, and **worktrees off ⇒ limit by total agents only**.

Three independently revertable commits. No limiter is added anywhere;
one is deleted, one is made structurally absent, and one is pinned.

---

## 1. Merge concurrency ratcheted at 1 (test-only)

I was asked to add a limiter if merge concurrency could be raised. **It
cannot** — there is no setting, workflow property, pool or trait config
anywhere that raises it, so this adds no code and pins what already
holds.

Serialization lives in the **pump**: `drainMergeQueue`’s `mergeRunning`
re-entrancy latch, `activeMergeTaskId` as a single-slot identity, the
`mergeBodyInFlight` next-generation latch, and one `ProjectEngine` per
projectId.

**Not** in the merge-queue lease, which is a per-task ROW (`primaryKey
[projectId, taskId]`) — two tasks can hold leases simultaneously by
construction, and it has exactly one caller (the worktree-reuse
handoff). Ordinary merges never take it. A lease-level test would have
been describing an invariant that layer has never held.

The second half guards the other direction: a merge-concurrency
*setting* would not fail the pump ratchet — it would sit unread until
someone wired it up.

**Revert-proof:** deleting the latch → `expected 1 times, but got 2
times`; deleting the `finally` → latch-stuck; injecting
`maxConcurrentMerges: 2` → fails naming the key; injecting a
`maxParallelLanes` merge-trait field → fails naming the field. Sources
restored byte-identical after each injection.

## 2. `worktreesEnabled` — off means the worktree limit cannot bind

No worktrees-off mode existed (no
`worktreesEnabled`/`useWorktrees`/`worktreeMode` anywhere — only
worktree *configuration*).

**Why not `maxWorktrees: 0`, which needs no new key:** it deadlocks. `??
4` keeps `0` (not nullish), the gate is `used >= limit`, so `0 >= 0`
holds **on an empty board** and nothing ever dispatches — while the
operator-visible reason reads `gate=maxWorktrees; used=0/0`, a limiter
that looks like it is working while the board is dead. It also needs the
Command Center `{min:1}` clamp relaxed. So `0` costs the gate rewrite
*and* the clamp change *and* encodes a mode as a magic value.

**Off is absence, not a big number.** `resolveWorktreeCapacityLimit`
returns `number | null`; `ConcurrencyGateDiagnostic.maxWorktreesGate` is
now optional, so consulting a worktree limit in OFF mode does not
type-check. A gate holding `Infinity` can start binding again the moment
someone "fixes" a comparison; an absent gate cannot.

That paid for itself immediately: making it nullable surfaced a
**second, independent** worktree gate (`activeWorktrees >= maxWorktrees`
early-return) that a skip-by-convention approach would have missed
silently.

**Scope, deliberately:** this is a statement about *counting*, not
isolation. It does not make concurrent agents safe to share one checkout
and builds nothing toward that — the non-worktree paths that exist today
are fallbacks to the operator’s own tree, one of which caused FN-8600.

**Revert-proof:** a resolver ignoring the flag turns both OFF scheduler
tests red while every ON test stays green — they reuse the *same*
fixture (5 in-progress, limit 4) that pre-existing tests prove blocks,
so the pair moves in opposite directions. Removing `disabled:` reddens
the UI test.

## 3. `maxTriageConcurrent` deleted — it controlled nothing

**Measured: zero enforcement reads.** The only `.maxTriageConcurrent`
reference in the repo was a route echoing it back in `/config`. FN-8453
removed the pool it gated and left the knob shipping in
`DEFAULT_SETTINGS`, the settings type, the section registry, the API
response and six i18n catalogs, doing nothing, for releases.

Historical FNXC comments are **updated, not deleted** — they explain a
real past incident; they now say "planning admission slot" so they stop
implying a live setting. Tombstoned so it cannot return.

`/config` loses a field; safe in-repo since `fetchConfig`’s own return
type never declared it.

---

## Two corrections worth recording

- I earlier reported `maxWorktrees` had **no** Settings UI. Wrong —
`WorktreesSection.tsx:47`; my grep was truncated by `head`. It changed
the placement (toggle beside it, rather than a duplicate key in
Scheduling).
- I planned to assert the queued-reason string is rewritten in OFF mode.
Measured that it is **unreachable**: when `maxConcurrent` binds, the
sweep bails before the per-task reason and logs nothing. The test
asserts absence instead.

Two near-misses caught before commit: a pre-existing FN-7505 guard
caught my *new* key missing a description mapping; and editing i18n via
`json.load/dump` silently dropped unrelated duplicate keys
(`autoUpdateAndRestart` in `fr`) — Python keeps only the last of a
duplicated key. Redone textually, every catalog re-validated.

## Verification

`pnpm lint` clean · core/engine/dashboard/i18n typecheck clean · `pnpm
test:gate` green (309 + 10 + 71) · capacity/worktree suites 11/11 ·
engine merge-invariant + scheduler 45/45 · dashboard settings 114/114.
Rebased onto current main and re-verified.

Nothing was booted at any point.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a project setting to enable or disable running tasks in
worktrees.
* Disabling worktrees removes worktree capacity limits from task
scheduling.
* The “Max Worktrees” setting is disabled when worktree execution is
turned off.

* **Changes**
* Removed the unused triage concurrency setting from configuration and
dashboard responses.
* Updated scheduling diagnostics and queue messages to reflect disabled
worktree capacity limits.
  * Added localized labels and help text for the new setting.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 18:52:31 -07:00

250 lines
12 KiB
TypeScript

/**
* Workflow capacity resolution (U6, KTD-10, R9 capacity half).
*
* WIP/capacity limits are trait *configuration*; their *enforcement* is a
* substrate capability that runs INSIDE `moveTaskInternal`'s transaction and is
* NEVER bypassable (not a guard — runs regardless of bypassGuards/recoveryRehome
* /moveSource). This module is the pure resolution layer shared by both the
* in-txn check (`store.ts`) and the hold/release sweep (`@fusion/engine`
* `hold-release.ts`): given a workflow IR + a column id + settings it answers
* - does this column have a `wip` (capacity) trait?
* - what is its effective limit (read-through to `settings.maxConcurrent` for
* the default workflow's in-progress column so the legacy knob keeps working
* — U6 scheduler-integration half)?
* - does its config opt into counting mid-`transitionPending` cards?
*
* It performs NO DB access and NO counting — the caller owns the count (the
* store counts in-txn; the sweep counts from a listTasks snapshot). Keeping the
* resolution pure means the two enforcement points can never disagree on what a
* limit *is*, only on the live count, which is exactly the serialization the
* in-txn check arbitrates (two holds, one slot → one wins).
*/
import type { Settings } from "./types.js";
import type { WorkflowIr, WorkflowIrV2, WorkflowIrColumn } from "./workflow-ir-types.js";
import { DEFAULT_WORKFLOW_COLUMN_IDS } from "./workflow-ir.js";
import { getTraitRegistry } from "./trait-registry.js";
/** The default-workflow column whose WIP limit read-through is
* `settings.maxConcurrent` (the legacy "N agents in-progress" gate). */
const DEFAULT_WIP_COLUMN_ID = "in-progress";
/** Fallback when `maxWorktrees` is unset. Matches `DEFAULT_SETTINGS.maxWorktrees`. */
const DEFAULT_MAX_WORKTREES = 4;
/*
FNXC:CapacityModel 2026-07-28-11:20:
THE one place "are worktrees a capacity dimension for this project?" is answered.
The capacity model is two configurable numbers per project:
1. total agents (`maxConcurrent`) — always binds
2. `maxWorktrees` — binds ONLY when worktrees are enabled
When `worktreeLimitEnabled === false` the operator asked for "limit via total agents
only". This returns `null` for that case, and callers construct NO worktree gate
at all — rather than a gate with a very high or infinite limit. That distinction
is the whole point: a limiter that still exists and merely happens not to bind is
the bug class this program keeps excavating (the pool-id sentinel that never
matched a real pool; the approval gate three surfaces re-derived; the always-true
flag whose "disabled" branch was the live one). An absent gate cannot silently
start binding again; a gate holding `Infinity` can, the moment someone "fixes" a
comparison. `ConcurrencyGateDiagnostic.maxWorktreesGate` is therefore OPTIONAL,
so consulting a worktree limit in OFF mode does not type-check.
Deliberately NOT expressed as `maxWorktrees === 0`. Zero is a legible number that
already means something to the gate (`used >= 0` is true on an empty board, so a
0 limit deadlocks dispatch rather than disabling it) and the Command Center
slider clamps it to a 1..50 range. Overloading a value as a mode is how sentinels
become defects; the boolean says what it means.
*/
export function resolveWorktreeCapacityLimit(
settings: Pick<Settings, "maxWorktrees" | "worktreeLimitEnabled"> | undefined,
): number | null {
if (settings?.worktreeLimitEnabled === false) return null;
const limit = settings?.maxWorktrees;
return typeof limit === "number" && Number.isFinite(limit) ? limit : DEFAULT_MAX_WORKTREES;
}
/** U6 (KTD-10): sentinel effective-workflow id for default-workflow
* (null-selection) tasks, so they all share one per-column capacity pool. It
* is not a real workflow row id (no `builtin:`/custom collision possible). */
export const DEFAULT_WORKFLOW_POOL_ID = "__default-workflow__";
/*
FNXC:WorkflowCapacity 2026-07-28-19:05 (pool-id sentinel fix):
THE one place the "no selection → which pool?" convention is expressed.
It exists because the convention was previously restated at each end of the
comparison, and the two restatements disagreed: the COUNTER bucketed
selection-less rows under `DEFAULT_WORKFLOW_POOL_ID` while `moves.ts` asked the
counter for pool `"builtin:coding"`. Nothing ever landed in the pool being asked
about, so the count came back 0 and a finite limit could never bind — the
in-transaction capacity gate was structurally dead for every default-workflow
task (Phase A3, R1).
A shared constant alone would NOT have prevented that: both sides had the
constant available and one of them still wrote a literal. Both sides now call
THIS function, so "what pool does a selection-less task belong to" has exactly
one answer and no call site is in a position to disagree with it.
NOT to be confused with `DEFAULT_WORKFLOW_ID` ("builtin:coding"). That is a real,
resolvable workflow row id and is the correct fallback when the value is used to
RESOLVE AN IR (as `scheduler.ts` does). This is a bucketing key that deliberately
cannot collide with any workflow id. Using either one in the other's role is the
bug this function exists to make unspellable.
*/
export function resolveCapacityPoolId(selectionWorkflowId: string | null | undefined): string {
return selectionWorkflowId ?? DEFAULT_WORKFLOW_POOL_ID;
}
/** Resolved capacity configuration for a single column. */
export interface ColumnCapacity {
/** True when the column carries a capacity (`wip`/`countsTowardWip`) trait. */
hasCapacity: boolean;
/** The effective max concurrent cards. `Infinity` means "no finite limit"
* (a capacity trait with no resolvable limit does not gate). */
limit: number;
/** Whether mid-`transitionPending` cards (holding their destination slot from
* commit time) count toward the limit. Defaults true: a card that has
* committed its move into the column holds the slot even before its
* post-commit hooks finish (KTD-10). */
countPending: boolean;
}
const NO_CAPACITY: ColumnCapacity = { hasCapacity: false, limit: Infinity, countPending: true };
function findColumn(ir: WorkflowIr, columnId: string): WorkflowIrColumn | undefined {
const v2 = ir as WorkflowIrV2;
if (!Array.isArray(v2.columns)) return undefined;
return v2.columns.find((c) => c.id === columnId);
}
/** True when the IR's column set is exactly the default-workflow column ids. */
function isDefaultWorkflowColumns(ir: WorkflowIr): boolean {
const v2 = ir as WorkflowIrV2;
if (!Array.isArray(v2.columns)) return false;
const ids = v2.columns.map((c) => c.id);
if (ids.length !== DEFAULT_WORKFLOW_COLUMN_IDS.length) return false;
const set = new Set(ids);
return DEFAULT_WORKFLOW_COLUMN_IDS.every((id) => set.has(id));
}
/**
* Resolve the capacity configuration for `columnId` under `ir`.
*
* Limit resolution order:
* 1. An explicit numeric `limit` in the column's `wip` trait config wins.
* 2. A `limitSetting: "maxConcurrent"` declaration reads through to the
* project setting, making the built-in workflow's capacity policy explicit.
* 3. Otherwise, for the DEFAULT workflow's `in-progress` column, read through
* to `settings.maxConcurrent` (default 2) so the legacy knob keeps working
* and flag-ON default-workflow scheduling matches flag-OFF (legacy parity).
* 4. Otherwise the column has a capacity trait but no resolvable finite limit
* → `Infinity` (does not gate; the trait is inert until configured).
*/
export function resolveColumnCapacity(
ir: WorkflowIr,
columnId: string,
settings?: Pick<Settings, "maxConcurrent"> | undefined,
): ColumnCapacity {
const column = findColumn(ir, columnId);
if (!column) return NO_CAPACITY;
const flags = getTraitRegistry().resolveColumnFlags(column);
if (!flags.countsTowardWip) return NO_CAPACITY;
// The capacity trait config (the `wip` trait carries `limit` + `countPending`).
// Find the first trait config whose trait sets countsTowardWip.
let configLimit: number | undefined;
let limitSetting: string | undefined;
let countPending = true;
for (const ct of column.traits) {
const def = getTraitRegistry().getTrait(ct.trait);
if (!def?.flags.countsTowardWip) continue;
const cfg = ct.config ?? {};
if (typeof cfg.limit === "number" && Number.isFinite(cfg.limit)) {
configLimit = cfg.limit;
}
if (typeof cfg.limitSetting === "string") {
limitSetting = cfg.limitSetting;
}
if (typeof cfg.countPending === "boolean") {
countPending = cfg.countPending;
}
break;
}
let limit: number;
if (configLimit !== undefined) {
limit = configLimit;
} else if (limitSetting === "maxConcurrent") {
const maxConcurrent = settings?.maxConcurrent;
limit = typeof maxConcurrent === "number" && Number.isFinite(maxConcurrent) ? maxConcurrent : 2;
} else if (columnId === DEFAULT_WIP_COLUMN_ID && isDefaultWorkflowColumns(ir)) {
// Read-through: legacy maxConcurrent maps onto the default workflow's
// in-progress WIP limit (U6 scheduler integration).
const maxConcurrent = settings?.maxConcurrent;
limit = typeof maxConcurrent === "number" && Number.isFinite(maxConcurrent) ? maxConcurrent : 2;
} else {
limit = Infinity;
}
return { hasCapacity: true, limit, countPending };
}
/*
FNXC:WorkflowCapacity 2026-07-19-02:20 (U4/KTD-9):
Multiple `wip` columns SHARE one budget when they resolve their limit the same
way — via a shared `limitSetting` (e.g. maxConcurrent) or the default-workflow
in-progress read-through. The scheduler's single counter must count occupants
across ALL columns sharing the target's budget, so operator-visible concurrency
does not silently multiply when a workflow has two wip columns. A column with an
explicit numeric `limit` is INDEPENDENT — its budget is itself alone. This pure
helper resolves that column set; both enforcement points (the in-txn check in
moves.ts and the hold/release sweep) sum their live counts across it, keeping one
budget authority (KTD-5).
*/
/** The budget "key" a wip column resolves its limit through. Two columns share a
* budget iff their keys are equal. `undefined` = not a capacity column. */
function resolveColumnBudgetKey(ir: WorkflowIr, columnId: string): string | undefined {
const column = findColumn(ir, columnId);
if (!column) return undefined;
const flags = getTraitRegistry().resolveColumnFlags(column);
if (!flags.countsTowardWip) return undefined;
for (const ct of column.traits) {
const def = getTraitRegistry().getTrait(ct.trait);
if (!def?.flags.countsTowardWip) continue;
const cfg = ct.config ?? {};
// An explicit numeric limit is an independent per-column budget.
if (typeof cfg.limit === "number" && Number.isFinite(cfg.limit)) return `col:${columnId}`;
// A shared setting (maxConcurrent, …) pools every column that names it.
if (typeof cfg.limitSetting === "string") return `setting:${cfg.limitSetting}`;
break;
}
// Default-workflow in-progress read-through pools with the maxConcurrent setting.
if (columnId === DEFAULT_WIP_COLUMN_ID && isDefaultWorkflowColumns(ir)) return "setting:maxConcurrent";
// Capacity trait but no resolvable shared source → independent (self only).
return `col:${columnId}`;
}
/**
* Resolve the set of column ids whose live WIP occupancy shares ONE budget with
* `targetColumn` (KTD-9). Returns `[targetColumn]` for an explicit-`limit` or
* otherwise-independent column, every column sharing a `limitSetting`/default
* read-through for a pooled budget, and `[]` when the target is not a capacity
* column. Deterministic (declared column order).
*/
export function resolveWipBudgetColumns(ir: WorkflowIr, targetColumn: string): string[] {
const targetKey = resolveColumnBudgetKey(ir, targetColumn);
if (!targetKey) return [];
// A truthy targetKey proves findColumn located targetColumn, which proves
// `columns` is an array — and the loop necessarily re-collects targetColumn
// itself, so the set is never empty. No defensive fallbacks needed.
const set: string[] = [];
for (const c of (ir as WorkflowIrV2).columns) {
if (resolveColumnBudgetKey(ir, c.id) === targetKey) set.push(c.id);
}
return set;
}