test(cli): four suites broke on mocks that predate the PG cutover and U11 (12 red → 0) (#2702)

## What was red

`project.test.ts` (8), `extension.test.ts` (2),
`project-lock-retry.test.ts` (1), `extension-workflow-tools.test.ts` (1)
— **12 failures** on clean main, all from full-suite shard 4/4. **No
product defect in any of them.**

### Cause 1 — an incomplete mock that the fail-soft catch disguised (9
cases)

`getTaskCounts` now **enriches** each row before counting, so a renamed
wip column still counts as running work (`FNXC:WorkflowLifecycleColumns
2026-07-30-12:20`). But `resolveWorkflowIrForTask` and
`enrichRunningAgentTaskShape` were absent from the `@fusion/core` mocks.
Calling an undefined function threw, and `getTaskCounts`'s
**deliberately fail-soft** catch converted that into `{ byColumn: {},
runningAgentCount: 0 }`.

So the failures presented as "task counts are zero" — indistinguishable
from a real counting bug. Worth flagging beyond this PR:
`project-lock-retry.test.ts` exists specifically to prove a transient
lock does *not* "silently masquerade as zero tasks" (FN-7731/FN-7740),
and an incomplete mock produced that exact symptom by a different route.
**That catch will hide the next real enrichment failure the same way.**

### Cause 2 — column-vocabulary drift (3 cases)

A column-less `createTask` used to land in `triage`; post-U11 it lands
in `todo`, the merged Planning column. Three assertions were pinned to
the old landing column, or to `COLUMN_LABELS` values **the product never
prints** — the stale mock said `"Triage"` / `"To Do"` where core says
`"Planning"` / `"Todo"`. A label assertion could pass against a string
that exists only in the mock. The mock's labels are now copied from the
real table.

## Measured

| Check | Result |
|---|---|
| the four files | 12 failed → **0** (107 passed) |
| whole `@runfusion/fusion` package | **122 of 126 files green, 1656
passed** |
| `pnpm test:gate` | **726 passed** |
| `pnpm lint`, CLI `tsc --noEmit` | clean |

**Census unaffected — 721 both with and against my diff**, verified by
reverting the four files and re-measuring rather than assuming test
files are unscanned. (I also caught that `--strict` *writes* the
baseline locally; that write is reverted, so this PR does not touch the
baseline.)

## Not touched

`task.test.ts`'s **5 failures**. That file is claimed by
`feature/tool-permission-gates`, and its 5 are exactly the ones shard
4/4 reports — so shard 4/4 goes 17 → 5, with the remainder belonging to
that PR.

## Two judgement calls recorded in-file rather than made silently

**The `extension-workflow-tools` guard is RE-PINNED, not deleted.** It
asserted "a task created on `builtin:coding` lands in `triage`" — a
byte-identical regression guard that fired because the change was
*intended* (U11's merge). Deleting it would remove the only check that
this landing column stays stable; leaving it pinned a column the product
no longer declares. Re-pinning to `todo` keeps it doing its job.

**The broad-listing assertion now names the group that actually leads
the output.** It asserted a *second* group header inside a listing that
truncates to a text budget. That only ever passed because `triage` sorts
before `todo` in `COLUMNS`, so its group fitted before truncation —
fixture ordering masquerading as a bounding assertion. Per-column
coverage of every group is still asserted by the three filtered cases,
which do not truncate. I also seeded those 8 tasks in an explicit third
column so that case still covers a three-way filter instead of
collapsing to two groups.
This commit is contained in:
gsxdsm
2026-07-30 04:10:06 -07:00
committed by GitHub
parent f6e460acdf
commit dbb53eebaa
4 changed files with 130 additions and 27 deletions

View File

@@ -304,7 +304,18 @@ pgTest("pi extension workflow authoring tools", () => {
expect(result.content[0].text).not.toContain("Column: triage");
});
it("still reports Column: triage for the default builtin:coding workflow (byte-identical regression guard)", async () => {
/*
FNXC:WorkflowResolvedColumns 2026-07-30-20:25:
RE-PINNED, not deleted. This guard existed to catch an unintended change to the column a task
created on `builtin:coding` lands in, and it fired — but the change was INTENDED: U11 merged the
two pre-implementation columns, so the default lineage's intake column is now `todo` (the merged
Planning column, carrying intake+hold+resetOnEntry) and declares no `triage` column at all.
Re-pinning to the new value keeps the guard doing its job. Deleting it would remove the only check
that this landing column stays stable, and leaving it on `triage` pinned a column the product no
longer has.
*/
it("reports the merged Planning column (`todo`) for the default builtin:coding workflow (byte-identical regression guard)", async () => {
const api = createMockApi();
registerExtension(api);
const createTask = api.tools.get("fn_task_create")!;
@@ -317,7 +328,7 @@ pgTest("pi extension workflow authoring tools", () => {
);
expect(result.isError).not.toBe(true);
expect(result.details.column).toBe("triage");
expect(result.content[0].text).toContain("Column: triage");
expect(result.details.column).toBe("todo");
expect(result.content[0].text).toContain("Column: todo");
});
});

View File

@@ -3004,10 +3004,19 @@ pgTest("fn pi extension (runnable structured-output regression slice)", () => {
dependencies: [todoFirst.id],
});
}
/*
FNXC:WorkflowResolvedColumns 2026-07-30-20:40:
Seeded explicitly in a THIRD distinct column. These used to be created with no `column` and
relied on landing in `triage`, which post-U11 no longer exists — they landed in `todo`
instead, so the board became Todo (20) / Done (6) and the "Planning (8)" group the
assertions looked for never existed. Naming a real column keeps this case covering a
three-way column filter rather than collapsing it to two groups.
*/
for (let i = 1; i <= 8; i += 1) {
await store.createTask({
title: `${realisticTaskTitle("triage", i)} ${"x".repeat(1_000)}`,
description: `Realistic triage task ${String(i).padStart(3, "0")}`,
title: `${realisticTaskTitle("in-progress", i)} ${"x".repeat(1_000)}`,
description: `Realistic in-progress task ${String(i).padStart(3, "0")}`,
column: "in-progress",
});
}
for (let i = 1; i <= 6; i += 1) {
@@ -3030,7 +3039,18 @@ pgTest("fn pi extension (runnable structured-output regression slice)", () => {
);
expectSingleBoundedTextBlock(broadResult);
expect(broadResult.content.some((block: { type: string }) => block.type === "image")).toBe(false);
expect(broadResult.content[0].text).toContain("Planning (8):");
/*
FNXC:WorkflowResolvedColumns 2026-07-30-20:55:
Assert the group that actually LEADS the broad listing, not one truncation may cut. This
previously asserted "Planning (8):" and passed only incidentally: `triage` sorts before `todo`
in COLUMNS order, so its group fitted before the text budget truncated the rest. Post-U11
`todo` leads, fills the budget, and the later groups collapse into the "... and N more tasks"
notice — so a second group header here is a fixture-ordering assertion, not a bounding one.
Per-column coverage of every group is asserted by the three filtered cases below, which do not
truncate.
*/
expect(broadResult.content[0].text).toContain("Todo (12):");
expect(broadResult.content[0].text).toContain("truncated to fit; narrow with column/limit");
expect(broadResult.details.count).toBe(26);
for (const { callId, params, header, ids } of [
@@ -3041,9 +3061,9 @@ pgTest("fn pi extension (runnable structured-output regression slice)", () => {
ids: ["FN-001", "FN-002"],
},
{
callId: "list-realistic-triage",
params: { column: "triage", limit: 8 },
header: "Planning (8):",
callId: "list-realistic-in-progress",
params: { column: "in-progress", limit: 8 },
header: "In Progress (8):",
ids: ["FN-013", "FN-014"],
},
{
@@ -3094,7 +3114,13 @@ pgTest("fn pi extension (runnable structured-output regression slice)", () => {
expect(result.content.some((block: { type: string }) => block.type === "image")).toBe(false);
expect(text).toBeTruthy();
expect(text.length).toBeLessThanOrEqual(MAX_TASK_LIST_TEXT_CHARS);
expect(text).toContain("Planning (15):");
/*
FNXC:WorkflowResolvedColumns 2026-07-30-20:45:
These 15 are created with no `column`, so they land in the DEFAULT lineage's intake column,
which post-U11 is `todo` — labelled "Todo". The old expectation "Planning (15):" came from
`triage`'s label back when a column-less create landed there; `triage` is no longer declared.
*/
expect(text).toContain("Todo (15):");
expect(text).toContain("FN-001");
expect(text).toContain("truncated to fit; narrow with column/limit");
expect(result.details.count).toBe(15);

View File

@@ -66,6 +66,15 @@ vi.mock("@fusion/core", () => ({
})),
TaskStore: makeConstructibleMock(makeTaskStore),
countRunningAgentTasks: () => 0,
/*
FNXC:CliTests 2026-07-30-20:10:
`getTaskCounts` enriches each row before counting. Both helpers were missing here, so
`resolveWorkflowIrForTask` was undefined and the throw was absorbed by getTaskCounts's fail-soft
catch — which reported zero tasks, the exact symptom this FN-7740 test exists to catch. The test
could not distinguish "lock retry failed" from "the mock is incomplete".
*/
resolveWorkflowIrForTask: vi.fn(async () => undefined),
enrichRunningAgentTaskShape: vi.fn((task: unknown) => task),
ensureMemoryFileWithBackend: vi.fn(),
readProjectIdentity: vi.fn().mockReturnValue(undefined),
writeProjectIdentity: vi.fn(),

View File

@@ -42,9 +42,21 @@ const mockTaskStoreInit = vi.fn();
const mockTaskStoreListTasks = vi.fn();
const mockTaskStoreClose = vi.fn();
const mockEnsureMemoryFileWithBackend = vi.fn();
const mockGetTaskWorkflowSelectionAsync = vi.fn(async () => null);
const mockGetWorkflowDefinition = vi.fn(async () => undefined);
// Mock @fusion/core
vi.mock("@fusion/core", () => ({
/*
FNXC:CliTests 2026-07-30-23:10 (greptile P2 — mocks bypassed the enrichment):
Spread the REAL module and override only the parts that must be faked (the central/global stores and
the store factory). Previously this factory hand-rolled `countRunningAgentTasks`, `COLUMN_LABELS`,
`resolveWorkflowIrForTask` and `enrichRunningAgentTaskShape`, so the suite exercised a reimplemented
legacy-literal count instead of the workflow-aware enrichment whose absence caused the failures — a
regression in the resolver-to-enrichment integration could leave it green while renamed workflow
columns produced wrong in-flight counts. With the real implementations in place the enrichment path
is genuinely under test, and the label table can no longer disagree with the product.
*/
vi.mock("@fusion/core", async (importOriginal) => ({
...(await importOriginal<Record<string, unknown>>()),
CentralCore: makeConstructibleMock(() => ({
init: mockInit.mockResolvedValue(undefined),
close: mockClose.mockResolvedValue(undefined),
@@ -73,6 +85,10 @@ vi.mock("@fusion/core", () => ({
taskStore: {
init: mockTaskStoreInit,
listTasks: mockTaskStoreListTasks,
/* Workflow readers so the REAL `resolveWorkflowIrForTask` can resolve a custom IR; the
renamed-column case below drives these. Default: no selection -> the default IR. */
getTaskWorkflowSelectionAsync: mockGetTaskWorkflowSelectionAsync,
getWorkflowDefinition: mockGetWorkflowDefinition,
},
shutdown: vi.fn(async () => {}),
})),
@@ -85,23 +101,9 @@ vi.mock("@fusion/core", () => ({
// identity surface so runProjectAdd reaches its forced registration behavior.
hasProjectIdentity: vi.fn(() => false),
isValidSqliteDatabaseFile: vi.fn(() => false),
countRunningAgentTasks: (tasks: Array<{ column: string; status?: string; paused?: boolean }>) => tasks.filter((task) => (
task.column === "in-progress" ||
(task.column === "triage" && task.status === "planning" && !task.paused) ||
(task.column === "in-review" && ["merging", "merging-pr", "merging-fix", "reviewing", "fixing"].includes(String(task.status ?? "")) && !task.paused)
)).length,
ensureMemoryFileWithBackend: mockEnsureMemoryFileWithBackend,
readProjectIdentity: vi.fn().mockReturnValue(undefined),
writeProjectIdentity: vi.fn(),
COLUMNS: ["triage", "todo", "in-progress", "in-review", "done", "archived"],
COLUMN_LABELS: {
triage: "Triage",
todo: "To Do",
"in-progress": "In Progress",
"in-review": "In Review",
done: "Done",
archived: "Archived",
},
}));
vi.mock("node:readline/promises", () => ({
@@ -127,6 +129,10 @@ describe("project commands", () => {
beforeEach(() => {
vi.resetModules();
vi.clearAllMocks();
/* `clearAllMocks` clears calls but NOT queued resolved values, so the renamed-workflow case
would otherwise leak its custom IR into every later test and mis-count their columns. */
mockGetTaskWorkflowSelectionAsync.mockReset().mockResolvedValue(null as never);
mockGetWorkflowDefinition.mockReset().mockResolvedValue(undefined as never);
consoleSpy = vi.spyOn(console, "log").mockImplementation(() => {});
consoleWarnSpy = vi.spyOn(console, "warn").mockImplementation(() => {});
consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {});
@@ -350,7 +356,9 @@ describe("project commands", () => {
expect(mockTaskStoreListTasks).toHaveBeenCalled();
const output = consoleSpy.mock.calls.map((call) => String(call[0])).join("\n");
expect(output).toContain("Total: 3");
expect(output).toContain("To Do: 2");
// FNXC:CliTests 2026-07-30-20:10: the real COLUMN_LABELS entry for `todo` is "Todo"; this
// asserted "To Do", a string the product never prints (the old mock invented it).
expect(output).toContain("Todo: 2");
expect(output).toContain("In Progress: 1");
});
@@ -386,6 +394,55 @@ describe("project commands", () => {
expect(output).toContain("Completed: 10");
});
/*
FNXC:CliTests 2026-07-30-23:30 (greptile P2 — the enrichment had no coverage):
The case the previous stubs made impossible. A card sits in a RENAMED wip column (`building`), so
the legacy `column === "in-progress"` fallback answers NO and the count is only correct if the real
`resolveWorkflowIrForTask` -> `enrichRunningAgentTaskShape` -> `countRunningAgentTasks` chain runs
and reads `countsTowardWip` off the resolved IR.
This is the regression the finding described: with the resolver stubbed to `undefined` and
enrichment stubbed to identity, a break anywhere in that chain left the suite green while
`fn project` reported zero in-flight agents on a renamed board — the number an operator reads to
decide whether the board is busy.
*/
it("runProjectShow counts a RENAMED wip column as In-Flight via real trait enrichment", async () => {
mockGetProject.mockResolvedValue({
id: "proj-1",
name: "demo",
path: "/tmp/demo",
status: "active",
isolationMode: "in-process",
createdAt: "2026-01-01T00:00:00.000Z",
updatedAt: "2026-01-02T00:00:00.000Z",
});
mockGetSettings.mockResolvedValue({});
mockGetProjectHealth.mockResolvedValue(undefined);
mockGetTaskWorkflowSelectionAsync.mockResolvedValue({ workflowId: "custom:renamed", stepIds: [] } as never);
mockGetWorkflowDefinition.mockResolvedValue({
ir: {
version: "v2",
id: "custom:renamed",
nodes: [],
edges: [],
columns: [
{ id: "drafting", label: "Drafting", traits: [{ trait: "hold", config: { release: "capacity" } }] },
{ id: "building", label: "Building", traits: [{ trait: "wip", config: { limitSetting: "maxConcurrent" } }] },
{ id: "shipped", label: "Shipped", traits: [{ trait: "complete" }] },
],
},
} as never);
// `building` is NOT a legacy wip id, so only trait enrichment can classify it as running.
mockTaskStoreListTasks.mockResolvedValue([{ id: "FN-900", column: "building", status: null }]);
const { runProjectShow } = await import("../project.js");
await runProjectShow("proj-1");
const output = consoleSpy.mock.calls.map((call) => String(call[0])).join("\n");
expect(output).toContain("In-Flight Agents: 1");
});
it("runProjectShow derives In-Flight Agents from live executors, triage planners, and in-review agents when central health is stale", async () => {
mockGetProject.mockResolvedValue({
id: "proj-1",