Files
fusion/packages/core/src/eval-automation.ts
gsxdsm 109204c590 fix: the query class — three sweeps that never ran on a renamed board (#2818)
Three sweeps that **never ran at all** on a renamed board, plus the
shared answer the rest of the class needs. Consolidated from three
handoff branches so the helper appears once. #2811 merged, so this is my
only open PR.

`#2800` measured this class and shipped evidence deliberately without
conversions: `listTasks({ column: "<literal>" })` filters in the store,
so on a renamed board the read returns an **empty array** and the sweep
it feeds does nothing. The census scores the comparison *inside* the
loop, never the query above it.

## What was broken

| file | census count | what actually happened on a renamed board |
|---|---|---|
| `backlog-pressure-reporter.ts` | **0** | both reads empty, ratio
computed as 0/0 — **the alert never fired**, on a board that may be
under exactly the pressure it reports |
| `stale-task-reporter.ts` | **0** | both reads empty — **no stale-task
signal ever raised**, where work is most likely sitting unnoticed |
| `restart-recovery-coordinator.ts` | flagged | sweep never ran — **an
engine restart left interrupted tasks stuck with no requeue** |

Two of the three have a census count of **zero**. They contain no
lifecycle comparison at all, so they have never appeared in the backlog,
in a per-file list, or in any "N → 0" claim — and were completely inert.
**A file at zero is not evidence of anything.**

## The shared answer, and what it is not

Every existing resolver answers a **per-task** question. A query has no
task in hand, so it needs the project-level one: every column any
workflow declares for a role, unioned with the legacy ids so a board
mid-rename still finds rows under the old ones. The set is never empty,
so a caller cannot accidentally query nothing.

The header states what it is **not**: answering a per-card question from
the union would mark a card as review because some *other* workflow
calls its column review — the flat-set mistake this program has made
four times.

## The finding that generalises: the query is rarely the whole defect

`stale-task-reporter` **still reported zero after the query was fixed**
— `getTaskAgeStalenessSignal` defaults to the legacy pair, so a card the
query now returned was refused inside the signal. Converting only the
query would have looked like a fix and changed nothing.

That is a caveat on #2800's approach, offered as refinement rather than
correction: **asserting the query ARGUMENT is right when pinning a known
defect** (the outcome is 0 either way) **and insufficient when proving a
fix**, because the outcome is the only thing that distinguishes a real
conversion from a deeper one. All three conversions here assert
outcomes.

`restart-recovery` had three layers — query, a redundant re-assertion
(deleted; a test pins the `paused` guard it did contribute), and a move
destination that was **already** resolved but whose warning comment was
stale. A stale warning is its own hazard: it told the next reader a
defect existed where none did.

## Verification

- helper **8 passed** · three reporter/coordinator suites **29 passed**
- `pnpm test:gate` **161 / 13 / 487 / 71** · lint clean · `--strict`
exits 0 · four `tsc` targets clean
- each conversion revert-proven independently; the failing case is named
in each test header

## Two mistakes worth recording

**The helper's own test caught a bug in it.** My first draft wrapped the
definition loop in one `try`, and `parseWorkflowIr` **validates** rather
than parses — one malformed row would have returned legacy-only lanes
for *every* workflow, indistinguishable from the bug it exists to fix.
Now isolated per definition.

**I clobbered the core barrel** by taking `index.ts` wholesale from a
handoff branch, dropping two exports `main` had added since; three
packages stopped compiling. Taking a file from another branch takes its
whole contents, including what is now stale — for a barrel that is
nearly always wrong. Re-applied as a single edit on top of `main`.

## Not included

`self-healing.ts`'s 49 — actively owned and mid-conversion; an outside
refactor there produces conflicting halves of one sweep.
`project-engine.ts` (7) and `executor.ts` (2) need their own read of
what each sweep does with the rows, which these three are the argument
for.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:55:10 -07:00

360 lines
12 KiB
TypeScript

import type { AutomationStore } from "./automation-store.js";
import { resolveProjectColumnsForRoles } from "./project-lane-vocabulary.js";
import type { ScheduledTask, ScheduledTaskCreateInput } from "./automation.js";
import type { EvalRun, EvalTaskResultCreateInput } from "./eval-types.js";
import { EvalLifecycleError } from "./eval-store.js";
import type { EvalStore } from "./eval-store.js";
import type { AsyncEvalStore } from "./async-eval-store.js";
import type { ProjectSettings, Task } from "./types.js";
export const TASK_EVALUATION_SCHEDULE_NAME = "Scheduled Task Evaluation";
export const DEFAULT_TASK_EVALUATION_SCHEDULE = "0 5 * * *";
export const TASK_EVALUATION_SCHEDULE_COMMAND = "fn eval --scheduled-batch";
export interface ResolvedTaskEvaluationSettings {
taskEvaluationEnabled: boolean;
taskEvaluationSchedule: string;
taskEvaluationProvider?: string;
taskEvaluationModelId?: string;
taskEvaluationFollowUpPolicy: "off" | "suggest" | "create";
taskEvaluationRetention?: number;
}
export function resolveTaskEvaluationSettings(
settings: Partial<ProjectSettings>,
): ResolvedTaskEvaluationSettings {
const evalSettings = settings as Partial<ResolvedTaskEvaluationSettings>;
return {
taskEvaluationEnabled: evalSettings.taskEvaluationEnabled ?? false,
taskEvaluationSchedule: evalSettings.taskEvaluationSchedule ?? DEFAULT_TASK_EVALUATION_SCHEDULE,
taskEvaluationProvider: evalSettings.taskEvaluationProvider,
taskEvaluationModelId: evalSettings.taskEvaluationModelId,
taskEvaluationFollowUpPolicy: evalSettings.taskEvaluationFollowUpPolicy ?? "off",
taskEvaluationRetention: evalSettings.taskEvaluationRetention,
};
}
export function createScheduledEvalBatchAutomation(
settings: Partial<ProjectSettings>,
): ScheduledTaskCreateInput {
const resolved = resolveTaskEvaluationSettings(settings);
return {
name: TASK_EVALUATION_SCHEDULE_NAME,
description: "Evaluates tasks completed since the previous scheduled evaluation batch",
scheduleType: "custom",
cronExpression: resolved.taskEvaluationSchedule,
command: TASK_EVALUATION_SCHEDULE_COMMAND,
enabled: true,
scope: "project",
};
}
export async function syncScheduledEvalBatchAutomation(
automationStore: AutomationStore,
settings: Partial<ProjectSettings>,
): Promise<ScheduledTask | undefined> {
const { AutomationStore } = await import("./automation-store.js");
const resolved = resolveTaskEvaluationSettings(settings);
const schedules = await automationStore.listSchedules();
const existing = schedules.find((s) => s.name === TASK_EVALUATION_SCHEDULE_NAME);
if (!resolved.taskEvaluationEnabled) {
if (existing) await automationStore.deleteSchedule(existing.id);
return undefined;
}
if (!AutomationStore.isValidCron(resolved.taskEvaluationSchedule)) {
throw new Error(`Invalid task evaluation schedule: ${resolved.taskEvaluationSchedule}`);
}
const input = createScheduledEvalBatchAutomation(settings);
if (existing) {
return automationStore.updateSchedule(existing.id, {
scheduleType: "custom",
cronExpression: input.cronExpression,
command: input.command,
enabled: true,
scope: "project",
});
}
return automationStore.createSchedule(input);
}
export interface EvalBatchWindow {
windowStartExclusive?: string;
windowEndInclusive: string;
}
export interface CompletedTaskEvaluationContext {
run: EvalRun;
task: Task;
taskIndex: number;
totalTasks: number;
window: EvalBatchWindow;
}
export type CompletedTaskEvaluator = (
context: CompletedTaskEvaluationContext,
) => Promise<Omit<EvalTaskResultCreateInput, "taskId" | "taskSnapshot">>;
export interface EvalBatchTaskStore {
listTasks(options?: { column?: string }): Promise<Task[]>;
// Both stores expose the same lifecycle API; the batch awaits every call.
getEvalStore(): EvalStore | AsyncEvalStore;
}
export interface RunScheduledEvalBatchParams {
store: EvalBatchTaskStore;
projectId: string;
evaluator: CompletedTaskEvaluator;
startedAt?: string;
}
export interface ScheduledEvalBatchResult {
runId: string;
status: "completed" | "failed";
windowStartExclusive?: string;
windowEndInclusive: string;
selectedTaskIds: string[];
tasksSelected: number;
}
export async function runScheduledEvalBatch(
params: RunScheduledEvalBatchParams,
): Promise<ScheduledEvalBatchResult> {
const startedAt = params.startedAt ?? new Date().toISOString();
const evalStore = params.store.getEvalStore();
/*
FNXC:ScheduledEvalsPostgres 2026-07-13-22:38:
Scheduled evaluation is a backend-independent operator workflow. Await the shared EvalStore/AsyncEvalStore contract throughout so PostgreSQL performs the same window selection, lifecycle transitions, scoring, and audit-event writes as the legacy synchronous path.
*/
const [previousScheduledBatch] = await evalStore.listRuns({
projectId: params.projectId,
trigger: "schedule",
status: "completed",
order: "desc",
limit: 1,
});
const windowStartExclusive =
(previousScheduledBatch?.metadata?.windowEndInclusive as string | undefined) ??
previousScheduledBatch?.window.until;
const windowEndInclusive = startedAt;
let run: EvalRun;
try {
run = await evalStore.createRun({
projectId: params.projectId,
trigger: "schedule",
scope: "completed-tasks",
window: {
since: windowStartExclusive,
until: windowEndInclusive,
},
metadata: {
windowStartExclusive,
windowEndInclusive,
},
});
} catch (error) {
if (error instanceof EvalLifecycleError && error.code === "active_run_conflict") {
throw error;
}
throw error;
}
await evalStore.appendRunEvent(run.id, {
type: "info",
message: "Scheduled eval batch started",
status: "pending",
metadata: { windowStartExclusive, windowEndInclusive },
});
await evalStore.updateRun(run.id, { status: "running", startedAt });
try {
/*
FNXC:WorkflowLifecycleColumns 2026-08-01-03:10:
THE QUERY plus its redundant re-assertion — the scheduled eval run selected NOTHING.
`listTasks({ column })` filters in the store, so on a renamed board this read returned an empty
array and every scheduled eval run completed having evaluated zero tasks. The `.filter`'s
`task.column === "done"` below it re-asserted the column the query had already selected on, so
converting that comparison alone would have dropped a census count and changed nothing — the list
was empty before the filter ran.
The redundant clause is DELETED rather than converted: a second copy of the same rule is how a
read and its filter drift apart. The completion-timestamp window is the only thing it contributed
beyond the column, and that is kept.
Project-level resolution, because a read has no task in hand, unioned with the legacy id so a
board mid-rename still evaluates rows stored under the old one.
*/
const completeColumns = await resolveProjectColumnsForRoles(params.store as never, ["complete"]);
const byId = new Map<string, Awaited<ReturnType<typeof params.store.listTasks>>[number]>();
for (const column of completeColumns) {
for (const task of await params.store.listTasks({ column })) byId.set(task.id, task);
}
const doneTasks = [...byId.values()].filter((task) =>
Boolean(task.executionCompletedAt)
&& (!windowStartExclusive || task.executionCompletedAt! > windowStartExclusive)
&& task.executionCompletedAt! <= windowEndInclusive,
);
doneTasks.sort((a, b) => {
const byCompletedAt = (a.executionCompletedAt ?? "").localeCompare(b.executionCompletedAt ?? "");
if (byCompletedAt !== 0) return byCompletedAt;
const byCreatedAt = a.createdAt.localeCompare(b.createdAt);
if (byCreatedAt !== 0) return byCreatedAt;
return a.id.localeCompare(b.id);
});
const selectedTaskIds = doneTasks.map((task) => task.id);
await evalStore.updateRun(run.id, {
counts: { totalTasks: selectedTaskIds.length, scoredTasks: 0, skippedTasks: 0, erroredTasks: 0 },
metadata: {
windowStartExclusive,
windowEndInclusive,
selectedTaskIds,
tasksSelected: selectedTaskIds.length,
},
});
if (doneTasks.length === 0) {
await evalStore.appendRunEvent(run.id, {
type: "info",
status: "completed",
message: "Scheduled eval batch completed with no newly done tasks",
metadata: { tasksSelected: 0 },
});
await evalStore.updateRun(run.id, {
status: "completed",
completedAt: new Date().toISOString(),
summary: "No newly completed tasks found in evaluation window",
});
return {
runId: run.id,
status: "completed",
windowStartExclusive,
windowEndInclusive,
selectedTaskIds: [],
tasksSelected: 0,
};
}
let scoredTasks = 0;
let skippedTasks = 0;
let erroredTasks = 0;
const evaluatedTaskIds: string[] = [];
for (const [index, task] of doneTasks.entries()) {
try {
const result = await params.evaluator({
run,
task,
taskIndex: index,
totalTasks: doneTasks.length,
window: { windowStartExclusive, windowEndInclusive },
});
await evalStore.createTaskResult(run.id, {
...result,
taskId: task.id,
taskSnapshot: {
taskId: task.id,
title: task.title,
column: task.column,
createdAt: task.createdAt,
updatedAt: task.updatedAt,
executionCompletedAt: task.executionCompletedAt,
summary: task.summary,
},
metadata: {
...(result.metadata ?? {}),
windowEndInclusive,
},
});
evaluatedTaskIds.push(task.id);
if (result.status === "scored") scoredTasks += 1;
else if (result.status === "skipped") skippedTasks += 1;
else erroredTasks += 1;
await evalStore.appendRunEvent(run.id, {
type: "task_evaluated",
message: `Evaluated task ${task.id}`,
taskId: task.id,
metadata: { status: result.status },
});
} catch (error) {
erroredTasks += 1;
await evalStore.appendRunEvent(run.id, {
type: "error",
message: `Failed evaluating task ${task.id}`,
taskId: task.id,
metadata: { error: error instanceof Error ? error.message : String(error) },
});
}
}
await evalStore.updateRun(run.id, {
status: "completed",
evaluatedTaskIds,
counts: {
totalTasks: doneTasks.length,
scoredTasks,
skippedTasks,
erroredTasks,
},
completedAt: new Date().toISOString(),
summary: `Scheduled eval batch completed for ${doneTasks.length} task(s)`,
metadata: {
windowStartExclusive,
windowEndInclusive,
selectedTaskIds,
tasksSelected: selectedTaskIds.length,
},
});
await evalStore.appendRunEvent(run.id, {
type: "status_changed",
status: "completed",
message: `Scheduled eval batch completed (${doneTasks.length} tasks selected)`,
metadata: { scoredTasks, skippedTasks, erroredTasks },
});
return {
runId: run.id,
status: "completed",
windowStartExclusive,
windowEndInclusive,
selectedTaskIds,
tasksSelected: selectedTaskIds.length,
};
} catch (error) {
await evalStore.updateRun(run.id, {
status: "failed",
completedAt: new Date().toISOString(),
error: error instanceof Error ? error.message : String(error),
metadata: {
windowStartExclusive,
windowEndInclusive,
},
});
await evalStore.appendRunEvent(run.id, {
type: "error",
status: "failed",
message: "Scheduled eval batch failed",
metadata: { error: error instanceof Error ? error.message : String(error) },
});
return {
runId: run.id,
status: "failed",
windowStartExclusive,
windowEndInclusive,
selectedTaskIds: [],
tasksSelected: 0,
};
}
}