Files
fusion/packages/engine/src/evaluator.ts
gsxdsm 0da19f7963 fix(core): a renamed archive lane was recorded as done in the eval corpus; flag the scheduler's two honest literals (#3100)
Two pieces, both about the same distinction: which literals are worth
**converting** and which are worth **naming**.

## Converted — the eval corpus was mislabelling renamed archive lanes

`collectDeterministicSignals` writes `column` as a two-value eval-record
field. Against the `archived` literal, a card resting in a renamed
archive lane was recorded as `"done"`.

No crash, no lifecycle decision — a **mislabelled row in the eval
corpus**, which is a dataset every later comparison reads. That is the
expensive kind of quiet: nothing fails, the numbers just drift.

The collector is sync and pure (no store, no workflow), so the lane
answer arrives as an optional parameter.
`HybridEvaluatorService.evaluateTask` is async and already holds an
optional store, which is where the resolution is paid; a store-less
evaluator degrades to the legacy literal rather than failing.

**Only the archived arm was ever wrong.** A renamed *complete* lane was,
and remains, recorded as `"done"` — which is correct. So only that
answer is resolved, and a third case pins that the widening did not turn
every renamed lane into `"archived"`.

## Flagged, not converted — the scheduler's two honest literals

These are the two `scheduler.ts` literals the sync-lane pass did not
take, and **nothing in the file said why**. That silence is the problem:
the obvious next move is to "finish the job" the way the other ten were
converted, and that would make them **inert, not fixed**.

`getTaskWorkflowSelectionImpl` returns `undefined` unconditionally under
PostgreSQL, so `resolveTaskWorkflowIrSync` always answers with the
default builtin IR — proved in
`postgres/sync-workflow-ir-is-always-default.pg.test.ts`, and
`check-inert-sync-lane-conversions` already baselines **twenty** guards
in that state in this same file.

They stay literal and **counted**, which is the honest state. An
unconverted literal is visible to the census; an inert conversion leaves
the backlog and takes the evidence with it. The note names the real
blocker — a sync-capable workflow-selection reader — so the next pass
does not spend a cycle discovering this the way I did.

## Measured

- 3 new cases in `eval-signal-collector.test.ts` — file **5/5 pass**.
- **MUTATION**: restoring the `archived` literal fails the renamed case
and leaves **both** the legacy control and the renamed-complete negative
green. The negative matters here: the fix must not turn every renamed
lane into `"archived"`.
- core eval suites — **4 files / 20 tests**; engine scheduler +
evaluator — **14 files / 143 tests**.
- `tsc --noEmit` clean in both packages; census `--strict`,
`check-lane-wiring`, `check-inert-sync-lane-conversions`,
`check-fnxc-future-dates` clean.

## Census

Both files keep their counts, deliberately:

- `eval-signal-collector.ts` — the remaining entry is the new
parameter's documented default, which is the fallback doing its job.
- `scheduler.ts` — the two literals this PR deliberately leaves visible.

A census that fell here would mean the flags had been marked exempt,
which would assert the code is fine. It is not fine; it is blocked, and
those are different claims with different expiries.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 04:29:21 -07:00

330 lines
12 KiB
TypeScript

import {
collectDeterministicSignals,
computeOverallScore,
normalizeCategoryScore,
resolveScoreBand,
resolveValidatorSettingsModel,
resolveProjectColumnsForRoles,
EVAL_SCORE_CATEGORIES,
type DeterministicSignals,
type EvalScoreCategory,
type EvalTaskResultCreateInput,
type EvaluationEvidenceRef,
type FollowUpDraft,
type Settings,
type TaskDetail,
type TaskEvaluationEvidenceBundle,
type TaskStore,
} from "@fusion/core";
import { collectTaskEvaluationEvidence } from "./evaluator-evidence.js";
import { materializeEvalFollowUps, normalizeEvalFollowUps, resolveEvalFollowUpPolicyMode } from "./eval-followups.js";
import { createFnAgent, promptWithFallback } from "./pi.js";
import { createLogger } from "./logger.js";
import { resolveMcpServersForStore } from "./mcp-resolution.js";
const log = createLogger("evaluator");
export interface EvalRunContext {
runId: string;
startedAt: string;
}
export interface EvaluatorModelOverride {
provider?: string;
modelId?: string;
}
export interface EvaluatorDeps {
cwd: string;
store?: TaskStore;
runPrompt?: (prompt: string, provider?: string, modelId?: string) => Promise<string>;
collectEvidence?: (params: { task: TaskDetail; runId: string; cwd: string; store: TaskStore }) => Promise<TaskEvaluationEvidenceBundle>;
}
interface EvaluatorAiCategoryResponse {
score: number;
rationale: string;
evidence: EvaluationEvidenceRef[];
}
interface EvaluatorAiResponse {
categories: Record<EvalScoreCategory, EvaluatorAiCategoryResponse>;
overallRationale: string;
followUpDrafts: FollowUpDraft[];
}
export function resolveEvaluatorModel(
settings: Partial<Settings>,
override?: EvaluatorModelOverride,
): { provider?: string; modelId?: string } {
if (override?.provider && override?.modelId) {
return { provider: override.provider, modelId: override.modelId };
}
// Temporary fallback until FN-3393 introduces dedicated evaluator settings.
return resolveValidatorSettingsModel(settings);
}
export class HybridEvaluatorService {
constructor(private readonly deps: EvaluatorDeps) {}
async evaluateTask(
task: TaskDetail,
run: EvalRunContext,
settings: Partial<Settings>,
modelOverride?: EvaluatorModelOverride,
): Promise<Omit<EvalTaskResultCreateInput, "taskId" | "taskSnapshot">> {
/*
FNXC:WorkflowResolvedColumns 2026-07-31-23:55:
Resolve the archive lanes here, where an await is legal, and hand them to the sync collector. The
store is OPTIONAL on this service, so a store-less evaluator degrades to the legacy `archived`
literal — the collector's documented default — rather than failing.
*/
const archivedColumns = this.deps.store
? await resolveProjectColumnsForRoles(this.deps.store, ["archived"]).catch(() => undefined)
: undefined;
const deterministicSignals = collectDeterministicSignals(task, run, { archivedColumns });
const model = resolveEvaluatorModel(settings, modelOverride);
const evidenceBundle = this.deps.store
? await (this.deps.collectEvidence ?? collectTaskEvaluationEvidence)({
store: this.deps.store,
task,
runId: run.runId,
cwd: this.deps.cwd,
})
: undefined;
const prompt = buildEvaluationPrompt(task, run, deterministicSignals, evidenceBundle);
const responseText = await this.runPrompt(prompt, model.provider, model.modelId);
const ai = parseAiResponse(responseText);
const categoryScores = EVAL_SCORE_CATEGORIES.map((category) => {
const aiCategory = ai.categories[category];
return normalizeCategoryScore({
category,
deterministicScore: deriveDeterministicCategoryScore(category, deterministicSignals),
aiScore: aiCategory.score,
rationale: aiCategory.rationale,
evidence: aiCategory.evidence.map((ev) => ({
type: "other",
ref: `${ev.kind}:${ev.label}`,
excerpt: ev.value,
metadata: { source: ev.source },
})),
});
});
const overallScore = computeOverallScore(categoryScores);
const followUpPolicyMode = resolveEvalFollowUpPolicyMode(settings.taskEvaluationFollowUpPolicy);
const followUps = this.deps.store
? await materializeEvalFollowUps({
parentTaskId: task.id,
runId: run.runId,
policyMode: followUpPolicyMode,
overallScore,
store: this.deps.store,
followUps: await normalizeEvalFollowUps({
parentTaskId: task.id,
runId: run.runId,
overallBand: resolveScoreBand(overallScore),
drafts: ai.followUpDrafts,
store: this.deps.store,
policyMode: followUpPolicyMode,
}),
})
: [];
return {
status: "scored",
overallScore,
categoryScores,
rationale: ai.overallRationale,
summary: ai.overallRationale,
evidence: categoryScores.flatMap((categoryScore) => categoryScore.evidence),
evidenceBundle,
deterministicSignals: deterministicSignalsToEvalSignals(deterministicSignals),
followUps,
metadata: {
runId: run.runId,
evaluatorModel: model,
evaluatorRationale: ai.overallRationale,
hybridEvaluation: {
deterministicSignals,
ai,
},
},
};
}
private async runPrompt(prompt: string, provider?: string, modelId?: string): Promise<string> {
if (this.deps.runPrompt) {
return this.deps.runPrompt(prompt, provider, modelId);
}
let text = "";
// FNXC:McpConfig 2026-06-25-23:05: Evaluator sessions are an AI lane and receive the store-resolved MCP set at session creation; createFnAgent applies runtime support gating without logging plaintext env/header secrets.
const { session } = await createFnAgent({
cwd: this.deps.cwd,
systemPrompt: "You are a strict evaluator. Reply with JSON only.",
tools: "readonly",
defaultProvider: provider,
defaultModelId: modelId,
mcpServers: this.deps.store ? (await resolveMcpServersForStore(this.deps.store)).servers : undefined,
onText: (delta) => {
text += delta;
},
});
try {
await promptWithFallback(session, prompt);
return text;
} finally {
try {
session.dispose();
} catch (error) {
log.warn(`Evaluator session disposal failed: ${error instanceof Error ? error.message : String(error)}`);
}
}
}
}
function deriveDeterministicCategoryScore(category: EvalScoreCategory, signals: DeterministicSignals): number {
const workflowPassRate = signals.workflowSummary.total > 0
? (signals.workflowSummary.passed / signals.workflowSummary.total) * 100
: 50;
const errorPenalty = Math.min(signals.logSummary.errorCount * 20, 60);
const warningPenalty = Math.min(signals.logSummary.warningCount * 5, 25);
const commitScore = Math.min(signals.commitSummary.commitCount * 15, 100);
const reviewScore = signals.reviewStatus === "approved" ? 100 : signals.reviewStatus ? 70 : 50;
switch (category) {
case "agentPerformance":
return Math.round(Math.max(0, Math.min(100, (workflowPassRate * 0.5) + (reviewScore * 0.3) + (commitScore * 0.2) - warningPenalty)));
case "taskOutcomeQuality":
return Math.round(Math.max(0, Math.min(100, (workflowPassRate * 0.6) + (commitScore * 0.2) + (100 - errorPenalty) * 0.2)));
case "processCompliance":
return Math.round(Math.max(0, Math.min(100, (workflowPassRate * 0.5) + (reviewScore * 0.3) + ((signals.logSummary.timingEntries > 0 ? 100 : 50) * 0.2) - errorPenalty)));
default:
throw new Error(`Unsupported eval score category: ${String(category)}`);
}
}
function deterministicSignalsToEvalSignals(signals: DeterministicSignals): Array<{ signalId: string; kind: string; name: string; value?: string | number; passed?: boolean }> {
return [
{
signalId: "workflow-summary",
kind: "workflow",
name: "workflow-summary",
value: `${signals.workflowSummary.passed}/${signals.workflowSummary.total}`,
passed: signals.workflowSummary.failed === 0,
},
{
signalId: "timing-ms",
kind: "timing",
name: "timed-execution-ms",
value: signals.timedExecutionMs,
},
{
signalId: "commit-count",
kind: "commit",
name: "commit-count",
value: signals.commitSummary.commitCount,
},
];
}
function formatEvidenceForPrompt(evidenceBundle: TaskEvaluationEvidenceBundle): string {
return JSON.stringify({
sourceOrder: evidenceBundle.sourceOrder,
taskMetadata: evidenceBundle.taskMetadata,
commits: evidenceBundle.commits,
workflow: evidenceBundle.workflow,
reviews: evidenceBundle.reviews,
documents: evidenceBundle.documents,
taskActivity: evidenceBundle.taskActivity,
agentLogs: evidenceBundle.agentLogs,
runAudit: evidenceBundle.runAudit,
}, null, 2);
}
export function buildEvaluationPrompt(
task: TaskDetail,
run: EvalRunContext,
deterministicSignals: DeterministicSignals,
evidenceBundle?: TaskEvaluationEvidenceBundle,
): string {
return [
"Evaluate the completed task and respond with strict JSON.",
"Scores must be integers between 0 and 100.",
"When citing evidence, use labels that include evidence IDs from the ## Evidence section.",
`Run: ${run.runId}`,
"Schema:",
JSON.stringify({
categories: {
agentPerformance: { score: 0, rationale: "", evidence: [{ kind: "task", label: "", value: "", source: "" }] },
taskOutcomeQuality: { score: 0, rationale: "", evidence: [{ kind: "task", label: "", value: "", source: "" }] },
processCompliance: { score: 0, rationale: "", evidence: [{ kind: "task", label: "", value: "", source: "" }] },
},
overallRationale: "",
followUpDrafts: [{ title: "", description: "", reason: "", evidenceRefs: [] }],
}, null, 2),
"Task:",
JSON.stringify({
id: task.id,
title: task.title,
column: task.column,
status: task.status,
summary: task.summary,
}, null, 2),
"Deterministic signals:",
JSON.stringify(deterministicSignals, null, 2),
"## Evidence",
evidenceBundle ? formatEvidenceForPrompt(evidenceBundle) : JSON.stringify({ sourceOrder: [], note: "No evidence bundle available" }, null, 2),
].join("\n\n");
}
export function parseAiResponse(raw: string): EvaluatorAiResponse {
const candidate = extractJson(raw);
let parsed: unknown;
try {
parsed = JSON.parse(candidate);
} catch (error) {
throw new Error(`Evaluator AI response was not valid JSON: ${error instanceof Error ? error.message : String(error)}`);
}
const record = parsed as Partial<EvaluatorAiResponse>;
if (!record.categories || typeof record.categories !== "object") throw new Error("Evaluator response missing categories");
if (typeof record.overallRationale !== "string" || !record.overallRationale.trim()) throw new Error("Evaluator response missing overallRationale");
const categories = {} as Record<EvalScoreCategory, EvaluatorAiCategoryResponse>;
for (const category of EVAL_SCORE_CATEGORIES) {
const entry = (record.categories as Record<string, EvaluatorAiCategoryResponse>)[category];
if (!entry) throw new Error(`Evaluator response missing category ${category}`);
if (!Number.isInteger(entry.score) || entry.score < 0 || entry.score > 100) {
throw new Error(`Evaluator category ${category} score must be an integer in 0..100`);
}
if (typeof entry.rationale !== "string" || !entry.rationale.trim()) {
throw new Error(`Evaluator category ${category} rationale is required`);
}
if (!Array.isArray(entry.evidence) || entry.evidence.length === 0) {
throw new Error(`Evaluator category ${category} evidence is required`);
}
categories[category] = entry;
}
return {
categories,
overallRationale: record.overallRationale,
followUpDrafts: Array.isArray(record.followUpDrafts) ? record.followUpDrafts : [],
};
}
function extractJson(raw: string): string {
const trimmed = raw.trim();
if (trimmed.startsWith("```")) {
return trimmed.replace(/^```(?:json)?\s*/i, "").replace(/\s*```$/, "").trim();
}
const first = trimmed.indexOf("{");
const last = trimmed.lastIndexOf("}");
if (first >= 0 && last > first) return trimmed.slice(first, last + 1);
return trimmed;
}