FN-8955: harden reviewer verdict extraction

Recover structured reviewer verdicts from prose containing malformed brace or quote characters.

- Scan and recover JSON verdict candidates after parser desynchronization.
- Prevent unreadable structured verdicts from being treated as prose approvals.
- Cover malformed prose and multi-finding verdict recovery across review gates.

Files changed:
 .changeset/fn-8955-verdict-extractor.md            |  7 ++
 docs/workflow-steps.md                             |  4 +-
 packages/engine/src/__tests__/reviewer.test.ts     | 65 +++++++++++++++
 .../workflow-malformed-verdict-gate.test.ts        | 14 ++++
 .../workflow-step-verdict-parsing.test.ts          | 86 ++++++++++++++++++--
 packages/engine/src/execution/reviewer.ts          | 92 ++++++++++++++++++++--
 .../engine/src/executor/workflow-step-verdict.ts   | 16 ++--
 7 files changed, 267 insertions(+), 17 deletions(-)

Fusion-Task-Id: FN-8955

Fusion-Task-Lineage: 84107fd2-b7f3-4162-9e8c-2ec24600d9d5

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
This commit is contained in:
gsxdsm
2026-08-11 15:01:52 -07:00
parent f2c729bf77
commit 43141467fb
7 changed files with 267 additions and 17 deletions

View File

@@ -0,0 +1,7 @@
---
"@runfusion/fusion": patch
---
summary: Reviewer verdicts and findings no longer drop when review prose contains stray braces.
category: fix
dev: Harden `extractJsonObjectCandidates` recovery and share structured-verdict-key guards across review parsers.

View File

@@ -610,7 +610,7 @@ Use a JSON object with this schema:
- Valid `verdict` values are exactly: `APPROVE`, `APPROVE_WITH_NOTES`, `REVISE`.
- `notes` is optional and defaults to `""` when missing or non-string.
- The parser checks fenced and inline JSON candidates, and the **last valid candidate wins**.
- The parser checks fenced and inline JSON candidates, and the **last valid candidate wins**. Inline scanning keeps every balanced object (including nested objects) and is string-aware. If that scan detects desynchronization from arbitrary prose (an unpaired brace, quote, or stray close), it also recovers recent full JSON lines and `JSON.parse`-arbitrated brace slices across a generous trailing window. Fairly allocated close/open-anchor budgets preserve payloads with braces or escaped quotes in strings, many findings, and brace-bearing prose after the payload; recovery is based on scanner desync, not whether the payload happens to be near the end.
Accepted shapes:
@@ -630,7 +630,7 @@ Additional example:
#### Prose Fallback
Legacy prose is still supported when structured JSON is missing:
Legacy prose is still supported when structured JSON is missing. A visible but unreadable quoted JSON `"verdict":` key is never converted into a prose approval: prompt-gate parsing reports `malformed`, while reviewer and plan-review lanes return retryable `UNAVAILABLE`. Plain prose with no structured verdict key remains lenient:
- Output beginning with `REQUEST REVISION` (case-insensitive) maps to `REVISE`.
- Remaining prose becomes `notes`.

View File

@@ -376,6 +376,71 @@ describe("reviewStep — model settings threading", () => {
expect(result.verdict).toBe("APPROVE");
});
it("extracts a trailing REVISE payload after unpaired prose braces", async () => {
mockedCreateFnAgent.mockResolvedValue(createMockSession(
"Implementation looks solid overall. IndexOfAny('{','[') needs one change.\n" +
'{"verdict":"REVISE","notes":"fix the parser"}',
));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("REVISE");
});
it("recovers a quote-desynced multi-finding payload beneath brace-bearing trailing prose", async () => {
const payload = JSON.stringify({
verdict: "REVISE",
notes: 'full { note } with "quotes"',
findings: Array.from({ length: 12 }, (_, index) => ({ id: `finding-${index}`, title: "t", body: "b" })),
}, null, 2);
mockedCreateFnAgent.mockResolvedValue(createMockSession(
`{"example":true}\nodd quote "\n${payload}\nuse } to close\n} } } } } }\n{"example":1}`,
));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("REVISE");
});
it("lets an explicit verdict heading beat a JSON example", async () => {
mockedCreateFnAgent.mockResolvedValue(createMockSession(
'## Verdict: REVISE\nExample: {"verdict":"APPROVE"}',
));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("REVISE");
});
it.each([
'looks good\n{"verdict":"REVISE","notes":"truncated',
'looks good\n{"verdict":"PASS"}',
])("returns UNAVAILABLE rather than laundering unreadable structured verdict intent", async (review) => {
mockedCreateFnAgent.mockResolvedValue(createMockSession(review));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("UNAVAILABLE");
});
/*
* FNXC:ReviewLeniency 2026-08-11-19:37:
* A real explicit verdict heading remains authoritative even when an unrelated
* truncated JSON payload exposes a quoted verdict key. The anti-laundering
* guard applies only to Strategy 4 prose approval, never Strategies 1–2.
*/
it("keeps an explicit APPROVE heading authoritative over truncated JSON verdict intent", async () => {
mockedCreateFnAgent.mockResolvedValue(createMockSession(
'## Verdict: APPROVE\nlooks good\n{"verdict":"REVISE","notes":"truncated',
));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("APPROVE");
});
it("preserves lenient prose approval without a structured verdict key", async () => {
mockedCreateFnAgent.mockResolvedValue(createMockSession("looks good"));
const result = await reviewStep("/tmp/worktree", "FN-100", 1, "Test Step", "plan", "# prompt");
expect(result.verdict).toBe("APPROVE");
});
});
describe("reviewStep — spec review type", () => {

View File

@@ -50,6 +50,20 @@ describe("workflow malformed-verdict gate", () => {
expect(parseWorkflowStepOutput("native skill output", { requireVerdict: false })).toEqual({ output: "native skill output" });
});
/* FNXC:ReviewLeniency 2026-08-11-18:44: Visible but unreadable structured
* verdict intent is malformed, never a lenient prose approval. */
it("does not launder unreadable structured verdicts into prose approval", () => {
for (const output of [
'looks good\n{"verdict":"REVISE","notes":"truncated',
'looks good\n{"verdict":"PASS"}',
]) {
expect(parseWorkflowStepOutput(output)).toEqual({ output, malformed: true });
expect(parseWorkflowStepOutput(output, { requireVerdict: false })).toEqual({ output });
}
expect(parseWorkflowStepOutput("looks good")).toMatchObject({ verdict: "APPROVE" });
expect(parseWorkflowStepOutput("my verdict: looks good")).toMatchObject({ verdict: "APPROVE" });
});
it("extracts only validated findings from the selected trailing verdict JSON", () => {
expect(parseWorkflowStepOutput('prose {"verdict":"REVISE","notes":"old"}\n{"verdict":"REVISE","notes":"new","findings":[{"id":"a","title":"Issue","body":"Fix it","line":3,"severity":"high"},{"id":"a","title":"Second","body":"Also fix"},{"title":"bad","body":""}]}')).toMatchObject({
verdict: "REVISE",

View File

@@ -68,12 +68,44 @@ describe("parseWorkflowStepVerdict", () => {
expect(parseWorkflowStepVerdict(out)).toEqual({ verdict: "REVISE", notes: "tighten the type" });
});
it("prefers the LAST JSON object when several appear", () => {
const out = 'Example format: {"verdict":"REVISE"}. My actual verdict follows.\n' +
it("prefers the LAST JSON object when several appear despite an unpaired prose brace", () => {
const out = 'Example format: {"verdict":"REVISE"}. IndexOfAny(\'{\',\'[\') actual verdict follows.\n' +
'{"verdict":"APPROVE","notes":"ok"}';
expect(parseWorkflowStepVerdict(out)).toEqual({ verdict: "APPROVE", notes: "ok" });
});
it("preserves the reported APPROVE_WITH_NOTES verdict, notes, and findings after unpaired prose", () => {
const out = "prose with IndexOfAny('{','[') in it\n\n" +
'{"verdict":"APPROVE_WITH_NOTES","notes":"n","findings":[{"id":"x","title":"t","body":"b","severity":"low"}]}';
const expected = { verdict: "APPROVE_WITH_NOTES", notes: "n", findings: [{ id: "x", title: "t", body: "b", severity: "low" }] };
expect(parseWorkflowStepVerdict(out)).toEqual(expected);
expect(parseWorkflowStepOutput(out)).toEqual({ output: "n", ...expected });
});
it.each([
["a stray closing prose brace", 'stray } in prose\n{"verdict":"REVISE","notes":"n"}'],
["an odd prose quote with no primary candidate", 'odd quote "\n{"verdict":"REVISE","notes":"n"}'],
])("recovers a REVISE payload after %s", (_scenario, output) => {
expect(parseWorkflowStepVerdict(output)).toEqual({ verdict: "REVISE", notes: "n" });
});
it("recovers a REVISE payload from quote desync, dense findings, and brace-bearing trailing prose", () => {
const findings = Array.from({ length: 12 }, (_, index) => ({ id: `finding-${index}`, title: "t", body: "b" }));
/* FNXC:ReviewLeniency 2026-08-11-18:57: Keep the payload multi-line so no
* trailing-line fast path can rescue it; recovery must reach its outer opening after many inner objects. */
const payload = JSON.stringify({
verdict: "REVISE",
notes: 'full { note } with "quotes"',
findings,
}, null, 2);
const out = `{"example":true}\nodd quote "\n${payload}\nuse } to close\n} } } } } }\n{"example":1}`;
expect(parseWorkflowStepVerdict(out)).toEqual({
verdict: "REVISE",
notes: 'full { note } with "quotes"',
findings,
});
});
// "Any approved" — approval-family verdict tokens all map to an approve pass.
it.each([
['{"verdict":"APPROVED"}', "APPROVE"],
@@ -185,7 +217,7 @@ describe("proseSignalsClearApproval", () => {
});
describe("extractJsonObjectCandidates", () => {
it("returns balanced top-level objects in document order", () => {
it("returns balanced objects in document order", () => {
expect(extractJsonObjectCandidates('a {"x":1} b {"y":2} c')).toEqual(['{"x":1}', '{"y":2}']);
});
@@ -195,8 +227,52 @@ describe("extractJsonObjectCandidates", () => {
]);
});
it("captures a nested object as one top-level candidate", () => {
expect(extractJsonObjectCandidates('prose {"a":{"b":2}} tail')).toEqual(['{"a":{"b":2}}']);
/* FNXC:ReviewLeniency 2026-08-11-18:44: Nested candidates are intentional so a
* poisoned outer prose span cannot hide a later valid payload; the parent closes last. */
it("emits nested objects followed by their parent", () => {
expect(extractJsonObjectCandidates('prose {"a":{"b":2}} tail')).toEqual(['{"b":2}', '{"a":{"b":2}}']);
});
it("keeps a trailing payload after unpaired prose braces of either direction", () => {
const payload = '{"verdict":"APPROVE","notes":"n"}';
expect(extractJsonObjectCandidates(`IndexOfAny('{','[')\n${payload}`).at(-1)).toBe(payload);
expect(extractJsonObjectCandidates(`stray } in prose\n${payload}`).at(-1)).toBe(payload);
});
it("recovers a payload after quote desync with no primary candidate", () => {
const payload = '{"verdict":"REVISE","notes":"full"}';
const candidates = extractJsonObjectCandidates(`odd quote "\n${payload}`);
expect(candidates.at(-1)).toBe(payload);
});
it("recovers a payload after quote desync even when a bogus primary candidate exists", () => {
const payload = '{"verdict":"REVISE","notes":"full"}';
const candidates = extractJsonObjectCandidates(`{"example":true}\nodd quote "\n${payload}`);
expect(candidates).toContain('{"example":true}');
expect(candidates.at(-1)).toBe(payload);
});
it("recovers a brace- and quote-bearing multi-finding payload beneath brace-dense trailing prose", () => {
const payload = JSON.stringify({
verdict: "APPROVE_WITH_NOTES",
notes: 'has { and } plus "quoted" text',
findings: Array.from({ length: 12 }, (_, index) => ({ id: `f-${index}`, title: "t", body: "{ body }" })),
}, null, 2);
const trailing = ["use } to close", "} } }", "example {\"x\":1}", "} } }"].join("\n");
const candidates = extractJsonObjectCandidates(`{"example":true}\nodd quote "\n${payload}\n${trailing}`);
expect(candidates).toContain(payload);
expect(JSON.parse(candidates.find((candidate) => candidate === payload)!)).toEqual(JSON.parse(payload));
});
/* FNXC:ReviewLeniency 2026-08-11-21:39: Primary candidate retention is capped
* so brace-dense reviewer prose cannot grow memory without bound; retain the tail
* because callers prefer the final authoritative verdict. */
it("bounds brace-dense primary candidates while retaining the trailing verdict", () => {
const noise = Array.from({ length: 500 }, (_, index) => `{"example":${index}}`).join(" ");
const payload = '{"verdict":"APPROVE_WITH_NOTES","notes":"tail"}';
const candidates = extractJsonObjectCandidates(`${noise}\n${payload}`);
expect(candidates).toHaveLength(200);
expect(candidates.at(-1)).toBe(payload);
});
});

View File

@@ -1110,13 +1110,25 @@ export function proseSignalsClearApproval(rawOutput: string): boolean {
/*
FNXC:ReviewLeniency 2026-07-01-23:30:
Some models emit PROSE followed by a trailing JSON payload — e.g. a paragraph of reasoning, then `{"verdict":"APPROVE","notes":"..."}` at the very end. Extract balanced top-level `{...}` objects in document order, string/escape aware so a brace inside prose or a notes string does not miscount. Callers prefer the LAST candidate as the authoritative trailing verdict. Shared by extractVerdict (reviewer/plan-review) and parseWorkflowStepVerdict (code-review + browser-verification gate).
Some models emit PROSE followed by a trailing JSON payload — e.g. a paragraph of reasoning, then `{"verdict":"APPROVE","notes":"..."}` at the very end. Extract every balanced `{...}` object in close order, string/escape aware so a brace inside prose or a notes string does not miscount. Callers prefer the LAST candidate as the authoritative trailing verdict. Shared by extractVerdict (reviewer/plan-review) and parseWorkflowStepVerdict (code-review + browser-verification gate).
FNXC:ReviewLeniency 2026-08-11-18:44:
Balance tracking moved the prose-brace failure mode to unpaired braces or quotes. Emit every balanced object and, only after scanner desync, recover JSON.parse-arbitrated slices. Recovery must use scanner state rather than an empty-candidate or trailing-anchor test: desync can emit bogus candidates and prose after the payload can move its close far from the end. Quote-blind rescans are forbidden because braces inside JSON strings truncate payloads. Generous fair close/open budgets prevent prose closes or multi-finding opens from starving the payload; callers must keep preferring the last candidate.
*/
export function extractJsonObjectCandidates(text: string): string[] {
const out: string[] = [];
const MAX_PRIMARY_CANDIDATES = 200;
const MAX_TRAILING_LINES = 50;
const MAX_RECOVERY_WINDOW = 64_000;
const MAX_CLOSE_ANCHORS = 200;
const MAX_OPEN_ANCHORS = 500;
const MIN_OPENS_PER_ANCHOR = 8;
const MAX_RECOVERY_CANDIDATES = 4_000;
const primary: string[] = [];
const starts: number[] = [];
let inString = false;
let escaped = false;
let ignoredStrayClose = false;
for (let i = 0; i < text.length; i += 1) {
const ch = text[i];
if (inString) {
@@ -1129,10 +1141,71 @@ export function extractJsonObjectCandidates(text: string): string[] {
else if (ch === "{") starts.push(i);
else if (ch === "}") {
const start = starts.pop();
if (start !== undefined && starts.length === 0) out.push(text.slice(start, i + 1));
if (start === undefined) {
ignoredStrayClose = true;
} else {
primary.push(text.slice(start, i + 1));
if (primary.length > MAX_PRIMARY_CANDIDATES) primary.shift();
}
}
}
return out;
if (!inString && starts.length === 0 && !ignoredStrayClose && primary.length > 0) return primary;
const recoveryPreferred: string[] = [];
const seen = new Set<string>();
const add = (candidate: string) => {
if (candidate.length > MAX_RECOVERY_WINDOW || seen.has(candidate)) return;
seen.add(candidate);
recoveryPreferred.push(candidate);
};
// A full JSON payload on any recent non-empty line needs no brace interpretation.
let nonEmptyLines = 0;
for (const line of text.split(/\r?\n/).reverse()) {
const candidate = line.trim();
if (!candidate) continue;
nonEmptyLines += 1;
if (candidate.startsWith("{") && candidate.endsWith("}")) add(candidate);
if (nonEmptyLines >= MAX_TRAILING_LINES) break;
}
const windowStart = Math.max(0, text.length - MAX_RECOVERY_WINDOW);
const closes: number[] = [];
for (let i = text.length - 1; i >= windowStart && closes.length < MAX_CLOSE_ANCHORS; i -= 1) {
if (text[i] === "}") closes.push(i);
}
const opensByClose = closes.map((close) => {
const opens: number[] = [];
const earliest = Math.max(windowStart, close - MAX_RECOVERY_WINDOW);
for (let i = close - 1; i >= earliest && opens.length < MAX_OPEN_ANCHORS; i -= 1) {
if (text[i] === "{") opens.push(i);
}
return opens.reverse(); // widest slices first
});
let spans = 0;
const addSpan = (open: number, close: number) => {
if (spans >= MAX_RECOVERY_CANDIDATES) return;
spans += 1;
add(text.slice(open, close + 1));
};
// Phase A gives each close anchor a chance before a bogus one consumes the cap.
for (let closeIndex = 0; closeIndex < closes.length && spans < MAX_RECOVERY_CANDIDATES; closeIndex += 1) {
for (const open of opensByClose[closeIndex].slice(0, MIN_OPENS_PER_ANCHOR)) addSpan(open, closes[closeIndex]);
}
// Phase B deepens every anchor in the same newest-first order.
for (let closeIndex = 0; closeIndex < closes.length && spans < MAX_RECOVERY_CANDIDATES; closeIndex += 1) {
for (const open of opensByClose[closeIndex].slice(MIN_OPENS_PER_ANCHOR)) addSpan(open, closes[closeIndex]);
}
// Callers iterate last→first, so append recovery candidates in reverse preference.
return [...primary, ...recoveryPreferred.reverse()];
}
/** Detects visible JSON verdict intent without treating ordinary prose as structured output. */
export function textHasStructuredVerdictKey(rawOutput: string): boolean {
return /"verdict"\s*:/.test(rawOutput);
}
/*
@@ -1196,7 +1269,16 @@ function extractVerdict(review: string): ReviewVerdict {
// (and carries no revise/reject/negated-approval signal). Treat as APPROVE so
// an imperfectly-structured approval passes instead of collapsing to a
// synthetic UNAVAILABLE retry/block. See proseSignalsClearApproval.
if (proseSignalsClearApproval(review)) {
/*
FNXC:ReviewLeniency 2026-08-11-18:44:
A quoted JSON verdict key that Strategy 3 could not classify must not be laundered
into a prose APPROVE. This shared detector also guards workflow-step gates; here
suppression deliberately falls through to retryable UNAVAILABLE. Explicit heading
and line strategies above remain authoritative, and fail-closed gates are untouched.
*/
if (textHasStructuredVerdictKey(review)) {
reviewerLog.warn(`Structured verdict key was present but unparseable (${review.length} chars). Returning UNAVAILABLE.`);
} else if (proseSignalsClearApproval(review)) {
reviewerLog.log(`Verdict extracted via lenient prose approval (${review.length} chars) → APPROVE`);
return "APPROVE";
}

View File

@@ -6,7 +6,7 @@
* CLOSE_NO_OP is Plan Review only (FN-8841). Exact match + optionalGroupId gate so unrelated
* review groups and prose cannot open a terminal lifecycle path.
*/
import { proseSignalsClearApproval, extractJsonObjectCandidates } from "../execution/reviewer.js";
import { proseSignalsClearApproval, extractJsonObjectCandidates, textHasStructuredVerdictKey } from "../execution/reviewer.js";
import { normalizeSupersededFindingIds, normalizeWorkflowReviewFindings, PLAN_REVIEW_GROUP_ID, type WorkflowReviewFinding } from "@fusion/core";
/** Machine-readable workflow-step verdicts, including Plan Review CLOSE_NO_OP. */
@@ -94,7 +94,7 @@ export function parseWorkflowStepVerdict(
}
/*
FNXC:ReviewLeniency 2026-07-01-23:30:
Prefer a balanced, string-aware object scan over a greedy `\{[\s\S]*\}` match: models that emit reasoning PROSE (which may itself contain braces) followed by a trailing `{"verdict":...}` payload broke the greedy span into invalid JSON. extractJsonObjectCandidates returns each top-level object in document order; iterating last→first prefers the trailing verdict payload.
Prefer a balanced, string-aware object scan over a greedy `\{[\s\S]*\}` match: models that emit reasoning PROSE (which may itself contain braces) followed by a trailing `{"verdict":...}` payload broke the greedy span into invalid JSON. extractJsonObjectCandidates returns every balanced object in close order; iterating last→first prefers the trailing verdict payload.
*/
candidates.push(...extractJsonObjectCandidates(trimmed));
@@ -144,7 +144,10 @@ export function parseWorkflowStepVerdict(
return null;
}
export function inferWorkflowStepVerdictFromProse(rawOutput: string): { verdict: "APPROVE" | "APPROVE_WITH_NOTES" | "REVISE"; notes: string } | null {
export function inferWorkflowStepVerdictFromProse(
rawOutput: string,
options: { suppressLenientApprovalForStructuredVerdict?: boolean } = {},
): { verdict: "APPROVE" | "APPROVE_WITH_NOTES" | "REVISE"; notes: string } | null {
const trimmed = rawOutput.trim();
const revisionMatch = trimmed.match(/^REQUEST REVISION\s*\n*/i);
if (revisionMatch) {
@@ -167,8 +170,11 @@ export function inferWorkflowStepVerdictFromProse(rawOutput: string): { verdict:
/*
FNXC:ReviewLeniency 2026-07-01-22:15:
A gate review (code-review, browser-verification) whose text clearly approves must PASS even when it is not perfectly structured. Delegate to the shared proseSignalsClearApproval detector so this parser and the reviewer/plan-review parser agree on what "clearly approved" means, and so a prose rejection ("not approved", "please revise", "reject") is never promoted to APPROVE. Replaces the prior narrow approve/approved/looks good/no issues/out of scope regex (now a subset of the shared detector).
FNXC:ReviewLeniency 2026-08-11-18:44:
When the gate parser was unable to classify a visible structured verdict, do not invent an APPROVE from nearby prose. The reviewer lane uses this same detector before its lenient branch, while fail-closed merge/PR/mission gates remain untouched because they never used prose leniency.
*/
if (proseSignalsClearApproval(trimmed)) {
if (!options.suppressLenientApprovalForStructuredVerdict && proseSignalsClearApproval(trimmed)) {
return { verdict: "APPROVE", notes: "" };
}
return null;
@@ -205,7 +211,7 @@ export function parseWorkflowStepOutput(rawOutput: string, options: { requireVer
};
}
const inferred = inferWorkflowStepVerdictFromProse(trimmed);
const inferred = inferWorkflowStepVerdictFromProse(trimmed, { suppressLenientApprovalForStructuredVerdict: textHasStructuredVerdictKey(trimmed) });
if (inferred) {
return {
output: inferred.notes || trimmed,