Restores FN-6582's rule, which a later operator request had relaxed — deleting its test
along with it.
Operator decision, now carrying the reason the first reversal lacked: the only legitimate
reason to stop a task is an LLM problem; everything else is fixed at the source, or the AI
is made unable to return anything but what is expected — and if it does anyway, restart
cleanly.
Restarting cleanly already happens, twice, inside executeWorkflowStep: a malformed primary
retries on the fallback model, or self-retries once on the primary when no fallback is
configured. So `malformed` reaching this decision does not mean "one fumbled response" — it
means the reviewer failed to return a usable verdict across every attempt. That IS the
LLM-class condition an operator accepts as a legitimate stop. What it must never mean is
approval.
Measured: a reviewer reported in prose that the deliverables were absent, carried no verdict
JSON, and the gate recorded success. Unreviewed work merged on a rejection nobody could see.
A prose classifier cannot close this — that text held no rejection marker at all ("revise",
"reject", "must fix" all absent) because it was a factual statement of absence. Only the
ABSENCE of a verdict is detectable, so absence must not approve.
Advisory gates keep the relaxation: a step that was never allowed to hold a card does not
start holding one, which is where the original operator ask actually applies.
pnpm lint 0 errors, test:gate green, engine typecheck clean, pipeline-smoke 90/90.