Commit Graph

3 Commits

Author SHA1 Message Date
gsxdsm
9f24a517cf FN-8956: track resolved review findings
Add durable, scoped resolution states for workflow review findings.

- Persist reviewer-applied and superseded finding receipts without making them actionable.
- Scope supersession claims to a named prior workflow result, preserving duplicate IDs in other review lanes.
- Render informational resolution badges and reject resolved items from revision requests.

Files changed: .changeset/fn-8956-review-finding-resolution.md    |   7 +
 docs/dashboard-guide.md                            |   2 +-
 docs/workflow-steps.md                             |   8 +-
 .../src/__tests__/review-severity-gate.test.ts     |  25 ++++
 .../src/__tests__/workflow-step-results.test.ts    |  45 +++++-
 packages/core/src/index.gate.ts                    |   6 +
 packages/core/src/index.ts                         |   6 +
 packages/core/src/types.ts                         |   2 +
 packages/core/src/types/task/task-review.ts        |   4 +
 packages/core/src/types/workflow/workflow-steps.ts |  15 +-
 .../src/workflows/builtin-code-review-group.ts     |   2 +-
 .../src/workflows/builtin-plan-review-group.ts     |   2 +-
 .../core/src/workflows/review-severity-gate.ts     |  37 ++++-
 .../core/src/workflows/workflow-step-results.ts    |  54 ++++++-
 packages/dashboard/app/api/agents/run-audit.ts     |   1 +
 .../dashboard/app/components/TaskReviewTab.css     |  22 +++
 .../dashboard/app/components/TaskReviewTab.tsx     |  36 +++--
 .../components/__tests__/TaskReviewTab.test.tsx    |  46 ++++++
 .../dashboard/src/__tests__/routes-tasks.test.ts   |  48 ++++++
 .../src/routes/register-task-workflow-routes.ts    |  14 +-
 .../__tests__/review-finding-supersession.test.ts  | 163 +++++++++++++++++++++
 .../__tests__/review-findings-injection.test.ts    |  22 +++
 .../workflow-step-verdict-parsing.test.ts          |  16 +-
 .../engine/src/executor/execute-workflow-graph.ts  | 129 ++++++++--------
 .../engine/src/executor/execute-workflow-step.ts   |  25 +++-
 .../engine/src/executor/run-graph-custom-node.ts   |   7 +
 .../executor/workflow-step-failure-injection.ts    |   8 +-
 .../engine/src/executor/workflow-step-verdict.ts   |  17 ++-
 .../src/workflows/workflow-graph-executor.ts       |  15 ++
 packages/i18n/locales/en/app.json                  |   4 +-
 packages/i18n/locales/es/app.json                  |   4 +-
 packages/i18n/locales/fr/app.json                  |   4 +-
 packages/i18n/locales/ko/app.json                  |   4 +-
 packages/i18n/locales/pt-BR/app.json               |   4 +-
 packages/i18n/locales/zh-CN/app.json               |   4 +-
 packages/i18n/locales/zh-TW/app.json               |   4 +-
 36 files changed, 703 insertions(+), 109 deletions(-)

Fusion-Task-Id: FN-8956

Fusion-Task-Lineage: 80568280-85aa-4a49-a60a-99b75f88f486

Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai>
2026-08-11 13:34:27 -07:00
gsxdsm
963dba6f80 feat: gate review verdicts on finding severity and preserve remediation sessions
Review remediation loops were the dominant cost of task wall-clock: over 14 days,
tasks with >=5 post-review fix rounds were 22% of tasks but consumed 78% of all
task active time, and 311 of 331 recorded findings were spec-internal-consistency
complaints that changed no delivered behavior.

Two causes compounded. Plan/Code Review remediation was unbounded by default, and
the review policy ordered a full re-derivation of the artifact after every edit
("distrust the edit ... fresh holistic pass"), so each round surfaced a fresh crop
of previously-acceptable observations as new blockers.

Make the already-persisted WorkflowReviewFinding.severity load-bearing instead of
decorative: a REVISE only blocks when it carries a finding at or above the review
kind's threshold (plan: P0+P1, code: P0). Non-blocking findings are still parsed,
persisted, and handed to the implementer as advisory notes in PROMPT.md. Fails
closed — a REVISE with no findings, or with any unclassified finding, still blocks,
so prose-only and custom reviewers keep full blocking power. The gate only ever
relaxes a verdict, never promotes one.

Reviewer prompts now request the structured findings schema (Plan Review emitted
none before), define severity by consequence as P0/P1/P2, omit nits entirely rather
than filing them as low-severity findings, and use an incremental re-review contract.
Remediation renders findings grouped by priority and sanctions an explicit decline
with rationale, so a disputed finding has a terminal state.

Also preserve the implementation session across a review bounce: sendTaskBackForFix
no longer nulls sessionFile when preserving resume state, and the executor's finally
no longer clears it on a review handoff. Remediation rounds continue the conversation
instead of re-reading the repo and re-deriving the change they just wrote. The resume
prompt now directs a PROMPT.md re-read, without which a resumed agent would never see
the new findings.

New per-workflow settings planReviewBlockingSeverity / codeReviewBlockingSeverity;
set either to "any" to restore the previous behavior.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 11:28:32 -07:00
gsxdsm
1cf86baa1c refactor: package code organization wave 18 (executor pure peels) (#3317)
## Summary

Wave 18 continues the package code-organization program after wave 17
domain folders (U4 Slice A from
`docs/plans/2026-07-14-001-refactor-package-code-organization-plan.md`).

### What changed
Peel **pure, behavior-preserving** helpers out of
`packages/engine/src/executor.ts` into domain modules under
`packages/engine/src/executor/`, with **stable re-exports** from
`executor.ts` so deep imports and `vi.mock("../executor.js")` keep
working.

| New module | Symbols |
|------------|---------|
| `executor/task-done-refusal.ts` | `evaluateTaskDoneRefusal`,
`determineRevisionResetStart`, skip-bypass refusal helper |
| `executor/workflow-feedback-paths.ts` |
`extractReferencedPathsFromWorkflowFeedback`,
`isAlwaysAllowedScopeLeakPath`, `workflowPathMatchesDeclaredScope` |
| `executor/workflow-step-verdict.ts` |
`FUSION_WORKFLOW_STEP_CONVENTIONS_PREAMBLE`, `parseWorkflowStepVerdict`
/ `parseWorkflowStepOutput`, step outcome types |
| `executor/await-input-parse.ts` | `parseAwaitInputSentinel`,
`parseAwaitInputQuestionToolCall` |
| `executor/no-commit-eligibility.ts` | `getNoCommitEligibilityReason`
(+ prompt heuristics) |

`executor.ts` live LOC ~**22817 → ~22427** (first pure-peel batch; more
peels needed to approach the 2k cap).

### Shims
- `old path` `executor.ts` public exports → `new path` `executor/*.ts` →
delete-when consumer deep-imports are re-pointed (not this PR)

### Test plan
- [x] `@fusion/engine` typecheck
- [x] Oracle: task-done refusal, skip-bypass, workflow malformed
verdict, scope-leak allowlist, executor-step-session, executor-prompt
- [x] `vitest --project=engine-core` (merge-gate curated suite)
- [ ] CI merge gate

**Stack:** wave17 (merged) → **this PR**

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Improved recognition of workflow outcomes from structured and
conversational responses.
* Added support for extracting questions from await-input responses and
tool calls.
* Improved workflow feedback handling for referenced files and declared
scope patterns.
* Added clearer guidance for task execution, approvals, verification,
and available tools.

* **Bug Fixes**
* Prevented completion when required review approvals are missing or
revisions remain pending.
* Improved handling of workflows that legitimately require no code
changes.
  * Added clearer refusal messages and more reliable revision restarts.
  * Sanitized repository paths in Git remediation instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-09 15:46:09 -10:00