dca20496f4c5cc03d86906d63f5536ee545e2bbf
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2a4013b723 |
consolidate/u9 — review+merge lane: E2E evidence, re-greens, and the conversion blocker (#2646)
**U9's consolidation branch.** Supersedes nothing — #2637 and #2643 are green with zero threads and left for your sweep per rule 3. ## Contents | File | Change | Before → After | |---|---|---| | `__tests__/executor-step-numbering-zero-based.test.ts` | isolate the review-handoff `moveTask` call so the assertion is attributable | **1 failed / 3 passed → 4 passed** | | `__tests__/ce-workflow-step-executor.test.ts` | re-green against the block-first merge boundary | **3 failed / 48 passed → 51 passed** | | `__tests__/goal-anchoring-audit.test.ts` | swallow path reports at debug, not `console.warn` | **1 failed / 6 passed → 7 passed** | Triage-guard counts: **no change**. My lane has no remaining column receivers — the rest belong to the capacity/U7/U8/U11/U12 workers, or are deliberate compat retentions I verified individually (`spec-staleness.ts` carries its own "U11 proof" block; `live-agent-count.ts`'s literal fallback is reachable by flag-less callers). Commits kept small and separated: signature fix, then attribution fix, then the boundary re-green, then the debug-channel fix. **Census reconciliation:** `node scripts/lifecycle-column-census.mjs` reports **11** triage guards on main, and **none are in the review/merge lane** — they are the `moves.ts` flag-OFF branch plus the dashboard cluster. Nothing in this branch moves that number, and I am not chasing the 779 non-triage guards per your instruction. ## 1. The review-handoff assertion (and a lesson) The handoff gained a third argument (workflow move provenance), so a two-arg `toHaveBeenCalledWith` failed on the extra options object while the card moved correctly. My first fix used `expect.anything()` — and I *documented in the comment* that six mutations couldn't make it fail, then shipped it anyway. Greptile (P2) correctly called that out: this flow records two `moveTask` calls, so the assertion is satisfied by the boundary move even if the handoff regresses. **Documenting a weakness is not removing it.** Now the test selects the handoff call by its own marker (`workflowMoveMetadata.reason === "workflow-review-handoff"`), asserts exactly one such call, and asserts its target column: | Mutation | Before | After | |---|---|---| | change the seam's `reason` | green | **NEW=1**, this test only | | retarget the seam to `"done"` | green | **NEW=1**, this test only | ## 2. The merge boundary changed shape `ensureWorkflowMergeBoundaryTask` (`executor.ts:7808`) now **refuses** a foreach step-execute region with incomplete pre-merge node proof — logging `"Workflow merge boundary blocked: <reason>"` and returning **without moving**. The move-then-check sequence this file pinned is gone: `"Workflow merge boundary moved task to in-review before requesting merge"` no longer exists anywhere in production. Three fixes, one per failure: 1. **negative case** pinned the retired move-first log. Now pins the *stronger* property the new order gives: an unproven card is **not moved into review at all**. The old assertion could only say "it was moved, then blocked". Log text asserted by stable prefix — the reason clause enumerates missing instance ids, which is legitimately volatile. 2. **"moves direct-to-merge tasks into in-review"** got zero calls: its fixture recorded no node results, so the gate blocked it. Added one `steps#0:step-execute` pre-merge result. 3. **"completes graph-native checklist projection"** also got zero calls. Its existing `plan` result proves *some* pre-merge node ran but not the per-instance work; the gate additionally requires an instance per foreach step-execute. Added the two matching its two steps. (2) and (3) are the same class as the lifecycle E2E `seedTask` fix in #2634: a fixture that never modelled completed work, asking the engine to advance it, and reading the correct refusal as a failure. Proof shape matched to the evaluator (`source: "node"`, `phase: "pre-merge"`, terminal = `passed`/`skipped`) rather than guessed. Verified the gate is what these fixtures exercise: disabling the boundary proof check fails the negative case (`NEW=1`, that test only). `pnpm test:gate` green, `pnpm lint` clean. ## Where U9 actually stands The conversion (S06/S07/S08) is **not** done, and is now precisely characterised rather than "blocked on U8": `workflow-graph-executor.ts:310` short-circuits every `MERGE_REGION_KINDS` entry to the legacy merge seam, so `merge-gate`, `merge-attempt`, `manual-merge-hold`, `retry-backoff`, `recovery-router` and both `branch-group-*` handlers **never execute**. `createMergeGateHandler` does read `task.autoMerge` and emit auto-on/auto-off — and is never called. The builtin IR's `outcome:auto-*` edges are unreachable. **U9's conversion, concretely: stop short-circuiting `MERGE_REGION_KINDS` and let those nodes run.** S06/S07/S08 all hang off that one change. Safeguard 2 has no node-level representation today, so enabling the region without carrying the `autoMerge` contract into it would let an `autoMerge:false` card merge on PR-readiness alone. Full write-up in `docs/plans/workflow-owned-merge-stack/u9-safeguard-baseline.md` (#2634). ## 3. A recurring class worth a shared helper `goal-anchoring-audit`'s swallow path now reports via `log.debug` (a deliberate demotion of log noise), and `debug` is FUSION_DEBUG-gated so vitest emits nothing — the test asserted a channel that was both wrong *and* disabled. I kept both halves of the contract (swallowed **and** reported) by enabling the flag for that case, rather than deleting the awkward assertion. **This is the third instance this session** — `worktree-pool`, `self-healing`'s auto-archive line, and now this. If a fourth appears it deserves a shared test helper rather than three bespoke fixes. ## Two failing files I could NOT responsibly take — flagged, not touched **`executor-prompt.test.ts` (3 failures) — I ESCALATED THIS AND I WAS WRONG. Retracting.** I flagged these as a possible real pause-contract violation: an agent session spawning while an operator has globally paused the engine. I then finished the diagnosis, and the evidence goes the other way. Recording the retraction with the same detail as the alarm, because a false alarm aimed at another unit costs them a chase. **The discriminator I asked for, resolved.** Six tests in that file assert `expect(mockedCreateFnAgent).not.toHaveBeenCalled()` during global pause; 3 fail. Splitting them by what they drive: | Assertions | Drives | Result | |---|---|---| | `does not resume unpaused in-progress task while global pause is active` (+2 siblings) | no executor method — `task:updated` / resume paths | **pass** | | `parks todo tasks in in-progress when fn_task_done…` (+2 siblings) | `executor.execute(...)` **directly** | **fail** | So the guard holds on every event-driven path and is absent only from the direct `execute()` entry. **And `execute()` is not the guard site — the scheduler is.** `scheduler.ts:1491` is an explicit hard stop (*"Global pause (hard stop): halt all scheduling activity"*), with a second gate at `:1055`, and the scheduler never calls `.execute(` at all — dispatch routes through the runtime. In production a global pause halts scheduling before anything reaches the executor. **Conclusion: the pause contract is intact in production.** The 3 failing tests call `execute()` directly, bypassing the upstream gate, and assert a defence-in-depth check *inside* `execute()` that is not there. They are testing a path production does not take during a pause. What that leaves is a real but much smaller question, and a design one rather than a defect: should `execute()` carry its own pause check as defence-in-depth, given non-scheduler callers exist (self-healing, manual retry)? If yes, add the guard and all six assertions pass. If no, the 3 direct-`execute` assertions are asserting a guarantee the architecture places elsewhere and should be retired. **I have not changed either the code or the tests** — but nobody needs to hunt a pause-contract regression, because there isn't one. **`executor-fast-mode-workflows.test.ts` (1 failure) — mechanism not isolated.** `visitedNodeIds` is `['review']` where the test expects `['start','review']`. Three probes failed to explain it: giving the review node an explicit `column: "in-progress"` changed nothing (so it is not column-based entry resolution), and swapping `seam: "review"` for a plain prompt config did not isolate it either. Two structurally identical sibling tests in the same file still pass with `['start', ...]`, so something in graph traversal distinguishes them that I did not find. That is U8/graph-executor territory; I am not asserting a `visitedNodeIds` shape I cannot explain. ## Still open and green - **#2637** — `task-delete-notice` 21 failed → 34 passed. - **#2643** — shellout allowlist re-pin. **Merge early:** it re-drifts whenever `executor.ts`/`self-healing.ts` line counts shift, with no git conflict to warn you. It already drifted once while open (`executor.ts:17106 → 17198`) and I re-pinned it. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b007de5f94 |
fix(ci,tests): repair binary release pipeline and re-green the full suite
Binary Release (v0.73.0-beta.5 was fully red): - bun compile: mark chromium-bidi external — playwright-core@1.60 (feature-video) optionally requires it and bun fails closed on unresolvable requires. - Windows desktop EXE: quote -c.publish.channel=beta in release.yml; PowerShell tokenizes the bare flag into `-c` + a path and electron-builder ENOENTs on it. Full suite (all 4 shards red from stale-test drift, no product bugs found): - engine: align mock stores/assertions with atomic store.moveTaskIf dispatch (#2371), the fail-closed non-empty PROMPT.md artifact gate (#2390), oldest- first admission (FN-8453), alreadyClaimed graph routing (#2393), startStep step projection (#2403/FN-8464), structured retry presentation (FN-8503), provider-lane pause reasons (#2339), typed column-boundary entry (#2378), Type.Integer in CAS document schemas (#2375), bounded model-registry refresh. - engine-no-blocking-shellout: re-pin 17 drifted allowlist line numbers and drop the stale REBASE_HEAD entry whose execSync was removed. - core: schema-applier expectations track migrations 0033-0035 (96 tables) and the synthetic 0000 fixture gains workflow_work_items/mission_contract_assertions; work-item terminal state is "succeeded" post-#2378. Known follow-up (not addressed here): self-healing starved-refinement escalation bumps task.priority, which FN-8453 oldest-first admission no longer consults. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3f7c32c95c |
refactor(cutover 2/3): engine — graph-owned lifecycle, legacy execution deleted (#2342)
Part **2 of 3** of the IR-driven lifecycle cutover (stacked on #2341; top is #2335). **Scope (80 files, packages/engine + cli/pi skill docs + AGENTS/architecture):** graph-driven column moves via the column-boundary controller (R1), single-mover scheduler/hold-release trait cutover (KTD-2/KTD-9), trait re-keyed self-healing + merger with the R7b confirmed-merge-must-finalize guarantee, graph-exclusive Plan Review with leased dedup (R4/R5), the executeCore body-lift — zero legacy re-entry — with fn_review_step + interceptor machinery deleted and tombstone-ratcheted (R9), builtin workflow runtime fixes (missing hold handler, unseamed-node column inheritance, no-merge completion mover), the 6-column benchmark acceptance suite (11 tests) + 12-builtin lifecycle sweep (94 assertions), and the executor test-harness modernization. Also retires core's interpreter-cutover scaffolding whose last consumer (the authoritative driver) dies here. **Merge order:** #2341 → this → #2335. After #2341 merges, retarget this to main. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4441b72bbc | fix(FN-7273): prevent stale step resume regressions | ||
|
|
a013bc0309 |
FN-6607: align step tools with zero-based prompt steps
Align executor step tools and review bookkeeping with the 0-based Step N labels agents see in PROMPT.md. - Treat fn_task_update and fn_review_step step parameters as 0-indexed values, including validation, logs, checkpoints, and review verdict maps. - Update executor/reviewer/step-runner guidance and generated tool docs to describe Step 0 semantics consistently. - Adjust affected executor and reliability tests and add coverage proving Step 0 progress, review, and revise handling work without off-by-one shifts. - Add a patch changeset for the published Fusion CLI package. Files changed: .changeset/fn-6607-step-numbering.md | 5 + .../cli/skill/fusion/references/engine-tools.md | 4 +- .../engine/src/__tests__/executor-pause.test.ts | 2 +- .../executor-review-step-indexing.test.ts | 18 +- .../src/__tests__/executor-review-verdicts.test.ts | 18 +- .../executor-step-numbering-zero-based.test.ts | 196 +++++++++++++++++++++ .../src/__tests__/executor-step-session.test.ts | 24 ++- ...executor-task-done-revise-verdict-guard.test.ts | 4 +- .../executor-pending-review-skip-retry.test.ts | 8 +- .../task-done-refusal-x-invariant.test.ts | 2 +- packages/engine/src/__tests__/step-runner.test.ts | 4 +- packages/engine/src/executor.ts | 57 +++--- packages/engine/src/reviewer.ts | 3 + packages/engine/src/step-runner.ts | 6 +- 14 files changed, 284 insertions(+), 67 deletions(-) Fusion-Task-Id: FN-6607 Fusion-Task-Lineage: 1b1fb1d8-07ca-4a33-84f5-1dab83388c01 |