17fdf2f0ae25a068499e2bebece2a8d7a0cdbecf
3550 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
17fdf2f0ae |
U4: scope the user-pause safeguard to lifecycle MUTATION, not observation (re-ratified) (#2486)
Stacked on #2482. Base is `feature/workflow-vocabulary-u4-override-layer` — do not merge before it. Implements the coordinator's **re-ratification** of the user-pause safeguard with a narrower definition. ## The invariant, written into the code > The user-pause safeguard means **NEVER MUTATE LIFECYCLE STATE** of a user-paused card. It does **NOT** mean never observe one. Respecting a pause exists to stop the engine acting on a card *behind* the operator who paused it — moving, rebounding, archiving, resuming. A read-only diagnostic does the opposite: it tells that same operator what their paused card is doing. Blinding them to their own paused work is not safety; it is the engine deciding they should not be told. The sentence is in the source, because the distinction is the whole point and a future reader will otherwise re-broaden it. ## Why it needed narrowing Ratified broadly first, that reading was caught suppressing the very sweeps it was meant to protect. `surfaceStalePausedTodos` exists to report cards that have sat paused too long — routing it through a reconciler that suppresses paused cards turns a diagnostic into one that **silently reports nothing**. Measured, not argued: #2484 proves that sweep surfaces user-paused cards on current main. ## Scoped by action, never by sweep `OBSERVATIONAL_ACTIONS` is an **allow-list**, so a newly added mutating action is suppressed by default — the scoping fails closed. A sweep cannot opt itself out. `RecoveryActionKind` deliberately names actions the policy vocabulary cannot yet author (`rebound`, `archive`, `requeue`, `resume`). The scoping is only testable if those exist as values, and **a rule that cannot be tested is a rule that erodes**. `parseWorkflowIr` keeps a closed action list, so nothing becomes authorable by being named. ## Which field — chosen, not inherited The gate reads `userPaused`, **not** `paused`. The two diverge (`branch-group-ops.ts:128` says so outright), and `paused` also covers engine-authored automation pauses like dispatch-storm, which carry no operator intent to respect — gating on it would suppress recovery from the engine's own throttles. The safeguard defers to a **human** decision, so it keys on the field that records one. ## Both halves kept The broad case is **narrowed, not deleted**: one test proves mutation is still suppressed, one proves observation is now permitted. A future reader must be able to tell the scoping was *deliberate* rather than eroded by someone who found the broad rule inconvenient. **Mutation-verified in both directions**, since either error is silent: | mutation | result | |---|---| | re-broaden (suppress observation) | **2 tests fail** | | over-narrow (`rebound` treated observational) | **3 tests fail** | ## Verification 46 tests green (28 safety + 18 inheritance); tsc clean; lint clean; merge gate green (299+10+71). No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
89284df85e |
E2E: table-driven converted-sweep coverage, two new sites, and an honest unproven-sites ledger (#2485)
Follow-up to #2475 (merged). Test-only, plus one test-utility seam. ## Why a table #2475 proved one converted sweep. The count has since gone to **three**, twice while this work was open — `surfaceStalePausedTodos` appeared during #2475's review, and #2478 landed `recovery-reconciler.ts` while this branch was open. A suite with a bespoke `describe` per sweep is a coverage claim that quietly becomes false. Replaced with a table of `(seed, run, acted, roles, observability)`. The driver derives four assertions per entry: | | positive | negative | |---|---|---| | **renamed vocabulary** | acts on the card | inert in a non-target column | | **default vocabulary** | acts (regression floor) | inert | Adding a converted sweep is **one entry** — the #2478 site proved that in practice, not in principle. `actsOnRole`/`inertRole` are keys of `Vocabulary`, not column strings, so an entry cannot hardcode `todo` and pass for the wrong reason. ## Two findings, both from mutation rather than reading **1. The census was wrong about `recovery-reconciler.ts:198.`** It was flagged as a `resolveLifecycleColumns` site, so the row was first labelled as covering it. **Destroying that role resolution leaves all 18 tests green** — `decideRecovery` looks policy up by *column id* and never consults a role. The row is relabelled to what it actually proves, and mutation-verified against that instead: keying the reconciler's policy lookup on the `todo` literal fails exactly its renamed test. **2. `resolveRoleRecovery` is an unreachable export.** It is the only use of `resolveLifecycleColumns` in that file and has **no production caller anywhere** in engine, core, or dashboard. So that census line is not a live converted site — it is a helper written ahead of its consumer. **Not fixed here:** it is production code owned by the U4 slice, and whether the consumer is still to land or it should be deleted is its author's call. ## Observability is now explicit in the type `persisted-row` is the strong form. `returned-decision` is recorded as **weaker evidence** and the reconciler row uses it, because `reconcileRecovery` decides and does not apply — there is no row to read. Naming it in the type is what stops a return-value assertion from quietly passing as observed state, and it is what keeps the ledger truthful per site. ## Harness seam `PgTestHarness` now exposes its raw admin SQL client. The store **stamps** `updatedAt`/`columnMovedAt` on every write, so `updateTask` cannot express an aged row at all — the patch is accepted and the value silently replaced with `now`. **Found by the new case failing on BOTH vocabularies**, which is what distinguishes a broken fixture from a broken guard. Seeding only; assertions still read back through the real `getTask` path. ## Mutation verification | Mutation | Result | |---|---| | revert **only** `recoverStrandedCompletedTodoTasks`'s resolution | **exactly** that row's renamed test fails | | revert **only** `surfaceStalePausedTodos`'s resolution | **exactly** its own renamed test fails | | reconciler policy lookup keyed on `todo` | exactly the reconciler row's renamed test fails | | `resolveRoleRecovery` role resolution destroyed | **nothing fails** → finding #2 | | `hold-release` `isHeldTask` keyed on `todo` | 5 of 18 fail; default spine survives | | `markMoveInFlight` dropped | both spine tests fail | Per-site verification matters here: three rows could all be riding one guard. They are not. ## The honest number **Proven end to end: 5** (two self-healing sweeps, the reconciler's policy lookup at the weaker observability, hold-release's capacity release, and the graph boundary + `moveTask` + post-commit bus). **Not proven: 11 call sites** — `merger.ts:324-326`, `merger-ai.ts:1022,1039`, `auto-merge-finalization.ts:20-22`, `executor.ts:1763,6339,6341`, `self-healing.ts:713,6732`, `mesh-lease-manager.ts:61`, `task-agent-sync.ts:59`, `core/task-store/reads.ts:130`, `core/live-agent-count.ts:63-75`, and four dashboard route sites. The ledger lives in the file, not just here, so it stays with the code. ## Where the table does not fit — reported, not papered over The **merge/rebound family** cannot be a table row: those sweeps have no observable persisted effect without a real git repository, so `acted` cannot be written against the row at all. They need an engine-slow real-git lane. The dashboard sites need an HTTP route test with a live store. Both are different lanes, not missing entries. ## Verification - 18/18 green; engine + core `tsc --noEmit` clean; `pnpm test:gate` green (299 + 10 + 71) - full core PG suite run (the harness is shared): 1036 passed, 3 failed in `central-archive-secrets` and `workflow-settings-project-identity` — **reproduce identically with this change stashed**, pre-existing 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2dfac47917 |
U4: recovery policy as an OVERRIDE LAYER (unset defers to the operator setting) (#2482)
Stacked on #2478. Base is `feature/workflow-vocabulary-u4-reconciler-slice` — do not merge before it. Implements the ratified precedence: **declared explicitly → policy wins; left unset → defer to the project/global setting, exactly as today.** ## The design `resolveEffectiveRecovery(declared, inherited)` composes the two **per field**, so a workflow may declare a threshold while inheriting the action. It mirrors the two-tier merge `effective-settings.ts` already implements for workflow settings (a stored value overrides the base; a declaration default only fills an absent key) rather than inventing a fourth precedence system beside model selection, project settings, and workflow settings. **Absence stays absent.** `??` treats an explicitly-`undefined` field as unset, so a policy is never normalized into a built-in default. The distinction a naive implementation gets wrong: > **equal-to-default is not the same as unset** A declaration whose value happens to equal the legacy literal is a *deliberate choice* and must still override a customized operator setting. Only true absence defers. An effective policy requires **both** halves — a threshold with no action never fires, an action with no threshold has nothing to fire on — so a half-resolved policy yields `undefined` rather than something present but inert. ## The test that matters was written first, and failed > a project with a CUSTOMIZED threshold and the policy key UNSET must observe the customized value This is where a green suite lies. "Read the policy, else use the built-in default" passes every obvious test while silently resetting an operator who tuned `stalePausedTodoThresholdMs` — no error, nothing in any diff, the sweep just starts firing on a schedule nobody chose. **Mutation-verified in both directions:** | mutation | result | |---|---| | substitute a built-in default for the inherited setting | **5 tests fail** | | invert precedence (inherited beats declared) | **3 tests fail** | ## Upgrade guarantee Asserted as a property over several operator values: an undeclared workflow observes *exactly* the operator's value. That is what makes landing the policy table a zero-behavior-change upgrade that touches no project. ## What is NOT here **`surfaceStalePausedTodos` is not retired.** Migrating it surfaced a safeguard-semantics collision I escalated rather than resolved unilaterally: the sweep exists to surface cards that have been **paused** too long, but the reconciler's ratified user-pause safeguard suppresses `surface` on user-paused cards — so migrating it as-is would suppress a large part of what the sweep is for. `paused` and `userPaused` are distinct fields that can diverge (see `branch-group-ops.ts:128`). The sweep is untouched pending that decision. ## Verification - 34 tests green (10 new inheritance + 24 safety) - `tsc --noEmit` clean, `pnpm lint` clean No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Recovery decisions now correctly inherit operator settings when a workflow does not specify a recovery policy. - Workflow-specific recovery settings override inherited values, including when matching built-in defaults. - Recovery settings can now be applied independently by field, allowing thresholds and stale-item actions to inherit separately. - Recovery is suppressed safely when no complete policy is available, preventing unintended recovery actions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
3578903b16 |
U4: characterize surfaceStalePausedTodos before the policy migration (regression floor + safeguard evidence) (#2484)
Based on `main`, independent of #2482 — mergeable on its own. Pins what `surfaceStalePausedTodos` does **today**, so the first real policy migration is judged against *observed* behavior rather than against what the sweep looks like it should do. **No production code changed.** Passes against current main; must still pass after the migration. ## 1. Threshold source — the case you asked to see proven first The sweep reads `settings.stalePausedTodoThresholdMs` directly, so under the override layer an undeclared workflow must keep observing exactly that value. A migration that reaches for a declaration default instead finds the card fresh and **silently stops surfacing it** — resetting a deliberate operator choice with no error and nothing in any diff. Both directions are asserted, deliberately: - a customized **1h** threshold surfaces a 2h-old card; - the **same card is NOT surfaced** under the 24h built-in. Without that discriminator, "always surface" would pass the first test and prove nothing about which threshold was used. ## 2. User-paused cards — evidence for the open safeguard question This replaces my argument with a measurement. `getStalePausedTodoSignal` gates on `paused === true` and **never consults `userPaused`**, so a user-paused card that has sat too long **is surfaced today**. That test passes on current main. This is what blocks the migration: the reconciler's ratified user-pause safeguard suppresses the `surface` action for `userPaused` cards, so routing this sweep through it unchanged would **stop surfacing them** — a silent behavior change dropping a diagnostic operators rely on. `paused` and `userPaused` are distinct fields that can diverge (see the comment at `branch-group-ops.ts:128` about `userPaused` remaining true while legacy `paused` is false), so these are genuinely separable states rather than one condition spelled two ways. **The decision needed:** does the user-pause safeguard mean *"never act on a user-paused card"* or *"never mutate lifecycle state of one"*? It generalizes — `surfaceStalePausedReviews` and `surfaceInReviewStalls` are the same shape. ## 3. Things easiest to lose in a rewrite Also pinned: inert under `globalPause` and `enginePaused`; a non-positive threshold disables it entirely; an unpaused card is never surfaced; automation-paused and user-paused cards are not distinguished today. ## Verification 8 tests green against unmodified main; lint clean. No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
b133d521c4 |
U4 vertical slice: recovery-policy reconciler + ratified safety invariant (measured: engine ~780, real cost is a settings migration) (#2478)
Stacked on #2477. Base is `feature/workflow-vocabulary-u4-delete-dep-blocked` — do not merge before it. The smallest end-to-end slice of the U4 reshape, built to **measure** the real cost before committing to the full policy table. The survey's ~900-line reconciler figure was reasoned, not prototyped; this replaces it with numbers. **Everything here is additive and unwired. No behavior changes.** ## What lands | | lines | what | |---|---:|---| | `WorkflowColumnRecovery` (IR) | 42 | one key — `stalenessMs` + `onStale`. Optional and omitted when unset, so existing workflows serialize byte-identically. | | `recovery-reconciler.ts` | 176 | one engine: walks live cards, resolves each card's policy from **its own** workflow (per task, shared `irCache` — a 400-card board across three workflows reads three IRs), returns decisions. Decision and application are separate so the safety boundary is assertable without running an engine. | | `recovery-policy-safety.test.ts` | 156 | one-time. The **ratified invariant**. | ## Measured cost vs the ~900 estimate **Engine + IR types = 222 lines** for one action (`surface`) and one safeguard. Extrapolating the rest — `rebound` (target resolution, attempt budgets, backward-move proof, five more safeguards) ≈ +350, `archive` ≈ +50, the `budgets`/`dependencies` keys ≈ +150 — lands near **780**. So **~900 was a good estimate for the engine**, and the vertical slice does not move it much. That is the answer to the question asked. ## But the estimate's real miss is not lines **16 of the 34 POLICY sweeps read an operator setting today** — ~17 distinct policy-threshold keys, including `stalePausedTodoThresholdMs`, `inReviewStalledThresholdMs`, `taskStuckTimeoutMs`, `doneAutoArchiveDays`, `maxPostReviewFixes`. Moving those sweeps into workflow policy is **not a code refactor — it is a settings migration with operator-visible blast radius**, and it needs three decisions the line estimate never surfaced: 1. Does workflow policy **override** the global setting, or defer to it? 2. What happens to **existing projects** that already configured those settings? 3. Does an **unset** policy inherit the setting, or the built-in default? That is the gating question for the full table — not the reconciler's size. ## Why the sweep is not retired here Retiring `surfaceStalePausedTodos` requires builtin:coding to declare the policy **and** `stalePausedTodoThresholdMs` to migrate — or the behavior silently disappears for every existing project. That is the settings migration above, and it belongs behind its own decision rather than smuggled into a measurement slice. The reconciler is therefore **unwired — deliberately dead code**, for exactly as long as it takes to get that decision. ## The ratified safety invariant The six safeguards (user pause, `autoMerge:false`, dependency, capacity, merge-proof, at-most-once) live **outside** the policy table. A workflow must never be able to author a safety invariant away. Encoded two ways, because either alone is defeatable: - **structural** — the policy exposes only an allow-listed key set; adding a key requires editing the test and re-stating the safety argument (the friction is the point); - **behavioral** — a policy attempting every spelling of "ignore the user pause" has no effect. **Both halves mutation-verified**, because a safety test that cannot fail is worse than none: - making the reconciler honor a policy field that disables the user-pause safeguard → **fails** - adding an unreviewed key to the policy schema → **fails** A third test asserts the reconciler still **acts** on an unpaused card, so a reconciler that suppressed everything cannot pass by doing nothing. ## Scope limits stated rather than implied Only the `surface` action is implemented, so only its relevant safeguard is wired. `surface` mutates no lifecycle state; the other five gate lifecycle-**mutating** actions that do not exist yet, and wiring them now would be untestable dead code. A test records this so the absence reads as deliberate and must be updated when `rebound` lands. ## Verification - `tsc --noEmit` clean in core and engine; `pnpm lint` clean - merge gate green (299 + 10 + 71) - 23 safety tests green; `workflow-lifecycle-traits` green No changeset: `@fusion/core` and `@fusion/engine` are private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
2dce642ccc |
E2E validation: run a RENAMED-column workflow against a live engine (real graph + real PostgreSQL) (#2475)
Stacked on #2472 (`feature/workflow-vocabulary-b3-stranded-todo`). Test-only. No production file is touched. ## Why Every slice of this program has closed with the same caveat: *no renamed workflow was run against a live engine; all evidence is unit-level*. That caveat is load-bearing — eight times this session a test passed without exercising its subject. This PR removes it for the lifecycle spine. ## What actually runs `packages/engine/src/__tests__/workflow-lifecycle-live-e2e.pg.test.ts` drives the REAL pieces: - a **real PostgreSQL `TaskStore`** on a throwaway per-file database (shared PG harness; never the operator's DB, never port 4040), - the **real graph interpreter** (`WorkflowGraphTaskRunner`) with the **real column-boundary controller** wired to the **real `store.moveTask`** — all of its guards, traits, capacity reservation, and post-commit emission, - the **real scheduler release** (`runHoldReleaseSweep`), - the **real post-commit lifecycle bus** (`getWorkflowEventBus`), - the **real converted self-healing sweep** (`SelfHealingManager.recoverStrandedCompletedTodoTasks`, slice B3.1). Only the AI **seams** are scripted — the same boundary `testMode`/`mock` draws in production. **Assertion rule:** every lifecycle claim is asserted on **persisted state** (a fresh `getTask` with the store's task cache defeated, `run_audit_events` rows, `workflow_work_items` rows), never on "a function was called". The one spy — the event-bus subscriber — is asserted on the **received payload**, because the bus silently drops events that fail its shape check, so "emit was called" proves nothing. **Differential design:** the default-vocabulary (`todo`/`in-progress`/`in-review`/`done`) and renamed-vocabulary (`backlog`/`building`/`checking`/`shipped`) workflows come from ONE builder and differ ONLY in their four column ids. Any behavioral delta is attributable to the vocabulary alone. ## Coverage (9 tests, all green) | Scenario | What is proven | |---|---| | Default vocabulary, full spine | planning runs in the hold column, the card parks (graph does not self-promote), the **scheduler** performs hold→wip, the resumed run walks exec → review → merge-gate → end, persisted column is `done` | | **Renamed vocabulary, full spine** | identical, and no leg of the run touches any legacy column id | | Audit differential | the graph-owned boundary crossings are the same crossings node-for-node on both vocabularies; no legacy id appears in the renamed trail | | Event seam | a real subscriber **receives** a well-formed `TaskTransitioned` for the renamed `backlog`→`building` release and for the terminal move; `NodeEntered` arrives for every traversed node including `end` | | Crash / restart | exactly one durable continuation row at `exec`; a brand-new runner resumes from the row and the already-completed `planning` seam does **not** re-run; no duplicate continuation | | Converted sweep (B3.1) | a completed card in a **renamed** hold column is promoted (asserted on its persisted column), a card in the renamed **wip** column is not, and the default `todo` case still works | ## Mutation verification (both directions) Green suites are not evidence in this codebase, so both halves were falsified: 1. Keying `hold-release`'s `isHeldTask` on the `todo` literal → **5 of 6 spine tests fail, and the one that survives is the default-vocabulary one.** That is the exact signature the conversion program cares about. 2. Reverting slice B3.1's per-task hold-column resolution to the literal → **only the renamed stranded-todo test fails**; the default regression floor stays green. ## Findings surfaced by running it 1. **The IR validator refuses a `merge-blocker` column with no reachable merge-class node** ("the gate can never clear without one"). Kept rather than worked around — it means the review column here is genuinely gated. 2. **Entry into the merge region collapses to the legacy `merge` seam** (`MERGE_REGION_KINDS`), so a `merge-gate` node reaches the merge lane. Documented in the fixture. 3. **The transition policy refuses a direct hold → review move**, and it refuses it *workflow-resolved*: on the renamed board the only legal target is its own `building`, not `in-progress`. The recovery callback therefore promotes hold → wip → review rather than bypassing the policy. 4. **`moves.ts` still special-cases the `done` literal** (`if (toColumn === "done") clearNearDuplicateReferencesTo...`) after the post-commit emit. Not converted here and not in this PR's scope — flagged for the Phase B owner. ## Not driven end to end (stated plainly) - **Triage / specification.** The lifecycle starts from a task already bound to a workflow; `triage.ts` was not driven. The `planning` seam is scripted. - **Real merge.** No git worktree, no branch, no squash. `merge-gate` is pure policy; the `merge` seam is scripted. - **Lightweight / self-healing-off workflow.** The Tier 1 policy keys do not exist on this tip — there is no `policies` surface on the IR to set. Not drivable; not substituted with a unit test. - **Process-level crash.** The restart is an in-process one: a brand-new runner resuming from the persisted `workflow_work_items` row with no carried-over memory. No OS process was killed, so this proves durable-state resumption, not signal handling. ## Lane `.pg.test.ts` under the engine-default include glob, gated by `pgDescribe` so it skips cleanly with no PostgreSQL. The merge gate is untouched. Engine `tsc --noEmit` is clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added comprehensive live PostgreSQL workflow lifecycle coverage, including graph execution, suspension and resume, scheduler capacity release, crash recovery, and durable continuation. * Added validation for renamed workflow column configurations and columnless task movements. * Added event delivery checks for task transitions and node entry events. * Added self-healing recovery for stranded completed tasks in valid hold columns. * **Refactor** * Centralized workflow boundary handling, including task moves, continuation state, audit events, and diagnostics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
710d56b2db |
U4 trim: delete the dependency-blocked-todo feature (unreachable in production) and revert 5a2de7d (#2477)
Stacked on #2474. Base is `feature/workflow-vocabulary-u4-dead-code` — do not merge before it. Deletes an **entire feature that has never executed in production**, and reverts `5a2de7d`, which only threaded resolved lifecycle columns through it. ## Reachability evidence — the whole basis for this ``` surfaceDependencyBlockedTodos ← in NEITHER sweep registry; no caller in └─ getDependencyBlockedTodoReporter() engine/dashboard/cli — only tests └─ engine/dependency-blocked-todo-reporter.ts ← sole caller of ↓ └─ core/computeDependencyBlockedTodoReport ``` self-healing owns two name-based sweep registries (`runStartupRecovery`, 58 entries; `runMaintenance`, 76). `surfaceDependencyBlockedTodos` is in **neither**, so nothing ever invoked the chain below it. Its four tests passed while proving nothing about production. ## Why delete rather than wire it up Wiring was the tempting option and is the riskier one. Switching on a 450-line path that has never run — whose tests therefore establish nothing about its behavior against real data — is a **behavior change with unquantified blast radius**. This program already refused exactly that move for the **pool-id sentinel**, a one-line change that would switch on dormant enforcement across every project. This is the same class of move at ~450× the size. Deleting is also the recoverable direction: git keeps the feature, and it can be resurrected deliberately — with tests that prove it *runs* — if dependency-blocked reporting is actually wanted. ## The settings keys go with it `dependencyBlockedTodoReportEnabled` defaulted `true` while driving nothing. A schema/API-visible switch that lies about what the system does is worse than no switch. (It had no dashboard UI field — the dashboard test allowlist already recorded it as *"no UI field"*.) Four sibling tuning keys are removed with it. ## Against my own earlier work `5a2de7d` threaded resolved lifecycle roles into `computeDependencyBlockedTodoReport` and its reporter, answering a review finding I confirmed as real. **The code was correct; the impact claim was not**, because the path never executes. Neither the reviewer nor I checked *reachability* before agreeing the defect mattered — only correctness. A correction is posted on that thread in #2470. **Scope limit on that admission:** the same finding also described *incorrect scheduler ordering*. That half runs through `buildUnblockWeightMap` in `task-priority.ts`, which is **live** and was already threading `terminalColumns` (B1, `434b385`). Scheduler ordering was never affected, before or after. ## What survives `blocker-fanout.ts` **stays** — it is live via `task-priority.ts`. Only the plural `holdColumns` option added by `5a2de7d` is reverted, since the deleted report was its sole consumer. `holdColumn` (singular, from B1) remains. ## Net **1,244 deletions / 5 insertions across 15 files** — ~450 production lines, ~684 test lines, 5 settings keys. ## Verification - `tsc --noEmit` clean in **core, engine, and dashboard-app**; `pnpm lint` clean - merge gate green (299 + 10 + 71) - self-healing suite: 411 passed, 1 **pre-existing** failure (`archiveStaleDoneTasks`) - dashboard settings-descriptions suite green - `settings-parity.test.ts` has one **pre-existing** failure (`agentToolOutputMaxChars` overlap) that fails identically with these changes stashed — unrelated to this deletion No changeset: `@fusion/core` and `@fusion/engine` are private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added quiet-window backlog health diagnostics for stalled items in review, with repeat-alert suppression. * Added default thresholds for backlog-pressure alerts. * **Changes** * Removed dependency-blocked todo reporting and related alerts. * Removed the dependency-blocked todo enable/disable setting; remaining tuning options are no longer active. * Updated the workflow hold classification to use a single todo column. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
923a7c0fc9 |
U4 trim: delete the dead resetStepsIfWorkLost duplicate in self-healing (#2474)
Stacked on #2472 (Phase B slice B3.1). Base is `feature/workflow-vocabulary-b3-stranded-todo` — do not merge before it. First output of the **U4 reshape survey**: dead code removed with reachability evidence, not a redundancy argument. ## What is deleted **`SelfHealingManager.resetStepsIfWorkLost`** — 50 lines including its docblock and a section header left with no other member. Evidence: - `private`, with **zero in-file references** beyond its own declaration. - TypeScript already reported it as `declared but its value is never read` — the compiler has been flagging this. - The **live** implementation is `executor.ts:20080`, an independent copy that `executor.ts` actually calls (12667, 14664). The self-healing copy is an orphaned duplicate of it. - Not in either sweep registry, no test of its own, no caller anywhere in engine / dashboard / cli. Safe against every U4 constraint: it enforces none of user pause, `autoMerge:false`, dependency, capacity, merge-proof, or an at-most-once safeguard. ## Why the second approved deletion is NOT here `surfaceDependencyBlockedTodos` was approved alongside this one as a 28-line orphan. It is not. On inspection it is the **tip of an entire unreachable feature**: ``` surfaceDependencyBlockedTodos (in NEITHER registry, no production caller) └─ getDependencyBlockedTodoReporter() ← sole caller └─ engine/dependency-blocked-todo-reporter.ts 223 lines ← sole caller of ↓ └─ core/dependency-blocked-todo-report.ts 184 lines ``` ≈ **450 production lines across three files, plus 684 lines of tests in four files.** The operator-visible setting `dependencyBlockedTodoReportEnabled` defaults `true` in `settings-schema.ts` and drives nothing. Deleting only the approved 28-line tip would be **strictly worse than leaving it** — it orphans the getter and field and strands 407 lines of module with no remaining reference to explain why. Escalated for a decision (delete the subtree / wire the feature up / leave it) rather than resolved unilaterally. ## Verification - `tsc --noEmit` clean, `pnpm lint` clean - self-healing suite: **415 passed**, 1 failure **pre-existing** (`archiveStaleDoneTasks` — fails identically before this change) No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Behavior Changes** * Removed automatic detection and reset of task steps when no unique work is found on a task branch. * Tasks with completed or in-progress steps will no longer be automatically returned to pending based on this condition. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
a4dee1162e |
Phase B slice B3.1 (U4): resolve the hold column in recoverStrandedCompletedTodoTasks — query and guard together (#2472)
Stacked on #2471 (Phase B slice B2). Base is `feature/workflow-vocabulary-b2` — do not merge before it. **First landable slice of U4 (self-healing.ts).** One sweep, one PR, per the phase's sub-split rule. ## The finding: the guard and the query must convert together `recoverStrandedCompletedTodoTasks` promotes a card whose steps are all done/skipped but which is still sitting in the hold column — finished work that never handed off to review. It decided *"is this card in the hold column?"* **twice**, and both were literal: | | was | |---|---| | the QUERY | `listTasks({ column: "todo", slim: true })` | | the GUARD | `task.column !== "todo"` | **Either half alone is a green diff with zero behavior change.** A correct guard behind a literal query never runs; a converted query behind a literal guard rejects every row it just fetched. This is the shape that made B1's stale-paused-todo fix cosmetic, and the phase brief predicted more of it here — correctly. I proved it rather than asserting it: - literal **QUERY** restored (converted guard kept) → **3 tests fail** - literal **GUARD** restored (converted query kept) → **2 tests fail** Neither half passes the suite alone. ## Falsification came first Per the brief I tried to prove the work unnecessary before doing it. It is necessary, and the evidence is empirical, not assumed: the 7 tests were written against unmodified code and 3 failed. Unlike B2's hold-release — which turned out already converted — **self-healing is uniformly unconverted at the query level**: 53 of its sweeps carry a hardcoded `column:` filter (survey in the worker report). ## Negative half, per the brief A completed card resting in a WIP or review column is **not** promoted. Dropping a column filter without a per-task hold check would promote finished cards out of every column — laundering work past review, a louder bug than the silent one being fixed. ## Test-harness hazard (will recur in every remaining U4 slice) The pre-existing self-healing store mock returns its fixture from `listTasks` **regardless of arguments**. A renamed-hold test on that harness passes while the query stays hardcoded, because the mock hands the sweep rows the real store never would. The new harness **honors** the column filter, and one test asserts the query is no longer scoped to the literal. This is documented in the new file's header for whoever writes the next slice. ## Cost The column filter is gone, so the cheap non-column rejections (paused / executing / incomplete steps / errored / no-commits / skip-bypass taint) run **first and synchronously**; only survivors pay an IR resolution, shared through an `irCache`. A board spanning three workflows resolves three IRs regardless of card count. `includeArchived: false` preserves what the column filter did implicitly. The hold column resolves **per task** — a board spans workflows, and a card in *another* workflow's hold column must not be promoted. ## One pre-existing assertion changed, deliberately `self-healing.test.ts` pinned `listTasks` being called with `{ column: "todo", slim: true }`. That query shape changed on purpose; the assertion now pins the new one. The behavioral assertions either side of it (one qualifying card, promoted exactly once) are untouched and still pass. ## Carried A3 questions — both answered **Q1 — does the sync/SQLite counter have the same pool-id mismatch?** **Not applicable: there is no sync counter.** `occupantsByColumnForWorkflowImpl` and `listWorkflowOccupantTaskIds` are async/PG-only and throw without an initialized `AsyncDataLayer`; the sync twin went with the PG cutover. There is no second counter that could mismatch. The surviving pool-id sentinel sites are `project-store-ops.ts:767/819` and `moves.ts` — both parked by operator decision, untouched here. **Q2 — are custom workflows with an explicit numeric limit affected?** **No, by design.** `resolveColumnCapacity` gives `config.limit` top precedence (`configLimit` → `limitSetting` → default-workflow read-through → `Infinity`), and `resolveWipBudgetColumns` documents that a column with an explicit numeric limit is **independent — its budget is itself alone**. Such a column never pools, so there is no pool id to mismatch. Read-only analysis; no code changed for either question. ## Remaining U4 scope (not in this PR) 214 literal occurrences across ~70 methods; **53 sweeps carry a query-level column filter**. Hold-gated sweeps still to convert: `clearStaleBlockedBy`, `reclaimSelfOwnedBranchConflicts`, `reconcileCompletedTask`, `recoverMergedReviewTasks`, `recoverStuckMergeDeadlocks`, plus non-query `todo` guards in `recoverPausedAbortFailures`, `reconcileDependencyBlockingLeases`, and others. `surfaceStalePausedTodos` was already converted (B1 follow-up) and is verified intact on this branch. ## Verification - 7 new tests green; **both mutations kill the suite** - self-healing suite: 415 passed, **1 failure pre-existing** (`archiveStaleDoneTasks` — confirmed identical by stashing my changes) - merge gate green (299 + 10 + 71) - `tsc --noEmit` clean, `pnpm lint` clean No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved recovery of completed tasks stranded in workflow-specific hold columns, including renamed hold columns. * Preserved recovery for built-in workflows while correctly handling boards with mixed workflow configurations. * Prevented recovery for tasks in non-hold columns or with paused, incomplete, or errored states. * Added fallback handling when workflow details cannot be resolved. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
02b0f4f860 |
Phase B slice B2: U5 small movers — 12 literal sites converted, plus a negative result on hold-release (#2471)
Stacked on #2470 (Phase B slice B1). Base is `feature/workflow-vocabulary-conversion` — do not merge before it. ## What this is Phase B slice B2 — the U5 small movers. **12 literal sites converted, plus one negative result.** The plan estimated 36 sites. A survey found 12 genuinely-convertible ones, and separately found that the plan's headline hold-release scenario **was already fixed**. Both are reported below rather than padded into a bigger-looking diff. ## The negative result (commit 1) The plan named hold-release.ts as a target on the scenario *"release readiness must hold and release identically for a RENAMED hold column."* I wrote that test first, to prove it broken. **It is not broken.** All five assertions passed against unmodified `hold-release.ts`. U6/KTD-5 had already converted the module — `isHeldTask`, `resolveReleaseTarget`, and `dependencySatisfied` each resolve the task's IR. **`hold-release.ts` has no production change in this PR.** The tests are kept as a regression floor: the invariant rests on three independent trait resolutions any of which could be "simplified" back to a literal, and nothing else covered a renamed vocabulary end-to-end through the sweep. **I verified the tests can actually fail.** Mutating `isHeldTask` back to `task.column === "todo"` kills all five. Without that check, a green run against unmodified code is indistinguishable from a test asserting something trivially true. Two drafting notes kept in the file: the renamed ids deliberately avoid colliding with any legacy literal, and the first draft's two dependency tests used a `capacity` hold — which never consults dependencies at all, so one passed **vacuously**. Both now use a `dependency` hold. ## The 12 conversions Each was red-green: the renamed-workflow test written first and **observed failing**, then made to pass. | Site | Was | Now | |---|---|---| | `task-agent-sync` CLEAR_COLUMNS | `{done,archived,todo,triage}` | resolved complete+archived+hold+intake | | `task-agent-sync` isParkedTaskColumn | `{todo,triage}` | `parkedColumns` param (hold+intake) | | `task-agent-sync` handler branch | `to === "todo" \|\| "triage"` | resolved parked set | | `mesh-lease` parked guard | `task.column !== "todo"` | resolved rebound column | | `mesh-lease` rebound move | `moveTask(id,"todo")` | resolved rebound column | | `mesh-lease` audit decisionPath | `=== "todo" ? … : …` | same resolved column | | `mesh-lease` audit newColumn | `… : "todo"` | same resolved column | | `merger-ai` already-finalized | `=== "done" \|\| "archived"` | resolved complete+archived | | `merger-ai` ×4 rebounds | `moveTask(id,"todo")` | shared `resolveFinalizeReboundColumn` | Rebound targets all use KTD-10 `resolveReboundTarget` (hold → intake → first column), the helper `self-healing.ts:714` already uses — reused, not invented. ## Three findings worth reading **1. The mesh-lease bug was in the AUDIT, not the move.** The guard and the audit were *independent* `=== "todo"` comparisons, so `newColumn` asserted the card landed in `todo` regardless of what the move did. For a workflow with no `todo` column that produced a lease-recovery trail naming a nonexistent column — and run-audit is the only post-hoc record of a lease recovery. Now resolved once and threaded to both, so they are structurally incapable of disagreeing. **2. The merger-ai failure mode was not what I predicted.** I expected the already-finalized guard to fail open and re-merge a finished card. The red run showed it actually throws `Cannot merge FN-1: task is in 'shipped', must be in 'in-review'` — a hard error blaming the column, on a task whose real state is "already done". The thing preventing the re-merge is *itself* a literal in core's `getTaskMergeBlocker`, outside this slice. Two bugs coinciding, not a design. **3. A fourth site had to move that wasn't on the list.** `evaluateParkedAgentTaskLink` calls `isParkedTaskColumn` internally. Converting only the handler would have left the preservation branch on legacy ids after the caller resolved a renamed workflow — trading a stale-link bug for a **worse** dropped-link bug (a live agent's link cleared mid-run). ## Deliberately NOT converted Both keep their literals with the reason recorded at the site under a greppable `DELIBERATE-LITERAL` tag: - **`hold-release.ts:326` `legacyDependencySatisfied`** — the FN-5719 dual-accept half. Converting makes both halves compute the same answer, deleting the compatibility signal *and* its divergence detector while looking like a cleanup. - **`replan-target.ts` final fallback** — its value is precisely that it is *not* trait-resolved; resolving it against the workflow is the stranded-card bug it was written to fix. ⚠️ **The U12 literal ratchet does not exist in the tree yet.** The brief assumed an allowlist to add entries to; there is none. `grep -rn DELIBERATE-LITERAL packages/*/src` enumerates the sites it must admit. ## What I could NOT verify - **One of the four merger-ai rebound sites is untested.** The `landWorkspaceTask` rebound is verified by inspection and the shared resolver's unit tests only — `landWorkspaceTask` is only ever *mocked* (project-engine.test.ts), never executed. Covering it needs a multi-repo git fixture and a full land run. The **other three are genuinely exercised** by pre-existing merger-ai.test.ts (lines 676/716/777/895 assert `moveTask("FN-1","todo",…)` through a real git repo) and pass unchanged — real wiring proof for those. - **3 of the 9 task-agent-sync tests passed before the conversion too**, vacuously — the literal handler early-returned and cleared nothing. They assert nothing about the old code; they are guardrails against the conversion over-clearing. - **No renamed workflow was run against a live engine.** All evidence is unit-level. ## Call sites outside this slice — NOT converted, byte-identical They keep the legacy defaults: `scheduler.ts:1273`, `agent-heartbeat.ts:1169/3642`, `self-healing.ts:11600/11665` (all `evaluateParkedAgentTaskLink`), and `merger.ts:6585` (the sibling terminal guard). Each is its own Phase C/D surface. ## Behavior changes (not a pure refactor) For a **renamed** workflow: agent links now actually get cleared on terminal moves (they never were); lease rebounds land in the resolved hold column; finalize-blocked rebounds land in the resolved hold column and their operator-facing task-log lines name the real column; already-finalized cards short-circuit cleanly instead of throwing. For **builtin:coding** and any unresolvable workflow: byte-identical. Every new parameter defaults to the legacy set, and both merger-ai resolvers fail *soft* to legacy ids in opposite directions — the terminal guard keeps `done`/`archived` (losing it sends a finished card into the merge path), the rebound keeps `todo` (abandoning it strands the card in the merge lane with no owner). ## Verification - Merge gate **green**: 299 + 10 + 71 tests - Slice suites **green**: 100 tests across 8 files (new + all pre-existing neighbours) - Existing merger suites **green**: 82 tests across 5 files, unchanged - `tsc --noEmit` clean, `pnpm lint` clean No changeset: `@fusion/engine` is private. 🤖 Generated with [Claude Code](https://claude.com/claude-code) |
||
|
|
5d0f1ef631 |
Phase B slice B1: lifecycle column roles in the U6 policy modules (4 guards, red-green) (#2479)
**Stacked on #2469** → #2468 → #2467. Base is `feature/workflow-capacity-ground-truth`. This is **slice B1 of Phase B, not all of Phase B.** Sizing escalation sent separately; the census is below. ## Why this is a slice Measured census of code lines referencing a lifecycle column literal (comments excluded): | Unit | Files | Sites | |---|---|---:| | U4 | `self-healing.ts` | 203 | | U5 | `executor.ts` 171, `scheduler.ts` 55, `replan-target.ts` 20, `merger-ai.ts` 5, `hold-release.ts` 4, `mesh-lease-manager.ts` 4, `task-agent-sync.ts` 3 | 262 | | U6 | `moves.ts` 34, `default-workflow-hooks.ts` 13, `board-config.ts` 9, `blocker-fanout.ts` 6, `task-priority.ts` 5, `dependency-blocked-todo-report.ts` 2, `stale-paused-todo.ts` 1 | 70 | | | **Total** | **535** | The plan's "~207" counts the guard category only. Under the phase's non-negotiable rule — a test that **fails before** conversion, per guard — that is ~200 red-green cycles. Doing it as one sweep would reproduce exactly the failure this phase exists to prevent: converted guards nobody proved still fire. `moves.ts` and `default-workflow-hooks.ts` stay **parked** per the dispatch constraint (move-path convergence and the pool-id sentinel are on an operator decision). ## Guards converted (4), each red-green Every case below was written **first** and observed failing against the literal implementation. | Module | Guard | Before → After | |---|---|---| | `stale-paused-todo.ts` | stall detection | `column !== "todo"` → resolved **hold** column | | `blocker-fanout.ts` | active | `ACTIVE_COLUMNS.has(col)` → `!terminalColumns.has(col)` | | `blocker-fanout.ts` | hold-wait metric | `col === "todo"` → resolved **hold** column | | `task-priority.ts` | unblock active | `UNBLOCK_ACTIVE_COLUMNS` **deleted**, folded into the terminal set | Three of the seven new cases are **regression floors** that pass before and after. One of them earned its keep immediately: it failed on my own fixture (`activeCount` vs the public `totalCount`), catching a bad test rather than bad code — which is the point of asserting the default path alongside the renamed one. ### The `task-priority` finding `UNBLOCK_ACTIVE_COLUMNS` and `DONE_COLUMNS` encoded **one concept twice**, two lines apart, and disagreed for any custom column: dependency counting treated a `drafting` card as unmet (correct) while the active check treated it as inactive (wrong), zeroing the blocker's unblock weight. The enumeration wasn't just legacy-shaped — it contradicted its own neighbour. ## ⚠️ Behavior change, not a pure refactor Inverting active from enumeration to exclusion means **a card in a column that is neither terminal nor in the legacy enum now counts as active where it previously did not.** That is the plan's stated intent, but it is a real change for any project already using a custom column — **Coding (Ideas)' `ideas` column is the in-tree case.** Fan-out counts and unblock weights for such cards will rise. ## Verification - Four affected suites green (45 tests), each conversion observed red→green. - `pnpm lint`, `tsc --noEmit` (core) green. **Not verified / not done, stated plainly:** - **Call sites are not wired.** These modules now *accept* resolved roles; every parameter still defaults to the legacy set, so at the call sites the vocabulary is unchanged. A caller that cannot resolve a workflow keeps literal behavior. Threading `resolveLifecycleColumns` through `reads.ts` and `self-healing.ts` is follow-on work — until then the guards are *convertible*, not *converted end-to-end*. - `dependency-blocked-todo-report.ts` and `board-config.ts` are untouched in this slice. - 19 core-suite failures exist on this branch; all confirmed **pre-existing** by stashing and re-running on a clean tree (`duplicate-guard`, `log-severity-spam-contract`, `settings-parity`, `task-delete-caller-attribution`, `settings-defaults`). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --- **Supersedes #2470**, which GitHub force-closed when its base branch was deleted by the merge of #2469 and refuses to reopen. Same head branch, same commits (rebased onto `main`), now based on `main` directly. The two P1 review threads on #2470 were resolved there — one of them with a correction noting the threading half landed in code that was subsequently deleted as a dead feature in #2477. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Dependency and blocker reports now correctly recognize custom hold, active, and terminal workflow columns. * Blockers in renamed terminal columns are no longer incorrectly reported as active. * Stale paused-task badges and self-healing now work with workflow-specific hold columns. * Mixed boards with different workflow column names are handled consistently. * Existing default workflow behavior remains compatible, including fallback handling when workflow details cannot be resolved. * **Enhancements** * Reporting and task-priority calculations now support configurable single or multiple hold and terminal columns. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
4158cf1ab7 |
Phase A: workflow-owned lifecycle foundation (U1, U2, U3) (#2467)
Phase A (Foundation) of
`docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md`.
Three units, one commit each. No operator-visible behavior change.
## U1 — Lifecycle-column resolution seam
`resolveLifecycleColumns(ir)` returns `{ intake, hold, wip, review,
complete, archived }` — the first column carrying each trait,
`undefined` for a role no column carries.
`resolveTaskLifecycleColumns(store, taskId, cache?)` is the store-aware
form; the cache is caller-owned so a sweep reads one IR per workflow
rather than one per card.
A v1/column-less IR resolves to `undefined` for the **whole struct**
rather than a struct of undefined roles. A caller must be able to
distinguish "this workflow declares no hold column" (a real shape to
honor) from "no column vocabulary at all" (skip and log) — only the
second licenses conservative fallback.
Nothing consumes the seam yet; Phases B–D convert the ~207 hardcoded
column literals onto it.
## U2 — Delete the pre-cutover parity machinery (delete-only)
**`workflow-columns-settings.ts`** — `isWorkflowColumnsEnabled` had the
body `return true`. Six live call sites branched on it, so every
flag-OFF arm was dead code that read as a supported configuration.
Deleted; surviving side inlined at self-healing's transitionPending
sweep, the scheduler's per-column capacity diagnostic, merge-trait's
policy resolver, the board-workflows payload, two task-workflow routes,
and the CLI TUI's column enrichment.
**`workflow-parity.ts`** — asserted the default workflow's adjacency
*equals* the legacy `VALID_TRANSITIONS`. U11 deliberately breaks that
equality by merging Todo into Planning, so this is not a stale assertion
to update; it is a contract against the target state. Its emitter
(`workflow-parity-observer.ts`) is already a tombstone, so
`getWorkflowParitySummary` and `computeWorkflowColumnsGraduationReport`
aggregated run-audit rows nothing writes and had no caller outside
`TaskStore`. Both store methods go with it.
`flagEnabled` stays on the board-workflows **wire** as a constant `true`
— shipped dashboard clients still branch on it, and changing the
response shape is not a deletion. U10 retires the field once no client
reads it.
The `legacy-tombstones` ratchet is extended to both files plus seven
symbols, each with the reason it is gone.
### ⚠️ Finding: the third listed deletion was NOT dead
The plan also lists "the flag-off inline move path" in
`task-store/moves.ts`. It is **not** deleted, per U2's execution note
("any behavior change found while removing a branch means the branch was
not dead").
That path is gated on `isWorkflowColumnsCompatibilityFlagEnabled`
(`store.ts:38`) — a **different** function from the always-true public
helper. It reads the raw `experimentalFeatures.workflowColumns` setting,
which nothing in production sets (`settings-schema.ts:396` — "no default
flags are emitted"; zero non-test writers; the operator's own
`~/.fusion/settings.json` has no such key). So `useWorkflow` is false
for effectively every real project: the flag-OFF inline side effects are
the **live** default move path and the flag-ON `default-workflow-hooks`
path is the dead one. The code says so itself at `moves.ts:638`.
Deleting that branch would swap every project onto an untravelled code
path — a behavior change, not a deletion.
**Carry this into Phases B and C, stated plainly so the plan's error is
not repeated:**
> **The inline move path in `moves.ts` is LIVE.
`default-workflow-hooks.ts` (the trait-hook path) is DEAD.** KTD-6
asserted the inverse. Until the convergence unit lands, **nothing may
assume trait hooks run** — a guard, sweep, or subscriber written against
`applyDefaultWorkflowMoveEffects` would never fire in production and
would still pass its tests.
Convergence is **not** attempted here. It is its own unit (Phase A2)
with a proper equivalence proof, per operator decision.
### U3's emit point is on the LIVE path — the seam is not born dead
Worth stating explicitly because it is the failure mode that would make
every later subscriber silently never fire: the `TaskTransitioned` emit
is **not** inside the `if (useWorkflow)` branch. That block closes at
`moves.ts:1212`; the emit sits at `:1214`, beside the existing
`store.emit("task:moved", …)`, on the unconditional post-commit path. It
therefore fires on **both** the live inline path and the dead hooks
path, and the convergence unit inherits the obligation to keep it firing
on whichever path survives — same events, same order, same payloads.
The graph-side emitters (`NodeEntered`, `RunSuspended`) carry the same
risk from a different direction: the bus refuses an invalid payload
*silently* by design, so an emitter regression would stop the event with
no test failure. They are asserted end-to-end through the real bus —
"did a subscriber actually receive it", not "was emit called" — because
a spy passes on a refused payload. The `moveTaskInternalImpl` emit does
**not** yet have that end-to-end assertion against a real store move;
that proof belongs to the convergence unit, which has to build the
both-paths fixture anyway.
## U3 — Post-commit event seam with a transactional outbox
**The bus is not a queue, not a transaction participant, and not a
delivery guarantee.** Durable follow-on work uses the transactional
outbox — a `workflow_work_items` row written *inside* the transition
transaction (the shape `createCompletionHandoffWorkflowWork` already
uses). "Emit after commit, let a subscriber enqueue the work" has a
crash window where a process dies between commit and subscriber, leaving
no event *and* no work-item row, so required work is skipped permanently
with nothing to recover from. Post-commit subscribers therefore carry
only losable reactions.
Emission is consequently lossy and isolated by design: a throwing or
rejecting subscriber is caught and logged, cannot roll back the
transition, and cannot stop the others. Deliveries append to one serial
chain, so two transitions on a task deliver in commit order.
The ids/outcomes-only rule is **mechanised, not documented** —
run-audit's equivalent lives only in prose and has been violated
repeatedly. A payload carrying an object body or a prose string is
refused at the emit boundary and never reaches a subscriber or log sink.
It degrades rather than throws: the emitter is post-commit, so a shape
bug must not become a lifecycle failure.
Emit points: `TaskTransitioned` from the single post-commit point in
`moveTaskInternalImpl`; `NodeEntered` and `RunSuspended` from the graph
column boundary, the latter *after* the durable continuation is
persisted so an observed suspension implies a resumable run.
`registerWorkflowEventSubscribers` (engine) is empty on purpose —
U7/U8/U10 move real reactions onto it, each with the characterization
test proving the reaction was non-authoritative first.
## Verification
- `pnpm test:gate` — green (2/10, 16/299, 1/71).
- `pnpm lint`, `pnpm build`, `tsc --noEmit` on core and engine — green.
- U1: 20 tests in `workflow-lifecycle-traits.test.ts`, including the
fully-renamed-workflow case (fails if the resolver falls back to a
literal) and a shared-cache read-count assertion.
- U2: `legacy-tombstones.test.ts` green with the extended ratchet;
`board-workflows`, `merge-trait`, `workflow-graph-executor-parity`, and
move-hook suites green with no expectation edits.
- U3: 20 bus-invariant unit tests (isolation, ordering, the allowed-key
and required-key halves of the ids-only rule, lossiness) plus 3
end-to-end emitter-delivery tests; 5 outbox tests against a **real
PostgreSQL** work-item table (crash survival, rollback, at-least-once
redelivery on lease expiry, idempotent handler → one effect,
dropped-subscriber vs. durable work). A hand-written fake of the lease
predicate would only prove the fake redelivers.
**Not verified:** the `moveTaskInternalImpl` emit is confirmed on the
unconditional post-commit path by structure and by the surrounding
tests, but is *not* yet asserted end-to-end against a real store move on
both flag settings — that is Phase A2's fixture. The engine subscriber
registry ships empty by design, so no production subscriber exercises
the bus end-to-end yet. `settings-defaults.test.ts` has one pre-existing
failure on `main` (a logger-prefix mismatch in the
`mergeIntegrationWorktree=cwd-main` warning) — confirmed present on a
clean tree, unrelated to this branch.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Workflow lifecycle columns are now derived from workflow definitions,
supporting renamed and custom workflows.
* Added post-commit lifecycle events for task transitions, node entry,
and run suspend/resume with validated payloads.
* Follow-on processing for lifecycle emissions is now more robust
(rollback-safe, at-least-once delivery, idempotent handling).
* **Bug Fixes**
* Workflow board responses, task enrichment, and promotion no longer
depend on workflow-columns feature-flag gating.
* Subscriber failures no longer impact committed workflow transitions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
|
||
|
|
8b039a543e |
fix(desktop): advance Pi runtime pin to 0.82.1 for packaging PR lane (#2465)
## Summary - Advance the matched Pi runtime pin (`pi-ai`, `pi-coding-agent`, `pi-agent-core`, `pi-tui`) from **0.82.0 → 0.82.1** so electron-builder's production-dependency walk accepts `pi-agent-core`'s `pi-ai@^0.82.1` requirement. - Fixes the Desktop packaging PR-lane failure: `Production dependency @earendil-works/pi-ai not found for package @earendil-works/pi-agent-core` (required `^0.82.1`). - Keep the workspace override guard; update pin-policy fixtures and CLI package-config expectations. - Tighten the advisory packaging step-order test so it asserts against the real `electron-builder --dir` step (not a missing release-only step name that previously passed via `indexOf === -1`). - Run `pnpm dedupe` so the packaging lane's lockfile dedupe early-warning is clean. ## Context #2439 pinned the full Pi closure at 0.82.0 and made recent main-based packaging runs green. This advances to the current upstream patch so deploy + electron-builder stay aligned with `pi-agent-core@0.82.1`'s declared dependency range. ## Test plan - [x] `node scripts/check-pi-versions-pinned.mjs` - [x] `node --test scripts/__tests__/check-pi-versions-pinned.test.mjs` - [x] `pnpm --filter @runfusion/fusion exec vitest run src/__tests__/package-config.test.ts` - [x] `pnpm --filter @fusion/desktop exec vitest run src/__tests__/release-workflow.test.ts` - [x] `pnpm dedupe --check` - [ ] GitHub: Desktop packaging (should run full packaging walk — lockfile/package.json touched) - [ ] GitHub: PR Checks (Lint, Typecheck, Build, Gate) |
||
|
|
99c9f14ee0 |
feat: run Plan Review in the planning lane with a Plan Review badge (#2462)
## What Plan Review, planning, and the replan loop move from the implementation column into the **planning lane** (`todo`), so a task under specification never holds a WIP slot. The card crosses into `in-progress` exactly once, at `parse`, released by the scheduler. Operators also finally see a **Plan Review** badge while the gate runs — it was previously invisible on the default workflow. ## The part that made it possible Moving the node is ten lines. It was attempted three times and reverted each time, because a graph run with no durable continuation replayed from `start` and dragged an in-progress card *backward* out of the WIP column, firing `abort-on-exit` and stranding it in a pre-WIP column with no releaser. So this PR adds the graph **entry contract** — `resolveColumnResumeNode`: | Card is in | Resumes at | |---|---| | `triage` | `start` | | `todo` | `plan` | | `in-progress` | `parse` — never re-plans, never moves backward | | `in-review` | first review node — gates are not skipped | `ir.columns` is ordered and that order is the lifecycle order; rework and failure edges are excluded so the entry point is always the main path. The proof it's the right fix: **`executor-task-done-invariant` passes unmodified** after failing every previous attempt. ## Also in here - **Release gate narrowed twice.** `isUnplannedForExecution` applies its pre-release plan-review gate only when the node's column equals the card's column *and* the group is enabled for the task. The enablement check fixes a real deadlock — a task with Plan Review toggled off was held forever waiting for evidence nothing would ever write. - **Badge cleanup.** Gate badge reads "Plan Review" instead of the ambiguous "Reviewing" and no longer hides behind a lane restriction; the status badge stops duplicating it; `planning` renders as "Planning" instead of the raw engine token. - **Coding (Ideas)** renames its planner column to "Planning" (id `todo` unchanged) and loses its private planning-node re-home — the graph it clones is already plan-in-place. - **New sweep** `reconcileUndeclaredTaskColumns` re-homes a row whose column its workflow no longer declares. Written for a follow-up, kept because it makes any column edit survivable. ## Test changes Scheduler and release fixtures now model a card whose Plan Review passed — the state every real card is in when the capacity sweep sees it. A held unreviewed card is the gate working, and that path stays owned by `pre-release-plan-review.test.ts`. New `workflow-graph-entry-contract.test.ts` covers the invariant at every lifecycle position, plus the gap-column and remediation-node cases. ## Verification Gate 299 + 70 + 10, dashboard badge suites 672, engine workflow/entry/executor suites 147, core 122. Lint and typecheck clean. Full engine suite sits at the pre-existing baseline (notifier / plugin-runner / notification-service, untouched by this). ## Follow-up Removing the Todo column entirely is a separate ~207-site lifecycle-vocabulary refactor — planned in `docs/plans/2026-07-26-001-refactor-workflow-owned-lifecycle-plan.md` (companion docs PR). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Plan Review now runs in the Planning lane before implementation begins. * Cards resume from their current workflow column without replaying earlier steps. * Added automatic recovery for cards stranded in outdated workflow columns. * **Improvements** * Renamed the Coding (Ideas) planner column to “Planning.” * Refined Plan Review gating to respect enabled settings and the card’s current column. * Updated planning and Plan Review badges for clearer, consistent labels across cards and lists. * **Bug Fixes** * Improved workflow transitions and release behavior around planning, review, and execution. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5ae6332563 |
refactor: collapse dead SQLite dual-path code; keep migration-only readers (#2454)
# Remove dead SQLite dual-path code; keep migration-only readers ## Summary PostgreSQL cutover left hundreds of production dual-path branches (`backendMode ? PG : SQLite/store.db`) whose SQLite arms only hit throwing `Database`/`ArchiveDatabase`/`CentralDatabase` stubs. This change mechanically collapses those unreachable arms so production authority is AsyncDataLayer/PostgreSQL only, while preserving the six authorized read-only migration/recovery `DatabaseSync` seams. ## Dual-path mass removed | Metric | Before | After | |---|---|---| | `if (…backendMode)` (non-test) | ~328 | ~70 | | `store.db` / `this.db` refs in core (non-test) | ~570+ | ~375 (mostly pure legacy MissionStore/eval/insight SQLite classes + thin getters) | | Net diff | — | **~6.7k lines removed** across 41 files | Remaining `backendMode` checks are intentional (incomplete-PG sync safe-defaults, settings-sync disabled-on-PG, symbol-lock PG-only gates, “requires PostgreSQL” config versioning throws), not live SQLite authority. ## Subsystems cleaned - **Core TaskStore / task-store/***: collapsed if/else and early-return dual-path across reads, moves, lifecycle, mutations, workflow, archive, branch/PR, artifacts, comments, audit, project ops, etc. `initImpl` is PostgreSQL-only (SQLite startup tail deleted). - **Satellite stores**: automation, agent, routine, plugin, secrets, approval-request, central-core dual-path arms collapsed. - **Plugins**: reports async methods, compound-engineering pipeline + session stores, CLI Printing Press store — SQLite fallbacks removed; PG required. - **Engine**: no functional dual-path change beyond whitespace (settings-sync / peer-exchange PG-disabled behavior kept). ## Six migration-only readers retained (allowlist unchanged) 1. `packages/core/src/postgres/sqlite-migrator.ts` 2. `packages/core/src/project-identity.ts` 3. `packages/core/src/sqlite-validation.ts` 4. `packages/core/src/postgres/startup-factory.ts` 5. `packages/cli/src/commands/db.ts` 6. `scripts/lib/start-local-project.mjs` Plus low-level `sqlite-adapter` and migrator/startup-import tests. Inventory ratchet still requires exactly these six `new DatabaseSync(` production sites, all `readOnly: true`. ## Not treated as SQLite - `.fusion/project.json`, `task.json`, `agent-log.jsonl` file storage - AsyncDataLayer / Drizzle PG paths - Incomplete-PG sync safe-default stubs (still return empty/false/null under backend without consulting SQLite) ## Verification - `sqlite-production-reader-inventory.test.ts` — 15/15 pass - `incomplete-pg-ports.pg.test.ts` — 6/6 pass - Targeted PG tests (create-task, move, handoff, runtime-persistence, agent, mission, insight, central-core) — green - `tsc --noEmit` for `@fusion/core`, `@fusion/engine`, `@fusion/dashboard` — green - `scripts/check-no-getdatabase.mjs` — clean <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Improved end-to-end consistency by making PostgreSQL/async persistence the standard across core task/workflow, automation, agents, plugins, routines, secrets, approvals, central operations, and session storage. * Unified scheduling, settings, configuration revision writes, run/workflow selection, queues/leases/transitions, and audit/lifecycle updates around consistent async transaction behavior. * **Bug Fixes** * Fixed edge cases for archived/deleted reads, unarchive/recovery flows, not-found handling, and task/artifact/document/log/comment operations, including more reliable emissions and hydration across search/list and lifecycle operations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
52d64fa66e |
fix(engine): project CE steps after review handoff (#2464)
## Summary - reconcile successful graph-native workflow results with pending task checklist steps even when review handoff already moved the card into the merge column - preserve terminal, paused, and no-redundant-move behavior - cover the real Compound Engineering post-review-handoff state with a regression test ## Root cause Compound Engineering runs `review-handoff` before `merge`. Review handoff moves the task to `in-review`, which is also the merge column. `ensureWorkflowMergeBoundaryTask()` returned immediately for cards already in that column, before projecting successful `workflowStepResults` onto legacy `Task.steps[]`. The merger then saw `0/N` and rejected approved work with `task has incomplete steps`. ## Verification - RED: regression test failed before the fix because `store.updateTask` was never called - GREEN: `executor-graph-boundary.test.ts` — 6 passed - relevant non-PostgreSQL set — 31 passed, 5 PostgreSQL tests explicitly skipped - `@fusion/engine` typecheck passed - changeset format passed - `git diff --check` passed ## Baseline note `ce-workflow-step-executor.test.ts` currently has three failures on clean `origin/main` after FN-8601 foreach-proof hardening. The same failures reproduce without this patch and are not regressions from this change. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved reconciliation after review handoff by projecting completed step results onto the legacy checklist when reaching the merge column. * Prevented tasks from being marked approved with incomplete step counts (including “0/N” style states). * Reduced unnecessary merge failures and deadlock/pause scenarios when merge-column progress was already recorded. * **Tests** * Added coverage for execute-and-merge workflows, ensuring merge-boundary resolution updates pending steps without moving the task. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
0e3d2a2265 |
refactor: delete meta-task auto-archive and automated recovery follow-ups (#2461)
Deletes two pieces of automated "meta" machinery that filed and garbage-collected cards restating state already on the task that failed. Net **-1015 lines**. ## Why **Automated recovery follow-ups.** `createAutomatedFollowup` and its dedup engine (289 lines of signature matching, 1h recurrence rate-limiting, 24h supersedes windows) existed to file recovery cards for verification-cap and merge-conflict give-ups. In both cases the parent is *already* parked `failed` with a descriptive `error` and a log entry carrying the failing command, branch, and output — the card was a second copy of that. **Meta-task auto-archive.** The sweeps that garbage-collected those cards were worse than redundant: the regex classifier matched ordinary feature work, and its positional fallback bound cards to unrelated tasks, so **live work could be archived**. They are removed together, because the auto-archive sweeps only existed to clean up after the follow-up engine. ## What changed ### Deleted - `packages/engine/src/verification-followup-dedup.ts` in full — `createAutomatedFollowup`, `decideAutomatedFollowup`, `AutomatedFollowupKind`, `computeVerificationFailureSignature`, `extractFailingTestFiles`. - `findActiveRecoveryFollowUp` — dead code, defined and never called (`tsc` independently flagged it `6133 declared but its value is never read`). - The meta-task auto-archive sweeps `autoArchiveResolvedMetaTasks` / `autoArchiveStalledMetaTasks` and helpers `classifyMetaTask` / `resolveMetaTargetTaskId` / `computeMetaChainDepth` / `archiveMetaTask` / `evaluateMetaAutoArchiveGuards`, plus settings `metaTaskStallAutoCloseMs` and `metaTaskActiveExecutionGraceMs`. - Run-audit types `task:auto-archived-meta-resolved`, `task:auto-archived-meta-stalled`, `task:auto-archive-meta-resolved-skipped`, `task:auto-archive-meta-stalled-skipped`, `verification:followup-created`, `verification:followup-deduped`. The two signature helpers were **deleted rather than relocated** — once the three call sites went they were provably unreachable: `buildVerificationFailureSignature` had exactly one caller, and it was the only caller of `extractFailingTestFiles`. ### Call sites 1 and 2 — park kept, card dropped Verification-cap and merge-conflict give-ups keep their park, audit event, operator comment, and log entry. Site 1's `error` string was reworded off `"See follow-up task for investigation."` (no follow-up will exist) to carry the guidance itself. `autoResolveDisabled` was **kept** — it still drives the outer park guard and the `reason` string; only the inner branch that guarded card creation is gone. ### Call site 3 — autostash orphan, replaced not deleted This one is a genuine data-loss guard, so it keeps a durable trail. A `live`-classified orphan is a merger stash holding **real uncommitted work**, and unlike sites 1–2 there is no parked parent — the parent may already be `done` and merged, so nothing else on the board would ever mention the stash. The card is replaced by a `logEntry` **and** an `addTaskComment` on the parent, preserving every fact the old description carried: the sha, `record.label` (the handle `git stash` recovery needs), `record.detectedByTaskId`, and `sourcePhase`. New truthful run-audit event `task:autostash-orphan-live-detected` replaces the borrowed `verification:followup-*` name, with ids/outcomes-only metadata per AGENTS.md. ### Kept unchanged: the two real product features Eval follow-ups (`eval-followups.ts`) and PR-comment follow-ups (`pr-comment-handler.ts`) only borrowed the shared engine for its dedup pass. Both keep their exact behavior, column, priority, `sourceType`, and log lines, with dedup inlined as a `listTasks` scan on `suggestionId` / `prNumber` respectively. Both fail open (create) if the listing throws, matching the old engine. ## Test changes — read this one Two tests asserted the *deleted* engine's rate-limited `"[verification recurrence]"` logEntry. Those assertions were removed, **not loosened**: both tests still assert no duplicate card is created, and the eval test still asserts the existing id is reported back. No coverage of surviving behavior was weakened. The three `meta-*` test files were deleted along with the sweeps they covered. ## Verification ``` $ pnpm test:gate Test Files 2 passed (2) Tests 10 passed (10) # core Test Files 16 passed (16) Tests 299 passed (299) # engine-core Test Files 1 passed (1) Tests 70 passed (70) # ci-shape GATE_EXIT=0 $ pnpm --filter @fusion/engine --filter @fusion/core exec tsc --noEmit -p tsconfig.json TSC_EXIT=0 (no output) ``` Plus a file-scoped run over the touched surfaces (`eval-followups`, `pr-comment-handler`, `merger-autostash-orphan-surface`, `merger-autostash-cleanup`, `run-audit`, `run-audit-secret-taxonomy`, `project-engine`, `project-engine-manager`): **213/213 passed**. A repo-wide grep confirms no surviving references to any deleted symbol, module, or audit event. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Failed tasks now retain recovery and verification details directly on the original task instead of generating separate follow-up cards. * Live autostash issues now preserve stash information in task comments and activity logs. * Existing evaluation and pull-request follow-ups continue to be reused when appropriate. * **Changes** * Removed automatic archival of meta-tasks. * Removed obsolete meta-task timing settings. * **Documentation** * Updated architecture and settings documentation to reflect these workflow changes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
232e17d1cb |
test(engine): complete runtime logger mock (#2458)
## Summary - add the missing `debug` method to the runtime-resolution logger mock - prevent logger calls from short-circuiting runtime selection assertions ## Test plan - `pnpm --filter @fusion/engine exec vitest run src/__tests__/runtime-resolution.test.ts` - `pnpm --filter @fusion/engine typecheck` - `pnpm check:changesets` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Updated the runtime-resolution test suite’s mocked logger to also support debug-level messages, alongside existing log, warn, and error handling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> |
||
|
|
00778cb10f |
test(engine): refresh shellout allowlist (#2451)
## Summary - refresh the engine synchronous-shellout allowlist after recent self-healing and executor source additions shifted audited call sites - keep the guard's path, primitive, and signature checks unchanged ## Test plan - `corepack pnpm --filter @fusion/engine exec vitest run src/__tests__/engine-no-blocking-shellout.test.ts --silent=passed-only --reporter=dot` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Updated static validation allowlists for audited synchronous shell command operations so matching stays accurate with the latest call-site locations. * Kept safeguards that prevent unapproved blocking shell commands from passing validation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: gsxdsm <gsxdsm@users.noreply.github.com> |
||
|
|
256c64a7bd |
chore(release): v0.74.0-beta.5
Version bump via changesets. |
||
|
|
beebd270bd |
fix: make Queued to plan / Ready badges agree with the planning lane
TaskCard inferred "unplanned" from steps.length === 0 while triage's todo-discovery and the scheduler's dispatch filter both decide from PROMPT.md seed-ness, so the badges disagreed with the engine in both directions: a real spec that parsed to zero steps read as "Queued to plan" while the scheduler already treated it as a WIP-slot candidate, and a re-seeded card still carrying old steps read as "Ready" while triage was about to plan it. Either way the badge sent operators to the wrong cap. Adds the shared isTaskAwaitingPlanning predicate (replan park, missing spec, seed-vs-real content) used by both triage's discovery and a new best-effort `awaitingPlanning` enrichment on GET /api/tasks. TaskCard derives both badges from that one value — strict complements — and keeps the step count only as a fallback for SSE payloads and older servers. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5ea98f7d4b |
fix: release admission claims when triage evicts a hung planner
evictStaleProcessing cleared `processing` but left the task in `coordinatorAdmittedTaskIds`, which is only cleared by specifyTask's finally — the path a hung promise never reaches. The card stayed eligible (so the throttle branch never logged or emitted `task:plan-admission-throttled`) while admitOldest's refresh filtered it out, leaving it on the "Queued to plan" badge with free slots and no diagnostic until engine restart. Also drop an untransferred pre-held host slot, which otherwise waits out the 600s stale-excess valve. Regression tests assert the invariant on the real production candidate source: an evicted card is re-offered and its host slot returned, while a still-live stale task keeps both claims. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0022621d22 |
chore(release): v0.74.0-beta.4
Version bump via changesets. |
||
|
|
2bb8537352 |
FN-8616: make agent tool-output limits configurable
Expose the shared agent tool-output budget as a scoped operator setting with an explicit no-limit option. - Resolve global and project output caps with a safe finite default and zero sentinel. - Propagate configured budgets through PI and plugin runtime tool wrappers. - Add settings controls, localized labels, documentation, and regression coverage. Files changed: .changeset/fn-8616-tool-output-budget-setting.md | 7 ++++ docs/agents.md | 4 +- docs/settings-reference.md | 1 + .../core/src/__tests__/tool-output-budget.test.ts | 23 ++++++++--- packages/core/src/index.gate.ts | 2 + packages/core/src/index.ts | 2 + packages/core/src/settings-schema.ts | 12 ++++++ packages/core/src/tool-output-budget.ts | 31 +++++++++++++-- packages/core/src/types/settings-scope.ts | 8 ++++ .../app/components/settings/save-split.ts | 1 + .../sections/GlobalGeneralSection.search.ts | 20 ++++++++++ .../settings/sections/GlobalGeneralSection.tsx | 26 ++++++++++++ ...lobalGeneralSection.tool-output-budget.test.tsx | 46 ++++++++++++++++++++++ .../settings-default-descriptions.test.tsx | 1 + .../src/__tests__/agent-session-helpers.test.ts | 20 ++++++++++ .../src/__tests__/runtime-resolution.test.ts | 15 +++++++ .../__tests__/tool-output-budget-wrapper.test.ts | 45 ++++++++++++++++----- packages/engine/src/agent-runtime.ts | 2 + packages/engine/src/agent-session-helpers.ts | 18 +++++++-- packages/engine/src/pi.ts | 29 ++++++++++---- packages/engine/src/runtime-resolution.ts | 10 ++++- packages/i18n/locales/en/app.json | 4 ++ packages/i18n/locales/es/app.json | 6 ++- packages/i18n/locales/fr/app.json | 6 ++- packages/i18n/locales/ko/app.json | 6 ++- packages/i18n/locales/zh-CN/app.json | 6 ++- packages/i18n/locales/zh-TW/app.json | 6 ++- packages/i18n/src/resources.d.ts | 4 ++ 28 files changed, 323 insertions(+), 38 deletions(-) Fusion-Task-Id: FN-8616 Fusion-Task-Lineage: 3ca99a61-d6ae-48ff-98d2-f14a153aa2b7 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
30f81ac0cd |
fix(engine): compact verification failure output
Keep successful verification responses quiet and return bounded, high-signal diagnostics for failures without hiding zero-work or green-while-red warnings. |
||
|
|
07c8c95b10 |
FN-8614: cap agent tool output
Bound every engine-injected tool result to preserve agent context capacity. - Add shared 16,000-character total text budgets with deterministic truncation markers and validated overrides. - Apply outermost output clamps to Pi and non-Pi plugin tool paths, with semantic caps for high-volume reads. - Cover budget behavior and document the operator-facing configuration contract. Files changed: .changeset/fn-8614-tool-output-budget.md | 7 ++ docs/agents.md | 8 ++ .../core/src/__tests__/tool-output-budget.test.ts | 58 +++++++++++++ packages/core/src/index.gate.ts | 7 ++ packages/core/src/index.ts | 7 ++ packages/core/src/tool-output-budget.ts | 97 ++++++++++++++++++++++ .../src/__tests__/agent-artifact-tools.test.ts | 10 +++ .../src/__tests__/agent-document-tools.test.ts | 10 +++ .../__tests__/agent-task-logs-read-tools.test.ts | 8 ++ .../__tests__/tool-output-budget-wrapper.test.ts | 67 +++++++++++++++ packages/engine/src/agent-session-helpers.ts | 7 +- packages/engine/src/agent-tools.ts | 43 ++++++++-- packages/engine/src/pi.ts | 54 +++++++++++- 13 files changed, 374 insertions(+), 9 deletions(-) Fusion-Task-Id: FN-8614 Fusion-Task-Lineage: b6a76ccd-d7b4-4b43-af7e-cfd16ffb7fc8 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
15a2fb18cc |
Merge branch 'fix/incomplete-pg-ports'
Wire incomplete PostgreSQL ports for archive, reconcile, health, settings cache, agent cache, and async prompt overrides. |
||
|
|
ab87d0d803 |
fix(api): return 404 for missing tasks, and make task deletions attributable
Three related fixes, all originating from a `[api:error] Request failed` log line showing a 500 on `GET /api/tasks/FN-8610/runtime-fallback`. 1. Missing/deleted tasks now return 404 instead of 500. `getTaskImpl` signalled a miss with a bare `Error`, and route catches only mapped errno `ENOENT` to 404 — a leftover from the file-backed storage era. In Postgres mode nothing sets an errno code, so every unknown/missing/soft-deleted/wrong-project read returned 500. Adds a typed `TaskNotFoundError` (message byte-identical) plus a shared `task-lookup-error` mapper applied across the task, session-diff, git/GitHub, workflow and file-workspace route registrars. The same bare throw existed on both archive-lifecycle delete paths, so `DELETE /tasks/:id` was affected too. 2. 5xx logs now carry the origin stack. `rethrowAsApiError` constructed a fresh `ApiError` from the message and discarded the original, so the `FNXC:ApiErrorDiagnostics` contract logged the rethrow site rather than the throw site — the reported log entry had no stack at all. Threads `cause` through the error factories and walks the chain (bounded, cycle-guarded). 3. Task deletions are attributable, and non-operator deletes notify. `task:deleted` audit rows recorded `agentId: "system"` for every HTTP delete, making an operator click indistinguishable from a script or an agent; the calling agent's task id was accepted by the store and then never persisted. Adds a `callerKind` union recorded in audit metadata, tags every delete call site, and stamps a self-reported `x-fusion-client` header from the dashboard client. When the caller is `agent-tool` or `api-unattributed`, a best-effort notice is sent to the operator mailbox; operator and engine deletes stay silent. `x-fusion-client` is attribution, not authentication — anything can send it. No delete-blocking, gating or permission logic is added here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b55077546 |
fix: wire incomplete PostgreSQL ports for archive, reconcile, health
Replace empty backendMode stubs with real AsyncDataLayer paths: archive ID reservation and isTaskArchivedAsync, orphaned task.json re-import, health snapshots via checkPostgresHealth, settings/agent memory caches for sync readers, async builtin prompt overrides, and self-healing audit/health callers that previously used dead sync SQLite fallbacks. |
||
|
|
f01461a70e |
feat(engine): thread the node id onto review-gate leases, activating pre-boot reclaim
Completes
|
||
|
|
3b83282273 |
feat(engine): attribute review-gate leases to a node so dead local leases reclaim fast
Groundwork for FN-8603's remaining ~14-minute wait. Liveness for a pending review gate is judged purely by a 15-minute staleness floor because a lease records WHO took it (`leaseOwner` = run id) but not WHERE, and under multi-node every engine sees every other engine's leases. A fresh-but-unknown lease might be running on a peer, so the floor was the only safe test -- and a lease left by this node's own crashed process is indistinguishable from it. Adds `WorkflowStepResult.leaseNodeId` plus an optional `LocalNodeLeaseIdentity` argument to `classifyReviewLease`. One narrow new case: a lease stamped with the caller's OWN node id whose `startedAt` predates the caller's process boot is provably dead -- the process that could have owned it is gone -- so it classifies as `reclaim` immediately rather than aging out. Deliberately narrow, because widening it is a double-dispatch risk: absent (legacy) or peer node ids keep the floor, and a lease taken by this process after boot is still adopted. InProcessRuntime.start() resolves the local node id from CentralCore (fail-soft; on error it stays undefined and floor-only semantics apply) and passes it to SelfHealingManager. The graph executor stamps the field when deps.localNodeId is set. NOT YET WIRED, so this is inert in production and behavior is unchanged end to end: `localNodeId` is not threaded from WorkflowGraphTaskRunner / WorkflowTaskRuntime down into the executor deps, so no lease actually carries a `leaseNodeId` yet. The reader is ready; the writer needs that pass-through (WorkflowGraphTaskRunnerDeps gains the field, the runner forwards it, and the runtime supplies this.localNodeId). Stopping here rather than half-threading it. Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green (299 + 70), core workflow-step-results suite green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00011b0113 |
fix(engine): recover restart-orphaned review steps in one cycle, raise fix budget
FN-8603 sat in-review for ~36 minutes after an engine restart killed its Code Review session 34 seconds in. It did recover on its own; the cost was latency, not a terminal park. Sweep ordering. reconcile-orphaned-pending-step-results PRODUCES the failed results that recover-failed-pre-merge-steps CONSUMES, but in the periodic maintenance list it ran ~15 entries after it. A step orphaned in cycle N was therefore rewritten to failed only after recovery had already scanned, so nothing re-ran it until cycle N+1. Moved it immediately before its consumer and removed the now-duplicated later entry. Startup recovery already ordered the two correctly. Post-review fix budget. Default raised 3 -> 10 per operator request. Three passes is below the observed convergence length for the gates this fallback actually governs -- Browser Verification and custom optional gates -- since Plan Review and Code Review already resolve to "unbounded" when unset, and exhausting the budget parks the card for a human. The declaration default and five inline `settings.maxPostReviewFixes ?? 3` call sites in executor.ts/self-healing.ts had drifted into separate literals, so raising one alone would have left every unset-settings path on the old value; they now share the exported DEFAULT_MAX_POST_REVIEW_FIXES. Not done, and why. Re-dispatching a restart-orphaned lease immediately at startup is the change that would close the remaining ~14-minute wait, but it is unsound as specified: liveness is judged by a 15-minute lease-staleness floor because leases carry no node attribution, so treating a pre-boot lease as dead would let one node orphan another node's genuinely running review. Needs a node id on the lease record first. Left the floor intact. Verified: tsc clean on core and engine, pnpm lint clean, pnpm test:gate green, self-healing orphaned-pending-step-results and optional-step-revision suites green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cca13737b6 |
FN-8603: reduce steady-state diagnostic log noise
Route routine core, engine, and dashboard diagnostics through debug-gated shared loggers. - Demote steady-state diagnostic sites while preserving warnings and errors for actionable failures. - Add cross-package severity contracts and manifest coverage for demoted log sites. - Document logging severity guidance and add a patch changeset. Files changed: .changeset/fn-8603-log-severity.md | 7 ++ docs/diagnostics.md | 20 ++++-- .../__tests__/log-severity-spam-contract.test.ts | 71 ++++++++++++++++++ packages/core/src/activity-analytics.ts | 5 +- packages/core/src/ai-summarize.ts | 61 +++++++--------- packages/core/src/async-mission-store.ts | 5 +- packages/core/src/async-secrets-store.ts | 7 +- packages/core/src/central-core.ts | 17 ++--- packages/core/src/docker-provisioning.ts | 13 ++-- packages/core/src/index.ts | 1 + packages/core/src/master-key.ts | 9 ++- packages/core/src/memory-compaction.ts | 29 ++++---- packages/core/src/memory-insights.ts | 7 +- packages/core/src/migration-orchestrator.ts | 7 +- packages/core/src/mission-store.ts | 5 +- packages/core/src/node-discovery.ts | 7 +- packages/core/src/notification/dispatcher.ts | 9 ++- .../core/src/plugins/bundled-plugin-install.ts | 11 +-- packages/core/src/reflection-store.ts | 5 +- packages/core/src/secrets-store.ts | 7 +- packages/core/src/task-store/agent-logs.ts | 21 +++--- packages/core/src/task-store/async-events.ts | 5 +- packages/core/src/task-store/async-maintenance.ts | 7 +- packages/core/src/task-store/comments-ops.ts | 7 +- packages/core/src/task-store/task-mutation-ops.ts | 11 +-- packages/core/src/task-store/workflow-integrity.ts | 9 ++- packages/core/src/types/merge-policy.ts | 5 +- packages/core/src/usage-events.ts | 5 +- .../__tests__/log-severity-spam-contract.test.ts | 48 +++++++++++++ packages/dashboard/src/ai-refine.ts | 5 +- packages/dashboard/src/ai-session-diagnostics.ts | 10 +-- packages/dashboard/src/chat.ts | 8 ++- packages/dashboard/src/devserver-manager.ts | 9 ++- packages/dashboard/src/file-service.ts | 5 +- packages/dashboard/src/github-tracking-comments.ts | 7 +- .../dashboard/src/github-tracking-reconciler.ts | 5 +- packages/dashboard/src/github-tracking-state.ts | 5 +- packages/dashboard/src/gitlab-lifecycle.ts | 5 +- packages/dashboard/src/insights-routes.ts | 9 ++- packages/dashboard/src/issue-image-attachments.ts | 5 +- packages/dashboard/src/knowledge-index.ts | 5 +- packages/dashboard/src/plugin-routes.ts | 7 +- packages/dashboard/src/routes/board-workflows.ts | 5 +- packages/dashboard/src/routes/context.ts | 5 +- .../dashboard/src/routes/register-auth-routes.ts | 13 ++-- .../routes/register-docker-provisioning-routes.ts | 7 +- .../dashboard/src/routes/register-git-github.ts | 21 +++--- packages/dashboard/src/routes/register-gitlab.ts | 7 +- .../src/routes/register-session-diff-routes.ts | 9 ++- .../src/routes/register-settings-memory-routes.ts | 7 +- .../src/routes/register-setup-activity-routes.ts | 7 +- .../dashboard/src/routes/register-signal-routes.ts | 5 +- .../src/routes/register-task-workflow-routes.ts | 11 +-- packages/dashboard/src/runtime-logger.ts | 11 +-- packages/dashboard/src/server.ts | 7 +- packages/dashboard/src/sse.ts | 8 ++- packages/dashboard/src/terminal-service.ts | 34 ++++----- packages/dashboard/src/view-chunk-manifest.ts | 5 +- .../engine/src/__tests__/log-severity-manifest.ts | 83 ++++++++++++++++++++++ .../__tests__/log-severity-spam-contract.test.ts | 40 ++++++++++- .../src/__tests__/logger-debug-gating.test.ts | 7 +- packages/engine/src/goal-anchoring-audit.ts | 5 +- packages/engine/src/plugin-runner.ts | 44 ++++++------ packages/engine/src/pty-native.ts | 9 ++- .../engine/src/runtimes/child-process-worker.ts | 4 +- packages/engine/src/self-healing.ts | 12 ++-- packages/engine/src/worktree-hooks.ts | 10 ++- 67 files changed, 632 insertions(+), 250 deletions(-) Fusion-Task-Id: FN-8603 Fusion-Task-Lineage: 53901db6-1af2-4bd7-b5ea-49507e048ef2 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
9ff1587b84 |
fix(engine): keep the task worktree across replan bounces
`moveTask`'s reopen-to-todo/triage block clears `task.worktree` but leaves `task.branch` intact, and `moveTaskToReplanColumn` called it with no options. A replan bounce therefore left the row split-brained: no worktree pointer, but still owning `fusion/<id>`, which was still checked out in the worktree it had just orphaned. The next planning acquisition skipped its resume branch (gated on `task.worktree`), re-created the same branch, collided, and fell into `cleanupConflictingWorktree` — force-remove + `git branch -D` + fresh `git worktree add` + init command, on every bounce. Observed on FN-8603: two Plan Review REVISE bounces burned two full teardown/rebuild cycles for nothing, since planning writes its spec to the task store, not the worktree. Pass `preserveWorktree: true` at the shared seam, so this covers every replan mover — Plan Review REVISE, required-artifact recovery, and the executor and scheduler spec-staleness and filesystem-validation rebounds. Acquisition still re-validates, so a preserved pointer to a removed checkout self-heals as before; the rest of the replan contract (steps reset, status/error cleared) is unchanged. Regression coverage asserts the invariant across both replan-column shapes (triage and plan-in-place todo) and all three reopen origins. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
71279ed042 |
fix(FN-8600): recover a duplicate verdict the planner reported in its reply
The prompt fix stops planners writing the verdict in prose, but it relies on every model reading one sentence correctly. This closes the hole underneath it. When the finalize read finds no spec at all, the planner's streamed reply is searched for a line that is exactly `DUPLICATE: FN-NNNN`. If found, the engine writes the canonical marker file and continues — so marker parsing, keep/delete resolution, and the sourceMetadata.nearDuplicateOf that renders the operator's decision all run on the unchanged file contract rather than a second code path that could drift from it. Deliberately narrow. The marker must occupy a whole line, only the first counts, and recovery is gated on the plan being genuinely absent — a planner that wrote a real spec is never overridden by something it said in passing. The text tail is bounded because the verdict lands in the closing summary, and it tees off onText rather than reading AgentLogger, whose buffer is flushed on a timer. Verified both directions: the tests fail without the recovery block, and the "wrote a real spec while mentioning a marker" case keeps its spec. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a00f2633ce |
fix(engine): demote more TUI chatter across merger, self-heal, and ntfy
Route foreach/merger/worktree/self-healing skips, ntfy send bookkeeping, session-purpose runtime picks, planning using-model, and checkpoint rewind lines to debug so recoveries and failures stay visible in the operator log. |
||
|
|
c7fa02f370 |
FN-8597: restore executor task-done invariant coverage
Restore the quarantined executor graph-completion invariant suite with real foreach projections. - Exercise complete and partial expanded workflow-step projections at the merge boundary. - Remove the rescued invariant suite from Vitest quarantine and clear its ledger entry. - Extend the shared executor logger mock with the debug method required by the integration tip. Files changed: .../__tests__/executor-task-done-invariant.test.ts | 267 +++++++++++++++++++-- .../engine/src/__tests__/executor-test-helpers.ts | 7 + packages/engine/vitest.config.ts | 7 - scripts/lib/test-quarantine.json | 8 +- 4 files changed, 254 insertions(+), 35 deletions(-) Fusion-Task-Id: FN-8597 Fusion-Task-Lineage: 05a08e31-7da0-4c93-86a0-9baf8db7ce52 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
9bad0e1233 |
fix(engine): demote high-frequency TUI log spam to debug
Route process spawn/exit, verification success paths, MCP connect, skill info listings, createFnAgent/session bookkeeping, and executor dispatch chatter through FUSION_DEBUG so the operator log pane keeps real lifecycle outcomes. |
||
|
|
ae512aec2b |
FN-8601: enforce foreach merge proof
Require complete foreach execution evidence before workflow merge review. - Add reusable foreach instance coverage proof evaluation. - Block checklist projection and merge admission on incomplete or failed node results. - Cover core proof logic and PostgreSQL merge-boundary behavior. - Add a patch changeset for the merge safeguard. Files changed: .changeset/fn-8601-foreach-merge-proof.md | 7 ++ .../src/__tests__/workflow-merge-proof.test.ts | 43 ++++++++ packages/core/src/index.gate.ts | 2 + packages/core/src/index.ts | 2 + packages/core/src/workflow-merge-proof.ts | 74 +++++++++++++ ...xecutor-merge-boundary-foreach-proof.pg.test.ts | 111 +++++++++++++++++++ packages/engine/src/executor.ts | 117 +++++++++++++-------- 7 files changed, 314 insertions(+), 42 deletions(-) Fusion-Task-Id: FN-8601 Fusion-Task-Lineage: 40578171-0b13-4538-8f38-3948ed1e92c0 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
2d263acc49 |
fix(FN-8600): keep self-healing from pausing live planners and unstick queued planning
Planning moved into the task's own worktree but never published that path to activeSessionRegistry, so the self-owned-branch reclaim sweep's FN-4819 liveness guard was blind to a live planner. A zero-commit task branch trivially reads as tip-already-merged, so the sweep ran `git worktree remove --force` on the tree a planning session was using, the removal failed, and the failure escalated to branch-conflict-unrecoverable — parking a healthy card paused with no operator action. Planning now claims its worktree through acquireActiveSessionPath (new "planning" session kind) and releases it only while it still owns the record, so a live executor that took over the same path mid-teardown is never cleared. Also fixes planning starvation and its diagnosability: - admitOldest walks past candidates whose lane declines instead of ending the pass on candidates[0], unwinding each declined attempt's pre-held executor slot and reservation exactly so a decline cannot leak capacity past maxConcurrent. - Withheld planning admission emits a deduped task:plan-admission-throttled run-audit event (ids/counts only), written fire-and-forget with the dedupe marker set only after the write lands. Previously the binding gate lived only in a log line that is persisted nowhere, so "why did this card sit queued to plan?" was unanswerable after the fact. Reviewed by 8 review agents; every finding acted on or recorded. A proposed STALE_SEMAPHORE_EXCESS_REPAIR_MS 600s->180s reduction was reverted under review — nested runs are already excluded from the reclaim floor, so the window guards uncounted top-level holders such as a merge body, and shortening it would trade a bounded visible stall for an unbounded silent cap breach. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
795a38c018 |
fix(engine): quiet graph review-entry audits and label engine aborts truthfully
Recognise workflow-graph moves into in-review so gate entry no longer emits handoff-invariant violations, and split pause-abort provenance so engine teardowns are engine-abort instead of hard-cancel. |
||
|
|
af897d9e3c |
FN-8596: isolate cross-root plugin MCP discovery
Prevent cross-root MCP discovery from unloading active plugin runtimes. - Isolate discovery loader lifecycle and runtime-state persistence. - Preserve shared plugin owners when non-owner loader participants stop. - Cover core, dashboard, and engine cross-root discovery behavior. Files changed: .changeset/fn-8596-plugin-discovery-isolation.md | 7 ++ .../plugin-loader-lifecycle-scope.test.ts | 12 +++ .../plugin-mcp-servers-discovery-isolation.test.ts | 115 +++++++++++++++++++++ packages/core/src/plugin-loader.ts | 34 +++++- packages/core/src/plugin-mcp-servers.ts | 8 +- .../context-plugin-mcp-discovery-isolation.test.ts | 48 +++++++++ packages/dashboard/src/routes/context.ts | 39 ++++++- ...-runtime-plugin-mcp-discovery-isolation.test.ts | 69 +++++++++++++ packages/engine/src/runtimes/in-process-runtime.ts | 47 +++++++-- 9 files changed, 364 insertions(+), 15 deletions(-) Fusion-Task-Id: FN-8596 Fusion-Task-Lineage: 231e53b6-a9a3-4a65-9732-3dabe44da198 Co-authored-by: Fusion (runfusion.ai) <noreply@runfusion.ai> |
||
|
|
d47d44c669 |
fix(engine): demote residual routine/peer-exchange poll chatter
Close log-spam skeptic gaps: routine-scheduler pause and re-entrance no-ops, and peer-exchange zero-work sync cycles, move to debug with contract-test locks. |
||
|
|
cfa84781d6 |
fix(engine): demote high-frequency TUI log spam to debug
Route steady-state chatter (maintenance batch, skill listings, activity heartbeats, stuck polls, SSE connect/disconnect, heartbeat timer skips, cron/routine de-dupe skips, hold-release capacity races) through FUSION_DEBUG so the operator log pane keeps real state changes and failures. |
||
|
|
beb83a1c1b |
fix(engine): close the unowned-card strand and harden the planning path
Second FN-8596 strand, found after the first fix shipped. Clearing the
stale `planning` status moved the card into a state owned by NOBODY:
- planning excluded it: stale `firstExecutionAt` from its first pass made
hasAdvancedPastPlanning true, and the previous fix only rescued cards
that still carried a planning-stage status;
- recoverAdvancedTriageTasks — the designated owner of that
"stranded-advanced" class — also excluded it, because it bails on
`workflowIrPinColumnId === "triage"`: it cannot resume a card into the
column it already occupies (the pin was plan-replan, which lives in
triage).
So the card sat indefinitely with no sweep, log, or audit event naming it.
hasAdvancedPastPlanning now decides on arrival order alone for any card in
the planner column: a stamp written BEFORE the card reached triage belongs
to a previous pass, whatever the status is now. A card that genuinely
advanced is still caught by the column check at the top, and one claimed by
execution AFTER landing here has a stamp newer than its arrival, so it
still reads advanced and stays with advanced-recovery. This flips one case
I added in the previous commit — production proved that classification
stranded the card.
Hardening, so this class cannot hide again:
- detectStalledCards: a detect-only watchdog emitting
`task:stall-watchdog-detected` for any non-terminal, unpaused card idle
past 30m with no live session and no queued continuation. Deduped per
shape. It deliberately does NOT mutate — a generic mutator racing the
specialized sweeps is the bug class this file keeps re-fixing, so
recovery stays with the sweep that owns each shape and this guarantees
visibility.
- The silent skips are now loud: runIfStillPlanningUnderTaskLock (all
four callers inherit it), the planning handoff moveTaskIf, and the four
requestPreMergeOptionalStepFix refusals now log why nothing was
scheduled and that the card was left parked.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f005cee885 |
fix(engine): surface silent stalls and add stalled-card watchdog
Make planning-guard and remediation no-ops emit warnings, and detect idle non-terminal cards with no session or continuation so FN-8596-class strands show up in logs and run-audit. |
||
|
|
08f69745f9 |
fix(engine): quiet routine maintenance batch logs in TUI
Route per-step maintenance success/skip chatter to debug so the operator log pane keeps real recovery events instead of 50+ no-op lines every cycle. |
||
|
|
4633c6441b |
fix(engine): stop stranding replan cards on stale execution stamps
Root cause of the FN-8596 strand (card sat in Planning, doing nothing, until an engine restart). Plan Review returned REVISE, the graph rebounded the card to `triage` with `needs-replan`, and triage claimed it — overwriting the status with the TRANSIENT `planning`. `needs-replan` is a durable park that outranks the execution timestamps, but `planning` is deliberately excluded from REPLAN_PARK_STATUSES, so the card fell through to the stamp check. Those stamps were written when it entered `in-progress` on its FIRST pass and are never cleared, so the replanning card read as "advanced past planning" for the rest of the session. From there everything was a silent no-op: updatePlanningStateIfStillCurrent returned false and its callers returned with no log, no audit and no requeue. The revision session wrote the revised PROMPT.md (via the store tool, which bypasses the guard) and the finalize refused to hand the card off — "prompt written, then total silence", status frozen at `planning`. Stale stamps are now discriminated from a live claim by arrival order: a stamp written BEFORE the card arrived in the planner column belongs to a previous pass, while one written after arrival means execution genuinely won the FN-8361 race and recovery must not clear the status out from under it. A missing/unparseable columnMovedAt keeps the prior answer, so this can only narrow the strand, never widen the race. The PR #2360 stranded-advanced class (stamps with no planning status) is untouched — all 30 pre-existing guard cases still pass. Also warns when a planning finalize declines to hand off. That path was completely silent, which is why this strand left nothing in any log. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
581b7d0a49 |
fix(triage): clear stale planning statuses periodically, not only at startup
Observed on FN-8596: a plan-review REVISE routed to `plan-replan`, triage claimed the card with `status:"planning"` and ran the revision session, and the session wrote the revised PROMPT.md then died without finalizing. The card sat in `triage` with `status:"planning"`, no live planner, and no workflow continuation. That status makes the card invisible to triage rediscovery (it looks claimed), and the only sweep that cleared it ran at STARTUP — so the card was unrecoverable short of an engine restart. The leaked-slot reaper then reclaimed its concurrency slot, which made it look idle without making it runnable. Adds a periodic counterpart in the poll loop. Clearing the status is the whole repair: the card is back in triage with a real spec, so ordinary rediscovery re-picks it. It does not move, pause, or fail the card. Guards against racing a healthy planner: the in-process `processing` set, plus a 20-minute staleness floor that also covers a planner owned by another node this process cannot see. Operator parks are never touched. This fixes the recovery gap, not the trigger — why that session failed to finalize is still under investigation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
26dcccb7c3 |
fix(workflow): harden review-gate lifecycle interactions in In review
Follow-ups to running the pre-merge review gates in `in-review`. Each was verified against the code before being fixed; one reported issue was refuted and is noted below. 1. Symbol locks (packages/core/src/task-store/moves.ts) FN-8306 made the lifecycle transition the symbol-lock RELEASE authority but wrote no counterpart. That was harmless while a task only left WIP at handoff/terminal; the gate crossing now releases the task's declared symbols and the remediation node re-enters `in-progress` to edit the same files in the same live worktree with its locks gone. Neither acquire site (scheduler dispatch, claimDueWorkflowWorkItem) is on the graph re-entry path. Adds a symmetric re-acquire on `!wip -> wip`. Best-effort by design: a contended symbol logs and proceeds, which is exactly the pre-fix posture, rather than parking the remediation behind another holder and re-creating the stranding this change set removed. 2. Premature merge (packages/engine/src/self-healing.ts) `recoverMergeableReviewTasks` was the only in-review sweep with no liveness gate. The graph commits the column crossing at node entry and writes the gate's pending lease two DB round trips later, and `getTaskMergeBlocker` has no notion of "enabled but resultless", so in that window the sweep could enqueue a merge with Code Review never run. Filters `executingIds`, matching recoverGhostReviewTasks. 3. Orphan sweep (packages/engine/src/self-healing.ts) The reported restart hazard is REFUTED: nothing re-attaches an in-review graph run, so those leases are genuinely dead and marking them failed is correct FN-8492 behavior. But the sweep also runs from periodic maintenance in the same live process, where a tick between the lease write and session registration could fail a gate that just started. Honors a within-floor `classifyReviewLease`, matching the semantics Plan Review already had. Cleanup of dead leases is delayed by the staleness floor, not defeated. The audit event gains `needsOperatorBypass` for `autoMerge:false` rows, which self-healing deliberately skips and only fn_task_bypass_review can clear — previously indistinguishable from an auto-recoverable rewrite. 4. Stall detection (packages/engine/src/planner-overseer.ts) The `reviewer` and `merger` stages had no time-based check at all and returned `progressing` unconditionally, so a hung gate produced no signal however long it sat. Adds gate-anchored detection on both (a plain in-review card with no reviewState resolves to `merger`, not `reviewer`), keyed on the pending lease's own `startedAt` rather than `columnMovedAt` so it cannot fire during a legitimate human merge-wait. `cumulativeActiveMs` is documented, not changed: it now excludes gate runtime, but adding the `timing` trait to `in-review` would count arbitrary human merge-wait as active work — a worse distortion than the omission. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |