Files
fusion/packages/engine
gsxdsm a4dee1162e Phase B slice B3.1 (U4): resolve the hold column in recoverStrandedCompletedTodoTasks — query and guard together (#2472)
Stacked on #2471 (Phase B slice B2). Base is
`feature/workflow-vocabulary-b2` — do not merge before it.

**First landable slice of U4 (self-healing.ts).** One sweep, one PR, per
the phase's sub-split rule.

## The finding: the guard and the query must convert together

`recoverStrandedCompletedTodoTasks` promotes a card whose steps are all
done/skipped but which is still sitting in the hold column — finished
work that never handed off to review. It decided *"is this card in the
hold column?"* **twice**, and both were literal:

| | was |
|---|---|
| the QUERY | `listTasks({ column: "todo", slim: true })` |
| the GUARD | `task.column !== "todo"` |

**Either half alone is a green diff with zero behavior change.** A
correct guard behind a literal query never runs; a converted query
behind a literal guard rejects every row it just fetched. This is the
shape that made B1's stale-paused-todo fix cosmetic, and the phase brief
predicted more of it here — correctly.

I proved it rather than asserting it:

- literal **QUERY** restored (converted guard kept) → **3 tests fail**
- literal **GUARD** restored (converted query kept) → **2 tests fail**

Neither half passes the suite alone.

## Falsification came first

Per the brief I tried to prove the work unnecessary before doing it. It
is necessary, and the evidence is empirical, not assumed: the 7 tests
were written against unmodified code and 3 failed. Unlike B2's
hold-release — which turned out already converted — **self-healing is
uniformly unconverted at the query level**: 53 of its sweeps carry a
hardcoded `column:` filter (survey in the worker report).

## Negative half, per the brief

A completed card resting in a WIP or review column is **not** promoted.
Dropping a column filter without a per-task hold check would promote
finished cards out of every column — laundering work past review, a
louder bug than the silent one being fixed.

## Test-harness hazard (will recur in every remaining U4 slice)

The pre-existing self-healing store mock returns its fixture from
`listTasks` **regardless of arguments**. A renamed-hold test on that
harness passes while the query stays hardcoded, because the mock hands
the sweep rows the real store never would. The new harness **honors**
the column filter, and one test asserts the query is no longer scoped to
the literal. This is documented in the new file's header for whoever
writes the next slice.

## Cost

The column filter is gone, so the cheap non-column rejections (paused /
executing / incomplete steps / errored / no-commits / skip-bypass taint)
run **first and synchronously**; only survivors pay an IR resolution,
shared through an `irCache`. A board spanning three workflows resolves
three IRs regardless of card count. `includeArchived: false` preserves
what the column filter did implicitly. The hold column resolves **per
task** — a board spans workflows, and a card in *another* workflow's
hold column must not be promoted.

## One pre-existing assertion changed, deliberately

`self-healing.test.ts` pinned `listTasks` being called with `{ column:
"todo", slim: true }`. That query shape changed on purpose; the
assertion now pins the new one. The behavioral assertions either side of
it (one qualifying card, promoted exactly once) are untouched and still
pass.

## Carried A3 questions — both answered

**Q1 — does the sync/SQLite counter have the same pool-id mismatch?**
**Not applicable: there is no sync counter.**
`occupantsByColumnForWorkflowImpl` and `listWorkflowOccupantTaskIds` are
async/PG-only and throw without an initialized `AsyncDataLayer`; the
sync twin went with the PG cutover. There is no second counter that
could mismatch. The surviving pool-id sentinel sites are
`project-store-ops.ts:767/819` and `moves.ts` — both parked by operator
decision, untouched here.

**Q2 — are custom workflows with an explicit numeric limit affected?**
**No, by design.** `resolveColumnCapacity` gives `config.limit` top
precedence (`configLimit` → `limitSetting` → default-workflow
read-through → `Infinity`), and `resolveWipBudgetColumns` documents that
a column with an explicit numeric limit is **independent — its budget is
itself alone**. Such a column never pools, so there is no pool id to
mismatch. Read-only analysis; no code changed for either question.

## Remaining U4 scope (not in this PR)

214 literal occurrences across ~70 methods; **53 sweeps carry a
query-level column filter**. Hold-gated sweeps still to convert:
`clearStaleBlockedBy`, `reclaimSelfOwnedBranchConflicts`,
`reconcileCompletedTask`, `recoverMergedReviewTasks`,
`recoverStuckMergeDeadlocks`, plus non-query `todo` guards in
`recoverPausedAbortFailures`, `reconcileDependencyBlockingLeases`, and
others. `surfaceStalePausedTodos` was already converted (B1 follow-up)
and is verified intact on this branch.

## Verification

- 7 new tests green; **both mutations kill the suite**
- self-healing suite: 415 passed, **1 failure pre-existing**
(`archiveStaleDoneTasks` — confirmed identical by stashing my changes)
- merge gate green (299 + 10 + 71)
- `tsc --noEmit` clean, `pnpm lint` clean

No changeset: `@fusion/engine` is private.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved recovery of completed tasks stranded in workflow-specific
hold columns, including renamed hold columns.
* Preserved recovery for built-in workflows while correctly handling
boards with mixed workflow configurations.
* Prevented recovery for tasks in non-hold columns or with paused,
incomplete, or errored states.
  * Added fallback handling when workflow details cannot be resolved.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 14:36:32 -07:00
..
2026-07-26 18:11:47 -07:00