Commit Graph

10 Commits

Author SHA1 Message Date
gsxdsm
20e3731eb1 docs(workflow-learnings): a deferral's stated blocker is a claim, and it decays like a measurement (#3026)
Two pieces of work were filed rather than fixed in one session, each
with a specific technical reason. **Both reasons were wrong**, and in
both cases the real obstacle was smaller than the stated one.

| filed rationale | reality |
|---|---|
| "the plugin has no scaffolding for faking its stores" (#3020) |
`_harness.ts` builds a real `PluginContext` over a live PostgreSQL
layer; the gap was **two missing readers on a stub** — fixed in #3022 |
| "supplying this needs a published-API change" (#3003) | the type is
dashboard-internal, `@fusion/plugin-sdk` is `private: true`; the actual
obstacle is stale type declarations between two in-repo packages |

The first one matters most: the filed issue was a **pipeline that stalls
forever** on a renamed board. The cost of that excuse would have been a
real stall sitting open behind a plausible-sounding note.

## The shape

Both times the blocker was asserted **from the shape of the problem**
rather than tested. *"This needs infrastructure that doesn't exist"* and
*"this crosses a published boundary"* are each checkable in about five
minutes, and neither was checked before I wrote a paragraph explaining
why the work couldn't proceed.

## Why it's worth writing down

Filing is often right — someone else owns the contract, the fix needs a
decision, the data genuinely isn't there. What makes it wrong is filing
on an **untested** blocker, because a filed issue with a confident
rationale is the one thing nobody re-derives. It reads as settled.

That's the same mechanism as a stale "do not re-probe" note (which this
document already records, and which I had to correct in #3018), one
level up: there a *measurement* went stale, here a *decision* did.

## The rule

**Before writing the blocker down, spend five minutes trying to hit
it.** If it's real you'll hit it immediately and can describe it
precisely — which makes the issue more useful. If it isn't, you have the
fix instead of the issue.

Docs only. No code, no baselines.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 01:17:19 -07:00
gsxdsm
eecc87c31e docs(workflow-learnings): the "named legacy-id collections are clean" entry was wrong (#3018)
It hid two real defects — and it explicitly told the next reader not to
re-probe them.

## What the entry did

Counted **declarations** (48, then 49) and concluded the population was
benign because each one is a fallback vocabulary, a builtin column list,
or an already-converted seam.

All true of the declarations. **The declaration isn't where the defect
lives.**

## Measure the use, not the declaration

A collection used as a **membership gate against a column**. Nine exist,
and two were live user-visible defects sitting inside a population this
doc had marked clean:

| site | defect |
|---|---|
| `TIME_INDICATOR_COLUMNS.has(task.column)` — `TaskCard` | elapsed-time
indicator never rendered on a renamed board (#3014) |
| `PLANNER_ACTIVITY_COLUMN_IDS.has(task.column)` — `useTasks` | planning
border and pulsing badge never appeared (#3017) |

The other seven are genuinely fine, and the reasons are kept because
they're the shapes worth recognising: the no-flags fallback *inside* a
role helper, a seam that seeds the legacy pair then unions resolved
lanes, a marked `DELIBERATE-LITERAL` fallback chain, and a plugin with
no trait source at all.

## The tell

One question separates the two groups: **does a flags path exist in this
file at all?**

Both defects had none — the gate was the only decision, with nothing to
degrade from. Every benign case had a resolved path sitting right next
to the literal.

## Why this is worth its own PR

A "do not re-probe" note that is wrong is **worse than no note**: it
converts one person's incomplete measurement into everybody's blind
spot. That's the same failure this document already records for
`sortTasksForDisplayColumn`, one level up — there an annotation told
readers to skip a *row*, here it told them to skip a *population*.

I wrote the original entry, and I'd read past it twice myself before
#3014 forced the re-measurement.

Docs only. No code, no baselines.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 00:49:56 -07:00
gsxdsm
e9f587c363 docs(workflow-learnings): correct the "bounded" heuristic — a clock-shaped dep is not a fast one (#3012)
The severity heuristic I wrote in #2998 sorted dependencies **by name**,
and #3007 is the counterexample.

## What I got wrong

I classified `lifecycleDates` as *bounded* because its dep list contains
`lifecycleNowMs`, and deferred it in #3001 with the line *"any wrong
answer there survives only until the next update."*

That value is driven by a **local-midnight boundary timer** — one tick
per card per day. So a finished card shows no completion date for up to
**twenty-four hours**. @gsxdsm found it after I'd written it off.

`nowMs`, `Ticker` and `lastFetchTimeMs` span a live 30-second ticker, a
per-fetch stamp, and a daily boundary. Sorting them by name puts a
day-long defect in the same bucket as a 30-second one.

## The sharper half

A card in a **completion lane doesn't subscribe to the shared live
ticker at all** — that's exactly what the ticker's eligibility check is
for, and what #2996 fixed. So the "fast" dependency that would have
rescued this population is the one thing that population never receives.

The corrected question is: **which dependencies refresh *for this
population*** — not which ones appear in the list. Two of my three
severity calls in that sweep leaned on a dep that the affected cards
structurally never get.

## Why this is worth a PR rather than a quiet edit

The doc is what the next person triages against. #3001 explicitly told
them the four "bounded" sites were deprioritised **by design** — on
reasoning that was wrong for at least one of them. Leaving that in place
means someone defers a day-long defect on my say-so.

Docs only. No code, no baselines.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 00:33:25 -07:00
gsxdsm
e78bf20d55 docs(workflow-learnings): a sixth shape — the resolved value arrives after a memo has answered (#2998)
## The shape

Three defects this session, all the same, none visible to any instrument
here:

A lane value resolved **asynchronously** (the board fetches workflow
traits after first paint) is read inside a `useMemo`/`useCallback` whose
dependency list omits it. The first computation runs with the flags
`undefined`, the role helpers correctly fall back to legacy ids, and on
a **renamed** board that answer is wrong. When the flags arrive nothing
in the dep list changed, so the memo never recomputes.

| defect | severity |
|---|---|
| blocker fan-out trait index (#2993) | permanent — empty index for the
mount |
| card live elapsed-time indicator (#2996) | permanent — never
subscribes |
| near-duplicate chip (#2997) | bounded — self-heals on the next task
refresh |

A legacy board hides all three: there the fallback already answers
correctly on the first paint, so the stale list costs nothing. **Every
instance is renamed-board-only**, which is why they accumulated — and
this repo has no `react-hooks/exhaustive-deps` rule, so the class is
invisible to lint.

## Two properties decide severity, both readable off the dep list

1. **Does any dependency refresh quickly?** `allTasks`, a live clock, a
task identity — any of them rebuilds the closure on the next update,
making the wrong answer a bounded window. The chip keys on `allTasks`
and recovers; the indicator keys on `task.column`, which never changes,
so it never does.
2. **Is the value covered transitively?** A dependency that itself lists
the flags gets a new identity when they arrive, and that propagates.

## A gate was built and rejected — the part worth writing down

The scanner reports **19 sites; two were real.** Property 2 is why:
transitive coverage is invisible to any purely syntactic check and would
need a real dependency graph.

`TaskCard`'s context-menu memo omits all three role flags and is
**nonetheless correct** — it depends on `taskActionMenuModel.actions`,
and that model lists `taskColumnFlags`, so the whole chain recomputes. I
checked that before filing it, which is the only reason this PR isn't a
bug report about missing Archive/Revert menu entries.

Freezing 19 would have baselined mostly noise and trained everyone to
skip the report — the exact failure this document already records for
`sortTasksForDisplayColumn`, where an annotation saying "ignore these"
hid a real defect for days. **A good investigative tool is not
automatically a good ratchet**, and the next person deserves to know the
turn was considered rather than missed.

The triage that does work is cheap: run the scan, then ask the two
questions above. Nine of nineteen survive question 1; hand-checking
those is an afternoon, not a project.

Docs only — no code, no baselines.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 00:00:06 -07:00
gsxdsm
6a465e1006 docs(workflow-learnings): probe harnesses lie more often than the gates do (#2983)
## What

Probing four gates with unimagined shapes this session (#2979, #2980,
#2981) produced **two rounds of silently invalid results** — both from
the harness rather than the instrument, and both agreeing with what I
expected, which is why neither was noticed on the spot.

1. **`node gate.mjs | tail` then `echo $?` reads *tail's* exit status.**
Every probe reported "caught". The gate was in fact failing on `main`
for an unrelated reason, so the runs proved nothing. That fictional
evidence nearly shipped a double-counting change to the SQL gate.
2. **A gate that lists files with `git ls-files` cannot see an untracked
probe file.** Six census probes reported "missed" — including the shape
the census is explicitly built for, which was the tell.
Filesystem-walking gates (`check-sql-column-literals`,
`check-inert-flag-seams`) see untracked files; the census does not.

The rule that catches both in one step, now written down:

> **A probe run needs its own control.** Include one shape the
instrument is known to catch and one it must not flag. If the known-good
shape doesn't come back caught, stop — you're measuring your harness.

Worth stating plainly because the two failure modes have opposite costs:
a probe that wrongly reports *caught* retires a real hole; one that
wrongly reports *missed* sends you rewriting an instrument that was
already correct.

## Two measured negative results, recorded so nobody re-runs them

Added to the existing "Surfaces that were checked and are CLEAN"
section:

| shape | population |
|---|---|
| `switch (task.column)` with legacy `case` labels | **0 sites** |
| a legacy id hoisted into a single const, then compared | **1 site —
and it is correct code** |

The one site is `self-healing.ts:2992`, which seeds `let holdColumn =
"todo"` as its documented legacy floor and then overwrites it from
`resolveLifecycleColumns(...).hold`. The census is right not to flag it;
a naive version of this probe reports it as a defect.

The second shape was worth measuring precisely because **the same shape
had a real population in SQL** — it's what #2980 fixed. It did not
transfer. Population is a property of how people write that particular
kind of code, so each instrument has to be measured on its own rather
than by analogy to a sibling that just turned something up.

## Why this is docs and not a gate change

The census's comparison-only scope is adequate for this codebase: every
blind shape I could construct has an effectively empty real population.
Demanding new detection would have forced a large baseline change across
the program's central instrument for **zero defects** — the same mistake
as filing "48 uncounted sites" that the existing section already warns
about.

Docs only. No code, no baselines touched. All five gates green.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 23:16:26 -07:00
gsxdsm
b1bd571682 batch-sql-ratchet: the census / gate-ratchet family — collection branch, fold here (#2941)
## Family branch for consolidation directive item 4

`batch-sql-ratchet` did not exist and ~10 open PRs are waiting for a
collection point, so this establishes it. **Fold your census/ratchet
commit here and close your own PR as superseded.**

```bash
git fetch origin batch-sql-ratchet
git checkout -B batch-sql-ratchet origin/batch-sql-ratchet
git cherry-pick <your-sha>
# verify scoped, not full suite:
pnpm --filter @fusion/core exec vitest run src/__tests__/archived-column-gate-parity.test.ts --silent=passed-only --reporter=dot
git push origin HEAD:batch-sql-ratchet
```

**Candidates I can see open right now** (owners: please fold + close):

| PR | branch |
|---|---|
| #2938 | `fix/comments-ops-sentinel` |
| #2935 | `fix/task-artifacts-sentinels` |
| #2933 | `chore/commit-tightened-census-baseline` |
| #2931 | `fix/async-comments-sentinels` |
| #2928 | `fix/audit-ops-sentinel-marker` |
| #2925 | `live-task-column-lanes` |
| #2923 | `fix/task-id-integrity-sentinel` |
| #2921 | `fix/plugin-store-migration-marker` |
| #2894 | `gate/sql-literals-match-census-placement` |

That is **10 → 1** once folded. I have not cherry-picked anyone else's
commits — folding someone's work without them verifying it is how a
batch lands broken.

---

## What is in it so far (mine, from #2924)

**Clears a live main red:** `archived-column-gate-parity` fails on
`origin/main` today.

```
AssertionError: TypeScript encoding changed.
  async-comments-attachments.ts: 8 → 5
```

#2886 fixed a real bug — archived-document guards failing in *opposite*
directions on a renamed lane — by replacing three `column ===
"archived"` comparisons with `isArchivedLane(column, archivedColumns)`.
The AST scan counts raw comparisons, so the tally dropped.

**What I did not do is record it as three sites converted**, because
measured, it is not:

```
grep -rn "archivedColumns:" packages/core/src packages/engine/src --include="*.ts" | grep -v __tests__
→ (no matches)
```

No caller passes it. The parameter defaults to `LEGACY_ARCHIVED_LANES =
new Set(["archived"])`, so every call resolves to the literal it
replaced — byte-identical behaviour, resolved branch dead.

That matters for this guard's whole argument: its header warns that
converting the TypeScript half while the Drizzle and raw-`sql` halves
still compare the string is a split brain *"no test would catch, because
every builtin workflow spells the column `archived` so the two halves
agree by accident on every board we ship."* **There is no split brain
today precisely because the resolved half is unwired** — it becomes one
the moment a caller threads real lanes in without the SQL sides moving.
Recorded inline so `5` cannot be read as "3 sites done"; flagged on
#2886.

Verified not a split brain: the Drizzle and raw-sql inventories are
unchanged and both pass — worth stating because those assertions run
*after* the TypeScript one, so a plain red says nothing about them.

Scoped edit to `AUDITED_TS_SITES` by line range: these paths appear in
more than one inventory here, and an unscoped replace would quietly edit
the raw-sql side too, making the parity guard agree with itself (the
trap I hit in #2817).

Guard still bites: appending a real `task.column === "archived"` to an
audited file fails it. Core **4852 passed / 0 failed**, lint clean,
test-only.

Closing #2924 as superseded by this.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved task delegation messages when workflow pickup cannot be
confirmed.
* Delegation results now clearly indicate when a task has not been
verified for pickup.

* **Quality Improvements**
* Added validation checks to catch future-dated markers and inconsistent
SQL-column usage.
* Refined workflow checks to distinguish stale configuration from
incomplete configuration.

* **Documentation**
* Updated lifecycle conversion guidance with more accurate audit
findings and limitations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 19:59:14 -07:00
gsxdsm
7b68f20501 batch(docs): fold the three workflow-learnings / annotation PRs into one (#2942)
## Family batch — replaces #2926, #2892, #2887

Per the consolidation directive: the u9/e2e **docs family**, folded into
one branch and one CI run. Three PRs, five commits, **five files,
comment and markdown only**.

| folded PR | commits |
|---|---|
| #2892 `docs/union-vs-per-task` | the project union and the per-task
answer are not ranked; date correction |
| #2926 `docs/date-my-measured-claims` | date the measured claims (one
was wrong); date the grep-vs-AST measurement in the SQL gate header |
| #2887 `docs/archived-state-literals` | mark the three archived STATE
literals as deliberate |

Cherry-picked in original order with authorship preserved; all five
applied clean, no conflicts.

## Scope is provably comment-only

```
docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md
docs/solutions/workflow-learnings/project-union-versus-per-task-lanes.md
packages/core/src/task-store/async-maintenance.ts        ← FNXC DELIBERATE-LITERAL annotation
packages/core/src/task-store/workflow-definitions.ts     ← FNXC DELIBERATE-LITERAL annotation
scripts/check-sql-column-literals.mjs                    ← header prose only
```

Every added line in `packages/` and `scripts/` is inside a comment —
checked by filtering the diff for declarations, conditionals and
returns, which returns nothing. The two core files gain
`DELIBERATE-LITERAL` markers explaining that `'archived'` is a **state**
marker there, not a lane: the sweep collects rows Fusion itself archived
or soft-deleted, so widening to the resolved archived set would pull
live cards into a cleanup pass.

## Verification (scoped, per the directive — not the full suite)

- `pnpm lint` — clean
- `check-sql-column-literals` — exit 0 (the file it annotates)
- `check:lifecycle-columns` — exit 0 (the markers it adds are
census-visible)
- `sync-workflow-ir-callsite-allowlist.test.ts` — 3/3

## A correction worth recording

Mid-fold I saw a changeset, `self-healing.ts` and a test file in `git
diff origin/main..HEAD` and nearly reported the batch as impure. They
were **main's own commits** — `origin/main` advanced between branch
creation and the diff, so the comparison was against a stale base.
Rebasing onto current `main` reduced it to the five files above. Worth
flagging for anyone else folding a family today: with `main` moving this
fast, diff the branch **after** rebasing or the file list will lie to
you.

## Closing the originals

#2926, #2892 and #2887 are superseded by this and are being closed. I
hold no PRs of my own in this family — all mine merged — so this fold is
on behalf of the family rather than a rollup of my own work.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 19:12:26 -07:00
gsxdsm
5792452f0a docs(workflow-learnings): mutation testing has one blind spot — your own imagination (#2858)
The most transferable thing this lane produced, and it is a correction
to advice I wrote earlier in the same document.

## The gap

Every other section here says *"watch the guard go red before you trust
it."* That rule is necessary and **not sufficient**, and the way it
fails cost the most.

Three instruments were written during this program. Each was
mutation-tested in both directions before shipping. Each was green.
Reviewers then found, in those same instruments:

- a **file-level pre-filter** that skipped whole files, so a forbidden
site added to a file with no other SQL was invisible;
- an **anchored pattern** that missed qualified and compound fragments
(`t."column" = 'done'`);
- a scan over **SOURCE text**, where a double-quoted TS string still
spells `\"column\"` with the backslashes in it;
- an operator list of `= != <>` that never considered **`IN (...)`**;
- and worst, a template scan that joined only the **static spans** — so
a Drizzle query, which puts the COLUMN in the interpolation hole and the
legacy id in the static text, matched nothing. **That gate was blind on
the exact files it was built to freeze.** Enabling that one shape took
the population from 14 to 31 and revealed five previously invisible
files.

## Why the mutation tests could not catch any of them

All five are **false negatives**, and the reason is structural rather
than sloppy:

> A mutation you write is a mutation you already imagined, so it lands
inside the space your scanner understands.

Reintroducing a defect the checker was designed around proves the
checker still handles that defect. It says nothing about shapes you
never modelled.

## What does find them

1. **Run the instrument against the code it was written for and read the
hits by hand.** The Drizzle blindness was obvious the moment someone
asked *"why is the merge-queue query — the reason this exists — not in
the output?"*
2. **Prefer one unanchored pattern over a fast pre-filter plus a precise
one.** Every false negative above came from two patterns disagreeing
about whether to run the real check at all. A pre-filter is a second,
weaker specification of the thing you are testing.
3. **Treat a guard's own count as a claim to verify, not a result.** "14
sites" read as coverage for days; it was the subset one scanner happened
to model.

A false positive is loud and gets fixed. **A false negative prints a
baseline and reads as coverage.**

Docs only — no changeset.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 15:55:00 -07:00
gsxdsm
be12ca905b docs(workflow-learnings): the fifth shape — converted consumer, literal-passing producer (#2835)
Found by the operator reviewing **my own** shipped fix (#2823). It
belongs in this document because it is the only shape so far that every
instrument in the program reported as done.

## The defect

`clearNearDuplicateReferencesTo` was converted to resolve the
canonical's column flags, and my test proved that by supplying `column:
"shipped"`. Both production call sites in `moves.ts` gated on the
**resolved** complete lane and then passed the literal `column: "done"`.
The consumer was correct; the producer was never converted; and the
hand-supplied fixture value is exactly what hid it.

Why nothing caught it:

- the **census** counts comparisons — a value passed as an argument is
not a comparison;
- the **seam check** asks whether callers *supply* the argument — these
did, with a literal;
- the **test** supplied the interesting value itself, so it exercised
the consumer and never the producer.

And the part I would have got wrong: driving the flow end to end still
does **not** distinguish them. The consumer looks the passed column up
in the canonical's IR, finds nothing for `done` on a renamed board, and
falls through to the legacy predicate where `done` *is* terminal — right
answer, wrong reason. It only bites on a board that declares a `done`
column **without** the complete trait.

## The rule

> A differential test must vary the value the **production** code
computes, not one the test hands in.

If the fixture passes the lane name, it has tested the consumer. Who
computes that argument in production, and do they compute it or spell
it, is a separate question — and the one that was wrong here.

## Measured, so nobody builds the wrong instrument

An AST probe for call arguments shaped `{ column: "<legacy id>" }` finds
**79 sites** across `packages/`. It correctly flags the two real
`moves.ts` offenders — but most of the rest are legitimate: `set({
column: "archived" })` writing the archive state, `listTasks({ column:
"todo" })` filtering a query.

So a blocking gate on this shape needs a curated list of consumers that
interpret a column as a **role**, as opposed to storing or filtering it
— a per-consumer judgment call, not a mechanical check. **Recorded
rather than built**, and deliberately not attempted on top of five
unmerged PRs in packages I do not own.

I also verified one nearby call site that a naive version of this probe
flags and which is **not** a defect: `archive-lifecycle-2.ts:353` passes
`column: "archived"`, but that path sets `task.column = "archived"`
unconditionally, so it is passing the column the card actually reached.
Exactly what the operator's fix is about.

Docs only — no changeset.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 14:25:59 -07:00
gsxdsm
ba40942a10 batch-dashboard-app: 75 → 2 across packages/dashboard/app — the last two are deliberate, not missed (#2772)
**Batch branch is live: `batch-dashboard-app`.** Push conversions here
as commits rather than opening per-file PRs — that is the CI-run
bottleneck this model removes.

**One-line ownership note for you to arbitrate:** you have addressed me
as U11, U12 and U7 at different points, so the `u12 worker ->
batch-dashboard-app` mapping is ambiguous from my side. I claimed it
because `dashboard/app` is where I have done the most work this session
(TaskContextMenu, Column, TaskCard, TaskDetailModal, columnRoles,
taskActivity) and I know which of its guards are load-bearing fallbacks.
**If another worker is the intended owner, say so and I will hand the
branch over rather than both of us pushing to it** — two workers on one
shared branch is exactly what silently discarded a reviewed fix in #2645
today.

## The work order (measured at branch point, tests excluded)

**75 guards across 32 files.** Largest: `TaskContextMenu.tsx` 9 ·
`Column.tsx` 7 · `ListView.tsx` 6 · `TaskDetailModal.tsx` 4 · then a
long tail of 3s, 2s and 1s. Full per-file list is in the committed work
order so feeders can claim without re-measuring.

## Two rules this surface keeps tripping on

**1. A literal after `??`, or in the `else` of a `flags ?` ternary, is a
DEGRADED-MODE answer — not an unconverted guard.** Two real states reach
it: the **pre-load window** (board renders before the workflows fetch
resolves) and a card stranded on an id its workflow no longer declares.
In both, `columnFlagsById` has no entry at all. Deleting the fallback
does not remove a decision — it substitutes "no role" silently, and
affordances vanish during first paint.

Those sites reach 0 by **marking**, not deleting. Expect
`TaskContextMenu.tsx` and the `utils` files to be **mostly marks**. A "9
→ 0" that deleted 9 fallbacks is a regression wearing a green census.

**2. A marker excuses ONLY the construct it is attached to** — the
statement or function holding the literal, not a sibling declaration.
This has cost three passes, two of them mine; my first attempt on
`reliability-metrics.ts` scored **1 of 6**. **Verify by the count
moving, not by the comment existing.** With the ratchet gate-blocking, a
mis-marked batch either wedges the gate or locks the miss into a
re-recorded baseline.

## Status

Opening commit is the work order only — **0 of 75 converted so far.** I
am near the end of my context, so I am establishing the branch and the
shared list rather than starting conversions I cannot finish cleanly.
Feeders can begin immediately; I will keep the branch rebased.

My other PR **#2762** (`live-agent-count.ts` 6 → 0) is green and
unconflicted — per your rule it should land rather than fold into a
batch, and it is `packages/core` so it belongs to batch-core anyway.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Task UI now resolves workflow “column roles” per task to drive
diffs/merge details, routing/steering, progress/runtime visibility, and
review badges.
* Right-dock/overflow views and dev-server now use per-task column
traits for “executing” behavior and dependency-based “Up Next”
eligibility.
* **Bug Fixes**
* Fixed bulk action selection/delete/archive eligibility and prevented
cross-workflow role leakage.
* Made in-review/stale-paused-review, stuck, and effective
executor/validator model logic role-aware.
* **Tests**
* Added regression coverage for degraded-flag behavior and ensured
resolved-flag props aren’t ignored.
  * Added a static check to fail builds on inert optional flag seams.
* **Documentation**
  * Updated batch work-order and mega-batch branch guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

## Late addition: the seam gate was masking a real offender

`scripts/check-inert-flag-seams.mjs` matched call sites by NAME, so two
same-named functions in
different modules were conflated. I had documented that as a known
false-positive source and moved
on — reports mentioning `sortTasksForDisplayColumn` are noise, read past
them.

That annotation was the damage. Core's `sortTasksForDisplayColumn`
genuinely never receives its
`columnFlags` argument outside its own tests. The dashboard's separate
function of the same name
(`app/components/taskSorting.ts`), called with up to five arguments from
`Lane`/`Board`/`ListView`,
was raising the arg-count max and clearing core's seam. The offender was
behind a row everyone had
been told to skip.

The gate now records the module each callee is imported from and matches
it against the seam's
declaring module.

**Measured, by reverting the change:** the scan prints `17 seams, all
supplied` and emits **no row**
for the function. With the change, it is reported. Both directions
watched.

Reported on #2783 rather than fixed from outside — core owns it, and
"wire the flags" vs "drop the
parameter and let the literal stay counted" is their judgment call.
TEMPORARY allow-list entry
carries it meanwhile; the existing staleness check fails the moment the
site becomes supplied, so
the entry cannot outlive the fix.

Two known limits remain, both inherent to name matching and both
documented in the script: the
one-supplier floor, and the `__tests__` exclusion (hence the two
permanent `ALLOWED` entries).


## And the one-supplier floor, closed the same way

I wrote in the section above that the floor "hasn't cost anything yet."
That is verbatim the
reasoning that kept the imported-shadow bug alive, so I closed it
instead of leaving the note.

`best < arity` asked only whether SOME caller supplied the argument. One
correct call site cleared
the seam while every sibling took the legacy fallback — the
`isTaskStuck` defect class, where two of
three sites omitted the flags and the gate stayed green because the
third was right. Review caught
that one. A partially-supplied seam is the harder of the two:
wholly-unsupplied is uniformly wrong,
this works on the board you tested and degrades on the column you did
not.

**Measured:** dropping the flags argument at `Column.tsx`'s supplied
call site produces
`supplied by 5/6 call sites; omitted at
packages/dashboard/app/components/Column.tsx:1 (of 2)`;
restoring returns `all supplied at every call site`. Red and green both
watched.

Two real omissions found, both on `isNearDuplicateCanonicalInactive`:

- **`TaskDetailModal.tsx`** — deliberate, and it **corrects a note I
left at that site**. The old
note said hoisting the flags state was "the actual fix." It is not, for
this call: the flags in
scope describe the *modal's* task, and the canonical is a **different
task** on a column this
component never resolves. Passing them would type-check, read as a
conversion, and answer about
the wrong task — exactly what `column-role-degraded-flags.test.ts`
exists to catch. Supplying it
  correctly needs a fetch, which is a data change and out of scope.
- **`core/task-store/branch-group-ops.ts`** — genuinely wireable (the
impl is async and already
holds `store` and `canonicalId`). Reported on #2783, not edited from
outside.

Exemptions for this class are keyed by **call site**
(`<file>::<function>`), not by function name.
A name-level entry would waive every site of a partially-supplied seam,
which is backwards — its
other sites are correct and are the reason the omission is worth
reporting. Both entries carry the
same staleness check as the name-level list and cannot outlive their
fix.

Remaining known limit, now the only one: the `__tests__` exclusion,
which makes a test-only export
read as having no callers. That is what the two permanent `ALLOWED`
entries are.


## The `__tests__` exclusion, and two allow-list entries built on false
reasons

Named as the "last remaining limit" above, so it got closed too. The
scan now reads test files for
call sites — but counts them **separately**, and a test never clears a
seam. That direction is the
dangerous one: counting test callers as suppliers would have re-hidden
core's
`sortTasksForDisplayColumn`, whose only suppliers are its own tests.
Measured by lifting its
exemption: still reported.

Both permanent allow-list entries claimed the scanner couldn't see their
callers. **Both reasons
were false**, and reading tests is what proved it:

- **`evaluateMergeBlockerGuard`** — zero callers in tests either. Its
only reference in the repo is
its own declaration; never registered as a trait hook; the
`evaluateDefaultWorkflowGuards` reader
its file header credits does not exist. The `lifecycleColumns`
conversion went onto dead code, and
its note describes a crossing the guard cannot make. Reported on #2783,
including the two things I
am explicitly *not* concluding (no `"guard"` hook is registered in
production; whether that is
  residue or a dropped registration needs core's intent).
- **`isRecoverableMissingWorktreeReviewFailure`** — 5 test call sites.
It wraps
`...WithProgress`/`...NoProgress`, the live pair called from
`self-healing.ts`, both supplying
  `reviewColumns`. Entry kept, true reason recorded.

### A wrong turn, recorded because it is the failure mode this PR is
about

I first classified no-production-caller seams as *informational* when
they weren't re-exported from
a package index, reasoning that a public export might be called
externally. That silently downgraded
`sortTasksForDisplayColumn` — a confirmed real offender — from failing
to a footnote. Publication
status has nothing to do with whether there is production behaviour to
be wrong. Reverted to the
simple rule: no production caller means inert, and it fails.

It is worth stating plainly because it is the exact shape of everything
else in this PR: a change
that made the gate read *cleaner* while making it catch *less*, and it
type-checked, passed every
test, and would have reviewed fine.

### Where that leaves the check

Every blind spot named in this PR has now been closed, and **each one
produced a real defect within
minutes of closing it** — imported shadows, the one-supplier floor, the
`__tests__` exclusion. Four
verified findings went to core, one to engine. I would not read the
remaining ~240 guards' green
gates as evidence that they are clean; I would read them as untested.


## Two guards for one question, one of them worse

Having hardened the script, I checked its older twin rather than
assuming it was fine.
`resolved-flags-seams-have-suppliers.test.ts` carried its own copy of
the trailing-flags-parameter
check — written before the script existed — with **all three** holes the
script has since closed.

**Measured on one reintroduced defect** (dropping the flags argument at
`Column.tsx`'s supplied
`isNearDuplicateCanonicalInactive` call):

| | result |
|---|---|
| `scripts/check-inert-flag-seams.mjs` | `supplied by 5/6 call sites;
omitted at .../Column.tsx:1 (of 2)` |
| this test's arity half | **3 passed** |

Deleted the arity half. Redundancy between a strong and a weak check
isn't redundancy — it's a green
result available to whoever runs the weak one, and there was no signal
at the call site telling you
which you were looking at.

The **props-shape half stays**: it has no twin in the script, and I
confirmed it still fires by
reintroducing the original `PrPanel` defect (outer component stops
destructuring `taskColumnFlags`)
— it reports `PrPanel declares taskColumnFlags but never takes it`.

Dashboard app suite: **113 files / 3921 tests** (was 3922 — the deleted
case is the difference).


## The gate started catching defects as they landed

Syncing with main brought in three fresh conversions from other workers.
The hardened check flagged
all three immediately — the first time these guards have fired on
someone else's landed code rather
than on my own.

- **`TaskCard`** — `getRunningOptionalGateBadge(task)` omitted flags
while *both* `ListView` sites
supplied. Fixed, and `taskColumnFlags` added to the `useMemo` deps: no
`exhaustive-deps` rule here,
so a memo that reads flags without listing them keeps the first-paint
`undefined` answer and
  reproduces the bug through staleness instead of omission.
- **`TaskTokenStatsPanel`** — `getTotalAgentActiveMs` omitted while
`TaskCard` supplied, so the same
runtime number came from the real column on a card and from legacy ids
in the detail modal. Now
takes `columnFlags`, supplied from `detailColumnFlags` — correct here
because the panel renders the
  modal's **own** task, unlike the near-duplicate canonical above.
- **`ListView` ×2** — passed `columnFlagsById.get(task.column)`, the
cross-workflow **union**. A task
whose own workflow doesn't declare that column gets a *neighbour
workflow's* traits. The landed
comment justified it as "this list already owns `columnFlagsById`" —
exactly the reasoning
`column-role-degraded-flags.test.ts` exists to reject. It failed on
merge and is how I found this.

Also: the `getTotalAgentActiveMs` exemption I was carrying
**self-retired**. Main wired the seam, the
staleness check failed the entry, and I removed it. That mechanism has
now paid for itself once.

### Pre-existing, NOT from this PR: `App.test.tsx` is red on main

`app/components/__tests__/App.test.tsx` fails **10 of 141** identically
with my changes, with my
changes stashed, and with main's own `App.tsx` restored. Not mine, and
not in the merge gate.

**Bisected on clean `main` checkouts, so this is measured rather than
inferred:**

| commit | date | result |
|---|---|---|
| `main~400` (`41d60f0355`) | 2026-07-25 | **140 passed** (140 tests) |
| `main~275` (`74d6513fae`) | 2026-07-27 | 3 failed / 141 |
| `main~210` (`d2ce1ba8b5`) | 2026-07-29 | 10 failed / 141 |
| `main` (`6fc98fd6c7`) | 2026-07-30 | 10 failed / 141 |

So it is **not one regression** — it degraded in two stages across
2026-07-25 → 07-29, and the test
file itself changed in that window (140 → 141 tests). Three commits
touched it there:
`73b2a32e2b`, `f26cbedf4f`, `f157bf7460`. That window overlaps the
workflow-owned lifecycle
migration, which is suggestive but not something I confirmed.

The failures are render-level, not assertion-level — `Unable to find an
element with the text: + New
Task`, `Unable to find role="dialog"`, `Unable to find ... Back nav
task`. The board appears to
render nothing. That reads like a real regression or a harness mismatch
after the lifecycle
migration, not a flake, so I have deliberately **not** quarantined it —
quarantine is for flakes, and
using it here would hide the signal. Flagging for whoever owns
`App.tsx`.

My suites: `app/__tests__` **113 files / 3921 tests** green, `tsc` 0,
lint 0, census `--strict` 0,
seam gate 0.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:46:12 -07:00