The active-worktree slot-accounting fix removed two deliberate scheduler
literals (done/archived: 3 -> 2); re-record so the ratchet follows the count
down.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Retained directories on queued, paused, blocked, or terminal tasks no longer
consume scheduler slots. Agent concurrency and worktree capacity now count the
same canonical live-task population through one project admission ceiling
(resolveActiveTaskCapacityLimit) with an atomic reserveIfAvailable claim, so
planning, execute, and merge lanes cannot each observe and claim the final
worktree slot independently.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The field was defined in persistence and serialization, the executor's Plan
Review replan-cap park wrote it, the triage manual gate null-cleared it, and
the dashboard special-cases it (isReviewBudgetExhaustedApproval badge + detail
explanation) — but updateTask's field-by-field merge never applied the key, so
every writer silently dropped it. FN-8647's 15-cycle non-converging Plan Review
loop therefore parked with a generic 'needs approval' and no hint it was a cap
escalation.
Merge contract, pinned by tests with a measured revert proof (3/4 fail
pre-fix): set persists, explicit null clears, a status write that leaves
awaiting-approval without addressing the reason auto-clears it so an approved
or replanned card cannot carry a stale escalation reason into its next park,
and unrelated updates leave it untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The right-padding bug recurred three times because 'Task Detail modal' names ONE
surface with THREE shells and fixes kept landing in the wrong file. Disambiguate:
- Rename the just-introduced .floating-window--tablet marker to
.floating-window--tablet-viewport and document the naming contract next to it:
--tablet-viewport = viewport MODE classifies tablet (touch or not, styling
surface); --touch-geometry = tablet AND touch (enlarged 44px targets only).
- Add SHELL NAMING MAP breadcrumbs at the top of TaskDetailModal.css and
TerminalModal.css pointing inset/padding fixes at the FloatingWindow shell
that tablet popups and floating terminals actually render through.
Comment wording deliberately avoids dot-prefixed class tokens because
FloatingWindow.test.tsx scans raw CSS (comments included) with selector regexes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Allow queued tasks to reuse worktrees they already hold when the durable worktree ledger is at or above capacity. Preserve independent agent and semaphore limits, and avoid releasing a worktree slot that a rejected transfer never acquired.
Third recurrence of the Task Detail right-padding bug (FN-8630/FN-8634): those
fixes only covered the .modal-overlay shells, while every tablet task popup and
floating terminal renders through FloatingWindow, whose shared body carries
FN-8015's margin-inline-end scrollbar gutter — a 16px right border with a 0px
left one. Tablet-mode windows (.floating-window--tablet, keyed on viewport MODE
so non-touch tablet widths match too) now zero the gutter; touch never grabs
scrollbar thumbs, so the desktop hot-zone conflict FN-8015 solves cannot occur.
GitHub-import's detail panel, which used the gutter as its right inset,
compensates locally. Desktop keeps the FN-8015 contract.
Also per operator request: the floating terminal is draggable from the empty
strip space behind the tabs and anywhere in the top toolbar, not only the
FN-8633 grip. Tab presses keep stopPropagation (scoped to .terminal-tab) so
they never start a window drag; the tablet floating header supersedes the
pan-x contract with touch-action: none (an overflowing strip is replaced by
the mobile-tabs dropdown, so no visible strip pans horizontally).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## Census
**Before: `COLUMN guards (the backlog): 2`, `--strict` RED. After:
`BACKLOG ZERO`, all five gates green.**
Two commits from last night's `maxWorktrees` rollout copied the same
holder ledger, both with literals:
| commit | file | gate |
|---|---|---|
| `374956ef23` | `triage.ts` | planning admission |
| `6c7467a78d` | `executor.ts` | `fn_spawn_agent` |
```ts
t.column !== "done" && t.column !== "archived"
```
## What it costs
Both exclude terminal lanes because a finished card's worktree is
**cleanup-owned, not capacity**. On a renamed board neither literal
matches, so every finished card keeps counting as a live holder. The
count only grows, the gate reaches zero room on a board with free slots,
and planning admission is withheld forever / every spawn is refused.
That is the **mirror** of the breach these commits fixed, and strictly
worse: 8 planners on a 4-slot board is visible; a permanent stall is
silent. The recorded reason even names the worktree budget, which the
operator then checks and finds has room.
## The conversion
`resolveProjectColumnsForRoles(store, ["complete", "archived"])` —
project-level, because the ledger spans the whole board with no single
task to resolve against. Matches triage's existing use in
`sweepStalePlanningStatuses` and executor's at the wip gates.
Legacy-seeded, so a default board still excludes exactly `done` and
`archived` — byte-identical there.
## Both conversions were UNCOVERED when written
Measured with #3214's blinding procedure **before** writing tests:
reverting either to the literals left **all 19 tests in the capacity
suites green**. Nothing in the tree could tell the conversion from what
it replaced — which is how the literals got there in the first place.
Each now has a renamed-board case that fails when blinded:
```
triage converted 2 passed | BLINDED 1 failed | 1 passed | restored 2 passed
executor converted 8 passed | BLINDED 1 failed | 7 passed | restored 8 passed
```
## The pairing earned itself immediately
Both new cases assert an **absence** (no throttle / no refusal), so each
is paired with a positive proving the gate still fires on the same
renamed board when a card genuinely holds the last worktree.
That caught a real defect in my own fixture: the candidate scan resolves
each task's **own workflow selection**, not `listWorkflowDefinitions`,
so my first version fell back to the default board where `drafting`
isn't a hold lane. No card was eligible, nothing throttled, and the
absence assertion **passed for the wrong reason**. The positive failed
and exposed it. Recorded at the fixture so the next reader doesn't
reintroduce it.
## Verification
```
42 tests across 6 capacity suites pass
check-fnxc-future-dates green
check-inert-sync-lane-conversions green
check-lane-wiring green
check-sql-column-literals green
census --strict green (BACKLOG ZERO restored)
```
No changeset: internal engine fix, no published-package surface change.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Completes the family. `check-fnxc-future-dates` was #3287,
`lifecycle-column-census` is #3289, and this is the third and last gate
that rewrote its baseline during a plain check.
## Reproduction
```
inflated one allowance by 6, ran the gate with NO flags
rc=0
entry RESET to 1 ← the check modified the tree it was checking
```
## Why it matters
The tightening is right in substance — an allowance nobody spends is a
hole a literal can be regrown into. Performing it as a **side effect of
checking** handed every worker a byte-identical uncommitted diff they
had not authored, which they then reasonably committed.
Measured across the family: nine PRs chased three defects on
2026-07-31/08-01, two of them (#3283/#3285, five minutes apart, `+0/-1`
each) deleting the **same line neither author wrote**.
#3289 states the class best — *a check that writes turns every reader
into an author*. Two of us separately mis-attributed a gate-written
baseline to our own work while debugging something else.
## Measured, all four directions
| scenario | result |
|---|---|
| plain run, stale baseline | `rc=0`, reports `allowed 7, now 1`;
**inflation survived** — read-only |
| a new SQL literal added | **`rc=1`**, names `__sql_probe.ts` —
regression detection intact |
| `--update-baseline` | `rc=0`, entry written |
| clean tree, plain run | `rc=0`, **zero files dirty** |
Row 2 is the one worth checking: a read-only change to a gate is
worthless if it also stops catching the thing it exists for. The rise
path is untouched.
`census --strict` 0, `check-fnxc-future-dates` 0, eslint clean.
## Correcting my own delay
I measured this defect family on #3267 and then **declined to fix two of
the three**, reasoning that the census *"deliberately fails on a drop"*
so the port might be unsafe. That was wrong: it tightened and exited
`0`, exactly as its own test asserts — *"TIGHTENS on a drop and exits 0,
so somebody else's merge cannot redden the gate."* I had read the
`--exact` contract and attributed it to the default path.
The caution cost hours and prevented nothing. #3289 was written by
someone else in the meantime; this finishes what I should have finished
then.
## What
My blind-spot table in #3251 audited **one axis**. Adds the one that
missed three defects. Docs only.
That table records what each of the five lifecycle ratchets can and
cannot **see**. I probed that carefully — several spellings per tool —
and then wrote *"nothing found; sound"* for two of them.
Within a day, three of those same tools turned out to share a completely
different defect: **they wrote to the tree they were checking**,
auto-tightening their own baseline during a plain check run.
| gate | wrote during a check | fixed by |
|---|---|---|
| `check-fnxc-future-dates` | yes | #3287 |
| `lifecycle-column-census` | yes, under `--strict` | #3289 |
| `check-sql-column-literals` | yes | #3292 |
**No number of detection probes could have surfaced that.** The table
asserted one property carefully and said nothing about the other *while
reading as comprehensive* — which is precisely the failure it documents
in the tools it audits.
## The rule it adds
1. **What can it see?** — probe each spelling of the thing it claims to
catch.
2. **Can it fail at all?** — invoke it as `package.json` does; a
report-only run exits 0 forever (#3255).
3. **Does it write?** — `git status --porcelain` before and after, on a
clean tree.
With the trap on the third spelled out: these gates write only when a
tightening is **available**, so a clean tree after a run proves the
*trigger* is absent, not that the tool is read-only. Inflate a baseline
entry first, then run it. I hit exactly this while reviewing #3292 — ran
all three gates on main, saw a clean tree, and had to stop myself
concluding the SQL gate was fine.
## Why the pattern, not the people
Three tools converged on write-during-check independently. That argues
the design is **attractive**, not that three authors were careless: the
tightening is correct, the write saves a step, and the message even
tells you to commit it. It only becomes a defect at the moment a second
person runs the same gate — which is invisible from inside any one of
them.
What it cost, measured: #3283 and #3285 are the same `+0/-1`, five
minutes apart, by two authors, **neither of whom wrote that line**.
```
lint clean; fnxc-future-dates clean
```
Last loose thread from the 2026-07-31 stamp-repointing wave. Two comment
lines.
```
line 1027 ConcurrencyAdmission 2026-07-31-09:00 -> 2026-07-21-22:30 (eef5eb751e)
line 1471 WorkflowLifecycleColumns 2026-07-31-05:00 -> 2026-07-30-20:55 (109204c590)
```
## Not a revert — neither value was ever right
| stamp | originally | after #3280 | authoring commit (UTC) |
|---|---|---|---|
| `ConcurrencyAdmission` | `2026-08-06-09:00` (16 days ahead) |
`2026-07-31-09:00` (10 days late) | **2026-07-21 22:30** |
| `WorkflowLifecycleColumns` | `2026-08-01-05:00` (~1.5 days ahead) |
`2026-07-31-05:00` (~1 day late) | **2026-07-30 20:55** |
Both were written **ahead of their own commits** to begin with. Three
lanes then repointed stamps to turn `main` green (#3261, #3269, #3280),
moving the **date** back a day while keeping the clock time — which
converts an hours-off stamp into a days-off one in the opposite
direction. #3282 reverted the batch it owned; these two were outside its
scope.
So restoring the originals would be wrong too. The defensible values are
the authoring commits' UTC timestamps, per the `date -u` rule #3281
settled.
## Why now
`scheduler.ts` is a hot file. This survived two successive claimants — I
flagged it on #3262 and again on #3288 rather than opening a conflicting
PR, and said I'd take it once the file was unclaimed.
`check-file-claimed` now reports UNCLAIMED, so here it is.
## Scope
The gate is **green either way** — #3277 fixed the comparison, so
nothing is blocked by this. It is purely about the FNXC trail recording
when the work actually happened, which is the only reason the trail
exists. A stamp that satisfies a check while misstating the date by ten
days is worse than no stamp.
## Verification
```
check-fnxc-future-dates green
check-inert-sync-lane-conversions green
check-lane-wiring green
check-sql-column-literals green
census --strict green
```
Diff is two comment lines — no executable change. No changeset.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extends the doc from #3255/#3273 with the failure that cost the most in
a single session: **one stale install produced five wrong reports on one
issue** (#3264).
## What happened
A `node_modules` that had drifted from the lockfile — `jsdom@29.0.1`
installed, `29.1.1` pinned — generated failures that existed on no CI
machine and no other checkout. They were not subtle: deterministic,
reproducible on demand, with plausible stack traces and real-looking
assertion diffs.
Each round of triage got **more precise about the wrong data**:
| round | claim | why it was wrong |
| --- | --- | --- |
| 1 | "4 deterministic failures" | measured in a 4-file batch, called it
isolation |
| 2 | "3 deterministic, 2 order-dependent" | isolated correctly, but a
race is not deterministic |
| 3 | "TaskCard is broken" | stale jsdom; the CSS assertion was correct
|
| 4 | "no contamination" | true of four app files; published unqualified
|
| 5 | "quarantine these two" | never read the failure text — both were
timeouts |
The through-line is not carelessness about the code. **The environment
was never treated as part of the claim**, so no amount of care about the
analysis could recover it.
## The checks, in the order they cost the most
```bash
pnpm install --frozen-lockfile # node_modules is not evidence until it matches the lockfile
<run the file ALONE, 3+ times> # isolation and repetition answer different questions
<read the failure TEXT> # a timeout and an assertion failure need opposite responses
uptime # a loaded box manufactures timeouts that mean nothing
```
## Why the load check earned its place
Two tests "failing" in a full-suite run were `Test timed out in 15000ms`
on a box at **load average 9.7 with 84 users**. Under AGENTS.md's
quarantine-on-sight rule that reads as a flake to quarantine — and the
ledger's **14-day deletion ratchet would have made the lost coverage
permanent**.
The rule presumes the failure is a property of the test, not of the
machine. A wall-clock budget crossed under local contention is evidence
about the hardware. I was one comment away from deleting healthy
coverage on that basis.
## The tell
A finding is environment-derived when it is **local, recent, and
unshared**: nobody else has reported it, CI is green, and it appeared
without a commit that could explain it. Any two of those should stop a
report before it is written. All three applied here, and the report went
out anyway — five times.
## Verification
Docs only; no code paths change. `fnxc-future-dates`,
`lifecycle-columns`, `quarantine-ledger` exit 0. No changeset — internal
docs are excluded.
Context: the one finding in #3264 that survived all five rounds is #3286
(merged), and it survived because it was verified by **reverting the
product change** rather than by trusting a red — 3/3/2 failures without
the fix, 27/27 across four runs with it.
## What
Pins the **worktree-capacity arithmetic** — the gap #3262 measured,
named, and explicitly left for someone to claim. Two commits: a
behaviour-preserving seam extraction, then the test.
#3262's own scope note:
> blinding this predicate to `false` leaves all 22 scheduler suites
green (365 tests). The capacity logic it feeds has no behavioural
coverage at all.
## Two live defects, opposite directions
Both came out of these few lines:
- **UNDER-COUNT admits work over the cap.** `maxWorktrees=4`, four
planning sessions each holding a worktree, and a replan dispatch
admitted as the **fifth** — the ledger counted WIP cards only and never
learned to count planners.
- **OVER-COUNT self-deadlocks.** A planned Ready card *retains* its
planning worktree for execution reuse, so counting it as a holder blocks
its own release: `2 wip + 3 idle-held = 5/4`, and the first unpause
released 2 of 4 slots' worth of work.
**Both are pinned, and the asymmetry is why.** Under-counting breaks the
cap and lets real work over it; over-counting only starves dispatch. A
test covering the "safe" direction alone would leave the expensive one
open.
## Mutation-tested — all four caught
| mutation | result |
|---|---|
| drop the terminal exclusion | 1 failed / 7 passed |
| count WIP cards twice | 2 failed / 6 passed |
| count cards holding no worktree | 1 failed / 7 passed |
| drop the self-slot subtraction | 1 failed / 7 passed |
```
clean: 8 passed (8)
scheduler suite: 16 files / 151 tests passed (behaviour preserved by the extraction)
typecheck, lint: clean
```
## Scope, stated rather than implied
The terminal predicate is **injected**, not resolved here. Which lanes
are terminal is #3262's test; resolving it in this file would make it
fail for that reason instead of this one. This pins the **set
arithmetic** — who is excluded, and how the total is formed.
Still not covered, and I am not claiming otherwise: the *stateful* half
of the ledger — the `+= 1` on dispatch and the `Math.max(0, … - 1)` on
failure inside `schedule()`'s loop. Extracting that would mean
restructuring dispatch itself, which is a different change from this
one.
## Process note
I claimed this on #3262 **before** starting rather than after, because
`scheduler.ts` is the hottest file in the tree and I produced three
duplicate PRs earlier tonight by picking up small shared-surface work
someone else already had in flight. Announcing first cost one comment;
the duplicates cost three PRs and two closes.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved scheduler capacity calculations for tasks that retain
existing worktrees.
* Prevented WIP tasks from being counted twice.
* Excluded completed and worktree-less tasks from reserved capacity.
* Corrected candidate capacity calculations when no worktree capacity is
reserved.
* **Tests**
* Added coverage for worktree reservation totals and candidate reuse
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## What
**Port of #3287 to the sibling tool.** `lifecycle-column-census.mjs
--strict` called `writeBaseline()` during a plain **check**, so running
the gate modified the tree it was checking.
```
clean: 0 files dirty
$ node scripts/lifecycle-column-census.mjs --strict # no --update-baseline
rc=0
after: M scripts/lib/lifecycle-column-census-baseline.json
```
## Why it matters — measured by #3287, reproduced here
#3287 established what this costs: every worker who runs the gate
receives a **byte-identical uncommitted diff they did not author**, and
reasonably commits it. #3283 and #3285 are the same `+0/-1`, five
minutes apart, by two different authors, **neither of whom wrote that
line** — the gate wrote it in both checkouts.
I hit this one the same way, which is the part worth recording: I saw a
modified baseline on my own branch and started reasoning about where
*my* change had touched it. It had not. A check that writes turns every
reader into an author.
The tightening is right in substance, and this tool's `COMMIT IT`
message made the diff *explained* rather than mysterious — better than
fnxc's was. **Neither addresses the mechanism.**
## The shape, matching #3287
Still computed, still reported loudly, written only under an explicit
`--update-baseline` (which has its own path above and is untouched):
```
lifecycle-column-census --strict: baseline CAN BE TIGHTENED — the tree has fewer guards than it allowed
packages/engine/src/scheduler.ts: allows 1, tree has 0
Not written. Record it deliberately, so the diff has one author:
node scripts/lifecycle-column-census.mjs --strict --update-baseline
```
**A plain run stays green rather than failing.** Guard counts drop when
someone *else's* merge removes a literal, so failing on a tightening
would redden main on a change the author never made. Report, don't
enforce — same reasoning #3287 gives for stamps aging into the past.
## Measured, all three directions
| scenario | result |
|---|---|
| plain `--strict`, stale baseline | reports + hint; **tree clean**
(was: 1 file dirty) |
| `--strict --update-baseline` | writes, rc=0 |
| a new guard added | **rc=1** — regression detection intact |
```
lint clean
```
## Note
Claimed on #3287 before starting, since it is that author's fix and they
may have had the port in flight. The two differences from the fnxc case
are noted there: this one fires under `--strict` rather than a bare run
(but `--strict` is what `package.json` and CI invoke, so it is the
common path), and its message was already loud.
#3286 fixed a real user-facing regression and **shipped no test**, so
nothing stops it returning. The regression was mine.
## The bug
#3215 (mine) added `isArchivedColumn` to an effect's dependency list to
keep the task fetch honest. That effect **also owned four `setState`
calls**, and `isArchivedColumn` is a `useMemo` over
`useBoardWorkflows()` — which revalidates asynchronously.
Every revalidation re-ran the reset over whatever the operator had
typed. A title entered before the workflows settled silently reverted to
`Research: <heading>`, and the task was created with a title nobody
wrote.
## Why my own four tests could not see it
Every existing case in this file asserts the **filtered task list** —
render, await `fetchTasks`, read the datalist. **None types into the
form.**
I tested what I added and not what I touched. That is why the regression
belongs in this file rather than a new one: the gap is this file's.
## The case
Renders with `boardWorkflows: null` — the state when an operator opens
the modal and starts typing — types a title with per-character
`userEvent`, then rerenders with a resolved workflow set (a **new object
identity**, which is the entire mechanism) and asserts the typed text
survived.
Two details that each cost a cycle, recorded at the site:
- **`fetchTasks` is not awaited.** It runs only in enrich mode, while
the title field exists only in create mode — so the reset effect, not
the fetch, is under test. My first version waited on it and failed for
the wrong reason.
- **`userEvent.type`, not `fireEvent.change`.** The documented failure
is state overwritten between renders; a single synthetic change event
can land after the reset and mask it.
## Measured both directions, on main `6834ba35bd`
| state | result |
|---|---|
| fixed main | **5 passed** |
| dependency re-added to the reset effect (my bug) | **1 failed / 4
passed** — and only that case |
The second row is the point: it fails on precisely the mutation that
recreates the defect, and leaves the four archived-lane cases green — so
it pins the regression without duplicating what is already covered.
`eslint` clean, `check-fnxc-future-dates` 0. Test-only.
## Note on provenance
Getting this measurement took three attempts: a `git checkout` of the PR
branch silently failed (stderr suppressed), so I twice ran against the
wrong tree and nearly concluded the fix did not work. HEAD and
dirty-count are printed beside every number above for that reason.
Found by chasing the deterministic half of #3264 (dashboard red on
`main`). **The tests were right; the product is broken.**
## The bug
Open **Create Task** from a research finding, type a title before the
board workflows settle, and the field silently reverts to the derived
default `Research: <heading>`. The task is then created with a title the
operator did not write. `description`, `priority` and `taskId` reset the
same way.
`ResearchTaskActionModal` reset those four fields in the same effect
that fetched the task list, and that effect's dependency list carried
`isArchivedColumn`:
```ts
const isArchivedColumn = useMemo(() => { … }, [boardWorkflows]); // useBoardWorkflows() — async
useEffect(() => {
setTitle(`Research: ${finding.heading || run.title}`); // ← re-runs on every revalidation
…
}, [open, mode, projectId, finding.heading, preview, run.title, isArchivedColumn]);
```
`useBoardWorkflows` resolves and revalidates asynchronously, so the
memo's identity changes and the reset re-runs over whatever the operator
has typed.
Introduced by #3215, which correctly added the archived-column filter
but hung its dependency on an effect that also owns form state. Same
class as the documented
`docs/solutions/ui-bugs/skill-autocomplete-highlight-reset-on-swr-revalidation.md`.
## The fix
Split into two effects: the reset depends only on what it derives from;
the fetch keeps `isArchivedColumn`. No behaviour change to the archived
filter — #3215's guard is untouched.
## Verification, both directions
The three standing `ResearchView` tests fail without this and pass with
it:
```
isArchivedColumn back on the reset effect: 3 failed | 24 passed (27)
as committed: 27 passed (27)
```
## What I tried and removed, because it matters
I wrote a dedicated invariant test (per "fix the invariant, not the
repro") asserting that *all* typed fields survive a revalidation. **I
deleted it, because it did not work.**
- First draft used `mockImplementationOnce` to defer
`fetchBoardWorkflows`. `ResearchView` resolves board workflows on mount,
so that once-implementation was consumed before the modal opened.
Reverting the product fix left the test **green** — it proved nothing.
- Second draft deferred *every* call. `beforeEach` uses
`vi.clearAllMocks()`, which clears calls but **not implementations**, so
the deferral leaked into later tests and left `fetchBoardWorkflows`
permanently pending — masking two of the three genuine failures. The
revert then showed `1 failed` instead of `3`, i.e. my test was hiding
real bugs.
Rather than ship a regression test that cannot regress, I removed it.
The three existing tests already fail without the fix, which is real
coverage; a broader invariant test needs a modal-level harness that
resets implementations between cases, and that is worth doing properly
rather than badly here.
## Scope
Also in #3264: `TaskCard.badge-wrap` (1 deterministic failure, unrelated
— CSS/layout), and `useChat` / `WorkflowNodeEditor` /
`PlanningModeModal`, which pass standalone and are cross-file
contamination, not product bugs. Untouched here; the issue has the
per-file matrix.
No changeset — `@fusion/dashboard` is private.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Preserved form entries during workflow revalidation in the research
task modal.
* Limited task selection to active workflow columns when enriching
findings.
* Prevented outdated task results from replacing newer selections.
* Improved loading and task-list behavior when source findings or modal
state changes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Root-cause fix for the duplicate-PR pileup tracked in #3267. **Running
the check modified the working tree.**
## Reproduction
```
clean: 0 files dirty
$ node scripts/check-fnxc-future-dates.mjs # no flags, no --update-baseline
exit 0
after: 1 file dirty → M scripts/lib/fnxc-future-dates-baseline.json
```
`:227` auto-tightened and `writeFileSync`'d on every run.
## Why that produced nine PRs
The tightening is **right in substance** — the comment above it explains
why banking a stale allowance is worse than re-recording. Doing it as a
*side effect of checking* is what hurt: every worker who ran the gate
received an identical uncommitted diff they had not written, and
reasonably committed it.
The clearest evidence is #3283 and #3285 — five minutes apart, `+0/-1`
each, both deleting the same baseline line. **Neither author wrote that
line.** The gate wrote it, in both of their checkouts.
I also mis-attributed my own dirty tree to leftover work while
retracting a measurement on #3277/#3278. The dirt was this script.
## The change
Still computed, still reported loudly — only **written** under
`--update-baseline`:
```
[check-fnxc-future-dates] baseline CAN BE TIGHTENED for 1 file(s):
packages/cli/src/__tests__/cli-active-count-lanes.test.ts: 10 -> 5
run `pnpm check:fnxc-future-dates --update-baseline` to record it (one commit, one author)
```
**A plain run stays green rather than failing on a tightening.** Stamps
age into the past on their own, so failing would redden main on a clock
tick — which is precisely why the auto-write existed. Report, don't
enforce.
## Measured, both directions
| scenario | result |
|---|---|
| stale allowance, plain run | reports + hint; baseline **unchanged**
(verified still inflated at 10) |
| stale allowance, `--update-baseline` | `baseline written: 122 stamp(s)
in 63 file(s)`; value reset to 5 |
| clean tree, plain run | exit 0, **zero files dirty** |
| `census --strict` / eslint | 0 / clean |
The first row is the one that matters: I inflated an allowance, ran the
check, and confirmed the file was **still inflated afterwards**.
Asserting only "exit 0, no diff" would have passed even if the write had
silently succeeded and produced no net change.
## Scope
One script. CI is unaffected — it never committed the side-effect write,
so that write was always discarded there. The only behaviour change is
that an interactive run no longer edits your tree.
This is a smaller intervention than the claim-protocol I proposed
earlier in #3267, and I now think that one was treating a symptom:
workers were not colliding because they lacked a protocol, but because
the tool handed each of them the same diff.
**Rebased. The census claim in my original title was overtaken — this is
now a correction, not a conversion.**
## Census: 0 before, 0 after
#3261 got there first, by **recording** the fallback rather than
converting it. Its `DELIBERATE` reasoning is correct and I kept it
verbatim.
## What this corrects
That note says:
> Recorded rather than converted because **there is nothing to convert
TO**.
There is. **`isTerminalColumnRole` in core is this predicate, term for
term** — verified against `column-roles.ts` rather than assumed:
| | hand-rolled | `isTerminalColumnRole` |
|---|---|---|
| flags present | `flags.complete === true \|\| flags.archived === true`
| same, via the two role helpers |
| flags undefined | `columnId === "done" \|\| columnId === "archived"` |
same, via `LEGACY_COMPLETE/ARCHIVED_COLUMN_ID` |
And the helper's own doc names this exact case — it exists *"because the
pattern `column !== \"done\" && column !== \"archived\"` is the single
most repeated shape in the backlog"* and *"keeps callers from
re-deriving it and from accidentally dropping one half."*
**The rest of #3261's argument stands and is preserved.** The
undefined-flags arm is a **live** path, and treating an unreadable
workflow as non-terminal would count a finished card's retained worktree
against live capacity. That reasoning is about the *fallback's
existence*, not about *where the predicate lives* — and the shared
helper carries the identical fallback.
Second time in this file: `isWipColumnTask` two lines up records that it
was itself once *"a hand-rolled copy of `isWipColumnRole`"*. That's an
argument for the helper being easy to miss, not for anyone being
careless.
## Coverage, stated rather than implied
**Blinding this predicate to `false` leaves all 22 scheduler suites
green (365 tests)** — the capacity logic it feeds has no behavioural
coverage at all.
The added test pins the **lane vocabulary** (both renamed terminal
lanes, the non-terminal lanes, the legacy fallback). It does **not** pin
the capacity arithmetic, which stays unguarded and belongs to that
gate's owner. Under-counting is the dangerous direction: the commit
adding the gate reports `maxWorktrees=4` with **a fifth worktree
admitted**.
## Verification
- 23 scheduler suites — **368 green**; `tsc` clean
- Census 0 → 0; DELIBERATE count unchanged at 148; inert ratchet green
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**The date gate passes on `main` — but not because this was fixed.**
## What actually happened
UTC rolled over to `2026-08-01`. The gate compares against the later of
local and UTC, so two `2026-08-01` stamps in `scheduler.ts` became valid
on their own. That is the ratchet's normal drop path and is fine.
`#3278` then re-recorded the baseline "after the UTC rollover", which
set `scheduler.ts` to **allow 1** — exactly enough to absorb the one
stamp that did *not* age out:
```
FNXC:ConcurrencyAdmission 2026-08-06-09:00 ← six days out, wrong on any calendar
```
So the gate reports `123 known future-dated stamp(s), none added` and
exits 0, with a stamp inside it that will not be valid until next week.
## Why this is the failure the gate exists to catch
A blanket re-record cannot distinguish **aged out** from **still
wrong**, so it launders the second past the first. The sibling ratchet
states the rule outright:
> Do NOT re-record the baseline to clear this — that is the same false
green one layer up.
This is that, one layer up again: not a guard cleared by a baseline, but
a *baseline refresh* clearing a guard as a side effect.
## The fix
- stamp repointed to `2026-08-01` — today in UTC, which is the calendar
the gate actually compares against
- **allowance removed**, not left at 1, so the entry cannot be regrown
into
**Mutation-verified**: with the allowance gone, restoring `2026-08-06`
exits **1**. Before this change the same stamp exited **0**. That is the
whole point — the ratchet can now see it.
## One thing worth carrying forward
A six-days-out stamp is not a timezone slip. Neither the old `date -u`
guidance nor the current local-date guidance in AGENTS.md would have
prevented it, and CI-only checking cannot catch it before merge. This is
the concrete case for running the date check at author time, which I
have flagged but not landed since it changes the gate's contract.
## Verification
- `check-fnxc-future-dates` — exit 0, allowance removed
- `scheduler` suites — **148 pass**
- `tsc --noEmit` (engine) — 0 errors
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#3279. Undoes the damage my #3261 did, now that #3277 has landed
and made it safe.
## What went wrong
#3277 established that last night's "future-dated" stamps were
**correct** — the author's local date in a UTC+1 container, written
minutes before their commits. The gate compared against the runner's
local calendar (PDT) and called them tomorrow.
I diagnosed it as author error and repointed seven stamps to turn main
green. The values I wrote were **neither the author's local time nor
UTC** — invented times chosen to satisfy a broken check. The FNXC record
is this project's why-does-this-exist trail, so those stamps misstated
when the work happened.
## Restored verbatim
| file | mine (wrong) | restored |
|---|---|---|
| `workflow-column-boundary-capacity.test.ts` | `22:30` |
`2026-08-01-00:30` |
| `runtimes/in-process-runtime.ts` | `22:20` | `2026-08-01-00:20` |
| `scheduler.ts` (`MissionReconciliation`) | `22:00` |
`2026-08-01-00:00` |
| `workflow-column-boundary-hooks.ts` | `22:20` | `2026-08-01-00:20` |
| `workflow-column-boundary.ts` ×2 | `22:20` | `2026-08-01-00:20` |
| `workflow-graph-task-runner.ts` | `22:20` | `2026-08-01-00:20` |
## The check that mattered
Sequencing was deliberate — #3277 had to land first or this would have
re-reddened main. The real question is whether the gate now accepts the
**originals**, measured across the rollover boundary at local
`2026-07-31 17:23 PDT` / UTC `2026-08-01 00:23`:
```
America/Los_Angeles exit 0 Europe/Paris exit 0
UTC exit 0 Asia/Tokyo exit 0
```
`123 known future-dated stamp(s), none added`. **No baseline change
needed** — #3278's pruning already re-recorded `scheduler.ts`, and these
are known stamps rather than new ones.
Stamps only: `git diff` shows **zero** non-FNXC lines, 7 insertions / 7
deletions across 6 files. `census --strict` 0, `pnpm test:gate` 0.
## The part worth keeping
I argued against exactly this on #3263 — *"it rewrites stamps whose
authors are not us"* — and then did it myself six lines later, because I
was confident about a cause I had not checked. The commits' timestamps
were available the entire time; I read the runner's clock and never
asked what timezone the **author** was in.
Four of last night's seven PRs were fixing something that was not
broken. This is the cleanup for my share of that.
#3277 fixed the gate correctly and I have no argument with the code:
`today = max(localToday, utcToday)`, so a stamp is future only if ahead
of **both** calendars. That is the right shape for a fleet spread across
timezones.
It reversed the **authoring rule** along with it, and that part is
backwards:
> **Write your own local date and a real clock time.** … Do NOT reach
for `date -u`
Under #3277's own comparison, that reintroduces the failure it just
fixed.
## Measured against the merged gate, on current main
```
stamp 2026-08-02 (a UTC+2 author's local date at 22:00 UTC) gate exit=1 REJECTED
stamp 2026-08-01 (the same author using date -u) gate exit=0 ACCEPTED
```
Run today, 2026-08-01, with the runner in PDT. Probe file added and
removed; tree clean after.
## Why `date -u` is the only safe rule
The bound is `max(localToday, utcToday)`, and **UTC only moves forward
between writing a stamp and checking it**. So a `date -u` stamp has
already been passed by the bound at check time, from every timezone,
always. No other rule has that property.
Writing your own local date is safe *only if you are not east of UTC*.
During a UTC+2 author's evening their local date is already tomorrow in
UTC, and the stamp is rejected until UTC catches up hours later — which
is exactly the "five reds in two hours" incident #3277 diagnosed. The
gate change widens the window enough that CI usually catches up before
anyone looks, but "usually, after a delay" is a race, not a rule, and it
fails hardest for the authors furthest east.
The prior instruction (`date -u`) was correct; what was wrong was its
stated *rationale* ("validates against UTC"), which is what I was fixing
in #3276 before #3277 landed. This PR keeps #3277's both-directions
history — the part that explains why neither naive rule works on its own
— and restores the prescription.
## Why not just comment on #3277
It is merged, and AGENTS.md is the file every agent reads before writing
a stamp. Leaving the inverted rule in place for a review cycle means
every east-of-UTC author in the fleet follows it. Filed as a PR so it
can be judged on the measurement rather than on my say-so — if the
numbers above are wrong, this should be closed.
**Supersedes #3276**, which documented the pre-#3277 mechanism and is
now stale. I will close it once this is judged.
Docs only; no changeset (AGENTS.md is excluded).
## What
**Three separate commits landed a future-dated FNXC stamp this
evening**, each turning this blocking gate red on main (#3261 fixed six
across five engine files; #3270 fixes a third in core). This makes the
failure message actionable. Tooling only.
The message said *"Use the current date and a real clock time."* That
tells the author to use the value they already believed they had. It now
prints the exact stamp:
```
Current UTC stamp to use: 2026-07-31-23:39
```
Copy-paste instead of a second judgement call, computed only on a path
that has already failed.
## Why a message change rather than a rule
**The offsets are the evidence.** `00:50` against `23:34`; the engine
batch similar — consistently **1–2 hours into tomorrow**, not wrong
dates. That is the shape of a clock or timezone difference, not
carelessness, and no amount of restating the rule fixes a clock.
AGENTS.md already says to take the stamp from `date -u`; three actors
violated it in one evening anyway.
**I am one of those actors** — I have broken this rule twice today. So
this is not a complaint about anyone's diligence; it is an argument that
the instruction is doing less work than a printed value would.
## I made the same mistake inside this change
The FNXC comment documenting the fix was stamped **two minutes ahead**
of the real UTC time. Corrected from `date -u`.
**The gate did not catch it** — it scans `packages/` and not `scripts/`,
so FNXC stamps in the tooling itself are entirely unchecked. That is a
genuine scope gap and I am reporting rather than closing it: pulling
`scripts/` into scope would surface existing stamps across the tooling
and needs its own baseline pass, which does not belong in a message fix.
There is something clarifying about writing a future-dated stamp *in the
fix for future-dated stamps*, in a file the checker cannot see. It is
the same lesson this whole session kept producing — **an instrument's
blind spot is invisible in exactly the way its subject is** — and I
walked into it while holding the flashlight.
```
lint clean; gate prints the stamp on failure
```
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Improved validation feedback for future-dated entries by showing the
exact UTC timestamp format and value to use when corrections are needed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
**`check-fnxc-future-dates` exits 1 on `origin/main`.**
```
packages/engine/src/scheduler.ts: 3 future-dated stamps, baseline allows 2
FNXC:ConcurrencyAdmission 2026-08-06-09:00 (six days out)
FNXC:WorkflowLifecycleColumns 2026-08-01-05:00
FNXC:WorkflowScheduling 2026-08-01-01:05
```
All three repointed to `2026-07-31`, times preserved. Gate now exits 0.
## This red has outlived three owners
#3270, #3272 and #3274 were each opened against it and each **closed
without merging**. Main has been red on this gate for hours while three
fixes came and went.
Claimed with `check-file-claimed.mjs` before starting — only #3262
touches `scheduler.ts`, and it is a terminal-role refactor rather than a
stamp fix, so this was genuinely unowned.
## Why this keeps recurring
Seven incidents in roughly two hours. The mechanism, in one line: **the
date check runs only in CI** (`pr-checks.yml:66`, no pre-commit or
pre-push hook), so every PR is validated against main's baseline *at its
own CI time* and cannot see a concurrent or later change. Two PRs
stamping the same file both pass, then compose into a red main. One case
(#3273) was a stale branch **reverting** an already-merged fix.
Patching instances has not converged — this PR is the eighth attempt at
the same class. Two structural options, neither of which I am landing
unilaterally since the second changes the gate's contract:
- run the date check at **author time** (pre-push); it needs no baseline
for "is this date in the future", so it cannot be raced
- make the date rule **baseline-free** — a future-dated stamp is always
wrong, unlike a lifecycle literal that may be a deliberate fallback
`2026-08-06` being six days out also suggests these are not off-by-one
timezone slips but stamps written from an intended future date.
## Verification
- `check-fnxc-future-dates` — **exit 0** (was exit 1 on main)
- `scheduler` suites — **148 pass**
- `tsc --noEmit` (engine) — 0 errors
- comment-only diff, no behaviour change
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lifecycle-column-census test file spawned the census CLI ~14 times, each
parsing every tracked source file to a TypeScript AST (~2s for ~1960 files).
Many spawns were byte-identical, deterministic, read-only real-repo scans:
the --json census 4x, the plain report 2x, plus a repeated identical
--update-baseline tree-sync across the ratchet cases. Memoize each distinct
read-only spawn's output (keyed by argv) and reuse the synced baseline JSON,
collapsing duplicate full-AST scans without changing any assertion.
File wall-time: 30.5s -> 18.6s (-39%). 57/57 tests still pass.
Fusion-Task-Id: FN-slow-test-census
Root-cause fix for tonight's repeated red `main`, instead of repointing
stamps one at a time — **four PRs across three lanes did that in ninety
minutes** (#3261, #3269, and my #3263 and #3272, two of which I closed
as superseded by concurrent work).
## The defect
The fleet writes stamps from **many** machines; this gate evaluates them
on **one**.
#2941 fixed the case where the author sits **west** of the runner — a
correct 5pm-in-California stamp read as "tomorrow" under a UTC
comparison — by switching to the runner's **local** calendar. The mirror
case was left open, and that is what broke `main`:
| commit | landed (PDT) | = UTC | stamp written |
|---|---|---|---|
| `9094d1640e` | 16:12 | 23:12 Jul 31 | `2026-08-01-00:20` |
| `e52da740a5` | 16:32 | 23:32 Jul 31 | `2026-08-01-00:50` |
| `3f95c6d53e` | 16:40 | 23:40 Jul 31 | `2026-08-01-01:05` |
Those are **neither** the runner's local date **nor** UTC. They are the
*author's* local date in a UTC+1 container — and they are **correct** by
this project's own convention ("authors write the local date"). The
gate, running in PDT, called all three "tomorrow" and reddened `main`
for every other lane.
## The fix
A stamp is future only if it is ahead of **both** the local and UTC
calendar dates.
- Accepts both honest directions (author east or west of the runner).
- **Preserves #2941**, doesn't revert it — west of Greenwich the local
date is the earlier of the pair, so the 5pm-in-California case still
passes.
- Still catches an invented date: `scheduler.ts`'s `2026-08-06` stamp
(six days out) remains counted, and a mutation probe at `2026-09-15`
fails the gate.
## AGENTS.md corrected in the same commit
It still instructed **`date -u`**, which describes the *pre-#2941* gate.
That instruction is now the one that **produces** the failure from any
machine east of the runner — I followed it myself earlier tonight and
repointed stamps that were already correct. Rewritten to say: write your
own local date; the gate accepts anything not ahead of both calendars.
## Baseline
Auto-tightened for **37 files** — the gate's no-author drop path. Those
allowances were false positives carried since the UTC-only era, so this
**strengthens** the ratchet rather than widening it (`scheduler.ts` 2 →
1, keeping the genuinely-invented stamp counted).
## Verification
```
check-fnxc-future-dates green
check-inert-sync-lane-conversions green
check-lane-wiring green
check-sql-column-literals green
census --strict green
MUTATION: FNXC:MutationProbe 2026-09-15-10:00 → gate fails (real future dates still caught)
```
No changeset: tooling/gate + internal docs.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## What
Re-records the future-dates baseline after the UTC rollover. **92 → 64
file entries.** Tooling only, no source changes.
The gate's own instruction is *"If a count went DOWN, re-record the
baseline in the same commit."* UTC is now `2026-08-01`, so every stamp
dated `2026-08-01` is dated **today** rather than after it, and **28
files no longer carry any future-dated stamp at all.**
## Why pruning matters more than the number
Those 28 files kept a non-zero allowance they no longer need, and **an
allowance is a hole a new violation can hide in.** Four separate commits
landed a future-dated stamp last evening — in that environment, a stale
allowance on a file is exactly where the fifth would go unnoticed. With
the entries pruned, the next one in any of those files is caught on the
first run instead of being absorbed silently.
This is the same direction as #3168 (*"tightens the allowance 1 → 0"*)
and #3211, just triggered by the clock rather than by a fix.
## What this is not
- **Not a correction to anyone's stamp** — no source file is touched.
- **Not a loosening** — no entry increases and no file is added.
- The **123 stamps still dated beyond today** (e.g. `2026-08-06`) keep
their existing allowances untouched.
```
baseline file entries 92 -> 64
check-fnxc-future-dates rc=0
lint clean
```
## Related, deliberately not included
`scheduler.ts:2258` carries `FNXC:WorkflowScheduling 2026-08-01-01:05` —
dated today, one hour ahead of the current clock. I repointed it while
preparing this change and then reverted: after the rollover it is no
longer a gate violation, and mixing a cosmetic timestamp edit into a
baseline re-record would make both harder to review. Noting it so the
residual I flagged when closing #3270 does not get lost — it is now an
accuracy nit rather than a gate concern.
## The hazard
The column backlog is **0**. The two largest numbers the census prints
are now `ROLE (12)` and `STATUS (185)`, sitting directly beneath it,
labelled only `(not guards)`.
That is a verdict with no reason. For a worker under a directive to
drive a census down — finding the backlog line already at zero and two
bigger numbers underneath — "(not guards)" is thin protection. This PR
puts the reason where the numbers are.
## Why they are genuinely not backlog
Both classify by **receiver**, not by the literal
(`ROLE_RECEIVER_TOKENS` = `role, agentType, agent, lane, capability,
sessionPurpose, surface, purpose, agentRole`; status matches
`/status/i`). A legacy column id next to one of those is a different
domain that happens to share vocabulary with the old board. Sampled from
the current tree, not reasoned:
| site | receiver | what it actually is |
| --- | --- | --- |
| `packages/cli/src/commands/task.ts:529` | `outcome === "archived"` | a
task **outcome** |
| `.../routes/register-chat-routes.ts:894` | `type === "done"` | a chat
**message type** |
| `.../cli-agent/telemetry-hub.ts:304` | `kind === "done"` | a telemetry
**kind** |
| `packages/cli/src/commands/goals.ts:178` | `status === "archived"` | a
**goal's** status |
| `packages/cli/src/commands/mission.ts:145` | `status ===
"in-progress"` | a **mission's** status |
None is a task column, so none has a workflow lane to resolve against.
Converting a goal's `status === "archived"` to a column trait would not
remove a legacy id — **it asks the wrong object for a lane it does not
have**, and the resulting bug would be invisible on the default board
for precisely the reason every inert conversion is. That is 185
opportunities to inject a real defect while a number goes down.
## Output-only, verified
Counts, JSON, baseline comparison and exit codes are untouched:
```
--strict exit=0
json totals: {"column": 0, "role": 12, "status": 185, "deliberate": 150} # byte-identical
```
60 census tests pass (`lifecycle-column-census.test.ts`,
`census-reclassification-message.test.ts`). `lifecycle-columns`,
`move-target-literals`, `inert-sync-lanes`, `quarantine-ledger` all exit
0.
## Provenance
I raised "role/status have no inertness proof behind them" several times
as a reason not to touch them, which was too weak — it implied the work
might be valid pending proof. Rather than leave that hanging I went and
looked. They are not unproven conversions; they are **not conversions at
all**. Correcting my own earlier framing, and putting the finding where
the next person will hit it instead of in a report they will not read.
No changeset — internal tooling.
Extends the doc merged in #3255 with two more instances of the same
pattern, both found this session, **neither involving a ratchet**. Four
instances now, from four unrelated directions:
| what was read as "pass" | what the green actually meant |
| --- | --- |
| `node scripts/check-*.mjs` exits 0 | report-only mode — the failure
path needs `--strict` |
| a census reports 0 for a new file | the file is untracked, so it was
never scanned |
| a backgrounded `cmd > log; grep …` reports exit 0 | that is `grep`'s
status; the suite inside had 8 failures |
| a rebased branch's tests pass | the rebase never started, so it ran on
the **old** base |
The two new ones are worth writing down because they are not about
tooling anyone built here — they are about how results are read.
**Exit codes belong to the last command in the pipeline.** A
backgrounded `run_tests > log 2>&1; echo done; grep X log` exits with
`grep`'s status, so the harness reported "completed, exit code 0" for a
dashboard suite that had 8 failures. I nearly recorded that suite as
green. Read the summary out of the log; never infer a suite's result
from a wrapper's exit code.
**A failed rebase leaves you on the old base, and the tests still pass
there.** `git rebase` refused with `cannot rebase: You have unstaged
changes`, so the branch never moved. `git diff origin/main` then listed
20+ files including other workers' commits — which reads exactly like my
branch had reverted their work — and a full test run on that tree came
back green. Both signals were true about a tree nobody cared about.
```
git merge-base --is-ancestor origin/main HEAD
```
said STALE while the tests said pass. That is the only check that
separates the two, and it belongs before any claim of "verified on
current main".
The shared tell, stated once: **a result too clean, or too alarming, for
what changed.** Every probe shape passing including ones that obviously
should not; a two-file branch appearing to revert twenty. When the
answer does not fit the size of the question, find out what was actually
measured before believing it.
## Verification
Docs only; no code paths change. `lifecycle-columns`,
`move-target-literals`, `inert-sync-lanes`, `quarantine-ledger` all exit
0. No changeset — AGENTS.md excludes internal docs.
**Pre-existing red, not from this branch:** `check:fnxc-future-dates`
currently fails on main from a `2026-08-01-00:50` stamp in
`packages/core/src/task-store/lifecycle-ops.ts` (commit `e52da740a5`) —
a timezone-ahead clock writing tomorrow's date, at 23:45 UTC. Already
claimed by **#3269 and #3270**, so I have not touched it; flagging only
so this branch's CI result is not misattributed. It is the same
recurring class this doc's sibling rule addresses: take the stamp from
`date -u`, not the local clock.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added guidance for identifying misleadingly successful CI and test
results.
* Documented checks for report-only runs, untracked files, masked
failures, and tests running on an outdated code base.
* Included recommendations for reviewing logs and verifying branch
ancestry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>