Completes `proving-a-code-path-actually-runs.md` (merged as #2642) with
the rule its own author broke three times while writing it. **Docs
only.**
## Why this belongs in that document rather than a new one
Findings 1-5 are about proving **your own** claim: does this path run,
can this test fail, is this negative result observable. Finding 6 is the
mirror image — the claims we make against **other people's** work — and
it is the same underlying error pointed outward. Splitting them would
let a reader take the first five as "be rigorous about my code" and miss
that the identical discipline applies when reviewing someone else's.
## The three cases, all mine, all in one day
| What I claimed | What was actually true |
|---|---|
| The census undercounts triage guards, 13 vs 10 | `summarize()` counts
`byColumnId` only for `kind === "column"`. My patched counter summed
`role`, `status` and `deliberate` too. The three "missing" ones were
exactly the ones it classifies correctly — and I reported this against
the instrument the program had just adopted as authoritative. |
| `resolvePlannerLanesForTask` silently disables two recovery paths for
legacy cards — escalated across four messages | The file's own header
had already reasoned it through and documented why that answer is
correct. And `TaskStore` implements `getTaskWorkflowSelectionAsync`,
which the resolver prefers — so real projects never take the path my `{
getTask }`-only probe forced. |
| `executor.ts` is clean of triage guards | A receiver-specific grep
missed three under `from` and `originColumn`. Same error one step
earlier: trusting a reconstruction of the thing instead of the thing. |
Every one was: reconstruct behaviour from outside → compare to actual
output → find a difference → report a defect, **without reading the
implementation.**
## The rules it adds
- Read the implementation and its header comment before reporting
anything as wrong. On this codebase the reasoning is usually already
written down, and the FNXC note frequently answers the exact objection —
twice today it answered mine verbatim.
- **A fixture is not a measurement of production.** When a probe and the
real system disagree, suspect the probe: ask what it had to stub, and
whether production ever supplies that shape.
- Retract precisely and immediately. A false defect report against
shared infrastructure costs more than the bug would have — it sends
people to verify something already correct, and spends the credibility
needed for the next report that is real.
Also updates the count in the intro (five → six) and adds an
`applies_when` entry so the doc surfaces for "about to report a tool as
defective", which is when it is needed and not when someone is already
debugging.
`pnpm lint` clean. No changeset — internal documentation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Durable write-up of U8's verification findings. **Docs only — no code
change, no CI risk beyond lint.**
These currently exist only in PR bodies, which nobody greps.
`docs/solutions/` is where this project keeps exactly this kind of
thing, and every one of the five will recur: the handler-pair shape and
the resolved-vs-guessed fork both have more call sites than U8 touched.
## The five
1. **Two prompt-node handlers exist; only one runs.**
`createDefaultNodeHandlers` prefers the primitives handler whenever
`deps.primitives` is set, and `executeWorkflowGraph` always sets it — so
every seam entry in `createAuthoritativeWorkflowSeams` is unreachable
for prompt nodes. A lifecycle announcement sat there through two PRs. It
type-checked and its unit tests passed, because a seam-level test calls
the seam object directly and therefore always can.
2. **A negative instrumentation result is worthless without a control.**
No output from an instrumented seam is only evidence once you have shown
writes from that module are visible under the harness. One
`process.stderr.write` at module load separates "never ran" from "output
swallowed" — opposite conclusions.
3. **Source-string ratchets prove syntax, not behavior.** Three were
torn down in review. The sharpest guarded a never-executed-code bug with
a source search, reproducing the bug one level up; measured, the
behavioural version fails an inverted dispatch and the textual one
passes it. Includes the sub-rules paid for the hard way: use the AST not
regex (a brace in a string truncated an extraction to 13 lines and every
count read a *passing* zero), guard the guard, anchor by index rather
than a character window.
4. **A green test on first try, on a path with no prior coverage, is a
warning.** Two conversions were reverted in one day because their tests
passed with the change reverted. Negative assertions succeed trivially
when the method returns early — `recoverCompletedTask` has seven guards
before the converted line, and the fixture has to satisfy all of them.
5. **A named workflow selection is not a resolved one.** Provenance
cannot be inferred from the returned value, because a fallback IR and a
valid id-less IR are structurally identical — the resolver that knows
has to report it. This is the fork every remaining lifecycle-column
conversion hits.
## Why this rather than another conversion
Everything left in my area is now owned and further along than I could
take it: `executor.ts` → #2628 (which solved the `recoverCompletedTask`
fixture I could not), `self-healing.ts` → #2560 (independently hit all
three traps I catalogued), the dashboard cluster → #2625/#2626/#2636.
Duplicating that would be motion, not progress. Turning findings that
cost real cycles into something greppable is the useful thing I can
still add.
`pnpm lint` clean. No changeset — internal documentation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added a best-practices guide for verifying that workflow code paths
actually execute.
* Covers reliable behavioral assertions, instrumentation controls,
regression-proof tests, source validation, and detection of fallback
behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Document the refactor-vs-semantics merge resolution pattern (port the
upstream semantic change into the extracted helper, never pick a side)
and the parallel-bootstrap CONCEPTS.md add/add union, with the FN-5902 /
runFeatureValidation merge as the worked example.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>