docs(solutions): record "silence is not success" (#3241)
## What A new `docs/solutions` note recording a failure that hit **three different tools in one session**, each time reading as a pass. Docs only. ## The three costumes | what happened | looked like | was | |---|---|---| | `git stash --keep-index` swept the new test file out of the tree | "45 passed" | the pre-existing count; the new test never ran | | a blinding script hit an unmapped role and `sys.exit(2)` **with no message**; `&&` skipped the check, `;` let the run proceed | "375/375 green under blinding" | nothing blinded — run was against unmodified source | | a gate piped to `tail -1`, printing a blank line | "gate ran, no complaints" | exit code 1; the FNXC stamp check had failed, and **CI caught it in #3238** | ## Why it deserves its own note **A passing run and a run that never happened produce the same evidence: no failure text.** Every other bug announces itself; this one is defined by the absence of an announcement. The instinct that catches ordinary bugs — *"nothing looks wrong"* — is precisely the instinct that certifies this one. It gets worse under automation, where output is piped and skimmed. `| tail -1`, `| grep "Tests"`, `>/dev/null 2>&1` all discard the part that would have said `No test files found` or `command not found`. ## The five rules, each paid for above 1. **Assert the exit code before any pipe.** A pipeline's status is the *last* stage's — `cmd | tail -1` reports `tail`'s success, never `cmd`'s. 2. **Confirm the run did the work.** "Test Files 1 passed" when you expected 16 is a finding, not a pass. 3. **A tool that can no-op must say what it did** — print the substitution and location, fail loudly where it cannot act. 4. **Verify the mutation, not the tool's promise** — `git diff --stat`, not the exit code. 5. **Break the guard on purpose once** and watch it fail. A guard never observed failing has not been shown to work — the standard this repo already applies to product ratchets, turned on your own verification. ## The uncomfortable part, kept in The third instance was a rule **I added to AGENTS.md myself in #3174**, broken for the second time. I ran the gate. I read `tail -1`. I moved on. Writing a rule down does not make you follow it. The only reason it was caught is that **CI read the output when I did not** — an argument for the gate existing, not for me having been careful. Cross-linked from the resolver-audit note, whose every wrong reading came from a run that never happened rather than from the blinding itself. That connection is the point: I spent this session auditing a program whose subject is defects hiding behind green results, and reproduced the same class three times in my own tooling. ``` lint clean; fnxc-future-dates: none added (exit code checked before piping this time) ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added workflow guidance explaining why silent or seemingly successful output does not confirm that a test, script, or validation gate ran. * Documented verification practices including checking exit codes, work counts, no-op detection, post-run changes, and intentional failure checks. * Added a case study highlighting how filtered output can conceal verification failures. * Added cross-references connecting resolver interpretation, test execution, and conversion coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -8,6 +8,13 @@ applies_when: deciding whether a lane conversion is actually protected by a test
|
||||
|
||||
# Blind the resolver: the only way to know a conversion is covered
|
||||
|
||||
See also [silence-is-not-success](./silence-is-not-success.md) — every wrong reading this method
|
||||
produced came from the RUN, never from the blind itself. Two kinds: a run that never happened (the
|
||||
swept test file, the silent `exit 2`), and a run that happened but could not measure anything (a
|
||||
harness that made the blind unobservable, a suite that never reached the blinded site — both below).
|
||||
Absent output and invalid output read identically at a glance, which is why the rules there are about
|
||||
confirming what a run actually did rather than about whether it printed a failure.
|
||||
|
||||
Sibling to [a-falling-count-is-not-evidence](./a-falling-count-is-not-evidence.md), which records that
|
||||
a metric moving is not proof the system moved. This one records the **positive** procedure: how to
|
||||
find out whether a landed conversion is held by anything, and how to write a test that holds it.
|
||||
|
||||
72
docs/solutions/workflow-learnings/silence-is-not-success.md
Normal file
72
docs/solutions/workflow-learnings/silence-is-not-success.md
Normal file
@@ -0,0 +1,72 @@
|
||||
---
|
||||
category: workflow-learnings
|
||||
module: scripts
|
||||
tags: [verification, gates, measurement, false-green, tooling]
|
||||
problem_type: process-learning
|
||||
applies_when: reading the output of a test run, gate, or script as evidence it passed
|
||||
---
|
||||
|
||||
# Silence is not success
|
||||
|
||||
Sibling to [a-falling-count-is-not-evidence](./a-falling-count-is-not-evidence.md) (a metric moving
|
||||
is not proof the system moved) and
|
||||
[blind-the-resolver-to-find-uncovered-conversions](./blind-the-resolver-to-find-uncovered-conversions.md)
|
||||
(a green suite is not proof a conversion is held). This one is narrower and more embarrassing: **the
|
||||
absence of a failure message is not proof anything ran.**
|
||||
|
||||
Written after hitting the same root cause **three times in one session**, in three different tools,
|
||||
while auditing a program whose entire subject is defects that hide behind green results.
|
||||
|
||||
## The three costumes
|
||||
|
||||
| what happened | what it looked like | what it was |
|
||||
|---|---|---|
|
||||
| `git stash --keep-index` swept the new **untracked** test file out of the tree (it stashes tracked, unstaged changes and, with `-u`, untracked ones — an unstaged new file is not protected by `--keep-index`) | "45 passed" | the pre-existing count; the new test never ran |
|
||||
| a blinding script hit an unmapped role and `sys.exit(2)` with **no message**; `&&` skipped the check, `;` let the run proceed | "375/375 green under blinding" | nothing was blinded — the run was against unmodified source |
|
||||
| a gate piped to `tail -1`, which printed a blank line | "gate ran, no complaints" | exit code 1; the FNXC stamp check had failed and CI caught it later |
|
||||
|
||||
Each read as success. None was.
|
||||
|
||||
## Why this class is hard to see
|
||||
|
||||
A passing run and a run that never happened produce **the same evidence**: no failure text. Every
|
||||
other bug announces itself; this one is defined by the absence of an announcement. The instinct that
|
||||
catches ordinary bugs — "nothing looks wrong" — is precisely the instinct that certifies this one.
|
||||
|
||||
It is worse under automation, where output is piped, filtered, and skimmed. `| tail -1`, `| grep
|
||||
"Tests"`, `>/dev/null` and `2>&1` all discard the part that would have said "0 files matched" or
|
||||
"command not found".
|
||||
|
||||
## Rules
|
||||
|
||||
**1. Assert the exit code, not the absence of text.** `rc=$?` immediately after the command, before
|
||||
any pipe. A pipeline's exit status is the *last* stage's by default, so `cmd | tail -1` reports
|
||||
`tail`'s success, never `cmd`'s. (`set -o pipefail` changes that, and `PIPESTATUS`/`pipestatus`
|
||||
exposes each stage — but neither is on by default in the shells these commands are pasted into, so
|
||||
the rule stands unless you have explicitly enabled it.)
|
||||
|
||||
**2. Confirm the run did the work.** A test run must report a plausible file/test COUNT — "Test Files
|
||||
1 passed" when you expected 16 is a finding, not a pass. The exact strings quoted here are **Vitest's**
|
||||
(`Test Files N passed`, `No test files found`, which exits non-zero and is easy to mistake for a
|
||||
failing test); other runners word it differently, so match on the count you expected rather than on
|
||||
the phrasing.
|
||||
|
||||
**3. A tool that can no-op must say what it did.** Print the substitution and its location, and fail
|
||||
loudly on the paths where it cannot act. A silent `exit 2` is indistinguishable from success once
|
||||
piped.
|
||||
|
||||
**4. Verify the mutation, not the tool's promise.** `git diff --stat` after an edit script; `grep`
|
||||
for the changed line. The tool's exit code describes the tool, not the file.
|
||||
|
||||
**5. Make the check obviously fatal once.** If a guard is supposed to fail, break it on purpose and
|
||||
watch it fail. A guard never observed failing has not been shown to work — the same standard this
|
||||
repo applies to product ratchets, turned on your own verification.
|
||||
|
||||
## The uncomfortable part
|
||||
|
||||
The FNXC-stamp breach above was against a rule **I had added to AGENTS.md myself**, and it was the
|
||||
second time I broke it. I ran the gate. I read `tail -1`. I moved on.
|
||||
|
||||
Writing a rule down does not make you follow it. A gate whose output you do not read is not a gate,
|
||||
and the only reason this was caught is that CI read the output when the author did not. That is an
|
||||
argument for the gate existing — not for the author having been careful.
|
||||
Reference in New Issue
Block a user