From 851369a4806e2d126ea69f7b70beea94eceb3458 Mon Sep 17 00:00:00 2001 From: gsxdsm Date: Fri, 31 Jul 2026 13:47:08 -0700 Subject: [PATCH] docs(solutions): record "silence is not success" (#3241) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## What A new `docs/solutions` note recording a failure that hit **three different tools in one session**, each time reading as a pass. Docs only. ## The three costumes | what happened | looked like | was | |---|---|---| | `git stash --keep-index` swept the new test file out of the tree | "45 passed" | the pre-existing count; the new test never ran | | a blinding script hit an unmapped role and `sys.exit(2)` **with no message**; `&&` skipped the check, `;` let the run proceed | "375/375 green under blinding" | nothing blinded — run was against unmodified source | | a gate piped to `tail -1`, printing a blank line | "gate ran, no complaints" | exit code 1; the FNXC stamp check had failed, and **CI caught it in #3238** | ## Why it deserves its own note **A passing run and a run that never happened produce the same evidence: no failure text.** Every other bug announces itself; this one is defined by the absence of an announcement. The instinct that catches ordinary bugs — *"nothing looks wrong"* — is precisely the instinct that certifies this one. It gets worse under automation, where output is piped and skimmed. `| tail -1`, `| grep "Tests"`, `>/dev/null 2>&1` all discard the part that would have said `No test files found` or `command not found`. ## The five rules, each paid for above 1. **Assert the exit code before any pipe.** A pipeline's status is the *last* stage's — `cmd | tail -1` reports `tail`'s success, never `cmd`'s. 2. **Confirm the run did the work.** "Test Files 1 passed" when you expected 16 is a finding, not a pass. 3. **A tool that can no-op must say what it did** — print the substitution and location, fail loudly where it cannot act. 4. **Verify the mutation, not the tool's promise** — `git diff --stat`, not the exit code. 5. **Break the guard on purpose once** and watch it fail. A guard never observed failing has not been shown to work — the standard this repo already applies to product ratchets, turned on your own verification. ## The uncomfortable part, kept in The third instance was a rule **I added to AGENTS.md myself in #3174**, broken for the second time. I ran the gate. I read `tail -1`. I moved on. Writing a rule down does not make you follow it. The only reason it was caught is that **CI read the output when I did not** — an argument for the gate existing, not for me having been careful. Cross-linked from the resolver-audit note, whose every wrong reading came from a run that never happened rather than from the blinding itself. That connection is the point: I spent this session auditing a program whose subject is defects hiding behind green results, and reproduced the same class three times in my own tooling. ``` lint clean; fnxc-future-dates: none added (exit code checked before piping this time) ``` ## Summary by CodeRabbit * **Documentation** * Added workflow guidance explaining why silent or seemingly successful output does not confirm that a test, script, or validation gate ran. * Documented verification practices including checking exit codes, work counts, no-op detection, post-run changes, and intentional failure checks. * Added a case study highlighting how filtered output can conceal verification failures. * Added cross-references connecting resolver interpretation, test execution, and conversion coverage. --------- Co-authored-by: Claude Opus 5 (1M context) --- ...-resolver-to-find-uncovered-conversions.md | 7 ++ .../silence-is-not-success.md | 72 +++++++++++++++++++ 2 files changed, 79 insertions(+) create mode 100644 docs/solutions/workflow-learnings/silence-is-not-success.md diff --git a/docs/solutions/workflow-learnings/blind-the-resolver-to-find-uncovered-conversions.md b/docs/solutions/workflow-learnings/blind-the-resolver-to-find-uncovered-conversions.md index b3f172840e..53a82d98c6 100644 --- a/docs/solutions/workflow-learnings/blind-the-resolver-to-find-uncovered-conversions.md +++ b/docs/solutions/workflow-learnings/blind-the-resolver-to-find-uncovered-conversions.md @@ -8,6 +8,13 @@ applies_when: deciding whether a lane conversion is actually protected by a test # Blind the resolver: the only way to know a conversion is covered +See also [silence-is-not-success](./silence-is-not-success.md) — every wrong reading this method +produced came from the RUN, never from the blind itself. Two kinds: a run that never happened (the +swept test file, the silent `exit 2`), and a run that happened but could not measure anything (a +harness that made the blind unobservable, a suite that never reached the blinded site — both below). +Absent output and invalid output read identically at a glance, which is why the rules there are about +confirming what a run actually did rather than about whether it printed a failure. + Sibling to [a-falling-count-is-not-evidence](./a-falling-count-is-not-evidence.md), which records that a metric moving is not proof the system moved. This one records the **positive** procedure: how to find out whether a landed conversion is held by anything, and how to write a test that holds it. diff --git a/docs/solutions/workflow-learnings/silence-is-not-success.md b/docs/solutions/workflow-learnings/silence-is-not-success.md new file mode 100644 index 0000000000..9dd2439299 --- /dev/null +++ b/docs/solutions/workflow-learnings/silence-is-not-success.md @@ -0,0 +1,72 @@ +--- +category: workflow-learnings +module: scripts +tags: [verification, gates, measurement, false-green, tooling] +problem_type: process-learning +applies_when: reading the output of a test run, gate, or script as evidence it passed +--- + +# Silence is not success + +Sibling to [a-falling-count-is-not-evidence](./a-falling-count-is-not-evidence.md) (a metric moving +is not proof the system moved) and +[blind-the-resolver-to-find-uncovered-conversions](./blind-the-resolver-to-find-uncovered-conversions.md) +(a green suite is not proof a conversion is held). This one is narrower and more embarrassing: **the +absence of a failure message is not proof anything ran.** + +Written after hitting the same root cause **three times in one session**, in three different tools, +while auditing a program whose entire subject is defects that hide behind green results. + +## The three costumes + +| what happened | what it looked like | what it was | +|---|---|---| +| `git stash --keep-index` swept the new **untracked** test file out of the tree (it stashes tracked, unstaged changes and, with `-u`, untracked ones — an unstaged new file is not protected by `--keep-index`) | "45 passed" | the pre-existing count; the new test never ran | +| a blinding script hit an unmapped role and `sys.exit(2)` with **no message**; `&&` skipped the check, `;` let the run proceed | "375/375 green under blinding" | nothing was blinded — the run was against unmodified source | +| a gate piped to `tail -1`, which printed a blank line | "gate ran, no complaints" | exit code 1; the FNXC stamp check had failed and CI caught it later | + +Each read as success. None was. + +## Why this class is hard to see + +A passing run and a run that never happened produce **the same evidence**: no failure text. Every +other bug announces itself; this one is defined by the absence of an announcement. The instinct that +catches ordinary bugs — "nothing looks wrong" — is precisely the instinct that certifies this one. + +It is worse under automation, where output is piped, filtered, and skimmed. `| tail -1`, `| grep +"Tests"`, `>/dev/null` and `2>&1` all discard the part that would have said "0 files matched" or +"command not found". + +## Rules + +**1. Assert the exit code, not the absence of text.** `rc=$?` immediately after the command, before +any pipe. A pipeline's exit status is the *last* stage's by default, so `cmd | tail -1` reports +`tail`'s success, never `cmd`'s. (`set -o pipefail` changes that, and `PIPESTATUS`/`pipestatus` +exposes each stage — but neither is on by default in the shells these commands are pasted into, so +the rule stands unless you have explicitly enabled it.) + +**2. Confirm the run did the work.** A test run must report a plausible file/test COUNT — "Test Files +1 passed" when you expected 16 is a finding, not a pass. The exact strings quoted here are **Vitest's** +(`Test Files N passed`, `No test files found`, which exits non-zero and is easy to mistake for a +failing test); other runners word it differently, so match on the count you expected rather than on +the phrasing. + +**3. A tool that can no-op must say what it did.** Print the substitution and its location, and fail +loudly on the paths where it cannot act. A silent `exit 2` is indistinguishable from success once +piped. + +**4. Verify the mutation, not the tool's promise.** `git diff --stat` after an edit script; `grep` +for the changed line. The tool's exit code describes the tool, not the file. + +**5. Make the check obviously fatal once.** If a guard is supposed to fail, break it on purpose and +watch it fail. A guard never observed failing has not been shown to work — the same standard this +repo applies to product ratchets, turned on your own verification. + +## The uncomfortable part + +The FNXC-stamp breach above was against a rule **I had added to AGENTS.md myself**, and it was the +second time I broke it. I ran the gate. I read `tail -1`. I moved on. + +Writing a rule down does not make you follow it. A gate whose output you do not read is not a gate, +and the only reason this was caught is that CI read the output when the author did not. That is an +argument for the gate existing — not for the author having been careful.