From 5792452f0a385796fb06c6318965b22cd1dc7187 Mon Sep 17 00:00:00 2001 From: gsxdsm Date: Thu, 30 Jul 2026 15:55:00 -0700 Subject: [PATCH] =?UTF-8?q?docs(workflow-learnings):=20mutation=20testing?= =?UTF-8?q?=20has=20one=20blind=20spot=20=E2=80=94=20your=20own=20imaginat?= =?UTF-8?q?ion=20(#2858)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The most transferable thing this lane produced, and it is a correction to advice I wrote earlier in the same document. ## The gap Every other section here says *"watch the guard go red before you trust it."* That rule is necessary and **not sufficient**, and the way it fails cost the most. Three instruments were written during this program. Each was mutation-tested in both directions before shipping. Each was green. Reviewers then found, in those same instruments: - a **file-level pre-filter** that skipped whole files, so a forbidden site added to a file with no other SQL was invisible; - an **anchored pattern** that missed qualified and compound fragments (`t."column" = 'done'`); - a scan over **SOURCE text**, where a double-quoted TS string still spells `\"column\"` with the backslashes in it; - an operator list of `= != <>` that never considered **`IN (...)`**; - and worst, a template scan that joined only the **static spans** — so a Drizzle query, which puts the COLUMN in the interpolation hole and the legacy id in the static text, matched nothing. **That gate was blind on the exact files it was built to freeze.** Enabling that one shape took the population from 14 to 31 and revealed five previously invisible files. ## Why the mutation tests could not catch any of them All five are **false negatives**, and the reason is structural rather than sloppy: > A mutation you write is a mutation you already imagined, so it lands inside the space your scanner understands. Reintroducing a defect the checker was designed around proves the checker still handles that defect. It says nothing about shapes you never modelled. ## What does find them 1. **Run the instrument against the code it was written for and read the hits by hand.** The Drizzle blindness was obvious the moment someone asked *"why is the merge-queue query — the reason this exists — not in the output?"* 2. **Prefer one unanchored pattern over a fast pre-filter plus a precise one.** Every false negative above came from two patterns disagreeing about whether to run the real check at all. A pre-filter is a second, weaker specification of the thing you are testing. 3. **Treat a guard's own count as a claim to verify, not a result.** "14 sites" read as coverage for days; it was the subset one scanner happened to model. A false positive is loud and gets fixed. **A false negative prints a baseline and reads as coverage.** Docs only — no changeset. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) --- ...ifecycle-conversions-that-score-as-wins.md | 43 +++++++++++++++++++ 1 file changed, 43 insertions(+) diff --git a/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md b/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md index cbeef3b5f2..1d12057f8c 100644 --- a/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md +++ b/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md @@ -199,6 +199,49 @@ wired upstream. That is the point at which a ratchet has earned its keep. Until a guard has failed on code you did not write, you know it encodes your habits; you do not yet know it encodes the invariant. +## Mutation testing has one blind spot: your own imagination + +Everything above says "watch the guard go red before you trust it." That rule is necessary and it is +not sufficient, and the way it fails is worth stating because it cost the most. + +Three instruments were written during this program, each mutation-tested in both directions before +shipping, each green. Reviewers then found, in those same instruments: + +- a file-level pre-filter that skipped whole files, so a forbidden site added to a file with no other + SQL was invisible; +- an anchored pattern that missed qualified and compound fragments (`t."column" = 'done'`); +- a scan over SOURCE text, where a double-quoted TS string still spells `\"column\"` with the + backslashes in it; +- an operator list of `= != <>` that never considered `IN (...)`; +- and worst, a template scan that joined only the STATIC spans — so a Drizzle query, which puts the + COLUMN in the interpolation hole and the legacy id in the static text, matched nothing. **That gate + was blind on the exact files it was built to freeze.** Enabling that one shape revealed five + previously invisible files and took the population from 14 to 31. + +Every one of those is a FALSE NEGATIVE, and none was catchable by the mutation tests that were run, +for a structural reason: **a mutation you write is a mutation you already imagined, so it lands +inside the space your scanner understands.** Reintroducing a defect the checker was designed around +proves the checker still handles that defect. It says nothing about the shapes you never modelled. + +What actually finds these: + +1. **Run the instrument against the code it was written for, and read the hits by hand.** The Drizzle + blindness was visible the moment someone asked "why is the merge-queue query, the reason this + exists, not in the output?" +2. **Separate the two causes, because they have different fixes.** Only the first defect above is a + pre-filter problem: two patterns disagreeing about whether to run the real check, where the cheap + one is a second, weaker specification of the thing you are testing. Delete it and run the real + pattern everywhere. The other four — the anchored fragment, the raw-source scan, the missing `IN`, + the static-span join — are single-pattern defects: the pattern was the only specification and it + simply did not describe the shape. No amount of restructuring finds those; only running the + instrument against real code does (see 1). Conflating them is tempting because they all present as + "the guard is green and the site is there", and it sends you refactoring when you should be + reading output. +3. **Treat a guard's own count as a claim to verify, not a result.** "14 sites" read as coverage for + days; it was the subset one scanner happened to model. + +A false positive is loud and gets fixed. A false negative prints a baseline and reads as coverage. + ## The rule that produced every fix above **A green guard is evidence only once you have watched it go red.**