diff --git a/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md b/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md index cbeef3b5f2..1d12057f8c 100644 --- a/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md +++ b/docs/solutions/workflow-learnings/lifecycle-conversions-that-score-as-wins.md @@ -199,6 +199,49 @@ wired upstream. That is the point at which a ratchet has earned its keep. Until a guard has failed on code you did not write, you know it encodes your habits; you do not yet know it encodes the invariant. +## Mutation testing has one blind spot: your own imagination + +Everything above says "watch the guard go red before you trust it." That rule is necessary and it is +not sufficient, and the way it fails is worth stating because it cost the most. + +Three instruments were written during this program, each mutation-tested in both directions before +shipping, each green. Reviewers then found, in those same instruments: + +- a file-level pre-filter that skipped whole files, so a forbidden site added to a file with no other + SQL was invisible; +- an anchored pattern that missed qualified and compound fragments (`t."column" = 'done'`); +- a scan over SOURCE text, where a double-quoted TS string still spells `\"column\"` with the + backslashes in it; +- an operator list of `= != <>` that never considered `IN (...)`; +- and worst, a template scan that joined only the STATIC spans — so a Drizzle query, which puts the + COLUMN in the interpolation hole and the legacy id in the static text, matched nothing. **That gate + was blind on the exact files it was built to freeze.** Enabling that one shape revealed five + previously invisible files and took the population from 14 to 31. + +Every one of those is a FALSE NEGATIVE, and none was catchable by the mutation tests that were run, +for a structural reason: **a mutation you write is a mutation you already imagined, so it lands +inside the space your scanner understands.** Reintroducing a defect the checker was designed around +proves the checker still handles that defect. It says nothing about the shapes you never modelled. + +What actually finds these: + +1. **Run the instrument against the code it was written for, and read the hits by hand.** The Drizzle + blindness was visible the moment someone asked "why is the merge-queue query, the reason this + exists, not in the output?" +2. **Separate the two causes, because they have different fixes.** Only the first defect above is a + pre-filter problem: two patterns disagreeing about whether to run the real check, where the cheap + one is a second, weaker specification of the thing you are testing. Delete it and run the real + pattern everywhere. The other four — the anchored fragment, the raw-source scan, the missing `IN`, + the static-span join — are single-pattern defects: the pattern was the only specification and it + simply did not describe the shape. No amount of restructuring finds those; only running the + instrument against real code does (see 1). Conflating them is tempting because they all present as + "the guard is green and the site is there", and it sends you refactoring when you should be + reading output. +3. **Treat a guard's own count as a claim to verify, not a result.** "14 sites" read as coverage for + days; it was the subset one scanner happened to model. + +A false positive is loud and gets fixed. A false negative prints a baseline and reads as coverage. + ## The rule that produced every fix above **A green guard is evidence only once you have watched it go red.**