Flaky Tests: Why They're Poison and How to Catch Them
Every team has a few “coin-flip” cases: no code change, red in the morning, green in the afternoon, green again on a re-run. Most people treat them as a minor nuisance — annoying, sure, but just re-run them. This article argues the opposite: the flaky test is the single most destructive thing in a test system — and how to fish it out of your suite.
The one-sentence version:
What a flake burns isn’t machine time — it’s the credibility of a red light. And the entire value of a test system rests on red lights being credible.
1. Why It’s Poison
1.1 It creates a “boy who cried wolf” culture
A test system’s value = the team’s first reaction to a red light being “go check the product.” Flakes corrode that reaction into “just re-run it.” Once re-running becomes the culture, a real bug’s red light gets re-run into green too — the most dangerous thing about flakes isn’t the flakes themselves; it’s the invisibility cloak they hand to real regressions.
1.2 It’s worse than having no test at all
With no test, the team knows it’s flying blind and stays alert. A flaky case lets the team believe it has a test when it doesn’t — and burns daily attention on “is this red real or fake?”
1.3 The cost structure degrades across the board
- Direct cost: re-run machine time (amplified 10–100× on HIL), CI queueing;
- Indirect cost: triage cost dwarfs execution cost — every failure needs a human to answer “is the product wrong, or is the test throwing a tantrum?”;
- Systemic cost: the gate gets jammed by flakes repeatedly → engineers start routing around the gate → once the gate is bypassed, the whole system collapses.
1.4 It’s contagious
One flake left alone tells the whole team “failures can be re-run away” → the bar for writing cases drops → more flakes pour in. Left alone, the flake rate only goes one way: up — you have to actively govern it.
1.5 Automotive’s two amplifiers
- HIL rig time is expensive: one re-run burns the whole team’s queueing resource;
- The ISO 26262 evidence chain: at audit time, “this case passed 2 out of 3 runs” is a work product you’ll have to explain — flakes pollute the traceability chain (execution records are compliance evidence, not just engineering data).
2. Root-Cause Taxonomy (the Prerequisite for Detection)
A flake isn’t a random ghost — every flake has a root cause. Know the cause type, and you know which fishing method to use:
| Root cause | Typical symptom | Fishing method |
|---|---|---|
| Timing assumptions | sleep(100) waits for a result and fails on a slow machine; assertions assume zero latency |
Run under perturbation (CPU saturated / bus loaded) |
| Environment leakage | Shared state between cases not cleaned up; order-dependent | Shuffle runs (--gtest_shuffle) |
| Uncontrolled randomness | No fixed seed; depends on clock/date/network | Re-run the same commit repeatedly |
| Resource boundaries | Timeout threshold set right at the normal duration; a leak makes only the Nth case fail | Long-haul repetition + resource monitoring |
| A real product bug | Concurrency races, state-machine corner cases | ⚠️ See the warning below |
Important warning: a flake is sometimes the early signal of a real product bug — concurrency races and state-machine corner cases show up in tests precisely as “sometimes passes, sometimes fails.” Blaming the test environment by default throws away the product’s early alarm as noise. The governance process must include a step that rules out the product as a suspect.
3. Five Detection Methods (Ranked by Power)
- The statistical method (gold standard): re-run the same commit N times (10 times overnight); whatever disagrees with itself is a flake. With the code variable eliminated, any remaining inconsistency can only come from the test or the environment;
- Historical-data analysis: look at each case’s red-green flip count (flips) — a real regression’s failure curve is a step (100% red from some commit, green once fixed); a flake’s failure curve is scatter (red and green alternate at random, uncorrelated with commits). Sort by flips and take the top N — that’s your flake candidate list;
- Pass-on-re-run = a strong signal: in CI, anything that fails and then passes on an in-place re-run gets auto-tagged (
retried-pass). Many CI platforms natively support retry-with-recording — let this data flow away uncollected and it’s gone for good; - Admission testing for new cases: before a new case enters the library, run it 50–100 times (varying loads, varying seeds) — flake detection moves upstream into case review, which is far cheaper than fishing it out after admission;
- Perturbation / chaos: run with the CPU saturated, network delay injected, the HIL bus loaded to the max — timing-assumption flakes surface in bulk under perturbation.
4. Three Governance Moves
- Quarantine: on discovery, move the flake out of the gate, tag it
quarantine, fix it by a deadline, delete it if it can’t be fixed. A flake left in the gate burns trust every day — a bigger loss than deleting it; - Aim the fix at the root cause: replace
sleepwith condition waits / event-driven checks; fix the random seed; clean shared state in teardown; set timeout thresholds at p99×3, not hugging the mean; - Report the flake rate: track it as a test-system health KPI — the trend may only go down, never up.
Back to the opening line: what a flake burns is the credibility of a red light. The goal of flake governance isn’t “a greener suite” — it’s making red mean “check the product” again. The day that happens is the day the test system is actually working.
Appendix: Glossary (in order of appearance)
| Term | Plain explanation |
|---|---|
| flaky test / flake | A case that passes and fails without any code change — a coin flip |
| re-run | Running the same failure again in place, no code or environment change |
| HIL | Hardware-in-the-Loop testing: a real controller hooked to a simulation rig; rig time is expensive |
| CI | Continuous Integration: the pipeline that builds and tests every commit automatically |
| quality gate | An automated checkpoint in the CI pipeline: fail it and you can’t merge |
| ISO 26262 | The automotive functional-safety standard; execution records are compliance evidence examined at audit |
| evidence chain / traceability chain | The whole set of records that lets a third party recompute test conclusions and trace them to requirements |
| flips | How many times a case’s result flips between adjacent runs; the ranking metric for flake detection |
| retried-pass | The tag for a case that fails and then passes on re-run — a strong flake signal |
| quarantine | The tag mechanism that moves flakes out of the gate into a supervised area with a fix deadline |
| p99 | The 99th-percentile duration: the 99th value when 100 runs are sorted fastest to slowest — the baseline for setting timeout thresholds |
| KPI | Key Performance Indicator; the flake rate is one of a test system’s health KPIs |
Next step:
View all notes