Flaky Tests: Why They're Poison and How to Catch Them

Every team has a few “coin-flip” cases: no code change, red in the morning, green in the afternoon, green again on a re-run. Most people treat them as a minor nuisance — annoying, sure, but just re-run them. This article argues the opposite: the flaky test is the single most destructive thing in a test system — and how to fish it out of your suite.

The one-sentence version:

What a flake burns isn’t machine time — it’s the credibility of a red light. And the entire value of a test system rests on red lights being credible.

1. Why It’s Poison

1.1 It creates a “boy who cried wolf” culture

A test system’s value = the team’s first reaction to a red light being “go check the product.” Flakes corrode that reaction into “just re-run it.” Once re-running becomes the culture, a real bug’s red light gets re-run into green too — the most dangerous thing about flakes isn’t the flakes themselves; it’s the invisibility cloak they hand to real regressions.

1.2 It’s worse than having no test at all

With no test, the team knows it’s flying blind and stays alert. A flaky case lets the team believe it has a test when it doesn’t — and burns daily attention on “is this red real or fake?”

1.3 The cost structure degrades across the board

1.4 It’s contagious

One flake left alone tells the whole team “failures can be re-run away” → the bar for writing cases drops → more flakes pour in. Left alone, the flake rate only goes one way: up — you have to actively govern it.

1.5 Automotive’s two amplifiers

2. Root-Cause Taxonomy (the Prerequisite for Detection)

A flake isn’t a random ghost — every flake has a root cause. Know the cause type, and you know which fishing method to use:

Root cause Typical symptom Fishing method
Timing assumptions sleep(100) waits for a result and fails on a slow machine; assertions assume zero latency Run under perturbation (CPU saturated / bus loaded)
Environment leakage Shared state between cases not cleaned up; order-dependent Shuffle runs (--gtest_shuffle)
Uncontrolled randomness No fixed seed; depends on clock/date/network Re-run the same commit repeatedly
Resource boundaries Timeout threshold set right at the normal duration; a leak makes only the Nth case fail Long-haul repetition + resource monitoring
A real product bug Concurrency races, state-machine corner cases ⚠️ See the warning below

Important warning: a flake is sometimes the early signal of a real product bug — concurrency races and state-machine corner cases show up in tests precisely as “sometimes passes, sometimes fails.” Blaming the test environment by default throws away the product’s early alarm as noise. The governance process must include a step that rules out the product as a suspect.

3. Five Detection Methods (Ranked by Power)

  1. The statistical method (gold standard): re-run the same commit N times (10 times overnight); whatever disagrees with itself is a flake. With the code variable eliminated, any remaining inconsistency can only come from the test or the environment;
  2. Historical-data analysis: look at each case’s red-green flip count (flips) — a real regression’s failure curve is a step (100% red from some commit, green once fixed); a flake’s failure curve is scatter (red and green alternate at random, uncorrelated with commits). Sort by flips and take the top N — that’s your flake candidate list;
  3. Pass-on-re-run = a strong signal: in CI, anything that fails and then passes on an in-place re-run gets auto-tagged (retried-pass). Many CI platforms natively support retry-with-recording — let this data flow away uncollected and it’s gone for good;
  4. Admission testing for new cases: before a new case enters the library, run it 50–100 times (varying loads, varying seeds) — flake detection moves upstream into case review, which is far cheaper than fishing it out after admission;
  5. Perturbation / chaos: run with the CPU saturated, network delay injected, the HIL bus loaded to the max — timing-assumption flakes surface in bulk under perturbation.

4. Three Governance Moves

  1. Quarantine: on discovery, move the flake out of the gate, tag it quarantine, fix it by a deadline, delete it if it can’t be fixed. A flake left in the gate burns trust every day — a bigger loss than deleting it;
  2. Aim the fix at the root cause: replace sleep with condition waits / event-driven checks; fix the random seed; clean shared state in teardown; set timeout thresholds at p99×3, not hugging the mean;
  3. Report the flake rate: track it as a test-system health KPI — the trend may only go down, never up.

Back to the opening line: what a flake burns is the credibility of a red light. The goal of flake governance isn’t “a greener suite” — it’s making red mean “check the product” again. The day that happens is the day the test system is actually working.


Appendix: Glossary (in order of appearance)

Term Plain explanation
flaky test / flake A case that passes and fails without any code change — a coin flip
re-run Running the same failure again in place, no code or environment change
HIL Hardware-in-the-Loop testing: a real controller hooked to a simulation rig; rig time is expensive
CI Continuous Integration: the pipeline that builds and tests every commit automatically
quality gate An automated checkpoint in the CI pipeline: fail it and you can’t merge
ISO 26262 The automotive functional-safety standard; execution records are compliance evidence examined at audit
evidence chain / traceability chain The whole set of records that lets a third party recompute test conclusions and trace them to requirements
flips How many times a case’s result flips between adjacent runs; the ranking metric for flake detection
retried-pass The tag for a case that fails and then passes on re-run — a strong flake signal
quarantine The tag mechanism that moves flakes out of the gate into a supervised area with a fix deadline
p99 The 99th-percentile duration: the 99th value when 100 runs are sorted fastest to slowest — the baseline for setting timeout thresholds
KPI Key Performance Indicator; the flake rate is one of a test system’s health KPIs

Next step:

View all notes