75% of Flaky Failures Cluster: Flaky Test Detection Without Reruns
75% of Flaky Failures Cluster: Flaky Test Detection Without Reruns

A flaky test is one that passes and fails on the same code without any real change in behavior, and the fastest way to detect it is to capture failure metadata on every red build and run a differential coverage check before you waste cycles on a rerun. That single habit turns a guessing game into a repeatable signal. Once a test is confirmed flaky, triage or quarantine it immediately so it stops eroding trust in your CI pipeline.
TL;DR:
- Most flaky tests are caused by external dependency instability, timing issues, shared state, or environment differences, often forming systemic clusters.
- Differential coverage can identify flaky tests efficiently by comparing failure code coverage with recent code changes, reducing reliance on costly reruns.
- Quarantining flaky tests should follow a strict process, with clear criteria such as three failures in 30 days and set deadlines for fixing or removal.
- Small suites may benefit from reruns, whereas large-scale projects should prioritize coverage-based detection and clustering to manage high flaky rates effectively.
- Combining pre-merge risk detection with CI flaky detection tools helps prevent flaky conditions from reaching production, saving time and improving test trustworthiness.
Table of Contents
- What a flaky test is and why it matters
- Common causes of flakiness: a practical taxonomy
- Detection methods: reruns, coverage diffs, and clustering
- A practical detection checklist you can run in CI
- Triage and remediation: quarantine, fix, or suppress
- Tooling and automation options for managing flaky tests
- Veridical perspective: catching flaky conditions before they merge
- Engineering recommendation: prioritize systemic fixes first
- Reduce flaky surface before it reaches CI
- FAQ
- Sources
What a flaky test is and why it matters
A flaky test fails intermittently against unchanged code, which separates it from a regression, where a failure traces to a real code change. The distinction matters because teams that treat flaky failures as regressions waste engineering hours chasing bugs that do not exist, while teams that ignore them entirely start merging around red builds, which defeats the purpose of having a test suite at all.
Flaky tests erode CI trust in a specific way: once a developer sees a test fail and pass again on retry without explanation, they stop reading failures as signal. Merges slow down because reviewers manually rerun suites instead of trusting the first result, and legitimate regressions get lost in the noise.
Three metrics are worth tracking on a rolling basis:
- Flaky rate: flaky test occurrences per 100 runs, tracked weekly.
- Triage time: average time from failure detection to root-cause classification.
- Time-to-merge impact: added minutes or hours a flaky failure adds to the average pull request lifecycle.
Common causes of flakiness: a practical taxonomy
Mapping a failure to its likely cause before you start debugging saves hours. Most flaky tests fall into one of four buckets.
- External dependency instability: calls to third-party APIs, DNS resolution delays, or network timeouts that succeed most of the time but fail under load.
- Timing and concurrency issues: race conditions between threads, or tests that rely on a fixed
sleep()instead of waiting for an actual state change. - Test order dependencies and shared state: a test that passes in isolation but fails when run after another test that left global state, a database row, or a cache entry behind.
- Resource leaks and platform differences: file handles, ports, or memory that are not released between runs, or behavior that differs between a developer’s machine and the CI runner’s operating system.
Pro Tip: Tag every flaky failure with its suspected cause category the first time you see it. Patterns across tags often reveal a shared root cause faster than debugging each test individually.
Detection methods: reruns, coverage diffs, and clustering
The oldest detection method is the rerun: run a failing test again, in isolation, with a fresh process or JVM fork, and if it passes, call it flaky. This works but it is expensive at scale, since exhaustive rerun strategies multiply CI minutes across every suite, and a test can fail flaky-style on one rerun and pass cleanly on the next, giving a false sense of certainty either way.
A cheaper and more precise approach is differential coverage, the technique behind DeFlaker. Instead of rerunning, DeFlaker checks whether the newly failing test actually executed any of the code changed in the current commit. If the coverage has not changed, the failure cannot be a regression, so it gets marked as suspected flaky without a single rerun. Deployed across 96 Java projects, this method surfaced many previously unknown flaky tests with a low false alarm rate.
Static and machine learning classifiers add another layer: vocabulary-based features and similarity-distance measures between test runs can predict which tests are likely to be systemically flaky before they even fail repeatedly.
- Reruns: simple, but costly at scale and sometimes inconclusive on a single pass.
- Differential coverage: avoids reruns entirely by comparing failure location to code changes.
- Static and ML classifiers: predict flakiness risk from test code patterns and historical distance metrics.
- Clustering: groups co-occurring failures to expose systemic root causes rather than isolated bugs.
Roughly 75% of flaky tests belong to clusters with a mean cluster size of 13.5, according to research on systemic flakiness covering 24 projects and 10,000 test-suite runs. That means most flakiness is not random noise scattered across your suite. It is a small number of shared causes, mainly intermittent networking and external dependency instability, manifesting as dozens of individually “random” failures. Static ML models using distance features achieved an R-squared up to 0.74 in predicting which runs belong together.
Choose reruns when you have small suites and spare CI budget, differential coverage when you want to cut rerun costs at scale, and clustering when your flaky rate is high enough that individual fixes have stopped moving the needle.

A practical detection checklist you can run in CI
Detecting flaky tests reliably does not require a research team. It requires discipline about what you capture on every failure.
- On failure, record the commit SHA, job environment, test node ID, full logs, and a coverage diff against the previous passing run.
- Check whether the coverage diff shows any overlap with the current change set. No overlap means the failure is a strong flaky candidate before you spend a single rerun.
- Apply a controlled rerun policy: 2 to 3 retries maximum, each in an isolated process or fresh JVM fork, never a shared one.
- If a test fails intermittently across three or more runs without a clear fix, quarantine it and open a ticket rather than letting it block merges indefinitely.
- Instrument tests to emit random seeds, timestamps, and environment flags, so a failure six weeks from now can be compared against today’s.
- Enable co-occurrence windowing: track which tests fail together in the same run or the same hour, and flag clusters for review.
Pro Tip: Store failure artifacts for at least 90 days. Clustering and ML-based flakiness prediction only work with enough historical data to compare against.
This sequence keeps cost low because you only spend rerun budget on tests that survive the coverage check, and it builds the dataset you need for clustering later.
Triage and remediation: quarantine, fix, or suppress
Once a test is flagged, the triage flow should be mechanical: reproduce locally with the same seed and environment flags, gather the failure evidence already captured in CI, and assign an owner within a set window, typically one sprint.

Quarantine criteria should be explicit rather than ad hoc. A reasonable gate: quarantine after three flaky occurrences in 30 days, with a hard deadline (two weeks is common) before the test is either fixed or deleted. Quarantining without a deadline tends to become permanent, which quietly lowers the bar for your entire suite.
Common fix patterns worth trying first:
- Isolate shared state: reset databases, caches, and global variables between tests rather than relying on run order.
- Mock unstable external services instead of hitting real APIs or third-party endpoints during the test run.
- Replace sleeps with wait-for predicates that poll for the actual condition rather than guessing a fixed delay.
- Use resource pools for ports, files, and connections so parallel tests cannot collide.
After a fix, verify stability over multiple runs, not just one green build, and track whether the fix reduced the size of the cluster it belonged to rather than just silencing one test.
Tooling and automation options for managing flaky tests
CI platforms increasingly ship flaky-test management natively. Azure Pipelines, for example, documents built-in detection, quarantine options, and configuration for handling flaky results without custom scripting, which is a reasonable baseline for teams that do not want to build this from scratch.
Beyond built-in CI features, a few categories matter when you are choosing what to add:
- Analytics and grouping tools that correlate failures across runs and segment by environment, branch, or node.
- Historical trending dashboards that show whether your flaky rate is improving or getting worse over time.
- Ticketing and ownership integration so a flagged flaky test automatically creates a tracked, assigned ticket instead of sitting in a log.
- Audit trails for quarantine decisions, so nobody has to guess why a test has been skipped for three months.
Teams managing complex data pipelines face a similar observability challenge outside of test suites; the same instinct to capture environment and correlation data applies to CI/CD patterns for Power BI deployment pipelines, where catching drift early avoids expensive downstream fixes. The right tool depends on scale: a small team might get by with CI-native retry and quarantine settings, while a large suite with hundreds of flaky tests needs dedicated correlation and clustering to make triage tractable.
Veridical perspective: catching flaky conditions before they merge
A lot of flakiness is introduced, not discovered. A pull request that adds an uninitialized global, a new external service call, or time-dependent logic without a wait condition creates the exact conditions that show up as flaky failures weeks later.
Our AI code review for GitHub looks at these risk patterns directly in the diff, with verified findings and an F1-based advisory score that flags risky changes before merge rather than after they start failing intermittently in CI.
- We surface new external calls, shared-state writes, and timing-dependent code paths with inline evidence.
- Our advisory score is designed to reflect risk rather than relying on a generic linting heuristic.
- Combining pre-merge checks with CI-side flaky detection narrows the surface area for new flakiness while catching what already exists.
Engineering recommendation: prioritize systemic fixes first
When clustering shows that a handful of root causes account for most of your flaky failures, fix the cluster, not the symptom. A shared mock contract or a fixed network timeout config often resolves a dozen “unrelated” tests at once. Keep quarantine and retries as short-term stopgaps, but measure success by cluster size reduction and falling triage time, not by how many individual tests you have patched this month.
— Łukasz
Reduce flaky surface before it reaches CI
We built Veridical to catch the risky diffs that turn into flaky tests weeks later, with verified evidence instead of guesswork at every pull request.

Pricing starts with a free allowance for open-source projects, and paid plans run from $39 a month on the Standard tier to $69 a month on Pro for teams that need higher review volume and advanced features.
FAQ
How do you detect flaky tests?
Capture failure metadata, including coverage diffs, on every red build, then check whether the failing test touched any changed code. If coverage did not change, treat it as suspected flaky before spending a rerun, following the DeFlaker approach.
What does “Nx detected a flaky task” mean?
It means a build tool identified that the same task produced different pass or fail results across runs without a code change, which is the same signal underlying most flaky test detection: inconsistent outcomes against unchanged inputs.
What causes test flakiness?
The most common causes are unstable external dependencies, timing and concurrency race conditions, test order dependencies with shared state, and environment differences between machines. Research on systemic flakiness found that intermittent networking and external dependency instability drive the largest clusters of related failures.
How do you fix flaky tests?
Isolate shared state between tests, mock unstable external services, and replace fixed sleeps with predicates that wait for an actual condition. Verify any fix across multiple runs and confirm it shrank the related failure cluster rather than just silencing one test.
How much does flaky testing actually cost teams?
Industrial case study data estimated developers spend at least 2.5% of productive time on flaky-test related work, with manual investigation, not reruns, accounting for most of that cost.
Sources
- Systemic flakiness research (arXiv preprint)
- DeFlaker: automatically detecting flaky tests (ACM, 2018)
- Cost of flaky tests in CI: an industrial case study (ICST 2024)
- Manage flaky tests - Azure Pipelines (Microsoft Docs)