---
title: "Flaky tests: classify failures by cause | qventis blog"
description: Flaky is a symptom, not a diagnosis. How to classify every failed check by cause, from data and environment to timing and vendor change, using evidence.
---

[Skip to content](https://qventis.ai/blog/flaky-tests-classify-failures-by-cause#main)

[![qventis](https://qventis.ai/hubfs/raw_assets/public/qventis-one/images/qventis-logo.svg)![qventis](https://qventis.ai/hubfs/raw_assets/public/qventis-one/images/qventis-logo-reverse.svg)](https://qventis.ai/)

- Product
  
  [**Qventis Engine**Any app, any stack, in plain English. Zero AI tokens at run time.WebDesktopTerminalAPIDataOne plain-English suite](https://qventis.ai/engine)
  
  [**Qventis One**One pane of glass across every app](https://qventis.ai/one) [**App Model**Your app, learned once](https://qventis.ai/one/app-model) [**Quality View**Live quality signals and alerts](https://qventis.ai/one/quality-view) [**Agent gateway**Coding agents create and update tests](https://qventis.ai/platform/agent-gateway)
  
  [**Studio**Write tests in plain English](https://qventis.ai/platform/studio) [**Eight engines**Functional to accessibility](https://qventis.ai/one#engines) [**Integrations**CI, ALM and chat](https://qventis.ai/platform/integrations) [**Trust and security**People approve every AI change](https://qventis.ai/platform/trust)
  
  How it works[Detect](https://qventis.ai/release-intelligence)→[Assess](https://qventis.ai/one/app-model)→[Automate](https://qventis.ai/engine)→[Assure](https://qventis.ai/one/quality-view)
- Enterprise apps
  
  [**Release Intelligence**Every vendor update read, traced to your processes and tested before it lands.PreviewTestProdNextImpact known before the window opens](https://qventis.ai/release-intelligence)
  
  [**Oracle**Quarterly updates](https://qventis.ai/platforms/oracle)[**Workday**R1 and R2](https://qventis.ai/platforms/workday) [**SAP**S/4HANA releases](https://qventis.ai/platforms/sap)[**Salesforce**Three a year](https://qventis.ai/platforms/salesforce) [**ServiceNow**Family releases](https://qventis.ai/platforms/servicenow)[**Dynamics 365**Release waves](https://qventis.ai/platforms/dynamics-365) [**nCino**Two calendars](https://qventis.ai/platforms/ncino)[**Coupa**Major releases](https://qventis.ai/platforms/coupa) 
  
  [**All 40+ apps**ERP, HCM, CRM, ITSM, banking, insurance](https://qventis.ai/engines/packaged-apps)
  
  [**Release testing**Test only what a change affects](https://qventis.ai/solutions/release-testing) [**Industries**Regulated and complex businesses](https://qventis.ai/industries) [**Legacy desktop**Apps automation never reached](https://qventis.ai/solutions/legacy-desktop) [**Quality CoE**One standard across the business](https://qventis.ai/solutions/quality-coe)
- Learn
  
  [**Blog**Field notes on releases, quality and AI](https://qventis.ai/blog) [**Briefs**Platform briefs by email](https://qventis.ai/resources) [**Customers**Example scenarios by role](https://qventis.ai/customers) [**Support**Guides for customers](https://qventis.ai/support)
- Company
  
  [**About qventis**Why we built one platform](https://qventis.ai/about) [**Editions**Start with one need, grow](https://qventis.ai/editions) [**Contact**Send us a request](https://qventis.ai/contact)

Search platforms, releasesCtrl K

[Book a demo](https://qventis.ai/contact#demo)

[Blog](https://qventis.ai/blog)/STLC pitfalls

# Flaky tests: classify failures by cause

Flaky is a symptom, not a diagnosis. How to classify every failed check by cause, from data and environment to timing and vendor change, using evidence.

qventis teamAugust 21, 2026

Every large suite has a list of tests that fail now and then and pass on a second try. The usual response is a rerun policy: retry failed tests once or twice, and count them as passed if they go green. It keeps the pipeline moving. It also teaches the team to ignore failures, and sooner or later a real defect gets retried into a pass.

"Flaky" describes what a failure looks like from the outside. It says nothing about why it happened. Treating it as a diagnosis is how teams end up with a growing quarantine list and no idea which entries matter.

## Seven causes behind most failures

In enterprise suites, failed checks almost always fall into one of a small number of causes. Each has signals that set it apart, and each belongs to a different owner.

| Cause | Typical signal | Owner |
| --- | --- | --- |
| Application defect | Same check fails the same way, reproduces by hand | Development or vendor |
| Environment | Many unrelated tests fail in the same window | Platform or environment team |
| Test data | Fails on one record, passes on a fresh one | Test data owner |
| Timing | Fails at a wait, passes on rerun | Test author |
| UI change | Control not found after a deployment | Test author, via a heal |
| Test logic | Wrong expected result or order dependence | Test owner |
| Vendor change | Starts after an update, tied to a feature | Release lead |

Only some of these are truly intermittent. Timing, environment and shared data produce failures that come and go. Defects, UI changes and vendor changes usually fail consistently once they appear, but get mislabeled as flaky because a rerun happened to land before or after a fix or a deployment.

## Classify at first sight

The time to classify a failure is the first time it happens, while the evidence is fresh. Each failed check should carry enough to decide its cause without running it again.

- The step that failed, with what it expected and what it saw
- A screenshot or response body from the moment of failure
- Timings for the failing step and the steps around it
- The state of the test data record it used
- Environment health for the same window, including integrations
- Recent change events: deployments, configuration changes and vendor updates

With that in hand, most failures sort themselves. A timeout on a step that usually takes a second, while ten other suites also timed out, is environment. A record already in "Approved" status when the test expected "Draft" is data. A failure that began the morning the quarterly update reached the test environment, on a page named in the release content, is a vendor change.

## Stop rerunning to green

Reruns are a reasonable way to gather evidence. They are a poor way to decide a result. If you keep them, put limits around them:

- Record both the failure and the pass, never only the final state
- Classify the original failure even when the rerun passes
- Never rerun a failure whose expected result was actually checked and missed
- Report the rate of pass-on-rerun by cause, so the trend is visible

The third rule matters most. A check that compared a value and got the wrong one did not flake. It found something.

## Quarantine with an expiry date

Sometimes a test has to be pulled from a gate while its cause is fixed. That is fine, if the quarantine is managed like any other open issue.

- Every quarantined test has a classified cause and a named owner
- Every quarantine has an expiry date, after which the test returns or is retired
- Quarantined tests still run and report, outside the gate
- The quarantine list is reviewed with the same weight as open defects

A quarantine without an owner and a date becomes a place where coverage quietly disappears.

## Fix causes, not tests

Once failures are classified, look at the totals by cause rather than by test. The pattern usually points to a few system-level fixes that remove many failures at once.

If shared test data accounts for a large share, the answer is data isolation or reservation, not editing each test. If environment failures cluster at certain times, the answer may be a scheduling conflict with batch jobs. If timing failures cluster on one kind of page, the waits for that page type need a shared fix. Working test by test hides these patterns and spends effort where it helps least.

## Make the classification part of the record

A classified failure is useful long after the run. It tells the next person what to check first. It shows release owners which failures are about the application and which are about the test system. And it gives an honest view of suite health, separating "the product has problems" from "our testing has problems", which are very different conversations.

## How Qventis One helps

The Qventis Engine records every check with what it expected and what it saw, plus screenshots, logs and timings, and runs the same way every time, so a failure points to the application or the environment rather than to a model's interpretation. Quality View groups failures by cause, separating product defects from test and environment issues, tracks flaky tests over time, and shows recent change events, including vendor updates from Release Intelligence, next to each result. [Book a demo](https://qventis.ai/contact#demo) to see failures sorted by cause.

**The Qventis Engine** automates any app in plain English with zero AI tokens at run time. Qventis One scales it to one view of quality.

[Book a demo](https://qventis.ai/contact#demo)

More in STLC pitfalls

### [Oct 9, 2026 End-to-end handoffs: the accrual that never clears Every module passed its tests, yet the receipt accrual never cleared. Why handoffs between modules and systems go untested, and how chain tests catch them.](https://qventis.ai/blog/end-to-end-handoffs-accrual-that-never-clears)

### [Sep 22, 2026 Regression suites that only grow Suites grow because adding feels safe and deleting feels risky. Why tag filters do not fix it, and how to select tests as a plan with a purpose and a budget.](https://qventis.ai/blog/regression-suites-that-only-grow)

### [Sep 8, 2026 Test data bottlenecks in shared environments Shared test environments turn data into the slowest part of testing. Six provisioning strategies, when each fits, and how to stop tests colliding.](https://qventis.ai/blog/test-data-bottlenecks-shared-environments)

[![qventis](https://qventis.ai/hubfs/raw_assets/public/qventis-one/images/qventis-logo.svg)![qventis](https://qventis.ai/hubfs/raw_assets/public/qventis-one/images/qventis-logo-reverse.svg)](https://qventis.ai/)

[Support](https://qventis.ai/support)[Trust](https://qventis.ai/platform/trust)[Privacy and terms](https://qventis.ai/legal)Contact

© 2026 qventis.aiThird-party product names are trademarks of their respective owners.

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "qventis team",
    "url" : "https://qventis.ai/blog/author/qventis-team"
  },
  "dateModified" : "2026-10-11T09:03:35.897Z",
  "datePublished" : "2026-08-21T15:00:00.000Z",
  "headline" : "Flaky tests: classify failures by cause | qventis blog",
  "mainEntityOfPage" : {
    "@id" : "https://qventis.ai/blog/flaky-tests-classify-failures-by-cause",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject"
    }
  }
}
```