Engineering7 min readSeptember 28, 2026

Test Data Management: Why Your 'Flaky' Tests Might Not Be Flaky At All

TL;DR

A test that fails intermittently isn't necessarily flaky — it might be data-dependent, failing because shared state changed under it, not because it's internally nondeterministic. Rerun the failing test alone: if it passes in isolation but fails in the full suite, look for shared accounts, missing teardown, or order-dependent fixtures before touching timeouts. Fix with per-worker data isolation, factories over shared fixtures, and API-based teardown.

A team spends months blaming "flaky tests" for intermittent CI failures — rewriting waits, adding retries, tweaking timeouts — and the failure rate barely moves. Then someone finally digs into a specific failure and finds the real cause: two tests running in parallel both wrote to the same seeded account. It was never nondeterminism. It was test data.

Flaky vs. Data-Dependent: Different Problems, Same Symptom

A genuinely flaky test fails nondeterministically for reasons internal to itself — a race condition against an animation, a network call that resolves at a slightly different time each run. See our reliability playbook for that category. A data-dependent failure looks identical from the CI log — red, then green on rerun — but the cause is external: the test depended on data state that something else in the run changed out from under it. The fix for one does nothing for the other, which is why teams that misdiagnose this spend months adding retries that never actually reduce the failure rate.

The Usual Suspects

  • Shared seeded accounts across parallel workers. Two tests running concurrently both log in as the same demo user and both mutate its cart. Neither test is broken in isolation — together, they race.
  • Tests relying on production-like data that drifts. A test asserts on a specific product name or price that existed in the seed data six months ago and has since been updated by someone else's migration.
  • No teardown between runs. A test creates an order and never cleans it up. The next run's "first order in the list" assertion now points at yesterday's leftover record.
  • Order-dependent fixtures. Test B only passes if test A ran first and left the database in a particular state — a dependency that's invisible until someone reorders the suite or runs B in isolation.

How to Tell the Difference in 10 Minutes

Rerun the failing test completely alone, outside the full suite. If it passes reliably in isolation but fails intermittently inside the full run, the cause is almost certainly shared state, not nondeterminism internal to the test. Next, grep the suite for other tests touching the same record, account, or resource — if two tests both mutate demo@example.com's cart, you've found it.

Fixing It

  1. Isolate data per worker. Each parallel worker gets its own account or tenant, generated fresh, so no two tests can collide on the same record. This is the data-layer equivalent of the isolation strategy in parallel execution without chaos.
  2. Prefer factories over shared fixtures. Generate the exact data a test needs at the start of that test, rather than relying on a pre-seeded dataset every test quietly depends on.
  3. Reset state through the API, not the UI. Tearing down via a direct API call is faster and more reliable than clicking through a delete flow, and it removes one more place timing can introduce flakiness.
  4. Make seed data deterministic. If a baseline dataset is necessary, version it and treat changes to it like any other schema migration — reviewed, not silently overwritten.

Diagnosing this well is most of the work in fixing flaky tests for good — the mechanical fixes only help once the root cause is correctly identified. If your team has been chasing flakiness without progress, book a demo and we'll help you find out whether it's really flakiness at all.

Tags

test dataflaky testsparallel testingCI/CD

See QA Guardian in action

Everything we write about is what we build and run every day. Book a demo and we'll show you on your own codebase.