Create A/B tests by chatting with AI and launch them on your website within minutes.

Try it for FREE now
CONTENTS
How to
6
Min read

Test Collisions: What Happens When Two A/B Tests Run on the Same Page at Once

Mida Team
Mida Team
|
5-star rating
4.7
Reviews across G2 & Capterra
Test Collisions: What Happens When Two A/B Tests Run on the Same Page at Once

Quick answer

When two experiments touch the same users, page, or flow at the same time, their effects don't simply add together. Researchers call this a statistical interaction — formally, "a statistical interaction between two treatments A and B exists if their combined effect is not the same as the sum of two individual treatment effects" (Kohavi et al., 2013). When that happens, both tests can read out results that are quietly wrong, and neither one will throw an error to tell you.

Key takeaways

  • Test collisions aren't rare — for teams without dedicated experimentation infrastructure, the practical ceiling is 2–4 concurrent tests before interaction effects start corrupting data.
  • The four-group check (both variants of both tests) is the same logic Statsig and Harness recommend, just at a smaller scale — you don't need Bing's daily interaction-detection job to run it.
  • Prevention comes down to a shared registry of what's live, mutual exclusivity where it's cheap, and re-reading any result that moves right when another test ships.

What exactly is a "test collision"?

Picture two tests running at once: one changing your hero headline, another changing the CTA button color directly below it. Individually, each might have a clean, measurable effect. Together, they can produce something different — a bolder headline might make the new button color pop in a way the old one never would have, or clash with it. As Harness's engineering team puts it, "interaction effects occur when two experiments modify the same page element or user flow, creating combined impacts that differ from isolated effects." Neither test's dashboard will tell you this is happening — both will just report numbers, and the numbers will be wrong by an amount neither team can see from inside their own results.

Is this a real risk, or an edge case for companies running thousands of tests?

It's tempting to assume this is a Big Tech problem. Statsig data scientist Kane Luo notes that "a study by Microsoft shows that interaction effects are rare" — but that's rare given Microsoft's infrastructure, which includes dedicated tooling most teams don't have. Bing's experimentation platform, for instance, "runs a daily job that tests each pair of experiments for additivity of their treatment effects" across hundreds of concurrent tests (Kohavi et al., 2013). A team without that safety net isn't automatically safer — it just can't see the problem.

Mida's own test prioritization guide puts a number on it for teams without that infrastructure: "there's a real ceiling on how many concurrent tests a single team can run before interaction effects start corrupting the data — for many teams it's somewhere in the 2–4 range." That's a much lower bar than "hundreds," and it's the number that actually applies to most CRO teams. Even hyperscale companies take the risk seriously enough to build real infrastructure around it — LinkedIn, for example, runs paired "two-sided randomization" experiments to separate production-side and consumption-side changes specifically to avoid this kind of contamination (Gupta et al., 2019).

Free A/B Testing Tool

Run your next A/B test the right way

Visual editor, 15 KB script, GA4-native — and free forever up to 100,000 monthly visitors. No developer required.

✓ Visual editor✓ 15 KB script✓ GA4 integration✓ Free up to 100k visitors
Try Mida free →

How do you actually tell if two tests are contaminating each other?

At platform scale, the answer is statistical: "regression models with interaction terms are your friend here… you can actually quantify whether Test A is messing with Test B's results" (Statsig, 2025). Statsig's own detection method splits exposed users into the four combinations of both tests' variants and checks whether the differences between groups are larger than chance would produce (Luo, 2025) — the same four-group logic Harness recommends teams check manually when they suspect a collision.

You don't need a dedicated data science team to use a rougher version of this. If a test that had been flat for weeks suddenly moves the same week another test ships nearby, or a result reverses direction right after an unrelated test ends, that's the practical version of the same signal: it's disagreeing with itself in a way the single-test model can't explain.

What do you actually do about it?

There are two real approaches to prevention, and they trade off against each other. Mutually exclusive experiment groups are the simple option — Statsig calls them "your first line of defense: each user gets assigned to exactly one test" — but they cost you sample size and testing throughput, since every experiment is competing for the same pool of users. The more scalable option is what Google and Microsoft call "numberlines" or "layers": infrastructure that guarantees experiments on the same layer get an exclusive random sample, so unrelated tests can run in parallel without touching each other (Gupta et al., 2019). That's a real engineering investment most CRO teams won't build for a handful of concurrent tests, which is exactly why Mida's own guidance on server-side vs. client-side testing recommends a lighter middle ground: a shared experiment registry so every team can see what's live and where, before two tests ever land on the same element.

Whichever you pick, the one option statistician John D. Cook warns against is picking none: "if you only test one factor at a time, you're betting that interaction effects don't matter." That bet is invisible right up until it isn't.

The takeaway

A test collision isn't a rare, exotic failure mode reserved for companies running thousands of experiments — it's a structural risk of stacking any two tests on the same page or flow, and the "2–4 concurrent tests" ceiling means most CRO teams are closer to this problem than they think. You don't need Bing's daily interaction-detection job to manage it. A shared registry of what's currently live, mutual exclusivity where it's cheap to enforce, and a habit of double-checking any result that moves right when another test ships, will catch the overwhelming majority of collisions before they quietly cost you a decision.

Related reading

Free A/B Testing Tool

Run your next A/B test the right way

Visual editor, 15 KB script, GA4-native — and free forever up to 100,000 monthly visitors. No developer required.

✓ Visual editor✓ 15 KB script✓ GA4 integration✓ Free up to 100k visitors
Try Mida free →

Sources

Run Your First A/B Test in Minutes — 100,000 MTU Free

Visual editor, AI-powered variant creation with MidaGX, GA4 integration, and more. No credit card required, no time limit.

Decorative graphicDecorative graphicDecorative graphicDecorative graphic