Mida's agent reads your own results and tells you what is worth testing next.

Try it for FREE now

How we built these benchmarks

Every rate comes from tests real teams ran on Mida — not a survey, not a literature review, not an estimate. This page states exactly how the numbers are produced, including the parts weaker than we would like.

Where the data comes from

Concluded and stopped experiments run on Mida between August 2023 and August 2026, across thousands of tests and hundreds of accounts. Each variant contributes one record: what changed, how much traffic it saw, its conversion rate against its control, and how it resolved.

Which records reach a published figure

Not every recorded variant is publishable, and the largest cut is for traffic rather than anything to do with the result. Of every variant we hold, 61% met the sample floor, and 55% survive every filter to reach a published rate. The rest are dropped for the reasons below.

The sample floor

A variant has to have seen enough traffic for the result to mean anything. Below that bar a flat result is not evidence the change failed — it is evidence the test never ran long enough to say. 39% of non-control variants are excluded on this basis alone, which is the main reason our win rates are higher than a raw count over all tests would give.

Removing measurement artefacts

Some recorded wins are setup faults. The clearest signature is a variant converting above 50% while its control converts under 5% — which happens when a conversion goal fires on the variant's own page, so arriving is converting. Split URL tests are where this concentrates, and the worst examples record lifts in the tens of thousands of percent. Every record with that signature is removed before anything else is counted. They barely move win rates but badly distort lift.

What counts as a win

A variant that beat its control on that test's own primary goal, and cleared one bar we apply to every test: at least a 95% chance to beat the original, or for the few tests using frequentist stats, a one-sided p-value of 0.05 or lower. Flat and losing variants are both in the denominator.

Why one bar: every Mida test picks its own confidence level, and they do not all pick the same one. In this data, 43.1% were set to 95%, 32% to 80%, 22.8% to 90% and 1.8% to 99%. Taking each test's own verdict would count an 80% result as a win next to a 95% one, and inflate the categories where the looser settings happen to cluster. So a win here can be stricter than the verdict a team saw on its own dashboard, which still uses the team's own setting.

Losses are the one exception. A losing variant's chance to beat the original is stored as zero, with no size attached, so the same bar cannot be turned around and applied to it. A loss here is the loss each test reported.

Multiple variants in one test

Just under a third of variants share a test with at least one sibling. A test with four variants has more chances for one to clear its bar, so the rates carry some inflation from that. We report each variant as one observation rather than correcting for it, because these are descriptive rates rather than decisions.

What these tests measured, and why it matters

Every test here counts something. What it counts turns out to predict whether it "wins" almost as strongly as what was changed.

measured outcomeshare of testswin rate95% interval
Purchase / order complete45%12%10–13%
Signup / registration12%10%8–13%
Lead / form submission11%15%12–18%
On-page click / engagement9%19%16–23%
Product or content page reached5%17%13–22%
Booking / appointment4%7%4–11%
Add to cart4%14%10–19%
Reached a later funnel step4%11%7–16%
Outbound click to a provider3%15%10–22%
Scroll or time on page1%20%12–31%
Subscription1%3%0–14%

Read the bottom of that table against the top. A test measured on scroll depth or an on-page click wins around one time in five. A test measured on a completed purchase wins about one time in eight, and one measured on a booking or a subscription less often still. The intervals for on-page clicks and purchases do not overlap, so this is a real gap and not sampling noise.

None of that means the softer metrics are wrong to use. A click can be the right thing to optimise, and an early-funnel test often has no purchase to point at. It does mean a win rate is only worth as much as the outcome behind it, and that two rates measured on different outcomes are not comparable. Every page here shows what its tests measured for exactly that reason.

Outcomes are classified from the goal's configured value and the name the customer gave it, in two passes: deterministic rules where the signal is unambiguous, and a language model over the remainder, which is often not in English. 0.4% could not be placed and are excluded from the table. The goal names themselves are customer content and are never published — only the category.

Why you should not compare categories directly

This is the biggest caveat on this data, and it is one we got wrong in our first publication.

Every rate here is partly a statement about what those tests chose to measure, not only about what they changed. Goals divide into business outcomes such as a purchase, funnel steps such as reaching a checkout page, and proxies such as a click or a scroll. They do not clear the bar equally often:

what the test measuredwin rate
A proxy — a click, a scroll, time on page13%
A funnel step — reaching a later page10%
A business outcome — a purchase, a signup9%

If the mix were the same across categories this would not matter. It is not: proxy goals account for anywhere from 2% to 41% of a category's tests. Some kinds of change are simply more often measured against a soft signal than others.

The consequence is direct. Comparing only tests measured against a business outcome — like for like — the gap between the highest and lowest category falls from 9 percentage points to 6, and at that point the confidence intervals overlap. Layout and headline, the comparison our pages leaned on hardest, go from 15% versus 8% to 12% versus 10% — a difference we cannot distinguish from noise at these sample sizes.

So the per-category rates are sound as descriptions and weak as a ranking. "15% of layout tests won" is true. "Layout beats headline, so test layout" is not supported once you account for what each was measured against. We have softened those recommendations across the site rather than leaving them standing, and every comparison table now carries this note.

The honest version of the advice: pick the change you have a real hypothesis for, and if you want to compare categories yourself, compare ones measured the same way.

What counts as a change type

Labels are assigned automatically from the variant definition, not by hand.

how the label was derivedrecords
Code diff3617
Text only2120
Redirect target1658
Screenshot comparison533
Structural diff6

Records the labeller could not classify are dropped rather than filed under a catch-all. Tests changing several things at once are labelled by their dominant change, so a "headline" test here may have carried minor styling changes too. Single-variable discipline is a property of the teams running the tests, not something we can enforce retrospectively.

Industry comes from the verified domain

Our first publication had no industry breakdowns, because the only industry field available was a self-reported onboarding answer. Industry is now derived from the domain the Mida pixel actually fired on, covering 97.8% of published variants.

That change was worth making: where we hold both values, the self-reported industry and the verified one disagree 65% of the time. The substitutions are not near misses — sites self-reporting as Professional Services that are plainly e-commerce, sites self-reporting as SaaS that are agencies. Any industry benchmark built on the survey field would have been wrong about roughly two rows in three.

Domains are classified in two passes with an explicit "unclear" option at each stage. The first reads the domain and its real page paths; the second fetched the unplaced homepages and classified them from each site's own title and description. A small remainder is still unresolved and are excluded from every industry figure — some no longer respond, some sit behind a bot wall, some return a page that still does not say what the business does. A bot wall is recorded as blocked rather than fed to the classifier, since an interstitial would otherwise be labelled SaaS every time. This is a classification, not a registry lookup: treat a single label as good-not-certain.

What we do not publish

  • No customer copy. No headline, button label or body text from any account appears here. Descriptions of what changed are machine-rewritten to remove verbatim wording, brand names, product names, prices, selectors and domains.
  • No URLs, no account names, no test names.
  • Nothing unreviewed. 58 descriptions the rewrite left with residual identifying detail are withheld from every example. They remain in the denominators, because a description we cannot quote is still a test that ran.
  • No cell under 30 tests. Thinner categories are not published at all rather than published with a caveat.

What these numbers cannot tell you

They do not tell you what will work on your site. A 15% win rate for layout changes is a base rate across hundreds of sites, not a prediction about yours. The useful move is to read it as odds: at 15%, most layout tests do not win, so plan for several before one lands, whichever category you pick.

Goal quality varies. About 41% of these tests measured a business outcome such as a purchase, 38% a funnel step and 21% a proxy such as a click. A proxy win is weaker evidence than a purchase win, and the headline rate on each page mixes all three, which is why the comparison tables carry the note above.

Survivorship runs through everything. These are tests that got set up, launched and left running long enough to read. Abandoned tests, and ideas nobody built, are invisible.

The corpus skews to what people test, not what works. Styling is the most-run category. That tells you it is cheap to build, not that it is important.

Updates

Figures are recomputed periodically, and every benchmark page shows the date its figures were last updated. We do not revise past numbers silently. Where a recomputation moves a rate, the page says so.

Revision, August 2026. Industry re-derived from the verified domain the pixel fired on, including a homepage fetch for domains a name alone could not place. That made industry benchmarks publishable for the first time, at 97.8% coverage.

Who ran these tests

This is every Mida account that ran a readable test — in-house marketers, founders, product teams, and agencies working on client sites. Nothing here is filtered by who ran the experiment or how experienced they are.

Low win rates are normal in experimentation, including at the top end. Microsoft's experimentation team, reporting on its own platform, found that only about one third of ideas improve the metric they were designed to improve — and that roughly another third actively hurt it. That is a dedicated experimentation organisation with research, prioritisation and review behind every test.

A mixed population like this one runs below that. The gap is roughly what disciplined practice buys you: ideas grounded in research rather than opinion, one variable at a time, and tests built so the result can actually be read.

A win is a variant that beat its control on that test's primary goal with a statistically significant result. Tests that never got enough traffic to say anything either way are excluded. The methodology has the full detail, including what these numbers cannot tell you. Figures last updated .