Cited from real sources 6 min read Written by Arun Agrahri Updated September 2026

A framework by Ronny Kohavi

Ronny Kohavi's Sample Ratio Mismatch Check: Test the Split Before the Result

A sample ratio mismatch is a gap between the traffic split you designed an A/B test to get and the split it actually got. Ronny Kohavi worked on experimentation at Amazon, Microsoft and Airbnb, and he calls it by far the most common way an experiment breaks. Check the split before you read a single metric. A 50/50 test that comes back 50.2/49.8 on a million users is broken, however good the result looks.

Why a clean-looking test can still be wrong

8% of experiments, invalid.

Roughly that share of Microsoft's experiments failed the sample ratio check, on a platform running about 20,000 tests a year.

Ronny Kohavi on Lenny's Podcast The ultimate guide to A/B testing Watch at 57:20

The framework

A split that is off by 0.2% is a red flag

You design a test to send half your users to control and half to treatment. A random hash decides who goes where, so the counts that arrive should match the design. Kohavi's rule is blunt:

if you get something off from 50 percent it's a red flag
Kohavi on what counts as a mismatch Watch at 56:09

The gap does not have to look big. With a million users, 50.2 against 49.8 sounds like noise. Put the counts into the formula and it says a split that uneven should turn up about once in half a million experiments. Kohavi keeps a spreadsheet that does the sum. In a 2016 talk he read out an alert from that week: 821,588 users in control and 815,582 in treatment. The ratio was 50.2%, and the odds of that by chance were one in 500,000.

This is also one of the few places where a p-value, the chance of seeing a result if nothing real is going on, means what people think it means. The design fixes the split at 50/50. So a tiny p-value really is the chance of seeing that split by luck.

The split matters because the traffic you lose is usually not a random slice. Often it is the people who viewed zero pages, or the heavy clickers who got filtered out as bots. Kohavi, on the first mistake his team made:

you tend to get very extreme results because the traffic that you're missing is usually highly biased in some way
Kohavi on why a broken split distorts the result Watch at 33:27

A failed split makes results look either super good or super bad. The bad ones kill ideas that might have worked. The good ones get celebrated. So the check comes first. If it fails, the metrics are not evidence of anything.

How to apply it

How do you check an A/B test for sample ratio mismatch?

The first three steps run on every test. The last three make the check stick.

  1. 1

    Write down the split before launch.

    The check compares what arrived against what you designed, so the design has to exist first. Kohavi's examples run from 50/50 to 20/20, and the same test works for unequal splits.

  2. 2

    Run the numbers on the split.

    Compare the users in each variant against the designed ratio with a goodness-of-fit test. Chi-squared is the standard one. If the probability of that split is tiny, stop.

  3. 3

    Hide the scorecard when the check fails.

    Microsoft first showed a warning banner, and people presented the results anyway. The team blanked the scorecard behind an OK button. Then it marked every number with a red line, so a screenshot still carried the warning.

  4. 4

    Look in the data pipeline before the randomizer.

    The hash almost always works. What breaks is a step that removes data unevenly. Kohavi names bot filtering, fraud removal and bad-traffic rules that hit one variant harder. A marketing campaign that sends its visitors into only one variant does the same.

  5. 5

    Run A/A tests as well.

    Split traffic between two identical versions. At a 0.05 threshold, the system should find no significant difference about 95% of the time. Kohavi calls this one of the most useful tools for finding bugs in the software, the method and the variance estimates.

  6. 6

    Count your past invalid results.

    When vendors added the check, they found about 8% of the experiments they had already reported to users were invalid. Teams running it for the first time often find more.

if you run experiments and you do not check for sample ratio mismatch copy this slide carefully make sure you're on it it is the most common alarm that fires for us at Microsoft
Kohavi at CXL Live 2016 Watch at 15:20

Boundary conditions

When does the SRM check mislead you?

Works best when

  • Users are assigned by a hash, and every assignment is logged.
  • Traffic is large enough that a fraction of a percent shows up.
  • The platform runs the check automatically on every test.
  • Someone owns the call to stop reading a failed test.

Fails when

  • A warning banner is all that stands between the team and the result.
  • You debug the randomizer and never look at what the pipeline removed.
  • A passing check is read as proof that the whole test is sound.
  • Traffic is so small that only a large imbalance shows up.

Kohavi tells the banner story in two of his recordings. Telling people not to trust a result did not stop them presenting it:

we just put a banner saying you have a sample ratio mismatch do not trust these results and we noticed that people ignored it
Kohavi on why the warning had to become a block Watch at 59:21

Passing the check leaves every other bug in place. Right after SRM on Lenny's Podcast, Kohavi moved to Twyman's law. Say your experiments normally move a metric by under 1%, and one suddenly moves it by 10%. Hold the celebration dinner and investigate. His Overall Evaluation Criterion covers the other half: deciding what the test should optimize before it runs.

Sources

Where Kohavi discusses this

Of the seven long Kohavi recordings we pulled for this page on 2026-09-30, five raise the sample ratio check, from a 2016 conference talk to a 2024 lesson.

Where experts disagree

Where operators disagree: check every split, or ship the winners fast?

Ronny Kohavi

says no result counts as evidence until the split passes the sample ratio check. Lost traffic is biased traffic, so a failed split makes results look super good or super bad. His team at Microsoft blanked the scorecard rather than let anyone read one.

John McEvoy

ran about ten paywall and onboarding experiments on his own app over two or three months. Conversion went from 0.5% to 8%, and revenue from around $8K to over $30K a month. His rule: A/B test paywalls and onboarding flows relentlessly (Starter Story, 9:12).

The deciding variable is whether anything outside the test can confirm the win. McEvoy had that check, because monthly revenue nearly quadrupled and a broken split cannot fake revenue. A team shipping small lifts it can see only in the scorecard has no outside check. At the rates Kohavi reports, about one test in twelve has a broken split, so the SRM check is what stands between that team and a roadmap built on noise.

Useful? Pass it to whoever reads your A/B test results.

Want the full playbook?

Get 128 product management frameworks.

34 frameworks 21 rules 65 heuristics & principles 49 operators

From Stewart Butterfield, Ami Vora, Codebase Guidelines, and 46 more. Drop one .md into Claude, Cursor, or ChatGPT. Your AI cites practitioners, not guesses.

See the pack

Instant .md download · One-time purchase · No subscription

New experts every week

Know when the next expert lands.

Gavel adds new operators to the database every week, each one with cited frameworks you can check and a note on where they disagree with the others. You found this page by searching. Get the next one by email instead.

57 experts 76 cited frameworks

Latest: Sam Parr on Presell Calls

One email a week, only when new experts shipped. Unsubscribe with one click. We never sell or share email.

If a test came back better than you expected, paste the user counts per variant into Gavel. Ask whether the split holds up, and the answer cites Kohavi's recordings.

Related frameworks