Cited from real sources 7 min read Written by Arun Agrahri Updated September 2026

A framework by Ronny Kohavi

Ronny Kohavi's Minimum Detectable Effect Rule: Size the Test Before You Trust It

The minimum detectable effect, or MDE, is the smallest change an A/B test is sized to catch. Statistical power is the chance the test catches a change that size when it is real. Ronny Kohavi ran experimentation at Amazon, Microsoft and Airbnb, and he sets both before launch: an MDE of 5% or less, and 80% power. A test that skips this step misses most real effects, and the wins it does report are exaggerated.

When an A/B test starts to work

About 200,000 users.

Kohavi's number for a site converting at 5% that wants to catch a 5% lift, at the standard error rates. The formula asks for 121,000 users per variant.

Ronny Kohavi A/B Testing Myths (2025) Watch at 24:17

The framework

Decide the smallest lift worth catching, then count the users

Power is set before a test runs. Kohavi's definition:

the statistical power of your test is the probability that if there is a difference say B is better by 5% that you will detect it
Kohavi on what power means Watch at 10:30

The power formula takes three inputs: the variance of your metric, which last week's data gives you; the MDE; and the error rates you will accept. Out comes the number of users you need. His worked example is a store that converts 5% of visitors and wants to detect a 5% relative lift, a quarter of a percentage point. At the industry defaults of a 5% false positive rate and 80% power, the formula asks for 121,000 users per variant.

Most people expect far fewer users. Kohavi explains why:

the formula is such that smaller changes are hard to detect and in fact it grows quadratically so if you want to be able to detect something half as big you need four times as many users
Kohavi on why small effects are expensive Watch at 18:44

It works in reverse too. A vaccine trial looks for a 50% effect, so the same formula asks for 1,216 users per variant. That is why those trials run on a few thousand people.

Low power also distorts the wins you do see. At 20% power, a real effect comes back significant only two times in ten, and those two overstate it:

It is true that most of the time you will not see a statistically significant effect. When you do it will be exaggerated.
Kohavi on the winner's curse Watch at 14:01

He calls this the winner's curse. At 10% power, a significant result is exaggerated four times on average, so a reported 8% gain is likely really 2%. Near 5% power the factor is about 25. One test shared on Guess the Test claimed a 337% lift from 80 users per variant. Kohavi ran the power calculation: about 5% power, where the exaggeration is undefined.

How to apply it

How do you choose a minimum detectable effect?

The first four steps happen before launch. The last two are for when the traffic is not there.

  1. 1

    Set the MDE at 5% or below.

    Kohavi treats 5% as the upper bound for online tests. With enough users, go for 2% or 1%. An MDE of 20% sizes the test for a lift that almost never happens, and what comes back is noise with a large exaggeration factor.

  2. 2

    Check what similar changes have moved.

    Search repositories of past tests, such as GoodUI and Guess the Test, for your change. They show the effect sizes similar changes produced.

  3. 3

    Keep the threshold at 0.05 and power at 80%.

    The p-value threshold is the cutoff for calling a result significant. Loosening it looks like a cheap way to get results. By Kohavi's numbers, a 0.05 threshold already leaves around 22% of significant results false. At 0.10, about a third are.

  4. 4

    Run to the user count, then read it once.

    A test does not reach power partway through. Run the formula before you start, run until you hit that count, and do not peek unless the platform uses sequential testing built for it. Power computed after the fact is not valid. Extra users are fine, because they shrink the confidence intervals.

  5. 5

    Short on traffic? Test bigger changes.

    Double the MDE, from 1% to 2%, and you need a quarter of the users. Or measure a metric closer to the change, such as clicks into the next funnel step instead of checkout, and then check that the rest of the funnel holds.

  6. 6

    Lower the metric's variance.

    Use conversion instead of revenue when the change should not move order size. Cap outliers, so anyone who spends over $100 counts as $100. At Airbnb, moving from number of bookings to booked or not helped. CUPED, which adjusts the metric using pre-experiment data, always lowers variance; at Bing it helped by 50%.

Asked about teams whose results mostly come back inconclusive, Kohavi started with power:

so if you're telling me that most of your experiments come back inconclusive it may be that you don't have enough statistical power
Kohavi at a Statsig AMA Watch at 15:08

A flat result from a well-powered test still teaches you something. If the test could detect a 3% change and found none, the idea moves the metric by less than 3% either way. At Bing, with enough traffic to detect a one-millisecond slowdown, a third of experiments still came back flat.

Boundary conditions

When does a power calculation mislead you?

Works best when

  • The team agrees on the MDE before anyone builds the variant.
  • The metric's variance comes from recent data, not a guess.
  • Users, not page views, are the unit being randomized.
  • The test runs its planned length and is read once.

Fails when

  • The MDE is picked to fit the traffic you have.
  • A flat result is read as proof the change does no harm.
  • A flat test is sliced into segments until one lights up.
  • Power is computed after the result is in.

The second failure is the common one. Teams ship a feature because the test came back flat, not worse. Kohavi's objection:

You can never accept the null that there is no treatment effect. It could very well be that the test is what's called underpowered.
Kohavi on shipping flat results Watch at 16:50

His example comes from The Cartoon Guide to Statistics: a polluter who must show the discharge is not significantly worse, so it measures a few ducks. Segments fail another way. Each one has fewer users, and slice enough of them and one will show a large effect by chance. Kohavi's rule for a segment result: use it to design the next test, never to launch.

Below 200,000 users, Kohavi does not say stop. Start at tens of thousands of users and look only for large effects. Copy what Amazon, which tests everything, already ships. On Lenny's Podcast:

so you ask for rule of thumb 200,000 users you're magical below that start building the culture start building the platform
Kohavi on testing before the traffic arrives Watch at 27:47

A power calculation also assumes the test itself is sound. Run the sample ratio mismatch check before you read any result. And fix the Overall Evaluation Criterion first, because the MDE is a change to that metric.

Sources

Where Kohavi discusses this

We pulled nine long Kohavi recordings for this page on 2026-09-30. Eight of them raise statistical power, from a 2016 conference talk to a March 2026 lesson.

Where experts disagree

Where operators disagree: 200,000 users, or five?

Ronny Kohavi

says size the test before it runs: an MDE of 5% or less at 80% power, which for a store converting at 5% means about 200,000 users. Below that a significant result is likely exaggerated, so build the culture and the platform, and test only for large effects.

Jake Knapp

tests a prototype with five target customers on the Friday of a design sprint, not with internal stakeholders. In Sprint, five customers reveal the major patterns without statistical overhead, so there is no reason to wait for a large sample.

They answer different questions. Five sessions show whether people understand and want the thing, because confusion repeats as a pattern. They cannot show whether a shipped change moved conversion by 3%, which is the question Kohavi's power formula answers. Use five customers to decide what to build, and the power formula before you claim a change moved a metric.

Useful? Pass it to whoever sizes your A/B tests.

Want the full playbook?

Get 128 product management frameworks.

34 frameworks 21 rules 65 heuristics & principles 49 operators

From Stewart Butterfield, Ami Vora, Codebase Guidelines, and 46 more. Drop one .md into Claude, Cursor, or ChatGPT. Your AI cites practitioners, not guesses.

See the pack

Instant .md download · One-time purchase · No subscription

New experts every week

Know when the next expert lands.

Gavel adds new operators to the database every week, each one with cited frameworks you can check and a note on where they disagree with the others. You found this page by searching. Get the next one by email instead.

57 experts 76 cited frameworks

Latest: Sam Parr on Presell Calls

One email a week, only when new experts shipped. Unsubscribe with one click. We never sell or share email.

Before your next test, give Gavel your conversion rate and weekly traffic. Ask what MDE you can afford, and the answer cites Kohavi's recordings.

Related frameworks