Cited from real sources 6 min read Updated August 2026

A framework by Ronny Kohavi

Ronny Kohavi's OEC: Deciding What an A/B Test Optimizes For

The Overall Evaluation Criterion is the metric an experiment is judged against, chosen before the experiment runs. Ronny Kohavi built the experimentation platforms at Microsoft and Amazon and led experimentation at Airbnb, and his warning is that most teams skip this step. They ship a test, watch revenue, and never notice that revenue went up because the product got worse.

The question nobody answers first

What are you optimizing for?

Answer "revenue" and the experiment will deliver it, by making the product worse in ways the metric cannot see.

Ronny Kohavi Lenny's Podcast Watch at 28:13

The framework

Every metric you can move, you can also game

Teams new to experimentation treat the success metric as an afterthought. Kohavi treats choosing it as the hard part, harder than the statistics.

the OEC or the overall evaluation criteria is something that I think many people that start to dabble in A/B testing miss
Kohavi, on the step teams skip Watch at 28:13

The obvious answer is the dangerous one. Optimize for revenue and you will get revenue, because there are always ways to extract more of it in the short term at the expense of the people paying you.

it's very easy to say we're going to optimize for money, revenue, but that's the wrong question
Kohavi, on the obvious metric Watch at 28:32

His example is search advertising. Put more ads on the results page and you will make more money, with no doubt and no delay. The question the revenue metric cannot answer is what that does to the person who came to search, and whether they come back next week. So the OEC is not one number. It is a number plus a guard.

there has to be some countervailing metric that tells you how do I improve revenue without hurting the user experience
Kohavi, on the shape of an OEC Watch at 28:45

How to apply it

How do you define an OEC for your test?

Six steps. The first three define the criterion, the last three keep you honest about the result.

  1. 1

    Write down the success metric before you build.

    Deciding after you see the data is how a losing test gets re-scored against whichever metric happened to move.

  2. 2

    Ask how you would cheat it.

    Name the change that would move the metric while making the product worse. More ads, more emails, a harder-to-find cancel button. If cheating is easy, the metric is incomplete.

  3. 3

    Add the countervailing metric that catches the cheat.

    Pair the thing you want to grow with the thing you refuse to damage. A win has to clear both, otherwise it is not a win.

  4. 4

    Expect most ideas to lose.

    Kohavi's framing is that you are deliberately trying things that will probably fail, in exchange for the occasional one that is a home run. A team that wins most of its tests is testing only safe changes.

  5. 5

    Distrust the enormous result.

    Twyman's law: a figure that looks remarkable is usually a mistake. A 30% lift from a button colour is a bug in the instrumentation until proven otherwise.

  6. 6

    Keep experimenting when things get hard.

    A crisis is the moment teams abandon measurement and go on instinct. Kohavi's counter is that your instincts were wrong most of the time in calm conditions, so a crisis is a strange moment to start trusting them.

things that are too extreme, your meter should go up and say hey I don't believe that
Kohavi, on Twyman's law Watch at 74:47

Boundary conditions

When does A/B testing fail you?

Works best when

  • You have enough traffic for a result to mean anything
  • The change is reversible and the decision is genuinely open
  • Someone owns the OEC and will not renegotiate it after the readout

Fails when

  • You are pre-product-market-fit, where the change you need is too big to A/B
  • The OEC is a single revenue number with nothing guarding it
  • A startling result gets shipped instead of investigated

The first boundary is the one founders should take seriously. Experimentation is a tool for improving something that already works, and early on the honest move is Sean Ellis's product-market fit test or Bob Moesta's switch interview, not a 50/50 split on a landing page nobody visits.

The failure rate is also worth internalising before you build a testing culture, because it is the number that makes the OEC matter. Kohavi's point about crisis decision-making is really a point about baseline humility:

if in peacetime you're wrong two-thirds to 80 of the time
Kohavi, on why a crisis is a bad time to trust instinct Watch at 49:17

Two-thirds to eighty percent of your ideas do not work when conditions are calm and you have data. That is the strongest argument for writing the OEC down in advance, and the strongest argument against deciding by seniority. It is the same discipline Annie Duke applies to kill criteria: commit to what would count as failure while you still have no stake in the answer.

The sources

Where Kohavi discusses this

Useful? Send it to whoever is about to ship a test with one success metric.

Want the full playbook?

Get 98 product management frameworks.

31 frameworks 12 rules 47 heuristics & principles 49 operators

From Chandra Janakiraman, Stewart Butterfield, Ami Vora, and 46 more. Drop one .md into Claude, Cursor, or ChatGPT. Your AI cites practitioners, not guesses.

See the pack

Instant .md download · One-time purchase · No subscription

Related frameworks