You design a test to send half your users to control and half to treatment. A random hash decides who goes where, so the counts that arrive should match the design. Kohavi's rule is blunt:
The gap does not have to look big. With a million users, 50.2 against 49.8 sounds like noise. Put the counts into the formula and it says a split that uneven should turn up about once in half a million experiments. Kohavi keeps a spreadsheet that does the sum. In a 2016 talk he read out an alert from that week: 821,588 users in control and 815,582 in treatment. The ratio was 50.2%, and the odds of that by chance were one in 500,000.
This is also one of the few places where a p-value, the chance of seeing a result if nothing real is going on, means what people think it means. The design fixes the split at 50/50. So a tiny p-value really is the chance of seeing that split by luck.
The split matters because the traffic you lose is usually not a random slice. Often it is the people who viewed zero pages, or the heavy clickers who got filtered out as bots. Kohavi, on the first mistake his team made:
A failed split makes results look either super good or super bad. The bad ones kill ideas that might have worked. The good ones get celebrated. So the check comes first. If it fails, the metrics are not evidence of anything.