Too small to A/B test

At a few hundred visitors a week, a split test is noise with a chart attached — the useful moves are different ones.

Here is a number worth sitting with: to tell a 3% signup rate apart from a 3.6% signup rate with any confidence, you need somewhere north of ten thousand visitors in each variant. Not ten thousand total — ten thousand each. Most of the sites reading this get ten thousand visitors in a good quarter, and they are being asked by every growth article ever written to split that traffic in half and run a test. The advice is not wrong in general. It is wrong at your size, and nobody says so.

The reason is not statistics being fussy. It is that small numbers wobble on their own. Say a page gets 200 visitors a week and converts at 3%, so six signups. Next week you change the headline and get nine. That is a 50% lift, it is the best week the page has ever had, and it is also exactly what you would expect to see now and then if you changed nothing at all — three extra people is three extra people. You cannot hear the change over the noise, because at this volume the noise is the same size as anything you are likely to do.

under a few hundred conversions, a split test is a coin you flip twice and call it a trend.

So the honest version of the rule of thumb: if a variant is not accumulating something like thirty to fifty conversions a week, a split test will not resolve inside any timeframe you care about. Run the arithmetic before you build the test, not after. Weekly conversions divided in half, times the number of weeks you are willing to wait — if that total is under a couple of hundred per side, you are not running an experiment, you are running a ritual. The good news is that knowing this early frees up the hours you were about to spend on tooling.

What replaces it is not "guess". It is three things, and the first is the biggest: test changes large enough to see with the naked eye. Button colours and headline synonyms need thousands of visitors to detect because their real effect is small. Changing the offer, the price, the audience you point at the page, or what the page is even about can move a rate by half or double it — and a change that big does show up in small numbers. Small traffic does not mean you cannot learn. It means you can only learn from big swings, which is a useful constraint: it stops you spending a month on a test of two words.

The second is sequential comparison with your eyes open. Change one thing, run it for a defined stretch, compare against the same length of time before, and look for a step rather than a wiggle. This is genuinely weaker than a split test, and it is worth naming why: anything else that changed in those weeks — a mention somewhere, a seasonal lull, a campaign you forgot was still running — lands entirely in your result. The discipline that makes it usable is the same one behind [reading a dip without panicking](/blog/reading-a-dip-without-panicking): write down what else was happening, and be suspicious of any result that arrives in the same week as something else you did.

The third is that at your size, five conversations beat five thousand impressions. Watch five people try the page. Email five who signed up and did not come back, and five who bought, and ask what nearly stopped them. This is not a smaller version of quantitative testing, it is a different instrument — it gives you reasons, not rates. You cannot conclude "version B converts 12% better" from five people. You can absolutely conclude "four of five could not tell what this costs", which is a better finding than most tests produce and takes an afternoon.

There is one place where testing does work at small scale, and it is worth knowing where the line sits. Push the test toward the metric with the highest base rate that is closest to the thing you changed. A list of ten thousand people testing two subject lines is measuring open rate, which runs around a third rather than around a thirtieth, so it resolves in one send. Two ad creatives can be told apart on click-through long before they can be told apart on purchases. That is a real result about the creative, and you should hold it loosely as a result about revenue — the variant that wins on clicks sometimes loses on customers. Still, it is a signal you can actually get, which is more than the signup-rate test will give you.

This changes what AI drafting is good for, too. Anything that writes copy — this product included — will happily hand you five versions of a landing page or six subject lines, and it is tempting to read that as five tests waiting to be run. It is not. Producing variants was never the bottleneck; distinguishing them is, and no amount of generated copy fixes that. What variants are actually good for at this size is escaping your own first instinct: read the five, notice the one that says something the others do not, ship that one on judgment. Then check every number in it before it goes out, because a fluent draft with an invented statistic is the one failure worth catching every time.

None of this is a reason to stop measuring. Keep the [single number you steer by](/blog/one-goal-one-number) and keep watching it — you just read it as a trend over months rather than as a verdict on Tuesday's change. The distinction that matters is between measuring, which works fine at any size, and testing, which needs volume you do not have yet.

So the rule is simple enough to say in one line: until the numbers are big, make decisions on judgment plus changes large enough to be visible, and say out loud that this is what you are doing. A decision made on judgment and labelled as judgment is fine. A decision made on judgment and dressed up as a test is worse than either, because now you believe it.

These notes come from building SiteOps

Get started