A test that is too small does not produce a wrong answer. It produces no answer — and that is a worse outcome, because a wrong answer is at least detectable.
The failure mode is familiar. Six weeks in, one version is ahead by a few percent, nobody is confident it means anything, and the test gets called on the strength of whoever argues hardest. Then the same variable gets re-tested a year later.
All of it is avoidable, because whether a test can produce an answer is calculable before it starts.
Your daily traffic to the tested page. Split across two variants.
Your current conversion rate. Lower rates need far more traffic, because each conversion carries more weight in a noisier signal.
The smallest lift worth detecting. Not the lift you hope for — the smallest one you would actually act on.
Roughly: at a 2% conversion rate, detecting a 20% relative lift within eight weeks needs on the order of 700 visitors a day to the tested page. Below about 470 a day, an onsite test of that size cannot be read at all in any reasonable timeframe.
Those figures move with your own conversion rate and the effect size you care about, but the shape holds: the bar is higher than most people assume, and small sites are usually below it.
If you have 200 visitors a day and you are testing a button color, you will never know. You will get a number, the number will drift, and you will make a decision from noise.
That is worse than not testing. It manufactures false confidence and consumes weeks that could have gone to something readable.
Test bigger changes. Detecting a 5% lift takes enormous traffic; detecting a 50% lift takes far less. If you cannot read small changes, stop making small changes — test a genuinely different page, a different offer, a different proposition.
Test higher up the funnel. More people see your homepage than your checkout. Traffic is largest at the top, so that is where readable tests live.
Use sequential rather than split testing, carefully. Run version A for four weeks and version B for the next four. Weaker — seasonality and traffic mix confound it — but at low volume it may be the only structure that produces a signal at all. Just do not pretend it is a controlled experiment.
Or accept that this is not your lever. If your traffic cannot support onsite testing, the honest answer is that conversion optimization is not currently available to you and your effort belongs in acquisition. That is a legitimate finding, not a failure.
Before running any test, work out the required traffic and duration. If your volume will not support it, say so and do something else.
This applies well beyond websites. A media test spread too thin across too many markets is the same error at larger scale — one that cost $200,000 in a single engagement before anyone worked out that the tests had never been large enough to register against the existing baseline.
The instinct to test is right. The discipline is in checking that the test can answer the question before you spend the money.
You can calculate before you start whether a test has enough traffic to produce an answer. Most don't, and they end in "we're not sure."

Branded and non-branded search are different businesses with different economics, and blending them hides the expensive one.

You can calculate before you start whether a test has enough traffic to produce an answer. Most don't, and they end in "we're not sure."

The variables with the largest effect on conversion are usually owned by finance or operations, and no one has priced the marketing consequence.