A/B testing has a credibility problem because most reported wins are noise. A test stopped when it looked good, running two variants of one weak idea, on traffic too thin to detect anything, produces a number rather than a fact.
Done properly it is the cleanest evidence in marketing. Hypotheses come from research rather than opinion, sample size is calculated before launch, the test runs full weekly cycles, and results are read with confidence intervals attached.
How testing runs here
- Research-led hypotheses drawn from analytics, recordings and form data
- Prioritisation by expected impact against implementation cost
- Sample size and duration calculated before anything launches
- Implementation that avoids flicker and does not slow the page
- Analysis with significance and interval, including honest “no difference” outcomes
- A test log, so the same idea is not relitigated every six months
If your traffic will not support valid testing, I will say so at the start and recommend research-led sequential changes instead. That is not a lesser method for small sites; it is the correct one.
