Direct answer
The short version.
A credible A/B test defines the hypothesis, primary outcome, assignment method, required sample or duration, guardrails and decision rule before exposure starts. Post-hoc definitions make ordinary noise look persuasive.
Key takeaways
Keep these three decisions.
- Pre-register the decision rule.
- Use one primary outcome.
- Treat uncertainty as part of the result.
The operating problem
Teams stop tests when a preferred variant moves ahead, inspect many metrics until one is positive or run simultaneous changes that contaminate the comparison. Low volume and conversion lag make the interface's early confidence especially easy to misread.
A practical system
Write a short protocol with control, treatment, unit of randomisation, eligible audience, primary metric, minimum detectable effect and analysis window. Check tracking before launch, avoid overlapping experiments that share the same population and record deviations rather than silently changing the plan.
What to measure next
Read the estimated effect with uncertainty and operational consequences, not a winner badge alone. Review sample ratio, data loss, novelty, guardrails and downstream quality. An inconclusive test is useful when it rules out a large effect and sharpens the next question.
Primary and authoritative references
Sources used for context.
Frequently asked questions
Two useful follow-ups.
How long should an A/B test run?
Long enough to reach the planned sample and cover meaningful business cycles and conversion lag. Duration cannot be chosen from a universal number of days.
Can multiple metrics be used?
Use one primary decision metric and a small set of guardrails. Additional metrics can diagnose the mechanism without redefining success.