Review the design before the result
Confirm that the comparison units were assigned appropriately and that users did not switch between versions in a way that undermines the design. Inspect missing data, unusual allocation and inconsistent event firing.
Define the primary outcome before launch. Selecting whichever metric improved afterward increases the risk of a misleading story, especially when many outcomes were inspected.
Read the evidence carefully
- Check that the test ran under the intended conditions.
- Inspect counts as well as percentages.
- Use an appropriate statistical method and its assumptions.
- Allow business outcomes and guardrails to mature.
- Decide whether the evidence supports rollout, revision or further observation.
Repeatedly checking a conventional test and stopping at a favorable result can distort inference. Use a valid analysis plan; seek qualified statistical help for consequential or complex designs.
Worked example: small counts
Illustrative example: version A produces 10 purchases from 100 eligible visitors, while version B produces 12 from 100. The observed rates are 10% and 12%. That arithmetic does not by itself establish a reliable improvement or justify claiming a proven 20% lift.
The team needs uncertainty, design quality and business context. A two-purchase difference may be useful as a question to investigate without being a rollout verdict.
Interpretation checklist
| Check | Why it matters |
|---|---|
| Allocation | Supports a fair comparison. |
| Event reliability | Prevents measurement changes from appearing as behavior changes. |
| Primary metric | Keeps the conclusion tied to the original question. |
| Guardrails | Protects margin, customer quality and experience. |
Explain the scope
A result applies to the tested population, conditions and period. A winning checkout explanation for one market may not transfer unchanged to another language or product. Describe the mechanism and limits rather than declaring a universal rule.
What if traffic is too low?
Use methods suitable to the question, such as interviews or usability work, and label their evidence appropriately. Do not pretend a tiny test has stronger certainty than it does.
Is an inconclusive test a failure?
It can reveal uncertainty, a weak effect or a design issue. Record what remains unknown and whether another test is worth the cost.
Put this into practice
Before launching a test, write the primary outcome, allocation unit, guardrails and analysis plan. Afterward, report the observed difference and uncertainty without stretching beyond the design.
Primary-source reading for platform details: NIST: hypothesis-testing concepts.
Related foundation: CAC vs ROAS: which metric should guide your marketing budget?. How these guides are prepared.
Related portfolio work: Etsy & Shopify. The worked examples in this guide are illustrative and are separate from the portfolio’s project evidence.
