A worked example, rendered from real sample data. Sign in to run the tool on your own input.
450, 1000
520, 1050═══ Verdict ═══
✓ Variant A beats Control: 10.05% lift, two-tailed p = 0.040307 < α = 0.0500
ℹ Single comparison at α = 0.050
⚠ Current power for the observed effect is only 53% — this test is underpowered, so a null result is not evidence of no effect
✓ Traffic split looks consistent with an even allocation
═══ Groups ═══
Control: 450/1,000 = 45.000% (CI 41.94%–48.10%)
Variant A: 520/1,050 = 49.524% (CI 46.51%–52.54%)
Intervals are Wilson score intervals, which stay valid at low conversion counts.
═══ Comparisons ═══
─── Variant A vs Control ───
Absolute difference: 4.524% (CI 0.205% to 8.843%)
Relative lift: 10.05% (CI 0.40% to 20.63%)
z: 2.0506
p-value (two-tailed): 0.040307
Significant at α 0.0500: yes
═══ Practical Significance ═══
ℹ Variant A: plausible true lift runs from 0.40% to 20.63% — decide whether the low end is still worth the change
Statistical significance says the difference is probably not zero. Practical significance asks whether it is big enough to act on. They are not the same question.
═══ Sample Size and Power ═══
Variant A: 1,911 per group needed for 80% power at this effect; you have 1,000 (52%), current power 53%
Required n is computed for the effect you actually observed, which is itself noisy — treat it as a rough guide.
═══ Stopping Rules and Peeking ═══
⚠ Checking the dashboard repeatedly and stopping when p first drops below 0.05 inflates the false-positive rate well past 5% — often to 20-30%.
Decide the sample size in advance, run to it, then look once.
If you must monitor mid-flight, use a sequential method (alpha spending, group seque
…
Calculate A/B test results and statistical significance. Part of the DevTools Surf developer suite. Browse more tools in the Statistics collection.