A worked example, rendered from real sample data. Sign in to run the tool on your own input.
Variant A: 23.5, 24.1, 22.8, 25.3, 23.9, 24.5
Variant B: 21.2, 22.1, 20.8, 21.9, 22.3, 21.5═══ Verdict ═══
✓ Significant by the recommended test (Welch's t-test): two-tailed p = 0.000338 < α = 0.05
ℹ Recommended because the data pass the normality screen, so the parametric test is the more powerful choice
ℹ Variant A mean 24.0167 vs Variant B mean 21.6333 (difference 2.3833)
ℹ Effect size: Cohen's d = 3.279 (large), rank-biserial 1.000
═══ Samples ═══
Variant A: n 6, mean 24.0167, sd 0.8542, median 24.0000
Variant B: n 6, mean 21.6333, sd 0.5715, median 21.7000
Difference of means: 2.3833
95% CI for the difference: 1.4297 to 3.3370
Comparison: independent samples
═══ Parametric Test ═══
Welch's t-test: t(8.73) = 5.6802, p = 0.000338
Student's pooled t-test: t(10) = 5.6802, p = 0.000204
Welch is the default because it does not assume the two groups share a variance; its df is fractional for that reason.
Cohen's d: 3.2794 (large)
═══ Rank-Based Test ═══
Mann-Whitney U test: U = 0.00
z (normal approximation, tie corrected): -2.8022
p-value: 0.005075
Rank-biserial correlation: 1.0000
This compares whole distributions by rank, so it is robust to outliers and does not need normality.
═══ Assumption Checks ═══
⚠ Normality of Variant A: not testable with n = 6; judge from the histogram and be cautious
⚠ Normality of Variant B: not testable with n = 6; judge from the histogram and be cautious
Equal variances: F(5, 5) = 2.2337, p = 0.398417
✓ No outliers beyond the 1.5×IQR fences
⚠ Smallest group has n = 6; tests have little power below about 15 per group
Independence cannot be tested from the numbers — it comes from how you collected them.
═══ Plain English ═══
Varia
…
Test statistical significance of two samples. Part of the DevTools Surf developer suite. Browse more tools in the Statistics collection.