A worked example, rendered from real sample data. Sign in to run the tool on your own input.
25.3 5.2 30
23.1 4.8 28═══ Verdict ═══
✗ Not significant: two-tailed p = 0.099423 ≥ α = 0.05
ℹ Welch's two-sample t-test, t(55.99) = 1.6754
ℹ Group 1 mean − Group 2 mean = 2.2000, 95% CI [-0.4304, 4.8304]
ℹ Effect size: Cohen's d = 0.439 (small)
═══ Groups ═══
Group 1: mean 25.3000, sd 5.2000, n 30
Group 2: mean 23.1000, sd 4.8000, n 28
Difference (Group 1 mean − Group 2 mean): 2.2000
Standard error of difference: 1.3131
Variance ratio (larger/smaller): 1.174
═══ Test ═══
Test performed: Welch's two-sample t-test
Tail: two-tailed
t statistic: 1.6754
Degrees of freedom: 55.99 (Welch-Satterthwaite, not n₁+n₂−2)
p-value: 0.099423
Two-tailed p: 0.099423
Critical t (α 0.05): ±2.0032
Decision: fail to reject H₀
H₀: μ₁ = μ₂
H₁: μ₁ ≠ μ₂
═══ Effect Size and Interval ═══
Cohen's d: 0.4390 (small)
Hedges' g (small-sample corrected): 0.4331
95% CI for the difference: -0.4304 to 4.8304
CI width: 5.2609
The interval contains 0, so no difference remains plausible.
d is the difference expressed in standard deviations, so it does not change when you collect more data.
═══ Assumption Checks ═══
F test for equal variances: F(29, 27) = 1.1736, p = 0.678604
ℹ Variances look similar; Welch is still safe and costs very little power
ℹ Summary input cannot be checked for normality or outliers — paste the raw values to get those checks
Observations must be independent within and between groups; repeated measures need the paired test.
═══ Plain English ═══
Group 1 averages 2.2000 higher than Group 2. A gap this size is within what sampling noise produces.
Practically: the true difference is plausibly anywhere from -0.
…
Calculate t-test results for comparing means. Part of the DevTools Surf developer suite. Browse more tools in the Statistics collection.