Skip to content
Popular Calculators
Browse Statistics calculators

T-Test Calculator

Run a one-sample, two-sample (Welch) or paired t-test, with the t statistic, degrees of freedom, p-value and a confidence interval.

What do you want to work out?

About the T-Test Calculator

The t-test is the most widely used test for comparing means. It answers questions such as: is the average weight of these packets different from the 500 g on the label? Did students taught with a new method score higher than those taught the old way? Did patients' blood pressure fall after treatment? In each case the t-test weighs the size of a difference against the amount of random variation in the data, and gives a p-value showing how surprising the difference would be if there were no real effect.

This t-test calculator runs all three common versions. The one-sample t-test compares a sample mean with a fixed value. The two-sample t-test compares the means of two independent groups, using Welch's method, which does not assume equal variances. The paired t-test compares two measurements on the same subjects. Each gives the t statistic, degrees of freedom, two-tailed and one-tailed p-values, and a 95 percent confidence interval for the difference.

How to Use the T-Test Calculator

Choose One sample, Two samples or Paired.

For one sample, paste your values and enter the hypothesised mean.

For two samples, enter each group's mean, standard deviation and size.

For paired data, paste the first and second measurements in the same order, so that each pair lines up.

The Formulas

  one sample:   t = (x̄ − μ₀) ÷ (s ÷ √n)                    df = n − 1
  two samples:  t = (x̄₁ − x̄₂) ÷ √(s₁²/n₁ + s₂²/n₂)
                df = (s₁²/n₁ + s₂²/n₂)² ÷ [ (s₁²/n₁)²/(n₁−1) + (s₂²/n₂)²/(n₂−1) ]
  paired:       one-sample test on the differences d        df = n − 1

Step-by-Step Example: One Sample

Is the mean of 2, 4, 4, 4, 5, 5, 7, 9 different from 3?

  x̄ = 5,  s = 2.1381,  n = 8
  SE = 2.1381 ÷ √8 = 0.7559
  t = (5 − 3) ÷ 0.7559 = 2.646,  df = 7
  p (two-tailed) = 0.033
  95% CI for the difference: 0.21 to 3.79

The mean is significantly different from 3 at the 5 percent level.

Step-by-Step Example: Two Samples

Group 1: mean 75, SD 10, n 30. Group 2: mean 70, SD 12, n 30.

  SE = √(100/30 + 144/30) = √8.133 = 2.852
  t = (75 − 70) ÷ 2.852 = 1.753
  Welch df ≈ 56.2
  p (two-tailed) = 0.085
  95% CI for the difference: −0.71 to 10.71

The 5-point difference is not significant at 5 percent; the interval includes zero.

Step-by-Step Example: Paired

Six people measured before (72, 75, 80, 68, 90, 77) and after (70, 74, 76, 65, 86, 75).

  Differences: 2, 1, 4, 3, 4, 2
  mean d = 2.667,  s_d = 1.211,  SE = 1.211 ÷ √6 = 0.4944
  t = 2.667 ÷ 0.4944 = 5.39,  df = 5
  p (two-tailed) = 0.003

Every person improved, and the paired test finds the change highly significant.

Why Pairing Matters

Had the before-and-after data above been analysed as two independent groups, the large differences between people — from 65 to 90 — would swamp the small, consistent changes: a two-sample test gives t = 0.63 and p = 0.54, finding nothing. Pairing removes the between-person variation by looking only at each person's change. Whenever data comes in natural pairs — the same subject twice, twins, matched patients, left and right eyes — use the paired test.

Welch or Student?

The original two-sample t-test, Student's, assumes both groups have the same variance and pools them. When variances and group sizes differ, it can give misleading p-values. Welch's version drops the assumption and adjusts the degrees of freedom instead. It loses almost nothing when the variances are equal and protects you when they are not, which is why many statisticians and software packages now use it by default.

Assumptions and Robustness

The t-test assumes independent observations and roughly normal data — or, for two samples, roughly normal data within each group. Thanks to the central limit theorem, it is fairly robust with moderate samples of 30 or more per group unless the data is very skewed. With small samples, check for outliers and strong skew; if they are present, consider a non-parametric alternative such as the Wilcoxon or Mann–Whitney test.

Effect Size: Cohen's d

A p-value says whether a difference is likely to be real; it does not say whether it is large. Cohen's d expresses the difference in standard deviations. For the two-sample example, the pooled standard deviation is √((10² + 12²) ÷ 2) = 11.05, so d = 5 ÷ 11.05 = 0.45. By Cohen's rough guide, 0.2 is a small effect, 0.5 medium and 0.8 large, so this is a small-to-medium difference. It failed to reach significance here not because it is trivial but because 30 per group is too few to detect an effect of that size reliably; about 80 per group would give a good chance of doing so. Always report the estimated difference and its confidence interval alongside the p-value, so readers can judge the size of the effect as well as its significance.

Understanding Your Result

The headline gives the t statistic and the two-tailed p-value.

The degrees of freedom line gives df, fractional for Welch's test.

The decision line says whether the result is significant at 5 percent.

The one-tailed line gives the one-sided p-value, for use only when the direction was predicted in advance.

The difference and 95% CI line gives the estimated difference and its interval.

When Should You Use This Calculator?

Use it to compare a sample average with a target or claim.

Use it to compare two groups, such as treatment and control.

Use it for before-and-after or matched-pair studies.

Use it for statistics coursework on hypothesis testing.

Common Mistakes

Using an independent test on paired data. It throws away the pairing.

Using a paired test on unrelated groups. The lists must match item by item.

Running many t-tests on several groups. Use ANOVA instead.

Choosing a one-tailed test after seeing the result. Decide in advance.

Ignoring the size of the difference. Report the confidence interval, not only p.

Frequently Asked Questions

How do I test whether 2, 4, 4, 4, 5, 5, 7, 9 has a mean of 3?

The sample mean is 5 and the standard error is 2.1381 ÷ √8 = 0.7559, so t = (5 − 3) ÷ 0.7559 = 2.65 with 7 degrees of freedom. The two-tailed p-value is about 0.033, significant at the 5 percent level.

How does a two-sample t-test work?

It divides the difference between the means by its standard error. For means of 75 and 70, SDs of 10 and 12, and 30 in each group, the standard error is 2.852, t = 1.75 and p is about 0.085, not significant at 5 percent.

What is Welch's t-test?

A two-sample t-test that does not assume equal variances. It uses each group's own variance and adjusts the degrees of freedom with the Welch–Satterthwaite formula, giving about 56.2 in the example. Many statisticians recommend it as the default.

When should I use a paired t-test?

When each value in one list is matched to a value in the other, such as the same people before and after a treatment. The test works on the differences. For the example the mean change is 2.67, t = 5.39 and p is about 0.003.

What assumptions does the t-test make?

The data should be roughly normally distributed, or the samples large enough for the central limit theorem to help, and observations should be independent. For paired tests, the differences should be roughly normal. Strong skew or outliers in small samples can mislead.

What is the difference between a t-test and a z-test?

A z-test assumes the population standard deviation is known; a t-test estimates it from the sample and uses the t distribution, which has heavier tails. With large samples the two give nearly the same answer, but for small samples the t-test is correct.

Last reviewed September 28, 2026 by the CalculatorPeak editorial team.