Statistics

One-Sample t-Test

Enter your sample statistics (mean, standard deviation, size), the hypothesized population mean μ₀, the significance level α and pick a one- or two-tailed alternative. The calculator computes the t statistic, the standard error, the degrees of freedom, the critical values and the p-value, then concludes whether to reject H₀.

One-Sample t-Test

Test a sample mean against a hypothesized μ₀ — t, df, p-value, conclusion.

Try:
Answert = 1, df = 24, p = 0.327287, fail to reject H₀
  1. Givenx̄ = 52, s = 10, n = 25, μ₀ = 50, α = 0.05
  2. Standard errorSE = s/√n = 10/√25 = 2
  3. Test statistict = (x̄ − μ₀)/SE = 1
  4. Degrees of freedomdf = n − 1 = 24
  5. Critical valueTwo-tailed critical t at α/2 = 0.025: ±2.0639
  6. p-value0.327287
  7. Conclusionp ≥ α — fail to reject H₀: μ = 50 at α = 0.05.

Why t rather than z

A sample mean will almost never equal the value you are testing it against, and that difference on its own proves nothing — samples vary. The question is whether the gap is larger than sampling noise would comfortably produce, and answering it needs the size of the gap measured against the size of the noise.

That ratio is the t statistic. This calculator computes it from your sample statistics, together with the degrees of freedom, the critical value, the p-value and the resulting conclusion.

When the population's spread is known, a sample mean is compared against the normal distribution. It almost never is: the standard deviation is estimated from the same small sample, which adds a second source of uncertainty the normal curve does not account for.

The t distribution accounts for it by being wider in the tails, and by exactly how much depends on the sample size through the degrees of freedom. That is why a small sample needs a larger t before anything is declared significant, and why the two distributions converge as the sample grows.

How to use this calculator

  1. Enter the sample statistics The mean, the standard deviation and the size. The standard deviation must be positive and the size a whole number of at least 2.
  2. Enter the hypothesised mean The value being tested against, which comes from the question rather than from the data. It is what the null hypothesis asserts.
  3. Set the significance level Strictly between 0 and 1. It fixes the threshold before the data is seen, which is what makes the conclusion meaningful.
  4. Choose the alternative Two-tailed unless there is a reason decided in advance to care about only one direction. That choice changes both the critical value and the p-value.

The formula, and where it comes from

SE = s/√n t = (x̄ − μ₀)/SE df = n − 1 reject when |t| exceeds the critical value

The standard error is how much a mean of this sample size varies from one sample to the next. Dividing the observed gap by it converts a difference in the data's own units into a count of standard errors, which is comparable across any measurement.

The degrees of freedom are one less than the sample size because the standard deviation was computed using the sample mean, which consumes one piece of information. That is the same reason the sample standard deviation divides by one less than the count.

The critical value is found by bisection on the t distribution rather than looked up in a table, so any significance level works rather than only the tabulated ones. For a two-tailed test the level is halved first, since the rejection region is split between the two tails.

The p-value comes from the same distribution and answers a different question from the critical value: not whether the threshold was crossed, but by how much. It is the probability of seeing a t at least this extreme if the null hypothesis were true, computed through a regularised incomplete beta function that stays accurate across all degrees of freedom.

What each input means

x̄, s, n Sample statistics — form field “Sample mean x̄”
The mean, standard deviation and size of the observed sample. The standard deviation is the spread of the observations, not the standard error.
μ₀ Hypothesised mean — form field “Hypothesized mean μ₀”
The value the null hypothesis asserts. It must come from the question, never from the data being tested.
α Significance level — form field “Significance level α”
The probability of rejecting a true null hypothesis that you are willing to accept. Chosen before the data is examined.
tail Alternative hypothesis — form field “Alternative hypothesis”
Two-tailed or one-tailed in a specified direction. A one-tailed test is more sensitive on its side and blind on the other.

Worked examples

Every number below is produced by the same calculation engine the tool above runs. Nothing here is typed by hand, so the walkthrough cannot drift from what you get when you enter the same values yourself.

A two-tailed test that does not reject

A sample of 25 with a mean two above the hypothesised value, tested in both directions.

Inputs Sample mean x̄ = 52, Sample standard deviation s = 10, Sample size n = 25, Hypothesized mean μ₀ = 50, Significance level α = 0.05, Alternative hypothesis = two

  1. Given x̄ = 52, s = 10, n = 25, μ₀ = 50, α = 0.05
  2. Standard error SE = s/√n = 10/√25 = 2
  3. Test statistic t = (x̄ − μ₀)/SE = 1
  4. Degrees of freedom df = n − 1 = 24
  5. Critical value Two-tailed critical t at α/2 = 0.025: ±2.0639
  6. p-value 0.327287
  7. Conclusion p ≥ α — fail to reject H₀: μ = 50 at α = 0.05.

Result t = 1, df = 24, p = 0.327287, fail to reject H₀

The standard error is 2, so the observed gap of 2 is exactly one standard error — a t of 1. That is well inside the critical value for 24 degrees of freedom, and the p-value comfortably exceeds the level.

Failing to reject is not evidence that the means are equal. It says the data is consistent with the hypothesis, which is a much weaker claim and the one the conclusion line is careful to state.

A one-tailed test that rejects

A sample of 30 with a mean five above the hypothesised value, tested only for an increase.

Inputs Sample mean x̄ = 105, Sample standard deviation s = 15, Sample size n = 30, Hypothesized mean μ₀ = 100, Significance level α = 0.05, Alternative hypothesis = right

  1. Given x̄ = 105, s = 15, n = 30, μ₀ = 100, α = 0.05
  2. Standard error SE = s/√n = 15/√30 = 2.73861
  3. Test statistic t = (x̄ − μ₀)/SE = 1.82574
  4. Degrees of freedom df = n − 1 = 29
  5. Critical value Right-tailed critical t at α = 0.05: 1.69913
  6. p-value 0.0391017
  7. Conclusion p < α — reject the null hypothesis H₀: μ = 100 at α = 0.05.

Result t = 1.82574, df = 29, p = 0.0391017, reject H₀

The whole rejection region sits in one tail, so the critical value is smaller than a two-tailed test at the same level would use. That extra sensitivity is bought by giving up any ability to detect a decrease.

The gap here is larger relative to the standard error than in the previous example, and the test rejects. Both the larger difference and the larger sample contribute — the standard error falls as the sample grows.

Reading the result

Failing to reject is not accepting

It means the evidence was insufficient to rule the hypothesis out, which is compatible with the hypothesis being false and the sample too small to show it. The conclusion is about the evidence, not about the truth.

What the p-value actually is

The probability of a result at least this extreme if the null hypothesis were true. It is not the probability that the hypothesis is true, and not the probability the result was chance — both readings are common and both are wrong.

Significance is not size

A large enough sample makes any difference significant, however small. The t statistic combines the gap with the precision, so reading the gap itself alongside the verdict is what distinguishes a real effect from a detectable one.

When you would use this

Testing a measured average against a target

Where a process is meant to hit a specification, comparing a sample of output against that value is exactly this test, and the significance level sets how much evidence is demanded before intervening.

Checking a before-and-after difference

Paired measurements reduce to a single sample of differences tested against zero, which is this same procedure applied to the differences rather than the raw values.

Assumptions and limitations

What this calculator assumes

  • The observations are independent and drawn from an approximately normal population.
  • The standard deviation is estimated from the sample rather than known in advance.
  • The significance level and the direction of the alternative are fixed before the data is examined.
  • The sample size is a whole number of at least 2, since the degrees of freedom are one less.

Where it stops being the right tool

  • One sample only: comparing two groups needs a two-sample test.
  • Summary statistics are entered rather than raw data, so the normality assumption cannot be checked here.
  • No confidence interval is reported alongside the test.
  • No effect size is computed, so significance and practical importance must be judged separately.

Common mistakes

Choosing the tail after seeing the data

Why it happens. A one-tailed test in the direction the sample happened to go gives a smaller p-value, and the choice feels like a presentational one.

How to avoid it. Fix the direction from the question before looking. Choosing afterwards doubles the real error rate while reporting the smaller one.

Entering the standard error instead of the standard deviation

Why it happens. Software often reports both, and both describe variability. Feeding the standard error in means dividing by the root of the sample size twice.

How to avoid it. Enter the spread of the observations. Check the standard-error line: it should be noticeably smaller than what you typed.

Reading failure to reject as proof of no difference

Why it happens. The two outcomes sound symmetric, so not rejecting reads as confirming.

How to avoid it. State it as insufficient evidence. A small sample fails to reject almost regardless of the truth, which is why the conclusion is worded that way.

Frequently asked questions

When do I use a t-test instead of a z-test?

When the population standard deviation is unknown and estimated from the sample, which is nearly always. The t distribution has heavier tails to account for that extra uncertainty, and the two converge as the sample size grows past about thirty.

What does two-tailed mean?

That the test rejects when the sample mean is significantly larger or smaller than the hypothesised value. A one-tailed test looks in only one direction, which makes it more sensitive there at the cost of being unable to detect a difference the other way.

How is the p-value computed?

From the Student t distribution with degrees of freedom one less than the sample size, evaluated through a regularised incomplete beta function so that it stays accurate at any degrees of freedom rather than only at tabulated values.

Why is the degrees of freedom one less than the sample size?

Because the standard deviation was computed using the sample mean, which uses up one piece of independent information. The same reasoning is why the sample standard deviation divides by one less than the count.