Z-Test for a Proportion
Enter the sample proportion p̂, sample size n, hypothesized proportion p₀, significance level α, and the direction of the alternative hypothesis. The calculator computes the standard error √(p₀(1 − p₀)/n), the z statistic, the critical z value and the p-value.
One detail that distinguishes this from an interval
A survey finds 58% support and the question is whether the true figure differs from half. The sample proportion will never land exactly on the hypothesised value, so the real question is whether the gap is bigger than sampling variation alone would produce.
This calculator answers it. It computes the standard error under the null hypothesis, the z statistic, the critical value and the p-value, and states the conclusion.
The standard error here is built from the hypothesised proportion, not the observed one. That is the difference between a test and a confidence interval: a test reasons about what would happen if the null hypothesis were true, so its variability must be the variability under that assumption.
A confidence interval makes no such assumption and uses the observed proportion instead. The two standard errors differ, sometimes noticeably, which is why a test and an interval can occasionally disagree at the margins.
How to use this calculator
- Enter the sample proportion A decimal between 0 and 1 — 58% is entered as 0.58. It is the observed fraction, not a count.
- Enter the sample size A positive whole number. It enters through the square root, so quadrupling it halves the standard error.
- Enter the hypothesised proportion Strictly between 0 and 1 — the values at either end are rejected, since they would make the standard error zero.
- Set the level and the direction Both must be fixed before the data is examined. The direction changes the critical value and the p-value together.
The formula, and where it comes from
SE = √(p₀(1 − p₀)/n) z = (p̂ − p₀)/SE reject when |z| exceeds the critical value
The standard error measures how much a sample proportion of this size varies around the hypothesised value. The product inside the root is largest at one half and shrinks towards either extreme, so a proportion near 0 or 1 is estimated more precisely from the same sample size.
Dividing the observed gap by that standard error converts it into a count of standard errors, which is the z statistic. Being a count rather than a proportion, it is comparable to the standard normal distribution regardless of what was being measured.
The critical value is found by bisection on the normal distribution rather than looked up, so any significance level works. For a two-tailed test the level is halved first, because the rejection region is divided between the two tails.
The p-value comes from the same distribution and reports how extreme the result is rather than merely whether it crossed the threshold. The two always agree on the verdict — rejecting exactly when the p-value falls below the level — but the p-value carries more information.
What each input means
- p̂ Sample proportion — form field “Sample proportion p̂”
- The observed fraction, between 0 and 1. Convert a count by dividing it by the sample size first.
- n Sample size — form field “Sample size n”
- A positive whole number. It enters only through a square root, so precision improves slowly with more data.
- p₀ Hypothesised proportion — form field “Hypothesized proportion p₀”
- Strictly between 0 and 1, since it determines the standard error. The endpoints are rejected.
- α Significance level — form field “Significance level α”
- The tolerated probability of rejecting a true null hypothesis, fixed before the data is seen.
Worked examples
Every number below is produced by the same calculation engine the tool above runs. Nothing here is typed by hand, so the walkthrough cannot drift from what you get when you enter the same values yourself.
A two-tailed test that rejects
A sample of 200 showing 58% against a hypothesised half, tested in both directions.
Inputs Sample proportion p̂ = 0.58, Sample size n = 200, Hypothesized proportion p₀ = 0.5, Significance level α = 0.05, Alternative hypothesis = two
- Given p̂ = 0.58, n = 200, p₀ = 0.5, α = 0.05
- Standard error SE = √(p₀(1 − p₀)/n) = √(0.5·0.5/200) = 0.0353553
- Test statistic z = (p̂ − p₀)/SE = 2.26274
- Critical value Two-tailed critical z at α/2 = 0.025: ±1.95996
- p-value 0.0236515
- Conclusion p < α — reject H₀: p = 0.5 at α = 0.05.
Result z = 2.26274, p = 0.0236515, reject H₀
The standard error under the null is about 0.035, so an eight-point gap is more than two standard errors — beyond the two-tailed critical value, and the test rejects.
Note the standard error uses the hypothesised half rather than the observed 0.58. Using the observed value would give a slightly different figure, which is exactly where a test and a confidence interval part company.
A one-tailed test that does not reject
The same sample size showing 52%, tested only for an increase above half.
Inputs Sample proportion p̂ = 0.52, Sample size n = 200, Hypothesized proportion p₀ = 0.5, Significance level α = 0.05, Alternative hypothesis = right
- Given p̂ = 0.52, n = 200, p₀ = 0.5, α = 0.05
- Standard error SE = √(p₀(1 − p₀)/n) = √(0.5·0.5/200) = 0.0353553
- Test statistic z = (p̂ − p₀)/SE = 0.565685
- Critical value Right-tailed critical z at α = 0.05: 1.64485
- p-value 0.285804
- Conclusion p ≥ α — fail to reject H₀: p = 0.5 at α = 0.05.
Result z = 0.565685, p = 0.285804, fail to reject H₀
The gap is a quarter of the previous one, and the resulting z falls well short of even the one-tailed critical value. Two hundred observations cannot distinguish 52% from 50%.
The one-tailed critical value is lower than the two-tailed one at the same level, and the test still fails to reject. Choosing the more sensitive test does not rescue a result this close to the hypothesis.
A small sample
Fifty observations showing 40%, tested for a decrease below half. The sample is at the edge of what the approximation supports.
Inputs Sample proportion p̂ = 0.4, Sample size n = 50, Hypothesized proportion p₀ = 0.5, Significance level α = 0.05, Alternative hypothesis = left
- Given p̂ = 0.4, n = 50, p₀ = 0.5, α = 0.05
- Standard error SE = √(p₀(1 − p₀)/n) = √(0.5·0.5/50) = 0.0707107
- Test statistic z = (p̂ − p₀)/SE = -1.41421
- Critical value Left-tailed critical z at α = 0.05: −1.64485
- p-value 0.0786497
- Conclusion p ≥ α — fail to reject H₀: p = 0.5 at α = 0.05.
Result z = -1.41421, p = 0.0786497, fail to reject H₀
The standard error is much larger here at around 0.07, because fifty observations pin a proportion down far less precisely than two hundred do. The ten-point gap amounts to well under two standard errors.
This sample only just satisfies the usual rule that the expected counts on both sides should reach ten. Below that, the normal approximation underlying the whole test starts to break down and an exact binomial test is the right tool.
Reading the result
Two standard errors, two purposes
The test uses the hypothesised proportion because it reasons under the null hypothesis. A confidence interval uses the observed proportion because it assumes nothing. Neither is wrong; they answer different questions.
Precision improves with the square root
Quadrupling the sample halves the standard error, so detecting a small difference costs disproportionately many observations. That is why a two-point gap remains undetectable at two hundred.
Failing to reject is not confirming
It means the evidence is insufficient to rule the hypothesised value out, which a small sample will produce almost regardless of the truth. The conclusion is about the strength of the evidence.
When you would use this
Testing a survey result against a threshold
Whether observed support exceeds a majority, or a defect rate exceeds a tolerance, is this test with the threshold as the hypothesised value.
Checking a rate against a published figure
Comparing an observed rate against an established benchmark tests whether the difference is more than sampling variation, which is the question a raw comparison of percentages cannot answer.
Assumptions and limitations
What this calculator assumes
- Observations are independent and each falls into one of two categories.
- The sample is large enough for the normal approximation, conventionally requiring the expected counts on both sides to reach about ten.
- The standard error is computed from the hypothesised proportion rather than the observed one.
- The significance level and direction are chosen before the data is examined.
Where it stops being the right tool
- One proportion only: comparing two groups needs a two-sample test.
- The normal approximation is used throughout, so a small sample or an extreme hypothesised proportion is better served by an exact binomial test.
- No confidence interval is reported alongside the test.
- No continuity correction is applied, which would slightly widen the p-value for small samples.
Common mistakes
Entering the proportion as a percentage
Why it happens. Survey results are quoted as percentages, so 58 is the number at hand. It falls outside the valid range and is rejected, which is at least a visible failure rather than a silent one.
How to avoid it. Divide by 100 first. Both proportions must be decimals between 0 and 1.
Using the observed proportion in the standard error
Why it happens. It is the value in front of you, and it is the correct choice for a confidence interval — so the habit transfers from one procedure to the other.
How to avoid it. A test assumes the null hypothesis while computing the statistic, so the hypothesised proportion is the right one. The standard-error line shows which was used.
Applying the test to a small sample
Why it happens. The calculation runs on any inputs and returns a plausible p-value, with nothing to indicate the approximation has stopped being reliable.
How to avoid it. Check that the sample size times each of the two hypothesised proportions reaches about ten. Below that, an exact binomial test is the appropriate method.
Frequently asked questions
When can I use a z-test for a proportion?
When the sample is large enough for the sampling distribution of the proportion to be approximately normal — conventionally, when the sample size multiplied by the hypothesised proportion and by its complement both reach at least ten.
Why is the standard error computed from p₀ rather than p̂?
Because a hypothesis test reasons about what would happen if the null hypothesis were true, so the variability it uses must be the variability under that assumption. A confidence interval makes no such assumption and uses the observed proportion instead.
Two-tailed or one-tailed?
Two-tailed rejects a difference in either direction; one-tailed only in the direction specified. Use one-tailed only when there is a substantive reason, decided before seeing the data, to care about one direction alone — choosing afterwards invalidates the stated error rate.
What does the p-value mean here?
The probability of observing a sample proportion at least this far from the hypothesised value, if that value were the truth. It is not the probability that the hypothesis is correct, which is a different quantity entirely.