P-Value Calculator
Result
P-value
A p-value calculator converts a test statistic into the probability of seeing a value at least that extreme if the null hypothesis were true. It handles the four statistics that carry most of applied statistics — z, t, chi-square and F — and needs only the statistic, its degrees of freedom and a choice of tails. The tails setting is the one that is got wrong most often: a two-tailed test counts both directions, a one-tailed test counts only the one that was predicted. The output is a single number, in per cent. What the page deliberately does not do is judge, because the significance level is yours to bring.
Formula
upper tail: p = P(T ≥ t) lower tail: p = P(T ≤ t) two-tailed: p = 2 · min( upper, lower ) — symmetric distributions only
- t
- The test statistic you already computed — a z from a large sample, a t from a small one, a chi-square from counted categories, or an F from a ratio of two variances. The page does not compute it; it converts it
- T
- The distribution the statistic would follow if the null hypothesis were true. Which one it is comes from the four-way selector, and the choice is not a preference: using the normal distribution where a t belongs makes every p-value too small
- df
- The degrees of freedom — a property of how the statistic was computed rather than of the data, and required by t, chi-square and F. F needs two: the numerator degrees of freedom, which is the number of groups or terms being compared, and the denominator one, which comes from the residual sample size
- α
- The significance level, the threshold the p-value will be compared against. It is not an input on this page because it is not part of the calculation: it belongs to the decision you are about to make, and printing a verdict against a default of 0.05 would be choosing that threshold on your behalf
- p
- The tail probability, reported as a percentage rather than as a decimal. A p-value of 0.05 is shown as 5.0000, so the figure on screen can be compared with a quoted significance level directly, without moving a decimal point first
Use it once you already have a statistic and want the probability attached to it — from a textbook exercise where the statistic is given, from software that printed the statistic but not the p-value, or from a hand calculation you want to check. The choice that matters most is the tails setting, and the rule is fixed before the data is seen rather than after: pick two-tailed when either direction would count as a departure from the null, which is the usual case, and one-tailed only when a result in the opposite direction would be of no interest at all and you said so in advance. The second choice is the distribution. Match it to how the statistic was built rather than to which p-value you are hoping for: a z belongs to a proportion or a known-variance mean, a t to a mean whose spread came from the same small sample, a chi-square to counted categories, and an F to a comparison of variances. Getting that pairing wrong is the most common way for a correct arithmetic to produce a meaningless number.
Worked examples
A z of 1.96, the classic two-tailed threshold
- Upper tail: P(Z ≥ 1.96) = 0.025002, computed from the standard normal distribution
- The normal curve is symmetric, so the lower tail is the same: P(Z ≤ −1.96) = 0.025002
- Two-tailed: 2 × 0.025002 = 0.050004
- As a percentage: 5.0000%, which is the 5% everyone quotes
1.96 is the value chosen precisely so that this comes out at 5%, and seeing 4.9996 rather than a flat 5.0000 is not an error — the true quantile is 1.959964, so 1.96 is itself a rounding. The three degrees-of-freedom fields are all filled in and only the statistic and the tails setting are read, because z needs neither degrees of freedom: the two unused fields are always on screen, and a blank or unused value in them changes nothing.
A t of 2.2281 with 10 degrees of freedom
- The t distribution with 10 degrees of freedom has heavier tails than the normal curve
- Upper tail: P(T ≥ 2.2281) = 0.0250015
- Two-tailed: 2 × 0.0250015 = 0.050003, or 5.0003%
Compare this with the z example and the point of the t distribution is visible: the same tail area of 5% sits at 2.2281 rather than at 1.96, because a small sample has to reach further out before the evidence is as unusual. That gap is exactly what the sample size controls — at 1000 degrees of freedom the t critical value is 1.962 and the two distributions are practically the same curve.
A chi-square of 20 on three degrees of freedom, upper tail only
- The chi-square distribution has no negative half, so only the upper tail is available here
- P(χ² ≥ 20) with 3 degrees of freedom = 0.00016975
- As a percentage: 0.017%
This is the statistic that the chi-square page produces from the counts 30, 10, 20 and 40 against an even expectation. The p-value came out small — 0.017% — so the four categories are not equally likely. Note what is missing and has to be asked for separately: this page says how unusual the statistic is, and says nothing about the smallest expected count, which is the check that decides whether the chi-square approximation was appropriate in the first place.
Limitations
A p-value answers one narrow question and is routinely read as answering several others it cannot. It is not the probability that the null hypothesis is true, not the probability that the result happened by chance, and not a measure of how large an effect is — a trivial difference becomes overwhelmingly significant if the sample is large enough, and an important one can sit at p = 0.20 in a sample of twelve. It is also only as good as the model behind the statistic: the four distributions here assume independent observations, and a statistic computed from clustered or repeated measurements carries fewer independent observations than its formula believes, which makes every p-value reported for it too small. Two guards are enforced rather than left to the reader. A negative statistic is refused for chi-square and F, because those distributions live on the positive half-line and a negative value means the statistic was miscomputed. And a two-tailed chi-square or F is refused, because those distributions are asymmetric and there is no single accepted two-sided convention for them.
Frequently asked questions
- What is a p-value?
- The probability of getting a test statistic at least as extreme as the one you have, assuming the null hypothesis is true and the model behind the statistic is right. It describes the data, not the hypothesis: a small p-value says the observed result would be unusual if nothing were going on, which is not the same as saying something is going on. The step from there to a conclusion runs through a significance level you choose in advance, plus everything you know about the subject that is not in the numbers.
- Should I use one tail or two?
- Two, unless you decided otherwise before seeing the data and can say which direction you predicted. A two-tailed test counts a result in either direction as evidence against the null, which matches how most questions are actually posed: a drug that makes things worse is a finding too. One-tailed testing is legitimate but it halves the p-value, and choosing the tail after looking at the data is the single most common way to manufacture a significant result out of nothing. If you are not sure, use two-tailed — it is the conservative choice and the one a reviewer will expect.
- What significance level should the p-value be compared against?
- Whichever one you fixed before collecting the data — by convention 0.05, but the convention is a habit rather than a result. It is deliberately not an input here, and no verdict is printed, because the same p-value of 0.04 is a discovery in one setting and noise in another: screening ten thousand genes for a signal, or testing one hypothesis that a decade of prior work supports, call for very different thresholds. The convention worth keeping is that the level is chosen in advance and reported alongside the p-value, rather than adjusted once the number is on screen.
- Why can I not choose two tails for chi-square or F?
- Because those two distributions are asymmetric, and there is no single accepted way to make them two-sided. The usual recipe — double the smaller tail — behaves badly where it matters most: at a chi-square of exactly 0, which is a perfect fit between observation and expectation and therefore the least significant result possible, it returns a p-value of 0%. A perfect fit reported as maximal significance is not a rounding artefact but a wrong answer, and it would appear on screen looking entirely normal. The two tails are unambiguous for a symmetric distribution, where doubling one of them is exact, and for an asymmetric one the upper tail alone answers the question that is actually asked: how surprising is a statistic this large?
- Why does the page reject a negative chi-square or F statistic?
- Because neither statistic can be negative, so a negative value means the number came from somewhere other than the test it is being attributed to. Both are built from squared quantities — a chi-square from squared differences between counts, an F from a ratio of two variances — and a squared quantity is never below zero. The page could return 100%, and it would be arithmetically consistent, but it would also accept a miscomputed input and dress it up as a real result. Refusing costs nothing here, because there is no legitimate negative chi-square for the refusal to block.
- Why is there no critical value table on this page?
- Because a p-value table would have to be indexed by the test statistic, and the statistic can be any number at all, so such a table would need infinitely many rows. The tables printed in the back of a textbook run the other way round: a handful of rows for the conventional significance levels and columns for the degrees of freedom you might need, which fits on a page only because the significance levels are few and agreed in advance. This page evaluates the tail area at exactly the statistic you typed, which is the one thing a table cannot do — a table can only tell you whether your statistic fell above or below the values it happens to list. The nearest table in this set is on the critical value page, where the rows are the significance levels and the columns give the matching z. The two pages are the same fact read from opposite ends, which is why each of them links to the other first among its related tools.
References
- 1.3.6.7.1. Cumulative Distribution Function of the Standard Normal Distribution — e-Handbook of Statistical Methods (the cumulative function a tail probability is read from, and the table it is normally looked up in) — National Institute of Standards and Technology (NIST)
- 6.2 Using the Normal Distribution — Introductory Statistics 2e (reading probabilities off a normal curve, including the upper and lower complements a two-tailed test is assembled from) — OpenStax, Rice University
- 1.3.6.1. What is a Probability Distribution — e-Handbook of Statistical Methods (a test statistic as a quantity with a distribution under a stated hypothesis, which is what makes a tail area meaningful) — National Institute of Standards and Technology (NIST)