Chi-Square Calculator
Result
Chi-square statistic
- P-value
- 0.0170%
- Degrees of freedom
- 3
- Smallest expected count
- 25.0000
A chi-square calculator compares counted categories against the frequencies you expected, and reports the statistic, its degrees of freedom and the p-value that goes with it. It is the test behind every question of the form are these groups the same size: whether a die is fair, whether four options were picked equally often, whether the observed frequency in each category matches the expected frequency you had a reason to predict. Leave the expected frequencies blank and the page assumes every category is equally likely, which is both the commonest case and the one that needs no arithmetic done first. The smallest expected frequency is printed alongside, because that is the number the test's own conditions are checked against.
Formula
χ² = Σ ( O − E )² / E df = k − 1 p = P( χ² ≥ χ²₀ )
- O
- The observed frequency in each category — a count, not a measurement, so the numbers are whole and their sum is the sample size. A count of zero is allowed: it means the category never came up, which is a result rather than a mistake
- E
- The expected frequency in each category, which is what the null hypothesis predicts rather than what you hope to see. Left blank, it becomes the total divided by the number of categories, so the test asks whether the categories are equally likely
- k
- The number of categories, read off the length of the list you typed. It sets the degrees of freedom and it is also what the expected frequencies are averaged over when they are left blank, so one list of counts determines both
- df
- The degrees of freedom, which is k − 1 rather than k. The subtraction is the part people forget: once the total and all but one of the counts are fixed, the last count is determined, so the categories carry one less piece of information than there are of them
- χ²₀
- The statistic this page computed, inserted into its own null distribution to get the tail area. It is the same operation as reading a value off a normal curve, with the chi-square curve in place of the normal one — and it is why the p-value needs the degrees of freedom as well as the statistic
Use it when the data is a set of counts falling into categories, and the question is whether the counts are consistent with some prediction. Counts are what make this test different from the ones on the mean and the proportion: the observations are not measured on a scale, and what matters is how many landed in each bucket rather than how large any of them was. The prediction can be an even split, or a set of expected frequencies built from previous data — a national rate applied to a local population, a Mendelian ratio, a model's output. The test is only as good as those expectations: it compares what you counted against what you said, and it will not notice that the expectation itself was wrong. Two conditions are worth checking before trusting the p-value, and the page prints the first of them for you, since the smallest expected frequency is the number the usual rule of thumb is about.
Worked examples
Four categories counted unevenly, with the expected frequencies left blank
- The total is 100, and with the blank expectation each of the 4 categories is expected to hold 100 / 4 = 25
- Squared differences from 25: (30 − 25)² = 25, (10 − 25)² = 225, (20 − 25)² = 25, (40 − 25)² = 225
- Divide each by its expected 25 and add: 1 + 9 + 1 + 9 = 20
- Degrees of freedom: k − 1 = 3, and the upper tail beyond 20 is 0.017%
This is the page's default, and it was chosen to be visibly uneven: 40 and 10 against an expectation of 25 each. The p-value comes out at 0.017%, so the counts are far from what an even split would produce, and the smallest expected frequency of 25 is comfortably above any of the thresholds in common use — so the approximation behind the p-value is on solid ground here. Note that the expected frequencies never had to be typed. Leaving the field blank is how an even split is expressed, and it is the reason the field is optional at all.
Three categories against expected frequencies taken from an earlier survey
- Both lists total 100, so the expectation is consistent with the counts and the test can proceed
- Squared differences: (50 − 60)² = 100, (30 − 30)² = 0, (20 − 10)² = 100
- Divide by each expectation and add: 100/60 + 0/30 + 100/10 = 1.6667 + 0 + 10 = 11.6667
- Degrees of freedom: 3 − 1 = 2, and the upper tail beyond 11.6667 is 0.2928%
The two lists have to describe the same number of observations, and here both come to 100 — which is what makes the second list a hypothesis about the first rather than a different experiment. Reading the terms one at a time is instructive: the middle category matched its expectation exactly and contributed nothing, while the third contributed ten of the eleven and a half points of the statistic despite being the smallest category, because it missed its expectation by 100%. The smallest expectation of 10 is still above the level usually required.
A coin flipped 20 times, tested against an even split
- Expected counts: 20 flips split evenly between two outcomes gives 10 and 10
- Squared differences: (12 − 10)² = 4 and (8 − 10)² = 4
- Divide and add: 4/10 + 4/10 = 0.8
- Degrees of freedom: 2 − 1 = 1, and the upper tail beyond 0.8 is 37.1093%
12 heads out of 20 looks like a biased coin until the tail probability is read: 37% of fair coins produce a split at least this uneven, so the result is unremarkable. This is the same arithmetic as the first example done on the smallest useful table, and it shows why the statistic is never read on its own — 0.8 and 20 are not comparable numbers, and only the p-value puts them on a scale. The two-category case also has a shortcut worth knowing: the chi-square statistic equals the square of the z a two-proportion test would use.
Limitations
The p-value rests on an approximation that needs the expected frequencies to be reasonably large, and the printed smallest expected frequency is the number to check that against — textbooks disagree on where the line sits, with at least 5 in every category and at least 1 in every category plus 5 in most of them both in wide use. When the expectations are too small the approximation fails quietly, returning a p-value that looks ordinary. The counts are also assumed to be independent, which fails for paired or repeated measurements and for clustered sampling, and each of those makes the statistic too large. The expectations themselves are taken as given, so a test against a wrong prediction answers a different question from the one intended. Finally, a small p-value says the counts are inconsistent with the expectation, not why: it does not say which category is responsible, and with several categories the differences that drive the statistic are worth inspecting separately.
Frequently asked questions
- What does the chi-square test actually tell me?
- Whether the counts you observed are consistent with the frequencies you expected, given how many observations there are. The statistic grows with the total squared mismatch between observation and expectation, scaled by how large each expectation was, so the same proportional gap counts for more when the expectations are small. The p-value then says how often a fair world would produce a mismatch at least that large. It is a test of the whole table rather than of any single category, and the two are easy to confuse when one category is obviously the odd one out.
- What happens if I leave the expected frequencies blank?
- The page assumes every category is equally likely and computes the expected frequency as the total divided by the number of categories. That is the commonest use of the test — a die, a spinner, four buttons that ought to get equal traffic — and it saves typing a list whose values are all the same anyway. It is also the only reading available when no prior data exists to predict the split. Type the expected frequencies whenever you have a real prediction to test against, such as last year's distribution or a theoretical ratio, because a test against an even split and a test against a specific model answer different questions, and only the second one uses the information you already have.
- Why must the two lists add up to the same total?
- Because expected frequencies in a goodness of fit test are derived from the total number of observations, not chosen freely — they describe how those same observations would have been distributed if the hypothesis were true. Two lists with different totals describe two different experiments, and the statistic computed from them is not a test of anything. When every expected frequency is written as an integer, each can be off by up to half a unit without changing the intent, so the page allows a tolerance of 0.5 per category: 20.5, 20.5 and 20.5 against a total of 60 is accepted, since those are rounded values for an exact even split of 20 each.
- Can a category have a count of zero?
- Yes for an observed count, no for an expected one, and the asymmetry comes straight out of the formula. An observed zero is an ordinary result — a category that never came up — and the squared difference simply comes out large, which is what should happen. An expected zero cannot be used because it is the divisor, and the term would be infinite: the prediction that a category can never occur is not a prediction this test can weigh against a count. If a real expectation is very small rather than zero, the test accepts it, and the smallest expected frequency printed below the result is where to see whether that is a problem.
- Where is the chi-square critical value table?
- Not on this page, for the same reason as the p-value page: the single value that matters is already printed, since the statistic and its degrees of freedom determine the tail area directly. A table would also have to be built at one significance level, so a reader working at 1% would be looking at a table contradicting the panel. The critical value tool is the place for the lookup, and there is one further reason this page cannot show a table even in principle: its degrees of freedom are k − 1, where k is however many categories the reader typed, so the row of the table that applies is not known until the counts are.
References
- 1.3.6.1. What is a Probability Distribution — e-Handbook of Statistical Methods (a test statistic as a quantity with its own distribution under the null hypothesis, which is what makes the tail area computable) — National Institute of Standards and Technology (NIST)
- 3.1 Terminology — Introductory Statistics 2e (relative frequency as an estimate of a probability, which is the step that turns a category's probability into an expected count of observations) — OpenStax (Rice University)
- 1.3.6.7.1. Cumulative Distribution Function of the Standard Normal Distribution — e-Handbook of Statistical Methods (the same reading of a tail area off a cumulative curve that this page performs with the chi-square curve instead) — National Institute of Standards and Technology (NIST)