To run a chi square goodness of fit test in SPSS you check your categories, type the expected proportions into the Expected Counts tab of the Nonparametric Tests dialog, and read the Pearson Chi-Square row in the output. On a clean dataset that takes about ten minutes.
The menu is not the hard part. The hard part is deciding what your expected distribution should be, and doing that before you look at the results. Get the expected proportions wrong and the whole test answers a question nobody asked.
This walkthrough covers the desktop version of IBM SPSS Statistics, the Analyze menu route, and the output tables you will actually see on screen. Menu wording shifts a little between releases, so I have named the tabs and settings as they appear in the standard desktop interface. If your window looks different, the tab names are the thing to look for.
Table of Contents
- 1What You Need
- 2The data
- 3A defensible expected distribution
- 4The assumptions, in one place
- 5Step-by-Step
- 6Step 1: Prepare the categorical variable
- 7Step 2: State the expected proportions before opening the test dialog
- 8Step 3: Run Analyze > Nonparametric Tests > Chi-Square
- 9Step 4: Check the output and interpret the test statistic
- 10Step 5: Report the result in an academic paper
- 11Common Mistakes
- 12Frequently Asked Questions
- 13When should I use a chi square goodness of fit test?
- 14What if my expected counts are lower than five in SPSS?
- 15Should I enter expected proportions or expected counts?
- 16What does a nonsignificant chi square goodness of fit result mean?
- 17Can I use a chi square goodness of fit test in R or Stata?
- 18How do I report a chi square goodness of fit result in APA style?
- 19Conclusion
What You Need
You need four things before you open the test dialog, and only the first one is about data. The rest are decisions that shape the result, so settle them first.

The data
One categorical variable with mutually exclusive, exhaustive categories. Every case contributes exactly one observation, and the categories are recorded as codes rather than as text where possible, because a text variable with inconsistent spelling will quietly split a category in two.
You also need an alpha level, usually 0.05, chosen before you see the p-value. And you need somewhere to write down your expected proportions by hand, because you are going to check SPSS’s numbers against yours.
A defensible expected distribution
This is the part students skip. The expected distribution can come from theory, from a previous study, from a historical baseline, from a published reference such as a Mendelian ratio, or from the null of equal proportions. Whatever you choose, you need a written reason, because the test only tests what you tell it to test.
An expected distribution that is invented after looking at the observed counts guarantees a significant result, and a reviewer who spots that will not read the rest of your paper.
The assumptions, in one place
The procedure assumes one categorical variable, independent observations, and expected counts large enough for the chi-square approximation to hold. That last one trips people up, so write the expected counts out before you run anything.
| Assumption | What it means | How to check it in SPSS |
|---|---|---|
| One categorical variable | Goodness of fit compares counts across the categories of a single variable, not across a two-way table. | Run Analyze > Descriptive Statistics > Frequencies on the variable and read the frequency table. |
| Independent observations | Each case must be separate. Repeated measures from the same person, or clustered responses, break this. | Check your study design. If responses are grouped by person, school, or site, the ordinary test is not the right one. |
| Mutually exclusive categories | A case belongs to exactly one category, and every case falls into one of them. | Compare the Valid N in the frequency table against your total N, then inspect the value labels. |
| Expected counts | Multiply your total N by each expected proportion. Keep every count at 5 or above, and keep them all reasonably close together. | SPSS prints the expected counts in its own output table. Read it before you read the p-value. |
One practical note on sample size. A chi square goodness of fit test with tiny expected counts will still run in SPSS and will still return a p-value, but that p-value is unreliable. With 20 cases spread over four categories, you have expected counts of 5 each, which is the floor rather than a comfortable margin.
Step-by-Step
The procedure below follows the order I would use, which is not quite the order the dialog forces on you. The two steps that people skip are the second and the fifth, and they are the two that decide whether the analysis holds up.
Step 1: Prepare the categorical variable
Open the file and go to the Data View tab, then check three things in order. First, look at the actual values in the column, not just the value labels, by widening the column or clicking the cell and reading the measure in the status bar. A variable coded 1, 2, 3, 4, 5 with a stray 9 hiding in row 87 will produce five categories instead of four.
Second, open Variable View and confirm the Measure column says Nominal for that variable. SPSS will happily run frequency tables and chi-square tests on a Scale variable, but treating a category code as a number is how people end up reporting means instead of counts.
Third, run the frequency table so you have a real count to work from. Go to Analyze > Descriptive Statistics > Frequencies, move the variable into the Variable box, and click Statistics if you want the mean and median. The Frequency, Valid Percent and Cumulative Percent columns give you the observed counts for every category, including any you forgot about.
Look at the Valid N at the bottom of the table. If it is lower than the number of cases in your file, you have missing values or out-of-range codes, and you need to decide now whether those cases are genuinely missing or whether they belong in a category. Silently dropping them changes your expected counts.
Finally, give your value labels real names. When the output table says 1, 2, 3, 4 instead of Hardware, Software, Network and User error, you cannot do the post-hoc work in the next step, and neither can your reader.
Step 2: State the expected proportions before opening the test dialog
Write your null hypothesis in a sentence, then write the expected proportions next to it. The null hypothesis of a chi square goodness of fit test says that the population follows the distribution you specified, not that the categories are empty or that the sample is random.
Expected count for each category equals total N multiplied by the expected proportion for that category. If your baseline is 30 percent Hardware, 45 percent Software, 15 percent Network and 10 percent User error, and you have 240 valid cases, the expected counts are 72, 108, 36 and 24. They have to add up to your N exactly.
Those four numbers are the analysis. Write them on paper before you go near the software, because in about twenty minutes you are going to want to check whether SPSS agrees with you, and the only way to catch a wrong proportion is to have written it down in advance.
Do not enter 30, 45, 15 and 10. Those are percentages, and SPSS expects the same form you give it: either proportions that sum to 1, or raw counts that sum to your N. Mixing the two is the single most common way this test goes wrong, and it is easy to do because the numbers look plausible.
If you estimated the expected proportions from your own sample rather than from outside theory, you must subtract one degree of freedom for each parameter you estimated. SPSS will not do this for you. I have covered that in the mistakes section below, because it is the error that quietly produces an anti-conservative p-value.
Step 3: Run Analyze > Nonparametric Tests > Chi-Square
Go to Analyze > Nonparametric Tests > Chi-Square. The dialog that opens has three tabs: Nonparametric Tests, Expected Counts, and Options. Work across them in that order.
On the Nonparametric Tests tab, move your categorical variable into the Test Variables box. Then tick Produce statistics, which is what makes SPSS print the Test Statistics table, and tick Display frequencies if you want the counts repeated in the output. If the Test Statistics table is missing from your output later on, an unticked Produce statistics is the reason nine times out of ten.
On the Expected Counts tab, tick Use values from data and enter your proportions in the grid, one per category, in the same order as the categories appear in the variable. Leave the Values must sum to box at 1. SPSS will refuse to run and warn you if your entries do not add up, which is a useful guard rail, but it cannot tell you that you entered percentages instead of proportions.
There are two alternatives on that tab worth knowing about. Equal proportions based on observed sets every category to 1 over k, which is the right choice when your null really is that all categories are equally likely, such as a die or a coin. Use range groups categories into equal-sized bins, which matters for scores rather than nominal categories. Choose one deliberately rather than leaving the default, because the default silently assumes equal proportions and that assumption may not be yours.
On the Options tab, tick Descriptives if you want mean and median, and tick Continue to produce residual tables. That last checkbox is the one people miss, and it is the only way SPSS will print the residuals you need to find out which categories are driving a significant result. Leave both ticked.
Click OK. The Output Viewer opens with up to four tables: a Frequencies table, an Expected Counts table, a Test Statistics table, and a Residuals table if you ticked the option above. The next step is reading them in that order.
Step 4: Check the output and interpret the test statistic
Read the Expected Counts table first, every time, before you look at the p-value. The values in the Expected column should match the numbers you wrote down in Step 2. If they do not, something has gone wrong with the proportions you entered, and a p-value computed from the wrong expectation is worse than no p-value at all.
Then check the residuals. In the Residuals table, the Std. Residual column is the observed count minus the expected count, divided by the square root of the expected count. It tells you how far a single category sits from expectation in standard-error units, so you can compare categories of different sizes. A standard residual above 2 or below negative 2 is the usual flag, and a squared residual above 5 is the more conservative version some disciplines use.
Working through the example: Hardware came in at 55 observed against 72 expected, a standard residual of negative 2.00. Software came in high at 118 against 108, a residual of 0.96. Network came in at 30 against 36, a residual of negative 1.00. User error came in at 37 against 24, a residual of 2.65. Software and Network are sitting close to expectation, and both Hardware and User error are far from it. The overall test is significant, and the residuals tell you the difference lives in the first and fourth categories.
| Category | Observed | Expected | (O − E) | (O − E) squared / E | Std. residual |
|---|---|---|---|---|---|
| Hardware | 55 | 72 | −17 | 4.014 | −2.00 |
| Software | 118 | 108 | 10 | 0.926 | 0.96 |
| Network | 30 | 36 | −6 | 1.000 | −1.00 |
| User error | 37 | 24 | 13 | 7.042 | 2.65 |
| Total | 240 | 240 | 0 | 12.98 | — |
The Test Statistics table gives you the Pearson Chi-Square value, the degrees of freedom, and the p-value. In the worked example, Pearson Chi-Square equals 12.98 with 3 degrees of freedom, and the reported significance is .005.
One thing trips up nearly everyone reading SPSS output: the column is headed Asymp. Sig. (2-sided). For this test that label is misleading. A chi-square statistic cannot be negative, so SPSS reports the probability in the right tail, and the two-sided label is a leftover from the shared output format. Compare it to your alpha exactly as if it were a one-tailed p-value, which is what it is.
Degrees of freedom for a goodness of fit test equal the number of categories minus one, so four categories give you 3. You can confirm that against a chi-square distribution table: at 3 degrees of freedom and alpha 0.05 the critical value is 7.815, and 12.98 is above it, so the result is significant. Reporting the critical value alongside the p-value is old-fashioned but harmless, and it lets a reader check you without a computer.
| Degrees of freedom | Alpha 0.10 | Alpha 0.05 | Alpha 0.01 |
|---|---|---|---|
| 1 | 2.706 | 3.841 | 6.635 |
| 2 | 4.605 | 5.991 | 9.210 |
| 3 | 6.251 | 7.815 | 11.345 |
| 4 | 7.779 | 9.488 | 13.277 |
| 5 | 9.236 | 11.070 | 15.086 |
Now write your conclusion in plain language. If the p-value is below alpha, you reject the null hypothesis and state that the observed distribution differs from the expected one, then say where using the residuals. If the p-value is above alpha, you fail to reject the null hypothesis and state that the data are consistent with the expected distribution. Do not write that you proved it, and more on that below.
Add an effect size. Cohen’s w is the square root of the chi-square value divided by your total N, so for this example the square root of 12.98 over 240 is 0.23. On the usual scale that is a small effect, which is worth saying out loud: a significant result on a large sample can be a difference too small to matter in practice, and stating both numbers protects you from that criticism.
Worth knowing, since SPSS is rarely the only tool in a project: the same test is a single formula in Excel with the CHISQ.TEST function applied to your observed and expected ranges, and one line in R with chisq.test(x, p = c(0.30, 0.45, 0.15, 0.10)). Both return the same Pearson statistic, the same degrees of freedom and the same p-value. Running one of them as a check on SPSS is a cheap way to catch a transcription error.
Step 5: Report the result in an academic paper
Most marks are lost in this step, not the calculation. An APA 7 report needs four things: the test statistic, the degrees of freedom, the p-value, and what you concluded in substance rather than in symbols.
Fill-in template: A chi-square goodness-of-fit test showed that the distribution of ticket causes [did not differ from / differed significantly from] last year’s baseline, chi-square(3, N = 240) = 12.98, p = .005, Cohen’s w = 0.23.
Add a sentence that shows you looked inside the result. In this example: The significant result was driven by hardware tickets, which fell below expectation, and user-error tickets, which rose above it, while software and network causes stayed close to the baseline.
A few formatting details. Write p as p = .005 with no leading zero, italicise p and N but not chi-square, and report exact p-values to three decimal places rather than writing p less than .001. If the p-value came out at .000, report it as p less than .001, because no software can actually produce a p-value of zero.
Always report the effect size alongside the p-value in an APA paper, and always say how you got your expected distribution. A reader who cannot tell whether your expected proportions came from theory or from a previous cohort cannot judge whether the test is worth anything.
Common Mistakes
Every one of these produces output that looks perfectly normal. None of them will raise an error message.
Using a test of independence instead. If you have two categorical variables, or if you are comparing the same categories across two independent groups, you need the Chi-Square Test of Independence, which lives under Analyze > General Linear Tables > Crosstabs. Goodness of fit answers one question only: does one variable follow a stated distribution? Independence answers a different one: are two variables associated? Using the wrong route gives you a plausible p-value about something you never asked.
Entering percentages as counts. Typing 30, 45, 15, 10 instead of 0.30, 0.45, 0.15, 0.10 makes SPSS compare your observed counts against an expectation that does not sum to your sample size, and the resulting chi-square value is meaningless. Enter proportions that sum to 1, or raw counts that sum to N, and pick one form for the whole table.
Choosing expected proportions after seeing the data. If you look at your frequency table, notice the counts are lopsided, and then set your expected distribution to match, the test can only ever return a nonsignificant result. That is not a finding, it is a rigged comparison. The fix is procedural rather than technical: write the expected proportions down and date them before you run the test, so the order of operations is visible.
Ignoring expected counts below 5. The chi-square approximation is unreliable when expected counts are small. The usual responses are to pool categories that have a defensible reason to be combined, to run an exact multinomial test, or to run a Monte Carlo version of the test that simulates the sampling distribution. Pooling on a substantive ground such as combining two rare failure modes into Other is legitimate; pooling to make the numbers work is not.
Deleting categories with small observed counts. Removing a category removes the very evidence of a departure from the expected distribution. If a rare category is genuinely part of your theory, keep it and deal with the small expected count properly. If it is a genuine data entry problem, correct the underlying records, not the analysis.
Forgetting the degrees of freedom correction. When you estimate the expected proportions from your own data rather than from theory, subtract one degree of freedom for each parameter estimated. With four categories and one estimated parameter, that is 3 minus 1, so you report 2. SPSS does not make this adjustment for you, and reporting the uncorrected degrees of freedom gives a p-value that is too small.
Reading a nonsignificant result as proof. Failing to reject the null hypothesis does not show that the expected distribution is correct. It shows that this sample did not provide strong enough evidence against it, and with a small sample that is easy. A helpful sentence: the result does not rule out small departures from the expected distribution, and a larger sample might detect them.
Reporting a p-value on its own. Significance tells you whether a difference is detectable, not whether it matters. Pair every chi-square result with Cohen’s w, and say in words which categories moved and by how much in raw counts, because that is the part your reader will act on.
Frequently Asked Questions
When should I use a chi square goodness of fit test?
Use it when you have counts for the categories of a single variable and a stated distribution to compare them against, such as equal proportions, a theoretical ratio, or a baseline from earlier data. Typical uses include checking whether a die is fair, whether offspring genotypes match a Mendelian ratio, or whether a year’s support tickets follow last year’s pattern. It is not the right test for continuous measurements, for two variables, or for samples that are not independent.
What if my expected counts are lower than five in SPSS?
SPSS will still run the test, but the chi-square approximation is unreliable with small expected counts, so treat the p-value with caution. Three options work. Pool categories that have a substantive reason to be combined. Run an exact multinomial test, which is more accurate at small counts. Or run a Monte Carlo version of the test that simulates the sampling distribution. Check the Expected Counts table in your SPSS output before you interpret anything.
Should I enter expected proportions or expected counts?
Enter proportions, on the Expected Counts tab under Use values from data, and let SPSS multiply them by your total N. Either form works as long as it is consistent and sums correctly: proportions must sum to 1 and counts must sum to your N. What does not work is percentages such as 30, 45, 15, 10, which look plausible, run without complaint, and produce a chi-square value that means nothing.
What does a nonsignificant chi square goodness of fit result mean?
It means the data were consistent with your expected distribution, not that you proved the distribution is correct. A small sample gives a nonsignificant result even when real differences exist, because there is not enough data to detect them. The correct wording is that you failed to reject the null hypothesis. If the question matters practically, report the observed counts for each category so readers can see the size of the gaps directly.
Can I use a chi square goodness of fit test in R or Stata?
Yes, and both give the same Pearson statistic, degrees of freedom and p-value that SPSS reports. In R the whole test is one line: chisq.test(x, p = c(0.30, 0.45, 0.15, 0.10)). In Stata, tabi and the user-written gfit command both do the job. Most tutorials on how to run a chi square goodness of fit test stop at SPSS, so the thing worth carrying across is the expected proportions, which must be specified in every package.
How do I report a chi square goodness of fit result in APA style?
Report four things: the test statistic, the degrees of freedom, the p-value, and the effect size. A template reads: A chi-square goodness-of-fit test showed that the observed distribution differed significantly from the expected distribution, chi-square(3, N = 240) = 12.98, p = .005, Cohen’s w = 0.23. Italicise p and N, omit the leading zero before p, and add a sentence naming the categories that drove the result.
Conclusion
Start with the expected proportions, not the software. Decide what distribution you are testing against, write those proportions and the resulting expected counts on paper, and check them against the Expected Counts table that SPSS prints. Then go to Analyze > Nonparametric Tests > Chi-Square, tick Produce statistics, enter the proportions under Use values from data, and read the residuals so you know which categories moved.
When you write it up, give the statistic, the degrees of freedom, the p-value and an effect size, then add one sentence of substance. That is a report a reader can check, which is the whole point of running the test by hand first.


