To run a Pearson correlation in SPSS, open your data file, go to Analyze > Correlate > Bivariate, move your two numeric variables into the Variables box, leave Pearson selected under Correlation Coefficients, and click OK. The output gives you r, a two-tailed significance value, and N in one table.
The whole procedure takes about two minutes once your file is loaded. The part that trips people up is not the clicking. It is reading the output afterwards and knowing whether Pearson was the right test in the first place.
I will use one small example throughout so the numbers stay consistent: 60 students, the hours they studied before an exam (study_hours) and the exam mark out of 100 (exam_score). You can follow along with any two numeric variables of your own.
The menu paths below match IBM SPSS Statistics v29 and v30. The Bivariate dialog has looked the same since roughly v20, so if you are on an older academic licence the labels should still be there.
Table of Contents
- 1What You Need
- 2Pearson, Spearman or Kendall: which one do you need?
- 3Step-by-Step: How to Run a Pearson Correlation in SPSS
- 4Check That Both Variables Are Numeric in SPSS
- 5Check Linearity With a Scatterplot
- 6Open the Bivariate Correlations Dialog
- 7Select the Two Variables and Set the Options
- 8Run the Analysis and Read the Output
- 9Common Pearson Correlation Mistakes in SPSS
- 10Running Pearson on ordinal or categorical variables
- 11Treating significance as effect size
- 12Ignoring outliers
- 13Skipping the assumption checks
- 14Forgetting missing value codes
- 15Claiming causation
- 16Reading the asterisk wrong
- 17How to Interpret a Pearson Correlation in SPSS
- 18How to Report the Result in APA Style
- 19Frequently Asked Questions
- 20How do I know if I should use Pearson or Spearman?
- 21What does Sig. (2-tailed) mean in the SPSS correlation output?
- 22Why does N differ between pairs in my SPSS correlation table?
- 23Is a significant Pearson correlation a strong relationship?
- 24How do I create a scatterplot in SPSS before correlating?
- 25Conclusion
What You Need
You need two things above everything else: two numeric variables measured on each of the same cases, and a saved dataset. If your file is still a comma-separated import rather than a .sav file, save it first so the variable types are stored properly.
Then work through these checks. They take two minutes and they decide whether the rest of the article is useful to you.
- Both variables are numeric. In Variable View, the Type column should read Numeric for both. String variables such as student ID or postcode cannot be correlated.
- Measurement level is Scale. Pearson needs interval or ratio data. If a variable is marked Ordinal (a single Likert item, a rank, a grade band), think twice.
- Missing values are declared. Use the Missing column in Variable View to register user-missing codes such as 99 or 999, otherwise SPSS treats them as real data points and quietly wreck your result.
- Each row is one case. One respondent, one participant, one observation. Correlating rows that represent different time points for the same person is a different analysis.
- You have a scatterplot ready. Ten seconds of plotting saves a paragraph of defending an r that was never going to be meaningful.
Pearson, Spearman or Kendall: which one do you need?
| Test | Use it when | What it assumes |
|---|---|---|
| Pearson | Two continuous (scale) variables, roughly linear relationship, no extreme outliers | Interval or ratio data, approximate normality, linearity |
| Spearman | Ordinal data, single Likert items, skewed or heavily tailed scores, or a monotonic but curved relationship | Ranked observations; ties are handled |
| Kendall’s tau | Same situations as Spearman, plus small samples or lots of tied ranks | Ranked observations; more conservative |
A quick rule from the SPSS forums that has held up well: if you are staring at a five-point satisfaction scale and wondering whether to treat it as continuous, run Spearman and Pearson and compare. If the two coefficients sit far apart, your data was not linear enough and the Pearson number is misleading. This is a question that comes up constantly on r/spss and r/AskStatistics, and it deserves a real answer rather than a shrug.
Step-by-Step: How to Run a Pearson Correlation in SPSS
Check That Both Variables Are Numeric in SPSS
Open the Data View grid and look at the type indicator in the column header. SPSS shows a tiny number icon for numeric variables and an Aa icon for string variables. If either of your variables is text, use Transform > Automatic Recode to turn it into numbers, or drop it from the analysis.
Then switch to Variable View and check four columns for both variables.
- Type should be Numeric.
- Label should say something meaningful. The label, not the variable name, is what appears in your output table and in the APA write-up.
- Missing should list any codes that mean “no data”. Click the cell, choose Discrete missing values, and type the code in the box.
- Measure should be Scale for a continuous variable.
Before you trust r, run Analyze > Descriptive Statistics > Explore, move both variables into Dependent List, click Plots, tick Normality plots with tests, then OK. A histogram that looks like a single spike and a long tail is a good reason to stop and use Spearman.
Check Linearity With a Scatterplot
Pearson only measures straight-line relationships. A strong curved pattern can produce an r near zero, which sends students looking for software bugs that do not exist.
- Go to Graphs > Chart Builder.
- Pick Scatter/Dot from the gallery and drag Simple Scatter into the canvas.
- Drag
study_hoursinto the X axis box andexam_scoreinto the Y axis box. - Click the element, tick Show regression lines, then click OK.
On an older interface, Graphs > Legacy Dialogs > Scatter/Dot gets you the same chart faster. What you are looking for is a cloud of points that hugs a line. Points stacked in a rainbow, or a few dots floating far from the rest, tell you to deal with outliers before correlating.
Open the Bivariate Correlations Dialog
Go to Analyze > Correlate > Bivariate. The dialog that opens has four areas worth knowing.
Variables is the empty box on the left where you drop your numeric variables. Available Variables shows everything in the file; Variables in the Source is empty until you move something across.
Correlation Coefficients holds three checkboxes: Pearson, Kendall’s tau and Spearman. Pearson is ticked by default, which is exactly why so many people run it without thinking.
Test of Significance offers Two-tailed or One-tailed. Two-tailed is the safe default because it tests for a relationship in either direction. Only choose one-tailed when your hypothesis states the direction in advance and you will report it as such.
Options opens a second dialog containing Means and standard deviations, Standardized covariances, Exclude cases pairwise or listwise, and Minimum cases to be computed. Tick the means checkbox if you plan to describe the variables in your results section.
Select the Two Variables and Set the Options
Select study_hours in the Available Variables box, click the arrow button, and it appears in the Variables list. Do the same for exam_score.
Confirm that Pearson is ticked, leave Two-tailed selected, and open Options. Two settings are worth your attention:
- Means and standard deviations adds a second table with the mean and standard deviation of each variable, useful for your write-up and harmless otherwise.
- Exclude cases is the missing data decision. Pairwise uses every case that has valid data for that specific pair. Listwise drops any case missing on any variable in the analysis.
| Setting | What happens | Use it when |
|---|---|---|
| Pairwise (SPSS default) | Each pair uses all cases valid for those two variables, so N can differ per row | You have scattered missing values and want maximum data use |
| Listwise | Any case missing on any listed variable is removed from every pair, so N is identical throughout | You need one consistent sample for interpretation or comparison with other tests |
If you are unsure, run it both ways. If the coefficients shift noticeably between pairwise and listwise, the missing data is doing real work and that belongs in your limitations section.
One optional checkbox, Flag significant correlations, puts an asterisk next to any coefficient below your chosen alpha level (0.01 or 0.05). It is a visual shortcut, not a substitute for reading the Sig. column.
Run the Analysis and Read the Output
Click OK. The Output Viewer opens a table titled Correlations and, if you asked for it, a second table of descriptive statistics.
Here is what the three columns mean.
| Column | Meaning |
|---|---|
| Pearson Correlation | The coefficient r, from -1.00 to +1.00. Sign gives direction, size gives strength. |
| Sig. (2-tailed) | The p-value: the probability of an r this large or larger arising by chance if the true correlation is zero. |
| N | The number of valid cases that went into that specific coefficient. |
In our example the cell reads .482, .000, 58. Two cases were missing an exam mark, so pairwise deletion left N at 58 rather than the 60 rows you entered.
Do not confuse r with the coefficient of determination. Squaring .482 gives .232, meaning about 23 percent of the shared variance in the two variables, which is a different claim from the one you make with r.
Common Pearson Correlation Mistakes in SPSS
Running Pearson on ordinal or categorical variables
Nominal variables such as hair colour or degree subject cannot be correlated. SPSS will either refuse or quietly code them, and the number it produces means nothing. Ordinal items need a decision: combine them into a scale, treat them as continuous after justifying it, or switch to Spearman.
Treating significance as effect size
This is the most damaging mistake on the list. With a large enough sample, p will come back below .001 for a correlation of .05, and students then describe that relationship as strong. The p-value tells you whether an effect is unlikely to be chance. The size of r tells you how much there is. Report both, never one alone.
Ignoring outliers
A single student who slept for 40 hours and scored 12 will drag r downwards hard, because Pearson uses the actual values rather than the ranks. Check your scatterplot, decide whether the point is a data error or a legitimate case, and either correct or remove it before you analyse. Removing cases because they weaken your result is a different act, and it needs disclosing.
Skipping the assumption checks
A rectangular scatterplot, a straight line, and no wildly outlying points are the minimum before you press OK. If the plot is curved, no significance value will rescue the analysis. Run Spearman instead.
Forgetting missing value codes
If a survey field lets people leave a score blank and your import filled it with 99, SPSS treats 99 as a genuine score unless you declare it as user-missing. The result is a real but meaningless negative correlation. This one causes more confusion on r/spss than any other output oddity.
Claiming causation
r = .48 says the two variables move together. It says nothing about which one drives which, and neither direction is established by the coefficient. If the mechanism matters for your argument, you need a design or an experiment that supports it.
Reading the asterisk wrong
An asterisk appears next to significant coefficients only, and in a symmetric table it shows up on one side of the diagonal, not both. A diagonal value of 1.000 for a variable against itself is not a finding. It just means the variable is perfectly correlated with itself.
How to Interpret a Pearson Correlation in SPSS
Start with the sign. A positive r means both variables rise together; a negative r means one rises as the other falls. Then look at the size.
| Value of r | Strength of the linear relationship |
|---|---|
| .00 to .10 | Negligible |
| .10 to .30 | Weak |
| .30 to .50 | Moderate |
| .50 to .70 | Strong |
| .70 to .90 | Very strong |
| .90 to 1.00 | Near perfect |
So r = .482 is a moderate positive relationship. Hours studied and exam marks move together, which matches what your scatterplot showed and what you would hope to see.
Next, read Sig. (2-tailed) against your chosen alpha level, usually .05. A Sig. value of .000 does not mean the probability is zero. SPSS simply cannot print a value that small, so you report it as p < .001.
Check N next. Degrees of freedom for a correlation equal N minus 2, so N = 58 gives you df = 56. That number is part of the APA report and it tells you how much the estimate could shift with a different sample.
Finally, ask what the coefficient means for your research question. A correlation answers whether two measured variables are associated in this sample. It does not tell you why, and it does not transfer automatically to a different population, a different measurement, or a later time point.
How to Report the Result in APA Style
Use the correlation symbol r with two decimal places, the degrees of freedom in brackets, the p-value with no leading zero, and the sample size in the text where it fits your style.
Template for a significant result:
There was a [moderate/strong/positive] correlation between [variable name] and [variable name], r([N minus 2]) = [.48], p < .001, 95% CI [lower] to [upper].
Worked example from the analysis above:
There was a moderate positive correlation between study hours and exam score, r(56) = .48, p < .001, N = 58.
Template for a non-significant result:
There was no significant correlation between [variable A] and [variable B], r([df]) = [.12], p = [.24], N = [58].
Formatting rules that cost students marks: never write p = .000, never write a leading zero before r or p, and always italicise the statistical symbols. SPSS gives you the raw numbers; the rounding and the p-value rewrite are your job.
If your department wants confidence intervals, request them with syntax rather than by hand:
CORRELATIONS
/VARIABLES=study_hours exam_score
/PRINT=TWOTAIL NOSIG
/MISSING=LISTWISE.
The /PRINT subcommand controls the significance table and the asterisk flag; /MISSING switches between pairwise and listwise deletion. Adding CI to the CORRELATIONS command is not valid, so if confidence intervals are required, bootstrap them in a separate procedure or report the interval from your software’s regression output, which produces the same r.
Frequently Asked Questions
How do I know if I should use Pearson or Spearman?
Use Pearson when both variables are continuous scale measures, the scatterplot looks linear, and there are no extreme outliers. Use Spearman when either variable is ordinal, a single Likert item, or badly skewed. A fast check is to run both correlations: if the coefficients differ sharply, the relationship was not linear and Pearson is overstating the story.
What does Sig. (2-tailed) mean in the SPSS correlation output?
It is the p-value for the correlation, testing the null hypothesis that the true correlation is zero. Two-tailed means SPSS looked for a relationship in either direction, which is the correct default unless you specified a direction beforehand. Compare it to your alpha level, usually .05. When SPSS shows .000, report it as p less than .001.
Why does N differ between pairs in my SPSS correlation table?
That is pairwise deletion, which is SPSS’s default. Each coefficient uses every case with valid data for that specific pair, so a case missing one variable still contributes to the pairs that do not involve it. Switching Exclude cases to listwise gives you one consistent N across the whole table, at the cost of discarding more cases.
Is a significant Pearson correlation a strong relationship?
No. Significance and strength are separate. With a large sample, a correlation of .05 can return p below .001, which means it is unlikely to be chance but almost worthless for prediction. Always report the size of r alongside the p-value, and treat r below about .30 as a weak relationship no matter how small the p-value is.
How do I create a scatterplot in SPSS before correlating?
Open Graphs, then Chart Builder, drag Simple Scatter into the canvas, and drop your predictor into the X axis box and the outcome into the Y axis box. Tick Show regression lines so the fitted line appears. On older versions, Graphs, then Legacy Dialogs, then Scatter/Dot does the same job. Plotting first is the quickest way to catch outliers before you correlate.
Conclusion
Start by confirming both variables are numeric, declared Scale, and free of undeclared missing codes. Then plot them, open Analyze > Correlate > Bivariate, keep Pearson selected, and run the analysis. Read r for size and direction, Sig. (2-tailed) against your alpha level, and N before you write anything down. That takes about five minutes and it is the difference between a result you can defend and one you have to redo.


