How to Run Stepwise Regression in SPSS (2026) – Why Be Careful

Stepwise regression is a procedure that builds a multiple regression model one variable at a time, adding or removing predictors according to a significance threshold until no change improves the model. In SPSS you run it from Analyze > Regression > Linear, pick Stepwise in the Method box, and read the output tables in the order they appear. The procedure is quick and produces a tidy short model, which is exactly why students reach for it and also exactly why most methodologists warn against it. This guide shows the full workflow and then the eight problems you need to know before you report anything.

The short version: treat stepwise as an exploratory screening tool, not as evidence for your theory. Every variable that stays in the final model was picked because of the same data you are now testing, so the p-values beside it are not honest tests.

Table of Contents
  1. 1What You Need
  2. 2A continuous outcome plus a set of candidate predictors
  3. 3Enough rows for the number of candidates
  4. 4Missing values dealt with deliberately
  5. 5A recent SPSS version, and a syntax file habit
  6. 6Step-by-Step: Running Stepwise Regression in SPSS
  7. 7Open the Linear Regression dialog
  8. 8Move your variables into the right boxes
  9. 9Set Method to Stepwise
  10. 10Set the entry and removal thresholds
  11. 11Paste the syntax instead of pressing OK
  12. 12Confirm the run worked
  13. 13How to Run Stepwise Regression in R
  14. 14Understand What Stepwise Regression Is Doing
  15. 15A short worked example, so the output stops being abstract
  16. 16How to Interpret the SPSS Output
  17. 171. Variables Entered/Removed
  18. 182. Model Summary
  19. 193. ANOVA
  20. 204. Coefficients
  21. 215. Excluded Variables
  22. 22Residual diagnostics
  23. 23Why to Be Careful With Stepwise Regression
  24. 24Better Alternatives to Stepwise Regression
  25. 25How to Check and Report the Final Model
  26. 26Rerun the selected variables with Enter
  27. 27Run a hold-out or cross-validation check
  28. 28Work through the diagnostics checklist
  29. 29Report the whole process
  30. 30Common Mistakes and How to Fix Them
  31. 31Frequently Asked Questions
  32. 32What does stepwise regression mean?
  33. 33How do I perform stepwise regression in R?
  34. 34What is the difference between forward and backward stepwise selection?
  35. 35Why is stepwise regression criticized?
  36. 36What alpha should I use for stepwise regression?
  37. 37How many observations do I need for stepwise regression?
  38. 38Conclusion

What You Need

Before you open SPSS, four things have to be in order. Fix them first, because any one of them will quietly wreck the result.

A continuous outcome plus a set of candidate predictors

Linear regression needs one dependent variable that is numeric and roughly continuous, and several predictor variables measured on the same rows. Age, income, a depression score, a treatment condition coded 0 and 1, survey items, clinical measures, any of these work. The dependent variable has to vary and not be a constant.

Enough rows for the number of candidates

A common working rule is 15 to 20 complete cases per candidate predictor as a floor for ordinary multiple regression. Stepwise deserves more headroom than that, because you will evaluate every candidate several times over. With 30 candidates in your list and 200 rows, you are fitting many more than 30 parameters in effect, and you should expect a model that looks better than it predicts. If you have fewer rows than parameters, standard multiple regression cannot run at all.

Missing values dealt with deliberately

SPSS drops any case with a missing value on any variable in the analysis by default, and it can drop a lot without telling you in a way you will notice. Check Analyze > Descriptive Statistics > Frequencies, look at the Missing row, and decide in advance whether listwise deletion is acceptable or whether you need Transform > Replace Missing Values or an imputation step. Also confirm your variable types: numeric variables that are actually categories need recoding before they enter the equation.

A recent SPSS version, and a syntax file habit

The menu path below has been stable across the last several releases of SPSS Statistics on Windows and macOS. What has changed most is the ribbon layout in versions 29 and later, where Analyze still leads to the same dialog. Use Paste rather than clicking OK. Saving the syntax means you can rerun the exact model later, hand your supervisor a reproducible file, and add one line to switch methods without rebuilding anything.

Step-by-Step: Running Stepwise Regression in SPSS

Open the Linear Regression dialog

Choose Analyze > Regression > Linear. If you use the ribbon in SPSS 29 or later, the Regression button sits in the Analyze group on the home tab and Linear is the first item in its menu.

Move your variables into the right boxes

Drag your outcome into Dependent. Put every candidate predictor into Independent, including the ones you expect to be dropped. Selection only means something if the algorithm is allowed to reject it. If any predictor is categorical with more than two levels, use Transform > Recode into Different Variables to create dummy variables first, or use Analyze > Regression > Binary Logistic, which handles the coding for you.

Set Method to Stepwise

The Method dropdown offers Enter, Forward, Backward, and Stepwise. Stepwise is the both-direction version: it starts from the null model, adds the most useful remaining predictor, then re-tests every variable already in the model and removes any that no longer earn their place. The loop repeats until nothing changes. Forward alone only adds, Backward alone only removes.

Set Method to Stepwise

Choosing Stepwise rather than Enter is the whole difference in the run. Enter puts your list in the model untouched; Stepwise makes the data decide.

Set the entry and removal thresholds

Expand Statistics and check Collinearity diagnostics, Estimates, Model fit, and R squared and adjusted. Under Residuals, check Predicted, Residual, and Standardized residual. Defaults for the thresholds are an entry value of 0.05 and a removal value of 0.10. The removal value has to be larger than the entry value, or SPSS complains. Setting both to 0.05 makes the procedure stingier and keeps more variables in.

Paste the syntax instead of pressing OK

Hit Paste. SPSS fills a syntax window with the whole command, which looks roughly like this:

REGRESSION /DEPENDENT outcome /METHOD=STEPWISE age income education score sleep /CRITERIA=PIN(.05) POUT(.10) /STATISTICS DESCRIPTIVES COV OUTR R ANOVA COLLIN /RESIDUALS (PRED, RESID) DURBIN.

Read it once. PIN is the alpha to enter, POUT is the alpha to remove, and COLLIN is what produces the Tolerance and VIF columns later. Run it with the green triangle and the Output Viewer opens.

Confirm the run worked

Four tables should appear under Regression: Variables Entered/Removed, Model Summary, ANOVA, and Coefficients. If you also asked for them, Excluded Variables and a Residual Statistics table follow. Check the Variables Entered/Removed table first: if it shows a single row saying the model was not run, or if the removal column is empty where you expected removals, something in your criteria settings is off.

How to Run Stepwise Regression in R

R users usually reach for stepwise through MASS::stepAIC(), which selects on AIC by default rather than on p-values. The pattern that surprises people is forward selection returning the null model, and the fix is the scope argument, which says which variables are already in the model when you begin.

library(MASS)

fit.full <- lm(outcome ~ age + income + education + score + sleep, data = df)

fit.null <- lm(outcome ~ 1, data = df)

step(fit.full, scope = list(lower = fit.null, upper = fit.full), direction = "both")

Swap direction = "forward" with "backward" to compare the two. Two errors come up constantly. If stepAIC is not found, you skipped library(MASS), and MASS is one of the packages bundled with every R install. If forward selection gives you an empty model, the scope argument was missing or defined wrongly, so the procedure had no legal way to enter a first variable.

For exhaustive search rather than sequential search, leaps::regsubsets() evaluates every possible model of each size:

library(leaps)

subsets <- regsubsets(outcome ~ age + income + education + score + sleep, data = df, k.max = 5, method = "exhaustive")

plot(subsets)

The plot draws adjusted R-squared, BIC and Cp against model size at once, which is a good way to see that the criteria disagree about how many variables you need.

Understand What Stepwise Regression Is Doing

Stepwise regression is a model-building procedure for multiple regression that adds or removes candidate predictors one at a time, in a series of steps, keeping whichever change most improves the model by a stated criterion until no further addition or removal is justified.

There are three strategies, and SPSS exposes all three in the Method dropdown. Forward selection begins with an empty model and adds the single variable that improves fit the most, repeating until nothing passes the entry threshold. Backward elimination begins with every predictor in and removes the least useful one at a time until all remaining variables clear the removal threshold. Both-direction stepwise does both, so a variable can enter early and be dropped later.

StrategyStarting modelCan it drop a variable it added?Typical use
ForwardNull model, outcome onlyNoMore candidates than rows, quick screening
BackwardAll candidatesNoModerate candidate lists, you suspect few matter
StepwiseNull modelYesSPSS default, screening with re-checks

The thresholds do the deciding. Alpha to enter is the p-value a variable must show to be admitted; alpha to remove is the level at which a variable already inside is kicked out. The rule in SPSS is that removal happens when a variable’s significance exceeds the removal criterion, and insertion happens when a remaining variable’s partial F-test falls below the entry criterion. The procedure stops when no addition or removal qualifies.

Two things selection is not. It is not a significance test, because the p-values in the final Coefficients table are computed on data that already chose the variables, so the usual error rate no longer holds. And it is not a theory. Selection answers which of these columns correlated with the outcome in this sample. Which variables should be in the model is a question about mechanisms, design and prior work, and no threshold can answer it for you.

A useful picture: think of it as a greedy search through model space. It walks one step at a time toward a better fit and stops at the first place it cannot improve. A greedy search is not guaranteed to land on the best model available. Adding variable A early can block the path to a model that would have fitted better through B and C instead.

A short worked example, so the output stops being abstract

Here is a hypothetical run on 200 cases with ten candidate predictors and default criteria, entry 0.05 and removal 0.10. The numbers are illustrative, but the shape is what real runs look like.

StepActionR SquareAdjusted R SquareSig. F Change
0Null model0.0000.000none
1Entered sleep quality0.3100.3060.000
2Entered total sleep hours0.4210.4130.000
3Entered caffeine intake0.4520.4370.012
4Entered screen time before bed0.4630.4410.089
5Entered bedroom temperature0.4690.4420.184

Two things to notice. Adjusted R Square peaks at step 4 and then falls, which is the quantitative form of the argument that the last variable is not earning its place. And the Sig. F Change column walks from 0.000 to 0.184, which tells you the process was close to stopping several times before it finished.

The same table is a warning when you read it against the data. Six of those ten candidates have now been tested at least once, several more than once, on the same 200 rows you will report. The final model’s p-values of 0.049 and 0.012 were picked out of that search, which is why the honest move is to refit those six variables with Enter and check whether they still hold on rows the search never saw.

How to Interpret the SPSS Output

Read the five tables in a fixed order. Skipping ahead to the significance column is how people end up defending a model that does not hold up.

1. Variables Entered/Removed

This is the audit trail: one row per step, listing what went in, what came out, the alpha used, the number of variables in the model, and R-squared change. If the method column says Forward rather than Stepwise, you picked the wrong Method. Read this table first because it tells you how many rounds happened. Ten or more rounds over a modest candidate list is a warning that the data are being mined rather than modelled.

2. Model Summary

R is the multiple correlation between the outcome and the predicted values. R Square is the share of outcome variance the model accounts for in sample. Adjusted R Square penalises the model for its parameters and is the number to quote. The Standard Error of the Estimate is the typical size of a residual in your outcome units, which is often more interpretable to a reader than R-squared. Drop one predictor and watch what happens to Adjusted R Square: if it rises, the variable was not earning its place.

3. ANOVA

The Regression row gives the degrees of freedom, the sum of squares, the mean square and the F ratio for the model as a whole. The Sig. Column shows the model summary; Sig. F Change gives the change from the previous model, and that is the value corresponding to the entry or removal that happened at that step.

4. Coefficients

Each surviving variable gets a row with B, the unstandardised coefficient in outcome units; its standard error; t; Sig., the p-value for that coefficient; and Beta, the standardised coefficient, which lets you compare predictors measured on different scales. If you selected Collinearity diagnostics you also get Tolerance and VIF. Tolerance above 0.2, or VIF under 5, is usually acceptable; VIF above 10 means serious overlap and a coefficient you should not interpret.

5. Excluded Variables

Every rejected predictor appears here with its partial correlation, its t, its p-value, and a Part B value showing what its partial correlation would become in the model without it. Part B below 0.3 with a p-value under 0.10 is the standard SPSS suppression pattern: the variable was blocked by collinearity rather than shown to be useless. Do not report these as non-predictors without saying so.

Residual diagnostics

The Residual Statistics table lists standardised residuals, cases with values beyond plus or minus 2 as outliers, and Cook’s distance, which flags influential cases. Plot residuals against predicted values: a random scatter means the linearity and constant variance assumptions look fine, while a funnel or a curve means they do not. Normality of residuals matters for the validity of the p-values, and stepwise makes those p-values fragile already.

Why to Be Careful With Stepwise Regression

Eight problems, each with what it costs you and what to do instead. This is the half of the topic that decides whether your analysis survives review.

1. R-squared is inflated. In-sample R-squared after stepwise is not comparable to R-squared from a model you specified in advance. Because the procedure searched many models and kept the best-looking one, the fit statistic is optimistically biased, and the bias grows with the number of candidates and the number of steps. Adjusted R Square and predicted R Square correct partly, not fully. Fix: report out-of-sample performance, not the in-sample number, and keep the candidate list small.

2. The p-values are not valid tests. Each coefficient’s p-value is calculated after selection already used that same variable and that same outcome. The nominal alpha is no longer the true error rate, so a p of .03 means less than three times in a hundred in the way a reader assumes. Fix: treat the final model as an exploratory finding, and test it on fresh data before making inferential claims.

3. The selected set is unstable. Change a few observations, lose a few cases to missingness, or resample the rows, and a different variable can replace another one. This is the pattern practitioners describe: a model with all coefficients significant and a high R-square that fails to replicate. Frank Harrell’s much-quoted line that stepwise variable selection has done incredible damage to science comes up in nearly every discussion of the problem. Fix: run forward, backward, and stepwise, then compare. If the final variable lists differ, your data do not support a single model.

4. Collinear predictors get picked at random. When two variables carry nearly the same information, whichever one has the smaller p-value by a hair enters, and the other sits in Excluded Variables with a small Part B. Coefficients in that situation can flip sign when the twin is swapped in. Fix: check VIF in the full Enter model before you start, and combine or drop overlapping measures first.

5. Type I error inflates across the search. Each round tests many variables and keeps the best. Run twenty steps and you have performed an implicit multiple-comparison exercise while reporting a single model-level test. Fix: prespecify the primary model and run stepwise only on the remainder.

6. Theory gets ignored. Stepwise is theory-free by construction. It cannot tell a mechanism from a proxy, and it happily drops a confounder that your design requires. Fix: put every design-mandated variable in the model first with Enter, then let stepwise work on the rest if you must.

7. No guarantee of the best model. Because the search is greedy and one-directional at each step, the stopping point is often not the model with the highest adjusted R-square. This surprises people who have seen a stepwise run stop at three variables while a four-variable model fits better. Fix: run best subsets with leaps::regsubsets() and compare against the stepwise result.

8. It does not scale, and it does not handle structure. When rows are fewer than candidate predictors, ordinary stepwise fails outright. And it will not find interactions or nonlinear terms unless you create them by hand first, so a model with an interaction between two predictors can miss it completely. Fix: for high-dimensional data use lasso or ridge; for interactions, specify them yourself.

One more practical warning. When forward and backward selection return different final models on the same data, that is not a menu choice to argue about in your discussion section. It is evidence that the set of predictors is not identified by this sample.

Better Alternatives to Stepwise Regression

MethodWhat it optimisesUse it when
Theory-driven specificationNothing, you decideYou are testing hypotheses or controlling for confounders. This is the default.
Best subsetsEvery model of a given sizeCandidate list is under about 20 and you want to see all combinations.
LassoPrediction, shrinks and zeroes coefficientsMany correlated predictors, you can live with some dropped to zero.
RidgePrediction, shrinks without zeroingHighly collinear predictors that you want to keep.
Elastic netBoth penalties at onceGroups of correlated variables, as in survey items or gene panels.
Cross-validated selectionHeld-out prediction errorYou care about prediction rather than inference.

The migration is easier than it sounds. In SPSS, lasso and ridge are not in the Regression dialog, so people stay with stepwise out of convenience. If your real question is prediction, move to R or Python with glmnet, where the penalty does the selection and cross-validation sets it.

Selection criteria differ in more than name, and the difference explains a lot of disagreement in the literature.

CriterionRule of thumbBehaviour
Entry p-value (PIN)0.05 or 0.10Ignores model size entirely, so it drifts toward larger models.
AICLower is betterPenalty grows with parameters; leaner than p-value rules.
BICLower is betterPenalty grows faster than AIC; picks smaller models, sometimes too eagerly.
Adjusted R-squareHighest winsOnly valid if adjusted values are close; stops only at 0.05 steps.
Mallows CpCp near the number of predictorsGood for small samples, less used in SPSS than in textbooks.

When these criteria disagree about model size, that disagreement is information about your data, not a problem to hide. Cross-validated error is the fairest tiebreaker because it is the only one measured out of sample.

How to Check and Report the Final Model

Stepwise gives you a candidate. These steps turn it into something you can defend.

Rerun the selected variables with Enter

Take the variables that survived and fit them in a conventional linear regression, with Method set back to Enter. Write down the Adjusted R Square from this clean model. It is always lower than or equal to the stepwise figure, and that gap is your measure of how much the selection inflated the fit. Quote the clean number.

Run a hold-out or cross-validation check

Split the data, run the selection on the training portion only, and evaluate prediction on the held-out portion. Repeat across ten folds if you have the rows for it. If cross-validated error is worse than the full model with every predictor, the selection bought you nothing but a shorter equation. This is the test that answers whether the pruning helped.

Work through the diagnostics checklist

Four checks, in this order: VIF and tolerance for multicollinearity, residual-versus-fitted for linearity and constant variance, a histogram or normal Q-Q plot of standardised residuals, and Cook’s distance for influential cases. Then confirm the sample size still supports the final parameter count at roughly 15 to 20 observations per predictor. Fix what fails and rerun; do not report around a problem.

Report the whole process

A reviewer reading a stepwise analysis wants five things, and omitting any of them invites the criticism you were trying to avoid.

  • The entry and removal criteria you used, by name and value.
  • The full candidate list, including every variable that never made it into the model.
  • Every variable removed at each step, and the step number.
  • The final model with unstandardised coefficients, standard errors, and adjusted R-square, not just the equation.
  • Whether you validated out of sample, and the result if you did.

Phrasing matters too. Say that the analysis was exploratory and that the retained variables are candidates for a confirmatory model, rather than presenting them as established. Students who ran stepwise and then described it as hypothesis-confirming get asked to redo the analysis; students who framed it as screening and validated the result usually do not.

Common Mistakes and How to Fix Them

MistakeWhy it hurtsFix
Treating the selected model as proof of causationSelection says nothing about mechanism or directionDescribe associations; reserve causal language for a designed experiment
Choosing predictors only by p-valueSingle p-values ignore design, measurement quality and confoundingStart from theory and literature, then let stepwise handle only the remainder
Ignoring multicollinearityCoefficients become unstable and may flip signRun the full Enter model first and check Tolerance and VIF
Skipping assumption checksInvalid p-values on top of already fragile onesResidual plots and a normality check before interpreting anything
Reporting only the final equationHides the search, and reviewers know it happenedReport criteria, candidate list, removals and validation
Quoting R-square from the stepwise runOptimistically biased by the searchQuote adjusted R-square from the clean Enter rerun, plus out-of-sample error
Too few rows for the candidate listUnstable coefficients, no meaningful validationCut candidates first or use a penalised method
Comparing forward to backward and picking a winnerChoosing the luckier result hides the instabilityReport that they differ, and treat the disagreement as a finding

One software habit prevents most of this. Paste your syntax, keep a version history, and never edit a model by clicking through dialogs you cannot later reproduce.

Frequently Asked Questions

What does stepwise regression mean?

Stepwise regression builds a multiple regression model by adding or removing candidate predictors one at a time. At each step the procedure tests every possible single change, applies the one that most improves the chosen criterion, and stops when no further addition or removal qualifies. SPSS offers three variants: forward selection, backward elimination, and both-direction stepwise.

How do I perform stepwise regression in R?

Fit the full model and the null model first, then pass both to MASS::stepAIC with the scope argument set to lower = fit.null and upper = fit.full. Set direction to forward, backward, or both. If you see ‘stepAIC not found’, load MASS first. If forward selection returns an empty model, your scope argument was missing, leaving it no legal way to enter a first variable.

What is the difference between forward and backward stepwise selection?

Forward selection starts from an outcome-only model and adds the variable that improves fit the most each step, never removing one. Backward elimination starts with all candidates and drops the least useful each step, never adding. SPSS Stepwise combines both, so a variable can enter and later be removed. When forward and backward end with different models, your data do not identify one set of predictors.

Why is stepwise regression criticized?

The search uses the same data you later test on, which inflates R-squared and makes the final p-values invalid as hypothesis tests. The retained variables are unstable, so small data changes produce different models. With collinear predictors the choice is close to arbitrary. Most methodologists treat stepwise as exploratory screening that must be confirmed on held-out data, not as a way to decide which variables matter.

What alpha should I use for stepwise regression?

SPSS defaults to 0.05 to enter and 0.10 to remove, and the removal value must exceed the entry value. Setting both to 0.05 keeps more variables and produces a larger, more stable model. Going below 0.01 makes the search stingy and reduces how much the fit statistic is inflated, but it can discard real effects in a modest sample. None of these settings repair the underlying problems.

How many observations do I need for stepwise regression?

Ordinary multiple regression needs roughly 15 to 20 complete cases per candidate predictor as a working floor. Stepwise needs more, because every candidate is tested repeatedly and the selected set must survive out-of-sample checking. With 30 candidates, plan for several hundred complete cases. If your rows fall short of the number of parameters, use lasso or ridge instead, since stepwise cannot run at all.

Conclusion

Run stepwise when you need to see how a modest set of candidates behaves, then treat what comes back as a shortlist rather than an answer. Rerun the surviving variables with Enter, check VIF and the residuals, compare forward against backward, and see whether the model predicts new rows better than the full one. Write down the criteria, the removals, and the validation result. That last part is what turns a quick SPSS run into something a reviewer will accept in 2026.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides