How to Run Post Hoc Tests After ANOVA (2026)

A post hoc test is a multiple-comparison procedure you run after a statistically significant ANOVA to find out which specific pairs of group means differ, while keeping the family-wise error rate at your chosen alpha level. In practice: run the omnibus F-test, check the assumptions, then apply a correction procedure such as Tukey’s HSD. Ten minutes of work in SPSS or R, and the result is the pairwise table your results section needs.

Most of the confusion I see with post hoc testing comes from one thing — treating it as a separate analysis. It is not. It is a follow-up question asked of data you have already analysed, and it inherits your ANOVA’s assumptions whether you like it or not.

How to run post hoc tests after ANOVA is a short sequence once you know it. This guide walks through the whole of it: the prerequisites, the seven-step procedure, the exact menu path and code for SPSS, R, Python, Stata and Excel, how to read the output table, and how to write the result in APA 7 style. It also covers two-way ANOVA, where “which groups differ” is the wrong question and simple effects analysis is the right one.

Table of Contents
  1. 1What You Need
  2. 2Step-by-Step
  3. 31. Confirm That the ANOVA Is Significant
  4. 42. Check the Assumptions Behind the ANOVA
  5. 53. Choose the Correct Post Hoc Test
  6. 64. Run the Selected Test in SPSS
  7. 75. Run the Selected Test in R, Python or Stata
  8. 86. Interpret Pairwise Differences
  9. 97. Report the Results in APA Style
  10. 10Common Mistakes
  11. 11Frequently Asked Questions
  12. 12When should I run post hoc tests after ANOVA?
  13. 13Which post hoc test should I use after a significant ANOVA?
  14. 14Is Tukey HSD appropriate when group variances are unequal?
  15. 15What is the difference between Bonferroni and Tukey post hoc tests?
  16. 16How do I report post hoc ANOVA results in APA style?
  17. 17Can I run post hoc tests after a non-significant ANOVA?
  18. 18Conclusion

What You Need

Four things have to be true before a post hoc comparison is worth running.

  1. The omnibus ANOVA is significant. A non-significant F-test means the data give you no basis for comparing pairs, and running comparisons anyway inflates the false-positive rate.
  2. Your factor has three or more levels. With two groups you run a t-test, not a post hoc procedure.
  3. You still have the raw data, not just the output. Every post hoc method needs group means, standard deviations and sample sizes. If all you have is the ANOVA table, you cannot reconstruct them.
  4. You have run at least one assumption check. Normality of residuals, homogeneity of variance, independence, and a look for outliers.

There are four post hoc families you will meet constantly, and the choice among them turns on two things: whether group variances are equal, and whether you are comparing every pair or only comparing each group against a control.

  • Tukey’s HSD — the default when variances are equal and you want every pairwise comparison.
  • Bonferroni — simple, works in every package, more conservative than Tukey.
  • Games-Howell — the unequal-variance option, and the right partner for Welch’s ANOVA.
  • Dunnett’s test — compares each group to a single control group only.

Why a correction matters at all is easy to show. Every pairwise comparison you run carries a nominal 5 percent chance of a false positive. Stack those and the family-wise error rate climbs fast.

GroupsPairwise comparisonsFamily-wise error rate at alpha = .05
215.0%
3314.3%
4626.5%
51040.1%
61553.7%

With six groups, unadjusted t-tests give you better than even odds of reporting at least one difference that is not really there. That is why every post hoc method exists.

Step-by-Step

1. Confirm That the ANOVA Is Significant

Look at the ANOVA table and find the row for your factor. The Sig. column (SPSS) or the Pr(>F) column (R) holds the p-value for the omnibus F-test. If it is below .05, you proceed. If it is not, stop.

A non-significant omnibus test is a statement that your data do not support any claim of group differences. Running a post hoc test anyway is the single most common error reviewers flag in submitted manuscripts. It does not create information; it manufactures false positives from noise.

There is one legitimate exception: if the post hoc comparison was planned before you looked at the data, it is a planned contrast and does not need the omnibus test to be significant. Label it as planned, not post hoc, and report it that way.

2. Check the Assumptions Behind the ANOVA

This is the step people skip, and it is the step that decides your test. Four checks matter.

Residual normality. In SPSS, tick Options > Normality/Equality of Variances and the Normal Q-Q plot of residuals. In R, car::qqPlot(residuals(model)). With fewer than about 20 observations per group, look at the plot rather than the Shapiro-Wilk p-value, which is unhelpful at small n.

Homogeneity of variance. Levene’s test is the standard check, in SPSS under Options > Homogeneity of variance test, or car::leveneTest(y ~ group) in R. A significant result (p < .05) means the equal-variance assumption is not supported, and Tukey’s HSD is no longer the right call.

Independence. Check that observations are independent — one participant contributes one score, and clustered or repeated data needs a mixed model instead.

Outliers. Look at a boxplot by group. A single extreme value can distort a mean badly in a small group.

When variances are clearly unequal, or sample sizes differ a lot, use Games-Howell. When residuals are badly non-normal and n per group is small, the non-parametric path applies: Kruskal-Wallis first, then Dunn’s test with a Bonferroni or Holm adjustment.

3. Choose the Correct Post Hoc Test

Use this table to pick the procedure. If you are unsure how to run post hoc tests after ANOVA with your own data, start here — it is the decision most readers need, so it is worth keeping.

Your situationUse this testWhy
Equal variances, all pairs, balanced groupsTukey’s HSDGood balance of power and error control for all-pair comparisons
Equal variances, all pairs, very unequal nGames-Howell, or Tukey-KramerTolerates unequal group sizes better
Levene’s test significant, or Welch’s ANOVA usedGames-HowellDoes not assume equal variances
One control group, several treatment groupsDunnett’s testOnly the comparisons you planned are tested, so power stays high
Many groups, controlling Type I error strictlyBonferroni or HolmSimple, conservative, available everywhere
Non-normal residuals, small groups, ordinal outcomeKruskal-Wallis, then Dunn’s testRank-based alternative to the whole family
One factor, you only need 2 or 3 specific contrastsPlanned contrasts / polynomial contrastsAvoids the multiple-comparison penalty for unplanned tests

If you are unsure, Tukey’s HSD after a Levene’s-clean ANOVA is the defensible default. Judges, reviewers and examiners recognise it, and it is close to Bonferroni in leniency for small numbers of groups.

4. Run the Selected Test in SPSS

SPSS runs the omnibus test and the post hoc in one dialog for the one-way case.

  1. Go to Analyze > Compare Means > One-Way ANOVA.
  2. Move your dependent variable (the score) into Dependent List and your grouping variable into Factor. Set the number of levels if the group codes are not 1, 2, 3.
  3. Click Options. Tick Descriptives, Homogeneity of variance test, Welch’s test and Effect size estimates. Add Normality/Equality of Variances for the residual diagnostics. Click OK.
  4. Click the Post Hoc button. Under Equal variances assumed, tick Tukey. Under Equal variances not assumed, tick Games-Howell. For many groups you may add Bonferroni, Sidak or Scheffe. Click OK.
  5. Click OK on the main dialog and read the output.

Click Paste instead of OK and SPSS drops the equivalent syntax into a syntax window, which is worth doing once you want a reproducible record. The one-way command looks like this:

ONEWAY score BY group
  /STATISTICS DESCRIPTIVES HOMOGENEITY LEVENE=WELCH
  /POSTHOC = TUKEY ALPHA(.05)
  /EMMEANS=TABLES(group)
  /PRINT.

If Levene’s test comes back significant, go back and tick Games-Howell instead of Tukey. Doing the check and then ignoring it is the second most common error I see.

5. Run the Selected Test in R, Python or Stata

R, Tukey’s HSD. Base R covers most of what you need:

fit <- aov(score ~ group, data = df)
summary(fit)
TukeyHSD(fit)                       # all pairs, Tukey adjusted
plot(TukeyHSD(fit))

Bonferroni, Games-Howell, Levene’s and effect sizes:

car::leveneTest(score ~ group, data = df)   # homogeneity check
pairwise.t.test(df$score, df$group,
                p.adjust.method = "bonferroni")
rstatix::gameshowell_test(score ~ group, data = df)
rstatix::dunn_test(score ~ group, data = df)  # after Kruskal-Wallis

For anything with covariates, factors of more than one level, or a mixed model, use estimated marginal means. This is the standard route for repeated measures and mixed designs:

library(emmeans)
m <- lmer(score ~ condition * time + (1 | participant), data = df)
emmeans(m, ~ condition | time) |> pairs(adjust = "tukey")

Python. scikit-posthocs is the quickest route, and statsmodels gives you Tukey directly:

import statsmodels.api as sm
from statsmodels.stats.multicomp import pairwise_tukeyhsd
import scikit_posthocs as sp

sm.stats.f_oneway(*[g["score"] for _, g in df.groupby("group")])
pairwise_tukeyhsd(df["score"], df["group"])

sp.posthoc_all(df, group_col="group", valu_col="score",
                p_adjust="holm")          # bonferroni / fdr_bh also accepted
sp.posthoc_gameshowell(df, val_col="score", group_col="group")

Stata. The omnibus test, then pwcompare for the pairs:

oneway score group, tabulate
pwcompare score group, sidak() tukey() bonferroni()
pwcompare score group, duncan() scheffe()
pwcompare score, group(group) idn(3)      // each group vs control
pwcompare score group, holm()

Stata’s oneway already reports Scheffe and Fisher’s LSD by default. Use pwcompare when you need Tukey, Bonferroni, Holm or Dunn.

Excel. There is no post hoc button. The Data Analysis ToolPak gives you the one-way ANOVA table (with the F and p-value) and nothing more, so you calculate the comparisons yourself:

  1. Run Data > Data Analysis > ANOVA: Single Factor to get the ANOVA output, then read Within-Group Sum of Squares and Within-Group df.
  2. Compute MSwithin = SSwithin / dfwithin.
  3. For equal group sizes, HSD = qalpha(k, N − k) × √(MSwithin / n), where n is the size of one group and q is the critical value of the studentised range distribution.
  4. Compare each mean difference to HSD. Anything larger differs at alpha.

For unequal group sizes, switch to the Tukey-Kramer formula with the √(MSwithin/2ni + MSwithin/2nj) term instead. For a class project this is fine; for a thesis, use real software, because you also get adjusted p-values and confidence intervals this way.

6. Interpret Pairwise Differences

Open the post hoc table and read five columns. The output is usually laid out this way, whether the label is SPSS, R or Stata.

ColumnWhat it tells you
Group pair (I, J)Which two means are being compared
Mean difference (I – J)Direction and size of the gap; the sign shows which group is higher
Std. errorPrecision of that difference
Sig. adjustedThe p-value after correction. Compare this to .05, never the raw p
95% confidence intervalThe plausible range for the true mean difference; if it excludes 0, the difference is significant
Sig. (Bonferroni) etc.Comparison-specific adjusted p-values when several were requested at once

In SPSS you will also see Homogeneous Subsets, and it is the most misread block in the whole output. Read it as columns, not rows: groups in the same column share a letter and are not significantly different from each other. Groups in different columns differ. A group appearing in two columns sits between the others.

A worked example makes the pattern obvious. Suppose three teaching methods were compared and Tukey’s HSD returned:

  • Method A (M = 72.4, SD = 8.1) vs Method B (M = 69.8, SD = 7.4): mean difference = 2.6, adjusted p = .312, 95% CI [−5.1, 10.3].
  • Method A vs Method C (M = 61.2, SD = 9.4): mean difference = 11.2, adjusted p = .004, 95% CI [4.9, 17.5].
  • Method B vs Method C: mean difference = 8.6, adjusted p = .017, 95% CI [2.5, 14.7].

Now read it correctly. C differs from both A and B. A and B do not differ from each other. In SPSS’s homogeneous subsets that appears as two columns — {A, B} and {C} — which is exactly the right conclusion. The mistake is reading the row instead of the column and reporting that “A, B and C were significantly different”.

Add effect size for each significant pair where it matters. A mean difference with a 95 percent confidence interval, or Cohen’s d for the pair, tells the reader whether the difference is practically meaningful. A difference of 2.6 points flagged as significant in a group of 300 is arithmetic, not a finding.

For a two-way ANOVA with a significant interaction, none of the above applies. The omnibus test told you the effect of one factor depends on the level of the other, so you split the data: hold one factor at each level in turn and run simple effects analysis on the other. In SPSS this is Analyze > General Linear Model > Univariate with Options > Estimated Marginal Means > Compare main effects. In R it is emmeans(m, ~ factor2 | factor1) |> pairs(). Compare the simple effects, not the marginal means, and state which factor you held constant.

7. Report the Results in APA Style

APA 7 wants the omnibus test first, then the follow-up by name, with each significant comparison carrying its own numbers.

Fill-in-the-blank template:

A one-way ANOVA showed a [significant / non-significant] effect of [factor] on [dependent variable], F([df1], [df2]) = [F], p = [p], η2 = [η²]. [Post hoc test name] indicated [significant / no significant] differences between [group A] (M = [mean], SD = [sd]) and [group B] (M = [mean], SD = [sd]), [adjusted p = [p]], 95% CI [[lower], [upper]]. [Brief one-sentence interpretation of the pattern.]

Complete example, using the data above:

A one-way ANOVA indicated that teaching method had a significant effect on test score, F(2, 57) = 8.42, p = .001, η2 = .23. Tukey’s HSD found that Method C (M = 61.2, SD = 9.4) differed significantly from both Method A (M = 72.4, SD = 8.1; adjusted p = .004, 95% CI [4.9, 17.5]) and Method B (M = 69.8, SD = 7.4; adjusted p = .017, 95% CI [2.5, 14.7]). Method A and Method B did not differ significantly (adjusted p = .312).

Four reporting rules that keep you out of trouble. Report the adjusted p-value, with the word “adjusted” or the test name attached, so nobody reads it as a raw p. State the number of comparisons the correction covered. Report non-significant comparisons when they matter to your argument — as in Method A versus Method B above. And name the test you used; “post hoc tests were conducted” is not acceptable in APA 7.

Common Mistakes

MistakeWhy it is wrongFix
Running post hoc after a non-significant ANOVAManufactures false positives and is routinely rejected in peer reviewReport the omnibus result and stop, or state that follow-ups were pre-planned contrasts
Using Tukey when Levene’s test is significantTukey assumes equal variances; with unequal n it loses real power and mis-sets the error rateSwitch to Games-Howell, or run Welch’s ANOVA and pair it with Games-Howell
Reporting raw p-values from the pairwise tableThe reader cannot tell whether a correction was appliedReport the adjusted p-value and name the procedure
Running several corrected post hoc tests on the same dataEach correction assumes it is the only one; stacking them is incoherent and over-conservativePick one procedure per analysis, or use emmeans to request one adjustment family consistently
Reading homogeneous subsets by rowLeads to claiming differences that do not existRead down the columns; same column means no significant difference
Following a significant interaction with marginal meansMarginal means average across levels and can hide the very effect that was significantRun simple effects analysis, holding one factor constant, and compare within each level
Stopping at the omnibus F-testANOVA only says at least one mean differs, so the reader learns nothing actionableAlways pair it with a named pairwise procedure and adjusted p-values
Significance with no magnitudeA tiny effect in a large sample is statistically significant and practically meaninglessAdd mean differences with confidence intervals, and effect size where the decision matters

A short checklist before you submit: omnibus test significant, assumption check run and acted on, one correction method named, adjusted p-values reported, effect size or confidence interval included, every H2 claim backed by a specific pair.

One more route worth knowing, if your data are ordinal or the residuals look bad with small groups. The Kruskal-Wallis test replaces the one-way ANOVA, and Dunn’s test replaces the post hoc step, with a Bonferroni or Holm adjustment for the ranks. The logic is identical: omnibus first, then corrected pairs.

Frequently Asked Questions

When should I run post hoc tests after ANOVA?

Run post hoc tests only after the omnibus one-way ANOVA is statistically significant, and only when your factor has three or more levels. The omnibus test tells you at least one group mean differs; the post hoc procedure tells you which. Running pairs after a non-significant F-test inflates the family-wise error rate and is the most common reason reviewers send a paper back.

Which post hoc test should I use after a significant ANOVA?

Tukey’s HSD is the default: equal variances, balanced groups, and you want every pairwise comparison. Switch to Games-Howell when Levene’s test is significant or group sizes differ sharply. Use Dunnett’s test when every group is compared against a single control, and Bonferroni or Holm when you want a conservative correction available in any package.

Is Tukey HSD appropriate when group variances are unequal?

No, not on its own. Tukey’s HSD assumes equal variances across groups, and with unequal group sizes the error rate drifts away from your chosen alpha. Use Games-Howell, which is designed for that case and is built into SPSS under Equal variances not assumed. If variances are strongly unequal, also consider running Welch’s ANOVA first and pairing it with Games-Howell.

What is the difference between Bonferroni and Tukey post hoc tests?

Both correct for the number of comparisons, but Bonferroni simply multiplies each p-value by the number of tests, which is conservative and becomes wasteful as groups multiply. Tukey’s HSD uses a studentised range distribution, so it is less conservative at small numbers of groups while still controlling the family-wise error rate. In practice the two often agree, and Tukey is preferred for balanced designs.

How do I report post hoc ANOVA results in APA style?

Report the omnibus ANOVA first with F, its degrees of freedom, the exact p-value and an effect size such as eta squared. Then name the follow-up procedure and give each significant pair its group means, standard deviations, adjusted p-value and 95 percent confidence interval for the mean difference. Report non-significant comparisons that your argument depends on, and label p-values as adjusted.

Can I run post hoc tests after a non-significant ANOVA?

Generally no. If the omnibus F-test is not significant, there is no statistical basis for comparing pairs, and doing so turns noise into apparent findings. The one exception is a comparison you specified in advance as a planned contrast, which is reported as a contrast rather than a post hoc test. Some journals and examiners allow exploratory work if it is labelled clearly as such.

Conclusion

Start with the first action: confirm the omnibus ANOVA is significant, then run the assumption check that decides everything — Levene’s test. Equal variances and you want every pair means Tukey’s HSD. Levene’s is significant and you need Games-Howell. One control group and you need Dunnett’s.

Then report the pairwise results rather than the F-test alone: group means, adjusted p-values, confidence intervals, and a plain sentence saying which groups differed. That is a result section a reader can actually use.

Leave a Comment

Practical guides to statistics, surveys and research data

Read the latest guides