Model fit indices in CFA tell you how closely the covariance structure your hypothesized factor model implies reproduces the covariance matrix your sample actually produced. You read them as a pattern, not a scoreboard: the model chi-square and RMSEA judge absolute fit, CFI and TLI compare your model against a null model where nothing is correlated, and SRMR summarises the average size of the leftover discrepancies.
Most students learn one cutoff, usually CFI above .95, and then stop. That is the fastest way to misread a fit table, because indices disagree for understandable reasons, and the disagreement is itself information about your model.
This guide walks through what each index measures, where its thresholds come from, how to read a real output block line by line, and how to write the results paragraph when fit is mixed.
Table of Contents
- 1What CFA Model Fit Indices Measure
- 2Fit is not validity
- 3Why indices are judged as a pattern
- 4Absolute Fit Indices: How Well the Model Reproduces the Data
- 5The model chi-square (CMIN) and CMIN/DF
- 6What does RMSEA stand for, and what is a good RMSEA score?
- 7SRMR
- 8The older indices: GFI, AGFI, NFI, RFI, IFI, PNFI
- 9Incremental Fit Indices: Does the Model Improve on a Simpler Model
- 10CFI, in plain terms
- 11TLI, IFI and the rest
- 12How to Interpret Model Fit Indices in CFA Without Overclaiming
- 13Common CFA Model Fit Patterns and What They Suggest
- 14How to Report CFA Fit Indices in a Paper or Thesis
- 15Frequently Asked Questions
- 16What are the most important CFA model fit indices?
- 17Are RMSEA, CFI, SRMR, and chi-square enough to judge a CFA model?
- 18Why do CFA fit indices conflict or fall into different fit categories?
- 19What sample size is needed for reliable CFA fit indices?
- 20Does acceptable CFA fit prove that my factor model is correct?
- 21Should I modify a CFA model when the fit indices are poor?
- 22Conclusion
What CFA Model Fit Indices Measure
Fit is a comparison between two covariance matrices. The first is your sample covariance matrix, computed from the raw responses. The second is the covariance matrix implied by the model you specified, reconstructed from the factor loadings, factor correlations and indicator residual variances.
The gap between those two matrices is the discrepancy the fit indices summarise. A perfect model would imply exactly the observed matrix; every real model leaves some gap, and the indices are different ways of scaling and weighting that gap.
Fit is not validity
Fit tells you the structure is compatible with the data. It does not tell you the factors are the right conceptual explanation, that the items were worded well, or that the model will replicate. A mislabelled but internally consistent structure fits beautifully and means something else entirely.
Why indices are judged as a pattern
Each index weights the discrepancy differently, and each was calibrated on simulation studies with particular models and sample sizes. When they disagree, no index outranks the others automatically. The convention most reviewers accept, following Hu and Bentler (1999) and MacCallum, Cheung and Rensvold (1992), is that the chi-square test, RMSEA, CFI and SRMR together give a defensible read, and that any single index in isolation does not.
Absolute Fit Indices: How Well the Model Reproduces the Data
Absolute fit indices evaluate the observed-versus-implied discrepancy directly, with no comparison model. They answer one question: how far off is this model, on its own terms?
The model chi-square (CMIN) and CMIN/DF
The chi-square statistic tests the null hypothesis that the population covariance matrix equals the matrix implied by your model. Significant means the model does not reproduce the data perfectly, which in a large sample is almost the default outcome. Chi-square rises with sample size, so a significant result at N = 800 with a large model is weak evidence of a genuinely wrong structure.
The CMIN/DF ratio divides the statistic by degrees of freedom to partially absorb that sensitivity. Values below about 3 are usually treated as unremarkable, and values near or above 5 point to a real mismatch. Ratio rules of thumb are crude, but they are more informative than the raw p-value alone.
What does RMSEA stand for, and what is a good RMSEA score?
RMSEA is the root mean square error of approximation: an estimate of the average discrepancy between observed and implied covariance, per degree of freedom, expressed on a 0-to-1 scale where 0 is perfect fit. The commonly cited thresholds are .06 or below for acceptable fit and .05 or below for close fit (Hu and Bentler, 1999), with .05 also proposed by Byrne (1994).
Always report the 90% confidence interval with it. An RMSEA of .058 with a 90% CI of [.041, .075] is a genuinely borderline result; the same point estimate with a CI of [.038, .079] from a small sample says less. Many programs also print a p value for close fit, which tests H0: RMSEA ≤ .05. A close-fit p below .05 means you can reject that narrow claim, which is not the same as concluding the model fails.
SRMR
The standardised root mean square residual is the average size of the standardised differences between observed and implied covariances, with .08 or below treated as good (Hu and Bentler, 1999) and .05 as a stricter target used by some. Because it reads like a root mean square, it is intuitive, and it tends to react less than RMSEA to small samples. It is the easiest index for a committee member to interpret, which is why it has become a default report.
The older indices: GFI, AGFI, NFI, RFI, IFI, PNFI
The goodness-of-fit index and its adjusted version penalise model complexity but are known to reward over-parameterised models, so most current guidance de-emphasises them. The normed fit index compares the model to the independence model without adjusting for complexity, which makes it unstable in small samples. Its non-normed sibling TLI fixes that problem, the relative fit index and incremental fit index are further variants, and the parsimony-adjusted PNFI divides by degrees of freedom. You may need them for a legacy instrument whose published validity evidence reports them, but new work rarely needs more than the core four plus chi-square.
| Index | What it measures | Good fit | Common cutoff | Source |
|---|---|---|---|---|
| CMIN chi-square | Total discrepancy, tested against perfect fit | Not significant, or small relative to df | p > .05; CMIN/DF under 3 | General |
| RMSEA | Average per-df approximate discrepancy | Lower is better | ≤ .06 acceptable, ≤ .05 close | Hu and Bentler 1999; Byrne 1994 |
| SRMR | Average standardised residual covariance | Lower is better | ≤ .08 good, ≤ .05 strict | Hu and Bentler 1999 |
| CFI | Gain over the independence model | Higher is better | ≥ .95 strict, ≥ .90 lenient | Bentler 1990; Hu and Bentler 1999; Byrne 1994 |
| TLI | Non-normed variant of CFI | Higher is better | ≥ .95 strict, ≥ .90 lenient | Tucker and Lewis 1973 |
| RMR | Average raw residual covariance | Lower is better | ≤ .08 | Joreskog and Yarnold 1990 |
| GFI / AGFI | Variance accounted for, complexity-adjusted | Higher is better | ≥ .90 / ≥ .80 | Legacy, now discouraged |
| NFI / RFI / IFI | Normed, relative and incremental variants of CFI | Higher is better | ≥ .90 | Bentler and Bonett 1980; Bollen 1986 |
| PNFI | Parsimony-adjusted incremental fit | Higher is better | No universal rule | Parsimony adjustment |
If you are running this in R, the lavaan package returns these as named values, and a single call is enough to pull the whole set:
fit <- cfa(model, data = df)
fitmeasures(fit, c("chisq", "df", "pvalue", "cfi", "tli",
"rmsea", "rmsea.ci.upper", "srmr"))
In AMOS the same values appear under Calculate, Estimate, Model Fit, Model Fit Summary, with CMIN, RMR and GFI in one block and the baseline comparison block beneath it.
Incremental Fit Indices: Does the Model Improve on a Simpler Model
Incremental or comparative indices do not judge your model in isolation. They measure how much better your model does than a baseline model, normally the independence model in which all observed variables are uncorrelated. That baseline is deliberately terrible, which is why incremental indices tend to look generous, especially when the variables in your model are strongly intercorrelated and the sample is large.
CFI, in plain terms
The comparative fit index takes the improvement your model makes over the baseline and scales it between 0 and 1, where 1 means your model reproduces the data as well as the data reproduce themselves. Values at or above .95 are conventionally called good, and .90 or above is often accepted, particularly for models with many parameters. Cheung and Rensvold warned that CFI rises toward 1 as sample size grows, so a .99 with N = 60 tells you less than a .96 with N = 400.
TLI, IFI and the rest
TLI, also called NNFI, adjusts the increment for degrees of freedom and behaves better than NFI in small samples; it sits a little below CFI, so seeing TLI at .94 when CFI is .96 is normal rather than contradictory. IFI and RFI are further variants that are rarely needed in modern reporting. All of them answer the same underlying question: is your factor structure doing real work, or would uncorrelated variables explain the data about as well?
How to Interpret Model Fit Indices in CFA Without Overclaiming
- Check the estimator and data type first. If your items are ordinal Likert categories, the default maximum likelihood output assumes continuous, normally distributed variables. Switch to categorical WLS (WLSMV in Mplus,
estimator = "WLSMV"in lavaan, or DWLS with robust corrections in AMOS) and read the scaled or shifted indices instead of the unscaled ones. - Read degrees of freedom. Zero degrees of freedom means a saturated model, which fits perfectly by construction and tells you nothing. Negative degrees of freedom mean an under-identified or misspecified model. Neither belongs in a paper without explanation.
- Read the chi-square with its ratio. Note the statistic, the df, the p value and CMIN/DF together. Decide whether the significance is an artefact of sample size or tracks a poor ratio.
- Read RMSEA with its 90% confidence interval. A point estimate inside a cutoff whose interval extends well past it is a borderline result, and you should describe it as borderline.
- Read CFI and TLI as a pair. Report both, and check them against sample size. Very high values in a small sample deserve a caveat.
- Read SRMR last as a cross-check. When it agrees with RMSEA and CFI, you can state that fit is acceptable. When it disagrees, look at local fit before concluding anything.
- Look beyond global fit. Global indices are averages. Inspect the standardised residual covariance matrix and look for a handful of large residuals, which point to specific mis-specified paths rather than a model that is wrong everywhere.
- Compare nested models where the theory allows it. When one model nests inside another, use the chi-square difference test or the information criteria (AIC and BIC, which penalise complexity) rather than eyeballing indices.
- State the limitation. Fit was evaluated with a maximum likelihood estimator on continuous indicators; results may differ under categorical estimation.
Common CFA Model Fit Patterns and What They Suggest

Conflicting indices are not a sign that your software is broken. They are a symptom, and the symptom tells you where to look next.
| Pattern | Likely cause | What to check |
|---|---|---|
| CFI .97, RMSEA .09 | Small sample, or a highly correlated item set where the baseline model is extremely poor | df and N; RMSEA 90% CI upper bound; whether items are ordinal |
| RMSEA .05, CFI .90 | Highly parsimonious model with few parameters, so the baseline comparison looks weak | Model complexity, item count per factor, TLI for confirmation |
| Chi-square significant, all other indices acceptable | Large sample detecting a small discrepancy | CMIN/DF; report chi-square with the other indices instead of on its own |
| All indices good, one or two huge residuals | A single misspecified cross-loading or indicator wording problem | Residual covariance matrix; modification indices |
| Fit improves sharply after item removal | The removed item was not measuring the construct, or was worded double-barrelled | Content validity before fit, not after |
| Everything acceptable but loadings weak or a factor with two indicators | Structural underidentification rather than misfit | Identification constraints; fix a factor loading or correlate residuals |
Two questions come up constantly on Cross Validated and r/AskStatistics. The first is whether using modification indices to fix poor fit turns a confirmatory analysis into an exploratory one. The honest answer is that it moves you that way, and if you modify enough to reach good fit you should say so plainly in your limitations, or treat the model as exploratory and describe it that way. The second is how many participants you need. There is no single answer, but Hu and Bentler’s simulation work suggested that around 250 cases gives adequate recovery for three-indicator factors, with the requirement dropping as the number of indicators per factor rises. A widely used starting point is 5 to 10 cases per free parameter, which is a planning heuristic rather than a rule.
How to Report CFA Fit Indices in a Paper or Thesis
Kline’s minimum report set is the safest baseline: the chi-square with its df and p value, the CFI, the RMSEA with its 90% confidence interval, and the SRMR. Add TLI if your field expects it, and add the estimator to all of it.
A results paragraph that reads defensibly looks like this, with the bracketed parts filled from your output:
The confirmatory factor analysis indicated that the [number]-factor model fit
the data acceptably, chi-square = [chisq], df = [df], p = [p], CFI = [cfi],
TLI = [tli], RMSEA = [rmsea], 90% CI [[lower], [upper]], SRMR = [srmr].
The chi-square was [significant/not significant]. Because the chi-square
statistic is sensitive to sample size ([N] participants), the fit indices were
interpreted as a set. Fit was estimated using [ML/categorical WLS] with [no
imputation/listwise deletion].
Then put the same numbers in an APA-style table with index, estimate, threshold and source as columns, so a reader can check your claims against the same rules you used. If a value misses its threshold, say so in the sentence rather than in a footnote: “RMSEA = .071, slightly above the .06 threshold, indicating modest misfit in the [scale name] subscale.” That sentence survives peer review; a footnote hoping nobody looks does not.
For the same reason, when you test measurement invariance across groups or time points, report the change scores rather than the raw values: a drop of .010 or less in CFI, .015 or less in RMSEA, and .03 or less in SRMR is the usual change-score rule of thumb, with a tighter SRMR standard for scalar invariance.
Frequently Asked Questions
What are the most important CFA model fit indices?
The four that matter most are the model chi-square with its degrees of freedom, the CFI, the RMSEA with its 90% confidence interval, and the SRMR. TLI is a useful fifth because it is less affected by small samples. GFI, AGFI, NFI, RFI, IFI and PNFI still appear in older published studies, so you may need them when comparing your results to an existing instrument, but they are not required for a new analysis.
Are RMSEA, CFI, SRMR, and chi-square enough to judge a CFA model?
For most theses and journal submissions, yes, that set is what reviewers expect. Report the chi-square with its df, p value and ratio, RMSEA with its confidence interval, CFI and SRMR, and state the estimator you used. Add TLI when your area of study consistently reports it, and add information criteria if you compared competing structures rather than testing one against a baseline.
Why do CFA fit indices conflict or fall into different fit categories?
Each index weights the model-versus-data discrepancy differently, and each was calibrated on different simulation conditions. RMSEA tends to be conservative with few degrees of freedom, while CFI and TLI become generous as sample size grows, especially when observed variables correlate strongly so the baseline model looks hopeless. Divergence usually points to a small sample, a parsimonious model, or a genuinely awkward stretch of the model, and the pattern itself is worth reporting.
What sample size is needed for reliable CFA fit indices?
It depends on the number of indicators and free parameters rather than on one threshold. Hu and Bentler (1999) found roughly 250 cases were needed to recover parameters reliably for three-indicator factors, with fewer needed as indicators per factor increased. As a planning rule, aim for 5 to 10 cases per free parameter and never fewer than about 100, because below that the chi-square loses power and the comparative indices become unstable.
Does acceptable CFA fit prove that my factor model is correct?
No. Fit only shows that the observed covariance pattern is compatible with the structure you specified. Several different structures can fit the same data acceptably, and a well-fitting model may still mislabel its factors or miss an equivalent alternative. Fit is necessary evidence for a structural claim, not sufficient evidence, which is why convergent and discriminant validity evidence belong alongside it.
Should I modify a CFA model when the fit indices are poor?
Look at the content first. Poor wording, a double-barrelled item, a cross-loading or an influential outlier explains more misfit than any respecification, and fixing those is ordinary data preparation. Use modification indices and Lagrange multiplier output to locate the problem, then change the model only when theory supports the change, one at a time, and re-checking fit after each. If you end up with many changes, describe the analysis as exploratory and say so in your limitations.
Conclusion
Start with the whole pattern rather than a single number: check the estimator, then the degrees of freedom, the chi-square and its ratio, RMSEA with its confidence interval, CFI with TLI, and SRMR as a cross-check. If they agree, describe the fit as acceptable and report it. If they disagree, treat the disagreement as a finding about sample size, model complexity or estimator choice, and say which explanation you checked. Whatever you conclude, write the limits into the paper itself.


