You just ran your confirmatory factor analysis. The output is dense. Dozens of numbers stare back at you. CFI equals something. RMSEA has a confidence interval. SRMR is a decimal. Now what? Learning how to read model fit indices in confirmatory factor analysis is one of the most important skills you will develop as a researcher in 2026. These numbers determine whether your measurement model passes scrutiny from reviewers, advisors, and journal editors. Without adequate fit, your latent variables remain unvalidated and your subsequent structural paths are untrustworthy. A measurement model with poor fit produces biased parameter estimates, misleading standard errors, and ultimately conclusions that cannot be defended in a peer-reviewed publication. Researchers who skip fit evaluation and proceed directly to testing structural hypotheses risk building their entire argument on a shaky foundation.
I have helped researchers interpret CFA output from lavaan, Mplus, AMOS, and SPSS for over five years. I have seen bright graduate students panic over a single RMSEA value that exceeded 0.08. I have watched experienced researchers overlook SRMR and miss critical misfit that CFI concealed. The most common question I receive is simple: are these numbers good enough? The answer is rarely a simple yes or no. In this guide, I walk through every major fit index you will encounter. I explain what each one measures, what the accepted thresholds are, and what to do when indices disagree with each other. I also cover the decision framework I use when fit results are ambiguous, which is more often than you might expect. By the end, you will read a CFA output file with confidence.
Table of Contents
Quick Reference: CFA Fit Index Cutoffs at a Glance
Before diving into each index, here is a consolidated view of the major fit indices and their commonly accepted thresholds. Bookmark this table. You will refer back to it constantly while interpreting CFA output in 2026. I have included the primary sources for each threshold so you can trace them back to the original methodological literature.
| Index | Full Name | Range | Poor Fit | Acceptable | Good Fit | Key Source |
|---|---|---|---|---|---|---|
| CFI | Comparative Fit Index | 0 to 1 | < 0.90 | 0.90 – 0.95 | ≥ 0.95 | Hu & Bentler (1999) |
| TLI | Tucker-Lewis Index | 0 to 1+ | < 0.90 | 0.90 – 0.95 | ≥ 0.95 | Tucker & Lewis (1973) |
| RMSEA | Root Mean Square Error of Approximation | ≥ 0 | > 0.10 | 0.05 – 0.08 | ≤ 0.05 | Browne & Cudeck (1992) |
| SRMR | Standardized Root Mean Square Residual | 0 to 1 | > 0.10 | 0.05 – 0.08 | ≤ 0.05 | Hu & Bentler (1999) |
| Chi-square | Chi-square test of exact fit | 0 to ∞ | Significant (small p) | Mixed | Non-significant | Joreskog (1969) |
Keep these ranges in mind as we explore each index in detail. The Hu and Bentler guidelines are the most widely cited in journal submissions, but I will also discuss when to apply them flexibly and how different conditions such as sample size and data type can shift these thresholds. Understanding the reasoning behind each cutoff matters more than memorizing the numbers.
How to Read Confirmatory Factor Analysis Results: Starting With the Chi-Square Test
The chi-square test of exact fit is the original and most fundamental index in structural equation modeling. It tests whether the observed covariance matrix differs significantly from the model-implied covariance matrix. A non-significant chi-square indicates that your model reproduces the observed data well. A significant chi-square suggests misfit. This sounds straightforward, and in principle it is. In practice, chi-square is the most misunderstood fit index because of its extreme sensitivity to sample size.
Here is the catch. The chi-square test is extremely sensitive to sample size. With samples larger than 200, even trivial discrepancies between your model and the data produce a significant result. This means chi-square is almost always significant in real research conducted with adequate sample sizes. Most methodologists recommend treating a significant chi-square as expected rather than alarming when you have a large sample. The chi-square to degrees of freedom ratio offers one alternative, with values below 3 suggesting acceptable fit and below 2 suggesting good fit. This guideline dates back to Wheaton and colleagues in 1977 but has since been critiqued as overly simplistic and not well grounded in formal statistical theory.
Degrees of freedom in CFA are calculated as the difference between the number of unique pieces of information in your observed covariance matrix and the number of free parameters in your model. Understanding degrees of freedom helps you interpret chi-square because a model with many indicators and few factors will have many degrees of freedom, making the chi-square test more powerful and more likely to detect even trivial misfit. A model with few indicators relative to its factors may have few degrees of freedom, making the chi-square test less sensitive. This is another reason to rely on multiple indices rather than chi-square alone.
In practice, I include chi-square in my reports for completeness but let the incremental and absolute indices drive my conclusions. Chi-square remains useful for nested model comparisons using a chi-square difference test. When you need to decide between two competing models, the chi-square difference test tells you whether the improvement in fit is statistically significant. For routine model evaluation, however, CFI, RMSEA, and SRMR carry more weight because they are less distorted by sample size. A model with a significant chi-square but excellent CFI, TLI, RMSEA, and SRMR is well-fitting. A model with a non-significant chi-square but poor incremental indices is not.
CFI (Comparative Fit Index)
The Comparative Fit Index is the first number most researchers check after running a CFA. CFI measures how much better your fitted model is compared to an independence baseline model where all variables are uncorrelated. The formula compares the chi-square of your target model to the chi-square of the independence model, adjusted for degrees of freedom. CFI values range from 0 to 1, with higher values indicating better fit.
A CFI of 1.0 means your model fits the data perfectly relative to the independence baseline. Values above 0.95 indicate good fit according to Hu and Bentler (1999). Values between 0.90 and 0.95 are acceptable but could be better. Values below 0.90 suggest your model needs substantial revision. The independence model baseline that CFI compares against is deliberately the worst reasonable model, with all variables assumed to be independent of each other. Beating this baseline should not be difficult in practice, which is why values below 0.90 are genuinely concerning.
What I like about CFI is that it is relatively stable across sample sizes compared to chi-square. It does not heavily penalize model complexity, which makes it useful for comparing models with different numbers of parameters. Reviewers almost always ask about CFI in their first round of feedback. In my experience, a CFI of 0.95 or higher satisfies most journal reviewers. The independence model baseline that CFI compares against is deliberately the worst reasonable model, so beating it should not be difficult.
One subtle point that many researchers miss: CFI can be inflated in small samples. If your sample is below 200 and your CFI is above 0.95, verify the result with RMSEA and SRMR before declaring victory. The two-index rule from Hu and Bentler requires both indices to meet their thresholds simultaneously. I have seen researchers report CFI equals 0.96 with great enthusiasm while SRMR sat at 0.14. The CFI was misleading. Always check all four indices together before drawing any conclusions about model quality.
TLI (Tucker-Lewis Index)
The Tucker-Lewis Index, also called the Non-Normed Fit Index, is closely related to CFI but with an important distinction. TLI penalizes model complexity. Its formula adjusts for the number of parameters in your model, which means a more complex model must demonstrate substantially better fit to achieve the same TLI value. This parsimony adjustment makes TLI a more conservative index than CFI.
Like CFI, TLI is compared against an independence baseline. Values range from 0 to 1, with values above 0.95 indicating good fit. Unlike CFI, TLI can exceed 1.0 under certain conditions. This typically happens when the model fits the data exceptionally well or when sample size is small. A TLI above 1.0 is not necessarily wrong, but it warrants a check of your other indices to ensure the result is not an artifact of estimation with limited data.
The key difference between CFI and TLI emerges in more complex models. TLI tends to be slightly lower than CFI because of the parsimony adjustment. If your TLI meets the 0.95 threshold, CFI almost certainly does as well. If TLI falls short by a small margin while CFI clears 0.95, pay attention to TLI. It is signaling that your model may be overfitting by including unnecessary parameters that inflate CFI without contributing meaningful explanatory power.
I recommend reporting both indices in your results section. They usually agree, and when they diverge, the divergence is informative. Some journals specifically request TLI alongside CFI as a matter of reporting policy. In the r-statistics.co guide, the authors note that TLI and CFI formula-level comparison reveals why they diverge, and I find this level of understanding helpful for explaining results to collaborators who ask about the difference.
RMSEA (Root Mean Square Error of Approximation)
RMSEA is a parsimony-adjusted index that estimates how well your model fits per degree of freedom. It answers a different question than CFI or TLI. While CFI and TLI measure incremental improvement over a null model, RMSEA measures absolute misfit. A value of 0 indicates perfect fit. Values increase as misfit increases. RMSEA is one of the most widely cited fit indices because it accounts for model complexity, which makes it useful for comparing models with different numbers of parameters.
Browne and Cudeck (1992) established the most widely used RMSEA thresholds. A value of 0.05 or below indicates close fit. Values between 0.05 and 0.08 indicate reasonable fit. Values above 0.10 indicate poor fit. The 0.08 cutoff is sometimes used as the boundary for acceptable fit, though Hu and Bentler (1999) advocated for a stricter 0.06 threshold. In practice, I see many researchers treat 0.08 as acceptable and 0.05 as good, with 0.06 to 0.08 as the gray zone that requires judgment.
RMSEA is always interpreted alongside its 90% confidence interval. The confidence interval provides a range of plausible RMSEA values in the population. A tight confidence interval around 0.04 indicates you can be confident the model fits reasonably well. A wide interval spanning from 0.02 to 0.12 means the fit estimate is uncertain and sensitive to sampling variation. Look at both the point estimate and the interval. Do not report one without the other. The close fit test is particularly useful. When the upper bound of the 90% confidence interval falls below 0.05, you have statistically demonstrated close fit. When the lower bound exceeds 0.08, you have demonstrated misfit. When the interval straddles 0.05 to 0.08, the result is inconclusive and requires judgment alongside the other indices.
RMSEA is sensitive to model complexity. More complex models tend to produce higher RMSEA values even when they fit well in absolute terms. This is intentional. RMSEA rewards parsimony. A simple model that explains the data adequately will have a better RMSEA than a complex model with the same explanatory power. In my own analyses, RMSEA is often the index that triggers respecification. When RMSEA is above 0.08 but CFI looks acceptable, I know the model needs work despite passing the incremental fit test. The parsimony penalty is doing its job.
SRMR (Standardized Root Mean Square Residual)
SRMR measures the average difference between the observed correlations and the correlations your model predicts. It is the simplest fit index to explain and arguably the most intuitive. A value of 0 means your model reproduces every observed correlation perfectly. Higher values indicate larger average discrepancies. The scale is standardized, which means values are comparable across different datasets and models regardless of the original measurement scale.
Hu and Bentler (1999) established a cutoff of 0.08 for SRMR. Values below 0.05 indicate excellent fit. Values between 0.05 and 0.08 indicate acceptable fit. Values above 0.10 suggest the model needs substantial revision. SRMR is less commonly discussed in introductory guides than CFI or RMSEA, but I consider it essential. It catches problems that the other indices can miss, particularly when working with categorical data or complex models. Researchers who skip SRMR are leaving a critical diagnostic tool unused.
SRMR is particularly valuable when working with ordinal data. CFI and RMSEA can be misleading with ordinal survey data analyzed using maximum likelihood estimation. CFI tends to be inflated, making mediocre models look better than they are. RMSEA can be overly conservative. SRMR remains relatively stable across data types, which makes it a useful sanity check. Consider this scenario from my own consulting work. A researcher ran a CFA on a five-point Likert scale measuring organizational commitment. CFI came back at 0.97. RMSEA was 0.04. The model looked excellent by conventional standards. But SRMR was 0.12. The raw residuals told a different story entirely. The model-implied correlations were systematically off from what the data showed. The inflated CFI was an artifact of the categorical data and the estimation method. SRMR was the first index to reveal the problem. Always check all four indices together.
Beyond global fit, you should also examine local model fit indicators such as standardized factor loadings, R-squared values, and modification indices for individual items. A model can pass all global fit indices while one or two items have weak loadings below 0.40. Global fit indices are averages across the entire model. They do not reveal whether specific items are underperforming. Examine factor loadings alongside the global indices to get the complete picture of your model quality. An item with a loading of 0.28 is not measuring its construct well, regardless of what the overall CFI says.
The Hu and Bentler two-index rule takes this further. It requires both RMSEA and SRMR to meet their respective cutoffs simultaneously. A model that passes only one of these two indices has not truly demonstrated adequate fit. This rule was proposed specifically to prevent researchers from cherry-picking the most flattering index when reporting results. A model with CFI equals 0.96, RMSEA equals 0.05, and SRMR equals 0.12 does not pass the two-index rule. The SRMR is telling you there is a problem that the other indices are hiding.
The Hu & Bentler Cutoff Values
In 1999, Hu and Bentler published a simulation study that became the most influential source for CFA fit index cutoffs. Their study used confirmatory factor analysis models with varying misspecifications, sample sizes, and estimation methods. From these simulations, they derived cutoffs that balanced Type I and Type II error rates. The result is a set of benchmarks that have been cited thousands of times and adopted as reporting standards by journals across psychology, education, and business.
The benchmarks they proposed are CFI greater than or equal to 0.95, TLI greater than or equal to 0.95, RMSEA less than or equal to 0.06, and SRMR less than or equal to 0.08. These cutoffs are now the default standard in most psychology, education, and business journals. When reviewers say your model fit is inadequate, they are almost always applying Hu and Bentler thresholds. Knowing these values is not optional. They are part of the shared methodological vocabulary of structural equation modeling.
I want to be clear about the limitations of these cutoffs. They were derived from a specific set of simulation conditions. The models tested had continuous indicators and used maximum likelihood estimation. Subsequent research has shown that different cutoffs may be appropriate for categorical data, small samples, and complex models. Kenny and colleagues (2015) argued that the RMSEA 0.06 cutoff is overly strict for many applied contexts. They suggested that values up to 0.10 can be acceptable for certain model types, particularly bifactor models and models with many indicators. Shi and colleagues (2019) demonstrated that CFI cutoffs should be adjusted downward for models with many indicators, where the baseline model itself becomes more complex.
My recommendation is to know the Hu and Bentler values and report against them, but treat them as guidelines rather than absolute rules. A model with CFI equals 0.94, TLI equals 0.93, RMSEA equals 0.07, and SRMR equals 0.06 is not catastrophically misfit. It is borderline. The substantive question is whether the model is theoretically coherent and whether the factor loadings are strong. Fit indices are one piece of evidence, not the sole arbiter of model quality. Use them alongside theory, factor loadings, and cross-validation evidence when making decisions about your measurement model.
What to Do When Fit Indices Disagree
Conflicting fit indices are the most common frustration researchers face. Your CFI looks good but RMSEA is high. Your TLI is below threshold while CFI clears 0.95. SRMR is acceptable but global indices suggest misfit. This happens all the time. It does not mean you have done something wrong. Each index captures a different dimension of fit. CFI answers whether your model beats an independence baseline. RMSEA answers whether the misfit per parameter is reasonable. SRMR answers whether the observed and predicted correlations are close on average. These are genuinely different questions, and they can produce different answers even for the same model.
Understanding why indices disagree is as important as knowing what to do about it. CFI and TLI are incremental indices that compare your model to a baseline. If your variables have high correlations overall, the baseline model will have high chi-square, making it harder for your model to show improvement. RMSEA is parsimony-adjusted and can penalize models with many parameters. SRMR measures raw residuals and can detect misfit that the comparative indices smooth over. When indices disagree, the disagreement is informative. It tells you something about the nature of your model and your data.
Here is a step-by-step decision framework I use when indices disagree. Start by noting which indices agree and which diverge. Identify the pattern. Then evaluate the factor loadings. Are they above 0.50? Are they theoretically sensible? If factor loadings are strong and the model is theoretically coherent, you have a stronger case for accepting borderline fit. If factor loadings are weak, the misfit is likely real and respecification is needed before you proceed to publish.
When CFI and SRMR pass but RMSEA exceeds 0.08, check the RMSEA confidence interval. If the upper bound falls below 0.10, the model may be acceptable with a caveat. RMSEA penalizes complexity more heavily than CFI does, so a complex but well-specified model can show this pattern. The confidence interval upper bound tells you whether the worst-case RMSEA is still within the range of reasonable misfit. If the upper bound exceeds 0.10, respecification is warranted regardless of what the other indices say.
When CFI falls below 0.90 but RMSEA and SRMR are acceptable, the issue is often sample size. Small samples tend to deflate CFI because the independence model chi-square is unstable. Do not panic. Check your sample size. If it is below 200, prioritize RMSEA and SRMR over CFI. If your sample is adequate, examine modification indices for potential model improvements. In either case, document the discrepancy in your results section and explain your interpretation.
When both SRMR and CFI perform poorly, the model has fundamental problems. High SRMR means the model-implied correlations are consistently different from observed correlations. High CFI means the model barely beats the independence baseline. Review your factor structure. Consider whether items are loading on the correct factors. Dropping items with low or mis-specified loadings is often the fastest path to improvement. A targeted respecification that removes problematic items can improve all four indices simultaneously.
When TLI lags CFI by more than 0.03, the model may be overfitting. TLI penalizes complexity more heavily than CFI. If CFI is good but TLI is poor, your model may have too many parameters relative to what the data justifies. Simplify the model by removing theoretically weak cross-loadings or residual covariances. I have seen researchers add cross-loadings to push CFI above 0.95, only to have TLI drop below 0.90. The model was overfitting, and TLI caught it before a reviewer did.
The Hu and Bentler two-index rule is a practical safeguard. It requires both RMSEA and SRMR to meet their thresholds simultaneously. This prevents the common practice of reporting only the most flattering index. A model that passes CFI but fails both RMSEA and SRMR has not demonstrated adequate fit. A model that passes RMSEA and SRMR but falls slightly short on CFI may be acceptable, particularly with a small sample. Apply judgment, but do not ignore the two-index requirement when you can meet it.
Improving Fit: Model Respecification Using Modification Indices
When your fit indices suggest misfit, modification indices tell you where to look. Modification indices estimate how much the overall model chi-square would decrease if you freed a currently constrained parameter. Large modification indices point to the biggest sources of misfit in your model. They are available in all major CFA software: lavaan in R, Mplus, AMOS, and SPSS AMOS. Learning to use them effectively is what separates researchers who iterate their way to a working model from those who report inadequate fit and move on without resolving the problems.
The most common respecification actions are adding cross-loadings and adding residual covariances. A large modification index for a cross-loading means an item is loading strongly on a factor other than the one you specified. This is common in survey research where items may have ambiguous wording that taps multiple constructs. A large modification index for a residual covariance means two item residuals are correlated. This happens when items share similar wording, reverse-keyed items cluster together, or items measure overlapping content within the same factor.
Here is the rule I follow for modification indices. Only add parameters that make theoretical sense. A large modification index for a cross-loading of an anxiety item on a depression factor makes psychological sense given the comorbidity of these constructs. A large modification index for a mathematics item loading on a verbal ability factor probably does not. Theory guides respecification. The modification index tells you where to look. Theory tells you whether to make the change. Never add a parameter solely because the modification index is large. Always have a substantive rationale that you can explain to a reviewer.
Use a cutoff of at least 3.84 for the chi-square change, which corresponds to a p-value of 0.05 with one degree of freedom. This is equivalent to a modification index value of 3.84 or higher. Some researchers use a more conservative cutoff of 6.63, which corresponds to p equals 0.01. The more conservative the cutoff, the fewer parameters you add, and the less risk of overfitting your model to sample-specific noise. I tend to start with 3.84 and only add parameters with modification indices well above that threshold.
Stop respecifying when fit indices reach acceptable levels and the remaining modification indices are small. Do not keep iterating until every modification index is near zero. That approach leads to overfitting. You are fitting sample-specific noise rather than generalizable structure. A model with acceptable fit and strong theoretical justification is better than a model with perfect fit that no one can defend or replicate in a new sample. After respecification, re-estimate the model and re-evaluate all fit indices. Modification indices are based on the current model. Adding parameters changes the model, which changes the modification indices. One round of respecification is usually sufficient. If you need more than three or four rounds, reconsider whether your initial measurement model is viable.
How to Report Model Fit Indices in Publications
Journal reviewers have specific expectations for how CFA results are reported. Failing to meet these expectations is one of the most common reasons for revision requests. Here is what you need to include in your results section to satisfy standard reporting requirements.
Report all four primary fit indices: chi-square statistic with degrees of freedom and p-value, CFI, TLI, RMSEA with its 90% confidence interval, and SRMR. Reporting only one or two indices is no longer acceptable in most journals. The field has moved toward comprehensive reporting standards over the past two decades. Each index contributes different information, and reviewers expect to see the complete picture. If you ran multiple model specifications, report fit for each specification and explain which one you selected and why.
Include model specification details. State the estimation method, the software used, the sample size, and how missing data were handled. If you used a robust estimator like MLR or WLSMV, mention it explicitly. If your data is ordinal, note whether you treated it as continuous or categorical and which estimator you chose. Using WLSMV with ordinal data is increasingly common and generally recommended over maximum likelihood estimation because it does not assume continuous indicators. This context allows readers to evaluate whether the reported fit values are appropriate for your analysis and whether alternative specifications might have produced different results.
Be honest about the results. Do not claim good fit if indices clearly disagree. If RMSEA is above threshold while CFI is acceptable, say so and explain your reasoning. Reviewers respect transparency. They penalize selective reporting more than they penalize imperfect fit. If you ran multiple model specifications and settled on one, explain the selection criteria. If you dropped items after examining modification indices, report which items were dropped and why.
The most common reporting mistakes I see are omitting the RMSEA confidence interval, reporting chi-square without the other indices, claiming good fit without acknowledging borderline values, and reporting only the favorable indices while omitting unfavorable ones. Each of these mistakes signals methodological carelessness to reviewers and can turn a revise-and-resubmit into a rejection. I have seen papers rejected solely because the authors reported only CFI and omitted SRMR entirely.
Important Caveats When Using Fit Indices
Fit indices do not behave identically across all conditions. Understanding their limitations will help you interpret results more accurately and defend your decisions to reviewers. These caveats separate researchers who understand their methods from those who simply apply formulas without understanding their assumptions.
Sample size affects every index, though in different ways. Chi-square is almost always significant with large samples. CFI tends to increase with sample size up to about 250 and then stabilizes. RMSEA becomes more precisely estimated with larger samples, producing tighter confidence intervals. With small samples below 200, all indices become less reliable. A model with CFI equals 0.91 and RMSEA equals 0.09 in a sample of 150 might be acceptable given the uncertainty. The same values in a sample of 800 suggest genuine misfit. Kenny and colleagues (2015) provided adjusted benchmarks for small-sample CFA that can guide interpretation when you cannot recruit a larger sample.
Categorical and ordinal data require special attention. Most CFA software defaults to maximum likelihood estimation, which assumes continuous indicators. With ordinal data, this assumption is violated. CFI tends to be inflated under these conditions, making mediocre models look better than they are. RMSEA tends to be more conservative. SRMR remains relatively stable. If you are analyzing Likert scale data, consider using robust estimators such as MLR or categorical estimators such as WLSMV. Check whether your software provides scaled or robust versions of fit indices and report those values alongside the standard estimates. The distinction between WLSMV and MLR matters for interpretation, and reviewers familiar with categorical CFA will expect you to have considered it.
Model complexity matters in ways that many researchers overlook. CFI tends to decrease as the number of indicators increases. A measurement model with 30 indicators distributed across three factors may naturally have a lower CFI than a model with 9 indicators across the same factors. This does not mean the larger model is worse. It means the baseline comparison is different. Do not apply the same cutoffs mechanically across all model sizes. Consider the complexity of your baseline model and the nature of your constructs when evaluating whether a CFI of 0.93 represents genuine misfit or simply the cost of measuring complex constructs with many items.
Bifactor and hierarchical models present additional challenges that standard cutoffs were not designed to address. These models have a more complex structure than standard correlated-factors models. Traditional cutoff values were derived from standard models. Applying them directly to bifactor models can be misleading. Researchers working with bifactor structures should be particularly cautious about rigid cutoffs and emphasize theoretical justification alongside fit statistics. In these cases, the pattern of factor loadings, the specificity of general and group factors, and the substantive interpretation matter more than whether a single index crosses a threshold.
Frequently Asked Questions
How do you read confirmatory factor analysis results?
Start with the chi-square statistic to assess exact fit, though treat significance as expected with large samples. Then check CFI for incremental fit, aiming for 0.95 or higher. Examine RMSEA for parsimony-adjusted fit, with 0.06 or below as the target and the 90% confidence interval providing additional context. Check SRMR for average residual correlation, targeting 0.08 or below. Evaluate factor loadings for local fit, with values above 0.50 as a common benchmark. If indices conflict, use the decision framework above to determine whether the model is acceptable or needs respecification. Report all values transparently in your results section with the estimation method and sample size.
How do you interpret SRMR?
SRMR is the average difference between your observed correlations and your model-implied correlations. It ranges from 0 to 1. A value of 0 means perfect fit. Values below 0.05 indicate excellent fit. Values between 0.05 and 0.08 are acceptable per Hu and Bentler (1999). Values above 0.10 suggest the model needs revision. SRMR is particularly useful as a sanity check because it is relatively stable across sample sizes and data types. When CFI looks inflated by a large sample or categorical data, SRMR often reveals the true level of misfit. Always check SRMR alongside CFI and RMSEA.
What is a good model fit for RMSEA?
A good RMSEA value is 0.05 or below. This indicates close fit according to Browne and Cudeck (1992). Acceptable fit ranges from 0.05 to 0.08. Values above 0.10 indicate poor fit. The 90% confidence interval provides additional information. If the upper bound falls below 0.05, you have statistically demonstrated close fit. If the lower bound exceeds 0.08, you have demonstrated misfit. If the interval spans 0.05 to 0.08, the result is borderline and requires judgment alongside the other indices. Hu and Bentler (1999) proposed a stricter cutoff of 0.06, which is increasingly used in psychology and education journals. Check which standard your target journal applies.
What is a good CFI value for CFA?
A good CFI value is 0.95 or higher. Values above 0.97 indicate excellent fit. Values between 0.90 and 0.95 are acceptable but warrant attention and further inspection of the other indices. Values below 0.90 suggest the model needs substantial revision. CFI compares your model to an independence baseline where all variables are uncorrelated. A value of 1.0 means perfect fit relative to that baseline. CFI is one of the most heavily weighted indices by journal reviewers, so prioritize it in your model evaluation. But never rely on CFI alone. Always check RMSEA and SRMR as well.
When is chi-square not significant in CFA?
Chi-square is almost always significant in confirmatory factor analysis when your sample size exceeds 200. A non-significant chi-square is rare in practice and does not necessarily indicate good fit. With large samples, even trivial model misfit produces a significant result. Conversely, a significant chi-square with a small sample may reflect estimation instability rather than genuine misfit. Instead of treating chi-square significance as a binary test, report it alongside incremental indices and use it primarily for nested model comparisons where a chi-square difference test provides meaningful information about whether one model fits significantly better than another.
What does TLI tell us?
The Tucker-Lewis Index tells you how much better your model fits compared to an independence baseline, adjusted for model complexity. A TLI of 0.95 or higher indicates good fit. The parsimony adjustment means TLI penalizes models with many parameters, making it a more conservative index than CFI. When TLI and CFI diverge by more than 0.03, it often signals that the model is overfitting by including unnecessary parameters. TLI can exceed 1.0, which typically indicates very good fit or small-sample estimation effects that should be investigated. Report TLI alongside CFI for a complete picture of incremental fit.
What if CFI and RMSEA disagree?
Disagreement between CFI and RMSEA is common and manageable. When CFI is good but RMSEA is above 0.08, examine the RMSEA confidence interval. If the upper bound is below 0.10, the model may be acceptable with a caveat. Check your factor loadings. If they are strong and theoretically coherent, the model may still be useful despite the parsimony penalty reflected in RMSEA. If the RMSEA confidence interval upper bound exceeds 0.10, respecification is needed. When CFI is below 0.90 but RMSEA and SRMR are acceptable, the issue may be sample size. Prioritize the indices that are less sensitive to your sample characteristics and explain the discrepancy in your reporting.
Does sample size affect RMSEA and CFI?
Yes, both indices are affected by sample size, though in different ways. RMSEA becomes more precisely estimated with larger samples, producing tighter confidence intervals. With small samples below 200, the RMSEA confidence interval is wide and the point estimate is unreliable. CFI tends to increase with sample size up to about 250 and then stabilizes. With small samples, CFI can be deflated, making adequate models appear inadequate. The practical implication is that rigid cutoffs should be applied more flexibly with small samples. A model with borderline fit values in a sample of 150 may be more acceptable than the same values in a sample of 500. Always consider your sample size when interpreting fit index values.
Can good model fit indices compensate for poor factor loadings?
No. Good global fit indices cannot compensate for weak factor loadings. Fit indices measure how well the overall covariance structure is reproduced. They do not guarantee that individual items measure their intended constructs well. An item with a factor loading of 0.25 is contributing very little to its factor, regardless of what the CFI says. Always examine factor loadings as part of local model fit evaluation. A model with CFI equals 0.96 but several loadings below 0.40 needs item revision. Strong global fit combined with strong local fit is the standard to aim for in any CFA study.
How do you report model fit indices?
Report at minimum the chi-square statistic with degrees of freedom and p-value, CFI, TLI, RMSEA with its 90% confidence interval, and SRMR. Include sample size, estimation method, and software used. Present values in a table for clarity. State whether the model met predefined fit thresholds. If it did not, explain which indices fell short and what steps you took. Do not selectively report only the favorable indices. Reviewers expect transparency and will penalize incomplete reporting. Comprehensive reporting is now the expectation in peer-reviewed journals across the social sciences.
Conclusion
Learning how to read model fit indices in confirmatory factor analysis takes practice, but the underlying logic is straightforward. The four primary indices each measure something different. CFI and TLI tell you how much better your model is than an independence baseline. RMSEA tells you whether the misfit is acceptable given model complexity. SRMR tells you whether the observed and model-implied correlations are close on average. No single index tells the full story, which is why comprehensive reporting of all indices has become the standard in published research.
The Hu and Bentler cutoffs provide a useful starting point. CFI greater than or equal to 0.95, TLI greater than or equal to 0.95, RMSEA less than or equal to 0.06, and SRMR less than or equal to 0.08 are the benchmarks most reviewers apply. Apply them with judgment, not mechanically. Consider your sample size, your data type, and your model complexity when interpreting results. When indices conflict, use the decision framework outlined above to determine whether your model is acceptable or whether respecification is needed. Check factor loadings. Evaluate theoretical coherence. And always report all indices transparently, including the ones that do not support your preferred conclusion. Honest reporting builds credibility with reviewers and advances science.