If you have ever run a confirmatory factor analysis (CFA) or a structural equation model (SEM), you have probably come across the average variance extracted, or AVE. Reviewers ask for it, software reports it, and yet many researchers are not fully confident about what the number actually means. Learning how to interpret an average variance extracted value is one of the most useful skills you can develop when validating a measurement model, because that single number tells you whether your items are doing their job.
Our team has spent years building and reviewing measurement models across psychology, information systems, and education research. In that time we have seen every possible AVE scenario, from clean values above 0.70 to borderline cases stuck at 0.42. This guide breaks down the interpretation process step by step, with real numerical examples, so you can read your AVE output with confidence.
We will cover the definition, the formula, a complete worked calculation, the interpretation scale, the connection to convergent and discriminant validity, and a troubleshooting checklist for when your AVE comes back low. By the end, you will know exactly what your AVE value is telling you and what to do about it.
Table of Contents
Quick Answer: What Is a Good AVE Value?
An AVE of 0.50 or higher is considered acceptable. This threshold, established by Fornell and Larcker in 1981, means your latent construct explains more than half of the variance in its indicators, leaving less than half attributable to measurement error. That is the standard rule of thumb for convergent validity in 2026.
Values below 0.50 indicate that your items contain more error variance than construct variance. In that situation, the measurement is considered weak. However, values between 0.40 and 0.50 can sometimes be tolerated when composite reliability (CR) is high enough, typically at 0.70 or above. We will walk through that nuance later in this article.
What Is Average Variance Extracted?
Average variance extracted (AVE) is a measure of convergent validity used in structural equation modeling and confirmatory factor analysis. It quantifies the average proportion of variance in a set of indicators that is captured by the underlying latent construct, relative to the variance caused by measurement error. In simple terms, AVE tells you how much of your items’ variation is actually about the thing you are trying to measure.
The concept was introduced by Claes Fornell and David Larcker in their 1981 paper, “Evaluating Structural Equation Models with Unobservable Variables and Measurement Error.” That paper remains the most widely cited source for AVE interpretation in 2026, and reviewers across social science journals still expect you to report it.
AVE applies to reflective measurement models, where indicators are treated as manifestations of the latent construct. It does not apply to formative models, where indicators define the construct rather than reflect it. If you are working with formative indicators, AVE is not meaningful and should not be reported.
The AVE Formula Explained
The AVE formula looks more intimidating than it is. Here is the standard form:
AVE = (Sum of squared standardized factor loadings) / (Sum of squared standardized factor loadings + Sum of error variances)
Each squared standardized factor loading represents the communality of that item, meaning the proportion of its variance explained by the construct. The error variance for each item is simply 1 minus the squared loading. So a cleaner version of the formula is:
AVE = (Sum of squared loadings) / (Sum of squared loadings + Sum of (1 – squared loadings))
For a congeneric measurement model where all items load on a single factor, you square each standardized loading, add those squared values up, then divide by the number of items. The denominator effectively accounts for both explained and unexplained variance. The result is a proportion between 0 and 1.
How to Calculate AVE Step by Step
No competitor we reviewed provides a fully worked numerical example, so we will fix that. Imagine you have a three-item construct called “Customer Satisfaction” with standardized factor loadings of 0.80, 0.70, and 0.65 from your CFA output. Here is how to calculate AVE manually.
Step 1: Square each standardized factor loading. Squaring gives you the communality, which is the variance in each item explained by the construct. For our example, 0.80 squared equals 0.64, 0.70 squared equals 0.49, and 0.65 squared equals 0.4225.
Step 2: Add the squared loadings together. This is the numerator of the AVE formula. In our case, 0.64 + 0.49 + 0.4225 equals 1.5525.
Step 3: Calculate each item’s error variance. Error variance is 1 minus the squared loading. For our three items that gives 0.36, 0.51, and 0.5775.
Step 4: Add the error variances together. The sum is 0.36 + 0.51 + 0.5775, which equals 1.4475. This represents the total measurement error across all three items.
Step 5: Add the numerator and the error sum. This is your denominator. Here, 1.5525 + 1.4475 equals 3.0000. Notice that this also equals the number of items, because each item contributes one unit of total variance.
Step 6: Divide the numerator by the denominator. AVE = 1.5525 / 3.0000 = 0.5175. Rounded, your AVE is approximately 0.52, which clears the 0.50 threshold for convergent validity.
That is the entire calculation. You can replicate this process with any set of standardized loadings from AMOS, lavaan, SmartPLS, or any other SEM software. The key is to use standardized, not unstandardized, loadings.
How to Interpret an Average Variance Extracted Value: The Complete Scale
Knowing how to interpret an average variance extracted value means understanding what each range tells you about your measurement quality. Here is the full interpretation scale we use in our own reviews.
AVE above 0.70 (Excellent): Your construct explains more than 70 percent of the variance in its indicators. This is unusually strong and suggests a tight, well-defined measurement model. You rarely see values this high outside of scales with very homogeneous items.
AVE between 0.50 and 0.70 (Acceptable to Good): Your construct explains more variance than measurement error. This is the comfortable zone. Most published scales in 2026 land somewhere in this range, and reviewers will not question your convergent validity.
AVE between 0.40 and 0.50 (Borderline): Your items contain slightly more error than construct variance. This range is conditionally acceptable. Many methodologists, including those writing for journals like Psychological Methods, accept AVE values down to 0.40 if composite reliability is at least 0.70 and the items are theoretically justified. You should explicitly acknowledge the borderline value in your manuscript and cite supporting sources.
AVE below 0.40 (Unacceptable): Your measurement model has a serious problem. The construct is not capturing enough variance from its indicators, which means your items are not measuring what you think they are measuring. You need to revisit your scale, and we cover exactly what to do in the troubleshooting section below.
AVE below 0.30 (Critical): At this level, the construct explains less than a third of the indicator variance. The measurement model is essentially noise. Dropping weak items rarely fixes this on its own. You likely need to reconceptualize the construct or redesign the items entirely.
AVE and Convergent Validity
Convergent validity is the degree to which items that should load on the same construct actually do load together. AVE is the primary numeric test for convergent validity in CFA and SEM. When your AVE is 0.50 or higher, you have statistical evidence that your items converge on their intended construct.
Reviewers typically expect you to report AVE alongside composite reliability when you describe your measurement model. Together, these two statistics paint a complete picture: CR tells you whether your items are internally consistent, while AVE tells you whether they capture enough true variance to be meaningful. A scale can have high CR but low AVE, which is a red flag we discuss in the comparison section below.
Convergent validity is also assessed through individual factor loadings. As a companion to AVE, each standardized loading should ideally be 0.70 or higher, meaning the item shares at least half its variance with the construct. Loadings between 0.40 and 0.70 are sometimes retained if they are theoretically important and if dropping them does not improve AVE.
AVE and Discriminant Validity: The Fornell-Larcker Criterion
AVE is not only about convergent validity. It also plays a central role in testing discriminant validity, which confirms that your constructs are empirically distinct from one another. The most common method for this is the Fornell-Larcker criterion.
The Fornell-Larcker rule states that the square root of each construct’s AVE should be greater than the correlation between that construct and any other construct in the model. In practice, you create a correlation matrix where the diagonal entries are the square roots of AVE values and the off-diagonal entries are inter-construct correlations. If every diagonal value exceeds the correlations in its row and column, discriminant validity is established.
Here is a concrete example. Imagine three constructs with AVE values of 0.55, 0.60, and 0.50. The square roots are 0.74, 0.77, and 0.71. If the highest correlation between any two constructs is 0.65, then every square root of AVE exceeds every inter-construct correlation, and discriminant validity holds.
One important caveat: recent research, including Henseler and colleagues’ 2015 paper, shows that the Fornell-Larcker criterion can fail to detect discriminant validity problems, especially in PLS-SEM. The HTMT (heterotrait-monotrait) ratio is now recommended as a more sensitive complement. We mention this because journal reviewers in 2026 increasingly ask for HTMT in addition to Fornell-Larcker.
What to Do When AVE Is Below 0.50
This is the situation that sends researchers to forums and sends students to their advisors. If your AVE is below 0.50, you have several options, and we recommend working through them in this order.
1. Check for low-loading items. Review your standardized factor loadings. If one or two items have loadings below 0.50, they are dragging your AVE down. Calculate AVE again after temporarily removing the weakest item. If AVE jumps above 0.50, you have found the culprit.
2. Consider cross-loadings. An item with a low loading on its intended factor may have a high cross-loading on another factor. Check the pattern matrix. If an item loads more strongly on a different construct, consider reassigning it or removing it.
3. Evaluate theoretical justification. Before dropping any item, ask whether it is theoretically central to the construct. If removing it changes the meaning of your scale, you may need to keep it and accept the borderline AVE. Document your reasoning in your manuscript.
4. Check for reverse-coded items. A common mistake is failing to reverse-code negatively worded items before running the CFA. This can deflate loadings and crush your AVE. Verify that all items are scored in the same direction.
5. Look at composite reliability. If your AVE is between 0.40 and 0.50 but your CR is above 0.70, you may be able to justify retaining the model. Cite methodological literature that supports the conditional acceptance of borderline AVE values.
6. Increase sample size. Standardized loadings can be unstable in small samples. If your sample is under 200, collecting more data may stabilize the loadings and improve AVE.
7. Reconsider the construct itself. If none of the above helps, your items may be tapping into more than one dimension. An exploratory factor analysis can reveal whether your construct should be split into two related but distinct factors.
AVE vs Composite Reliability: What Is the Difference?
This question comes up constantly on forums like Stack Exchange and the SmartPLS community. People assume AVE and composite reliability measure the same thing because both assess measurement quality. They do not, and confusing them leads to misinterpretation.
Composite reliability (CR) measures internal consistency, meaning how well your items hang together as a set. It is conceptually similar to Cronbach’s alpha but better suited to congeneric models where loadings vary. A CR of 0.70 or higher is considered acceptable.
AVE measures the amount of variance captured by the construct relative to error. It is possible to have high CR and low AVE at the same time. When that happens, your items are consistent with each other but they are collectively capturing more error than construct variance. This is why reporting both statistics together is standard practice.
Limitations of AVE
AVE is a useful metric, but it is not without problems. Understanding its limitations helps you interpret results more responsibly.
First, the 0.50 threshold is a rule of thumb, not a law. Recent critical papers argue that treating it as absolute leads to arbitrary item deletion and distorted constructs. Some constructs naturally produce lower AVE values because they are broad or multidimensional, and forcing them above 0.50 can do more harm than good.
Second, AVE is sensitive to the number of items. Adding more items to a construct can dilute AVE even if each item is individually adequate. Conversely, dropping items can artificially inflate AVE without genuinely improving measurement quality.
Third, AVE assumes a reflective measurement model. If your indicators are formative, AVE is not interpretable. Many researchers mistakenly report AVE for formative constructs, which misrepresents the measurement structure.
Fourth, the Fornell-Larcker criterion based on AVE has been shown to underperform in detecting discriminant validity issues, particularly in PLS-SEM. This is why HTMT is now recommended alongside or instead of Fornell-Larcker.
FAQs
What is the average variance extracted?
Average variance extracted (AVE) is a measure of convergent validity in structural equation modeling. It represents the average proportion of variance in a construct’s indicators that is explained by the latent construct itself rather than by measurement error. AVE is calculated as the mean of the squared standardized factor loadings across all items of a construct.
What is the acceptable value of AVE?
The acceptable value for AVE is 0.50 or higher, following the Fornell and Larcker (1981) standard. An AVE of at least 0.50 means the construct explains more than half the variance in its indicators. Values between 0.40 and 0.50 are sometimes accepted when composite reliability is 0.70 or above, though this is considered borderline and should be justified in your manuscript.
What does average variance mean?
In the context of AVE, average variance refers to the average amount of variance in a set of measurement items that is explained by the underlying latent construct, as opposed to random measurement error. It is not the same as the statistical variance of a distribution. AVE specifically measures explained variance averaged across items.
What is the purpose of the average variance extracted in SEM?
The purpose of AVE in SEM is to assess convergent validity and to support discriminant validity testing. For convergent validity, an AVE of 0.50 or higher confirms that items converge on their intended construct. For discriminant validity, the square root of AVE is compared against inter-construct correlations using the Fornell-Larcker criterion to verify that constructs are empirically distinct.
Conclusion
Knowing how to interpret an average variance extracted value gives you direct control over the quality of your measurement model. The core rule is simple: aim for 0.50 or higher, treat anything between 0.40 and 0.50 as conditionally acceptable with strong composite reliability, and treat anything below 0.40 as a signal that your scale needs work. Pair AVE with the Fornell-Larcker criterion or, better yet, the HTMT ratio to establish discriminant validity.
The next step is to run through your own CFA output and apply the interpretation scale and troubleshooting checklist from this guide. If your AVE is sitting in the borderline zone, work through the seven troubleshooting steps before you drop any items. Most AVE problems are fixable with targeted adjustments, not wholesale scale redesigns.