If you are planning a study that involves factor analysis, one of the first questions you will face is how many participants you need. The answer is not as simple as running a standard power analysis in G*Power and plugging in an effect size. In fact, traditional power calculations do not directly apply to factor analysis the way they do for t-tests or ANOVA, which leaves many researchers stuck.
I have seen this confusion play out repeatedly on forums like r/AskStatistics and r/psychometrics. Graduate students ask whether 61 responses are enough for a 12-item scale, or whether 100 participants can support a 65-variable factor analysis. The conflicting advice they receive from different textbooks only makes things worse. One source says you need 100 participants, another says 300, and yet another insists on 10 participants per variable.
This guide breaks down how to calculate the sample size you need for a factor analysis using clear, practical steps. I will walk you through every major rule of thumb, explain when each one applies, and show you how factors like communality and loading strength change the math. By the end, you will have a concrete process for determining whether your sample is large enough and what to do if it falls short.
The topic of sample size factor analysis deserves more than a single number answer, so let me walk you through the full picture.
Table of Contents
What Is Factor Analysis and Why Sample Size Matters
Factor analysis is a statistical method used to identify underlying latent constructs from a set of observed variables. Researchers use it for scale development, instrument validation, psychometric assessment, and exploring the dimensionality of measures. The goal is to find a smaller set of factors that explains the patterns of correlations among your variables.
Sample size matters because factor analysis relies on correlation matrices, and correlations estimated from small samples are unstable. A correlation computed from 50 people has a wide confidence interval, meaning the true relationship between variables could be quite different from what you observed. When you build a factor solution on shaky correlations, the entire structure can shift if you collect data from a new sample.
Small samples create specific problems in factor analysis. You risk overfitting, where the factor solution captures noise rather than true latent structure. You may encounter Heywood cases, where a variable has a communality above 1.0 or a factor loading above 1.0, which is mathematically impossible and signals a problem. Solutions from small samples often fail to replicate, which undermines the entire purpose of identifying stable latent constructs.
Statistical power, in the traditional sense of detecting a true effect, is not the main concern here. Instead, the concern is recovery: can your sample reproduce the true population factor structure? That question requires different thinking and different guidelines than a simple power calculation.
Absolute Sample Size Rules of Thumb (N = 100, 200, 300, 500)
The most widely cited approach to factor analysis sample size uses absolute numbers of participants. These recommendations come from decades of simulation research and practical experience. Comrey and Lee (1992) proposed one of the most recognized hierarchies, and it remains a useful benchmark today.
According to Comrey and Lee, 100 participants is poor, 200 is fair, 300 is good, 500 is very good, and 1,000 is excellent. These are not hard cutoffs but rather a gradient that reflects increasing confidence in your factor solution. Hatcher (1994) recommended a floor of either 100 participants or five times the number of variables, whichever is larger. Tabachnick and Fidell suggested a minimum of 300 cases for reliable factor solutions in most situations.
Here is a quick-reference comparison of the major absolute sample size recommendations:
- Comrey and Lee (1992): 100 = poor, 200 = fair, 300 = good, 500 = very good, 1000 = excellent
- Hatcher (1994): Minimum of 100, or 5 times the number of variables (whichever is larger)
- Tabachnick and Fidell: 300 cases as a general minimum
- Field: At least 300 for reliable solutions, with more needed for complex structures
- Guilford: Minimum of 200 participants
Notice that these rules give you a range, not a single number. Where you fall in that range depends on characteristics of your data, which we will cover shortly. If you have 50 participants, you are below every established threshold. If you have 400, you are in solid territory for most applications.
The absolute N approach is easy to remember and quick to apply. But it has a limitation: it ignores how many variables you are analyzing. Running factor analysis on 5 variables with 200 participants is very different from running it on 50 variables with the same 200 participants. That is where the subjects-to-variables ratio comes in.
Subjects-to-Variables Ratio (10:1, 5:1, 3:1)
The subjects-to-variables ratio, sometimes called the STV ratio or N:p ratio, addresses the fact that more variables require more participants to estimate stable correlations. The idea is intuitive: each additional variable adds estimation uncertainty, so your sample needs to grow proportionally.
Several different ratios have been proposed over the years, and they range significantly in stringency:
- 10:1 ratio: Ten participants per variable. This is the most conservative rule and is often attributed to Nunnally. If you have 30 variables, you would need 300 participants.
- 5:1 ratio: Five participants per variable. This is probably the most commonly cited ratio in practice and is recommended by Hatcher and others. For 30 variables, you need 150 participants.
- 3:1 ratio: Three participants per variable. This is a more liberal minimum, sometimes cited as the absolute floor. For 30 variables, this means 90 participants.
- 2:1 ratio: Two participants per variable. Cattlett suggested this as an absolute minimum, but few methodologists endorse it today without major caveats.
To calculate your ratio, simply divide your sample size by the number of variables. For example, if you have 250 participants and 20 variables, your ratio is 12.5:1, which exceeds even the most conservative recommendation. If you have 100 participants and 40 variables, your ratio is 2.5:1, which falls below most guidelines.
However, simulation studies have shown that these ratio rules oversimplify the picture. A landmark study by MacCallum, Widaman, Zhang, and Hong (1999) found that the STV ratio is less important than characteristics like communality and factor overdetermination. In some cases, a 3:1 ratio with high communalities produces better factor recovery than a 10:1 ratio with low communalities. The ratio is a useful starting point, but it should not be the only factor in your decision.
Key Factors That Affect Required Sample Size
The most important insight from modern factor analysis research is that required sample size depends on data characteristics, not just on the number of variables. Three factors matter most: communality, loading magnitude, and overdetermination. Understanding these will help you make an informed judgment about whether your sample is adequate.
Communality Levels
Communality refers to the proportion of variance in a variable that is explained by the common factors. High communalities mean your variables are well-explained by the factor structure, while low communalities mean substantial unique variance remains unexplained.
MacCallum et al. (1999) found that communality level is one of the strongest determinants of factor solution quality. When all communalities are above 0.60, even relatively small samples can produce good factor recovery. When communalities are below 0.40, you need much larger samples to achieve the same quality.
Here is a practical breakdown:
- All communalities above 0.60: Sample sizes as low as 100 can be adequate
- Communalities in the 0.40 to 0.60 range: Plan for at least 200 to 300 participants
- Communalities below 0.40: You may need 500 or more participants for stable solutions
- Wide range of communalities (some high, some very low): Larger samples needed because the low-communality variables destabilize the solution
You will not know your communalities until you run the analysis, which creates a chicken-and-egg problem. In practice, researchers use prior research or pilot data to estimate expected communality levels before deciding on a sample size.
Factor Loading Magnitude
Factor loadings indicate how strongly each variable relates to its underlying factor. Strong loadings make factors easier to identify and reproduce, while weak loadings create ambiguity about factor structure.
When your variables have strong loadings (above 0.70), the factor structure is clear and can be recovered even with modest sample sizes. When loadings are moderate (0.40 to 0.60), you need larger samples to distinguish true factor structure from noise. Loadings below 0.30 are generally considered too weak to interpret meaningfully, and their presence inflates sample size requirements substantially.
Simulation studies consistently show that loading magnitude interacts with communality. Variables with high communalities tend to have high loadings, so these two factors reinforce each other. If you expect your scale items to load strongly onto their factors based on prior validation work, you can be more comfortable with a smaller sample.
Overdetermination of Factors
Overdetermination refers to having many variables per factor. A factor defined by 8 or 10 variables is overdetermined, meaning there is abundant information to identify that factor reliably. A factor defined by only 3 variables is underdetermined and much harder to recover consistently.
MacCallum et al. (1999) demonstrated that overdetermination matters as much as sample size. In their simulations, solutions with 6 or more variables per factor were stable even at smaller sample sizes. Solutions with 3 or fewer variables per factor required much larger samples.
Practical guidance based on these findings:
- 6 or more variables per factor: More forgiving on sample size, may work with 150 to 200 participants if communalities are decent
- 4 to 5 variables per factor: Moderate sample requirements, aim for at least 250 to 300 participants
- 3 or fewer variables per factor: High sample size requirements, consider 400 or more participants or rethink your measurement model
Number of Variables and Factors
The total number of variables and the number of factors you expect both influence sample size needs. More variables means more parameters to estimate, which requires more data. More factors means more complexity in the solution, which also demands more information from your sample.
This is why the absolute N rules can be misleading in isolation. A study with 60 variables and 8 expected factors needs more participants than a study with 15 variables and 3 factors, even if both follow a 5:1 STV ratio.
How to Calculate the Sample Size You Need for a Factor Analysis
Now let me walk you through a concrete, step-by-step process for determining your sample size. This approach combines insights from all the guidelines above rather than relying on a single rule.
Step 1: Count your variables. Identify every variable that will enter the factor analysis. If you are developing a scale with 25 items, your variable count is 25. Be precise here because every ratio calculation depends on this number.
Step 2: Apply the STV ratio. Multiply your variable count by 5 for a moderate recommendation, or by 10 for a conservative recommendation. For 25 variables, the moderate estimate is 125 participants and the conservative estimate is 250 participants.
Step 3: Check against absolute minimums. Compare your STV-based estimate against the absolute N guidelines. If your STV estimate falls below 100, round up to at least 100. If you can realistically collect 300 or more, do so, as this places you in the “good” range regardless of other factors.
Step 4: Estimate your communalities. Based on prior research, pilot data, or theoretical expectations, estimate whether your communalities will be high (above 0.60), moderate (0.40 to 0.60), or low (below 0.40). If you expect high communalities and strong loadings, your STV estimate is probably sufficient. If you expect low communalities, add 100 to 200 participants.
Step 5: Count variables per expected factor. Estimate how many variables will load onto each factor. If each factor will have 6 or more variables, you have a buffer. If some factors will have only 3 variables, increase your sample size target.
Step 6: Take the largest number. Your final sample size target should be the largest number produced by steps 2 through 5. This conservative approach ensures you have enough participants regardless of which guideline matters most for your specific situation.
Here is a worked example. Suppose you are developing a new measure of workplace engagement with 30 items expected to form 5 factors of 6 items each. Prior scales in this area show moderate communalities around 0.50.
- STV at 5:1 gives you 150 participants
- STV at 10:1 gives you 300 participants
- Absolute minimum: at least 200 based on Guilford, 300 based on Tabachnick and Fidell
- Communality adjustment: moderate communalities mean no reduction, aim for the conservative estimate
- Variables per factor: 6 items per factor is adequate, so no additional inflation needed
- Final target: 300 participants
This process gives you a defensible number that accounts for multiple factors rather than blindly following one rule.
EFA vs CFA Sample Size Requirements
Exploratory factor analysis (EFA) and confirmatory factor analysis (CFA) have different sample size considerations. EFA is typically the first step in scale development, where you explore the factor structure without specifying it in advance. CFA tests a pre-specified factor structure for fit, which requires enough participants to estimate model parameters reliably.
As a general rule, CFA requires larger samples than EFA. This is because CFA estimates more parameters, including factor variances, covariances, and residual variances. A common rule of thumb for CFA is 10 participants per estimated parameter, though this varies by model complexity.
For CFA, several specific guidelines exist:
- Minimum: 200 participants for simple models with strong loadings
- Standard recommendation: 300 or more for most CFA models
- Complex models: 500 or more when you have many factors and cross-loadings
- Jackson’s rule: 20:1 ratio of participants to estimated parameters as a target, 10:1 as a minimum
CFA also relies on model fit indices like RMSEA, CFI, and SRMR. These indices need adequate sample sizes to produce trustworthy values. At very small sample sizes, fit indices behave erratically and can give misleading signals about model quality.
If you plan to run EFA on half your sample and CFA on the other half (a recommended practice for scale development), you need to double your target. A development sample of 300 for EFA plus a validation sample of 300 for CFA means recruiting 600 participants total.
KMO Test and Bartlett’s Test: Checking Adequacy
Before running factor analysis, you should check whether your data are suitable for it. Two pre-analysis tests help with this: the Kaiser-Meyer-Olkin (KMO) measure and Bartlett’s test of sphericity. These tests do not replace proper sample size planning, but they provide additional information about data adequacy.
The KMO measure ranges from 0 to 1 and indicates the proportion of variance among variables that might be common variance. Kaiser’s original guidelines specify that KMO values above 0.90 are marvelous, 0.80 to 0.90 are meritorious, 0.70 to 0.80 are good, 0.60 to 0.70 are mediocre, 0.50 to 0.60 are poor, and below 0.50 is unacceptable.
If your KMO falls below 0.60, you likely have issues with sample size or variable quality. Low KMO values often occur when you have too few participants relative to the number of variables or when some variables share little common variance with the others.
Bartlett’s test of sphericity checks whether your correlation matrix is significantly different from an identity matrix, where variables would be completely uncorrelated. A significant result (p less than 0.05) means your variables are correlated enough to proceed with factor analysis. If Bartlett’s test is not significant, factor analysis is probably inappropriate for your data regardless of sample size.
What to Do When Your Sample Is Too Small
Sometimes you cannot collect more data. Maybe your population is small, your funding is limited, or your data are already collected. Forum discussions on r/psychometrics and ResearchGate reveal this is one of the most common pain points for researchers, especially PhD students working with hard-to-reach populations.
If your sample is smaller than recommended, consider these options:
- Reduce the number of variables. Drop items with low variance, high missingness, or poor theoretical justification. Fewer variables means a better STV ratio.
- Use parceling techniques. Combine related items into parcels to reduce the number of variables and improve communalities.
- Consider Bayesian factor analysis. Bayesian methods can incorporate prior information, which can help with smaller samples. This requires specialized software and expertise.
- Use shorter extraction methods. Some extraction methods, like image factoring or least squares, may perform better with small samples than maximum likelihood.
- Report limitations honestly. Acknowledge that your factor solution may not replicate and frame it as exploratory rather than definitive.
Be transparent about your sample size limitations in your writeup. Reviewers and readers respect honesty about constraints more than overblown claims based on inadequate data.
Common Mistakes to Avoid
Several recurring mistakes appear in the questions researchers ask on forums and in published studies. Recognizing these will help you avoid the same pitfalls.
The biggest mistake is blindly applying a single rule of thumb without understanding when it applies. A 10:1 ratio is conservative and may be unnecessary with high communalities and strong loadings. Conversely, a 3:1 ratio may be dangerously optimistic with weak variables. Always consider the context of your data.
Another common error is confusing EFA and CFA requirements. Using EFA guidelines for a CFA study will likely leave you underpowered. CFA models estimate more parameters and need correspondingly larger samples for stable fit index estimates.
Ignoring KMO and communality diagnostics is another frequent problem. Researchers collect what they consider enough participants, skip the pre-analysis checks, and then produce unstable factor solutions that fail to replicate. Running KMO and inspecting communalities takes minutes but can save you from publishing irreproducible results.
Finally, many researchers forget about cross-validation. If you run EFA and CFA on the same sample, you are not truly validating your structure. Plan for split-sample validation from the start and budget your sample size accordingly.
FAQs
How do you calculate the needed sample size for factor analysis?
To calculate sample size for factor analysis, first count your variables and multiply by 5 (moderate) or 10 (conservative) to get your subjects-to-variables ratio estimate. Then check this against absolute minimums of 200 to 300 participants. Adjust upward if you expect low communalities (below 0.40) or have fewer than 4 variables per factor. Take the largest number from all these calculations as your target.
What is the rule of thumb for CFA sample size?
For confirmatory factor analysis, the general rule of thumb is at least 200 participants for simple models, 300 or more for standard models, and 500 or more for complex models with many factors. Jackson recommended a 20:1 ratio of participants to estimated parameters as ideal and 10:1 as a minimum. CFA typically requires larger samples than EFA because it estimates more parameters.
What is the minimum sample size for exploratory factor analysis?
The absolute minimum for exploratory factor analysis is generally 100 participants, though this only works with high communalities (above 0.60) and strong loadings. Most methodologists recommend at least 200 for fair quality and 300 for good quality factor solutions. If your communalities are low or you have few variables per factor, you need 500 or more.
How many participants do I need for a 12-item scale factor analysis?
For a 12-item scale, the 5:1 subjects-to-variables ratio gives you 60 participants, and the 10:1 ratio gives you 120 participants. However, absolute minimum guidelines suggest at least 100 to 200 participants regardless of variable count. For a 12-item scale, aim for at least 150 participants to ensure stable factor solutions.
Can I run factor analysis with 50 participants?
Running factor analysis with 50 participants is generally not recommended and falls below every established guideline. You risk unstable solutions, overfitting, and results that will not replicate. If you cannot collect more data, consider reducing your variable count, using parceling, or exploring Bayesian methods. Be transparent about limitations if you proceed.
Conclusion
Determining the right sample size for factor analysis requires more than memorizing a single number. The guidelines from Comrey and Lee, Hatcher, Tabachnick and Fidell, and MacCallum et al. each contribute important pieces to the puzzle. Your final decision should account for the number of variables, expected communalities, loading strength, and the number of variables per factor.
Use the step-by-step process I outlined above to calculate a defensible target for your specific study. Run KMO and Bartlett’s tests before extraction, inspect communalities after extraction, and be honest about limitations if your sample falls short. When you understand how to calculate the sample size you need for a factor analysis, you set up your research for results that are stable, reproducible, and ready for publication.