How Parallel Analysis Improves Factor Retention Decisions? (2026 Guide)

Deciding how many factors to retain in exploratory factor analysis is one of the most consequential choices a researcher makes. Retain too few and you collapse distinct constructs together. Retain too many and you model noise as if it were meaningful structure. Parallel analysis has emerged as the most reliable, empirically grounded method for making this decision, yet many researchers still default to outdated rules that overestimate the true number of factors.

In this guide, I will walk you through exactly how parallel analysis improves factor retention decisions compared to traditional methods. You will learn why the Kaiser criterion falls short, how random data simulation works, and how to implement parallel analysis in your own research using accessible tools.

Whether you are developing a new psychometric scale, validating a survey instrument, or conducting dimensionality assessment for the first time, understanding parallel analysis will give you more confidence in your factor structure. Let us break down the methodology without drowning in heavy mathematical notation.

What Is Parallel Analysis in Factor Analysis?

Parallel analysis is a statistical method for determining the number of factors to retain in exploratory factor analysis by comparing observed eigenvalues from your real data to eigenvalues from randomly generated datasets with the same sample size and number of variables.

Here is the core idea in plain terms. When you extract factors from random data that contains zero true latent factors, you still get eigenvalues above zero. These are artifacts of sampling, not signals of real structure. Parallel analysis accounts for this by establishing an empirical threshold based on what random noise actually looks like for your specific data dimensions.

A factor from your actual data is retained only if its eigenvalue exceeds the corresponding average eigenvalue from the random datasets. This factor-by-factor comparison is what sets parallel analysis apart from methods that apply a single arbitrary cutoff to all factors.

Researchers in psychology, education, management, and the social sciences rely on parallel analysis during scale development, psychometric validation, and any project where understanding the latent structure of variables matters. The method has been studied extensively in simulation research and consistently ranks among the most accurate approaches for factor retention.

For those interested in parallel analysis applications in research, the methodology extends beyond psychology into any field that uses factor-based measurement models.

Traditional Factor Retention Methods and Their Limitations

Before parallel analysis became widely recommended, researchers relied on two dominant methods for deciding how many factors to keep: the Kaiser criterion and the scree plot. Both are still taught in graduate statistics courses and embedded as defaults in popular statistical software. Both have well-documented problems that lead to poor factor retention decisions.

The Kaiser Criterion: Simple but Flawed

The Kaiser criterion, also known as the eigenvalue-greater-than-one rule or the Kaiser-Guttman rule, says you should retain any factor with an eigenvalue above 1.0. The logic is that a factor should explain at least as much variance as a single original variable to be worth keeping.

The problem is that this rule systematically overestimates the number of factors. Simulation studies have shown that the Kaiser criterion tends to retain too many factors, especially when you have many variables or a moderate sample size. The reason is straightforward: eigenvalues from pure noise data routinely exceed 1.0, particularly in the first several positions. So when you use 1.0 as your cutoff, you are counting noise factors as real ones.

This over-retention has real consequences. You end up interpreting factors that represent random fluctuation rather than meaningful latent constructs. You may report a five-factor solution when the true structure has only three factors, leading to a measurement model that will not replicate in a new sample.

The Kaiser criterion also ignores the fact that the expected eigenvalue from random data varies depending on the position of the factor in the extraction order. The first random eigenvalue is always higher than the second, which is higher than the third. A single fixed cutoff of 1.0 fails to capture this declining pattern.

The Scree Plot: Visual but Subjective

The scree plot, introduced by Raymond Cattell in 1966, plots eigenvalues in descending order and asks the researcher to identify the point where the curve levels off into a shallow slope, often called the elbow. Factors above the elbow are retained, and factors below it are discarded.

This sounds reasonable in theory. In practice, identifying the elbow is notoriously subjective. Different researchers looking at the same scree plot routinely disagree about where the elbow falls. Some see it at the third factor, others at the fourth, and some argue there is no clear elbow at all.

The subjectivity problem becomes worse when eigenvalues decline gradually without a sharp break. Many real-world datasets produce scree plots with ambiguous elbow points, leaving researchers to make judgment calls that introduce researcher degrees of freedom into the analysis.

Another limitation is that the scree plot alone provides no benchmark for comparison. You see your eigenvalues declining, but you have no reference for what a random data eigenvalue curve would look like in the same position. Parallel analysis solves this by overlaying the random data eigenvalue curve directly onto your scree plot.

How Parallel Analysis Works: Step by Step

The parallel analysis procedure follows a clear, repeatable sequence. Here is how it works in practice.

Step 1: Extract eigenvalues from your actual data. Start by computing the correlation matrix from your real dataset of N observations and P variables. Extract the eigenvalues using either principal component analysis or principal axis factoring, depending on your analytical framework. Record these eigenvalues in descending order.

Step 2: Generate random datasets. Create a large number of random datasets, typically 1,000 or more, each with the same number of variables (P) and the same sample size (N) as your real data. Each random dataset consists of uncorrelated variables drawn from a normal distribution, meaning these datasets contain no true latent factors by construction.

Step 3: Extract eigenvalues from each random dataset. For every random dataset, compute the eigenvalues using the same extraction method you used for your real data. This gives you a distribution of eigenvalues for each factor position across all 1,000 random datasets.

Step 4: Compute the average random eigenvalues. For each factor position (first eigenvalue, second eigenvalue, third eigenvalue, and so on), calculate the mean eigenvalue across all random datasets. You can also compute percentiles, such as the 95th percentile, to set a more conservative threshold.

Step 5: Compare your actual eigenvalues to the random averages. Line up your real eigenvalues next to the average random eigenvalues at each position. Starting from the first factor, check whether your actual eigenvalue exceeds the corresponding random data eigenvalue.

Step 6: Retain factors where actual exceeds random. Count the number of consecutive factor positions where your actual eigenvalue is larger than the random data average. This count is your recommended number of factors to retain. Once your actual eigenvalue drops below the random data threshold, you stop, because subsequent factors are indistinguishable from noise.

For example, imagine you have 20 variables and 300 participants. Your first actual eigenvalue is 5.2, and the average first random eigenvalue is 1.6. You retain the first factor. Your second actual eigenvalue is 2.8 versus a random average of 1.4. Retain it. Your third actual eigenvalue is 1.3 versus a random average of 1.25. You might retain this one, but barely. Your fourth actual eigenvalue is 0.9 versus a random average of 1.15. Your actual eigenvalue is now below the random threshold, so you stop and retain three factors.

This factor-by-factor comparison is the heart of why parallel analysis works. Each factor position gets its own empirically derived threshold based on what noise looks like at that specific position, rather than a one-size-fits-all cutoff.

Key Advantages of Parallel Analysis Over Other Methods

Parallel analysis offers several distinct advantages that make it the preferred method for factor retention decisions among methodologists and applied researchers alike.

Superior accuracy in simulation studies. Across dozens of published simulation studies, parallel analysis consistently demonstrates higher accuracy rates than the Kaiser criterion and the unaided scree test. Studies typically find accuracy rates between 80 and 95 percent for parallel analysis, compared to 40 to 60 percent for the Kaiser rule. This accuracy advantage holds across different numbers of true factors, varying sample sizes, and different levels of factor loading strength.

Empirical rather than arbitrary thresholds. The Kaiser rule applies a fixed cutoff of 1.0 to every factor regardless of context. Parallel analysis derives its threshold from the actual sampling distribution of eigenvalues under the null hypothesis of no factor structure. This means the threshold adapts to your specific number of variables, your specific sample size, and the specific position of each factor.

Correction for positive bias in observed eigenvalues. Sample eigenvalues are positively biased estimates of population eigenvalues, meaning they tend to be larger than the true values. This bias is especially pronounced for factors derived from small samples or many variables. Parallel analysis naturally corrects for this bias because the random data eigenvalues capture the same upward bias. When you compare your actual eigenvalues to the random data eigenvalues, the bias cancels out.

Built-in protection against over-retention. The most common error in factor analysis is retaining too many factors, not too few. Both the Kaiser criterion and unaided scree plot interpretation tend toward over-retention. Parallel analysis is specifically designed to prevent this by establishing how large an eigenvalue needs to be to exceed what pure noise would produce.

Addresses sampling variability. Recent research has extended parallel analysis to account for sampling variability in the suggested number of factors itself. Rather than producing a single point estimate, modified approaches like PA* report the proportion of random samples suggesting each possible number of factors, giving researchers a sense of how stable the retention decision is across hypothetical resamples.

Practical Implementation: Tools and Software

You do not need to write complex code from scratch to run a parallel analysis. Several accessible tools handle the computation for you.

R. The R statistical environment offers the most robust parallel analysis tools. The paran package provides a straightforward implementation with options for both PCA and principal axis factoring. The psych package, maintained by William Revelle, includes the fa.parallel function, which generates a visual scree plot with the random data eigenvalues overlaid. A typical call looks like fa.parallel(mydata, fa=”fa”, n.iter=1000), which runs 1,000 iterations and plots your actual eigenvalues against the random averages.

SPSS. SPSS does not include parallel analysis as a built-in menu option, but researchers have developed syntax macros that perform the analysis. Brian O’Connor has published freely available SPSS syntax that generates random datasets and computes the comparison. You paste your correlation matrix or raw data, run the syntax, and receive a table of actual versus random eigenvalues.

SAS. SAS users can implement parallel analysis using PROC FACTOR combined with a data step that generates random correlation matrices. O’Connor has also published SAS macros for this purpose. The approach mirrors the R implementation in logic.

Online calculators. Several websites offer parallel analysis calculators where you input your sample size, number of variables, and actual eigenvalues. The tool generates the random data comparison and returns the recommended number of factors. These are useful for quick checks, but they limit your control over parameters like the number of iterations and the extraction method.

When choosing a tool, the key consideration is whether you need principal component analysis eigenvalues or principal axis factoring eigenvalues. Parallel analysis works with both, but the random data eigenvalues differ depending on the method because PCA and PFA treat communalities differently. Make sure your random data comparison uses the same extraction method as your actual factor analysis.

The psych package in R is the tool I recommend for most researchers. It handles the entire workflow, produces a publication-quality scree plot with the random data overlay, and gives you full control over iteration counts and extraction methods. Best of all, it is free and open source.

For those exploring parallel analysis applications in research across different domains, the same tools apply regardless of whether you are working with psychological scales, educational assessments, or organizational survey data.

Common Pitfalls and Misinterpretations

Despite its accuracy advantages, parallel analysis is sometimes misapplied or misinterpreted. Here are the most common errors I see researchers make, drawn from forum discussions and published critiques.

Treating the threshold as a single fixed value. Some researchers report a single eigenvalue cutoff from parallel analysis and apply it to all factor positions. This defeats the purpose of the method. Each factor position has its own random data threshold, and these thresholds decline from the first factor to the last. Always compare your actual eigenvalue to the random eigenvalue at the same position, not to a single aggregate number.

Confusing PCA and factor analysis criteria. Parallel analysis produces different random data eigenvalues depending on whether you use principal component analysis or common factor analysis as the extraction method. PCA eigenvalues from random data tend to be higher than PFA eigenvalues because PCA uses 1.0 on the diagonal of the correlation matrix while PFA uses estimated communalities. If your actual factor analysis uses PCA but your parallel analysis uses PFA thresholds (or vice versa), your retention decision will be off. Match the extraction methods.

Ignoring sample size effects. Parallel analysis is sensitive to sample size, which is actually a strength. With small samples, the gap between actual and random eigenvalues narrows, making it harder to distinguish true factors from noise. If you have fewer than 100 participants and many variables, even parallel analysis may struggle to detect the true structure. Researchers should interpret results cautiously in these conditions and consider collecting more data.

Including the elbow factor inconsistently. When the scree plot and parallel analysis disagree, some researchers try to split the difference by retaining factors up to and including the elbow identified visually. This introduces the subjectivity problem of the scree plot back into the analysis. If you choose to use parallel analysis, commit to its decision rule rather than overriding it with visual judgment.

Different software producing different results. Researchers sometimes panic when R, SPSS, and an online calculator give slightly different recommended factor counts. This usually stems from differences in the number of random datasets generated, the percentile used for the threshold (mean versus 95th percentile), or the extraction method. These differences are usually small and become negligible with 1,000 or more iterations. Standardize your parameters across tools.

Assuming parallel analysis never makes errors. While parallel analysis is highly accurate, it is not infallible. In conditions with very weak factor loadings, highly correlated factors, or unusually structured data, parallel analysis can under-retain or over-retain factors. Treat its recommendation as strong evidence rather than absolute truth, and cross-check with theory and substantive interpretation.

FAQs

What is parallel analysis in factor analysis?

Parallel analysis is a method for determining how many factors to retain in exploratory factor analysis. It works by comparing the eigenvalues from your actual data to the average eigenvalues from multiple randomly generated datasets that have the same sample size and number of variables but contain no true factor structure. A factor is retained only when its actual eigenvalue exceeds the corresponding random data eigenvalue at the same position.

What is the Kaiser criterion for retaining factors in a factor analysis?

The Kaiser criterion, also called the eigenvalue-greater-than-one rule, retains any factor whose eigenvalue exceeds 1.0. The rationale is that a meaningful factor should explain at least as much variance as one original variable. However, this rule tends to overestimate the number of factors because eigenvalues from pure noise data frequently exceed 1.0, especially in the first few positions.

What does a scree plot tell you?

A scree plot displays eigenvalues in descending order, helping you identify the point where the curve flattens into a shallow slope. Factors above this elbow represent meaningful structure, while factors below it represent noise. The method is subjective because different researchers may identify the elbow at different points, leading to inconsistent factor retention decisions.

What are the eigenvalues of a scree plot?

The eigenvalues shown on a scree plot represent the amount of variance in the original variables that each successive factor accounts for. The first factor always has the largest eigenvalue, and each subsequent factor explains progressively less variance. In parallel analysis, these actual eigenvalues are plotted alongside the average eigenvalues from random data so you can see exactly where your real data exceeds what noise would produce.

Conclusion

Parallel analysis improves factor retention decisions by replacing arbitrary cutoffs with empirically derived thresholds based on what random noise actually looks like for your data. It corrects for the positive bias in sample eigenvalues, adapts to your specific sample size and variable count, and consistently outperforms the Kaiser criterion and unaided scree plot interpretation in simulation studies.

If you are running an exploratory factor analysis, make parallel analysis your primary retention method. Use the Kaiser criterion only as a secondary reference point, and let the scree plot serve as a visual companion rather than your main decision tool. The R psych package and the paran package make implementation straightforward, even for researchers without programming experience.

The result is a factor structure you can defend with evidence, replicate across samples, and trust as a foundation for your measurement model. In a research landscape where factor retention decisions shape everything from scale validation to theoretical claims, parallel analysis factor retention methodology gives you the rigor your work deserves.

Leave a Comment