How Many Items You Should Write Per Construct? (2026 Guide)

When you develop a measurement instrument, one of the first questions you face is how many items you should write per construct. It is a question that trips up graduate students and experienced researchers alike, because the answer depends on multiple factors working together. The number you choose affects reliability, validity, and the overall quality of your final scale.

Our team has reviewed the most cited scale development literature and real-world research examples to give you a clear, practical answer. The short version: you should write two to three times more items than you plan to keep in your final scale. If your final instrument needs 10 items per construct, you start with an initial pool of 20 to 30 items for that construct.

This over-representation strategy is not optional. DeVellis, one of the most frequently cited authorities in scale development, recommends this ratio consistently across his work from 2003 through 2016. Clark and Watson’s 1995 framework for construct validation reinforces the same principle with a slightly different framing. The goal is to generate enough items that statistical reduction can occur without sacrificing measurement quality.

In this guide, we break down the exact numbers, the reasoning behind them, and the practical steps you should take. You can see an example of scale development in practice from published research to understand how these principles play out in real studies. Let’s get into the specifics.

How Many Items You Should Write Per Construct When Developing a Scale

Here is the direct answer most researchers are looking for: write 3 to 5 items per construct for your final scale, but generate an initial item pool of 2 to 3 times that number. This means your starting pool for a single construct should contain 6 to 15 items at minimum, often more for complex constructs.

The rule comes down to a simple ratio. For every 1 item you want in your final instrument, write 2 to 3 candidate items. Matt C. Howard, a methodologist who writes extensively on scale construction, puts it plainly: if you need 10 items, make 30. If you need 20, make 50. You can always remove items later, but you cannot add quality items after data collection is complete.

Ohio State University’s scale development guidance echoes this recommendation. Their item pool should contain two to three times as many items as will appear on the final instrument. This gives you enough material to work with during item reduction without running short.

How does this break down by construct complexity? For a narrow, well-defined construct like “math anxiety,” you might aim for a final scale of 4 to 6 items and start with 12 to 18 candidate items. For a broader construct like “self-directed learning skills,” you might need 10 to 15 final items across sub-dimensions, meaning your initial pool should include 30 to 45 items. Complex, multidimensional constructs require more items because they capture more facets of the underlying concept.

The key insight is that construct breadth drives item count. A construct with a clearly bounded content domain needs fewer items. A construct with fuzzy boundaries or multiple sub-components needs more items to adequately cover the full conceptual space.

Why You Need More Items Than You Think

Many first-time scale developers underestimate how many items they will lose during the development process. Items drop out at every stage, and the losses add up faster than you might expect.

Expert review typically eliminates 10 to 20 percent of your initial pool. Reviewers flag items that are ambiguous, redundant, or off-target. Some items seem clear to you but confuse subject-matter experts, which is a strong signal they will confuse respondents too.

Pilot testing removes another chunk of items. When you administer your draft scale to a small sample, certain items will show poor discrimination. They may correlate weakly with the total score or fail to load cleanly on the expected factor. These items have to go.

Factor analysis is where the largest reduction happens. During exploratory factor analysis, you identify which items cluster together and which ones cross-load onto multiple factors. Cross-loading items are typically removed. Items with factor loadings below 0.40 (or below 0.32 in more lenient standards) get cut. This process alone can eliminate 30 to 50 percent of your remaining pool.

Finally, reliability testing removes items that drag down Cronbach’s alpha. An item-total correlation below 0.30 is a red flag. If removing an item increases your alpha coefficient, that item is working against your scale.

When you add up these losses across expert review, pilot testing, factor analysis, and reliability checks, starting with only a few extra items is a recipe for disaster. You end up with too few items to form a reliable scale, and you have no replacements available. That is why the 2:1 to 3:1 ratio exists.

Defining Your Construct Before Writing Items

Before you write a single item, you need a clear construct definition. This step directly determines how many items you will need, because it defines the boundaries of what you are measuring.

A construct is the theoretical concept you want to measure, such as “employee engagement” or “math anxiety.” Clark and Watson’s 1995 framework, published in Psychological Assessment, argues that construct definition is the most important and most frequently skipped step in scale development. Without a clear definition, you cannot determine the content domain, and without a content domain, you cannot know how many items are needed to cover it.

Narrow constructs have small content domains. “Fear of public speaking” is fairly specific, so a 4 to 6 item final scale might capture it well. Broad constructs like “general anxiety” cover many facets including cognitive, physical, and behavioral dimensions. These need more items to adequately represent each facet.

You can explore attitude scale development methodology from published research to see how construct definition shapes the item generation process. In that case study, the researchers mapped out the full content domain before writing a single item, which directly informed their item count decisions.

We recommend creating a conceptual map or diagram of your construct before item writing. List the sub-dimensions, facets, and behavioral indicators. Each facet in your map should be represented by at least 3 to 5 items in your initial pool. This ensures the full content domain is covered.

Item Pool Over-Representation Strategy

The item pool over-representation strategy is the backbone of how many items you should write per construct when developing a scale. It is the principle that separates well-developed instruments from shaky ones.

DeVellis recommends generating an initial item pool that is roughly twice the size of the intended final scale. Other methodologists push this to three times the final count. The exact ratio depends on how confident you are in your items and how complex your construct is.

Here is a practical framework for determining your initial pool size. Take your target final scale length and multiply it. For a simple, unidimensional construct, multiply by 2. You need 8 items final, so write 16. For a moderately complex construct with 2 or 3 sub-dimensions, multiply by 2.5. You need 12 items final, so write 30. For a complex, multidimensional construct, multiply by 3 or more. You need 20 items final, so write 60 or more.

When should you go higher than 3 times? Consider expanding your initial pool beyond the 3x mark in several situations. If your construct is new and has no existing validated scales to draw from, write more items because you are exploring uncharted territory. If your target population is heterogeneous, you need more items to capture diverse perspectives. If you are unsure about the factor structure, additional items give you flexibility during analysis.

One researcher in a forum discussion described testing 128 items on 130 participants for a complex construct. While that is on the high end, it illustrates how quickly item counts can grow when you take a thorough approach to over-representation.

The downside of writing too many items is participant fatigue. Long surveys produce careless responses and missing data. You need to balance thoroughness with practicality. A 60-item initial pool administered to 200 participants is manageable. A 200-item pool is not, unless you split it across multiple forms or samples.

Common Item Writing Mistakes to Avoid

Writing more items only helps if those items are well-crafted. Poorly written items get removed during analysis regardless of how many you generate. Here are the most common mistakes that waste your item pool.

Double-barreled items ask about two things at once. “I feel confident and motivated at work” is double-barreled because confidence and motivation are separate constructs. If a respondent feels confident but not motivated, they cannot answer honestly. Always check each item for a single, clear referent.

Leading questions push respondents toward a particular answer. “Most experts agree that exercise improves mood, do you agree?” leads the respondent by implying consensus. Neutral wording is essential for valid measurement.

Negatively worded items can cause problems. Items like “I do not enjoy social gatherings” require respondents to reverse their thinking, which increases cognitive load and produces response errors. Research consistently shows that negatively worded items form their own factor in factor analysis, distorting your results. If you must use reverse-scored items for methodological reasons, keep them simple and unambiguous.

Ambiguous language creates unreliable responses. Words like “often,” “sometimes,” and “rarely” mean different things to different people. “I often feel stressed” could mean daily to one person and weekly to another. Use concrete behavioral references instead: “I felt stressed at least 3 days this week.”

Overly complex vocabulary excludes respondents with lower reading levels. Unless your target population consists of subject-matter experts, write items at a 6th to 8th grade reading level. This ensures comprehension and reduces response bias.

Double negatives confuse everyone. “It is not uncommon for me to feel anxious” forces the respondent to process two negations. Rewrite it as “It is common for me to feel anxious” for clarity.

Reviewing your initial pool against this checklist before expert review can save significant time. Every item you catch and fix now is one fewer item that gets removed later.

Response Format Selection: How It Affects Item Count

The response format you choose interacts with your item count in important ways. A Likert scale with more response points captures more variance per item, which can reduce the total number of items you need.

Five-point Likert scales are the most common choice in published research. They offer a balance between discrimination and respondent ease. Five points give respondents enough options to express nuanced agreement without overwhelming them.

Seven-point scales provide finer discrimination. They are useful when your construct has subtle variations that a 5-point scale might collapse. The tradeoff is that respondents sometimes struggle to distinguish between adjacent points on a 7-point scale, especially points 2 and 3 or 5 and 6.

Research from MeasuringU suggests that single-item measures can be adequate for very simple constructs, but multi-item measures are always better for complex constructs. The number of response points partially compensates for having fewer items, but it does not replace the need for multiple items per construct.

If you use a 10-point scale, you capture maximum variance per item. Some researchers argue this allows fewer items overall. However, 10-point scales can introduce their own biases, including central tendency bias and endpoint avoidance. The conventional recommendation is to stick with 5 or 7 points unless you have a specific reason to deviate.

For most scale development projects, a 5-point or 7-point Likert scale paired with 4 to 8 final items per construct produces reliable, valid results. The response format does not change the initial item pool size you need, but it does influence how much information each surviving item contributes.

The Item Reduction Process

Item reduction is where your over-represented item pool gets trimmed down to the final scale. Understanding this process helps you appreciate why you need so many initial items.

Step 1: Expert Review. Assemble a panel of 5 to 10 subject-matter experts. This panel size is the most commonly recommended in the scale development literature. Ask each expert to rate each item on relevance, clarity, and representativeness of the construct. Calculate a content validity index (CVI) for each item and for the overall scale. Items with an item-level CVI below 0.78 should be revised or removed.

Step 2: Pilot Testing. Administer the surviving item pool to a small sample of your target population, typically 30 to 50 participants. This step identifies items that are confusing, have restricted variance, or show ceiling and floor effects. Compute descriptive statistics for each item, including mean, standard deviation, skewness, and kurtosis. Items with extreme distributions may need revision.

Step 3: Exploratory Factor Analysis (EFA). Use a larger sample to conduct EFA. The minimum sample size depends on your item count and the strength of factor loadings, but a common guideline is at least 10 participants per item. Examine factor loadings, cross-loadings, and the number of factors extracted. Retain items with loadings of 0.40 or higher on a single factor. Remove items that cross-load on multiple factors at 0.32 or higher.

Step 4: Reliability Analysis. Compute Cronbach’s alpha for each subscale. Examine item-total correlations and alpha-if-item-deleted statistics. An item-total correlation below 0.30 is a candidate for removal. If deleting an item increases alpha substantially, remove it.

Step 5: Confirmatory Factor Analysis (CFA). Use a separate sample to confirm the factor structure identified in EFA. Evaluate model fit using indices like CFI (above 0.95), RMSEA (below 0.06), and SRMR (below 0.08). This step validates that your remaining items form a coherent measurement model.

You can explore advanced item response theory analysis methods for even more precise item selection. IRT models provide detailed information about item difficulty and discrimination, which is especially useful for ability and aptitude scales.

Throughout this process, expect to lose 40 to 60 percent of your initial item pool. This is normal and expected. It is exactly why you started with 2 to 3 times more items than your target final scale length.

Item-to-Participant Ratio: A Critical Guideline

The number of items you write directly determines the sample size you need. The item-to-participant ratio is one of the most common pain points researchers raise in forums and discussions.

The most widely cited guideline comes from Hair and colleagues: a minimum of 10 participants per item for factor analysis. If your initial item pool has 50 items, you need at least 500 participants for the factor analysis phase. For published scale development, many researchers aim for 15 to 20 participants per item to ensure stable factor solutions.

This ratio applies to the factor analysis stage, not to pilot testing. Pilot testing can use a smaller sample because you are checking for clarity and basic psychometric properties, not establishing a factor structure.

Some methodologists argue that the absolute minimum sample size matters more than the ratio. Comrey and Lee suggested that 100 participants is poor, 200 is fair, 300 is good, 500 is very good, and 1000 is excellent for factor analysis. Regardless of how many items you have, aim for at least 300 participants for a credible factor analysis.

This ratio creates a practical tension. More items mean you need more participants, which means more time and resources. This is another reason to be strategic about your initial item pool. Writing 100 items when 40 would suffice creates an unnecessary burden on your data collection plan.

The balance is this: write enough items to cover your construct thoroughly and allow for reduction, but not so many that your sample size requirements become impractical. The 2:1 to 3:1 ratio strikes this balance for most research contexts.

Practical Checklist for Item Generation

Before we move to frequently asked questions, here is a practical checklist you can follow when deciding how many items to write per construct for your scale.

First, define your construct in one or two clear sentences. Identify whether it is unidimensional or multidimensional. If multidimensional, list the sub-dimensions and decide how many final items each sub-dimension needs.

Second, decide on your target final scale length. Most constructs work well with 4 to 8 final items per dimension. If you have 3 sub-dimensions, your final scale might contain 12 to 24 items total.

Third, multiply your target by 2 to 3. This gives you your initial item pool size. For a 12-item final scale, write 24 to 36 items.

Fourth, map your items to your construct definition. Each facet or sub-dimension should have at least 3 to 5 candidate items in the initial pool. No facet should be underrepresented.

Fifth, review every item against the common mistakes checklist. Fix double-barreled items, leading questions, ambiguous language, and double negatives before expert review.

Sixth, plan your sample size. Calculate 10 to 20 participants per item in your initial pool. If your pool has 36 items, plan for 360 to 720 participants.

Seventh, document your decisions. Scale development requires transparency. Record why you chose your item count, how you mapped items to facets, and what criteria you will use for reduction. This documentation supports your content validity argument and makes your scale development process reproducible.

FAQs

How many items should a Likert scale have?

A Likert scale measuring a single construct should typically have 4 to 8 items in its final form. During development, you should write 2 to 3 times that number as your initial item pool. For example, if your final scale will have 6 items, start with 12 to 18 candidate items to allow for statistical reduction through factor analysis and reliability testing.

How many participants do I need for scale development?

The most common guideline is a minimum of 10 participants per item for factor analysis. If your initial item pool has 40 items, aim for at least 400 participants. For published research, many methodologists recommend 15 to 20 participants per item. Regardless of item count, a minimum of 300 participants is considered the floor for credible factor analysis results.

Should I use a 5 or 7 point Likert scale?

Both 5-point and 7-point Likert scales are widely used and accepted. Five-point scales are simpler for respondents and work well for most constructs. Seven-point scales provide finer discrimination and are useful when subtle differences in respondent attitudes matter. Research shows that reliability differences between 5 and 7 point scales are minimal, so choose based on your construct complexity and respondent population.

What is the minimum number of items per construct?

The practical minimum for a reliable unidimensional construct is 3 to 4 items in the final scale. Fewer than 3 items makes factor analysis unreliable and limits your ability to compute internal consistency statistics. However, you should write far more than 3 during the initial item generation phase, typically 6 to 12 candidate items, to allow for reduction.

How many items should survive after factor analysis?

Typically 40 to 60 percent of your initial item pool will survive the full reduction process including expert review, pilot testing, factor analysis, and reliability testing. If you start with 30 items, expect 12 to 18 to remain in the final scale. The exact number depends on item quality, construct complexity, and the statistical criteria you apply during reduction.

How do I know which items to remove during item reduction?

Remove items based on multiple criteria applied sequentially. Start with expert review using content validity index scores below 0.78. Then check descriptive statistics for restricted variance or extreme distributions. During factor analysis, remove items with loadings below 0.40 or cross-loadings above 0.32. Finally, check item-total correlations below 0.30 and items that increase Cronbach’s alpha when deleted.

Conclusion

Knowing how many items you should write per construct when developing a scale comes down to one core principle: over-generate, then reduce. Write 2 to 3 times more items than your final scale needs, review them with experts, pilot test, and let factor analysis and reliability statistics guide your final selection.

Start by defining your construct clearly, mapping its content domain, and generating an item pool that covers every facet. A final scale of 4 to 8 items per construct is typical, which means your initial pool should contain 8 to 24 items depending on construct complexity. Plan your sample size accordingly, with at least 10 participants per item for the factor analysis phase.

The researchers who produce the most reliable, well-validated scales are the ones who embrace over-representation from the start. Write more items than you think you need, apply rigorous reduction criteria, and trust the process to yield a measurement instrument that stands up to scrutiny.

Leave a Comment