What Counts as Acceptable Test-Retest Reliability for a Survey Instrument? (September 2026) Guide

Acceptable test-retest reliability for a survey instrument is generally an intraclass correlation coefficient (ICC) of at least 0.70 for research use, 0.80 or higher for clinical decisions, and 0.90 or above for individual-level judgments. These thresholds come from decades of psychometric work by Fleiss, Cicchetti, Koo, and others.

In this guide, I will explain what test-retest reliability measures, which ICC values count as acceptable, how to choose the right statistical model, and how long you should wait between administrations. Whether you are validating a new questionnaire or checking the stability of an existing scale, this article will give you practical, citation-backed standards you can use in 2026.

Our team has helped researchers evaluate dozens of survey instruments, and the same questions come up every time: Which coefficient should I report? What time interval is appropriate? Is Pearson correlation enough? I will answer all of these with clear numbers and real-world examples.

If you are still building your instrument, you may want to review the full survey instrument development process before you lock in your reliability study design. That article covers the broader psychometric pipeline and shows where reliability fits into the overall validation workflow.

Test-retest reliability is one piece of the larger psychometric picture, but it is often the piece that reviewers and dissertation committees scrutinize most closely. Getting it right can mean the difference between a published instrument and one that never leaves the pilot phase.

Throughout this article, I will cite the specific researchers and professional standards behind every threshold. My goal is to give you a reference you can cite in your own work, whether you are writing a methods section, responding to a reviewer comment, or designing a reliability study from scratch.

Table of Contents

What Is Test-Retest Reliability?

Test-retest reliability measures how stable survey scores are when the same respondents complete the instrument on two separate occasions. If scores are consistent over time, the survey is producing reliable measurements; if they fluctuate randomly, the reliability is weak.

This type of reliability is sometimes called temporal stability because it focuses on whether scores stay the same across time. It is different from internal consistency, which measures whether items within a single administration agree with each other, and from inter-rater reliability, which measures agreement between different observers.

To run a test-retest study, you administer the same survey to the same group of people twice, with a carefully chosen time interval in between. You then compare the scores from the two administrations using a statistic that captures agreement, not just correlation.

The result is a single coefficient that tells you how much of the variation in scores is due to true differences between people rather than measurement error. Higher coefficients mean less noise and more signal.

Why Test-Retest Reliability Matters

Unstable measurements make it impossible to trust research conclusions. If a depression scale gives one score today and a very different score next week for no real reason, you cannot tell whether a treatment helped or whether the instrument simply varied.

High test-retest reliability does not prove that a survey measures what it claims. It only shows that the survey measures something consistently. Validity is a separate question, but reliability is usually a prerequisite for validity.

Think of reliability as the floor of measurement quality. Without it, every other psychometric property becomes harder to evaluate because the scores themselves are noisy.

For applied settings, reliability also has ethical implications. If a clinical screening tool is unreliable, patients may be misclassified, denied services, or given unnecessary treatment. The stakes rise as the consequences of the measurement increase.

How Test-Retest Differs From Other Types of Reliability

Researchers often confuse test-retest reliability with related concepts. Internal consistency, measured by Cronbach’s alpha, asks whether items within a single survey form measure the same construct. It says nothing about stability over time.

Inter-rater reliability asks whether different observers or judges produce similar scores when rating the same person. It is relevant when the measurement depends on a human rater, but it is separate from temporal stability.

Parallel-forms reliability asks whether two versions of the same instrument produce similar scores. Split-half reliability asks whether the first half of a test agrees with the second half. All of these are useful, but none of them answers the question this article focuses on: does the same form produce the same scores across time?

That question matters because many constructs, such as personality, attitudes, and chronic conditions, are expected to remain relatively stable. If your scores shift dramatically over two weeks, and the underlying construct should not have changed, your measurement is probably unreliable.

What Counts as Acceptable Test-Retest Reliability?

For most survey research, an ICC value below 0.50 is considered poor, 0.50 to 0.75 is moderate, 0.75 to 0.90 is good, and above 0.90 is excellent. These cutoffs are the most widely cited standards in psychometric literature.

The table below combines thresholds from Fleiss (1986), Cicchetti (1994), and Koo and Li (2016). It gives you a single reference you can use when interpreting test-retest reliability for a survey instrument.

ICC RangeInterpretationCommon Use Case
Below 0.50PoorNot acceptable for most research
0.50 to 0.75ModerateAcceptable for exploratory research
0.75 to 0.90GoodAcceptable for most research instruments
Above 0.90ExcellentRequired for clinical or diagnostic decisions

Koo and Li recommend slightly stricter language than Fleiss. They label values above 0.90 as excellent and reserve that range for measurements that will guide individual-level decisions. Cicchetti uses similar language for clinical research: 0.90 and above is excellent, 0.80 to 0.89 is good, 0.70 to 0.79 is fair, and below 0.69 is poor.

These three sets of standards are consistent enough that most researchers cite them interchangeably. The small differences in wording should not change your interpretation in practice. If your ICC falls near a boundary, treat it conservatively.

When 0.70 Is Acceptable and When It Is Not

A reliability coefficient of 0.70 is often treated as the minimum acceptable value for group-level research. This means the instrument is stable enough to detect differences between groups or track changes across populations.

For individual decisions, such as diagnosing a patient or placing a student in a program, the standard rises to 0.80 or 0.90. When a single person’s score determines an outcome, measurement error has more serious consequences, so stronger reliability is required.

Some methodologists have argued that 0.70 is too lenient even for group research. They point out that reliability affects statistical power, effect size estimation, and the replicability of findings. From that perspective, aiming for 0.80 even in group studies is a reasonable goal.

The key insight is that acceptability depends on context. A coefficient that looks fine for a satisfaction survey may be unacceptable for a clinical screening tool. Always justify your threshold relative to how the scores will be used.

Historical Context Behind These Thresholds

The 0.70 cutoff has roots in the work of Nunnally, who suggested it for early-stage research instruments in the 1970s. Fleiss published his interpretation guidelines in 1986, focusing on agreement among raters but extending to stability. Cicchetti refined these categories in 1994 with specific attention to clinical instruments.

Koo and Li updated the guidance in 2016 to address modern ICC models and to emphasize the importance of reporting confidence intervals. Their work is now one of the most cited sources in reliability reporting, and their thresholds are widely used in peer-reviewed publications.

Together, these researchers built the framework most psychometricians rely on today. The thresholds have remained stable over time because they work well in practice, even as the statistical models used to estimate reliability have evolved.

Why ICC Is Better Than Pearson r for Test-Retest

Many researchers make the mistake of reporting Pearson correlation between Time 1 and Time 2 scores. Pearson r measures linear association, not agreement. Two scores can correlate perfectly even when one is consistently higher than the other, which is a problem for reliability.

ICC measures agreement. It evaluates whether the actual scores are similar, not just whether they move in the same direction. ICC also accounts for systematic shifts in scores across administrations, which Pearson r ignores.

For example, if every respondent scores 5 points higher on the retest, Pearson r could still be 1.00, suggesting perfect reliability. ICC would detect that shift and report a lower value, giving a more honest picture of stability.

This distinction matters because systematic shifts are common in real data. Respondents may become more comfortable with the items, reflect more carefully on their answers, or feel different mood states at retest. Pearson r treats all of these as random noise; ICC treats them as meaningful deviations from agreement.

Mathematical Explanation in Plain Language

Pearson r only looks at how scores rank relative to each other. If everyone moves up by the same amount, the rank order stays the same, and Pearson r reports a perfect relationship. But the scores themselves are not the same, so reliability is not actually perfect.

ICC decomposes the total variance in scores into three components: variance due to real differences between people, variance due to systematic differences between time points, and residual error. A high ICC means that between-person variance dominates, while time-point shifts and error are small.

This decomposition is what makes ICC the right tool for test-retest reliability. It captures the full picture of score stability, including systematic effects that Pearson r misses entirely.

When Pearson r Might Be Acceptable

There are a few situations where researchers still report Pearson r. If you are working with historical data, publishing in a field where Pearson r is the convention, or doing a quick exploratory analysis, Pearson r can be informative as long as you acknowledge its limitations.

However, for formal psychometric validation, ICC should always accompany or replace Pearson r. Reviewers in psychology, education, and health research increasingly expect ICC, and papers that report only Pearson r may face revision requests.

Our recommendation is to report ICC as your primary statistic and use Pearson r only as supplementary information if needed. This approach keeps your work aligned with current best practices while acknowledging the historical context of reliability research.

ICC Interpretation Guide for Survey Researchers

The intraclass correlation coefficient comes in several forms. Choosing the wrong one can change your results, so it is worth understanding the basic options.

ICC models differ along two dimensions: the type of effects (one-way, two-way random, or two-way mixed) and the type of agreement (single vs average measures, absolute agreement vs consistency). Each combination answers a slightly different question, and reporting the wrong one can mislead readers.

One-Way Random, Two-Way Random, and Two-Way Mixed

One-way random effects models treat each measurement as if it came from a larger pool of possible measurements. This model is rarely used in test-retest studies because it assumes that each occasion is randomly drawn from a population of occasions.

Two-way random effects models separate participant variance from measurement occasion variance. Both participants and occasions are treated as random samples from larger populations. This model is useful when you want to generalize beyond the specific occasions in your study.

Two-way mixed effects models treat the measurement occasions as fixed rather than random. This is the most common choice for test-retest reliability because researchers typically care about the specific time points they used, not a hypothetical population of occasions.

Absolute Agreement vs Consistency

Absolute agreement asks whether scores are identical across administrations. Consistency asks whether scores are proportional, allowing for an additive or multiplicative shift. For test-retest reliability, absolute agreement is usually preferred because stability means the same score, not just the same rank order.

Imagine a bathroom scale that reads 2 pounds heavier every time you use it. A consistency ICC would say the scale is reliable because the rank order of weights is preserved. An absolute agreement ICC would flag the systematic bias and report lower reliability.

For most survey applications, absolute agreement is the better choice. It matches what researchers usually mean when they say that scores should be stable over time.

Single Measures vs Average Measures

Single-measures ICC applies when you will use a single administration score in practice. Average-measures ICC applies when you will average across multiple administrations. Most survey research uses single-measures ICC because the final instrument is administered once.

Average-measures ICC is always higher than single-measures ICC because averaging reduces error. If you report average-measures ICC, make sure that averaging is what you actually plan to do with the instrument.

ICC Model Selection Summary

ModelTypeUse When
ICC(1,1)One-way random, single measuresEach participant measured by different raters or occasions treated as random
ICC(2,1)Two-way random, single measuresParticipants and occasions are both random samples from larger populations
ICC(3,1)Two-way mixed, single measuresOccasions are fixed; common for test-retest studies
ICC(2,k)Two-way random, average measuresScores will be averaged across k occasions in practice

Koo and Li recommend ICC(3,1) with absolute agreement for most test-retest studies of survey instruments. This model treats the two occasions as fixed and asks whether single-administration scores agree across them.

Sample Size Requirements

Reliability estimates become unstable with small samples. A common rule of thumb is to use at least 30 participants, but more is better. If you expect the ICC to be around 0.80, a sample of 50 to 100 respondents will usually produce a reasonably precise confidence interval.

For pilot work, smaller samples are acceptable, but the resulting confidence interval will be wide. Always report that interval so readers can judge how much uncertainty surrounds your point estimate.

Shoukri and colleagues published sample size tables for ICC estimation. Their work shows that for an expected ICC of 0.75 and a desired confidence interval width of 0.20, you need roughly 50 participants. For narrower intervals, the required sample size grows quickly.

If your sample is small, consider reporting the confidence interval alongside your point estimate. This is more informative than a single number and helps readers interpret the precision of your findings.

Time Interval Guidelines for Test-Retest Studies

The length of time between administrations affects the reliability coefficient. If the interval is too short, memory effects inflate reliability. If it is too long, real changes in the construct deflate reliability.

Choosing the right interval is one of the most consequential design decisions in a test-retest study. A poorly chosen interval can make a reliable instrument look unreliable, or an unreliable one look acceptable.

Short Intervals (1 to 3 Days)

Use a short interval when you want to minimize true change in the underlying trait. This works well for stable characteristics like personality, cognitive ability, or demographic knowledge. Memory effects are still a risk, so keep the gap long enough that respondents cannot recall their exact answers.

Researchers sometimes use alternate item orders or reword items slightly to reduce memory bias. These strategies help, but they also introduce their own sources of variation, so use them carefully.

Medium Intervals (1 to 4 Weeks)

Two to four weeks is the most common time frame for survey test-retest studies. It is short enough that stable traits remain stable, but long enough that most respondents forget their original responses. Many researchers consider two weeks a practical default.

This window is widely used in published reliability studies. It balances the risk of memory effects against the risk of true change, making it a reasonable starting point if you have no specific reason to choose a different interval.

Long Intervals (1 Month or More)

Long intervals are appropriate for traits that should remain stable over months, such as certain attitudes or chronic health conditions. They are also useful when you want to evaluate whether the instrument captures real change rather than random noise. Be aware that longer intervals usually produce lower ICC values.

Some constructs, such as political opinions or health symptoms, can shift meaningfully over a month. If you use a long interval, be prepared to interpret lower ICC values in light of possible true change rather than assuming the instrument is unreliable.

Time Interval Decision Framework

Construct StabilitySuggested IntervalExample Constructs
Highly stable1 to 2 weeksPersonality, basic skills, values
Moderately stable2 to 4 weeksAttitudes, self-efficacy, anxiety
Susceptible to change1 to 3 monthsMood, pain, acute symptoms
Seasonal or contextual2 to 6 weeks with controlsJob satisfaction, classroom climate

When in doubt, choose an interval that matches the expected rate of change in your construct. If the trait changes quickly, a short interval is safer. If it changes slowly, a longer interval is fine and may even be more informative.

Factors That Affect Test-Retest Reliability

Reliability is not a fixed property of a survey. It depends on the population, the conditions, and the way the study is run.

Construct Stability

Some traits are genuinely more stable than others. A measure of gender identity will usually be more stable over time than a measure of daily mood. Reliability must be interpreted with the construct in mind.

This is why the same ICC value can be impressive for a mood scale and disappointing for a personality test. The context of the construct sets the benchmark for what counts as acceptable.

Sample Heterogeneity

ICC values tend to be higher when participants vary more on the measured trait. A homogeneous sample compresses score variance, which can make the instrument look less reliable than it is.

This is a well-known issue in reliability research. If your sample includes only people who are very similar to each other, the ICC may underestimate the true reliability because there is less between-person variance to capture.

One way to address this is to recruit a diverse sample that spans the full range of the trait. The resulting ICC will be more representative of how the instrument performs in real-world use.

Number of Items

Longer scales usually produce more reliable scores because random error cancels out across items. If your test-retest reliability is low, check whether the scale has too few items before blaming the construct.

This is related to the Spearman-Brown prophecy formula, which predicts how reliability changes as items are added or removed. In practice, adding well-written items to a short scale can meaningfully improve both internal consistency and test-retest reliability.

Memory Effects and Practice Effects

If respondents remember their earlier answers, they may repeat them even if their true state has changed. This inflates reliability artificially. Using alternate item orders or a longer retest interval can reduce this bias.

Practice effects are similar but involve improvement due to familiarity rather than direct recall. Cognitive tests are especially vulnerable to practice effects, which is why many neuropsychological instruments use alternate forms for retesting.

Motivation and Context

Fatigue, distraction, and changes in motivation can all increase score variability. Administering the survey in similar conditions at both time points helps keep these factors constant.

Small details, such as time of day, testing location, and instructions given to respondents, can affect scores. Standardizing these conditions reduces extraneous variance and gives you a cleaner estimate of true reliability.

Test-Retest Reliability vs Internal Consistency

These two types of reliability answer different questions. Internal consistency, usually reported as Cronbach’s alpha, asks whether items within one administration measure the same thing. Test-retest reliability asks whether total scores are stable across two administrations.

Cronbach’s alpha values of 0.70 to 0.95 are generally considered acceptable for research. Values above 0.95 may indicate redundancy, not better measurement. Test-retest reliability should be reported alongside internal consistency because a survey can have high internal consistency but low temporal stability.

When you read that a scale has a Cronbach’s alpha of 0.90, you know its items hang together. When you read that its two-week ICC is 0.85, you know its scores are stable over time. Both pieces of information are useful.

Reporting only one type of reliability gives an incomplete picture. A scale with alpha 0.90 and ICC 0.45 is internally consistent but unstable over time. A scale with alpha 0.70 and ICC 0.88 has weaker item cohesion but stronger temporal stability. Both patterns have implications for how the instrument should be used.

How to Interpret Both Coefficients Together

Internal consistency tells you whether the items measure a single underlying construct. Test-retest reliability tells you whether scores on that construct remain stable over time. Together, they describe two different aspects of measurement quality.

If both are high, the instrument is performing well on both dimensions. If alpha is high but test-retest is low, the items may be well-written but the construct itself may be unstable, or the time interval may be too long. If alpha is low but test-retest is high, the items may be heterogeneous but the total score may still be stable.

Standard Error of Measurement and Reliability

The standard error of measurement (SEM) is another way to express reliability in practical units. SEM tells you how much a single score is likely to differ from the true score due to measurement error.

SEM is calculated from the standard deviation of the scores and the reliability coefficient. Higher reliability produces smaller SEM, which means scores are more precise.

Researchers sometimes prefer SEM over ICC because it is expressed in the same units as the original scores. A depression scale with a standard deviation of 10 and an ICC of 0.85 has an SEM of about 3.87 points. That number tells you the typical error around a single score.

Reporting SEM alongside ICC gives readers a sense of how much uncertainty surrounds a single measurement. This is especially useful in clinical settings where individual scores carry meaningful consequences.

Common Mistakes in Test-Retest Reliability Studies

After reviewing many reliability reports, our team sees the same errors repeatedly. Avoiding them will strengthen your study and make your results easier to interpret.

Mistake 1: Using Pearson Correlation Instead of ICC

Pearson r overestimates reliability because it ignores systematic shifts. Always use ICC for continuous test-retest data and Cohen’s kappa for categorical items.

This is the single most common error in published reliability studies. It persists because Pearson r is easy to calculate and familiar to most researchers, but it produces misleading results in the context of test-retest reliability.

Mistake 2: Choosing the Wrong Time Interval

A one-day retest may produce inflated reliability due to memory. A one-year retest may produce deflated reliability due to real change. Match the interval to the construct.

If reviewers question your interval choice, be prepared to justify it. Cite literature on the expected stability of your construct and explain why your interval minimizes both memory effects and true change.

Mistake 3: Ignoring ICC Model Selection

Not all ICC values are comparable. Report the exact ICC model and form so readers can interpret your coefficient correctly.

SPSS, R, and other statistical packages offer multiple ICC options, and the default may not be the one you want. Check the documentation and make a deliberate choice rather than accepting the software default.

Mistake 4: Inadequate Sample Size

Small samples produce wide confidence intervals. Aim for at least 50 participants for a stable estimate, and more if you expect the ICC to be below 0.80.

If your sample is small, acknowledge the limitation in your report. A wide confidence interval that spans from 0.50 to 0.90 is much less informative than a narrow one that spans from 0.80 to 0.88.

Mistake 5: Failing to Report Confidence Intervals

A point estimate like 0.82 means little without a confidence interval. Report the 95 percent confidence interval so readers can see the precision of your estimate.

Many journals now require confidence intervals for reliability estimates. Even if your target outlet does not, including them demonstrates methodological rigor and makes your results easier to interpret.

Mistake 6: Confusing Reliability with Validity

Reliability is about consistency. Validity is about whether the survey measures what it should. A reliable instrument can still be invalid, so report both when possible.

This distinction is sometimes lost in practice, especially when researchers are under pressure to show that their instrument is sound. Remember that reliability sets the ceiling for validity: a measure cannot be more valid than it is reliable.

How to Report Test-Retest Reliability

Transparent reporting helps other researchers evaluate and reproduce your work. Include at least the following information whenever you report test-retest reliability.

List the ICC model and form, such as ICC(3,1) with absolute agreement. State the time interval between administrations. Report the sample size and response rate. Provide the point estimate and the 95 percent confidence interval. Mention any countermeasures used to reduce memory effects, such as item randomization.

A complete report might read: “Test-retest reliability was assessed using a two-way mixed effects model with absolute agreement, ICC(3,1). The interval between administrations was 14 days. The ICC was 0.87 (95% CI: 0.81 to 0.92), based on 82 respondents.”

This level of detail lets readers judge the quality of your evidence and compare your results with other studies. It also satisfies the reporting standards recommended by Koo and Li and by the editors of major psychometric journals.

Reporting Checklist

Before submitting your manuscript, confirm that you have reported each of the following: the statistic used (ICC, kappa, or other), the ICC model and form, the type of agreement (absolute or consistency), the time interval, the sample size, the point estimate, the confidence interval, and any steps taken to reduce bias.

This checklist is not exhaustive, but it covers the items most commonly missing from reliability reports. Going through it before submission can save you from reviewer requests and improve the credibility of your instrument.

Frequently Asked Questions

What is an acceptable test-retest reliability?

An acceptable test-retest reliability is usually an ICC of at least 0.70 for group research, 0.80 or higher for clinical use, and above 0.90 for individual decisions. ICC values below 0.50 are considered poor, while 0.75 to 0.90 are good.

What is a good test-retest reliability value?

A good test-retest reliability value is an ICC between 0.75 and 0.90. Values above 0.90 are excellent, and values between 0.50 and 0.75 are moderate. The exact standard depends on whether the score is used for research, clinical screening, or individual diagnosis.

Is a Cronbach’s alpha of 0.6 acceptable?

A Cronbach’s alpha of 0.6 is generally considered too low for most research purposes. Alpha values of 0.70 or higher are commonly accepted, and 0.80 or higher is preferred for clinical or diagnostic instruments. However, alpha measures internal consistency, not test-retest reliability.

What is the acceptable range of reliability?

The acceptable range of test-retest reliability is generally 0.70 to 0.90 for research instruments. ICC values of 0.70 to 0.79 are fair, 0.80 to 0.89 are good, and 0.90 or above are excellent. Below 0.70 is usually considered inadequate for most survey instruments.

What happens in an assessment of test-retest reliability?

In an assessment of test-retest reliability, the same survey is given to the same participants on two occasions. The interval is chosen based on the construct being measured. Scores from both administrations are compared using ICC for continuous data or Cohen’s kappa for categorical data to determine temporal stability.

What is the threshold for ICC reliability?

The threshold for ICC reliability depends on application. Below 0.50 is poor, 0.50 to 0.75 is moderate, 0.75 to 0.90 is good, and above 0.90 is excellent. For most research, 0.70 is the minimum acceptable value, while individual-level decisions require 0.90 or higher.

What time interval should be used between test and retest?

The recommended time interval for test-retest reliability is usually 1 to 4 weeks, with 2 weeks being a common default. Shorter intervals risk memory effects, while longer intervals risk true changes in the construct. Match the interval to the expected stability of what you are measuring.

Why is Pearson correlation not appropriate for test-retest reliability?

Pearson correlation measures linear association, not agreement. It cannot detect systematic shifts in scores between administrations, so it overestimates reliability. ICC is preferred because it accounts for both rank order and absolute agreement between time points.

Conclusion

So what counts as acceptable test-retest reliability for a survey instrument? For most research, an ICC of 0.70 or higher is acceptable, 0.80 or higher is good, and 0.90 or higher is excellent. Use ICC instead of Pearson correlation, choose a retest interval that matches your construct, and always report the model, sample size, and confidence interval.

If you take one thing from this guide, let it be this: reliability is context-dependent. The same coefficient can be acceptable for one study and insufficient for another. Match your threshold to your purpose, report your methods transparently, and your readers will trust your instrument much more.

As research methods continue to evolve in 2026, the fundamentals remain unchanged. Stable measurements are the foundation of credible research, and test-retest reliability is one of the clearest ways to demonstrate that stability.

Leave a Comment