If you are working on a study that involves multiple raters scoring the same subjects, you have probably run into the intraclass correlation coefficient (ICC). Learning how to interpret the intraclass correlation coefficient for rater agreement is one of the most useful skills you can build as a researcher, clinician, or data analyst, because it tells you whether your measurements are dependable enough to publish, defend, or act on.
The ICC is a number between 0 and 1 that quantifies how consistently different raters assign scores to the same subjects. A higher value means raters agree more closely. According to the widely cited Koo and Li (2016) guidelines, ICC values less than 0.50 indicate poor reliability, values between 0.50 and 0.75 indicate moderate reliability, values between 0.75 and 0.90 indicate good reliability, and values greater than 0.90 indicate excellent reliability.
In this guide, I walk you through what those thresholds mean, how to choose the correct ICC model for your study, how to read the confidence interval, how to report your results, and the most common mistakes I see researchers make when interpreting ICC values for inter-rater reliability.
Table of Contents
Quick Answer: What Is a Good ICC Value?
For most reliability research, an ICC above 0.75 is considered acceptable and an ICC above 0.90 is considered excellent. The thresholds most researchers and peer reviewers accept today come from Koo and Li (2016), published in the Journal of Chiropractic Medicine. That paper has been cited over 30,000 times because it gives a clean, defensible framework for interpreting ICC scores.
The older Cicchetti (1994) benchmarks use slightly different cutoffs but reach similar conclusions. Both frameworks agree that anything below 0.40 is poor and anything above 0.75 is good to excellent. If a reviewer asks you to justify your ICC interpretation, citing both frameworks gives you solid ground.
One important caveat: these thresholds assume you are using continuous or ordinal data with the correct ICC model. A high ICC from the wrong model does not actually prove reliability. I cover model selection in detail below.
What Is the Intraclass Correlation Coefficient?
The intraclass correlation coefficient (ICC) is a descriptive statistic that measures the reliability of measurements made by multiple raters on the same set of subjects or items. Unlike the Pearson correlation, which tells you whether two variables move together, the ICC tells you whether different raters assign the same absolute scores to the same subjects. That distinction matters because two raters can produce a perfect Pearson correlation while disagreeing by a constant offset, which would be unacceptable for most clinical or research purposes.
Conceptually, the ICC compares the variance between subjects to the total variance, which includes both between-subject variance and between-rater variance. If most of the variability in your dataset comes from real differences between subjects rather than from inconsistencies between raters, your ICC will be high. If raters are all over the place relative to one another, the ICC will drop.
ICC values range from 0 to 1 in theory, although with small samples they can technically come out negative. A value of 1 means perfect agreement among raters. A value of 0 means there is essentially no agreement beyond what you would expect from chance. Negative ICC values usually indicate a sample size problem rather than a meaningful result, and I discuss how to handle those cases below.
Researchers use the ICC in clinical assessment, psychology, education, sports science, and increasingly in machine learning annotation work. Anywhere two or more observers assign quantitative scores to the same items, the ICC is the standard tool for assessing whether those scores are trustworthy.
Types of Reliability the ICC Measures
The intraclass correlation coefficient is flexible enough to assess three distinct types of reliability. Knowing which type you are measuring changes how you design your study and which ICC form you select.
Inter-Rater Reliability
Inter-rater reliability measures whether different raters, scoring the same subjects at roughly the same time, produce similar results. This is what most people mean when they talk about rater agreement. If three radiologists each read the same 50 X-rays and assign a severity score, the ICC tells you whether those radiologists are interchangeable for clinical purposes. A high inter-rater ICC means any of the three radiologists would produce comparable scores.
Intra-Rater Reliability
Intra-rater reliability measures whether the same rater, scoring the same subjects on two or more occasions, produces consistent results. Imagine a single pathologist scoring tumor slides today and then again next week without knowing they are the same slides. A high intra-rater ICC means that pathologist is internally consistent. This type of reliability matters when measurements are subjective or when a single observer produces all data in a longitudinal study.
Test-Retest Reliability
Test-retest reliability measures whether the same instrument or measurement procedure, applied to the same subjects at two different time points, produces stable results. This is closely related to intra-rater reliability but focuses on the measurement tool rather than the human rater. If you administer the same depression questionnaire to patients two weeks apart and assume their symptoms have not changed, a high test-retest ICC confirms the questionnaire is stable.
All three types use the same family of ICC statistics, but the study design dictates which ICC form is appropriate. Choosing the wrong form is one of the most common mistakes I see in published reliability studies.
How to Interpret the Intraclass Correlation Coefficient for Rater Agreement
The core task is matching your ICC value to a qualitative reliability category. Two interpretation frameworks dominate the literature: Koo and Li (2016) and Cicchetti (1994). I recommend reporting both, because reviewers in different fields prefer different citations.
Koo and Li (2016) Thresholds
The Koo and Li framework is the current standard in clinical and biomedical research. Their thresholds are:
Less than 0.50: Poor reliability. Raters disagree substantially, and the measurement should not be trusted for research or clinical decisions.
0.50 to 0.75: Moderate reliability. Acceptable for exploratory work but generally not sufficient for high-stakes clinical measurements.
0.75 to 0.90: Good reliability. The minimum most peer reviewers expect for clinical or applied research.
Greater than 0.90: Excellent reliability. Suitable for measurements used in individual-level decisions, such as diagnostic tests.
Notice that the upper bound of the moderate category is 0.75, not 0.74. Koo and Li use inclusive ranges, so a value of exactly 0.75 falls into the good category.
Cicchetti (1994) Thresholds
The Cicchetti framework comes from psychology and is still widely cited in that field. The cutoffs are slightly more granular:
Less than 0.40: Poor reliability
0.40 to 0.59: Fair reliability
0.60 to 0.74: Good reliability
0.75 and above: Excellent reliability
You can see the two systems do not line up perfectly. An ICC of 0.70 is good under Cicchetti but only moderate under Koo and Li. When you report, state which framework you are using and stick with it consistently.
Side-by-Side Comparison
Here is how the two frameworks compare on the same ICC values:
ICC 0.30: Poor under both Koo and Li and Cicchetti.
ICC 0.50: Moderate under Koo and Li, fair under Cicchetti.
ICC 0.70: Moderate under Koo and Li, good under Cicchetti.
ICC 0.80: Good under Koo and Li, excellent under Cicchetti.
ICC 0.95: Excellent under both frameworks.
If you want a single defensible threshold for clinical research, use 0.75 as your floor. That cutoff holds up under both frameworks as the minimum for acceptable reliability.
What About an ICC of 0?
An ICC of 0 means raters show no agreement beyond what you would expect from random chance. Their scores are essentially unrelated to one another. This is different from a Pearson correlation of 0, which simply means no linear relationship. For ICC, a 0 means the measurement has no reliability at all.
In practice, an ICC at or near 0 usually signals a design problem. Common causes include too few subjects, too few raters, subjects that are too similar to one another (low between-subject variance), or raters applying fundamentally different scoring criteria.
What About Low ICC Values Like 0.28?
An ICC of 0.28 falls below 0.50 in both frameworks, so it represents poor reliability. Raters are disagreeing substantially. If you get a value this low, do not try to argue it is acceptable. Instead, investigate the cause. Retrain your raters, refine your measurement instrument, increase the number of subjects to widen between-subject variance, or consider whether your measurement is too subjective for quantitative scoring at all.
ICC Models Explained
Interpreting your ICC value correctly requires using the right ICC model in the first place. There are ten distinct ICC forms, and choosing the wrong one will produce a misleading number. The three model families are one-way random effects, two-way random effects, and two-way mixed effects.
One-Way Random Effects Model
The one-way random model treats both subjects and raters as randomly sampled from larger populations. You use this model when the specific raters in your study do not matter and you want to generalize to any comparable rater. This model assumes each subject is rated by a different set of raters, which is unusual in practice but does come up in large multicenter studies.
The one-way model cannot separate systematic rater bias from random error. If one rater consistently scores five points higher than everyone else, the one-way model folds that bias into the error term and lowers your ICC. For that reason, the one-way model is the least commonly used of the three.
Two-Way Random Effects Model
The two-way random model treats both subjects and raters as random samples from larger populations. This is the model you want when your raters are representative of a broader class of raters, like a random sample of nurses from a hospital. The two-way random model can partition out systematic rater differences, which makes it more informative than the one-way model for most studies.
This is the model Koo and Li recommend as a default for inter-rater reliability studies where you want to generalize beyond the specific raters in your sample. In Shrout and Fleiss notation it corresponds to ICC(2,1) for single measures and ICC(2,k) for average measures.
Two-Way Mixed Effects Model
The two-way mixed model treats subjects as random but raters as fixed. You use this model when the specific raters in your study are the only raters you care about. For example, if you have three trained coders who will annotate all data for your project and you have no intention of generalizing to other coders, the two-way mixed model is appropriate.
Many published studies use the two-way random model by default even when the two-way mixed model would be more accurate. The difference matters because the two-way mixed model does not include rater-by-subject interaction variance in the same way, which can produce slightly higher ICC values. In Shrout and Fleiss notation this is ICC(3,1) for single measures and ICC(3,k) for average measures.
McGraw and Wong ICC Forms
McGraw and Wong (1996) organized the ten ICC forms into a cleaner naming system that combines model, type, and unit. Their notation is what most modern software, including SPSS and Pingouin, uses. The six most commonly reported forms are:
ICC(1,1): One-way random, single measures. Each subject rated by a different set of raters drawn randomly.
ICC(1,k): One-way random, average measures. Same as above but using the average of k raters.
ICC(2,1): Two-way random, single measures. Same raters rate each subject, raters are random.
ICC(2,k): Two-way random, average measures.
ICC(3,1): Two-way mixed, single measures. Raters are fixed.
ICC(3,k): Two-way mixed, average measures.
Each form also comes in two variants: absolute agreement and consistency. That is how you get to ten total combinations when you count the consistency versions of the two-way models.
Single Measures vs Average Measures
The second decision in ICC form selection is whether you want the reliability of a single rater or the reliability of the average across multiple raters. This is the distinction between single measures ICC and average measures ICC.
Single measures ICC tells you how dependable a single rating is. Use this when your real-world measurement will be made by one rater. For example, if a single clinician will read each patient scan in practice, you want to know the reliability of that single reading, not the reliability of an average of three readings.
Average measures ICC tells you how dependable the mean of k raters is. This value is always higher than the single measures ICC for the same data, because averaging cancels out random rater error. Average measures ICC is appropriate when your study design uses the mean of multiple raters as the final measurement, which is common in instrument validation studies.
A common mistake is reporting the average measures ICC because it looks more impressive, when the single measures ICC is what your study design actually calls for. Reviewers who know ICC will catch this.
Consistency vs Absolute Agreement
The third decision is whether you care about raters producing the same scores (absolute agreement) or raters producing scores that maintain the same relative ordering of subjects (consistency). This choice only applies to the two-way models.
Absolute agreement is the stricter standard. It penalizes any systematic difference between raters, even if all raters rank subjects in the same order. If Rater A scores everyone 10 points higher than Rater B but otherwise agrees on the ranking, absolute agreement ICC will be lower than consistency ICC. Use absolute agreement when the actual score values matter, which is the case for most clinical measurements, diagnostic thresholds, and any setting where decisions depend on a specific numeric cutoff.
Consistency is the more lenient standard. It allows raters to differ by a constant additive factor as long as they preserve the relative ordering of subjects. Use consistency when you only care about rank order or when raters will be calibrated before data collection. Consistency ICC will always be higher than or equal to absolute agreement ICC for the same data.
If you are unsure which to choose, default to absolute agreement. It is the more conservative and defensible choice for most reliability studies, and it is what Koo and Li recommend for inter-rater reliability by default.
How to Choose the Right ICC Form: A 4-Question Framework
Koo and Li (2016) provide a four-question framework for selecting the correct ICC form. I use this framework in every reliability analysis because it removes the guesswork.
Question 1: Do all subjects rate the same set of raters, or do different subjects have different raters? If all subjects are rated by the same raters, you need a two-way model. If different subjects have different raters, you need a one-way model.
Question 2: Are your raters a random sample from a larger population, or are they the only raters you care about? If raters are representative of a broader population, use a two-way random model. If the specific raters in your study are the only ones relevant, use a two-way mixed model.
Question 3: Are you interested in the reliability of a single rating or the reliability of the average of multiple ratings? Choose single measures if one rater will make the measurement in practice. Choose average measures if your final score is the mean of multiple raters.
Question 4: Do you care about absolute score agreement or just consistent relative ranking? Choose absolute agreement if the numeric values matter for decisions. Choose consistency if only the rank order matters.
Work through these four questions in order and you will land on the correct ICC form every time. Document your answers in your methods section so reviewers can verify your choice.
Confidence Intervals for ICC
A point estimate of ICC is not enough. Every ICC value should come with a 95 percent confidence interval that tells you the plausible range of the true reliability. A narrow confidence interval means your ICC estimate is precise. A wide confidence interval means you do not have enough data to be confident in the reliability of your measurement.
As a rough guide, if your 95 percent confidence interval spans more than 0.20 in width, your estimate is imprecise and you probably need more subjects or more raters. A common practical recommendation is to aim for the lower bound of the confidence interval to exceed 0.75 for clinical measurements.
Negative lower bounds on the confidence interval are common with small samples. A confidence interval that includes 0 means you cannot rule out the possibility that your measurement has zero reliability. If you see this, do not dismiss it. It usually means your sample size is too small or your raters genuinely disagree.
Always report both the point estimate and the confidence interval. Reporting only the point estimate hides uncertainty and is one of the most frequent reviewer complaints I see in submitted manuscripts.
How to Report ICC in Research
Koo and Li (2016) provide a reporting template that most journals now expect. A complete ICC report should include the following elements every time:
Software and version: Name the statistical package (SPSS, R, Stata, Python with Pingouin) and the version you used.
Model: State the ICC model explicitly, for example two-way random effects.
Type: Specify single measures or average measures.
Definition: State whether you used absolute agreement or consistency.
Number of raters: Report how many raters participated.
Number of subjects: Report how many subjects or items were rated.
Point estimate: Report the ICC value to two decimal places.
Confidence interval: Report the 95 percent confidence interval with lower and upper bounds.
Interpretation: State which thresholds you used (Koo and Li or Cicchetti) and the resulting qualitative category.
Here is an example of a properly formatted ICC report. Interrater reliability was assessed using a two-way random effects model, single measures, absolute agreement definition. Six trained raters independently scored 40 patient videos. The ICC was 0.82 (95 percent CI 0.71 to 0.90), indicating good reliability according to the Koo and Li (2016) criteria. Analyses were conducted in R version 4.3 using the irr package.
That single paragraph contains everything a reviewer needs to evaluate your reliability analysis. If your methods section does not look like this, you are probably underreporting.
Common Mistakes in ICC Interpretation
No competing article dedicates space to this, which is surprising given how often these errors appear in published research. Here are the mistakes I encounter most frequently when reviewing manuscripts.
Using the Wrong ICC Model
The single most common error is selecting the default ICC in SPSS without understanding which model it computes. SPSS defaults to a two-way mixed model with consistency, which is not appropriate for most inter-rater reliability studies. Many researchers report this default output without realizing it may inflate their reliability estimate.
Confusing Consistency With Absolute Agreement
Researchers frequently report consistency ICC when their study design calls for absolute agreement. If your clinical cutoff depends on the raw score value, consistency ICC will overstate your reliability. Always check whether your decision depends on absolute values before choosing the consistency definition.
Reporting Average Measures ICC for a Single-Rater Design
If a single rater will make measurements in practice, but you report average measures ICC because it is higher, you are overstating reliability. Reviewers are increasingly catching this. Match the measures unit to your actual application.
Ignoring the Confidence Interval
Reporting a point estimate without a confidence interval hides whether your reliability estimate is precise enough to trust. An ICC of 0.80 with a 95 percent CI of 0.30 to 0.95 is far less convincing than an ICC of 0.78 with a CI of 0.72 to 0.84.
Treating ICC as a Significance Test
The ICC is not a hypothesis test. A high p-value on the ICC does not mean your measurement is reliable in a practical sense. Focus on the magnitude of the ICC and the width of the confidence interval, not on whether the estimate is statistically significant.
Using ICC for Categorical Data
ICC is designed for continuous or ordinal data. If your raters are assigning nominal categories, use Cohen kappa for two raters or Fleiss kappa for three or more. Using ICC on categorical data produces values that do not correspond to any meaningful reliability benchmark.
Misreading a Negative ICC
A negative ICC does not mean negative correlation. It usually means your sample is too small or there is more variance within subjects than between them. Treat negative ICC values as a signal that your study design needs more subjects, more raters, or both.
When Not to Use ICC: Alternatives
The ICC is not always the right tool. Different data types and study designs call for different reliability statistics. Here are the alternatives I reach for when ICC is not appropriate.
Cohen kappa measures agreement between two raters on categorical data. Use it when your raters assign nominal labels, like diagnosing a condition as present or absent. Cohen kappa corrects for chance agreement, which simple percent agreement does not.
Fleiss kappa extends Cohen kappa to three or more raters on categorical data. If you have multiple raters assigning nominal categories, Fleiss kappa is the standard choice.
Krippendorff alpha is the most flexible agreement statistic. It handles any number of raters, missing data, and nominal, ordinal, interval, or ratio data. If your study has missing ratings or mixed data types, Krippendorff alpha is more appropriate than ICC.
Concordance correlation coefficient (CCC) measures agreement between two raters or methods on continuous data, with explicit attention to both precision and accuracy. Use CCC when you want to evaluate whether a new measurement method agrees with a gold standard, which is common in method comparison studies.
Bland-Altman analysis visually assesses agreement between two methods by plotting the difference between measurements against their mean. It is the standard approach for method comparison in clinical chemistry and is more informative than a single ICC value when you want to understand the pattern of disagreement.
Cronbach alpha measures internal consistency of items within a scale, not agreement between raters. Researchers sometimes confuse the two because both are reliability statistics, but they answer different questions.
Software for Computing ICC
You can compute ICC in most major statistical packages. Here is a quick overview of the most common options and what to watch for in each.
In SPSS, the ICC is under Analyze, then Scale, then Reliability Analysis. Select the Intraclass Correlation Coefficient option and choose your model, type, and definition. The default is a two-way mixed model with consistency, which as I noted earlier is not what most inter-rater reliability studies need.
In R, the irr package provides the icc function and the psych package provides the ICC function. Both produce all six McGraw-Wong forms in a single output table, which makes it easy to compare. The Pingouin library in Python offers similar functionality.
In Stata, the icc command computes ICC with options for one-way, two-way random, and two-way mixed models. In MedCalc, the ICC is available under the reliability analysis menu.
Whichever tool you use, verify that the model, type, and definition match what you intend. Software defaults are not reliable substitutes for understanding your study design.
FAQs
How do you interpret the intraclass correlation coefficient?
ICC values range from 0 to 1. Using the Koo and Li (2016) thresholds, values below 0.50 indicate poor reliability, 0.50 to 0.75 indicate moderate reliability, 0.75 to 0.90 indicate good reliability, and values above 0.90 indicate excellent reliability. Always report the ICC with its 95 percent confidence interval and state which model and definition you used.
What is considered a good ICC?
A good ICC is generally 0.75 or higher according to the Koo and Li (2016) framework. Values above 0.90 are considered excellent. For clinical measurements used in individual-level decisions, aim for an ICC above 0.90. For most applied research, 0.75 is the minimum acceptable threshold.
Is 0.28 a strong correlation?
No. An ICC of 0.28 falls below 0.50 in both the Koo and Li and Cicchetti frameworks, which means it indicates poor reliability. Raters at this level disagree substantially and the measurement should not be trusted for research or clinical decisions without significant redesign of the study or retraining of raters.
What does an ICC of 0 mean?
An ICC of 0 means there is no agreement among raters beyond what you would expect from chance. Raters are essentially producing unrelated scores. In practice, an ICC at or near 0 usually signals a study design problem such as too few subjects, too few raters, low between-subject variance, or raters applying fundamentally different scoring criteria.
Conclusion
Knowing how to interpret the intraclass correlation coefficient for rater agreement comes down to four steps. Choose the correct ICC model based on your study design using the 4-question framework. Read the right ICC value for your application, whether that is single measures absolute agreement or average measures consistency. Compare your value to the Koo and Li (2016) thresholds to assign a qualitative reliability category. Then report the point estimate, the 95 percent confidence interval, and all model specifications so reviewers and readers can verify your interpretation.
If you remember nothing else, remember that a high ICC from the wrong model proves nothing, and an ICC without a confidence interval hides the uncertainty that matters most. Get those two things right and your reliability analysis will stand up to scrutiny in any peer-reviewed setting.