How to Interpret an Item Discrimination Index? (2026 Guide)

If you have ever stared at an item analysis report and wondered whether a discrimination index of 0.18 means your test question is fine, broken, or somewhere in between, you are in the right place. Learning how to interpret an item discrimination index for a multiple-choice test is one of the most useful skills an educator, instructional designer, or test developer can build. The number itself is simple, but the story it tells about each question can completely change how you assemble and revise exams.

When I first started running item analyses on my own quizzes, the report looked like alphabet soup: D, p-value, rpb, KR-20. None of it meant anything until I connected each statistic to a specific decision about whether to keep, revise, or scrap a question. That decision-making lens is what separates people who glance at item analysis reports from people who actually use them to improve tests.

This guide walks through everything you need in plain language. You will get the definition, the formula, a complete interpretation table with thresholds you can act on, a worked numerical example, a distractor analysis checklist, and the cautions that competing resources skip. By the end you should be able to look at any item analysis report and confidently answer: is this question doing its job?

If you also handle standard setting work, the discrimination index pairs naturally with standard setting and passing score determination, since both depend on knowing which items are pulling their weight.

Quick Answer: What the Discrimination Index Tells You

The item discrimination index (often abbreviated D) is a statistic that shows how well a multiple-choice question separates students who know the material from students who do not. It ranges from -1.0 to +1.0. A positive value means high-performing students answer correctly more often than low-performing students. A value near zero means the question does not differentiate at all. A negative value means weaker students are actually outperforming stronger students on that item, which is usually a red flag.

As a quick rule of thumb, a discrimination index of 0.30 or higher is considered acceptable to good for most norm-referenced tests, 0.40 or higher is very good, and anything below 0.20 should be reviewed. Negative values almost always warrant revision or removal.

What Is the Item Discrimination Index?

The item discrimination index is a statistical measure that indicates how well a multiple-choice test question differentiates between students who understand the material and those who do not. Specifically, it checks whether students who scored well on the overall test also answered that particular item correctly, while students who scored poorly did not.

Discrimination is one of the core statistics produced by an item analysis report alongside item difficulty (the p-value) and distractor analysis. Together these three give you a complete picture of item quality. Anyone who builds or evaluates multiple-choice assessments, including classroom teachers, nursing faculty, licensure exam developers, psychometricians, and curriculum designers, needs to understand the discrimination index to ensure their tests are fair, reliable, and valid measures of student learning.

Discrimination matters because it directly affects test reliability and validity. Items with high discrimination improve the internal consistency of a test and increase the standard error of measurement we can tolerate. Items with low or negative discrimination drag reliability down, confuse students, and may indicate a question that is poorly written, ambiguous, mis-keyed, or measuring a different construct than the rest of the test.

It is also worth noting that the discrimination index is not a fixed property of a question. The same item can show strong discrimination in one group of students and weak discrimination in another, depending on how homogeneous the group is, how well the instruction aligned with the item, and how the test was administered. That is why the same statistics course item can look excellent in one semester and weak in the next.

The Discrimination Index Formula

The most widely taught formula for the discrimination index is Kelley’s upper-lower 27% method. The procedure compares the top 27% of test-takers (the upper group) with the bottom 27% (the lower group), then calculates the difference in the proportion who answered the item correctly.

D = (U / Nu) − (L / Nl)

Where U is the number of correct responses in the upper group, L is the number of correct responses in the lower group, and Nu and Nl are the total numbers of students in each group. Because the upper and lower groups are usually the same size (the top and bottom 27%), the result simplifies to the difference in proportions correct, ranging from -1.0 to +1.0.

Why 27%? Truman Kelley recommended that cutoff in 1939 because it maximizes the difference between the two extreme groups while keeping enough students in each group for a stable estimate. The 27% rule works best with larger samples; with very small classes the proportion is less stable and you may want to use the top and bottom thirds or even the upper and lower halves.

An alternative and increasingly common approach is the point-biserial correlation, which correlates each student’s score on the item (0 or 1) with their total score on the rest of the test. The point-biserial uses data from every student rather than just the extremes, so it tends to be more stable with moderate sample sizes. Many modern testing platforms, including ScorePak and most learning management systems, now report point-biserial as the default discrimination measure.

Both statistics answer the same question, but they are not numerically identical. The Kelley 27% D value is usually larger in absolute terms than the point-biserial, so check which formula your software uses before applying thresholds. For a deeper comparison of difficulty estimation approaches that pair with these formulas, see the work on item difficulty estimation methods.

How to Interpret Discrimination Index Values

The single most important table on this page is the interpretation table below. Use it as your starting point whenever you open an item analysis report. The thresholds synthesize the ranges used by Ebel, the University of Washington, Conestoga, and Professional Testing, all of which are broadly consistent but differ slightly at the boundaries.

  • D ≥ 0.40: Very good to excellent. The item discriminates strongly and is a strong candidate for the item bank. Keep it.

  • D = 0.30 to 0.39: Good to very good. Reasonably discriminating, generally acceptable, minor revision only if other data suggest an issue.

  • D = 0.20 to 0.29: Fair to marginal. The item discriminates weakly. Review the wording and distractors before reuse.

  • D = 0.00 to 0.19: Poor. The item does not differentiate well. Revise or replace.

  • D < 0.00: Negative. More low scorers answered correctly than high scorers. Almost always revise, re-key, or remove.

What is a good discrimination index for a test question? For most norm-referenced multiple-choice tests, anything at or above 0.30 is considered acceptable, and 0.40 or higher is considered good to excellent. For criterion-referenced or mastery tests, where the goal is to confirm students have met a standard rather than spread them out, lower discrimination values in the 0.20 to 0.29 range are often tolerated because the test is not designed to maximize score variance.

The key is to pair the DI value with the recommended action rather than treating the number in isolation. A question with D = 0.25, a clean distractor pattern, and a difficulty level near 0.75 might be perfectly fine for a mastery test. The same D value in a norm-referenced选拔 exam would push you toward revision.

Recommended Actions by Discrimination Index Range

  • D ≥ 0.40: Add to the item bank. Use as an anchor item.

  • D = 0.30 to 0.39: Keep. Optionally polish wording for clarity.

  • D = 0.20 to 0.29: Review distractors and stem. Revise before next use.

  • D = 0.00 to 0.19: Rewrite the stem and at least one distractor. Pilot the revision.

  • D < 0.00: Check for mis-keying first. If the key is correct, retire or completely rewrite the item.

Negative Discrimination Index: Causes and Fixes

A negative discrimination index is alarming but very diagnosable. It means the students who scored highest on the overall test were less likely to answer that specific item correctly than the students who scored lowest. The most common cause by far is a mis-keyed answer, where the wrong option was marked correct in the scoring key.

Other frequent causes include ambiguous wording that stronger students overthink, a trick question that rewards superficial reading, an item that tests a different construct than the rest of the test, or a double-keyed item where two options are defensibly correct. Whenever I see a negative D, my first step is to recheck the answer key, then read the stem and options aloud to a colleague who has not seen the item before. That two-minute check resolves the majority of negative discrimination cases.

The Relationship Between Discrimination Index and Item Difficulty

Discrimination does not exist in a vacuum. It interacts tightly with item difficulty, and you cannot interpret one without the other. The item difficulty index, usually written as the p-value, is simply the proportion of students who answered the item correctly. A p-value of 0.85 means 85% of students got it right.

The mathematical ceiling on discrimination depends on difficulty. An extremely easy item, where 100% of students answer correctly, cannot discriminate at all because there is no variance to exploit. The same is true for an extremely hard item where 0% answer correctly. Discrimination is maximized when the p-value sits near 0.50 for a two-option item, and shifts slightly higher for items with more options.

Ideal Difficulty Ranges by Question Type

  • 5-option MCQ: Ideal p-value around 0.60 to 0.70 (chance score is 0.20)

  • 4-option MCQ: Ideal p-value around 0.60 to 0.75 (chance score is 0.25)

  • 3-option MCQ: Ideal p-value around 0.65 to 0.75 (chance score is 0.33)

  • True/False: Ideal p-value around 0.75 to 0.85 (chance score is 0.50)

  • Mastery/CRT: Difficulty less relevant; focus on whether the p-value aligns with the cut score

The practical takeaway is that a question can only discriminate well when it sits in the middle of the difficulty range. If your item analysis shows a high p-value paired with a low discrimination index, the most likely explanation is that the question is too easy to differentiate students. If both p-value and discrimination are low, the question is both hard and possibly confusing, which calls for revision rather than celebration.

How do you interpret the item difficulty index? Read it as the proportion of students who answered correctly. Values above 0.85 suggest the item is very easy; values between 0.30 and 0.70 are typically ideal for norm-referenced tests; values below 0.30 indicate a difficult item that may need review unless difficulty was intentional. Always pair difficulty with discrimination to decide on the next step.

Distractor Analysis and Item Quality

Distractor analysis is the third leg of item analysis. Even when the discrimination index is acceptable, looking at how students distributed across the wrong options tells you whether the question is doing its job for the right reasons. A well-functioning multiple-choice item has distractors that pull more strongly on the lower-performing students than on the higher-performing students.

A functional distractor is a wrong option chosen by at least 5% of test-takers, and ideally more often by the lower group than the upper group. A non-functional distractor (NFD) is one that almost nobody picks, which means it is not doing any work. Items with one or more non-functional distractors effectively become easier than their option count suggests, and they tend to have weaker discrimination.

Distractor Effectiveness Checklist

  • Each distractor is selected by at least 5% of the total sample.

  • Each distractor is selected by a higher proportion of the lower group than the upper group.

  • The keyed correct answer is chosen by more upper-group than lower-group students.

  • No single distractor draws the majority of the lower group (suggesting a near-miss or double key).

  • Students are not clustering on one distractor across both groups equally (suggests a plausible but unintended answer).

When a distractor fails any of these checks, the fix is usually to rewrite that option. Common improvements include making the distractor more plausible by referencing a common misconception, tightening the wording, or replacing a non-functional option with one based on errors you have actually seen students make in class. The goal is to design distractors that specifically attract students who do not yet understand the concept, not just plausible-sounding wrong answers.

Items with efficient distractors tend to have lower p-values and stronger discrimination. Research on item analysis confirms that the relationship between distractor efficiency, difficulty, and discrimination is consistent and worth monitoring across every administration of an exam.

Point-Biserial Correlation Explained

The point-biserial correlation (rpb) is the correlation between a dichotomous item score (correct or incorrect) and a continuous total test score. It is mathematically equivalent to the Pearson product-moment correlation when one variable is binary, and it answers the same question as the Kelley discrimination index but uses every student’s data rather than just the top and bottom 27%.

Interpretation thresholds for the point-biserial are similar to but slightly lower than the Kelley D. A common convention is that rpb values at or above 0.30 are good, 0.20 to 0.29 are acceptable, and below 0.20 needs review. Because the point-biserial uses all students, it is more stable with smaller samples and is the default in most current item analysis software.

If your report shows rpb instead of D, do not panic. The interpretation logic is identical, just apply the slightly lower thresholds. Some reports include both, and when they disagree noticeably the discrepancy usually points to an item that behaves differently in the middle of the score distribution than at the extremes.

Step-by-Step Worked Example

To make the calculation concrete, here is a small worked example. Imagine a 40-item multiple-choice quiz given to 100 students. For one item we want to calculate the discrimination index.

Step 1: Rank all 100 students by total score. Step 2: Identify the upper group as the top 27 students and the lower group as the bottom 27 students. Step 3: Count correct responses. Suppose 22 of the 27 upper-group students answered this item correctly (U = 22), and 10 of the 27 lower-group students answered correctly (L = 10). Step 4: Apply the formula: D = (22/27) − (10/27) = 0.815 − 0.370 = 0.445.

That D of 0.45 falls in the very good to excellent range. The item is a strong discriminator and belongs in the item bank. If we had also calculated the p-value we would see the overall proportion correct, and if we ran distractor analysis we would confirm that each wrong option was pulling more from the lower group than the upper group. Three statistics, one coherent story.

Repeat this process for every item in the test. For a 60-item exam you would expect to spend about an hour on a full item analysis using a spreadsheet, and far less if your LMS or a tool like DataLink or ScorePak handles the math automatically. The decision time, not the calculation, is what eats your hours.

Sample Size and Cautions When Interpreting DI Values

The single biggest gap in most item analysis guides is sample size. None of the leading public resources on discrimination index interpretation discuss minimum sample sizes, and educators on Reddit routinely ask how many students they need before they can trust a DI value. The honest answer is that small classes produce unstable statistics.

As a guideline, the Kelley 27% method becomes reasonably stable around 30 to 50 students, and is most reliable above 100. With fewer than 30 students, a single student moving from the upper to the lower group can swing the discrimination index dramatically, so any value should be treated as a hint rather than a verdict. For very small classes, prefer the point-biserial correlation, which uses all the data and is more stable.

Beyond sample size, three cautions deserve attention. First, discrimination index thresholds vary by test type. Criterion-referenced mastery tests should not be held to the same expectations as norm-referenced选拔 exams. Second, discrimination can be inflated by a test that is internally consistent but narrow in construct; a high DI does not guarantee the test is measuring what you think it measures. Third, item statistics from one administration are not immutable. Re-check them every semester and look for drift.

If your test is used for high-stakes decisions, also consider running a differential item functioning analysis alongside your standard item analysis. DIF analysis catches items that discriminate differently across demographic subgroups and is the natural next step beyond classical DI interpretation.

FAQs

How do you interpret the discrimination index?

Interpret the discrimination index using threshold ranges: values at or above 0.40 are very good to excellent, 0.30 to 0.39 are good, 0.20 to 0.29 are marginal, 0.00 to 0.19 are poor, and any negative value indicates the item is working against the test. Always read the DI alongside item difficulty and distractor analysis before deciding to keep, revise, or remove the item.

What is a good discrimination index for a test question?

A discrimination index of 0.30 or higher is generally considered good for norm-referenced multiple-choice tests, and 0.40 or higher is considered very good to excellent. For criterion-referenced mastery tests, values in the 0.20 to 0.29 range are often acceptable because the goal is to confirm mastery rather than spread students out.

How do you interpret the item difficulty index?

The item difficulty index, or p-value, is the proportion of students who answered the item correctly. Values above 0.85 indicate a very easy item, 0.30 to 0.70 is typically ideal for norm-referenced tests, and below 0.30 indicates a difficult item. Pair the p-value with the discrimination index before deciding to keep or revise the item.

What does the discrimination index tell you?

The discrimination index tells you whether a test question differentiates between students who know the material and students who do not. A positive value means high-performing students answered correctly more often than low-performing students, a value near zero means the item does not differentiate, and a negative value means the item is working against the rest of the test.

Conclusion

Knowing how to interpret an item discrimination index for a multiple-choice test turns a confusing report into a clear action plan. The five threshold ranges, the relationship with item difficulty, the distractor analysis checklist, and the cautions about sample size give you everything you need to decide what to keep, revise, or retire. Pair DI with the p-value and distractor analysis every time, treat negative values as diagnosable, and adjust your thresholds for the type of test you are building.

Your next step is to pull up your most recent item analysis report and run every item through the interpretation table above. Flag anything below 0.20 for revision, recheck any negative values for mis-keying, and shore up non-functional distractors. Do that once per semester and your test quality will compound quickly. If your items are also feeding into cut-score work, the same attention to discrimination will pay off in more defensible passing standards.

Leave a Comment