Differential Item Functioning (DIF) is one of the most important concepts in educational measurement and test fairness. When I first encountered the ETS DIF classification categories during a psychometric review, the A, B, and C labels felt like a foreign language. If you are in the same boat, this guide breaks down each category in plain terms.
The ETS DIF classification categories give test developers a standardized way to flag items that perform differently across demographic groups. Understanding the ETS DIF classification categories matters because these labels directly influence whether a test question stays, gets revised, or is removed entirely. In this article, I walk through what each category means, the statistical thresholds behind them, and how to interpret them in practice for 2026.
Whether you are a psychometrician, a graduate student, or a test developer reviewing fairness data, you will leave with a clear framework for making sense of DIF output. Let us start with the foundation.
Table of Contents
What Is Differential Item Functioning (DIF)?
Differential Item Functioning (DIF) occurs when test items function differently for different groups of respondents, even after controlling for overall ability differences. In simpler terms, a DIF item is one where two people with the same underlying ability have different probabilities of answering correctly simply because they belong to different groups.
These groups are typically defined by demographic characteristics such as gender, ethnicity, language background, or age. The key point is that DIF is not about overall score differences between groups. It is about item-level differences that exist after you equalize for ability.
For example, imagine a math word problem that references a sport more popular in one cultural group than another. Two students with identical math ability might answer differently because the question itself introduces an unfair advantage or disadvantage. That is DIF in action.
DIF is not the same thing as item bias, although the two concepts are related. An item showing DIF is a statistical flag that warrants further review. Bias is a judgment about whether that statistical difference reflects unfair content. The ETS DIF classification categories help organize those statistical flags by severity.
Why ETS DIF Classification Categories Matter
The Educational Testing Service (ETS) developed its classification system to bring consistency to DIF flagging across large-scale assessment programs. Without a standardized framework, every test developer would set their own cutoffs, making fairness reviews impossible to compare.
The ETS DIF classification categories solve this by using a combination of statistical significance and effect size. This dual approach prevents the most common errors I see in practice: flagging trivial differences as problematic, or missing meaningful differences because they fall just short of significance.
For test developers, these categories drive concrete decisions. Category A items typically stay in the pool without modification. Category B items trigger content review and possible revision. Category C items usually require removal or substantial rewrite. That decision tree is why getting the classification right matters so much.
Beyond operational decisions, the ETS DIF classification categories provide documentation for legal and accreditation standards. Testing programs regularly face audits asking whether they conducted fairness analyses. Showing a clear A, B, C classification record demonstrates due diligence.
Reference Group vs Focal Group: The Foundation of DIF
Every DIF analysis compares two groups: the reference group and the focal group. Understanding this pairing is essential before the A, B, C categories make sense.
The reference group is typically the larger or majority group on the assessment. In many U.S. testing contexts, this has historically been White, native English-speaking examinees. The focal group is the group of interest for fairness review, such as Black, Hispanic, female, or older test takers.
The choice of which group serves as reference versus focal does not change the math, but it does change the sign of the DIF statistic. A positive value means the item favors the focal group. A negative value means it disfavors the focal group relative to the reference group.
The plus and minus signs attached to ETS DIF classifications (such as B+ or C-) indicate direction. I have seen many analysts overlook this detail, but it matters when reviewing whether content review should focus on wording, cultural references, or other construct-irrelevant factors.
The Mantel-Haenszel Method: How DIF Is Measured
The ETS DIF classification categories rely on the Mantel-Haenszel (MH) procedure as their primary statistical engine. Developed by Holland and Thayer in 1988, this method compares the odds of answering correctly between reference and focal group members at each level of total test score.
The MH method produces two key outputs. First, the MH chi-square statistic tests whether the difference between groups is statistically significant. Second, the MH common odds ratio is converted into the Delta metric, producing a value called MH D-DIF that serves as the effect size measure.
The Delta metric is a standardized scale where positive values indicate items favoring the focal group and negative values indicate items disfavoring the focal group. The magnitude of MH D-DIF is what drives the assignment into Category A, B, or C.
One reason the MH method became the standard is its simplicity. It works well with binary scored items (correct or incorrect), requires no complex model fitting, and produces interpretable output. For polytomous items, ETS uses a generalization called the standardized P-DIF or the Mantel procedure.
It is worth noting that MH assumes uniform DIF. That means the method assumes the difference between groups is constant across all ability levels. When DIF varies in direction or magnitude across ability levels, more advanced methods like logistic regression or item response theory (IRT) likelihood ratio tests are needed.
ETS DIF Classification Categories: A, B, and C
The ETS DIF classification categories divide flagged items into three tiers based on the MH D-DIF value and its statistical significance. Each category carries specific flagging rules and operational consequences.
Below I break down each category with its exact thresholds, the plus and minus sign conventions, and what actions each typically triggers. The rules differ slightly for binary items (scored 0 or 1) versus polytomous items (scored across multiple ordered categories), so I cover both.
Category A: Negligible or No DIF
Category A consists of items that show no meaningful differential item functioning. According to the ETS flagging rules, an item lands in Category A when the MH D-DIF is not significantly different from zero at the 0.05 level, or when its absolute value is less than 1.0.
In practical terms, Category A items function essentially the same way for both the reference and focal groups. Test developers treat these items as fair from a statistical standpoint, and they typically remain in the operational pool without any content review for fairness.
The threshold of 1.0 on the Delta scale is not arbitrary. ETS research showed that differences below this value rarely correspond to content issues identifiable by review panels. Setting the bar here prevents fairness committees from chasing statistical noise.
Most well-constructed test items fall into Category A. In my experience reviewing DIF output across standardized assessments, 80 to 90 percent of items land here. That is the expected outcome when items are written, reviewed, and field-tested with fairness in mind.
A plus or minus sign can still appear with Category A (A+ or A-), indicating the direction of the small difference. However, because the magnitude falls below 1.0, the direction carries no practical weight and these items are not flagged.
Category B: Moderate DIF
Category B captures items with moderate differential item functioning. An item receives a Category B classification when the MH D-DIF is significantly different from zero at the 0.05 level and its absolute value falls between 1.0 and 1.5 inclusive.
These items show a real, statistically detectable difference between groups, but the magnitude is not severe enough to trigger automatic removal. The standard operational response is a content review by a fairness committee or sensitivity panel.
During that review, committee members examine the item for construct-irrelevant variance. They ask whether cultural references, vocabulary load, or contextual assumptions might explain the group difference. If reviewers identify a plausible content issue, the item is revised or retired.
If the review finds no content explanation, the item may remain in the pool but is monitored closely in subsequent administrations. ETS tracks Category B items over time to see whether the DIF pattern persists, intensifies, or disappears.
The B+ and B- signs indicate direction. A B+ means the item favors the focal group, while B- means it disfavors the focal group. Direction matters here because it helps reviewers focus their attention on the right content concerns.
Category C: Significant DIF
Category C is the most serious classification. An item lands in Category C when the MH D-DIF is significantly different from zero and its absolute value is 1.5 or greater.
These items exhibit large differential functioning that almost always reflects construct-irrelevant variance. The standard practice is to remove Category C items from the operational pool. In rare cases, an item might be retained if content review finds a compelling construct-related explanation, but this is uncommon.
In practice, Category C items are rare on published tests because field testing catches most of them before they reach operational use. When I have seen Category C flags on live assessments, they almost always trace back to a specific content issue like regional vocabulary, dated references, or scoring rubric ambiguity.
The C+ and C- signs follow the same convention as the other categories. A C- item is particularly concerning because it disadvantages the focal group by a large margin. These items receive priority attention in fairness reporting and remediation.
It is worth emphasizing that Category C is a statistical classification, not a definitive judgment of bias. However, the combination of large effect size and statistical significance makes content problems highly likely, so the presumption favors removal.
Polytomous Item Classifications: AA, BB, and CC
The ETS DIF classification categories extend to polytomous items, which are items scored across more than two ordered categories. Think of constructed-response questions scored on a 0 to 4 rubric or Likert-type survey items.
For polytomous items, ETS uses the standardized P-DIF statistic rather than the MH D-DIF. The classifications follow a parallel structure but use double letters to distinguish them from binary item categories.
Category AA items show negligible or no DIF, with the standardized P-DIF not significantly different from zero or with an absolute value below 0.10 in some implementations. These items function comparably across groups and require no action.
Category BB items show moderate DIF, triggering content review. Category CC items show significant DIF and typically warrant removal. The logic mirrors the binary A, B, C system, just adapted for the polytomous scoring context.
If you work with mixed-format assessments, pay close attention to which classification system appears in your output. Confusing a BB flag with a B flag can lead to misapplying binary item thresholds to a polytomous item.
Uniform vs Non-Uniform DIF
Another distinction worth understanding is uniform versus non-uniform DIF. The ETS DIF classification categories I described above are designed primarily for uniform DIF.
Uniform DIF means the difference between reference and focal groups is consistent across all ability levels. If the focal group struggles more on an item, they struggle more regardless of whether they are high or low ability overall. The MH method detects this type well.
Non-uniform DIF means the direction or magnitude of the group difference changes across ability levels. For example, an item might favor the focal group at low ability levels but disfavor them at high ability levels. The MH method cannot detect this pattern reliably.
For non-uniform DIF, analysts turn to logistic regression, which includes interaction terms between group membership and total score. IRT-based methods like lordif or the mirt package in R also handle non-uniform DIF. These approaches model the item response function directly and can identify crossover effects.
Why does this matter for the ETS categories? Because an item can show no uniform DIF (and thus land in Category A) while still exhibiting problematic non-uniform DIF. This is one reason experienced psychometricians run multiple DIF methods rather than relying solely on MH.
How to Interpret DIF Analysis Results: Step by Step
Interpreting DIF output becomes straightforward once you know the sequence. Here is the step-by-step process I use when reviewing DIF analysis results.
Step 1: Confirm your sample sizes. ETS recommends at least 200 examinees per group for stable MH estimates. Smaller samples inflate Type I error rates and produce unstable D-DIF values. If your groups are small, treat all classifications with caution.
Step 2: Check the MH D-DIF value and its standard error. The value tells you the magnitude and direction of DIF on the Delta scale. The standard error tells you how precise that estimate is.
Step 3: Compare against the ETS thresholds. Absolute D-DIF below 1.0 with non-significance points to Category A. Values between 1.0 and 1.5 with significance point to Category B. Values of 1.5 or greater with significance point to Category C.
Step 4: Note the direction sign. Positive values favor the focal group, negative values disfavor it. This guides where to focus your content review.
Step 5: Send B and C flagged items to a fairness committee for content review. Provide the committee with the item, the statistics, and the group definitions so they can evaluate construct relevance.
Step 6: Document your decisions. Whether you keep, revise, or remove an item, record the rationale. This documentation supports fairness audits and accreditation reviews.
Common Misconceptions About ETS DIF Classification Categories
Several misconceptions about the ETS DIF classification categories circulate in graduate programs and online forums. I want to address the ones I encounter most often.
Misconception one: DIF equals bias. As I noted earlier, DIF is a statistical flag. Bias is a content judgment. A Category C item might reflect genuine bias, but it might also reflect a measurement artifact or sampling quirk. Always pair statistics with content review.
Misconception two: Category A means the item is perfectly fair. Category A means no detectable DIF at the group level with the current sample. It does not guarantee the item is free of all fairness concerns, especially for subgroups too small to analyze separately.
Misconception three: Non-significance means no problem. With small samples, an item can have a meaningful D-DIF value that fails to reach significance. This is why ETS uses both significance and magnitude thresholds together rather than relying on either alone.
Misconception four: The MH method catches all DIF. As I explained in the uniform versus non-uniform section, MH detects uniform DIF only. Supplementing with logistic regression or IRT methods provides a more complete picture.
FAQs
What is DIF in education?
DIF, or Differential Item Functioning, occurs when a test item performs differently for different demographic groups after controlling for overall ability. It is a statistical flag used in test fairness analysis to identify items that may contain construct-irrelevant variance affecting specific groups.
What is a DIF analysis?
A DIF analysis is a statistical procedure that compares item-level performance between a reference group and a focal group across ability levels. The most common method is Mantel-Haenszel, which produces a D-DIF statistic used to classify items into ETS categories A, B, or C based on severity.
What are the different types of DIF?
There are two main types of DIF. Uniform DIF means the group difference is constant across all ability levels and is detectable by the Mantel-Haenszel method. Non-uniform DIF means the difference changes in direction or magnitude across ability levels and requires methods like logistic regression or IRT-based likelihood ratio tests.
What does Category B DIF mean?
Category B DIF means an item shows moderate differential functioning, with an MH D-DIF value that is statistically significant and has an absolute value between 1.0 and 1.5 on the Delta scale. These items are typically sent for content review to determine whether a construct-irrelevant factor explains the group difference.
What sample size is needed for reliable DIF detection?
ETS recommends a minimum of 200 examinees per group for stable Mantel-Haenszel DIF estimates. Smaller sample sizes increase Type I error rates and produce unstable D-DIF values. Larger samples, ideally 500 or more per group, improve the precision of classification into the A, B, and C categories.
Conclusion
Understanding the ETS DIF classification categories gives you a structured framework for evaluating test fairness at the item level. Category A signals negligible DIF, Category B flags moderate differences requiring content review, and Category C identifies significant DIF that typically leads to item removal.
The Mantel-Haenszel method powers the binary item classifications, while the standardized P-DIF extends the system to polytomous items through the AA, BB, and CC labels. Pairing these statistical categories with content review and adequate sample sizes produces defensible fairness decisions.
If you are reviewing DIF output for the first time, work through the step-by-step interpretation process I outlined. Start with sample sizes, check the D-DIF values against the thresholds, note the direction signs, and document every decision. That discipline is what separates rigorous fairness analysis from a box-checking exercise.