Differential item functioning (DIF) is a statistical method used to identify test items that function differently for different groups of test-takers who have equal ability on the construct being measured. When two people with the same underlying skill level have different probabilities of answering a question correctly simply because they belong to different demographic groups, that item may be measuring something other than what the test intends to measure. In this guide, we break down what differential item functioning is and why it matters for fairness in educational and psychological assessment.
Testing organizations, school districts, licensing boards, and admissions committees all rely on DIF analysis as a routine part of building fair assessments. The conversation around test fairness has grown louder in 2026, with major universities reconsidering their reliance on standardized tests and policymakers scrutinizing whether exams measure real ability or demographic background. Understanding what DIF is, how it is detected, and what to do when it shows up is essential for anyone involved in assessment design, validation, or policy.
Our team has spent years working with psychometric data, and we have seen firsthand how a single poorly functioning item can disadvantage thousands of capable test-takers. We wrote this guide to bridge the gap between technical psychometric research and the practical questions that test developers, educators, and policymakers actually ask. Whether you are a graduate student learning psychometrics for the first time or a seasoned test developer looking for a refresher on DIF detection methods, you will find clear explanations, real examples, and actionable guidance ahead.
Table of Contents
At-a-Glance: DIF Essentials
Before diving into the details, here are the core facts about differential item functioning that every test developer and assessment professional should know.
Definition: DIF occurs when test-takers from different groups with equal ability have different probabilities of answering an item correctly.
Goal: Identify items that may measure construct-irrelevant knowledge, cultural assumptions, or language differences rather than the intended skill.
Key distinction: DIF is a statistical signal, not proof of bias. Bias is an interpretive judgment that requires expert review.
Two types: Uniform DIF means one group is consistently advantaged across all ability levels. Non-uniform DIF means the advantage crosses or changes at different ability levels.
Main detection methods: Mantel-Haenszel, logistic regression, item response theory (IRT), and SIBTEST are the most widely used.
Classification: The ETS A/B/C scheme classifies items by DIF magnitude: A (negligible), B (moderate), and C (large).
Why it matters: DIF analysis is a cornerstone of fairness evidence required by the Standards for Educational and Psychological Testing.
What differential item functioning is and why it matters for fairness
Differential item functioning is present when test-takers from different demographic groups who are matched on the ability the test is supposed to measure nevertheless have different probabilities of giving a correct response. The key phrase is “who are matched on ability.” DIF analysis does not simply compare raw scores between groups. It controls for overall ability first, then looks for residual differences at the item level.
Think of it this way. Imagine two students, one from an affluent suburb and one from an under-resourced rural school, both scoring 75 percent on a math exam overall. If the suburban student answers a specific word problem correctly far more often than the rural student at the same overall ability level, that item may be functioning differently across the two groups. The item might be tapping into vocabulary, cultural context, or reading comprehension rather than mathematical reasoning alone.
The technical vocabulary in DIF analysis uses two terms that we will use throughout this guide. The reference group is typically the majority or more advantaged group, often used as the baseline for comparison. The focal group is the group of interest, usually a marginalized or underrepresented group whose performance we want to examine for fairness. These labels are conventional, not evaluative. The reference group is not “better,” and the focal group is not “deficient.” The terms simply describe which group serves as the comparison baseline.
Several group comparisons are common in DIF analysis. Gender comparisons examine whether items function differently for male and female test-takers. Racial and ethnic comparisons look at performance across demographic categories. Language background comparisons are critical when some test-takers are not native speakers of the test language. Socioeconomic status, disability status, and age are also frequently examined. The choice of grouping variable depends on the test context, the population being assessed, and the fairness questions that matter most for that assessment.
DIF is fundamentally about construct-irrelevant variance. If an item is supposed to measure algebraic reasoning but also requires familiarity with a specific cultural context that some test-takers lack, the item is measuring something beyond the intended construct. That extra something creates systematic advantage or disadvantage unrelated to the skill the test claims to assess. Identifying and addressing this construct-irrelevant variance is the core purpose of DIF analysis.
DIF vs Item Impact: The Critical Distinction
One of the most common misunderstandings in psychometric fairness work is confusing DIF with item impact. These are different concepts, and the distinction matters enormously for interpreting fairness evidence correctly.
Item impact refers to raw differences in performance between groups on a test item. If 70 percent of male test-takers answer an item correctly compared to 55 percent of female test-takers, that raw gap is item impact. It tells you the groups performed differently, but it tells you nothing about why.
DIF goes further. It asks whether that gap persists after you match the groups on overall ability. If you compare only male and female test-takers who all scored the same on the test as a whole, and the gender gap on that specific item disappears, then there was no DIF. The original raw gap was simply a reflection of the fact that the two groups had different overall ability distributions. Item impact without DIF is not necessarily a fairness problem.
A helpful analogy comes from Holland and Wainer’s foundational 1993 work on DIF. Imagine you are comparing the heights of men and women at a basketball camp. Men are taller on average, which is a raw difference, analogous to item impact. But if you only compare men and women who are the same height, and then check whether they differ on some other measure like vertical jump, that residual difference would be analogous to DIF. The matching step is what separates genuine item-level unfairness from legitimate group-level ability differences.
This distinction matters because eliminating all raw score gaps between groups is not the goal of fair test design. Tests should accurately reflect real differences in the construct being measured. The goal of DIF analysis is to ensure that no item introduces differences that are unrelated to the construct. When practitioners confuse impact with DIF, they either dismiss real fairness problems or create false alarms based on normal ability variation.
Types of DIF: Uniform vs Non-Uniform
DIF is not a one-size-fits-all phenomenon. It comes in two main forms that require different detection approaches and carry different implications for test design.
Uniform DIF occurs when one group is consistently advantaged across all ability levels. Picture two item characteristic curves, one for the reference group and one for the focal group, plotted on the same graph. The curves are the same shape, but one sits shifted to the right or left of the other. This means the item is uniformly easier or harder for one group at every ability level. A vocabulary item that uses a word more familiar in one regional dialect might show uniform DIF, with speakers of that dialect consistently more likely to answer correctly at every ability level.
Non-uniform DIF is more complex. Here the advantage reverses or changes magnitude at different points along the ability spectrum. The item characteristic curves cross each other rather than running parallel. Low-ability members of the focal group might be disadvantaged while high-ability members are advantaged, or vice versa. Non-uniform DIF often indicates that the item is tapping into different skills or strategies at different ability levels, and that those skills are distributed differently across groups.
Non-uniform DIF is harder to detect and interpret than uniform DIF. The Mantel-Haenszel procedure, the most popular DIF detection method, was originally designed for uniform DIF and can miss non-uniform patterns. Logistic regression and IRT-based methods are better suited for detecting non-uniform DIF because they can model the interaction between group membership and ability level.
The practical implication is significant. If you only screen for uniform DIF, you may miss items that disadvantage test-takers at specific ability levels. A high-ability student from the focal group might be penalized on an item that appears fair when you look at average performance. This is why method choice matters and why many testing programs use multiple detection methods in combination.
How DIF Is Detected: Statistical Methods
DIF detection relies on statistical procedures that compare item performance across groups after controlling for overall ability. Four methods dominate the field, each with strengths, limitations, and typical use cases.
The Mantel-Haenszel Procedure
The Mantel-Haenszel (MH) procedure is the workhorse of DIF detection. It was adapted for psychometric use by Holland and Thayer in 1988 and remains the most widely taught and applied DIF method. The procedure works by dividing test-takers into ability strata, typically based on total test score, then computing an odds ratio that summarizes how much more likely the reference group is to answer correctly compared to the focal group within each stratum.
The MH statistic combines these stratum-level odds ratios into a single summary measure on the delta scale, where positive values indicate the item favors the reference group and negative values indicate it favors the focal group. The beauty of the MH approach is its simplicity and interpretability. The delta-diff value directly tells you the magnitude and direction of DIF in a metric that aligns with the ETS classification system.
The MH procedure has limitations. It is designed for dichotomous items, requires a reasonably large sample, and assumes uniform DIF. It cannot detect non-uniform DIF. Despite these limitations, it remains the first-line DIF screening method in most large-scale testing programs because it is computationally efficient, easy to explain, and well understood by the psychometric community.
Logistic Regression for DIF
Logistic regression extends DIF detection by modeling the probability of a correct response as a function of total score, group membership, and the interaction between the two. The procedure fits three nested models. The first includes only the total score as a predictor. The second adds group membership. The third adds the group-by-score interaction.
If adding group membership significantly improves model fit, you have evidence of uniform DIF. If adding the interaction term improves fit beyond that, you have evidence of non-uniform DIF. This makes logistic regression more flexible than the MH procedure because it can detect both types of DIF in a single analysis.
Logistic regression also handles polytomous items, those with more than two response categories like Likert scales, by using ordinal or multinomial variants. The R packages lordif and difNLR implement logistic regression-based DIF detection and are widely used in both academic research and operational testing programs. Practitioners on psychometrics forums frequently recommend these packages for anyone working with real assessment data.
IRT-Based DIF Detection
Item response theory provides the most theoretically grounded approach to DIF detection. In an IRT framework, each item is described by parameters like difficulty and discrimination. DIF is detected by comparing whether these item parameters are the same for the reference and focal groups. If the difficulty parameter differs between groups after placing both groups on the same ability scale, the item exhibits DIF.
IRT-based methods include likelihood ratio tests that compare models with and without group-specific item parameters. These methods can detect both uniform and non-uniform DIF because differences in the difficulty parameter indicate uniform DIF while differences in the discrimination parameter indicate non-uniform DIF.
The advantage of IRT-based DIF detection is its precision and its ability to work within the broader IRT framework that many testing programs already use for scoring and equating. The disadvantage is computational complexity and the need for larger sample sizes than MH or logistic regression require. IRT methods also assume the test data fits a specific parametric model, which is not always the case.
SIBTEST and Other Methods
The SIBTEST (Simultaneous Item Bias Test) procedure was developed by Shealy and Stout in 1993 as a method that controls for measurement error in the matching variable. It is particularly useful when the total test score is itself contaminated by DIF on multiple items, which can create a noisy matching variable. SIBTEST uses a regression-based approach to estimate the DIF effect while accounting for this measurement error.
Other specialized methods include the Dorans and Kulick standardization approach, which is particularly effective for small focal groups, and various Bayesian approaches that incorporate prior information about expected DIF patterns. Machine learning methods for DIF detection are an emerging area, with researchers exploring classification algorithms and tree-based methods to complement traditional statistical approaches.
For a quick comparison, here is how the four main methods stack up. The MH procedure is best for large-scale uniform DIF screening with dichotomous items. Logistic regression is the most flexible general-purpose method. IRT-based methods are ideal when the testing program already uses IRT for scoring. SIBTEST is valuable when measurement error in the matching variable is a concern. Most operational testing programs use at least two of these methods in combination to maximize detection coverage.
Effect Sizes and the ETS A/B/C Classification
Statistical significance alone is not enough to decide whether an item’s DIF is meaningful. With large samples, even trivially small DIF effects can reach statistical significance. With small samples, large effects might fail to reach significance. This is why DIF analysis relies on effect size measures and classification schemes to separate meaningful DIF from statistical noise.
The most widely used classification system comes from Educational Testing Service (ETS). Items are classified into three categories based on both the magnitude of the MH delta-diff statistic and its statistical significance. Category A items have negligible DIF. Either the delta-diff is less than 1.0 in absolute value, or it is not statistically significant. These items are considered free of meaningful DIF and require no action.
Category B items have moderate DIF. The delta-diff is between 1.0 and 1.5 in absolute value and is statistically significant. These items warrant review. They may be retained if the DIF is judged to be construct-relevant or if no better replacement item is available, but they should be flagged for monitoring in future administrations.
Category C items have large DIF. The delta-diff exceeds 1.5 in absolute value and is statistically significant. These items are considered serious candidates for removal or revision. Most testing programs automatically flag Category C items for content review and either remove them from the scoring or revise the item content before the next administration.
The ETS classification is not the only system, but it is the most influential. Other organizations have adapted it with different thresholds or additional categories. The key principle is the same regardless of the specific thresholds used. DIF magnitude, not just statistical significance, should drive decisions about item retention, revision, or removal.
Anchor Items and Iterative Purification
A technical challenge in DIF analysis is that the matching variable, usually the total test score, may itself contain DIF. If 20 percent of your test items exhibit DIF, then using the total score to match test-takers on ability means you are matching them on a contaminated measure. This can mask real DIF on some items and create false DIF signals on others.
The solution is anchor item selection and iterative purification. The process works as follows. First, you identify a subset of items that are unlikely to show DIF. These anchor items form the basis for a cleaner matching variable. Second, you run DIF analysis using only the anchor items to compute the matching score. Third, you check whether any previously clean items now show DIF under the purified matching variable. Fourth, you remove any newly flagged items from the anchor set and repeat the process until the anchor set stabilizes.
This iterative purification process typically converges within three to five cycles. The result is a matching variable that is substantially free of DIF contamination, which improves the accuracy of DIF detection for all remaining items. Most modern DIF software, including the difR package in R, implements automated purification as an option.
Practitioners on psychometrics forums often ask when to stop purifying. The general guidance is to continue until the anchor set no longer changes between iterations. If the process does not converge within about five iterations, it may indicate that DIF is so widespread in the test that a different analytical approach is needed, such as multi-group IRT or a complete test review.
DIF Is a Statistical Signal, Not a Bias Verdict
Perhaps the most important interpretive principle in DIF analysis is this: DIF is a statistical signal, not a determination of bias. This distinction is critical and is frequently misunderstood by practitioners who do not specialize in psychometrics.
When an item is flagged for DIF, it means that a statistical procedure detected differential performance between groups after ability matching. It does not mean the item is biased. Bias is a substantive judgment about whether the differential performance is unfair, meaning it stems from construct-irrelevant factors that should not influence test scores.
An item can show DIF for construct-relevant reasons. Consider a reading comprehension passage about baseball on a test designed to measure reading skill. If male test-takers consistently outperform female test-takers at matched ability levels because they have more prior knowledge of baseball, the item shows DIF. But if the test is specifically designed to measure the ability to comprehend unfamiliar technical passages, and baseball knowledge is not supposed to be a prerequisite, then the DIF reflects a fairness problem. If the test is explicitly about sports knowledge, the same DIF might be construct-relevant and not a bias concern.
This is why a DIF panel review is a standard part of the DIF follow-up process. The typical panel review follows these steps. First, a psychometrician presents the DIF statistics to a content review panel. Second, the panel examines the item content, including the stem, options, and any reading passages or visual materials. Third, panel members discuss whether they can identify a construct-irrelevant factor that might explain the differential performance. Fourth, the panel makes a recommendation: retain the item, revise it, or remove it from scoring. Fifth, all decisions are documented as fairness evidence.
The panel typically includes content experts, fairness specialists, and representatives from the communities potentially affected by the item. Including diverse perspectives on the panel helps ensure that subtle cultural or linguistic issues are identified that might not be obvious to a homogenous reviewer group. Forum discussions among practitioners consistently highlight the value of diverse panels and the danger of relying solely on statistical flags without expert content review.
Why DIF Analysis Matters for Test Fairness
Understanding what differential item functioning is and why it matters for fairness requires connecting the statistical machinery to real-world consequences for test-takers. DIF analysis matters for at least six concrete reasons.
First, it protects individual test-takers from construct-irrelevant disadvantage. A single DIF item on a high-stakes exam can move a test-taker’s score enough to change an admissions or licensing decision. When that item is measuring something other than the intended construct, the resulting decision is not just statistically noisy. It is unfair.
Second, it provides systematic fairness evidence. The Standards for Educational and Psychological Testing, jointly produced by AERA, APA, and NCME, explicitly require test developers to investigate item-level fairness. DIF analysis is the primary statistical tool for generating this evidence. Testing programs that skip DIF analysis are not meeting professional standards for test validation.
Third, it builds public trust in assessments. When testing organizations can demonstrate that every item has been screened for DIF and that flagged items undergo expert review, they provide tangible evidence that fairness is taken seriously. This is particularly important for high-stakes tests like college admissions exams, professional licensing tests, and state accountability assessments.
Fourth, it identifies content problems early in test development. DIF analysis during item tryout can catch problematic items before they appear in operational forms. This is far less costly than discovering fairness problems after results have been reported and decisions have been made. The reddit psychometrics community frequently discusses how DIF screening has become a standard part of the item development pipeline at major testing companies.
Fifth, it supports valid score interpretation. If items function differently across groups, the meaning of a score may differ depending on who took the test. A score of 500 might represent a slightly different mix of skills for different demographic groups. DIF analysis helps ensure that score interpretations are comparable across populations, which is the foundation of measurement invariance.
Sixth, it informs test adaptation and translation. When tests are translated or adapted for use in different linguistic or cultural contexts, DIF analysis is essential for verifying that items function the same way in the adapted version. Items that function differently across language groups may need revision to maintain the construct equivalence that valid cross-cultural comparison requires.
Practical Applications of DIF Analysis
DIF analysis is applied across a wide range of assessment contexts. Understanding where and how it is used helps illustrate its practical value for fairness.
In large-scale educational assessment, state testing programs routinely run DIF analysis on every operational item. The Smarter Balanced assessment consortium, for example, publishes DIF results as part of its technical reports. Items flagged with moderate or large DIF undergo content review before being used in subsequent administrations. This routine screening ensures that the assessments used for school accountability and student progress monitoring are fair across demographic groups.
In admissions testing, organizations like the College Board and ACT conduct extensive DIF analysis on every form of their exams. The SAT and ACT serve millions of test-takers annually, and even small item-level fairness problems can affect large numbers of students. DIF screening is one layer in a multi-step fairness review process that also includes sensitivity review of content during item writing.
In professional licensing and certification, exams for medicine, nursing, law, and other professions use DIF analysis to ensure that licensing decisions are not influenced by demographic background. Medical licensing exams, for instance, routinely screen for DIF across racial, ethnic, and language groups. Items showing large DIF are reviewed by content experts before being included in scoring.
In computer-adaptive testing (CAT), DIF analysis faces additional complexity because each test-taker sees a different set of items. The item selection algorithm adapts to the test-taker’s estimated ability, which means group differences in item exposure can complicate DIF detection. Specialized CAT-aware DIF methods account for the adaptive selection mechanism to ensure that fairness screening remains valid in this context.
In test translation and adaptation, DIF analysis verifies that translated items measure the same construct in the target language. Direct translation often introduces linguistic or cultural differences that change how an item functions. Back-translation and cognitive interviewing complement DIF analysis in the adaptation process, but DIF provides the statistical evidence that the adapted items perform equivalently across language groups.
In health outcomes and patient-reported measures, DIF analysis is increasingly used to verify that scales measuring depression, anxiety, quality of life, or functional status function the same way across demographic groups. A depression item that references symptoms differently expressed across cultures could introduce measurement bias that compromises clinical decisions.
Common Misconceptions About DIF
Forum discussions among practitioners reveal several persistent misconceptions about DIF that can lead to misinterpretation or misuse of DIF results. Addressing these directly is part of understanding what differential item functioning is and why it matters for fairness.
Misconception 1: DIF means the item is biased. As discussed above, DIF is a statistical signal. Bias is a substantive interpretation that requires content review. Many items flagged for DIF turn out to have construct-relevant explanations or are false positives that do not replicate across administrations.
Misconception 2: No DIF means the test is fair. DIF is item-level analysis. A test can have no items with significant DIF but still show group-level score gaps due to genuine ability differences, curriculum differences, or systemic factors outside the test’s control. DIF is necessary but not sufficient for fairness evidence.
Misconception 3: Larger DIF always means a worse item. DIF magnitude depends on both the size of the group difference and the sample size. A large DIF statistic with a small sample may be less reliable than a moderate DIF statistic with a large sample. Effect size interpretation must consider the context of the specific item and test.
Misconception 4: DIF analysis replaces content sensitivity review. DIF is a post-hoc statistical method. Content sensitivity review during item writing catches many potential fairness problems before data are even collected. The two processes are complementary, not substitutes. The best testing programs use both.
Misconception 5: Any item with DIF must be removed. Removal decisions depend on the magnitude of DIF, the content review outcome, the availability of replacement items, and the overall test blueprint. Blindly removing all DIF items can shorten the test, reduce content coverage, and introduce new problems. The ETS classification system exists precisely to calibrate response to DIF magnitude.
What to Do When DIF Is Detected: A Decision Framework
When DIF analysis flags an item, test developers face a decision about what to do next. The following framework, drawn from professional standards and practitioner experience, provides guidance.
For Category A items (negligible DIF), no action is needed. Retain the item in the scoring and continue using it in future forms.
For Category B items (moderate DIF), convene a content review panel. If the panel identifies a clear construct-irrelevant source of the DIF and a better replacement item exists, consider revising or replacing. If no clear explanation is found and the DIF magnitude is not extreme, document the review and retain the item with monitoring in future administrations.
For Category C items (large DIF), conduct an immediate content review. In most cases, Category C items should be removed from scoring for the current administration if the analysis is available before scores are finalized. If the analysis is post-hoc, document the finding, exclude the item from future forms, and investigate whether the current administration’s scores need adjustment.
Across all categories, document every decision as fairness evidence. The documentation should include the DIF statistics, the content review findings, the decision rationale, and the names and qualifications of the panel members. This documentation is essential for meeting professional standards and for defending testing decisions if they are ever challenged.
FAQs
What is differential item functioning?
Differential item functioning (DIF) is a statistical phenomenon that occurs when test-takers from different demographic groups who have equal ability on the measured construct nevertheless have different probabilities of answering a specific test item correctly. DIF analysis identifies these items by matching test-takers on overall ability and then comparing item-level performance between a reference group and a focal group within those matched ability strata.
What is the difference between DIF and item impact?
Item impact refers to raw differences in performance between groups on a test item, without controlling for overall ability. DIF goes further by matching groups on ability first and then checking whether the item-level difference persists. A raw score gap that disappears after ability matching is item impact without DIF and does not necessarily indicate a fairness problem.
What is the difference between uniform and non-uniform DIF?
Uniform DIF means one group is consistently advantaged across all ability levels, with item characteristic curves that are parallel but shifted. Non-uniform DIF means the advantage changes or reverses at different ability levels, producing item characteristic curves that cross. Uniform DIF is detectable with the Mantel-Haenszel procedure, while non-uniform DIF requires logistic regression or IRT-based methods.
How is DIF detected?
DIF is detected using statistical methods that compare item performance across groups after matching on ability. The four main methods are the Mantel-Haenszel procedure, logistic regression, item response theory based approaches, and SIBTEST. Each method has strengths and limitations, and most operational testing programs use at least two methods in combination.
Does DIF mean a test item is biased?
No. DIF is a statistical signal that indicates differential performance between groups after ability matching. Bias is a substantive interpretation that requires expert content review. An item can show DIF for construct-relevant reasons, and some DIF flags are false positives. A content review panel examines flagged items to determine whether the differential performance reflects construct-irrelevant unfairness.
What is the difference between DIF and DTF?
DIF (differential item functioning) examines fairness at the level of individual test items. DTF (differential test functioning) examines whether the test as a whole functions differently across groups, typically by summing or aggregating DIF effects across all items. An item with large DIF may have minimal impact on DTF if the test is long and other items cancel out the effect. Both analyses are part of comprehensive fairness evidence.
How can DIF analysis improve test fairness?
DIF analysis improves fairness by identifying items that may be measuring construct-irrelevant factors, by providing systematic statistical evidence required by professional testing standards, by catching problematic items early in test development, and by supporting valid score interpretation across demographic groups. When combined with content sensitivity review and diverse panel input, DIF analysis forms the backbone of a credible fairness program.
Conclusion: Why DIF Matters Now
Differential item functioning analysis is one of the most powerful tools available for building fair assessments. It provides a systematic, statistically grounded method for identifying test items that may disadvantage capable test-takers based on demographic background rather than genuine ability. Understanding what differential item functioning is and why it matters for fairness is essential for anyone involved in test development, validation, or educational policy.
The key takeaways are straightforward. DIF analysis controls for ability before comparing groups, which distinguishes genuine item-level unfairness from legitimate ability differences. DIF is a statistical signal, not a bias verdict, and requires expert content review to interpret. The ETS A/B/C classification helps calibrate response to DIF magnitude. Multiple detection methods should be used in combination for comprehensive coverage. And every DIF finding should be documented as part of the fairness evidence required by professional testing standards.
As assessment continues to play a central role in education, employment, and professional licensing, the stakes of fairness in testing only grow. DIF analysis is not a one-time checkbox. It is an ongoing practice that should be embedded in every stage of test development, from item writing through operational use. For practitioners looking to deepen their skills, the foundational work by Holland and Wainer (1993), the practical guidance from APA Division 5, and the continually updated R packages like lordif and difNLR provide excellent starting points. Fair tests build public trust, protect individual test-takers, and ensure that scores mean what they claim to mean for everyone who takes them.