EPPP: PART 1, KNOWLEDGE • DOMAIN 5: ASSESSMENT AND DIAGNOSIS

Test Bias — Identify sources of cultural bias and fairness issues in assessment instruments

Understanding how cultural, linguistic, and structural factors compromise the validity and equity of psychological assessments.

Historical Context & Motivation

The history of psychological testing is deeply intertwined with questions of fairness, equity, and cultural context. From the earliest intelligence measures developed in the early twentieth century, assessment instruments have been scrutinized for the extent to which they produce systematically different outcomes across racial, ethnic, linguistic, and socioeconomic groups. The concept of test bias refers to systematic error in measurement that differentially affects identifiable subgroups of examinees, leading to scores that do not accurately reflect the construct being measured for those groups. This is distinct from mere group differences in scores, which may or may not reflect bias depending on whether those differences correspond to genuine differences in the construct or to artifacts of the instrument itself.

The stakes of this issue in behavioral health are enormous. Assessment instruments guide diagnostic decisions, treatment planning, forensic evaluations, educational placements, and personnel selection. When an instrument is biased, it can lead to misdiagnosis, inappropriate interventions, denial of services, or unjust legal outcomes — all of which disproportionately burden historically marginalized communities. Understanding the sources and detection of test bias is therefore not merely an academic exercise but a core ethical competency for psychologists preparing for licensure.

1905
Binet-Simon Scale
Alfred Binet develops the first standardized intelligence test in France. When translated and adapted for American use by Lewis Terman (the Stanford-Binet, 1916), the normative samples overwhelmingly represented white, English-speaking populations, embedding cultural assumptions into the instrument from the outset.
1969
Jensen's Controversial Claims
Arthur Jensen publishes his provocative article arguing that racial differences in IQ scores reflect genetic factors. This galvanizes the field into rigorous examination of whether such score differences are attributable to test bias, environmental factors, or genuine construct differences — launching decades of psychometric research on fairness.
1979
Larry P. v. Riles
A landmark federal court case rules that IQ tests used for placing African American children into classes for the 'educable mentally retarded' are culturally biased. The ruling bans the use of standardized intelligence tests for this purpose in California, illustrating the legal consequences of biased assessment.
1999
Standards for Educational and Psychological Testing
The AERA, APA, and NCME publish updated joint Standards that include comprehensive guidelines on fairness in testing, formally codifying the responsibility of test developers and users to examine and mitigate sources of bias across cultural groups.
2014
Updated Standards & Modern Frameworks
The most recent edition of the Standards introduces an expanded chapter on fairness, distinguishing among multiple definitions of fairness and emphasizing validity evidence for diverse populations. Modern frameworks increasingly integrate intersectionality, recognizing that bias may operate differently across overlapping social identities.

The central question that this body of work addresses is deceptively simple: When members of different cultural groups obtain different scores on a psychological test, does the test itself contribute to those differences in ways that are unrelated to the construct being measured? Answering this question requires a sophisticated understanding of validity, reliability, item analysis, and the sociocultural contexts in which assessments are administered and interpreted.

Core Principles & Definitions

Before examining specific sources of bias, it is essential to distinguish among several related but distinct concepts. Test bias has a precise psychometric definition that differs from everyday usage. In psychometrics, a test is biased when it systematically over- or under-estimates the true score of members of a particular group. This is fundamentally a question of construct validity — whether the test measures the same construct in the same way across groups. Test fairness, by contrast, is a broader concept that encompasses bias but also includes considerations of equitable treatment, equal opportunity to learn, and appropriate use of test results in decision-making.

1

Construct Bias

The construct measured by the test is not equivalent across cultural groups. The test may tap into different psychological processes or domains of knowledge depending on the examinee's cultural background, rendering cross-group comparisons invalid at a fundamental level.
2

Method Bias

Systematic differences in scores arise from characteristics of the testing method rather than the construct. This includes differential familiarity with test formats (e.g., multiple-choice vs. oral), examiner–examinee dynamics, response styles, and environmental conditions of administration.
3

Item Bias (DIF)

Specific items function differently across groups even after controlling for overall ability. An item exhibits differential item functioning (DIF) when examinees of equal standing on the construct have unequal probabilities of endorsing or correctly answering that item as a function of group membership.
4

Predictive Bias

The test predicts criterion outcomes (e.g., job performance, academic achievement) differently for different groups. This is evaluated by examining whether the regression of the criterion on the test score yields different slopes or intercepts across groups — known as differential prediction.
5

Measurement Invariance

The statistical structure of the test (factor structure, factor loadings, intercepts) is equivalent across groups. Establishing measurement invariance through confirmatory factor analysis is the gold standard for demonstrating that a test measures the same construct in the same metric across populations.
KEY TAKEAWAY
Think of a psychological test like a ruler. If you use a ruler calibrated in inches to measure objects but the markings are systematically shifted by half an inch for every object painted blue, the ruler is biased against blue objects — not because blue objects are shorter, but because the instrument itself introduces error contingent on a property irrelevant to what you are measuring. Similarly, when a depression inventory's items resonate differently with certain cultural expressions of distress, the test introduces systematic error related to culture rather than to the severity of depression.

A critical distinction that the EPPP frequently tests is the difference between bias and unfairness. A test may show mean score differences between groups without being biased, provided it accurately reflects genuine differences in the construct. Conversely, a test with no mean score differences could still be unfair if, for example, access to the test or preparation materials is inequitably distributed. Understanding this distinction is essential for evaluating claims about assessment instruments in clinical, educational, and forensic contexts.

Visual Explanation — Sources of Bias in Assessment

This diagram illustrates the three primary categories of test bias — construct bias, method bias, and item bias (DIF) — along with their specific manifestations and the primary detection methods used for each. All three types ultimately threaten the construct validity of the assessment.

The diagram above organizes the primary sources of test bias using the widely cited taxonomy from van de Vijver and Tanzer (2004). At the highest level, an observed test score may be contaminated by any of three bias sources. Construct bias is the most fundamental threat because it means the test is not measuring the same psychological construct across groups — for example, when a measure of 'intelligence' actually taps into acculturation or socioeconomic knowledge for certain populations. Method bias operates at the level of the testing procedure and includes factors such as differential comfort with timed tests, the examiner's race or language matching, and cultural norms around guessing versus leaving items blank. Item bias is the most granular level and is assessed statistically through differential item functioning analyses that examine whether specific items perform differently across groups after controlling for overall ability.

Statistical Frameworks for Detecting Bias

While much of the discourse around test bias is conceptual and ethical, several rigorous statistical methods exist for detecting and quantifying bias. Clinicians preparing for the EPPP should be familiar with the logic underlying these approaches, even if they do not perform the computations themselves. The two most important frameworks are differential item functioning (DIF) analysis and differential prediction (also called predictive bias or slope/intercept bias). A third approach, measurement invariance testing via confirmatory factor analysis, provides the most comprehensive evidence of whether a test measures the same construct in the same metric across groups.

DIFFERENTIAL PREDICTION (CLEARY MODEL)
Ŷ = a + b(X)
Where Ŷ = predicted criterion score (e.g., GPA, job performance), X = test score, a = intercept, b = slope. A test shows predictive bias (per the Cleary definition) when the regression equation produces different slopes or intercepts for different groups, meaning the same test score predicts different criterion outcomes depending on group membership.

The Cleary model (1968) defines an unbiased test as one in which a common regression equation predicts criterion performance equally well for all groups. If separate regression lines for different groups have significantly different slopes or intercepts, the test is biased. Interestingly, research on major tests like the SAT and GRE has generally found that these instruments slightly overpredict the academic performance of minority students — meaning they predict higher performance than is actually achieved — which is the opposite of what many laypeople assume about test bias. This overprediction may itself reflect bias in the criterion (grades) rather than in the test.

DIFFERENTIAL ITEM FUNCTIONING (MANTEL-HAENSZEL)
αMH = Σₖ (Aₖ × Dₖ / Tₖ) / Σₖ (Bₖ × Cₖ / Tₖ)
The Mantel-Haenszel statistic compares the odds of a correct response for the focal group versus the reference group at each ability level k. If αMH ≈ 1.0, the item functions equivalently. Significant departures from 1.0, converted to a delta metric (Δ = −2.35 × ln(αMH)), indicate DIF. ETS classifies items as A (negligible), B (moderate), or C (large) DIF based on the magnitude and significance of Δ.

The concept of measurement invariance provides the most rigorous test of construct bias. Using confirmatory factor analysis, researchers test a hierarchy of invariance levels: configural invariance (same factor structure across groups), metric invariance (same factor loadings), scalar invariance (same intercepts), and strict invariance (same residual variances). Scalar invariance is typically required before mean comparisons across groups can be considered meaningful. Failure at any level suggests that the test is not measuring the construct equivalently across populations.

Clinical Implication
Even when a test has been shown to be statistically unbiased at the item and prediction levels, clinicians must remain alert to broader fairness concerns. A test may be psychometrically sound yet still produce inequitable outcomes if the criterion it predicts is itself contaminated by systemic factors. For example, if workplace performance ratings are influenced by supervisory prejudice, then a test that accurately predicts those ratings is validly predicting a biased criterion.

Detailed Classification of Bias Sources

Understanding the specific mechanisms through which bias enters assessment instruments is critical for both the EPPP and for ethical clinical practice. Bias can be introduced at every stage of the testing process — from the initial conceptualization of the construct, through item writing and norming, to administration, scoring, and interpretation. The following diagram and table provide a detailed breakdown of these sources organized by the stage at which they operate.

The assessment pipeline diagram shows how bias can enter at each stage of test development and use — from the initial definition of the construct through interpretation and decision-making — and the downstream consequences when bias goes unaddressed.
Major sources of cultural bias in psychological assessment
Bias SourceDescriptionClinical Example
Linguistic biasItems require English proficiency beyond what is needed for the measured construct, disadvantaging non-native speakers.A depression screener with complex syntax may underestimate depression in an immigrant patient with limited English who actually experiences significant depressive symptoms.
Content biasItems reference knowledge, situations, or objects that are more familiar to certain cultural groups than others.An IQ test item asking 'What should you do if you find a stamped, addressed envelope on the street?' assumes familiarity with the postal system, disadvantaging examinees from rural or non-Western backgrounds.
Norm sample biasThe standardization sample does not adequately represent the demographic diversity of the population to whom the test will be applied.Older versions of the MMPI were normed primarily on white Minnesotans, making cross-cultural comparisons problematic until the MMPI-2 and MMPI-3 expanded and restratified the normative sample.
Stereotype threatAwareness of negative stereotypes about one's group's performance on a test creates anxiety that depresses scores, independent of actual ability.Steele and Aronson (1995) demonstrated that African American students performed worse on verbal GRE items when the test was framed as diagnostic of intellectual ability compared to when it was described as a problem-solving exercise.
Response style differencesCultural norms influence tendencies toward acquiescence, extreme responding, social desirability, or modesty, systematically shifting scores.East Asian respondents tend to show greater midpoint responding on Likert scales, while Latin American respondents may show greater extreme responding, affecting personality and symptom inventories.

Worked Example — Evaluating an Assessment for Bias

Consider the following clinical scenario: A psychologist is asked to evaluate whether a newly developed anxiety screening measure is appropriate for use with a diverse urban population that includes substantial Latino, African American, and East Asian communities. The measure was developed and normed primarily with white, college-educated adults. The following worked example walks through the systematic evaluation process.

Evaluating the Cultural Fairness of an Anxiety Screener
1
Step 1 — Examine the Construct DefinitionBegin by examining whether the construct of anxiety as operationalized by the measure maps onto the ways anxiety is experienced and expressed across the target populations. For instance, the measure may emphasize cognitive symptoms (worry, rumination) that are more salient in Western, individualistic cultures, while underrepresenting somatic symptoms (headaches, gastrointestinal distress) that are more prominent in certain Asian and Latino populations. If the construct is defined too narrowly, construct bias is present regardless of how well individual items perform statistically.
Finding: The measure's item pool overrepresents cognitive-affective symptoms and underrepresents somatic and interpersonal manifestations of anxiety → possible construct bias.
2
Step 2 — Review the Normative SampleExamine the standardization sample to determine whether it is representative of the populations with whom the measure will be used. A norm sample that is 85% white and 90% college-educated does not provide an appropriate reference group for a diverse urban clinic. Scores compared to these norms may systematically mischaracterize the standing of minority group members, even if the items themselves are unbiased. Evaluate whether the test manual reports separate reliability and validity data for different demographic groups.
Finding: Normative sample does not match target population demographics → norm sample bias present. No group-specific psychometric data reported.
3
Step 3 — Conduct or Review DIF AnalysesExamine whether DIF analyses have been conducted during test development. If they have, review which items were flagged and how the developers responded. If no DIF analyses exist, this is a significant limitation. For the present example, imagine that DIF analysis reveals three items with large DIF: one item referencing 'feeling on edge at cocktail parties' (culturally specific social context), one item using the idiom 'butterflies in my stomach' (non-translatable metaphor), and one item about 'fear of losing control' that has different implications in collectivist versus individualist cultures.
Finding: Three items show significant DIF → item bias confirmed. These items should be revised or removed and the measure re-evaluated.
4
Step 4 — Evaluate Administration and Contextual FactorsConsider factors related to the testing context. Will examinees be tested in their primary language? Is the reading level appropriate? Are there provisions for oral administration for individuals with limited literacy? Is the testing environment culturally welcoming? Will stereotype threat be activated? For this scenario, the measure is available only in English and requires a 10th-grade reading level, which may pose barriers for some community members even if they are fluent in conversational English.
Finding: Language and literacy barriers introduce method bias. Recommendation: Develop translated and back-translated versions; provide oral administration options.
5
Step 5 — Formulate RecommendationsSynthesize findings across all levels of analysis. In this case, bias has been identified at the construct, item, norm, and method levels. The psychologist should recommend: (a) expanding the item pool to include culturally diverse symptom presentations; (b) conducting a new standardization study with a representative sample; (c) removing or revising DIF-flagged items; (d) developing translated versions using rigorous translation–back-translation and decentering procedures; and (e) providing clinicians with guidelines for culturally informed interpretation.
Overall conclusion: The measure, in its current form, is not appropriate for use with the target population without significant modification. Use of this measure without adaptation risks systematic misclassification and constitutes an ethical concern under APA standards.

Competing Models of Fairness in Testing

One of the most important complexities surrounding test bias is that fairness is not a unitary concept — multiple models of fairness exist, and they can yield contradictory conclusions about whether a given test is fair. The EPPP expects candidates to recognize these different models and understand why satisfying one model may preclude satisfying another. The four historically prominent models are summarized in the table below.

Four major models of test fairness and their trade-offs
Fairness ModelCore CriterionStrengthsLimitations
Cleary (Unbiased Prediction)A single regression equation predicts criterion scores equally well for all groups (no differential slopes or intercepts).Psychometrically rigorous; most widely accepted in research; testable with available data.Does not account for bias in the criterion; can perpetuate existing inequities if the criterion itself is contaminated.
Thorndike (Constant Ratio)The proportion of members from each group selected by the test should equal the proportion who would succeed on the criterion.Addresses outcome equity; ensures proportional representation of qualified individuals.Requires knowledge of base rates for success; may conflict with maximizing overall prediction accuracy.
Cole-Darlington (Equal Opportunity)Among those who would succeed on the criterion, the probability of being selected by the test should be equal across groups.Focuses on minimizing false negatives for qualified individuals from all groups.May increase false positive rates for some groups; difficult to implement without complete criterion data.
Culture-Fair / Culture-FreeTests should minimize cultural content entirely, relying on nonverbal, figural, or 'universal' tasks.Intuitively appealing; reduces obvious content bias.Widely considered unachievable — all tasks require some culturally mediated skills. 'Culture-reduced' is the preferred term. These tests often have lower validity for predicting real-world outcomes.
KEY TAKEAWAY
Think of test fairness like designing a fair race. One definition of fairness says the race is fair if everyone runs on the same track (Cleary — same prediction equation). Another says the race is fair if the proportion of winners from each team matches the proportion of fast runners on each team (Thorndike — proportional success). Yet another says fairness requires that every fast runner, regardless of team, has an equal chance of winning (Cole-Darlington — equal opportunity). These definitions can conflict: a track that is equally predictive of speed may still systematically disadvantage runners who trained on different surfaces. The key insight for clinicians is that fairness is a values-laden judgment, not purely a statistical property, and test users must choose which model of fairness is most appropriate for their specific context.

Ethical Standards and Modern Directions

The identification and mitigation of test bias is not merely a technical enterprise — it is deeply embedded in the ethical obligations of psychologists. The APA Ethical Principles and Code of Conduct, the Standards for Educational and Psychological Testing, and licensing laws all impose responsibilities on test developers and users regarding fairness. Modern approaches have moved beyond the question of whether a test is biased in a binary sense toward more nuanced understandings of how assessment practices interact with systems of privilege and oppression.

Traditional vs. modern approaches to assessment fairness
DomainTraditional ApproachModern/Emerging Approach
Definition of biasStatistical: differential prediction, DIF in items, factor structure differences.Expanded: includes consequential validity — examining downstream effects of test use on different groups, even when statistical bias is absent.
Unit of analysisGroup-level comparisons: racial, ethnic, gender groups.Intersectional: recognizes that bias may operate differently for Black women versus Black men, for instance, or for working-class immigrants versus affluent immigrants.
ResponsibilityPrimarily on test developers to create unbiased instruments.Shared responsibility: test users are equally responsible for selecting appropriate instruments, interpreting results within cultural context, and considering alternative assessment approaches.
Cultural frameworkCulture-free or culture-fair ideal: attempt to remove cultural content.Culture-responsive assessment: acknowledge that culture is inherent in all measurement; develop instruments that are sensitive to cultural diversity rather than trying to eliminate cultural context.
Diagnostic implicationsCompare individual scores to normative data with demographic corrections.Integrate quantitative data with qualitative cultural assessment, use Cultural Formulation Interview (DSM-5), consider idioms of distress and cultural concepts of disorder.
APA Ethics Code Connection
Standard 9.02 (Use of Assessments) states that psychologists must use assessment instruments 'whose validity and reliability have been established for use with members of the population tested.' Standard 9.06 requires that when interpreting results, psychologists 'take into account the purpose of the assessment as well as various test factors, test-taking abilities, and other characteristics of the person being assessed, such as situational, personal, linguistic, and cultural differences.' These are not aspirational guidelines — they are enforceable ethical standards with implications for licensure.

Looking forward, the field is increasingly integrating concepts from consequential validity (Messick, 1995), which examines the social consequences of test use as part of the validation process. This perspective holds that a test cannot be considered valid if its use produces systematically harmful outcomes for particular groups, even if the test meets traditional psychometric criteria for absence of bias. The Cultural Formulation Interview in DSM-5 represents one clinical tool for supplementing standardized assessment with systematic exploration of cultural factors that may influence presentation, help-seeking, and diagnostic accuracy.

Practice Problems

PROBLEM 1CONCEPTUAL
A psychologist notices that African American clients consistently score higher on a measure of paranoia than white clients. A colleague argues this proves the test is biased. How would you evaluate this claim using the psychometric definition of test bias?
PROBLEM 2BASIC APPLICATION
A DIF analysis of a 50-item depression inventory reveals that 5 items have statistically significant and large DIF when comparing English-speaking and Spanish-speaking examinees matched on total depression scores. What does this finding indicate, and what is the most appropriate course of action?
PROBLEM 3INTERMEDIATE
A researcher conducts a measurement invariance analysis of a widely used PTSD measure across four ethnic groups. The model achieves configural and metric invariance but fails at the scalar invariance level. What are the implications for clinical practice?
PROBLEM 4APPLIED
You are a psychologist working in a forensic setting where a Hmong-speaking individual has been evaluated for intellectual disability using a standard IQ test administered through an interpreter. The individual obtained a Full Scale IQ of 62. Defense counsel argues the score reflects cultural and linguistic bias rather than genuine intellectual disability. Identify at least four specific sources of potential bias in this scenario and describe what steps should be taken.
PROBLEM 5CRITICAL THINKING
Some scholars (e.g., Helms, 2006) have argued that the concept of test bias as traditionally defined in psychometrics is itself culturally biased because it accepts the majority culture's test performance as the standard against which minority performance is evaluated. Critically evaluate this argument, considering both its strengths and its limitations. How might this perspective change the way we approach assessment development and validation?

Summary — Test Bias and Fairness in Assessment

Test bias refers to systematic error in measurement that differentially affects identifiable subgroups and is fundamentally a threat to construct validity. It is distinct from mere group differences in scores and must be evaluated at multiple levels: construct bias (whether the test measures the same construct across groups), method bias (whether administration procedures introduce systematic error), and item bias (whether specific items function differently across groups via differential item functioning analysis). Key statistical tools include the Mantel-Haenszel procedure for DIF, differential prediction analysis using the Cleary model, and measurement invariance testing through confirmatory factor analysis (configural → metric → scalar → strict).

Multiple competing models of fairness exist — Cleary (equal prediction), Thorndike (constant ratio), and Cole-Darlington (equal opportunity) — and satisfying one may preclude satisfying another. Modern approaches emphasize consequential validity, shared responsibility between test developers and users, intersectional analysis, and culture-responsive assessment rather than the unattainable ideal of culture-free testing. Clinicians are ethically obligated under APA Standards 9.02 and 9.06 to select instruments validated for their specific populations, interpret results within cultural context, and integrate standardized measures with tools like the Cultural Formulation Interview to ensure equitable and accurate assessment.

Varsity Tutors • EPPP: Part 1, Knowledge • Test Bias — Identify sources of cultural bias and fairness issues in assessment instruments