EPPP: PART 1, KNOWLEDGE • DOMAIN 5: ASSESSMENT AND DIAGNOSIS

Test Construction — Apply item analysis, standardization, norming, sensitivity, and specificity concepts

Understanding how psychological tests are built, refined, and evaluated for diagnostic accuracy in clinical practice.

Historical Context & Motivation

The history of psychological testing is inseparable from the broader scientific demand for measurement rigor. Early attempts at mental measurement were informal and unstandardized, producing results that varied wildly depending on who administered the test, where it was given, and how responses were scored. As psychology matured into an empirical discipline, researchers recognized that a test is only as useful as the evidence supporting its construction—its items must discriminate meaningfully between individuals, its administration must be uniform, and its scores must be interpretable against a well-defined reference group. These requirements gave rise to the interconnected domains of item analysis, standardization, norming, and the diagnostic metrics of sensitivity and specificity.

1905
Binet-Simon Scale
Alfred Binet and Théodore Simon developed the first practical intelligence test in Paris, introducing the concept of age-graded norms and demonstrating that test items could be empirically selected based on how well they differentiated children of varying ability levels.
1917
Army Alpha & Beta Tests
The U.S. military's mass testing of recruits during World War I established standardized administration procedures on an unprecedented scale, proving that large-group testing demanded strict uniformity in instructions, timing, and scoring.
1950s
Classical Test Theory Matures
Researchers such as Gulliksen formalized classical test theory (CTT), providing mathematical frameworks for item difficulty, item discrimination, and reliability estimation that remain foundational to modern test construction.
1966
Standards for Psychological Tests
The APA published the first edition of the Standards for Educational and Psychological Testing, codifying expectations for norming, standardization, and evidence of validity and reliability in published instruments.
1990s–Present
Diagnostic Accuracy Metrics
Borrowed from epidemiology, concepts of sensitivity and specificity became integral to evaluating psychological screening instruments, allowing clinicians to quantify the accuracy of diagnostic cutoff scores.

These historical milestones converge on a central question that every test developer and clinician must answer: How do we ensure that a psychological test produces scores that are meaningful, consistent, and diagnostically useful? The concepts covered in this lesson provide the technical foundation for answering that question at every stage of test development and clinical application.

Core Principles & Definitions

Test construction rests on a set of interrelated principles that guide the developer from initial item writing through final norm-referenced score interpretation. Understanding each principle individually—and seeing how they interact—is essential for the EPPP and for competent clinical practice. The five core constructs addressed here form a coherent pipeline: items are written and then refined through item analysis; the test is administered under uniform conditions via standardization; scores are made interpretable through norming; and the test's diagnostic accuracy is quantified using sensitivity and specificity.

1

Item Analysis

A statistical procedure applied to individual test items to evaluate their difficulty (proportion of examinees who answer correctly) and discrimination (ability to differentiate high-performing from low-performing examinees). Items that are too easy, too hard, or non-discriminating are revised or removed.
2

Standardization

The process of establishing uniform procedures for test administration and scoring. Every examinee receives the same instructions, time limits, and scoring rules, ensuring that observed score differences reflect true differences in the construct being measured rather than procedural variability.
3

Norming

Administering the finalized test to a large, representative normative sample to create reference distributions. Individual raw scores are then converted to derived scores (e.g., percentiles, z-scores, T-scores, standard scores) that communicate standing relative to the norm group.
4

Sensitivity

The probability that a test correctly identifies individuals who have the condition of interest (true positive rate). A highly sensitive test produces few false negatives, making it valuable as a screening instrument when missing a case carries serious consequences.
5

Specificity

The probability that a test correctly identifies individuals who do not have the condition (true negative rate). A highly specific test produces few false positives, which is critical when a positive result triggers costly, invasive, or stigmatizing follow-up interventions.
KEY TAKEAWAY
Think of test construction like building a medical thermometer. Item analysis ensures each degree marking actually reflects a real temperature change (not random noise). Standardization means every thermometer is read the same way—under the tongue, for the same duration. Norming tells you what counts as a 'fever' by comparing the reading to a large reference group. And sensitivity/specificity tell you how often the thermometer catches real fevers versus how often it falsely alarms on healthy people.

Visual Explanation — The Test Construction Pipeline

The left panel illustrates the item difficulty spectrum from very hard (p = .10) to very easy (p = .95), with the optimal zone near p = .50 for maximum discrimination. The right panel shows the classic sensitivity–specificity trade-off: as the cutoff score moves in one direction, sensitivity rises while specificity falls, and vice versa. The green dot marks the optimal balance point.

The pipeline diagram at the top illustrates that test construction is a sequential, iterative process. Developers begin by writing a pool of items designed to tap the target construct, then subject those items to empirical scrutiny via item analysis. Items that survive this vetting are assembled into the final form, which is administered under standardized conditions to a large, demographically representative sample for norming. When the test is used for clinical screening or diagnosis, its accuracy is evaluated through sensitivity and specificity metrics that quantify the rate of correct identifications versus errors.

Mathematical Framework

Item Analysis Formulas

ITEM DIFFICULTY INDEX
p = (Number of examinees answering correctly) / (Total number of examinees)
p ranges from 0 to 1.0. A higher p value means the item is easier. For maximum discrimination on a four-option multiple-choice test, the optimal difficulty is approximately p = .625 (midpoint between chance level of .25 and 1.0). For tests without guessing corrections, p ≈ .50 is ideal.
ITEM DISCRIMINATION INDEX
D = p(upper) − p(lower)
p(upper) = proportion correct in the top 27% of total test scorers; p(lower) = proportion correct in the bottom 27%. D ranges from −1.0 to +1.0. Values ≥ .30 are generally considered acceptable; values ≤ .19 suggest the item should be revised or discarded. A negative D indicates the item is performing paradoxically—low scorers outperform high scorers on that item.
POINT-BISERIAL CORRELATION
r_pb = (M₊ − M_total) / SD_total × √(p × q)
rpb correlates item score (0 or 1) with total test score. M₊ = mean total score for those who answered correctly; q = 1 − p. Higher values indicate better discrimination. This is the most commonly used item–total correlation in classical test theory.

Sensitivity & Specificity Formulas

SENSITIVITY (TRUE POSITIVE RATE)
Sensitivity = TP / (TP + FN)
TP = true positives (correctly identified as having the condition); FN = false negatives (have the condition but were missed). Sensitivity answers: 'Of all people who truly have the disorder, what proportion does the test correctly detect?'
SPECIFICITY (TRUE NEGATIVE RATE)
Specificity = TN / (TN + FP)
TN = true negatives (correctly identified as not having the condition); FP = false positives (do not have the condition but were incorrectly flagged). Specificity answers: 'Of all people who truly do not have the disorder, what proportion does the test correctly rule out?'
Clinical Decision Rule
When the cost of missing a true case is high (e.g., suicidality screening), clinicians prioritize high sensitivity even at the expense of specificity—accepting more false positives to ensure no true case goes undetected. Conversely, when a positive result leads to invasive or stigmatizing intervention, high specificity is essential to minimize unnecessary harm from false alarms.

Detailed Breakdown — The 2 × 2 Diagnostic Table and Norm-Referenced Scores

The 2 × 2 Diagnostic Decision Table

The 2 × 2 matrix is the foundation for computing all diagnostic accuracy metrics. Sensitivity reads across the top row (TP + FN), specificity reads across the bottom row (TN + FP). Positive predictive value (PPV) reads down the left column, and negative predictive value (NPV) reads down the right column.

Common Norm-Referenced Score Transformations

Standard derived scores and their scale properties
Score TypeMeanSDTypical Use
z-score01Research; baseline for other conversions
T-score5010MMPI-2, MMPI-3, personality inventories
Standard Score (IQ-type)10015WAIS-IV, WISC-V, Stanford-Binet
Scaled Score103WAIS-IV/WISC-V subtests
Stanine5≈ 2Educational achievement; 9-point scale
Percentile Rank50thN/A (ordinal)Communicating results to clients/parents

All derived scores depend on the quality of the normative sample. A norm group must be large enough to produce stable estimates and representative of the population for whom the test is intended. When the norm group is outdated or demographically skewed, derived scores can be misleading—a phenomenon that prompted the periodic re-norming of major instruments such as the Wechsler scales. The conversion formula from a raw score to a z-score is z = (X − M) / SD, from which all other linear transformations are derived.

Worked Example — From Item Analysis to Diagnostic Accuracy

Evaluating a Depression Screening Instrument
1
Step 1 — Compute Item DifficultyA new 20-item depression screener is piloted with 200 participants. On Item 7, 140 out of 200 examinees endorse the item (answer in the keyed direction). The item difficulty index is: p = 140 / 200 = 0.70. This item is moderately easy—most examinees endorse it. While a difficulty of .50 provides maximum discrimination, items slightly above or below .50 remain acceptable. An item at p = .70 may still be retained if it shows adequate discrimination.
p = 0.70 (moderately easy)
2
Step 2 — Compute Item DiscriminationWe rank all 200 examinees by total test score and select the upper 27% (n = 54) and lower 27% (n = 54). In the upper group, 50 out of 54 endorsed Item 7 (p(upper) = 0.926). In the lower group, 20 out of 54 endorsed it (p(lower) = 0.370). The discrimination index is: D = 0.926 − 0.370 = 0.556. This exceeds the .30 threshold, indicating that Item 7 effectively distinguishes high from low scorers.
D = 0.556 (excellent discrimination)
3
Step 3 — Establish a Cutoff Score and Classify OutcomesAfter assembling the final 15 items (removing 5 poor performers through item analysis), the developers set a cutoff of ≥ 10 to flag likely depression. Against a structured clinical interview (the gold standard), the following results emerge from a validation sample of 500: TP = 90, FP = 30, FN = 10, TN = 370.
Total with depression: 100; Total without: 400
4
Step 4 — Calculate SensitivitySensitivity = TP / (TP + FN) = 90 / (90 + 10) = 90 / 100 = 0.90 (or 90%). This means the screener correctly identifies 90% of individuals who truly have depression. Only 10% of depressed individuals are missed (false negatives).
Sensitivity = 0.90 (90%)
5
Step 5 — Calculate SpecificitySpecificity = TN / (TN + FP) = 370 / (370 + 30) = 370 / 400 = 0.925 (or 92.5%). This means the screener correctly rules out 92.5% of non-depressed individuals. Only 7.5% are false positives—flagged as potentially depressed when they are not.
Specificity = 0.925 (92.5%)
6
Step 6 — Compute Positive Predictive Value (PPV)PPV = TP / (TP + FP) = 90 / (90 + 30) = 90 / 120 = 0.75 (75%). Among those the test flags as positive, 75% truly have depression. Note that PPV depends heavily on base rate (prevalence): in a population with lower prevalence, the PPV would be substantially lower even with the same sensitivity and specificity.
PPV = 0.75 (75%)
Base Rate Matters
A common EPPP tested concept: Positive predictive value (PPV) is strongly affected by the base rate of the condition in the tested population. A screener with 90% sensitivity and 92.5% specificity performs very differently in a community sample (low prevalence) than in an inpatient psychiatric unit (high prevalence). In low base-rate conditions, even highly specific tests produce many false positives relative to true positives.

Strengths, Limitations, and Comparisons

Key strengths and limitations of test construction metrics
ConceptStrengthsLimitations
Item Difficulty (p)Simple to compute; intuitive interpretation; directly informs item selection for test assemblySample-dependent—difficulty changes with the ability level of the sample; does not capture item quality beyond proportion correct
Item Discrimination (D)Identifies items that differentiate high and low performers; straightforward upper-lower comparisonUses only extreme groups (27%), discarding middle scorers; can be unstable in small samples
Point-Biserial (r_pb)Uses all examinees; correlational metric integrates difficulty and discriminationAssumes linearity; can be attenuated when item difficulty is extreme (very high or very low p)
SensitivityEnsures that true cases are detected; vital for screening where missing a diagnosis is costlyMaximizing sensitivity typically increases false positive rate; cannot be evaluated independently of specificity
SpecificityReduces false alarms and unnecessary follow-up; protects clients from unwarranted labelsMaximizing specificity may miss true cases (increase false negatives); trade-off with sensitivity is unavoidable
KEY TAKEAWAY
No single metric tells the full story. Think of these concepts as a panel of indicators—like vital signs in medicine. A clinician does not diagnose based on heart rate alone; similarly, a test developer never evaluates an item or a test based on a single statistic. Item difficulty without discrimination tells you how hard the item is but not whether it measures the construct meaningfully. Sensitivity without specificity tells you about detection but not about the cost of false alarms. A comprehensive evaluation always considers the full constellation of metrics.

Connection to Advanced Theory — IRT and ROC Analysis

The metrics discussed in this lesson belong primarily to classical test theory (CTT), which remains the dominant framework in applied psychology. However, modern psychometrics increasingly draws on item response theory (IRT) and receiver operating characteristic (ROC) analysis to overcome limitations of the classical approach. Understanding these advanced frameworks contextualizes the foundational metrics and prepares you for more nuanced EPPP questions.

CTT vs. IRT comparison
FeatureClassical Test Theory (CTT)Item Response Theory (IRT)
Item ParametersSample-dependent (p, D change with sample)Sample-independent (item difficulty and discrimination estimated as invariant parameters)
Ability EstimatesTest-dependent (raw or derived scores from a specific form)Item-independent (theta, θ, estimated from any set of calibrated items)
Key GraphicItem difficulty/discrimination tablesItem characteristic curves (ICC) showing P(correct) as a function of θ
ApplicationMost published psychological tests; simple to computeComputerized adaptive testing (CAT); large-scale licensure exams

Similarly, the sensitivity–specificity framework extends into ROC curve analysis, which plots sensitivity (y-axis) against 1 − specificity (x-axis) across all possible cutoff scores. The area under the ROC curve (AUC) provides a single index of a test's overall diagnostic accuracy, with values of .90–1.0 considered excellent, .80–.89 good, and .70–.79 fair. ROC analysis allows clinicians to select the cutoff that best balances detection and false alarm rates for a given clinical context, rather than relying on a single pre-set threshold.

📝 EPPP Tip
The EPPP may ask you to compare CTT and IRT conceptually. Remember: IRT's chief advantage is that its item parameters are sample-invariant and its ability estimates are test-invariant—meaning items and persons are measured on the same latent scale, enabling adaptive testing. CTT metrics like p and D, while simpler, are always tied to the specific sample used.

Practice Problems

PROBLEM 1CONCEPTUAL
A test developer finds that an item has a difficulty index (p) of .95 and a discrimination index (D) of .05. Explain why these two values are related and what action the developer should take regarding this item.
PROBLEM 2BASIC CALCULATION
A screening test for PTSD is administered to 400 individuals. Using a structured clinical interview as the gold standard, the following results are obtained: TP = 60, FP = 20, FN = 15, TN = 305. Calculate the sensitivity and specificity of the screener.
PROBLEM 3INTERMEDIATE
An intelligence test reports standard scores with M = 100 and SD = 15. A client obtains a raw score that converts to a z-score of −1.33. (a) What is the client's standard score? (b) What percentile rank does this approximate? (c) Would this score fall within the 'normal' or 'below average' classification range on most Wechsler-based systems?
PROBLEM 4APPLIED
You are a psychologist developing a brief anxiety screening tool for a primary care setting where the base rate of clinical anxiety is 8%. Your screener has a sensitivity of .85 and a specificity of .90. In a sample of 1,000 patients, calculate the expected number of TP, FP, FN, and TN, then compute the positive predictive value (PPV). Discuss the clinical implications of this PPV.
PROBLEM 5CRITICAL THINKING
A colleague argues that the best way to improve a test is to keep only items with the highest discrimination indices and eliminate all items with p values above .80. Critically evaluate this strategy, considering its impact on content validity, test length, reliability, and the potential for differential item functioning (DIF).

Lesson Summary

Test construction is a systematic, empirically driven process. Item analysis evaluates each item's difficulty (p) and discrimination (D or r_pb), ensuring only psychometrically sound items are retained. Standardization establishes uniform administration and scoring procedures so that score differences reflect true construct variance, not procedural artifacts. Norming transforms raw scores into interpretable derived scores (z-scores, T-scores, standard scores, percentiles) by referencing a large, representative sample.

When tests are used for diagnostic classification, sensitivity (the true positive rate) and specificity (the true negative rate) quantify the accuracy of cutoff scores. These metrics exist in an inherent trade-off: raising one typically lowers the other. Positive and negative predictive values (PPV and NPV) are further modulated by base rate, making prevalence a critical consideration in clinical screening contexts. Together, these concepts form the psychometric backbone that clinicians and researchers rely on to build, evaluate, and responsibly use psychological assessments.

Varsity Tutors • EPPP: Part 1, Knowledge • Test Construction — Apply item analysis, standardization, norming, sensitivity, and specificity concepts