Historical Context & Motivation
The history of psychological testing is inseparable from the broader scientific demand for measurement rigor. Early attempts at mental measurement were informal and unstandardized, producing results that varied wildly depending on who administered the test, where it was given, and how responses were scored. As psychology matured into an empirical discipline, researchers recognized that a test is only as useful as the evidence supporting its construction—its items must discriminate meaningfully between individuals, its administration must be uniform, and its scores must be interpretable against a well-defined reference group. These requirements gave rise to the interconnected domains of item analysis, standardization, norming, and the diagnostic metrics of sensitivity and specificity.
These historical milestones converge on a central question that every test developer and clinician must answer: How do we ensure that a psychological test produces scores that are meaningful, consistent, and diagnostically useful? The concepts covered in this lesson provide the technical foundation for answering that question at every stage of test development and clinical application.
Core Principles & Definitions
Test construction rests on a set of interrelated principles that guide the developer from initial item writing through final norm-referenced score interpretation. Understanding each principle individually—and seeing how they interact—is essential for the EPPP and for competent clinical practice. The five core constructs addressed here form a coherent pipeline: items are written and then refined through item analysis; the test is administered under uniform conditions via standardization; scores are made interpretable through norming; and the test's diagnostic accuracy is quantified using sensitivity and specificity.
Item Analysis
Standardization
Norming
Sensitivity
Specificity
Visual Explanation — The Test Construction Pipeline
The pipeline diagram at the top illustrates that test construction is a sequential, iterative process. Developers begin by writing a pool of items designed to tap the target construct, then subject those items to empirical scrutiny via item analysis. Items that survive this vetting are assembled into the final form, which is administered under standardized conditions to a large, demographically representative sample for norming. When the test is used for clinical screening or diagnosis, its accuracy is evaluated through sensitivity and specificity metrics that quantify the rate of correct identifications versus errors.
Mathematical Framework
Item Analysis Formulas
Sensitivity & Specificity Formulas
Detailed Breakdown — The 2 × 2 Diagnostic Table and Norm-Referenced Scores
The 2 × 2 Diagnostic Decision Table
Common Norm-Referenced Score Transformations
| Score Type | Mean | SD | Typical Use |
|---|---|---|---|
| z-score | 0 | 1 | Research; baseline for other conversions |
| T-score | 50 | 10 | MMPI-2, MMPI-3, personality inventories |
| Standard Score (IQ-type) | 100 | 15 | WAIS-IV, WISC-V, Stanford-Binet |
| Scaled Score | 10 | 3 | WAIS-IV/WISC-V subtests |
| Stanine | 5 | ≈ 2 | Educational achievement; 9-point scale |
| Percentile Rank | 50th | N/A (ordinal) | Communicating results to clients/parents |
All derived scores depend on the quality of the normative sample. A norm group must be large enough to produce stable estimates and representative of the population for whom the test is intended. When the norm group is outdated or demographically skewed, derived scores can be misleading—a phenomenon that prompted the periodic re-norming of major instruments such as the Wechsler scales. The conversion formula from a raw score to a z-score is z = (X − M) / SD, from which all other linear transformations are derived.
Worked Example — From Item Analysis to Diagnostic Accuracy
Strengths, Limitations, and Comparisons
| Concept | Strengths | Limitations |
|---|---|---|
| Item Difficulty (p) | Simple to compute; intuitive interpretation; directly informs item selection for test assembly | Sample-dependent—difficulty changes with the ability level of the sample; does not capture item quality beyond proportion correct |
| Item Discrimination (D) | Identifies items that differentiate high and low performers; straightforward upper-lower comparison | Uses only extreme groups (27%), discarding middle scorers; can be unstable in small samples |
| Point-Biserial (r_pb) | Uses all examinees; correlational metric integrates difficulty and discrimination | Assumes linearity; can be attenuated when item difficulty is extreme (very high or very low p) |
| Sensitivity | Ensures that true cases are detected; vital for screening where missing a diagnosis is costly | Maximizing sensitivity typically increases false positive rate; cannot be evaluated independently of specificity |
| Specificity | Reduces false alarms and unnecessary follow-up; protects clients from unwarranted labels | Maximizing specificity may miss true cases (increase false negatives); trade-off with sensitivity is unavoidable |
Connection to Advanced Theory — IRT and ROC Analysis
The metrics discussed in this lesson belong primarily to classical test theory (CTT), which remains the dominant framework in applied psychology. However, modern psychometrics increasingly draws on item response theory (IRT) and receiver operating characteristic (ROC) analysis to overcome limitations of the classical approach. Understanding these advanced frameworks contextualizes the foundational metrics and prepares you for more nuanced EPPP questions.
| Feature | Classical Test Theory (CTT) | Item Response Theory (IRT) |
|---|---|---|
| Item Parameters | Sample-dependent (p, D change with sample) | Sample-independent (item difficulty and discrimination estimated as invariant parameters) |
| Ability Estimates | Test-dependent (raw or derived scores from a specific form) | Item-independent (theta, θ, estimated from any set of calibrated items) |
| Key Graphic | Item difficulty/discrimination tables | Item characteristic curves (ICC) showing P(correct) as a function of θ |
| Application | Most published psychological tests; simple to compute | Computerized adaptive testing (CAT); large-scale licensure exams |
Similarly, the sensitivity–specificity framework extends into ROC curve analysis, which plots sensitivity (y-axis) against 1 − specificity (x-axis) across all possible cutoff scores. The area under the ROC curve (AUC) provides a single index of a test's overall diagnostic accuracy, with values of .90–1.0 considered excellent, .80–.89 good, and .70–.79 fair. ROC analysis allows clinicians to select the cutoff that best balances detection and false alarm rates for a given clinical context, rather than relying on a single pre-set threshold.
Practice Problems
Lesson Summary
Test construction is a systematic, empirically driven process. Item analysis evaluates each item's difficulty (p) and discrimination (D or r_pb), ensuring only psychometrically sound items are retained. Standardization establishes uniform administration and scoring procedures so that score differences reflect true construct variance, not procedural artifacts. Norming transforms raw scores into interpretable derived scores (z-scores, T-scores, standard scores, percentiles) by referencing a large, representative sample.
When tests are used for diagnostic classification, sensitivity (the true positive rate) and specificity (the true negative rate) quantify the accuracy of cutoff scores. These metrics exist in an inherent trade-off: raising one typically lowers the other. Positive and negative predictive values (PPV and NPV) are further modulated by base rate, making prevalence a critical consideration in clinical screening contexts. Together, these concepts form the psychometric backbone that clinicians and researchers rely on to build, evaluate, and responsibly use psychological assessments.