Historical Context & Motivation
The formal study of psychological measurement has its roots in the late nineteenth and early twentieth centuries, when pioneers of mental testing recognized that a test is only useful if it produces consistent scores and actually measures the construct it claims to measure. Early intelligence tests developed by Alfred Binet and later adapted by Lewis Terman brought these concerns to the forefront: clinicians needed assurance that scores were not merely artifacts of random fluctuation, and that the tests captured genuine cognitive ability rather than some confounding variable. This dual concern—consistency and accuracy—gave rise to the twin pillars of psychometric theory: reliability and validity.
For clinicians preparing for the EPPP, these developments are not merely historical footnotes. Every time you administer a standardized assessment—whether the MMPI-3, the WAIS-IV, or the Beck Depression Inventory—you must evaluate whether the instrument produces stable, replicable results (reliability) and whether those results actually reflect the psychological construct you intend to assess (validity). The central question this lesson addresses is: How do you differentiate among the major types of reliability and forms of validity evidence, and what does each tell you about the quality of a psychological test?
Core Principles & Definitions
Before distinguishing among specific reliability types and validity forms, it is essential to grasp the foundational distinction between the two constructs. Reliability refers to the degree to which test scores are free from measurement error—it is about consistency, stability, and reproducibility. Validity refers to the degree to which accumulated evidence and theory support specific interpretations of test scores for given purposes—it is about meaning, accuracy, and justification. A critical axiom in psychometrics is that reliability is necessary but not sufficient for validity: a test can produce perfectly consistent scores that are consistently wrong in what they claim to measure.
Classical Test Theory (CTT)
Sources of Measurement Error
Validity as Evidential Argument
Reliability Sets the Ceiling for Validity
Standard Error of Measurement (SEM)
Visual Explanation — Reliability Types at a Glance
The diagram above organizes the reliability landscape according to the specific source of measurement error each type addresses. Test-retest reliability examines whether scores remain stable across time, making temporal instability the targeted error source. Alternate-forms (parallel-forms) reliability assesses whether two versions of a test produce comparable scores, isolating content sampling error. Internal consistency reliability evaluates whether items within a single test cohere, and inter-rater reliability determines agreement across different scorers. For EPPP preparation, it is essential to match each reliability type with its error source, statistical index, and appropriate use case.
Mathematical Framework
Classical test theory provides the mathematical backbone for understanding reliability. The equations below formalize the relationship between observed scores, true scores, error, and derived reliability metrics. Mastery of these formulas is essential for interpreting test manuals and evaluating assessment instruments in clinical practice.
Detailed Breakdown — Forms of Validity Evidence
While the EPPP frequently references the traditional triad of content, criterion-related, and construct validity, the current Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) conceptualizes validity as a unitary construct supported by five categories of evidence. Understanding both frameworks is essential for the exam. The table and diagram below cross-reference the traditional labels with the modern evidence categories.
| Traditional Label | Modern Evidence Category | Key Methods | Clinical Example |
|---|---|---|---|
| Content Validity | Evidence based on test content | Expert panel review, content blueprints, item-objective alignment | EPPP items are mapped to a content domain blueprint to ensure adequate coverage of all knowledge areas. |
| Criterion-Related (Predictive) | Relations with other variables | Correlation with future criterion measured later in time | GRE scores correlate with first-year graduate GPA (criterion measured months/years later). |
| Criterion-Related (Concurrent) | Relations with other variables | Correlation with criterion measured at the same time | A new depression scale correlates .85 with the BDI-II administered simultaneously. |
| Construct Validity | All five evidence categories | Convergent/discriminant evidence, factor analysis, known-groups comparisons, MTMM matrix | An anxiety measure correlates highly with other anxiety measures (convergent) and weakly with unrelated constructs (discriminant). |
| Face Validity | Not a formal evidence category | Subjective judgment: does the test 'look like' it measures what it claims? | A depression questionnaire asks about mood and sleep. Face validity affects examinee buy-in but is not formal psychometric evidence. |
Worked Example — Evaluating a Depression Screener
Imagine you are a psychologist selecting a new depression screening instrument, the Hypothetical Depression Scale (HDS), for use in a community mental health center. The test manual reports the following psychometric data: the HDS has 20 items scored on a 4-point Likert scale, coefficient alpha of 0.88, test-retest reliability (2-week interval) of 0.82, a correlation of 0.79 with the BDI-II (concurrent criterion), and a standard deviation of 10. Let us walk through how to interpret these data.
Strengths, Limitations, and Common Confusions
| Reliability Type | Strengths | Limitations / Pitfalls |
|---|---|---|
| Test-Retest | Directly assesses temporal stability; easy to understand and compute. | Practice effects inflate estimates with short intervals; genuine change in the construct (e.g., mood states) deflates estimates with long intervals. Inappropriate for unstable constructs. |
| Alternate Forms | Controls for practice effects and memory; addresses content sampling error. | Creating truly parallel forms is difficult and expensive. If administered at different times, confounds temporal instability with content sampling. |
| Internal Consistency | Requires only one administration; efficient and widely reported. | Inappropriate for speeded tests (inflated estimates). Overly heterogeneous or multidimensional tests may show misleadingly low alpha. Does not address temporal stability. |
| Inter-Rater | Essential for subjectively scored assessments; Cohen's kappa corrects for chance agreement. | Percentage agreement alone is misleading because it ignores chance. Kappa can be paradoxically low when base rates are extreme. Requires training raters, which adds cost. |
| Split-Half | Single administration; Spearman-Brown correction adjusts for test length reduction. | Result depends on how the test is split; only one of many possible splits is evaluated. Superseded by Cronbach's alpha, which averages all possible split-halves. |
Connection to Advanced Psychometric Theory
Classical test theory (CTT), while foundational, has notable limitations that advanced psychometric frameworks seek to address. The EPPP does not require deep expertise in these advanced methods, but awareness of their existence and relationship to CTT concepts of reliability and validity is increasingly important. Two advanced frameworks deserve mention: Generalizability Theory (G-Theory) and Item Response Theory (IRT).
| Feature | Classical Test Theory (CTT) | Generalizability Theory (G-Theory) | Item Response Theory (IRT) |
|---|---|---|---|
| Error model | Single undifferentiated error term (E) | Multiple error facets (raters, items, occasions) simultaneously estimated via ANOVA | Item-level error; reliability varies by trait level (information function) |
| Reliability index | Single coefficient (α, r) | Generalizability coefficient (Eρ²) and dependability coefficient (Φ) | Test information function I(θ) — precision varies across the trait continuum |
| Sample dependency | Statistics are sample-dependent; reliability varies across populations | Also sample-dependent but can model multiple error facets in a single study | Item parameters are sample-independent (invariance); enables adaptive testing |
| Clinical application | Most test manuals report CTT indices; used for standard clinical instruments | Used when multiple error sources matter (e.g., OSCE medical exams with raters × stations) | Computer adaptive testing (e.g., CAT-MH); items tailored to examinee's ability level |
For the EPPP, the most critical advanced concept is understanding that Generalizability Theory extends CTT by partitioning error variance into multiple facets simultaneously, rather than treating all error as a monolithic term. In clinical settings, this is increasingly relevant when behavioral health assessments involve multiple raters observing multiple clients across multiple occasions—a scenario common in behavioral observation systems and competency-based clinical evaluations. IRT, meanwhile, offers the concept of measurement precision that varies along the trait continuum, unlike CTT's single reliability coefficient. This is especially valuable for instruments designed to discriminate at clinical cutoff points, where precision at the decision threshold matters most.
Practice Problems
Summary — Reliability and Validity in Psychological Assessment
Reliability refers to the consistency and stability of test scores, rooted in classical test theory's decomposition of observed scores into true score plus error (X = T + E). Four major types address distinct error sources: test-retest (temporal instability), alternate forms (content sampling), internal consistency (item heterogeneity, indexed by Cronbach's alpha, KR-20, or split-half methods), and inter-rater reliability (rater subjectivity, indexed by Cohen's kappa or ICC). The Standard Error of Measurement (SEM) translates the reliability coefficient into a clinically meaningful confidence interval around an individual's observed score.
Validity is the degree to which evidence supports the intended interpretation of test scores. The modern framework identifies five evidence sources: test content, response processes, internal structure, relations with other variables (encompassing convergent, discriminant, predictive, and concurrent evidence), and consequences of testing. The cardinal rule is that reliability is necessary but not sufficient for validity: a test cannot accurately measure a construct if its scores are inconsistent, but perfectly consistent scores can still miss the intended construct entirely. When selecting or evaluating assessment instruments, clinicians must examine both reliability and validity evidence in light of the specific purpose, population, and stakes of the assessment.