EPPP: PART 1, KNOWLEDGE • DOMAIN 5: ASSESSMENT AND DIAGNOSIS

Reliability And Validity — Differentiate reliability types and validity forms in test interpretation

Understanding how consistency and accuracy in psychological measurement underpin sound clinical assessment and diagnosis.

Historical Context & Motivation

The formal study of psychological measurement has its roots in the late nineteenth and early twentieth centuries, when pioneers of mental testing recognized that a test is only useful if it produces consistent scores and actually measures the construct it claims to measure. Early intelligence tests developed by Alfred Binet and later adapted by Lewis Terman brought these concerns to the forefront: clinicians needed assurance that scores were not merely artifacts of random fluctuation, and that the tests captured genuine cognitive ability rather than some confounding variable. This dual concern—consistency and accuracy—gave rise to the twin pillars of psychometric theory: reliability and validity.

1904
Spearman's Reliability Theory
Charles Spearman introduced classical test theory (CTT) and the concept of the reliability coefficient, establishing a mathematical framework for separating true score variance from error variance.
1937
Kuder-Richardson Formulas
Kuder and Richardson published their KR-20 and KR-21 formulas for estimating internal consistency reliability of dichotomously scored tests, laying groundwork for later generalized methods.
1951
Cronbach's Alpha
Lee Cronbach generalized the Kuder-Richardson formula to coefficient alpha, applicable to tests with items scored on continuous or polytomous scales, making it the most widely reported reliability index in behavioral health research.
1954–1955
APA Standards & the Validity Trifecta
The APA Technical Recommendations introduced the classic triad of content, criterion-related, and construct validity, formalizing how test developers must justify what their instruments measure.
1999–2014
Unified Validity Framework
Samuel Messick and subsequent Standards for Educational and Psychological Testing reconceptualized validity as a unitary construct—all evidence supports the overall validity of score interpretations, rather than separate 'types' of validity.

For clinicians preparing for the EPPP, these developments are not merely historical footnotes. Every time you administer a standardized assessment—whether the MMPI-3, the WAIS-IV, or the Beck Depression Inventory—you must evaluate whether the instrument produces stable, replicable results (reliability) and whether those results actually reflect the psychological construct you intend to assess (validity). The central question this lesson addresses is: How do you differentiate among the major types of reliability and forms of validity evidence, and what does each tell you about the quality of a psychological test?

Core Principles & Definitions

Before distinguishing among specific reliability types and validity forms, it is essential to grasp the foundational distinction between the two constructs. Reliability refers to the degree to which test scores are free from measurement error—it is about consistency, stability, and reproducibility. Validity refers to the degree to which accumulated evidence and theory support specific interpretations of test scores for given purposes—it is about meaning, accuracy, and justification. A critical axiom in psychometrics is that reliability is necessary but not sufficient for validity: a test can produce perfectly consistent scores that are consistently wrong in what they claim to measure.

1

Classical Test Theory (CTT)

Every observed score (X) is composed of a true score (T) plus error (E). Reliability is the proportion of observed score variance attributable to true score variance. The higher the reliability coefficient (ranging from 0 to 1), the less error contaminates the measurement.
2

Sources of Measurement Error

Different reliability types address different error sources: temporal instability (test-retest), item sampling variability (internal consistency), and rater subjectivity (inter-rater). Identifying the relevant error source guides your choice of reliability estimate.
3

Validity as Evidential Argument

Modern psychometrics treats validity not as a property of a test but as a property of test score interpretations. Evidence for validity accumulates from multiple sources—content, internal structure, relationships with other variables, response processes, and consequences.
4

Reliability Sets the Ceiling for Validity

The validity coefficient of a test cannot exceed the square root of its reliability coefficient. Unreliable tests attenuate correlations with criterion measures, artificially lowering apparent validity. This mathematical relationship underscores why reliability assessment is always the first step.
5

Standard Error of Measurement (SEM)

The SEM quantifies the expected range within which a person's true score falls. It links reliability to clinical decision-making: a test with low reliability produces wide confidence intervals, undermining the precision needed for diagnosis or treatment planning.
KEY TAKEAWAY
Think of reliability and validity like a bathroom scale. Reliability means the scale gives you the same reading each time you step on it within the same minute. Validity means the number it shows actually corresponds to your true weight. A scale that consistently reads 10 pounds too heavy is reliable but not valid. A scale that shows different numbers every time is neither reliable nor valid. You need both properties for the measurement to be clinically useful.

Visual Explanation — Reliability Types at a Glance

This diagram illustrates the four major types of reliability, each targeting a distinct source of measurement error. Note that all four yield a reliability coefficient between 0 and 1, and all feed into the overarching question of whether a test produces sufficiently consistent scores for its intended clinical use.

The diagram above organizes the reliability landscape according to the specific source of measurement error each type addresses. Test-retest reliability examines whether scores remain stable across time, making temporal instability the targeted error source. Alternate-forms (parallel-forms) reliability assesses whether two versions of a test produce comparable scores, isolating content sampling error. Internal consistency reliability evaluates whether items within a single test cohere, and inter-rater reliability determines agreement across different scorers. For EPPP preparation, it is essential to match each reliability type with its error source, statistical index, and appropriate use case.

Mathematical Framework

Classical test theory provides the mathematical backbone for understanding reliability. The equations below formalize the relationship between observed scores, true scores, error, and derived reliability metrics. Mastery of these formulas is essential for interpreting test manuals and evaluating assessment instruments in clinical practice.

CLASSICAL TEST THEORY EQUATION
X = T + E
X = observed score; T = true score (the score a person would obtain if there were no measurement error); E = error. Error is assumed to be random, with a mean of zero across repeated testing.
RELIABILITY COEFFICIENT
rₓₓ = σ²ₜ / σ²ₓ = 1 − (σ²ₑ / σ²ₓ)
rₓₓ = reliability coefficient; σ²ₜ = true score variance; σ²ₓ = observed score variance; σ²ₑ = error variance. Reliability is the proportion of observed variance that is attributable to true score variance.
STANDARD ERROR OF MEASUREMENT
SEM = SDₓ × √(1 − rₓₓ)
SEM = standard error of measurement; SDₓ = standard deviation of observed scores; rₓₓ = reliability coefficient. The SEM creates a confidence interval around an individual's observed score to estimate where the true score likely falls. A 95% confidence interval = Observed Score ± 1.96 × SEM.
CRONBACH'S ALPHA
α = (k / (k − 1)) × (1 − (Σσ²ᵢ / σ²ₓ))
α = Cronbach's alpha; k = number of items; σ²ᵢ = variance of item i; σ²ₓ = total test score variance. Alpha represents the average of all possible split-half reliability coefficients. For dichotomous items, this formula reduces to the Kuder-Richardson 20 (KR-20).
📏 Clinical Rule of Thumb
For individual clinical decisions (e.g., diagnosis, placement), a reliability coefficient of ≥ 0.90 is preferred. For research and group-level decisions, ≥ 0.80 is generally considered acceptable. The higher the stakes, the more stringent the reliability requirement.

Detailed Breakdown — Forms of Validity Evidence

While the EPPP frequently references the traditional triad of content, criterion-related, and construct validity, the current Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) conceptualizes validity as a unitary construct supported by five categories of evidence. Understanding both frameworks is essential for the exam. The table and diagram below cross-reference the traditional labels with the modern evidence categories.

This diagram illustrates the modern unitary validity framework. All five categories of evidence—test content, response processes, internal structure, relations with other variables, and consequences—feed into the overarching construct validity argument represented by the dashed ellipse. Traditional validity labels are mapped to their modern counterparts at the bottom.
Traditional Validity Labels Mapped to Modern Evidence Categories
Traditional LabelModern Evidence CategoryKey MethodsClinical Example
Content ValidityEvidence based on test contentExpert panel review, content blueprints, item-objective alignmentEPPP items are mapped to a content domain blueprint to ensure adequate coverage of all knowledge areas.
Criterion-Related (Predictive)Relations with other variablesCorrelation with future criterion measured later in timeGRE scores correlate with first-year graduate GPA (criterion measured months/years later).
Criterion-Related (Concurrent)Relations with other variablesCorrelation with criterion measured at the same timeA new depression scale correlates .85 with the BDI-II administered simultaneously.
Construct ValidityAll five evidence categoriesConvergent/discriminant evidence, factor analysis, known-groups comparisons, MTMM matrixAn anxiety measure correlates highly with other anxiety measures (convergent) and weakly with unrelated constructs (discriminant).
Face ValidityNot a formal evidence categorySubjective judgment: does the test 'look like' it measures what it claims?A depression questionnaire asks about mood and sleep. Face validity affects examinee buy-in but is not formal psychometric evidence.
🎯 EPPP Exam Tip: Convergent vs. Discriminant
Convergent and discriminant validity are subtypes of construct validity, often assessed using Campbell and Fiske's Multitrait-Multimethod (MTMM) matrix. Convergent validity is demonstrated when measures of theoretically related constructs correlate highly. Discriminant validity is demonstrated when measures of theoretically unrelated constructs show low correlations. The EPPP frequently tests your ability to distinguish between these two concepts.

Worked Example — Evaluating a Depression Screener

Imagine you are a psychologist selecting a new depression screening instrument, the Hypothetical Depression Scale (HDS), for use in a community mental health center. The test manual reports the following psychometric data: the HDS has 20 items scored on a 4-point Likert scale, coefficient alpha of 0.88, test-retest reliability (2-week interval) of 0.82, a correlation of 0.79 with the BDI-II (concurrent criterion), and a standard deviation of 10. Let us walk through how to interpret these data.

Evaluating the Reliability and Validity of the HDS
1
Step 1 — Evaluate Internal ConsistencyCronbach's alpha is reported as 0.88. This exceeds the 0.80 threshold for research purposes, suggesting that the 20 items are measuring a reasonably homogeneous construct. However, for individual clinical decisions (such as making a diagnosis), we ideally want α ≥ 0.90. Thus, the HDS is adequate for screening but may lack the precision needed for sole reliance in diagnostic decision-making.
α = 0.88 → Adequate for screening, marginal for high-stakes individual decisions.
2
Step 2 — Evaluate Temporal StabilityThe test-retest reliability over a 2-week interval is 0.82. Because depression symptoms can fluctuate, a moderate test-retest coefficient is expected for state-dependent constructs. If the construct were a stable trait (e.g., personality), we would expect higher test-retest values (≥ 0.85). The 2-week interval is appropriate—too short risks practice effects; too long introduces genuine construct change.
r = 0.82 over 2 weeks → Acceptable for a state-dependent measure.
3
Step 3 — Calculate the Standard Error of MeasurementUsing the SEM formula: SEM = SDₓ × √(1 − rₓₓ). With SDₓ = 10 and rₓₓ = 0.88, we calculate SEM = 10 × √(1 − 0.88) = 10 × √0.12 = 10 × 0.346 ≈ 3.46. This means a 95% confidence interval around any individual's score spans approximately ±6.8 points (1.96 × 3.46). If a client scores 25, their true score likely falls between 18.2 and 31.8.
SEM ≈ 3.46 → 95% CI for a score of 25: [18.2, 31.8]
4
Step 4 — Evaluate Concurrent Criterion ValidityThe HDS correlates 0.79 with the BDI-II, administered simultaneously. This constitutes concurrent criterion-related validity evidence. The correlation is strong, suggesting the HDS measures a construct highly related to what the BDI-II measures. To determine how much variance the HDS shares with the BDI-II, we square the correlation: r² = 0.79² = 0.624, meaning approximately 62% shared variance.
r = 0.79 with BDI-II → 62% shared variance → Strong concurrent validity evidence.
5
Step 5 — Consider the Validity CeilingThe maximum possible validity coefficient is limited by reliability: max r(validity) = √rₓₓ = √0.88 ≈ 0.938. The obtained concurrent validity coefficient of 0.79 is well within this ceiling (0.79 / 0.938 = 84% of the maximum possible validity), suggesting that measurement error is not unduly suppressing the validity estimate. Had the reliability been lower—say 0.60—the maximum validity would be only √0.60 ≈ 0.775, severely limiting the instrument's potential.
Max validity = √0.88 ≈ 0.94 → Obtained 0.79 uses 84% of theoretical maximum.

Strengths, Limitations, and Common Confusions

Comparison of Reliability Types: Strengths and Limitations
Reliability TypeStrengthsLimitations / Pitfalls
Test-RetestDirectly assesses temporal stability; easy to understand and compute.Practice effects inflate estimates with short intervals; genuine change in the construct (e.g., mood states) deflates estimates with long intervals. Inappropriate for unstable constructs.
Alternate FormsControls for practice effects and memory; addresses content sampling error.Creating truly parallel forms is difficult and expensive. If administered at different times, confounds temporal instability with content sampling.
Internal ConsistencyRequires only one administration; efficient and widely reported.Inappropriate for speeded tests (inflated estimates). Overly heterogeneous or multidimensional tests may show misleadingly low alpha. Does not address temporal stability.
Inter-RaterEssential for subjectively scored assessments; Cohen's kappa corrects for chance agreement.Percentage agreement alone is misleading because it ignores chance. Kappa can be paradoxically low when base rates are extreme. Requires training raters, which adds cost.
Split-HalfSingle administration; Spearman-Brown correction adjusts for test length reduction.Result depends on how the test is split; only one of many possible splits is evaluated. Superseded by Cronbach's alpha, which averages all possible split-halves.
KEY TAKEAWAY
Think of choosing a reliability estimate like selecting a diagnostic lab test in medicine. A blood glucose test taken once tells you about current levels (analogous to internal consistency from a single administration), while hemoglobin A1c assessed across months tells you about stability over time (analogous to test-retest reliability). Different clinical questions demand different reliability evidence, just as different medical questions demand different lab panels. No single reliability estimate answers every question about a test's consistency.

Connection to Advanced Psychometric Theory

Classical test theory (CTT), while foundational, has notable limitations that advanced psychometric frameworks seek to address. The EPPP does not require deep expertise in these advanced methods, but awareness of their existence and relationship to CTT concepts of reliability and validity is increasingly important. Two advanced frameworks deserve mention: Generalizability Theory (G-Theory) and Item Response Theory (IRT).

CTT vs. Advanced Psychometric Frameworks
FeatureClassical Test Theory (CTT)Generalizability Theory (G-Theory)Item Response Theory (IRT)
Error modelSingle undifferentiated error term (E)Multiple error facets (raters, items, occasions) simultaneously estimated via ANOVAItem-level error; reliability varies by trait level (information function)
Reliability indexSingle coefficient (α, r)Generalizability coefficient (Eρ²) and dependability coefficient (Φ)Test information function I(θ) — precision varies across the trait continuum
Sample dependencyStatistics are sample-dependent; reliability varies across populationsAlso sample-dependent but can model multiple error facets in a single studyItem parameters are sample-independent (invariance); enables adaptive testing
Clinical applicationMost test manuals report CTT indices; used for standard clinical instrumentsUsed when multiple error sources matter (e.g., OSCE medical exams with raters × stations)Computer adaptive testing (e.g., CAT-MH); items tailored to examinee's ability level

For the EPPP, the most critical advanced concept is understanding that Generalizability Theory extends CTT by partitioning error variance into multiple facets simultaneously, rather than treating all error as a monolithic term. In clinical settings, this is increasingly relevant when behavioral health assessments involve multiple raters observing multiple clients across multiple occasions—a scenario common in behavioral observation systems and competency-based clinical evaluations. IRT, meanwhile, offers the concept of measurement precision that varies along the trait continuum, unlike CTT's single reliability coefficient. This is especially valuable for instruments designed to discriminate at clinical cutoff points, where precision at the decision threshold matters most.

Practice Problems

PROBLEM 1CONCEPTUAL
A psychologist administers a personality inventory to the same group of clients two weeks apart and correlates the two sets of scores. The resulting correlation coefficient of 0.85 is an estimate of which type of reliability? Additionally, what specific source of measurement error does this reliability type address?
PROBLEM 2BASIC CALCULATION
A cognitive screening test has a reliability coefficient of 0.91 and a standard deviation of 15. Calculate the Standard Error of Measurement (SEM) and construct a 95% confidence interval around an observed score of 100.
PROBLEM 3INTERMEDIATE
A researcher develops a new anxiety measure and reports that it correlates 0.82 with the State-Trait Anxiety Inventory (STAI) and 0.15 with a measure of extroversion. The researcher claims these findings support the construct validity of the new measure. Evaluate this claim, specifying what types of construct validity evidence each correlation represents and whether the evidence is adequate.
PROBLEM 4APPLIED
You are a clinical psychologist working in a forensic setting. You must select a risk assessment instrument for predicting violent recidivism among parolees. You have narrowed your choice to two instruments: Instrument A reports high internal consistency (α = 0.92) but no predictive validity data, while Instrument B reports moderate internal consistency (α = 0.78) but strong predictive validity (r = 0.65 with violent recidivism over 5 years). Which instrument would you choose and why? What additional psychometric information would you request?
PROBLEM 5CRITICAL THINKING
A test manual reports that a 50-item depression measure has a Cronbach's alpha of 0.95. The test publisher uses this as evidence of excellent measurement quality. Critically evaluate this claim. Under what circumstances might an alpha of 0.95 actually signal a problem with the instrument? How does this relate to the distinction between reliability and validity?

Summary — Reliability and Validity in Psychological Assessment

Reliability refers to the consistency and stability of test scores, rooted in classical test theory's decomposition of observed scores into true score plus error (X = T + E). Four major types address distinct error sources: test-retest (temporal instability), alternate forms (content sampling), internal consistency (item heterogeneity, indexed by Cronbach's alpha, KR-20, or split-half methods), and inter-rater reliability (rater subjectivity, indexed by Cohen's kappa or ICC). The Standard Error of Measurement (SEM) translates the reliability coefficient into a clinically meaningful confidence interval around an individual's observed score.

Validity is the degree to which evidence supports the intended interpretation of test scores. The modern framework identifies five evidence sources: test content, response processes, internal structure, relations with other variables (encompassing convergent, discriminant, predictive, and concurrent evidence), and consequences of testing. The cardinal rule is that reliability is necessary but not sufficient for validity: a test cannot accurately measure a construct if its scores are inconsistent, but perfectly consistent scores can still miss the intended construct entirely. When selecting or evaluating assessment instruments, clinicians must examine both reliability and validity evidence in light of the specific purpose, population, and stakes of the assessment.

Varsity Tutors • EPPP: Part 1, Knowledge • Reliability And Validity — Differentiate reliability types and validity forms in test interpretation