Historical Context & Motivation
The systematic study of research validity threats grew out of a practical need: clinicians and policymakers required confidence that the effects reported in experiments were real and generalizable. As behavioral science expanded during the mid-twentieth century, researchers recognized that poorly controlled studies could yield misleading conclusions—conclusions that, when applied in clinical settings, might harm rather than help clients. The identification and classification of validity threats thus became a foundational concern for any scientist-practitioner who wished to evaluate the quality of evidence supporting a given intervention or theory.
The central question this framework addresses is deceptively simple: How can we determine whether a study's findings reflect genuine causal relationships, and whether those findings apply beyond the specific conditions under which the study was conducted? For EPPP candidates working in behavioral health, the ability to identify validity threats is not merely an academic exercise—it is the skill that underpins evidence-based practice and ethical clinical decision-making.
Core Principles & Definitions
At its core, research validity concerns the degree to which a study's conclusions are warranted by the design, data, and analytic methods employed. Campbell and colleagues distinguished two primary types that EPPP candidates must master. Internal validity asks whether the independent variable truly caused the observed change in the dependent variable, ruling out alternative explanations. External validity asks whether those causal conclusions can be generalized to other populations, settings, times, and operationalizations. A study may possess strong internal validity yet weak external validity, or vice versa—understanding the tension between these two forms of validity is essential for critically evaluating behavioral health research.
Internal Validity
External Validity
Confound
Construct Validity
Statistical Conclusion Validity
Visual Explanation: The Validity Threat Landscape
The diagram above provides a comprehensive visual map of the threats you must be able to recognize for the EPPP. Notice that the internal validity threats are organized to reflect the classic Campbell and Stanley taxonomy, while the external validity threats emphasize interaction effects—situations where the treatment effect depends on specific conditions of the study that may not replicate elsewhere. When evaluating a research scenario on the exam, your first step should be to determine whether the flaw compromises the causal inference (internal) or the generalizability of the inference (external), and then to identify the specific threat by name.
How Validity Threats Operate: Mechanisms and Logic
Internal Validity Threats in Depth
Each internal validity threat functions by introducing a plausible alternative explanation for the observed relationship between the independent variable and the dependent variable. When such an alternative explanation exists, researchers cannot confidently attribute the outcome to the treatment or manipulation. Understanding the mechanism of each threat allows you to identify them in novel research scenarios—a skill the EPPP frequently tests.
History refers to external events occurring between pretest and posttest that could affect the dependent variable independently of the treatment. For example, if a researcher is evaluating a depression intervention over six months and a major natural disaster occurs during the study period, observed changes in depressive symptoms may reflect the disaster's psychological impact rather than the intervention's efficacy. The threat is most potent in designs that lack a control group, because without a comparison condition, there is no way to distinguish treatment effects from historical events.
Maturation encompasses any systematic change occurring within participants as a function of time—biological growth, fatigue, cognitive development, or spontaneous recovery. In behavioral health research, natural symptom fluctuation is a particularly important form of maturation. A client presenting at peak distress will often improve over time regardless of intervention, a phenomenon related to regression to the mean and the natural course of many psychiatric conditions.
Testing (also called practice effects or reactivity to assessment) occurs when the act of taking a pretest alters performance on the posttest. Participants may become familiar with test content, learn strategies, or become sensitized to the constructs being measured. In clinical research, administering a baseline depression inventory may prompt self-reflection that itself has therapeutic value, potentially inflating posttest improvement irrespective of the intervention.
Instrumentation threats arise when the measurement instrument or procedure changes between assessments. This includes recalibrated equipment, different raters or interviewers at pre- versus post-assessment, or changes in scoring criteria. If clinician-rated outcome measures are used, raters may become more lenient or more stringent over time (observer drift), producing apparent changes in the dependent variable that reflect measurement artifacts rather than true change.
Statistical regression (regression to the mean) is a mathematical phenomenon whereby extreme scores on any measure tend to move toward the group mean upon retesting, purely as a function of measurement error. When participants are selected because they scored extremely high or low on a screening measure, their subsequent scores will likely be less extreme regardless of any intervention. This threat is especially relevant in behavioral health research where participants are often recruited based on clinical cutoff scores.
Selection bias occurs when groups differ systematically before the treatment is applied. In true experiments, random assignment minimizes this threat. In quasi-experimental designs—common in behavioral health settings where random assignment may be ethically or practically infeasible—selection bias is a primary concern. If one clinic's clientele is systematically different from another's, comparing treatment outcomes across clinics conflates treatment effects with pre-existing group differences.
Attrition (also called experimental mortality) occurs when participants drop out of the study differentially across conditions. If participants who are not improving leave the treatment group while those who are improving remain, the treatment group's average outcome will appear inflated—not because the treatment worked for everyone, but because the non-responders are no longer counted. Differential attrition is particularly problematic in lengthy behavioral health trials where treatment burden may be unevenly distributed across conditions.
External Validity Threats in Depth
External validity threats limit the degree to which findings can be extended beyond the specific study conditions. The interaction of selection and treatment is perhaps the most consequential for behavioral health: if a treatment was tested only on young, college-educated volunteers recruited through online advertisements, its efficacy among older, lower-income adults seeking treatment in community mental health centers remains uncertain. The reactive effects of the experimental setting (also termed ecological validity concerns) arise when laboratory or research-clinic conditions differ so substantially from real-world practice that treatment effects may not transfer. The Hawthorne effect and demand characteristics further complicate external validity, because participants who know they are being studied may behave differently than they would in routine clinical care.
Classifying Threats by Research Design
Different research designs are vulnerable to different constellations of validity threats. Understanding which threats a given design controls for—and which it leaves unaddressed—is essential for evaluating the strength of evidence in behavioral health research. The following diagram illustrates how three common research designs relate to the spectrum of validity threats.
| Threat | Definition | Primary Design Safeguard |
|---|---|---|
| History | External events co-occurring with the treatment that affect the DV | Control group exposed to same historical context |
| Maturation | Passage-of-time changes (growth, fatigue, spontaneous remission) | Control group matures at similar rate |
| Testing | Pretest exposure improves posttest performance | Solomon four-group design; posttest-only control group |
| Instrumentation | Changes in measurement tools, criteria, or observers over time | Standardized protocols; inter-rater reliability checks |
| Statistical Regression | Extreme scores regress toward the mean on retest | Avoid selecting participants based on extreme scores alone; use reliable measures |
| Selection | Pre-existing differences between groups | Random assignment; matching; ANCOVA |
| Attrition | Differential dropout across conditions | Intent-to-treat analysis; minimize burden; track dropouts |
| Diffusion of Treatment | Control group receives elements of the treatment | Physical separation of groups; treatment fidelity monitoring |
Worked Example: Identifying Validity Threats in a Clinical Trial
Consider the following research scenario, representative of the type you might encounter on the EPPP:
Internal vs. External Validity: Tensions and Trade-Offs
One of the most important conceptual tensions in research methodology is the trade-off between internal and external validity. Highly controlled laboratory experiments maximize internal validity by eliminating confounds, but they often accomplish this by creating artificial conditions that do not resemble real-world clinical practice. Conversely, naturalistic or community-based studies maximize external validity by studying real clients in real settings, but they sacrifice the control necessary for strong causal inference. The responsible researcher does not pursue one at the expense of the other; instead, they select designs that offer the best feasible balance for the research question at hand.
| Dimension | Internal Validity | External Validity |
|---|---|---|
| Central Question | Did the IV cause the change in the DV? | Can findings generalize to other people, settings, and times? |
| Strongest Design | True experiment with random assignment, double-blind procedures, and strict controls | Large-scale multi-site field study with diverse, representative samples |
| Key Threats | History, maturation, testing, instrumentation, regression, selection, attrition | Selection × treatment, reactive settings, Hawthorne effect, multiple treatment interference |
| Trade-Off | High control may reduce ecological validity and limit generalizability | Real-world conditions introduce confounds that weaken causal conclusions |
| Behavioral Health Example | RCT of CBT for PTSD in a university research clinic with manualized treatment | Effectiveness study of CBT for PTSD in community agencies with varied therapist training |
| EPPP Relevance | Frequently tested: identifying which threats a design controls | Frequently tested: recognizing limits on generalizability from study features |
Connection to Advanced Validity Theory and Modern Applications
The original Campbell and Stanley framework focused primarily on internal and external validity. The expanded Shadish, Cook, and Campbell (2002) model adds two additional validity types that EPPP candidates should understand. Statistical conclusion validity concerns whether the statistical analysis correctly identifies the presence or absence of covariation between the independent and dependent variables—threats include low power, violated assumptions, unreliable measures, and inflated Type I error from multiple comparisons. Construct validity concerns whether the study's operational definitions accurately capture the theoretical constructs of interest—threats include mono-operation bias (using only one operationalization), mono-method bias, treatment diffusion, and experimenter expectancies that shape how the construct manifests.
| Validity Type | Core Question | Example Threats |
|---|---|---|
| Statistical Conclusion | Is the statistical relationship correctly identified? | Low power, fishing / multiple comparisons, violated assumptions, unreliable measures, restricted range |
| Internal | Is the relationship causal? | History, maturation, testing, instrumentation, regression, selection, attrition |
| Construct | Are we measuring/manipulating what we think we are? | Mono-operation bias, mono-method bias, hypothesis guessing, evaluation apprehension, experimenter expectancy |
| External | Can findings generalize beyond this study? | Interaction of selection × treatment, reactive settings, multiple treatment interference, Hawthorne effect |
Modern behavioral health research increasingly employs strategies to address validity threats simultaneously. Multisite randomized controlled trials enhance both internal validity (through randomization) and external validity (through diverse sites and populations). Pragmatic trials prioritize real-world conditions, accepting some reduction in control in exchange for greater generalizability. Meta-analyses synthesize findings across studies to assess the robustness of effects across different samples and contexts, directly addressing external validity concerns. The EPPP may ask you to identify which validity type is most relevant to a given research flaw, so familiarity with all four types is essential.
Practice Problems
Summary & Review
Research validity threats are the systematic flaws that undermine confidence in a study's conclusions. Internal validity is the degree to which the independent variable—rather than confounds such as history, maturation, testing, instrumentation, statistical regression, selection, or attrition—caused the observed outcome. External validity is the degree to which findings generalize beyond the study's specific sample, setting, and conditions. Key external threats include the interaction of selection and treatment, reactive effects of the experimental setting, and the Hawthorne effect.
True experiments with random assignment offer the strongest protection against internal validity threats but may sacrifice external validity through restrictive conditions. The four-validity framework of Shadish, Cook, and Campbell (2002) adds statistical conclusion validity and construct validity to the original two-validity model. For the EPPP, practice identifying threats in research scenarios by first classifying whether the flaw compromises causal inference (internal) or generalizability (external), and then naming the specific threat. The distinction between efficacy and effectiveness research reflects the practical implications of this trade-off for evidence-based practice in behavioral health.