EPPP: PART 1, KNOWLEDGE • DOMAIN 7: RESEARCH METHODS AND STATISTICS

Validity Threats — Identify threats to internal and external validity in research scenarios

Understanding how confounds and design flaws compromise causal inferences and limit the generalizability of behavioral health research.

Historical Context & Motivation

The systematic study of research validity threats grew out of a practical need: clinicians and policymakers required confidence that the effects reported in experiments were real and generalizable. As behavioral science expanded during the mid-twentieth century, researchers recognized that poorly controlled studies could yield misleading conclusions—conclusions that, when applied in clinical settings, might harm rather than help clients. The identification and classification of validity threats thus became a foundational concern for any scientist-practitioner who wished to evaluate the quality of evidence supporting a given intervention or theory.

1957
Campbell's Seminal Framework
Donald T. Campbell published his influential paper distinguishing internal validity from external validity, establishing a taxonomy of threats that could undermine experimental conclusions in social science research.
1963
Campbell & Stanley's Handbook
Campbell and Julian Stanley formalized eight threats to internal validity and four threats to external validity in their landmark chapter on experimental and quasi-experimental designs for research, providing a checklist that became standard across the behavioral sciences.
1979
Cook & Campbell's Expansion
Thomas Cook and Donald Campbell expanded the framework to include statistical conclusion validity and construct validity, creating a four-validity system that offered a more nuanced understanding of how research conclusions could go wrong.
2002
Shadish, Cook & Campbell's Modern Synthesis
Shadish, Cook, and Campbell published the definitive modern text integrating the full validity framework, extending its application to field experiments, quasi-experiments, and the kinds of community-based behavioral health research common in contemporary practice.

The central question this framework addresses is deceptively simple: How can we determine whether a study's findings reflect genuine causal relationships, and whether those findings apply beyond the specific conditions under which the study was conducted? For EPPP candidates working in behavioral health, the ability to identify validity threats is not merely an academic exercise—it is the skill that underpins evidence-based practice and ethical clinical decision-making.

Core Principles & Definitions

At its core, research validity concerns the degree to which a study's conclusions are warranted by the design, data, and analytic methods employed. Campbell and colleagues distinguished two primary types that EPPP candidates must master. Internal validity asks whether the independent variable truly caused the observed change in the dependent variable, ruling out alternative explanations. External validity asks whether those causal conclusions can be generalized to other populations, settings, times, and operationalizations. A study may possess strong internal validity yet weak external validity, or vice versa—understanding the tension between these two forms of validity is essential for critically evaluating behavioral health research.

1

Internal Validity

The confidence that the independent variable—and not some confounding factor—produced the observed effect on the dependent variable. It is the sine qua non of causal inference in experimental research.
2

External Validity

The extent to which findings generalize across different persons, settings, treatment variables, and measurement approaches. Without external validity, a study's clinical utility remains uncertain.
3

Confound

Any variable that covaries with the independent variable and could plausibly explain the observed effect, thereby creating an alternative explanation and threatening internal validity.
4

Construct Validity

Whether the study's operational definitions (measures, manipulations) accurately represent the theoretical constructs they are intended to capture—a concern at the measurement level.
5

Statistical Conclusion Validity

Whether the statistical analysis supports the conclusions drawn—including appropriate power, effect size interpretation, and correct use of inferential tests.
KEY TAKEAWAY
Think of internal and external validity like the difference between a controlled lab recipe and a home kitchen. Internal validity is whether the recipe works as written under ideal conditions (the right ingredients, precise measurements, calibrated oven). External validity is whether that same recipe produces a good result when different cooks try it in different kitchens with slightly different ingredients and equipment. A recipe that only works in one specific lab kitchen has limited practical value—just as a therapy outcome that only replicates under narrow conditions has limited clinical utility.

Visual Explanation: The Validity Threat Landscape

This diagram contrasts threats to internal validity (left panel) with threats to external validity (right panel). Internal threats compromise causal inference within the study, while external threats limit the generalizability of findings to broader populations and settings.

The diagram above provides a comprehensive visual map of the threats you must be able to recognize for the EPPP. Notice that the internal validity threats are organized to reflect the classic Campbell and Stanley taxonomy, while the external validity threats emphasize interaction effects—situations where the treatment effect depends on specific conditions of the study that may not replicate elsewhere. When evaluating a research scenario on the exam, your first step should be to determine whether the flaw compromises the causal inference (internal) or the generalizability of the inference (external), and then to identify the specific threat by name.

How Validity Threats Operate: Mechanisms and Logic

Internal Validity Threats in Depth

Each internal validity threat functions by introducing a plausible alternative explanation for the observed relationship between the independent variable and the dependent variable. When such an alternative explanation exists, researchers cannot confidently attribute the outcome to the treatment or manipulation. Understanding the mechanism of each threat allows you to identify them in novel research scenarios—a skill the EPPP frequently tests.

History refers to external events occurring between pretest and posttest that could affect the dependent variable independently of the treatment. For example, if a researcher is evaluating a depression intervention over six months and a major natural disaster occurs during the study period, observed changes in depressive symptoms may reflect the disaster's psychological impact rather than the intervention's efficacy. The threat is most potent in designs that lack a control group, because without a comparison condition, there is no way to distinguish treatment effects from historical events.

Maturation encompasses any systematic change occurring within participants as a function of time—biological growth, fatigue, cognitive development, or spontaneous recovery. In behavioral health research, natural symptom fluctuation is a particularly important form of maturation. A client presenting at peak distress will often improve over time regardless of intervention, a phenomenon related to regression to the mean and the natural course of many psychiatric conditions.

Testing (also called practice effects or reactivity to assessment) occurs when the act of taking a pretest alters performance on the posttest. Participants may become familiar with test content, learn strategies, or become sensitized to the constructs being measured. In clinical research, administering a baseline depression inventory may prompt self-reflection that itself has therapeutic value, potentially inflating posttest improvement irrespective of the intervention.

Instrumentation threats arise when the measurement instrument or procedure changes between assessments. This includes recalibrated equipment, different raters or interviewers at pre- versus post-assessment, or changes in scoring criteria. If clinician-rated outcome measures are used, raters may become more lenient or more stringent over time (observer drift), producing apparent changes in the dependent variable that reflect measurement artifacts rather than true change.

Statistical regression (regression to the mean) is a mathematical phenomenon whereby extreme scores on any measure tend to move toward the group mean upon retesting, purely as a function of measurement error. When participants are selected because they scored extremely high or low on a screening measure, their subsequent scores will likely be less extreme regardless of any intervention. This threat is especially relevant in behavioral health research where participants are often recruited based on clinical cutoff scores.

Selection bias occurs when groups differ systematically before the treatment is applied. In true experiments, random assignment minimizes this threat. In quasi-experimental designs—common in behavioral health settings where random assignment may be ethically or practically infeasible—selection bias is a primary concern. If one clinic's clientele is systematically different from another's, comparing treatment outcomes across clinics conflates treatment effects with pre-existing group differences.

Attrition (also called experimental mortality) occurs when participants drop out of the study differentially across conditions. If participants who are not improving leave the treatment group while those who are improving remain, the treatment group's average outcome will appear inflated—not because the treatment worked for everyone, but because the non-responders are no longer counted. Differential attrition is particularly problematic in lengthy behavioral health trials where treatment burden may be unevenly distributed across conditions.

External Validity Threats in Depth

External validity threats limit the degree to which findings can be extended beyond the specific study conditions. The interaction of selection and treatment is perhaps the most consequential for behavioral health: if a treatment was tested only on young, college-educated volunteers recruited through online advertisements, its efficacy among older, lower-income adults seeking treatment in community mental health centers remains uncertain. The reactive effects of the experimental setting (also termed ecological validity concerns) arise when laboratory or research-clinic conditions differ so substantially from real-world practice that treatment effects may not transfer. The Hawthorne effect and demand characteristics further complicate external validity, because participants who know they are being studied may behave differently than they would in routine clinical care.

Classifying Threats by Research Design

Different research designs are vulnerable to different constellations of validity threats. Understanding which threats a given design controls for—and which it leaves unaddressed—is essential for evaluating the strength of evidence in behavioral health research. The following diagram illustrates how three common research designs relate to the spectrum of validity threats.

This diagram compares three research designs—the one-group pretest–posttest (weakest internal validity), the quasi-experimental design (moderate), and the true experiment (RCT) (strongest)—in terms of which internal validity threats they control. Notice the trade-off: stronger internal validity often comes at the cost of external validity.
Key Internal Validity Threats and Their Primary Design Safeguards
ThreatDefinitionPrimary Design Safeguard
HistoryExternal events co-occurring with the treatment that affect the DVControl group exposed to same historical context
MaturationPassage-of-time changes (growth, fatigue, spontaneous remission)Control group matures at similar rate
TestingPretest exposure improves posttest performanceSolomon four-group design; posttest-only control group
InstrumentationChanges in measurement tools, criteria, or observers over timeStandardized protocols; inter-rater reliability checks
Statistical RegressionExtreme scores regress toward the mean on retestAvoid selecting participants based on extreme scores alone; use reliable measures
SelectionPre-existing differences between groupsRandom assignment; matching; ANCOVA
AttritionDifferential dropout across conditionsIntent-to-treat analysis; minimize burden; track dropouts
Diffusion of TreatmentControl group receives elements of the treatmentPhysical separation of groups; treatment fidelity monitoring

Worked Example: Identifying Validity Threats in a Clinical Trial

Consider the following research scenario, representative of the type you might encounter on the EPPP:

📋 SCENARIO
A community mental health center wants to evaluate a new 12-week group therapy program for adults with generalized anxiety disorder (GAD). They recruit 60 clients who score above the clinical cutoff on the GAD-7. Clients at Site A receive the new group therapy; clients at Site B receive treatment as usual (TAU). Assessments are conducted at baseline and at 12 weeks. At follow-up, Site A shows significantly greater improvement. However, 15 of the 30 clients at Site A completed the program, while 25 of the 30 clients at Site B completed TAU. Additionally, a new statewide teletherapy initiative launched at week 6, providing free anxiety resources to all state residents.
Systematic Threat Analysis
1
Step 1 — Identify the DesignThis is a quasi-experimental nonequivalent control group design. Participants were not randomly assigned to conditions—they were assigned based on their site location. The absence of random assignment means we must immediately consider selection bias and related interaction threats.
Design = quasi-experimental; random assignment absent
2
Step 2 — Check for Selection BiasBecause clients at Sites A and B were not randomly assigned, they may differ on important baseline characteristics (e.g., severity of GAD, comorbidities, socioeconomic status, motivation for treatment). These pre-existing differences could account for the between-group outcome differences. This is a clear selection threat to internal validity.
Threat identified: Selection bias
3
Step 3 — Check for Statistical RegressionAll participants were selected because they scored above the clinical cutoff on the GAD-7. This means the sample was selected based on extreme scores. Due to measurement error, some participants may have scored above the cutoff by chance and would be expected to score lower upon retesting—regardless of treatment. This constitutes a statistical regression threat. While both groups are affected, differential regression could occur if Site A and Site B had different baseline score distributions.
Threat identified: Regression to the mean
4
Step 4 — Check for AttritionThis is arguably the most conspicuous threat in the scenario. At Site A, only 15 of 30 clients (50%) completed the program, whereas at Site B, 25 of 30 (83%) completed TAU. This dramatic differential attrition means the groups being compared at posttest are no longer equivalent to the groups that began the study. It is plausible that the 15 completers at Site A were the most motivated, least severe, or most responsive clients. The apparent treatment benefit may simply reflect the removal of non-responders from the treatment group.
Threat identified: Differential attrition (most serious threat in this scenario)
5
Step 5 — Check for HistoryThe statewide teletherapy initiative launched at week 6 constitutes a history threat. Because both sites are in the same state, both groups had access to the free anxiety resources. However, if clients at one site were more likely to access the initiative (e.g., due to differences in internet access or digital literacy), the history threat could interact with selection to produce differential effects. Notably, history is partially controlled in this design because both sites were exposed to the same external event—but only if the exposure was truly equivalent.
Threat identified: History (partially controlled by having two groups)
6
Step 6 — Evaluate External ValidityEven if the treatment effect were genuine, generalizability is limited. The sample was drawn from a single community mental health center in one geographic region (interaction of selection and treatment). The severe attrition at Site A means the findings generalize, at best, to a very specific subgroup of completers—not to all adults with GAD (interaction of selection and treatment). There may also be reactive effects of the experimental setting if participants knew they were receiving a new program and were consequently more engaged than typical clients would be.
External validity limited by: selection × treatment interaction, reactive setting effects, and completer-only analysis

Internal vs. External Validity: Tensions and Trade-Offs

One of the most important conceptual tensions in research methodology is the trade-off between internal and external validity. Highly controlled laboratory experiments maximize internal validity by eliminating confounds, but they often accomplish this by creating artificial conditions that do not resemble real-world clinical practice. Conversely, naturalistic or community-based studies maximize external validity by studying real clients in real settings, but they sacrifice the control necessary for strong causal inference. The responsible researcher does not pursue one at the expense of the other; instead, they select designs that offer the best feasible balance for the research question at hand.

Comparison of Internal vs. External Validity
DimensionInternal ValidityExternal Validity
Central QuestionDid the IV cause the change in the DV?Can findings generalize to other people, settings, and times?
Strongest DesignTrue experiment with random assignment, double-blind procedures, and strict controlsLarge-scale multi-site field study with diverse, representative samples
Key ThreatsHistory, maturation, testing, instrumentation, regression, selection, attritionSelection × treatment, reactive settings, Hawthorne effect, multiple treatment interference
Trade-OffHigh control may reduce ecological validity and limit generalizabilityReal-world conditions introduce confounds that weaken causal conclusions
Behavioral Health ExampleRCT of CBT for PTSD in a university research clinic with manualized treatmentEffectiveness study of CBT for PTSD in community agencies with varied therapist training
EPPP RelevanceFrequently tested: identifying which threats a design controlsFrequently tested: recognizing limits on generalizability from study features
KEY TAKEAWAY
The internal–external validity trade-off can be understood through the distinction between efficacy research and effectiveness research. Efficacy studies ask, 'Can this treatment work under optimal conditions?' (prioritizing internal validity). Effectiveness studies ask, 'Does this treatment work in the real world?' (prioritizing external validity). Both are necessary: efficacy without effectiveness means the treatment stays in the lab; effectiveness without efficacy means we cannot be sure the treatment—rather than some confound—is actually helping.

Connection to Advanced Validity Theory and Modern Applications

The original Campbell and Stanley framework focused primarily on internal and external validity. The expanded Shadish, Cook, and Campbell (2002) model adds two additional validity types that EPPP candidates should understand. Statistical conclusion validity concerns whether the statistical analysis correctly identifies the presence or absence of covariation between the independent and dependent variables—threats include low power, violated assumptions, unreliable measures, and inflated Type I error from multiple comparisons. Construct validity concerns whether the study's operational definitions accurately capture the theoretical constructs of interest—threats include mono-operation bias (using only one operationalization), mono-method bias, treatment diffusion, and experimenter expectancies that shape how the construct manifests.

The Four-Validity Framework (Shadish, Cook & Campbell, 2002)
Validity TypeCore QuestionExample Threats
Statistical ConclusionIs the statistical relationship correctly identified?Low power, fishing / multiple comparisons, violated assumptions, unreliable measures, restricted range
InternalIs the relationship causal?History, maturation, testing, instrumentation, regression, selection, attrition
ConstructAre we measuring/manipulating what we think we are?Mono-operation bias, mono-method bias, hypothesis guessing, evaluation apprehension, experimenter expectancy
ExternalCan findings generalize beyond this study?Interaction of selection × treatment, reactive settings, multiple treatment interference, Hawthorne effect

Modern behavioral health research increasingly employs strategies to address validity threats simultaneously. Multisite randomized controlled trials enhance both internal validity (through randomization) and external validity (through diverse sites and populations). Pragmatic trials prioritize real-world conditions, accepting some reduction in control in exchange for greater generalizability. Meta-analyses synthesize findings across studies to assess the robustness of effects across different samples and contexts, directly addressing external validity concerns. The EPPP may ask you to identify which validity type is most relevant to a given research flaw, so familiarity with all four types is essential.

💡 EPPP TIP
When an EPPP question describes a study flaw, ask yourself four questions in order: (1) Was the statistical analysis appropriate? (statistical conclusion validity). (2) Can we rule out alternative explanations? (internal validity). (3) Are the measures and manipulations capturing the right constructs? (construct validity). (4) Can findings apply beyond this specific study? (external validity). This sequence—often remembered as SICE—mirrors the logical order in which validity questions arise.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher finds that clients who completed a 16-week mindfulness-based stress reduction (MBSR) program showed reduced cortisol levels compared to baseline. There was no control group. A colleague argues that the improvement may reflect natural seasonal variation in cortisol. Which threat to internal validity does this critique represent, and why is the absence of a control group relevant?
PROBLEM 2BASIC CALCULATION
In a study of a new ADHD intervention, 40 children were selected because they scored at or above the 95th percentile on a behavioral rating scale. After 8 weeks of treatment, their mean score dropped significantly. The researcher concludes the treatment is effective. Identify the primary threat to this conclusion, and explain the statistical mechanism involved.
PROBLEM 3INTERMEDIATE
A substance abuse treatment program compares outcomes between clients who voluntarily enrolled in a new intensive outpatient program (IOP) and clients who continued in standard outpatient treatment. After 6 months, IOP clients have higher abstinence rates. However, the researchers note that 40% of IOP clients dropped out before completing the program, compared with 10% of standard treatment clients. Identify at least three validity threats in this scenario, specifying whether each threatens internal or external validity.
PROBLEM 4APPLIED
You are a psychologist on a grant review committee evaluating a proposed RCT of a telehealth PTSD intervention. The proposal includes: (a) random assignment of 200 veterans to telehealth CBT or in-person CBT; (b) pre- and post-treatment assessments using the PCL-5, administered by the treating clinician; (c) recruitment exclusively from one VA medical center; (d) clinicians who are aware of participants' treatment condition. Identify the validity threats present despite the use of random assignment, and recommend design modifications.
PROBLEM 5CRITICAL THINKING
A prominent meta-analysis of psychotherapy for depression synthesizes 50 RCTs and finds a moderate effect size (d = 0.62). A critic argues that this estimate is inflated due to 'the file drawer problem' and the predominance of efficacy trials using manualized treatment with graduate student therapists in university clinics. Evaluate this critique using the four-validity framework (statistical conclusion, internal, construct, and external validity).

Summary & Review

Research validity threats are the systematic flaws that undermine confidence in a study's conclusions. Internal validity is the degree to which the independent variable—rather than confounds such as history, maturation, testing, instrumentation, statistical regression, selection, or attrition—caused the observed outcome. External validity is the degree to which findings generalize beyond the study's specific sample, setting, and conditions. Key external threats include the interaction of selection and treatment, reactive effects of the experimental setting, and the Hawthorne effect.

True experiments with random assignment offer the strongest protection against internal validity threats but may sacrifice external validity through restrictive conditions. The four-validity framework of Shadish, Cook, and Campbell (2002) adds statistical conclusion validity and construct validity to the original two-validity model. For the EPPP, practice identifying threats in research scenarios by first classifying whether the flaw compromises causal inference (internal) or generalizability (external), and then naming the specific threat. The distinction between efficacy and effectiveness research reflects the practical implications of this trade-off for evidence-based practice in behavioral health.

Varsity Tutors • EPPP: Part 1, Knowledge • Validity Threats — Identify threats to internal and external validity in research scenarios