EPPP: PART 1, KNOWLEDGE • DOMAIN 6: TREATMENT AND INTERVENTION

Treatment Efficacy — Evaluate comparative efficacy and effectiveness of treatment modalities

Understanding how research distinguishes which therapies work, for whom, and under what conditions.

Historical Context & Motivation

The question of whether psychotherapy actually works — and which forms work best — has been central to clinical psychology since the mid-twentieth century. In 1952, Hans Eysenck published a provocative paper claiming that neurotic patients improved at roughly the same rate whether or not they received psychotherapy, sparking decades of research designed to demonstrate — or refute — the value of psychological interventions. This challenge forced the field to develop rigorous methods for evaluating treatment outcomes, ultimately giving rise to the concepts of treatment efficacy and treatment effectiveness that now anchor evidence-based practice.

1952
Eysenck's Challenge
Hans Eysenck argued that psychotherapy was no more effective than spontaneous remission, igniting a decades-long debate about treatment validation and pushing the field toward controlled outcome research.
1977
Smith & Glass Meta-Analysis
Mary Lee Smith and Gene Glass conducted the first large-scale meta-analysis of psychotherapy outcomes, analyzing 375 studies and finding an average effect size of 0.85 standard deviations favoring therapy over no treatment.
1995
APA Task Force on ESTs
The APA Division 12 Task Force published its first list of empirically supported treatments (ESTs), establishing criteria for well-established and probably efficacious treatments based on randomized controlled trial evidence.
2005
APA Evidence-Based Practice Policy
The APA adopted a formal policy defining evidence-based practice in psychology (EBPP) as the integration of the best available research with clinical expertise in the context of patient characteristics, culture, and preferences.
2010s–Present
Comparative Effectiveness & Personalized Medicine
Modern research increasingly emphasizes comparative effectiveness, dismantling studies, patient-treatment matching, and transdiagnostic approaches, moving beyond the binary 'does it work?' to 'for whom and how?'

The overarching question that treatment efficacy and effectiveness research seeks to answer is not merely whether a given intervention produces statistically significant change, but whether it produces clinically meaningful change that generalizes beyond the controlled research setting. As you prepare for the EPPP, understanding the distinction between internal validity (efficacy) and external validity (effectiveness) — and the methodologies that support each — is essential for evaluating the comparative merit of treatment modalities across diagnostic categories.

Core Principles & Definitions

To evaluate treatment modalities comparatively, clinicians and researchers rely on a set of foundational distinctions that govern how evidence is generated, classified, and applied. The most fundamental distinction is between efficacy — whether a treatment works under ideal, controlled conditions — and effectiveness — whether it works in real-world clinical practice. A treatment may demonstrate strong efficacy in a randomized controlled trial (RCT) using a homogeneous sample and manualized protocols, yet fail to produce comparable results when delivered in community mental health settings with diverse, comorbid populations and variable therapist adherence.

1

Efficacy

Demonstrated through randomized controlled trials (RCTs) with high internal validity. Uses manualized treatments, homogeneous samples, random assignment, and comparison to control conditions (waitlist, placebo, or active treatment).
2

Effectiveness

Evaluated in naturalistic or field settings with high external validity. Examines whether treatments generalize to diverse populations, community clinics, and non-manualized delivery by clinicians with varying training levels.
3

Clinical Significance

Goes beyond statistical significance to ask whether the magnitude of change is meaningful. Jacobson and Truax's reliable change index and movement into a normative range are standard criteria for determining clinically significant improvement.
4

Empirically Supported Treatments (ESTs)

Treatments meeting specific criteria established by the APA Division 12 Task Force: at least two well-designed group studies or a large series of single-case designs demonstrating superiority to placebo or equivalence to an already-established treatment.
5

Evidence-Based Practice (EBP)

A broader framework integrating research evidence, clinical expertise, and patient values/preferences. Unlike ESTs, which focus on specific treatments, EBP encompasses the entire clinical decision-making process.
KEY TAKEAWAY
Think of efficacy and effectiveness like testing a new aircraft engine: efficacy is what happens in the wind tunnel under perfectly controlled conditions (temperature, altitude, fuel grade), while effectiveness is how that engine performs in actual flight — with turbulence, varying weather, different pilots, and real-world maintenance schedules. A treatment that has not been tested in the wind tunnel (no RCT evidence) is suspect, but one that has only been tested there may still fail under real conditions. Comprehensive evaluation requires both.

Visual Explanation — The Efficacy-Effectiveness Continuum

This diagram illustrates the continuum from efficacy research (left, high internal validity) through benchmarking studies to effectiveness research (right, high external validity). At the bottom, evidence-based practice integrates all three pillars: Research (R), Expertise (E), and Patient context (P).

The diagram above captures a critical nuance for EPPP preparation: efficacy and effectiveness are not competing paradigms but complementary endpoints on a research continuum. Benchmarking studies occupy the middle ground, comparing real-world clinical outcomes against published RCT effect sizes to determine whether community-delivered treatments approximate laboratory results. Seligman's 1995 analysis of the Consumer Reports survey exemplifies this approach, arguing that effectiveness research captures variables — such as self-selection into treatment, flexible duration, and patient preference — that RCTs deliberately control away. The comprehensive framework of evidence-based practice synthesizes these streams, demanding that clinicians consider the totality of evidence alongside their own clinical judgment and the unique circumstances of each client.

Research Methodology & Quantitative Framework

Comparative treatment evaluation relies on a quantitative framework centered on effect sizes — standardized metrics that allow researchers to compare treatment outcomes across studies, measures, and populations. Unlike p-values, which merely indicate whether a difference is unlikely to have occurred by chance, effect sizes communicate the magnitude and practical significance of treatment differences. Three key quantitative constructs dominate the literature: Cohen's d for between-group comparisons, the number needed to treat (NNT) for clinical translation, and the reliable change index (RCI) for individual-level clinical significance.

COHEN'S D — BETWEEN-GROUP EFFECT SIZE
d = (M₁ − M₂) / S_pooled
Where M₁ = mean of treatment group, M₂ = mean of control group, and S_pooled = pooled standard deviation. Cohen's conventions: d = 0.2 (small), d = 0.5 (medium), d = 0.8 (large). Smith & Glass's landmark meta-analysis found a mean d ≈ 0.85 for psychotherapy vs. no treatment.
NUMBER NEEDED TO TREAT (NNT)
NNT = 1 / (CER − EER) or equivalently 1 / ARR
Where CER = control event rate (proportion not improving without treatment), EER = experimental event rate (proportion not improving with treatment), and ARR = absolute risk reduction. An NNT of 3 means that for every 3 patients treated, 1 additional patient improves beyond what would have occurred without treatment. Lower NNTs indicate greater clinical utility.
RELIABLE CHANGE INDEX (RCI)
RCI = (X₂ − X₁) / S_diff where S_diff = √(2 × (S₁ × √(1 − r_xx))²)
Where X₁ = pretest score, X₂ = posttest score, S₁ = standard deviation of the measure at pretest, and r_xx = test-retest reliability. An RCI > 1.96 indicates statistically reliable change at p < .05. Jacobson and Truax (1991) proposed that clinically significant change requires both reliable change AND movement into a normative range.
📝 EPPP Tip
The EPPP frequently tests the distinction between statistical significance and clinical significance. A treatment may produce a statistically significant effect (p < .05) with a small effect size (d = 0.15), meaning the difference, while real, is too small to matter clinically. Conversely, a study may lack statistical power to detect a medium effect size (d = 0.50) that would be clinically meaningful. Always consider effect size magnitude alongside significance testing.

Comparative Findings Across Treatment Modalities

One of the most debated findings in psychotherapy research is the Dodo Bird verdict — named after the Dodo Bird in Alice in Wonderland who declared 'Everybody has won, and all must have prizes.' Proposed by Rosenzweig (1936) and empirically supported by Luborsky, Singer, and Luborsky (1975), this verdict suggests that different bona fide psychotherapies produce roughly equivalent outcomes. Wampold's (2001) meta-analytic work estimated that specific treatment factors account for only about 1% of outcome variance, while common factors — such as the therapeutic alliance, empathy, and positive expectations — account for substantially more. However, this conclusion remains contested, particularly for specific disorders where targeted interventions have demonstrated clear superiority.

Left panel: Lambert's (2013) model of outcome variance, showing that client factors and common factors account for the majority of outcome. Right panel: Notable disorders where specific treatments have demonstrated superiority, challenging the Dodo Bird verdict. Bottom: Key meta-analytic effect size benchmarks.

The diagram above highlights a central tension in comparative treatment research. On one hand, Lambert's model and Wampold's meta-analyses support the view that common factors drive the lion's share of therapeutic change. On the other hand, the right panel identifies specific disorders — particularly anxiety disorders and personality pathology — where targeted interventions have demonstrated clear advantages. For EPPP preparation, it is essential to hold both perspectives simultaneously: the general trend favors equivalence among bona fide treatments, but clinically important exceptions exist, and the responsible clinician must know which treatments have the strongest disorder-specific evidence bases.

Summary of disorder-specific best-supported treatments and associated effect sizes
DisorderBest-Supported Treatment(s)Key Effect Size
Major Depressive DisorderCBT, IPT, Behavioral Activation, Antidepressants; roughly equivalentd ≈ 0.60–0.80 vs. control
Generalized Anxiety DisorderCBT (applied relaxation, cognitive restructuring)d ≈ 0.80–1.00 vs. control
PTSDProlonged Exposure (PE), CPT, EMDRd ≈ 1.00–1.50 vs. waitlist
OCDExposure and Response Prevention (ERP)d ≈ 1.00–1.50 vs. control
Borderline Personality DisorderDBT, MBT, TFP, Schema Therapyd ≈ 0.50–0.80 vs. TAU
Substance Use DisordersMI, CBT, CRA, Contingency Managementd ≈ 0.30–0.60 vs. control

Worked Example — Evaluating Comparative Treatment Evidence

Consider the following scenario: A clinical researcher conducts an RCT comparing Cognitive-Behavioral Therapy (CBT) and Interpersonal Therapy (IPT) for moderate depression. The study includes 120 participants randomly assigned to CBT (n = 60) or IPT (n = 60). At post-treatment, the CBT group shows a mean BDI-II score of 12.4 (SD = 6.8), and the IPT group shows a mean of 14.1 (SD = 7.2). We will walk through how to compute the between-group effect size, interpret its magnitude, and assess clinical significance.

Comparing CBT vs. IPT for Depression
1
Step 1 — Identify the Relevant ValuesCBT group: M₁ = 12.4, SD₁ = 6.8, n₁ = 60. IPT group: M₂ = 14.1, SD₂ = 7.2, n₂ = 60. We are computing Cohen's d to determine the magnitude of difference between the two active treatments.
2
Step 2 — Compute the Pooled Standard DeviationS_pooled = √[((n₁ − 1) × SD₁² + (n₂ − 1) × SD₂²) / (n₁ + n₂ − 2)] = √[((59 × 46.24) + (59 × 51.84)) / 118] = √[(2728.16 + 3058.56) / 118] = √[5786.72 / 118] = √49.04 ≈ 7.00
S_pooled ≈ 7.00
3
Step 3 — Calculate Cohen's dd = (M₁ − M₂) / S_pooled = (12.4 − 14.1) / 7.00 = −1.7 / 7.00 ≈ −0.24. The negative sign indicates CBT produced slightly lower (better) BDI-II scores, but the absolute magnitude is small.
d ≈ −0.24 (small effect)
4
Step 4 — Interpret the Effect SizeUsing Cohen's conventions, d = 0.24 falls in the small range (0.20–0.49). This suggests that CBT and IPT produce comparable outcomes for moderate depression — a finding consistent with the broader literature and the Dodo Bird verdict for depression treatment. The practical difference between the two groups is approximately 1.7 points on the BDI-II, which is well below the minimum clinically important difference (typically ≈ 5 points).
5
Step 5 — Assess Clinical SignificanceEven though the between-group comparison yields a small effect, we should also examine within-group change. If baseline BDI-II means were approximately 28 (moderate depression) and the BDI-II has test-retest reliability of 0.93 with SD ≈ 9.5, we can compute the RCI for individual patients. S_diff = √(2 × (9.5 × √(1 − 0.93))²) = √(2 × (9.5 × 0.265)²) = √(2 × 6.33) ≈ 3.56. A patient dropping from 28 to 12 shows a change of 16 points; RCI = 16 / 3.56 ≈ 4.49, which exceeds 1.96 and indicates reliable change. If 12 also falls within the normative range (BDI-II < 14), the change is clinically significant.
Conclusion: CBT and IPT are comparably effective for moderate depression; both produce clinically significant within-group change.

Strengths & Limitations of Comparative Methods

Comparison of major research designs for evaluating treatment modalities
Research MethodStrengthsLimitations
Randomized Controlled Trials (RCTs)Gold standard for causal inference; controls for confounds via randomization; permits standardized comparison with control conditions; replicable through manualized protocolsLimited external validity; excludes comorbid/complex cases; allegiance effects bias results toward researcher-preferred treatments; demand characteristics; high cost
Meta-AnalysisAggregates effect sizes across studies; increases statistical power; identifies moderators; provides quantitative summary of evidence baseGarbage in, garbage out — quality depends on included studies; publication bias inflates effects; heterogeneity may mask important differences; coding decisions are subjective
Effectiveness StudiesHigh ecological validity; diverse samples; naturalistic delivery; captures real-world moderators like patient preference and therapist flexibilityWeaker internal validity; selection bias; unmeasured confounds; harder to attribute causation; variable treatment fidelity
Dismantling StudiesIdentifies active ingredients of treatment; tests necessity of specific components; informs treatment refinement and efficiencyRequires very large samples for adequate power; may miss synergistic effects between components; components may not function independently
Process-Outcome ResearchIlluminates mechanisms of change; identifies mediators and moderators; links session-level processes to outcomes; informs treatment developmentCorrelational designs cannot establish causation; temporal precedence issues; measurement of in-session processes is complex and often retrospective
KEY TAKEAWAY
No single research design can answer every question about treatment efficacy and effectiveness. Think of evaluating treatments like evaluating a car: an RCT is the crash test — highly controlled and essential for safety certification. An effectiveness study is the long-term consumer report — how does the car actually perform after 100,000 miles with real drivers? A dismantling study takes the engine apart to find which components are essential. And meta-analysis aggregates all test results across models and years. Each method contributes a unique piece of the puzzle, and the strongest conclusions come from convergent evidence across multiple designs.

A crucial methodological concern in comparative treatment research is researcher allegiance. Luborsky and colleagues (1999) found that the theoretical orientation of the researcher was a strong predictor of which treatment 'won' in comparative trials — a finding that has been replicated and extended. Allegiance effects can operate through subtle mechanisms: choice of comparison conditions, selection of outcome measures that favor the preferred treatment, differential enthusiasm in training therapists, and selective reporting of results. When evaluating EPPP items about comparative treatment research, always consider whether the cited study controlled for researcher allegiance and whether comparison conditions were both bona fide treatments — that is, treatments delivered with genuine therapeutic intent by trained clinicians, not straw-man comparisons designed to fail.

Connection to Advanced Theory & Emerging Paradigms

The field of comparative treatment evaluation is evolving beyond the simple question of 'which therapy wins?' toward more nuanced questions about mechanisms, moderators, and personalization. Three advanced paradigms are reshaping how clinicians and researchers think about treatment comparison: the common factors model, the specific factors/EST model, and the emerging personalized treatment selection model. Understanding the tensions and complementarities among these perspectives is essential for advanced EPPP preparation and for the future of clinical practice.

Three competing/complementary paradigms for comparative treatment evaluation
FeatureCommon Factors ModelEST/Specific Factors ModelPersonalized Treatment Selection
Core QuestionWhat shared processes drive change across all therapies?Which specific treatments work best for specific disorders?Which treatment works best for this individual patient?
Key AdvocatesWampold, Norcross, LambertChambless, Barlow, HofmannDeRubeis, Cohen, Zilcha-Mano
Primary EvidenceMeta-analyses showing small between-treatment differences; alliance-outcome correlationsRCTs showing specific treatments outperform controls for specific disordersPatient-by-treatment interaction analyses; machine learning prediction models
Clinical ImplicationPrioritize therapeutic relationship and therapist qualities; adapt to patient preferencesMatch treatment to diagnosis; train clinicians in ESTs for specific populationsUse patient characteristics (not just diagnosis) to select optimal treatment from among ESTs
LimitationMay undervalue genuine technique-specific effects; less prescriptive for trainingMay overemphasize diagnosis as organizing principle; ignores patient preference dataCurrently limited by sample sizes and replication challenges; not yet ready for routine clinical use

The personalized treatment selection paradigm represents the cutting edge of the field. DeRubeis and colleagues have demonstrated that even in studies where two treatments produce equivalent average outcomes, individual patients often show large differential responses — that is, Patient A might respond much better to CBT while Patient B responds better to IPT. The Personalized Advantage Index (PAI) uses baseline patient characteristics (severity, comorbidity, personality, cognitive style) to predict which treatment a specific individual would benefit from most. This approach moves the field from 'what works in general?' to 'what works for whom?' — a question Gordon Paul articulated in 1967 and that researchers are only now developing the statistical tools to address rigorously.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher wants to determine whether cognitive-behavioral therapy works under ideal conditions with a carefully selected, diagnostically homogeneous sample using a manualized protocol. Which type of study is the researcher conducting — an efficacy study or an effectiveness study? Explain the defining features that distinguish these two types of research.
PROBLEM 2BASIC CALCULATION
In a meta-analysis comparing two bona fide treatments for social anxiety disorder, the overall between-treatment effect size is d = 0.15. Using Cohen's conventions (0.2 = small, 0.5 = medium, 0.8 = large), classify this effect size and explain what it suggests about the comparative efficacy of the two treatments.
PROBLEM 3INTERMEDIATE
A dismantling study of a multicomponent CBT protocol for panic disorder compares (a) full CBT (cognitive restructuring + interoceptive exposure + in vivo exposure), (b) cognitive restructuring alone, and (c) interoceptive exposure alone. The full package produces d = 1.20 vs. waitlist; cognitive restructuring alone produces d = 0.60; interoceptive exposure alone produces d = 0.95. What conclusions can be drawn about the active ingredients of the treatment? What are the limitations of this design?
PROBLEM 4APPLIED
You are a clinical psychologist in a community mental health center. A patient with moderate depression, comorbid generalized anxiety, and a strong preference for a structured, skills-based approach presents for treatment. Multiple RCTs support both CBT and IPT for depression, with similar effect sizes. The patient has limited insurance coverage allowing 12 sessions. Using the evidence-based practice framework (integration of research, clinical expertise, and patient values), describe how you would select a treatment approach and justify your decision.
PROBLEM 5CRITICAL THINKING
Critically evaluate the following argument: 'Because meta-analyses consistently show that the effect size difference between bona fide psychotherapies is near zero (d ≈ 0.00–0.20), there is no scientific basis for preferring one treatment over another for any disorder. Training clinicians in specific ESTs is therefore unnecessary, and resources should instead be directed entirely toward enhancing common factors such as the therapeutic alliance.' Identify the logical and empirical weaknesses in this argument.

Lesson Summary

Evaluating comparative treatment efficacy and effectiveness requires understanding a fundamental distinction: efficacy refers to whether a treatment works under controlled RCT conditions (high internal validity), while effectiveness refers to whether it works in real-world clinical settings (high external validity). Effect sizes (especially Cohen's d) are essential for quantifying the magnitude of treatment differences, and the reliable change index (RCI) and Jacobson-Truax criteria determine whether individual change is clinically meaningful. The Dodo Bird verdict — the finding that bona fide therapies produce roughly equivalent outcomes — is broadly supported but has important exceptions for disorders like OCD, specific phobias, and panic disorder, where exposure-based treatments demonstrate clear superiority.

Three paradigms compete and complement one another: the common factors model (emphasizing alliance, empathy, and expectancy), the empirically supported treatments (EST) model (matching specific treatments to diagnoses), and the emerging personalized treatment selection approach (using patient characteristics to predict differential treatment response). The APA's evidence-based practice framework integrates the best available research with clinical expertise and patient values — providing a comprehensive decision-making structure that transcends any single paradigm. For the EPPP, remember that researcher allegiance is a significant confound in comparative trials, and that the strongest conclusions emerge from convergent evidence across multiple research designs — RCTs, meta-analyses, effectiveness studies, and dismantling studies each contributing unique and essential evidence.

Varsity Tutors • EPPP: Part 1, Knowledge • Treatment Efficacy — Evaluate comparative efficacy and effectiveness of treatment modalities