EPPP: PART 1, KNOWLEDGE • DOMAIN 7: RESEARCH METHODS AND STATISTICS

Statistical Interpretation — Interpret effect size, power, clinical vs statistical significance

Understanding how to move beyond p-values to judge whether research findings truly matter for clinical practice.

Historical Context & Motivation

For much of the twentieth century, behavioral and social science research was dominated by a single question: Is the result statistically significant? Researchers routinely relied on the p-value as the sole arbiter of whether a finding mattered. This practice led to widespread misinterpretation: studies with massive sample sizes could yield trivially small differences that were nonetheless "significant" at p < .05, while clinically meaningful effects were dismissed when samples were too small to achieve that threshold. The history of effect size, statistical power, and the clinical-versus-statistical significance debate reveals how the field gradually learned that a p-value alone is never enough to guide clinical decision-making.

1925
Fisher's Null Hypothesis Framework
Ronald Fisher formalized the concept of the p-value and the null hypothesis significance test (NHST), establishing a framework that would dominate behavioral science for decades. Fisher viewed p-values as continuous measures of evidence, not rigid pass/fail thresholds.
1933
Neyman–Pearson Decision Theory
Jerzy Neyman and Egon Pearson introduced the concepts of Type I error (α) and Type II error (β), along with the notion of statistical power (1 − β). Their framework framed hypothesis testing as a decision procedure with specifiable error rates.
1962
Cohen's Seminal Power Analysis
Jacob Cohen published a landmark review showing that the average study in behavioral science was severely underpowered—often with only a 50% chance of detecting a medium effect. His work on effect size conventions (small, medium, large) transformed how researchers plan and interpret studies.
1994
APA Task Force on Statistical Inference
The American Psychological Association convened a task force recommending that researchers always report effect sizes and confidence intervals alongside p-values, acknowledging the insufficiency of NHST alone.
2010s
Replication Crisis & Reform
High-profile failures to replicate major findings in psychology intensified calls for reporting effect sizes, ensuring adequate power, and distinguishing between statistical significance and real-world clinical importance. The 6th and 7th editions of the APA Publication Manual now mandate effect size reporting.

The central question that emerges from this history is deceptively simple: How do we know whether a statistically significant finding actually matters for the clients and patients we serve? Answering this question requires an understanding of effect size, statistical power, and the crucial distinction between statistical significance and clinical significance—concepts that form the backbone of evidence-based practice in behavioral health.

Core Principles & Definitions

Before diving into the mathematics and applications, it is essential to establish clear definitions of the four interrelated concepts that anchor statistical interpretation in behavioral health research. Each concept addresses a different aspect of research findings: how big an effect is, how likely we are to detect it, whether a statistical test flags it, and whether it translates to meaningful change in clinical practice.

1

Effect Size

A standardized, scale-free metric that quantifies the magnitude of a treatment effect or relationship between variables. Unlike p-values, effect sizes tell you how much difference exists—not just whether a difference exists. Common measures include Cohen's d, Pearson's r, and η² (eta squared).
2

Statistical Power (1 − β)

The probability that a study will correctly reject the null hypothesis when a true effect exists. A study with low power is like trying to hear a whisper in a noisy room—the signal may be real, but you cannot detect it. Cohen recommended a minimum power of .80.
3

Statistical Significance

A finding is statistically significant when the observed data are sufficiently unlikely under the null hypothesis (typically p < .05). This means there is less than a 5% probability of obtaining the result by chance alone, assuming the null is true. It does not tell you the size or importance of the effect.
4

Clinical Significance

A finding is clinically significant when the observed change is large enough to be meaningful in real-world practice—for example, when a client moves from a clinical range to a normal range on a symptom measure. Jacobson and Truax (1991) operationalized this with the Reliable Change Index.
KEY TAKEAWAY
Think of these four concepts like evaluating a new medication. Statistical significance tells you that the drug did something beyond placebo. Effect size tells you how much it helped. Power tells you whether your study was even capable of detecting the drug's benefit. And clinical significance tells you whether patients actually feel better in their daily lives. A responsible clinician needs all four pieces of information.

Visual Explanation — Effect Size and Statistical vs. Clinical Significance

One of the most intuitive ways to understand effect size is to visualize the overlap between two distributions—for example, a treatment group and a control group on a symptom measure. When the effect size is zero, the two distributions overlap completely. As effect size increases, the distributions pull apart. The following diagram illustrates small, medium, and large effects as defined by Cohen's conventions, showing how the degree of separation maps onto the practical impact of a treatment.

Top row: As Cohen's d increases from 0.2 to 0.8, the two group distributions separate, reducing overlap. Bottom panel: Statistical significance (left) addresses whether an effect is real, while clinical significance (right) addresses whether that effect matters for patients.

As shown in the diagram, a small effect size (d = 0.2) produces approximately 85% overlap between distributions—meaning that most individuals in the treatment and control groups are indistinguishable from one another. A medium effect (d = 0.5) reduces overlap to about 67%, and a large effect (d = 0.8) further reduces it to 53%. In psychotherapy outcome research, effect sizes of 0.5 to 0.8 are common for established treatments of depression and anxiety, suggesting that effective therapies produce visible, though not absolute, separation between treated and untreated groups.

Mathematical Framework

Each of the core concepts in statistical interpretation has an associated mathematical formulation. Understanding these formulas is critical not only for calculating and interpreting research findings but also for the EPPP, where questions often require you to interpret relationships among effect size, sample size, power, and significance.

Cohen's d — Standardized Mean Difference

COHEN'S D
d = (M₁ − M₂) / S_pooled
where M₁ = mean of the treatment group, M₂ = mean of the control group, and Spooled = pooled standard deviation of both groups. Conventions: d = 0.2 (small), 0.5 (medium), 0.8 (large).

Eta Squared (η²) — Proportion of Variance Explained

ETA SQUARED
η² = SS_between / SS_total
where SSbetween = sum of squares between groups and SStotal = total sum of squares. Conventions: η² = .01 (small), .06 (medium), .14 (large). Commonly reported alongside ANOVA results.

Statistical Power

POWER
Power = 1 − β = P(reject H₀ | H₀ is false)
where β = probability of a Type II error (failing to detect a true effect). Power is determined by four factors: sample size (n), effect size (d), significance level (α), and directionality of the test (one-tailed vs. two-tailed). The conventional minimum acceptable power is .80.

Reliable Change Index (RCI)

RELIABLE CHANGE INDEX
RCI = (X₂ − X₁) / S_diff
where X₁ = pretreatment score, X₂ = posttreatment score, and Sdiff = standard error of the difference between two scores. If |RCI| > 1.96, the change exceeds measurement error and is considered reliable. This is the Jacobson-Truax method for assessing clinical significance.
📌 EPPP Exam Tip
Remember the four factors that influence power: sample size, effect size, alpha level, and directionality. Increasing sample size or effect size increases power. Using a more lenient alpha (e.g., .10 instead of .05) or a one-tailed test also increases power, but at the cost of increased Type I error risk.

Effect Size Measures — Classification and Interpretation

Effect sizes come in two broad families: the d-family (standardized mean differences) and the r-family (measures of association or proportion of variance explained). Different research designs call for different effect size metrics, and converting between families is sometimes necessary when conducting meta-analyses. The table below summarizes the most frequently encountered measures and their conventional benchmarks.

Common effect size measures and Cohen's conventional benchmarks
MeasureFamilySmallMediumLargeCommon Use
Cohen's dd-family0.20.50.8Two-group comparisons (t-tests)
Hedges' gd-family0.20.50.8Meta-analysis (corrects for small-sample bias)
Pearson's rr-family.10.30.50Correlational studies
r² / R²r-family.01.09.25Regression — proportion of variance explained
η² (eta squared)r-family.01.06.14ANOVA — proportion of variance
Odds Ratio (OR)Other1.52.54.3Logistic regression, case-control designs
The four determinants of statistical power converge on the probability of detecting a true effect. Increasing sample size or effect size improves power directly, while relaxing alpha or using a one-tailed test also increases power but introduces trade-offs.

A critical point for EPPP preparation is that Cohen's benchmarks are general guidelines, not universal standards. Cohen himself cautioned that "small," "medium," and "large" depend on context. In some areas of behavioral health, a small effect size may represent a clinically important change—for instance, a d of 0.2 in a suicide prevention intervention might save a meaningful number of lives at scale. Conversely, a large effect size might not be clinically significant if the outcome measure does not map onto functional improvement. Always interpret effect sizes in the context of the specific clinical domain and population.

Worked Example — Interpreting a CBT Outcome Study

A researcher conducts a randomized controlled trial comparing cognitive-behavioral therapy (CBT) for generalized anxiety disorder (GAD) against a waitlist control. After 12 weeks, the CBT group (n = 40) has a mean GAD-7 score of 7.2 (SD = 3.8), while the waitlist group (n = 40) has a mean of 11.6 (SD = 4.2). The independent samples t-test yields t(78) = 4.92, p < .001. The GAD-7 clinical cutoff is 10 (scores ≥ 10 indicate moderate anxiety). The researcher wants to determine the effect size, evaluate power, and assess both statistical and clinical significance.

Interpreting a CBT Outcome Study for GAD
1
Step 1 — Calculate the Pooled Standard DeviationThe pooled standard deviation combines the variability from both groups. With equal sample sizes, the formula simplifies to: Spooled = √[(SD₁² + SD₂²) / 2] = √[(3.8² + 4.2²) / 2] = √[(14.44 + 17.64) / 2] = √[16.04] ≈ 4.005.
S_pooled ≈ 4.01
2
Step 2 — Compute Cohen's dApply the formula: d = (M₁ − M₂) / Spooled = (11.6 − 7.2) / 4.01 = 4.4 / 4.01 ≈ 1.10. Note that we place the larger mean first so that the effect size is positive, indicating the direction of improvement.
d ≈ 1.10 — a large effect by Cohen's conventions
3
Step 3 — Assess Statistical SignificanceThe t-test yielded p < .001, which is well below the conventional alpha of .05. We reject the null hypothesis and conclude that the difference between the CBT and waitlist groups is statistically significant—extremely unlikely to be due to chance alone.
Statistically significant: p < .001
4
Step 4 — Evaluate Power (Post Hoc)Using a power table or calculator with d = 1.10, α = .05 (two-tailed), and n = 40 per group, the estimated power exceeds .99. This means there was a greater than 99% probability of detecting an effect this large, confirming the study was adequately powered. Had the researcher conducted an a priori power analysis for a medium effect (d = 0.5) with power = .80, the required sample size would have been approximately 64 per group.
Power > .99 — adequately powered study
5
Step 5 — Determine Clinical SignificanceThe GAD-7 clinical cutoff is 10. The CBT group's mean (7.2) fell below this threshold, placing the average treated patient in the non-clinical range, while the waitlist mean (11.6) remained above the cutoff. Using the Jacobson-Truax method, the clinician would also calculate the Reliable Change Index for individual patients to determine how many showed reliable improvement that exceeds measurement error. The combination of a large effect size and movement below the clinical cutoff provides strong evidence that this treatment produces clinically significant change.
Clinically significant: CBT group mean is below clinical cutoff

Strengths and Limitations of Each Interpretation

No single metric provides a complete picture of a study's findings. Each interpretive tool has strengths and limitations that must be understood in order to apply them judiciously in clinical practice and on the EPPP. The following table highlights the key trade-offs.

Strengths and limitations of each statistical interpretation tool
ConceptStrengthsLimitations
p-value (Statistical Significance)Provides objective criterion; universally understood; controls Type I error rateConfounded by sample size; says nothing about magnitude; dichotomous (sig/not sig) thinking; often misinterpreted as P(H₀ is true)
Effect SizeScale-free; comparable across studies; essential for meta-analysis; quantifies magnitudeConventions are arbitrary; can be inflated by restriction of range or unreliable measures; not interpretable without context
Statistical PowerInforms study design; prevents wasted resources on underpowered studies; links effect size to sample size planningPost hoc power analysis is controversial (circular reasoning when based on observed effect); requires a priori effect size estimate that may be inaccurate
Clinical SignificanceDirectly relevant to patient outcomes; individual-level assessment; bridges research and practiceRequires validated norms and cutoffs; varies by instrument; RCI depends on test-retest reliability; no single agreed-upon method
KEY TAKEAWAY
Think of these tools as different instruments in a diagnostic workup. A thermometer (p-value) can tell you that a patient's temperature deviates from normal, but it cannot tell you why or how sick the patient is. Effect size functions like a blood panel—it quantifies the extent of the problem. Power analysis is like ensuring the lab equipment is sensitive enough to pick up the relevant markers. And clinical significance is the physician's judgment: given all the data, is this patient actually getting better? You would never base a treatment decision on just one instrument.

Connections to Advanced Theory and Evidence-Based Practice

The concepts covered in this lesson connect directly to several advanced topics that appear on the EPPP and are central to modern behavioral health practice. Understanding how effect size, power, and clinical significance extend into these domains deepens your interpretive skill set and prepares you for complex exam questions that integrate across knowledge areas.

How foundational statistical interpretation concepts extend to advanced methods
This Lesson's ConceptAdvanced ExtensionEPPP Relevance
Effect size (d, r, η²)Meta-analysis — aggregating effect sizes across studies to estimate a population-level effect with greater precisionEPPP may ask about heterogeneity (Q statistic, I²), fixed vs. random effects models, and forest plots
Statistical powerA priori sample size calculation — using G*Power or similar tools to determine required n before data collection beginsExpect questions about how changing parameters (n, α, d, tails) affects power; know the direction of each relationship
Clinical significance (RCI)Evidence-based practice (EBP) — integrating research evidence, clinical expertise, and patient preferences, where clinical significance guides treatment decisionsQuestions may connect to APA's EBP framework, especially interpreting treatment outcome research
p-value interpretationConfidence intervals & Bayesian methods — alternatives to dichotomous NHST that provide richer information about effect magnitude and uncertaintyKnow that CIs around effect sizes are increasingly recommended by APA; understand that CI width reflects precision

The trajectory of the field is clear: behavioral health is moving away from sole reliance on p-values toward a multimethod approach to statistical interpretation. The APA Publication Manual (7th ed.) requires effect sizes and encourages confidence intervals, and funding agencies like NIMH increasingly expect power analyses in grant proposals. For clinical psychologists, this means that competent consumption of research—and competent practice—demands fluency in all four concepts presented in this lesson.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher finds that a new psychotherapy technique for PTSD produces a statistically significant reduction in symptoms compared to a control group (p = .03), but the effect size is d = 0.15. The researcher concludes that the therapy is highly effective. Evaluate this conclusion and explain why it may be misleading.
PROBLEM 2BASIC CALCULATION
In a study comparing a mindfulness intervention to a control condition on a depression measure, the treatment group (n = 30) has M = 12.4, SD = 5.0, and the control group (n = 30) has M = 16.8, SD = 5.4. Calculate Cohen's d and interpret the result using Cohen's conventions.
PROBLEM 3INTERMEDIATE
A clinical psychologist is planning a study to evaluate a new group therapy for social anxiety. Based on prior research, she expects a medium effect size (d = 0.5). She plans to use α = .05, two-tailed. If she wants power of .80, approximately how many participants does she need per group? If she can only recruit 30 per group, what happens to power, and what are the implications?
PROBLEM 4APPLIED
A patient enters treatment for panic disorder with a pre-treatment score of 28 on the Panic Disorder Severity Scale (PDSS). After 16 sessions of CBT, her post-treatment score is 9. The PDSS has a test-retest reliability of .87 and a standard deviation (based on the normative clinical sample) of 6.5. The clinical cutoff for the PDSS is 12 (scores below 12 are in the non-clinical range). Calculate the RCI and determine whether this patient shows both reliable change and clinically significant change.
PROBLEM 5CRITICAL THINKING
Consider two studies on the same treatment for alcohol use disorder. Study A (n = 500 per group) finds a statistically significant reduction in drinks per week (p < .001, d = 0.18). Study B (n = 25 per group) finds a non-significant result (p = .12, d = 0.45). Which study provides stronger evidence for the treatment's clinical utility, and why? Discuss how effect size, power, and clinical significance interact in your evaluation.

Lesson Summary

This lesson covered the four pillars of statistical interpretation essential for EPPP preparation and evidence-based behavioral health practice. Effect size quantifies the magnitude of a finding using standardized metrics such as Cohen's d (small = 0.2, medium = 0.5, large = 0.8), Pearson's r, and η². Statistical power (1 − β) is the probability of detecting a true effect and is determined by four factors: sample size, effect size, alpha level, and directionality. The conventional minimum acceptable power is .80.

Statistical significance (typically p < .05) tells us a result is unlikely under the null hypothesis, but it is heavily influenced by sample size and says nothing about practical importance. Clinical significance addresses whether the change is meaningful for patients—operationalized by the Reliable Change Index (RCI) and movement across clinical cutoffs using the Jacobson-Truax method. Competent evidence-based practice requires integrating all four of these perspectives: reporting effect sizes alongside p-values, conducting a priori power analyses to design adequate studies, and evaluating whether statistically significant findings translate to meaningful patient improvement.

Varsity Tutors • EPPP: Part 1, Knowledge • Statistical Interpretation — Interpret effect size, power, clinical vs statistical significance