Historical Context & Motivation
For much of the twentieth century, behavioral and social science research was dominated by a single question: Is the result statistically significant? Researchers routinely relied on the p-value as the sole arbiter of whether a finding mattered. This practice led to widespread misinterpretation: studies with massive sample sizes could yield trivially small differences that were nonetheless "significant" at p < .05, while clinically meaningful effects were dismissed when samples were too small to achieve that threshold. The history of effect size, statistical power, and the clinical-versus-statistical significance debate reveals how the field gradually learned that a p-value alone is never enough to guide clinical decision-making.
The central question that emerges from this history is deceptively simple: How do we know whether a statistically significant finding actually matters for the clients and patients we serve? Answering this question requires an understanding of effect size, statistical power, and the crucial distinction between statistical significance and clinical significance—concepts that form the backbone of evidence-based practice in behavioral health.
Core Principles & Definitions
Before diving into the mathematics and applications, it is essential to establish clear definitions of the four interrelated concepts that anchor statistical interpretation in behavioral health research. Each concept addresses a different aspect of research findings: how big an effect is, how likely we are to detect it, whether a statistical test flags it, and whether it translates to meaningful change in clinical practice.
Effect Size
Statistical Power (1 − β)
Statistical Significance
Clinical Significance
Visual Explanation — Effect Size and Statistical vs. Clinical Significance
One of the most intuitive ways to understand effect size is to visualize the overlap between two distributions—for example, a treatment group and a control group on a symptom measure. When the effect size is zero, the two distributions overlap completely. As effect size increases, the distributions pull apart. The following diagram illustrates small, medium, and large effects as defined by Cohen's conventions, showing how the degree of separation maps onto the practical impact of a treatment.
As shown in the diagram, a small effect size (d = 0.2) produces approximately 85% overlap between distributions—meaning that most individuals in the treatment and control groups are indistinguishable from one another. A medium effect (d = 0.5) reduces overlap to about 67%, and a large effect (d = 0.8) further reduces it to 53%. In psychotherapy outcome research, effect sizes of 0.5 to 0.8 are common for established treatments of depression and anxiety, suggesting that effective therapies produce visible, though not absolute, separation between treated and untreated groups.
Mathematical Framework
Each of the core concepts in statistical interpretation has an associated mathematical formulation. Understanding these formulas is critical not only for calculating and interpreting research findings but also for the EPPP, where questions often require you to interpret relationships among effect size, sample size, power, and significance.
Cohen's d — Standardized Mean Difference
Eta Squared (η²) — Proportion of Variance Explained
Statistical Power
Reliable Change Index (RCI)
Effect Size Measures — Classification and Interpretation
Effect sizes come in two broad families: the d-family (standardized mean differences) and the r-family (measures of association or proportion of variance explained). Different research designs call for different effect size metrics, and converting between families is sometimes necessary when conducting meta-analyses. The table below summarizes the most frequently encountered measures and their conventional benchmarks.
| Measure | Family | Small | Medium | Large | Common Use |
|---|---|---|---|---|---|
| Cohen's d | d-family | 0.2 | 0.5 | 0.8 | Two-group comparisons (t-tests) |
| Hedges' g | d-family | 0.2 | 0.5 | 0.8 | Meta-analysis (corrects for small-sample bias) |
| Pearson's r | r-family | .10 | .30 | .50 | Correlational studies |
| r² / R² | r-family | .01 | .09 | .25 | Regression — proportion of variance explained |
| η² (eta squared) | r-family | .01 | .06 | .14 | ANOVA — proportion of variance |
| Odds Ratio (OR) | Other | 1.5 | 2.5 | 4.3 | Logistic regression, case-control designs |
A critical point for EPPP preparation is that Cohen's benchmarks are general guidelines, not universal standards. Cohen himself cautioned that "small," "medium," and "large" depend on context. In some areas of behavioral health, a small effect size may represent a clinically important change—for instance, a d of 0.2 in a suicide prevention intervention might save a meaningful number of lives at scale. Conversely, a large effect size might not be clinically significant if the outcome measure does not map onto functional improvement. Always interpret effect sizes in the context of the specific clinical domain and population.
Worked Example — Interpreting a CBT Outcome Study
A researcher conducts a randomized controlled trial comparing cognitive-behavioral therapy (CBT) for generalized anxiety disorder (GAD) against a waitlist control. After 12 weeks, the CBT group (n = 40) has a mean GAD-7 score of 7.2 (SD = 3.8), while the waitlist group (n = 40) has a mean of 11.6 (SD = 4.2). The independent samples t-test yields t(78) = 4.92, p < .001. The GAD-7 clinical cutoff is 10 (scores ≥ 10 indicate moderate anxiety). The researcher wants to determine the effect size, evaluate power, and assess both statistical and clinical significance.
Strengths and Limitations of Each Interpretation
No single metric provides a complete picture of a study's findings. Each interpretive tool has strengths and limitations that must be understood in order to apply them judiciously in clinical practice and on the EPPP. The following table highlights the key trade-offs.
| Concept | Strengths | Limitations |
|---|---|---|
| p-value (Statistical Significance) | Provides objective criterion; universally understood; controls Type I error rate | Confounded by sample size; says nothing about magnitude; dichotomous (sig/not sig) thinking; often misinterpreted as P(H₀ is true) |
| Effect Size | Scale-free; comparable across studies; essential for meta-analysis; quantifies magnitude | Conventions are arbitrary; can be inflated by restriction of range or unreliable measures; not interpretable without context |
| Statistical Power | Informs study design; prevents wasted resources on underpowered studies; links effect size to sample size planning | Post hoc power analysis is controversial (circular reasoning when based on observed effect); requires a priori effect size estimate that may be inaccurate |
| Clinical Significance | Directly relevant to patient outcomes; individual-level assessment; bridges research and practice | Requires validated norms and cutoffs; varies by instrument; RCI depends on test-retest reliability; no single agreed-upon method |
Connections to Advanced Theory and Evidence-Based Practice
The concepts covered in this lesson connect directly to several advanced topics that appear on the EPPP and are central to modern behavioral health practice. Understanding how effect size, power, and clinical significance extend into these domains deepens your interpretive skill set and prepares you for complex exam questions that integrate across knowledge areas.
| This Lesson's Concept | Advanced Extension | EPPP Relevance |
|---|---|---|
| Effect size (d, r, η²) | Meta-analysis — aggregating effect sizes across studies to estimate a population-level effect with greater precision | EPPP may ask about heterogeneity (Q statistic, I²), fixed vs. random effects models, and forest plots |
| Statistical power | A priori sample size calculation — using G*Power or similar tools to determine required n before data collection begins | Expect questions about how changing parameters (n, α, d, tails) affects power; know the direction of each relationship |
| Clinical significance (RCI) | Evidence-based practice (EBP) — integrating research evidence, clinical expertise, and patient preferences, where clinical significance guides treatment decisions | Questions may connect to APA's EBP framework, especially interpreting treatment outcome research |
| p-value interpretation | Confidence intervals & Bayesian methods — alternatives to dichotomous NHST that provide richer information about effect magnitude and uncertainty | Know that CIs around effect sizes are increasingly recommended by APA; understand that CI width reflects precision |
The trajectory of the field is clear: behavioral health is moving away from sole reliance on p-values toward a multimethod approach to statistical interpretation. The APA Publication Manual (7th ed.) requires effect sizes and encourages confidence intervals, and funding agencies like NIMH increasingly expect power analyses in grant proposals. For clinical psychologists, this means that competent consumption of research—and competent practice—demands fluency in all four concepts presented in this lesson.
Practice Problems
Lesson Summary
This lesson covered the four pillars of statistical interpretation essential for EPPP preparation and evidence-based behavioral health practice. Effect size quantifies the magnitude of a finding using standardized metrics such as Cohen's d (small = 0.2, medium = 0.5, large = 0.8), Pearson's r, and η². Statistical power (1 − β) is the probability of detecting a true effect and is determined by four factors: sample size, effect size, alpha level, and directionality. The conventional minimum acceptable power is .80.
Statistical significance (typically p < .05) tells us a result is unlikely under the null hypothesis, but it is heavily influenced by sample size and says nothing about practical importance. Clinical significance addresses whether the change is meaningful for patients—operationalized by the Reliable Change Index (RCI) and movement across clinical cutoffs using the Jacobson-Truax method. Competent evidence-based practice requires integrating all four of these perspectives: reporting effect sizes alongside p-values, conducting a priori power analyses to design adequate studies, and evaluating whether statistically significant findings translate to meaningful patient improvement.