Historical Context & Motivation
Throughout the twentieth century, behavioral health researchers confronted a persistent challenge: individual studies often produced conflicting results, psychological constructs proved difficult to measure directly, and claims about causation remained contentious in non-experimental settings. Three families of statistical techniques emerged to address these problems. Meta-analysis provided a systematic way to aggregate findings across studies and estimate the true magnitude of an effect. Factor analysis offered a method for identifying the latent dimensions underlying batteries of observed measures—revealing, for example, that dozens of personality items could be distilled into a smaller number of coherent traits. Causal modeling (including path analysis and structural equation modeling) allowed investigators to specify and test theoretical networks of direct and indirect effects, moving beyond simple correlational statements toward rigorous causal inference.
Each of these techniques addresses a fundamental gap in the behavioral health evidence base. How large is an effect really, when dozens of studies disagree? What latent constructs account for the covariation among observed symptoms? And can we test whether a theorized chain of causation—stress → coping failure → depression—is consistent with the data? The sections that follow equip you to interpret findings from all three analytic families with the precision expected on the EPPP.
Core Principles & Definitions
Before diving into formulas and output tables, it is essential to grasp the conceptual architecture that underpins each technique. Although meta-analysis, factor analysis, and causal modeling serve different purposes, they share a common logic: each transforms observed data into estimates of quantities that cannot be measured directly—aggregate effect sizes, latent factors, or causal path coefficients. Understanding five foundational principles will anchor your interpretation of results across all three methods.
Effect Size as a Common Metric
Latent vs. Observed Variables
Heterogeneity & Moderators
Model Fit & Parsimony
Causal Inference Requires Theory
Visual Explanation — How the Three Methods Relate
The following diagram illustrates the conceptual flow of each analytic family—from its input data, through its core operation, to its primary output. Notice how meta-analysis begins with multiple study-level results, factor analysis begins with a correlation matrix among observed items, and causal modeling begins with a theorized path diagram plus raw or summarized data. Each technique ultimately produces a different kind of interpretive product: an aggregate effect estimate, a set of latent dimensions, or a system of causal coefficients.
In the diagram, dashed lines converging on the SEM box illustrate that structural equation modeling is not a separate technique so much as a superordinate framework. When you encounter SEM in the literature, recognize that its measurement model is essentially a confirmatory factor analysis, while its structural model is essentially a path analysis. The EPPP expects you to understand this integration and to interpret SEM output accordingly.
Mathematical Framework
Meta-Analysis: Weighted Mean Effect Size
The core computation in meta-analysis is a weighted average of study-level effect sizes, where each study's weight reflects the precision of its estimate (typically the inverse of its variance). Larger, more precise studies exert greater influence on the pooled estimate. Two weighting models are commonly used: the fixed-effect model (assumes one true effect) and the random-effects model (assumes effects vary across studies due to real differences in populations, treatments, or contexts).
Factor Analysis: The Factor Model
Factor analysis posits that each observed variable is a linear combination of a smaller number of latent factors plus unique variance (error). The central equation expresses an observed score as the sum of contributions from each factor, weighted by the variable's loading on that factor.
Causal Modeling: Path Coefficients
Detailed Breakdown — EFA vs. CFA and Fit Indices
A critical distinction for the EPPP is the difference between exploratory factor analysis (EFA) and confirmatory factor analysis (CFA). EFA is used when the researcher has no strong prior theory about factor structure; the algorithm freely assigns loadings and the investigator interprets the resulting pattern. CFA, by contrast, is a hypothesis-testing procedure in which the researcher specifies which items load on which factors in advance and then evaluates how well that a priori model fits the data. Because CFA embeds within the SEM framework, it produces formal fit indices.
| Fit Index | Good Fit Cutoff | What It Measures |
|---|---|---|
| χ² (Chi-square) | Non-significant p (ideally) | Exact fit between model-implied and observed covariance matrices. Sensitive to sample size—often significant even with minor misfit in large samples. |
| CFI | ≥ .95 | Comparative Fit Index: compares the hypothesized model to a baseline (null) model. Ranges 0–1; higher is better. |
| RMSEA | ≤ .06 | Root Mean Square Error of Approximation: penalizes model complexity. Lower is better; values > .10 indicate poor fit. |
| SRMR | ≤ .08 | Standardized Root Mean Square Residual: average discrepancy between observed and predicted correlations. |
| TLI (NNFI) | ≥ .95 | Tucker-Lewis Index: similar to CFI but adjusts for parsimony. Can exceed 1.0 in over-fitting models. |
Worked Example — Interpreting a Meta-Analysis of CBT for Depression
Suppose a meta-analysis of 20 randomized controlled trials compares cognitive-behavioral therapy (CBT) to wait-list control for major depressive disorder. The forest plot and summary statistics are provided. Walk through interpretation step by step.
Strengths, Limitations, and Comparisons
Each technique carries distinctive strengths and vulnerabilities. Recognizing these is essential for critically evaluating published research and for selecting the appropriate method when designing your own studies. The table below organizes the key considerations for each analytic family.
| Technique | Strengths | Limitations |
|---|---|---|
| Meta-Analysis | Increases statistical power by aggregating samples; provides precise effect-size estimates; identifies moderators of effects; reduces reliance on single studies; promotes evidence-based practice. | Vulnerable to publication bias (file-drawer problem); quality depends on quality of included studies ('garbage in, garbage out'); coding decisions introduce subjectivity; combining heterogeneous studies may obscure meaningful differences. |
| Factor Analysis (EFA) | Reveals hidden structure in complex data sets; reduces dimensionality; aids scale development and construct validation; guides theory building when structure is unknown. | Multiple criteria for number of factors (eigenvalue > 1, scree plot, parallel analysis) can yield different solutions; factor interpretation is subjective; sample size requirements are substantial (typically N ≥ 300 or 10:1 subjects-to-items ratio). |
| Factor Analysis (CFA) | Tests a priori theory about factor structure; provides formal fit indices; allows comparison of competing models; integrates into SEM framework. | Requires strong theory before data collection; χ² is sensitive to sample size; modification indices can lead to data-driven (atheoretical) model changes; requires large samples (N ≥ 200). |
| Causal Modeling (Path/SEM) | Tests complex multivariate theories simultaneously; decomposes total effects into direct and indirect components; accounts for measurement error (in SEM); can model latent variables and structural paths together. | Causation is assumed, not proven—cross-sectional data cannot establish temporal precedence; equivalent models may fit equally well; requires large samples and multivariate normality; misspecification can produce biased estimates. |
Connection to Advanced Theory — Modern Extensions
The foundational methods discussed so far have spawned sophisticated extensions that increasingly appear in the behavioral health literature and, occasionally, on the EPPP. Understanding how the basic techniques connect to their advanced counterparts strengthens your interpretive toolkit and prepares you for emerging trends in psychological research.
| Foundational Method | Advanced Extension | Key Difference / Application |
|---|---|---|
| Pairwise meta-analysis | Network meta-analysis (NMA) | Compares multiple treatments simultaneously using both direct and indirect evidence. Useful when no single RCT compares all available interventions (e.g., ranking 8 antidepressants). |
| EFA / CFA (correlated factors) | Bifactor models | Model a general factor (e.g., psychopathology 'p' factor) alongside group-specific factors. Disentangles general from domain-specific variance in symptom measures. |
| Path analysis (observed variables) | Full SEM with latent variables | Corrects for measurement error by using multiple indicators per construct. Provides more accurate path coefficients than path analysis with observed composites. |
| Cross-sectional SEM | Latent growth curve modeling (LGCM) | Models individual trajectories of change over time using repeated measures. Estimates latent intercept (starting level) and slope (rate of change) as latent variables. |
| Traditional causal assumptions | Directed Acyclic Graphs (DAGs) | Graphical framework for formalizing causal assumptions and identifying confounders, colliders, and mediators before running statistical models. Increasingly used in epidemiology and behavioral health. |
For the EPPP, focus primarily on demonstrating competence with the foundational methods and their standard output (effect sizes, factor loadings, path coefficients, fit indices). However, awareness of these extensions—particularly SEM and mediation/moderation analysis—is increasingly expected. The critical interpretive principle remains constant across all these methods: statistical models test the plausibility of theoretical claims, but they do not prove causation in the absence of experimental control and temporal ordering.
Practice Problems
Summary — Advanced Analysis for the EPPP
This lesson covered three indispensable analytic families in behavioral health research. Meta-analysis synthesizes findings across multiple studies by computing a weighted mean effect size, evaluating heterogeneity (Q and I²), conducting moderator analyses when variability is high, and assessing publication bias via funnel plots and statistical tests. Factor analysis identifies latent constructs from observed data: EFA explores structure when theory is underdeveloped, while CFA tests a priori models using fit indices (CFI ≥ .95, RMSEA ≤ .06, SRMR ≤ .08). Key decisions include the number of factors to retain, choice of rotation (orthogonal vs. oblique), and the threshold for meaningful factor loadings (≥ |.40|).
Causal modeling—including path analysis and structural equation modeling (SEM)—tests theorized networks of direct and indirect effects. SEM integrates a measurement model (CFA) with a structural model (path analysis) and evaluates overall model fit. The critical caveat is that causal arrows are theory-driven: good model fit does not prove causation, especially with cross-sectional data. Distinguish mediators (mechanism variables on a causal chain) from moderators (variables that change the strength or direction of an effect). For the EPPP, be prepared to interpret forest plots, evaluate fit indices, decompose direct and indirect effects, and critically appraise causal claims made from statistical models.