Historical Context & Motivation
Political science has long grappled with a fundamental question: how can scholars move beyond mere correlation to establish that one political phenomenon actually causes another? For much of the discipline's history, researchers relied on observational data—surveys, case studies, and statistical analyses of existing datasets—to draw inferences about political behavior, institutions, and policy outcomes. Yet these approaches consistently confronted the problem of confounding variables: unmeasured or uncontrolled factors that could explain away apparent causal relationships. The migration of experimental methods from the natural sciences and psychology into political science represented a paradigm shift, one that redefined what counts as credible evidence in the study of politics.
The intellectual roots of experimental reasoning trace back to the philosophy of science and the statistical innovations of the early twentieth century. R.A. Fisher's work on randomization in agricultural experiments laid the theoretical groundwork, but it took decades for political scientists to embrace these tools. The credibility revolution in the social sciences—a term popularized by economists Joshua Angrist and Jörn-Steffen Pischke—arrived in political science during the late 1990s and early 2000s, encouraging scholars to think rigorously about identification strategies and causal inference rather than relying solely on regression analysis with observational data.
The central question that experimental designs address is deceptively simple: What would have happened in the absence of a treatment or intervention? This is the counterfactual problem—we can never simultaneously observe a unit in both the treated and untreated state. Different experimental designs offer different strategies for constructing a credible comparison group that approximates this unobservable counterfactual, and the distinctions among true experiments, natural experiments, and quasi-experiments hinge on how successfully each design resolves this fundamental challenge.
Core Principles & Definitions
Before distinguishing among the three major experimental designs, it is essential to establish the conceptual vocabulary that underpins all causal inference in political science. These foundational principles determine how researchers evaluate whether a study's findings genuinely reflect cause and effect, and they provide the criteria by which we judge one design as stronger or weaker than another.
Causal Inference & the Counterfactual
Random Assignment
Internal Validity
External Validity
Treatment Effect
Visual Explanation of Experimental Design Types
The following diagram illustrates the structural logic of the three experimental designs. Each design begins with a population, identifies how units enter treatment and control conditions, and measures outcomes to estimate causal effects. The key differentiator across the three designs is the assignment mechanism—whether it is controlled by the researcher, generated by an external event, or driven by self-selection that must be addressed analytically.
Notice how the diagram reveals a gradient of researcher control from left to right. In a true experiment, the researcher exercises direct control over the assignment mechanism, which is the gold standard for establishing causality. In a natural experiment, the researcher identifies a naturally occurring event—such as a policy change, lottery, or geographic boundary—that generates variation resembling randomization. In a quasi-experiment, there is no randomization at all, and the researcher must rely on design-based strategies and statistical adjustments to construct a credible comparison. This progression from randomized to as-if random to non-random assignment maps directly onto decreasing confidence in internal validity, all else being equal.
The Logic of Causal Identification
While political science is not typically a math-heavy discipline, the formal framework for understanding causal inference provides precision that verbal arguments alone cannot achieve. The potential outcomes framework (also known as the Rubin Causal Model or Neyman-Rubin framework) formalizes the counterfactual logic underlying all experimental designs. For each unit i in a study, we define two potential outcomes: Yi(1) is the outcome if unit i receives treatment, and Yi(0) is the outcome if unit i does not. The fundamental problem of causal inference is that we observe only one of these outcomes for any given unit.
This formal decomposition reveals why experimental designs differ in their inferential power. In a true experiment with random assignment, the selection bias term equals zero by design—treated and control groups are statistically identical before the intervention. In a natural experiment, the researcher argues that an exogenous event made treatment assignment effectively random, so the bias term is approximately zero. In a quasi-experiment, the selection bias term is likely non-zero, so the researcher must employ strategies—matching, regression discontinuity, difference-in-differences, instrumental variables—to either eliminate or bound it analytically.
Detailed Breakdown of Each Design Type
True (Randomized) Experiments
A true experiment (also called a randomized controlled trial or RCT) is the design in which the researcher directly controls the assignment of units to treatment and control conditions through a random mechanism such as a coin flip, random number generator, or lottery. This design provides the strongest basis for causal inference because randomization ensures that, in expectation, all confounding variables—both observed and unobserved—are balanced across groups. In political science, true experiments take two primary forms: laboratory experiments, where subjects (often university students) are brought into a controlled environment and exposed to stimuli such as mock campaigns or framed policy descriptions; and field experiments, where randomization occurs in real-world settings such as actual elections, government programs, or community interventions.
Natural Experiments
A natural experiment occurs when some exogenous event—a policy change, natural disaster, historical accident, lottery, or institutional rule—creates variation in treatment exposure that is as if randomly assigned, even though no researcher manipulated the assignment. The term was popularized in political science by Thad Dunning's 2012 book Natural Experiments in the Social Sciences. A canonical example is the Vietnam draft lottery, in which birth dates were randomly drawn to determine military conscription, allowing researchers to study the effects of military service on later-life outcomes without the selection bias that would contaminate a comparison of volunteers versus non-volunteers. The critical assumption in any natural experiment is that the exogenous event is truly independent of the potential outcomes—what is often called the exclusion restriction.
Quasi-Experimental Designs
A quasi-experiment resembles a true experiment in that it compares treated and untreated groups, but it lacks random assignment. Instead, the researcher exploits a particular feature of the treatment assignment process—a threshold, a temporal break, or a matched comparison—to construct a counterfactual. The three most common quasi-experimental techniques in political science are difference-in-differences (DiD), which compares changes over time between a group that received treatment and one that did not; regression discontinuity design (RDD), which exploits a cutoff rule that assigns treatment based on a continuous variable (such as an election margin); and matching, which pairs treated units with observationally similar control units to approximate what randomization would have achieved.
Worked Example: Identifying the Right Design
Consider the following research question: Does receiving information about candidate policy positions increase voter turnout? We will walk through how each of the three experimental designs could be applied to this question, evaluating the strengths and limitations of each approach in practice.
Strengths, Limitations, and Trade-Offs
No single experimental design dominates across all dimensions. Each offers a distinctive combination of inferential leverage and practical constraints, and understanding these trade-offs is essential for both designing and evaluating political science research.
| Criterion | True Experiment | Natural Experiment | Quasi-Experiment |
|---|---|---|---|
| Internal Validity | Highest—random assignment eliminates all confounders in expectation | High if as-if random assumption holds; must be argued, not guaranteed | Moderate to low—depends on quality of design and plausibility of identifying assumptions |
| External Validity | Variable—lab experiments often low; field experiments can be high | Often limited to context of the exogenous event; LATE rather than ATE | Can be high if samples are large and representative |
| Feasibility | Often expensive and ethically constrained; cannot randomize war, poverty, or regime type | Opportunistic—depends on finding suitable exogenous variation; cannot be planned in advance | Most feasible; can use existing observational data with appropriate design |
| Ethical Constraints | Significant—withholding beneficial treatment raises fairness concerns; IRB oversight required | Minimal—researcher observes rather than intervenes; but studying harmful events raises its own concerns | Minimal—no manipulation involved; but causal claims require careful qualification |
| Key Assumption | SUTVA: no interference between units; compliance with assignment | Independence: exogenous event is unrelated to potential outcomes; exclusion restriction | Design-specific: parallel trends (DiD), continuity (RDD), selection on observables (matching) |
| Estimand | ATE (sample-level); ITT if noncompliance exists | LATE (effect for compliers affected by the instrument) | ATT or LATE depending on design |
Connections to Advanced Causal Inference
The three experimental designs introduced in this lesson form the foundation of the broader causal inference toolkit that has reshaped empirical political science. As you advance in research methods coursework, you will encounter increasingly sophisticated extensions of these designs that address more complex identification challenges.
| Foundational Concept | Advanced Extension |
|---|---|
| Random assignment (RCT) | Cluster-randomized designs, factorial experiments, adaptive experiments, and conjoint analysis for multi-dimensional preferences |
| Natural experiments with exogenous variation | Instrumental variables (IV) estimation using two-stage least squares (2SLS); fuzzy regression discontinuity |
| Difference-in-differences | Staggered DiD with heterogeneous treatment effects (Callaway & Sant'Anna, Sun & Abraham estimators); synthetic control method |
| Matching on observables | Propensity score methods, coarsened exact matching (CEM), sensitivity analysis for unobserved confounders (Rosenbaum bounds) |
| Regression discontinuity design | Geographic RDD, multi-dimensional RDD, bandwidth selection optimization (Calonico, Cattaneo & Titiunik) |
One particularly important development is the synthetic control method, developed by Abadie, Diamond, and Hainmueller (2010), which constructs a weighted combination of untreated units to serve as a counterfactual for a single treated unit—a technique widely used in comparative politics and policy analysis. Another frontier is the integration of machine learning methods with causal inference, including causal forests for estimating heterogeneous treatment effects and double/debiased machine learning for high-dimensional confounder adjustment. These methods build directly on the potential outcomes framework and design-based logic you have learned in this lesson, extending them to settings with greater complexity and richer data.
Practice Problems
Lesson Summary
This lesson examined three foundational approaches to causal inference in political science. True experiments (RCTs) use random assignment to eliminate selection bias, providing the strongest internal validity for estimating the Average Treatment Effect (ATE). Natural experiments exploit exogenous events—policy changes, lotteries, institutional rules—that create as-if random variation, allowing researchers to estimate causal effects when controlled randomization is impossible, though they require careful argumentation that the independence assumption holds. Quasi-experimental designs—including difference-in-differences, regression discontinuity, and matching—employ sophisticated analytical strategies to approximate experimental conditions when assignment is non-random.
The choice among these designs involves navigating trade-offs between internal validity and external validity, ethical constraints, and practical feasibility. The potential outcomes framework (τᵢ = Yᵢ(1) − Yᵢ(0)) provides the formal foundation for understanding why the counterfactual problem makes causal inference challenging, and why each design's credibility ultimately rests on the plausibility of its identifying assumptions. Mastering these designs equips political scientists to both conduct rigorous original research and critically evaluate the causal claims that shape our understanding of political phenomena.