COLLEGE POLITICAL SCIENCE • RESEARCH METHODS

Experimental Designs — Explain experiments, natural experiments, and quasi-experimental designs

Understanding how researchers isolate causal effects in political science through randomized, natural, and quasi-experimental strategies.

Historical Context & Motivation

Political science has long grappled with a fundamental question: how can scholars move beyond mere correlation to establish that one political phenomenon actually causes another? For much of the discipline's history, researchers relied on observational data—surveys, case studies, and statistical analyses of existing datasets—to draw inferences about political behavior, institutions, and policy outcomes. Yet these approaches consistently confronted the problem of confounding variables: unmeasured or uncontrolled factors that could explain away apparent causal relationships. The migration of experimental methods from the natural sciences and psychology into political science represented a paradigm shift, one that redefined what counts as credible evidence in the study of politics.

The intellectual roots of experimental reasoning trace back to the philosophy of science and the statistical innovations of the early twentieth century. R.A. Fisher's work on randomization in agricultural experiments laid the theoretical groundwork, but it took decades for political scientists to embrace these tools. The credibility revolution in the social sciences—a term popularized by economists Joshua Angrist and Jörn-Steffen Pischke—arrived in political science during the late 1990s and early 2000s, encouraging scholars to think rigorously about identification strategies and causal inference rather than relying solely on regression analysis with observational data.

1935
Fisher's Design of Experiments
R.A. Fisher published The Design of Experiments, formalizing the logic of randomization and establishing the conceptual foundation for all experimental research across disciplines.
1965
Campbell & Stanley's Validity Framework
Donald Campbell and Julian Stanley published their influential taxonomy of experimental and quasi-experimental designs, introducing the concepts of internal validity and external validity that remain central to research design today.
1998
Gerber & Green's Voter Mobilization Experiment
Alan Gerber and Donald Green conducted a large-scale randomized field experiment on voter turnout in New Haven, CT, demonstrating that door-to-door canvassing substantially increased participation and sparking a wave of experimental research in political science.
2004
Angrist & Pischke and the Credibility Revolution
The growing influence of Angrist, Pischke, and their students catalyzed political scientists to adopt natural experiments and quasi-experimental designs, prioritizing causal identification over complex structural models.
2012
EGAP and Pre-Registration
The Evidence in Governance and Politics (EGAP) network formalized pre-registration protocols for experimental research in political science, addressing concerns about publication bias and researcher degrees of freedom.

The central question that experimental designs address is deceptively simple: What would have happened in the absence of a treatment or intervention? This is the counterfactual problem—we can never simultaneously observe a unit in both the treated and untreated state. Different experimental designs offer different strategies for constructing a credible comparison group that approximates this unobservable counterfactual, and the distinctions among true experiments, natural experiments, and quasi-experiments hinge on how successfully each design resolves this fundamental challenge.

Core Principles & Definitions

Before distinguishing among the three major experimental designs, it is essential to establish the conceptual vocabulary that underpins all causal inference in political science. These foundational principles determine how researchers evaluate whether a study's findings genuinely reflect cause and effect, and they provide the criteria by which we judge one design as stronger or weaker than another.

1

Causal Inference & the Counterfactual

Causal inference requires comparing what actually happened under a treatment to what would have happened without it. Because the counterfactual is inherently unobservable, research designs must construct credible approximations through control groups.
2

Random Assignment

When units are randomly assigned to treatment and control conditions, all observed and unobserved characteristics are distributed equally in expectation across groups, eliminating selection bias and isolating the causal effect of the treatment.
3

Internal Validity

A study has high internal validity when observed differences between groups can be confidently attributed to the treatment rather than confounding factors. Randomized experiments maximize internal validity; observational studies often struggle to achieve it.
4

External Validity

External validity captures how well findings generalize to populations, settings, or time periods beyond the study. Laboratory experiments may achieve high internal validity but low external validity if their controlled conditions diverge sharply from real-world politics.
5

Treatment Effect

The Average Treatment Effect (ATE) is the expected difference in outcomes between treated and control units across the entire population. Researchers may also estimate the ATE on the treated (ATT) or local average treatment effects (LATE) depending on design constraints.
KEY TAKEAWAY
Think of experimental design like a pharmaceutical drug trial. In a true experiment, patients are randomly assigned to receive the drug or a placebo—this ensures the two groups are comparable in all respects except the drug itself. A natural experiment is like discovering that one hospital happened to prescribe the drug while a neighboring hospital did not, creating a comparison not by design but by circumstance. A quasi-experiment is like comparing patients who chose the drug with those who declined, then using statistical techniques to adjust for the fact that choice itself may reflect health differences. Each approach tries to answer the same question—'Does the drug work?'—but with varying degrees of confidence in the answer.

Visual Explanation of Experimental Design Types

The following diagram illustrates the structural logic of the three experimental designs. Each design begins with a population, identifies how units enter treatment and control conditions, and measures outcomes to estimate causal effects. The key differentiator across the three designs is the assignment mechanism—whether it is controlled by the researcher, generated by an external event, or driven by self-selection that must be addressed analytically.

The three panels show how assignment mechanism differs across designs. In true experiments (left), the researcher randomly assigns units. In natural experiments (center), an external event creates as-if random variation. In quasi-experiments (right), assignment is non-random, requiring statistical adjustments such as difference-in-differences (DiD) or regression discontinuity (RDD) to approximate the causal estimate.

Notice how the diagram reveals a gradient of researcher control from left to right. In a true experiment, the researcher exercises direct control over the assignment mechanism, which is the gold standard for establishing causality. In a natural experiment, the researcher identifies a naturally occurring event—such as a policy change, lottery, or geographic boundary—that generates variation resembling randomization. In a quasi-experiment, there is no randomization at all, and the researcher must rely on design-based strategies and statistical adjustments to construct a credible comparison. This progression from randomized to as-if random to non-random assignment maps directly onto decreasing confidence in internal validity, all else being equal.

The Logic of Causal Identification

While political science is not typically a math-heavy discipline, the formal framework for understanding causal inference provides precision that verbal arguments alone cannot achieve. The potential outcomes framework (also known as the Rubin Causal Model or Neyman-Rubin framework) formalizes the counterfactual logic underlying all experimental designs. For each unit i in a study, we define two potential outcomes: Yi(1) is the outcome if unit i receives treatment, and Yi(0) is the outcome if unit i does not. The fundamental problem of causal inference is that we observe only one of these outcomes for any given unit.

INDIVIDUAL TREATMENT EFFECT
τᵢ = Yᵢ(1) − Yᵢ(0)
Where τi is the causal effect of treatment for unit i, Yi(1) is the treated potential outcome, and Yi(0) is the untreated potential outcome. Since we can never observe both simultaneously, we estimate the average effect across units.
AVERAGE TREATMENT EFFECT (ATE)
ATE = E[Yᵢ(1) − Yᵢ(0)] = E[Yᵢ(1)] − E[Yᵢ(0)]
E[·] denotes the expected value (population average). The ATE measures the average causal effect of treatment across all units in the population. With random assignment, E[Y | Treatment] − E[Y | Control] is an unbiased estimator of the ATE.
SELECTION BIAS
E[Y | D=1] − E[Y | D=0] = ATE + {E[Y(0) | D=1] − E[Y(0) | D=0]}
D indicates treatment status (1 = treated, 0 = control). The term in braces is selection bias—the systematic difference in baseline outcomes between those who receive treatment and those who do not. Randomization drives this term to zero in expectation; quasi-experiments must analytically address it.

This formal decomposition reveals why experimental designs differ in their inferential power. In a true experiment with random assignment, the selection bias term equals zero by design—treated and control groups are statistically identical before the intervention. In a natural experiment, the researcher argues that an exogenous event made treatment assignment effectively random, so the bias term is approximately zero. In a quasi-experiment, the selection bias term is likely non-zero, so the researcher must employ strategies—matching, regression discontinuity, difference-in-differences, instrumental variables—to either eliminate or bound it analytically.

🎲 Why Randomization Works
Randomization does not guarantee that treatment and control groups are identical in any single study—it guarantees that differences are due to chance alone. This is why researchers conduct balance checks (comparing pre-treatment covariates) and why larger sample sizes yield more precise estimates. The law of large numbers ensures that random differences wash out as N increases.

Detailed Breakdown of Each Design Type

True (Randomized) Experiments

A true experiment (also called a randomized controlled trial or RCT) is the design in which the researcher directly controls the assignment of units to treatment and control conditions through a random mechanism such as a coin flip, random number generator, or lottery. This design provides the strongest basis for causal inference because randomization ensures that, in expectation, all confounding variables—both observed and unobserved—are balanced across groups. In political science, true experiments take two primary forms: laboratory experiments, where subjects (often university students) are brought into a controlled environment and exposed to stimuli such as mock campaigns or framed policy descriptions; and field experiments, where randomization occurs in real-world settings such as actual elections, government programs, or community interventions.

Natural Experiments

A natural experiment occurs when some exogenous event—a policy change, natural disaster, historical accident, lottery, or institutional rule—creates variation in treatment exposure that is as if randomly assigned, even though no researcher manipulated the assignment. The term was popularized in political science by Thad Dunning's 2012 book Natural Experiments in the Social Sciences. A canonical example is the Vietnam draft lottery, in which birth dates were randomly drawn to determine military conscription, allowing researchers to study the effects of military service on later-life outcomes without the selection bias that would contaminate a comparison of volunteers versus non-volunteers. The critical assumption in any natural experiment is that the exogenous event is truly independent of the potential outcomes—what is often called the exclusion restriction.

Quasi-Experimental Designs

A quasi-experiment resembles a true experiment in that it compares treated and untreated groups, but it lacks random assignment. Instead, the researcher exploits a particular feature of the treatment assignment process—a threshold, a temporal break, or a matched comparison—to construct a counterfactual. The three most common quasi-experimental techniques in political science are difference-in-differences (DiD), which compares changes over time between a group that received treatment and one that did not; regression discontinuity design (RDD), which exploits a cutoff rule that assigns treatment based on a continuous variable (such as an election margin); and matching, which pairs treated units with observationally similar control units to approximate what randomization would have achieved.

This diagram illustrates three common quasi-experimental techniques. Difference-in-differences compares trends in treatment and control groups before and after an intervention (the green bar τ shows the estimated treatment effect). Regression discontinuity exploits a cutoff in a running variable—units just above and just below the threshold are compared, with the vertical green line representing the causal jump. Matching pairs treated units (T) with similar control units (C) on observed covariates, discarding unmatched units.

Worked Example: Identifying the Right Design

Consider the following research question: Does receiving information about candidate policy positions increase voter turnout? We will walk through how each of the three experimental designs could be applied to this question, evaluating the strengths and limitations of each approach in practice.

Designing a Study on Information and Voter Turnout
1
Step 1 — Define the Treatment and OutcomeThe treatment is exposure to information about candidate policy positions (e.g., a nonpartisan voter guide). The outcome is whether the individual votes in the upcoming election, measured through validated voter files. Before choosing a design, we must clearly operationalize what 'information' means—is it a mailed brochure, an email, a text message?—and how turnout will be verified.
Treatment: Nonpartisan voter guide mailed to household. Outcome: Validated turnout in voter file (binary: voted/did not vote).
2
Step 2 — True Experiment ApproachIn a randomized field experiment, the researcher would obtain a list of registered voters in a target jurisdiction and randomly assign a subset to receive the voter guide (treatment group) while the remainder receives no guide (control group). Because assignment is random, any pre-existing differences between groups—age, education, political interest, prior voting history—are balanced in expectation. After the election, the researcher compares turnout rates: ATE = (turnout rate in treatment group) − (turnout rate in control group). This is exactly the design Gerber, Green, and Larimer (2008) used when they mailed social pressure messages to Michigan voters.
Random assignment eliminates selection bias. The observed difference in turnout rates is an unbiased estimate of the causal effect of receiving the voter guide.
3
Step 3 — Natural Experiment ApproachSuppose the researcher lacks the resources to conduct a field experiment but discovers that a local newspaper accidentally failed to deliver its election guide to one set of zip codes due to a printing error. If the affected zip codes were determined by an essentially arbitrary logistical failure rather than by the characteristics of their residents, this natural experiment creates as-if random variation in information exposure. The researcher would compare turnout in zip codes that received the guide versus those that did not, arguing that the printing error was exogenous to voter characteristics. The key assumption—called the independence assumption—is that the printing error was unrelated to the political composition or turnout propensity of those zip codes.
If the independence assumption holds, the turnout difference estimates a LATE. The researcher must provide evidence (balance tests on demographics) that affected and unaffected areas are comparable.
4
Step 4 — Quasi-Experimental ApproachImagine no natural experiment is available either. Instead, the researcher observes that one county adopted a mandatory voter guide mailing program in 2020 while a neighboring county did not. Using a difference-in-differences design, the researcher compares the change in turnout from 2018 to 2020 in the treated county with the change in turnout over the same period in the comparison county. The critical assumption is the parallel trends assumption: absent the voter guide program, turnout trends would have evolved identically in both counties. The researcher supports this assumption by showing that turnout trends were parallel in the years preceding the intervention (e.g., 2014–2018).
DiD estimate: (ΔTurnout_treated county) − (ΔTurnout_comparison county). Valid only if parallel trends assumption is credible, which must be demonstrated with pre-treatment data.
5
Step 5 — Evaluate and SelectThe true experiment provides the strongest internal validity but may be costly and logistically complex. The natural experiment offers a clever alternative if an exogenous source of variation can be identified, but its validity depends entirely on the credibility of the as-if random argument. The quasi-experiment is the most feasible when neither randomization nor a natural instrument is available, but it requires the strongest assumptions. In practice, the best design is the one whose identifying assumptions are most defensible given the specific context—there is no universally superior approach. Researchers often use multiple designs to triangulate findings, strengthening confidence in causal conclusions.
Decision rule: Choose the design whose identifying assumptions are most credible and defensible for the specific research context. When possible, triangulate across multiple designs.

Strengths, Limitations, and Trade-Offs

No single experimental design dominates across all dimensions. Each offers a distinctive combination of inferential leverage and practical constraints, and understanding these trade-offs is essential for both designing and evaluating political science research.

Comparative strengths and limitations across the three experimental design types
CriterionTrue ExperimentNatural ExperimentQuasi-Experiment
Internal ValidityHighest—random assignment eliminates all confounders in expectationHigh if as-if random assumption holds; must be argued, not guaranteedModerate to low—depends on quality of design and plausibility of identifying assumptions
External ValidityVariable—lab experiments often low; field experiments can be highOften limited to context of the exogenous event; LATE rather than ATECan be high if samples are large and representative
FeasibilityOften expensive and ethically constrained; cannot randomize war, poverty, or regime typeOpportunistic—depends on finding suitable exogenous variation; cannot be planned in advanceMost feasible; can use existing observational data with appropriate design
Ethical ConstraintsSignificant—withholding beneficial treatment raises fairness concerns; IRB oversight requiredMinimal—researcher observes rather than intervenes; but studying harmful events raises its own concernsMinimal—no manipulation involved; but causal claims require careful qualification
Key AssumptionSUTVA: no interference between units; compliance with assignmentIndependence: exogenous event is unrelated to potential outcomes; exclusion restrictionDesign-specific: parallel trends (DiD), continuity (RDD), selection on observables (matching)
EstimandATE (sample-level); ITT if noncompliance existsLATE (effect for compliers affected by the instrument)ATT or LATE depending on design
⚖️ THE VALIDITY TRADE-OFF
Think of internal and external validity as sitting on opposite ends of a seesaw. A tightly controlled lab experiment pushes internal validity up but may sacrifice real-world relevance. A large-scale observational study with a quasi-experimental design may reflect the messiness of actual politics—enhancing generalizability—but at the cost of greater uncertainty about whether the observed effect is truly causal. The art of research design in political science is finding the approach that achieves the best balance for the specific question being asked, and no amount of statistical sophistication can fully substitute for a design that is fundamentally well-matched to its research question.

Connections to Advanced Causal Inference

The three experimental designs introduced in this lesson form the foundation of the broader causal inference toolkit that has reshaped empirical political science. As you advance in research methods coursework, you will encounter increasingly sophisticated extensions of these designs that address more complex identification challenges.

How foundational designs connect to advanced methods in political science
Foundational ConceptAdvanced Extension
Random assignment (RCT)Cluster-randomized designs, factorial experiments, adaptive experiments, and conjoint analysis for multi-dimensional preferences
Natural experiments with exogenous variationInstrumental variables (IV) estimation using two-stage least squares (2SLS); fuzzy regression discontinuity
Difference-in-differencesStaggered DiD with heterogeneous treatment effects (Callaway & Sant'Anna, Sun & Abraham estimators); synthetic control method
Matching on observablesPropensity score methods, coarsened exact matching (CEM), sensitivity analysis for unobserved confounders (Rosenbaum bounds)
Regression discontinuity designGeographic RDD, multi-dimensional RDD, bandwidth selection optimization (Calonico, Cattaneo & Titiunik)

One particularly important development is the synthetic control method, developed by Abadie, Diamond, and Hainmueller (2010), which constructs a weighted combination of untreated units to serve as a counterfactual for a single treated unit—a technique widely used in comparative politics and policy analysis. Another frontier is the integration of machine learning methods with causal inference, including causal forests for estimating heterogeneous treatment effects and double/debiased machine learning for high-dimensional confounder adjustment. These methods build directly on the potential outcomes framework and design-based logic you have learned in this lesson, extending them to settings with greater complexity and richer data.

🔭 Looking Ahead
As political science increasingly embraces pre-registration, replication, and open science practices, the transparency of research design becomes paramount. Understanding the logic of experiments, natural experiments, and quasi-experiments equips you not only to conduct original research but also to critically evaluate the causal claims made in published work—a skill essential for any informed consumer of political science scholarship.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain the fundamental problem of causal inference and describe how randomized experiments address it. Why can we never directly observe a causal effect for a single individual?
PROBLEM 2BASIC CALCULATION
A field experiment randomly assigns 5,000 registered voters to receive a mobilization text message (treatment) and 5,000 to receive no message (control). After the election, 1,750 treated voters turned out compared to 1,500 control voters. Calculate the estimated Average Treatment Effect and interpret it substantively.
PROBLEM 3INTERMEDIATE
A researcher wants to study the effect of compulsory voting laws on political knowledge. Country A adopted compulsory voting in 2015; Country B, which is similar in demographics and political culture, did not. The researcher has survey data on political knowledge scores in both countries from 2012 and 2018. Describe the difference-in-differences approach she should use. What is the key identifying assumption, and how could she assess its plausibility?
PROBLEM 4APPLIED
In many U.S. state legislatures, candidates who win close elections (e.g., by less than 1% of the vote) gain access to incumbency advantages such as name recognition, staff resources, and franking privileges. Describe how you would use a regression discontinuity design to estimate the causal effect of incumbency on the probability of winning the next election. Identify the running variable, the cutoff, the treatment, and one potential threat to validity.
PROBLEM 5CRITICAL THINKING
A colleague proposes using a natural experiment to study the effect of democratic institutions on economic growth. She argues that the colonial origins of legal systems (British common law vs. French civil law) provide exogenous variation in institutional quality because colonial assignment was 'essentially arbitrary.' Critically evaluate this identification strategy. Under what conditions would this be a valid natural experiment, and what are the most serious threats to this claim?

Lesson Summary

This lesson examined three foundational approaches to causal inference in political science. True experiments (RCTs) use random assignment to eliminate selection bias, providing the strongest internal validity for estimating the Average Treatment Effect (ATE). Natural experiments exploit exogenous events—policy changes, lotteries, institutional rules—that create as-if random variation, allowing researchers to estimate causal effects when controlled randomization is impossible, though they require careful argumentation that the independence assumption holds. Quasi-experimental designs—including difference-in-differences, regression discontinuity, and matching—employ sophisticated analytical strategies to approximate experimental conditions when assignment is non-random.

The choice among these designs involves navigating trade-offs between internal validity and external validity, ethical constraints, and practical feasibility. The potential outcomes framework (τᵢ = Yᵢ(1) − Yᵢ(0)) provides the formal foundation for understanding why the counterfactual problem makes causal inference challenging, and why each design's credibility ultimately rests on the plausibility of its identifying assumptions. Mastering these designs equips political scientists to both conduct rigorous original research and critically evaluate the causal claims that shape our understanding of political phenomena.

Varsity Tutors • College Political Science • Experimental Designs — Explain experiments, natural experiments, and quasi-experimental designs