MATH 3 • STATISTICS & PROBABILITY

Observational Studies vs. Experiments — I can distinguish between observational studies and experiments and explain implications for causation.

Understanding why only controlled experiments can establish cause and effect.

Historical Context & Motivation

For centuries, humans have tried to understand why things happen. Does a new medicine actually cure a disease, or do patients just get better on their own? Does studying with music help you learn, or does it just feel that way? The key to answering these questions lies in how we collect data. The distinction between simply watching what happens and deliberately controlling conditions has shaped modern science, medicine, and public policy.

1747
First Clinical Trial
Scottish naval surgeon James Lind tested six different treatments for scurvy on groups of sailors aboard HMS Salisbury. By deliberately assigning treatments, he discovered that citrus fruits cured the disease — one of the first known controlled experiments in medicine.
1854
Snow's Cholera Study
John Snow mapped cholera cases across London and noticed clusters near contaminated water pumps. He could not assign people to drink from different pumps, so this landmark investigation was an observational study — yet it strongly suggested contaminated water was the cause.
1948
Randomized Controlled Trials
The British Medical Research Council conducted the first modern randomized controlled trial (RCT) to test streptomycin for tuberculosis. Random assignment became the gold standard for establishing causation.
1964
Smoking & Lung Cancer Report
The U.S. Surgeon General concluded that cigarette smoking causes lung cancer, relying heavily on observational studies because it would be unethical to randomly assign people to smoke. This case showed that observational evidence can be powerful even without a true experiment.
2000s
Big Data & Modern Studies
With massive datasets and advanced computing, researchers now run observational studies on millions of people using electronic health records and social media data, while A/B testing (a form of experiment) is used daily by tech companies.

The central question that this lesson addresses is straightforward but powerful: When can we say one thing actually causes another, and when can we only say two things are related? The answer depends entirely on whether the data came from an observational study or an experiment.

Core Principles & Definitions

Before diving deeper, let's establish the foundational ideas that separate observational studies from experiments. These principles will guide your thinking every time you encounter a research claim in the news, in class, or on a standardized test.

1

Observational Study

A study in which researchers observe and measure subjects without attempting to influence or change any variables. The researcher simply records what naturally happens. Example: surveying students about their sleep habits and GPA.
2

Experiment

A study in which researchers deliberately impose a treatment on subjects and then measure the response. The researcher controls who gets the treatment. Example: randomly assigning students to sleep 8 hours or 6 hours and comparing test scores.
3

Confounding Variable

A confounding variable (or lurking variable) is a hidden factor that influences both the explanatory variable and the response variable, creating a misleading association. Example: students who sleep more may also have less stress — stress could be the real driver of GPA.
4

Random Assignment

In experiments, random assignment means using a chance process (like flipping a coin) to decide which subjects receive the treatment. This balances out confounding variables across groups, allowing researchers to isolate the effect of the treatment.
5

Causation vs. Association

An association (or correlation) means two variables tend to change together. Causation means one variable directly produces a change in another. Only well-designed experiments can establish causation.
KEY TAKEAWAY
Think of the difference like this: an observational study is like watching a football game from the stands — you can see patterns but you can't control what the players do. An experiment is like being the coach — you decide who plays what position and then see how the game turns out. Because the coach controls the setup, the coach can figure out what actually caused the win. The spectator can only guess.

Visual Explanation

The diagram below illustrates the fundamental structural difference between an observational study and an experiment. Pay close attention to where the researcher's role diverges: in one path, the researcher merely watches; in the other, the researcher actively assigns treatments.

The left path shows an observational study: subjects sort themselves into groups, and confounding variables may lurk behind any association. The right path shows an experiment: the researcher randomly assigns subjects to treatment or control, balancing confounders and enabling a causal conclusion.

Notice the critical fork in the diagram. In the observational study path, subjects end up in groups based on their own characteristics or choices — the researcher has no say. This means that the groups may differ in ways beyond the variable of interest, and those differences are confounding variables. In the experiment path, random assignment ensures that, on average, the groups are alike in every way except the treatment. That is why experiments — and only experiments — can support a claim of causation.

How Confounding Works & Why Random Assignment Fixes It

To understand the implications for causation, you need to see exactly how a confounding variable can trick you into thinking one thing causes another. Consider this scenario: a researcher notices that students who eat breakfast tend to earn higher grades. Does breakfast cause better grades? Maybe — but students who eat breakfast may also come from families that emphasize healthy routines and academic support. The family environment is a confounding variable that is associated with both breakfast-eating and academic performance.

The Confounding Triangle

Statisticians often visualize confounding as a triangle. The explanatory variable (breakfast) and the response variable (grades) sit at two corners, and the confounding variable (family environment) sits at the third. Arrows from the confounding variable point toward both the explanatory and response variables, showing that it influences both.

The confounding triangle: the amber confounding variable influences both the pink explanatory variable and the cyan response variable. The dashed arrow between the explanatory and response variables represents a potentially misleading observed association.

How Random Assignment Breaks the Triangle

When researchers use random assignment, they break the link between the confounding variable and the explanatory variable. Because a coin flip (or a random number generator) decides who gets the treatment, the confounding variable can no longer systematically "choose" who ends up in each group. On average, the treatment and control groups will have similar family environments, similar stress levels, similar everything — except the one variable the researcher is testing. When you remove the confounders, any remaining difference in outcomes can reasonably be attributed to the treatment itself.

Important Distinction
Don't confuse random assignment with random sampling. Random sampling is how you select subjects from a population — it affects whether your results generalize to the population (external validity). Random assignment is how you divide subjects into treatment and control groups within a study — it affects whether you can claim causation (internal validity). Both are desirable, but they serve different purposes.

Classifying Study Designs

Not every study fits neatly into one box. Within the broad categories of observational studies and experiments, there are several common designs. Learning to recognize them will help you quickly evaluate research claims and identify the level of evidence they provide.

Types of Observational Studies

  • Sample survey: Researchers collect data at a single point in time, often through questionnaires. Example: a poll asking teens how many hours they spend on social media per day and their self-reported anxiety level.
  • Prospective study: Researchers identify a group of subjects and follow them forward in time. Example: tracking 1,000 high school freshmen over four years to see if exercise habits predict college acceptance rates.
  • Retrospective study: Researchers look backward in time, using existing records or asking subjects about past behavior. Example: interviewing college students about their high school study habits.

Key Features of Well-Designed Experiments

  • Control group: A group that does not receive the treatment, providing a baseline for comparison.
  • Random assignment: Using a chance process to assign subjects to groups, minimizing the effect of confounders.
  • Replication: Using enough subjects so that the results are not due to chance or individual variation.
  • Blinding: Keeping subjects (single-blind) or both subjects and researchers (double-blind) unaware of group assignments to prevent bias.
  • Placebo: A fake treatment given to the control group so that any psychological effect of receiving treatment is equalized.
Comparison of key features between observational studies and experiments
FeatureObservational StudyExperiment
Researcher assigns treatment?NoYes
Random assignment used?NoYes (in well-designed experiments)
Can establish causation?NoYes
Confounders controlled?Not reliablyBalanced across groups by randomization
Ethical flexibilityCan study harmful exposures ethicallyCannot assign harmful treatments

Worked Example: Identifying Study Type & Drawing Conclusions

Let's walk through a realistic scenario step by step. A school administrator wants to know whether a new tutoring program improves math test scores. Two hundred students volunteer for the study. A coin is flipped for each student: heads means the student joins the tutoring program; tails means the student does not. After one semester, test scores for both groups are compared.

Does the Tutoring Program Improve Math Scores?
1
Step 1 — Identify the Study TypeAsk the key question: Did the researcher impose a treatment? Yes — the researcher used a coin flip to decide which students joined the tutoring program. Because a treatment was deliberately imposed on subjects, this is an experiment.
Study type: Experiment
2
Step 2 — Identify Key ComponentsThe explanatory variable (treatment) is participation in the tutoring program. The response variable is the math test score. The treatment group consists of the students assigned to tutoring, and the control group consists of those who were not.
Explanatory: tutoring (yes/no); Response: test score
3
Step 3 — Check for Random AssignmentA coin was flipped for each student, so group membership was determined by chance. This means random assignment was used. Confounding variables — such as prior math ability, motivation, or socioeconomic background — should be roughly balanced across the two groups.
Random assignment: Yes (coin flip)
4
Step 4 — State the Conclusion AppropriatelySuppose the tutoring group scored an average of 12 points higher than the control group. Because this was a well-designed experiment with random assignment, we can say the tutoring program caused the improvement in scores. If instead students had simply chosen whether to attend tutoring (no random assignment), this would be an observational study, and we could only say there is an association between tutoring and higher scores — motivated students might have signed up AND studied more.
Conclusion: The tutoring program caused higher math scores.
📝 Language Matters
When writing conclusions, use causal language ("causes," "leads to," "results in") only for experiments with random assignment. For observational studies, use association language ("is associated with," "is correlated with," "tends to"). Getting this language right is one of the most common — and most important — skills tested on exams.

Strengths & Limitations of Each Approach

If experiments are the only way to establish causation, why don't researchers always run experiments? The answer involves practical constraints, ethical boundaries, and the nature of the questions being studied. Both observational studies and experiments have important roles in building knowledge.

Strengths and limitations of observational studies vs. experiments
CriterionObservational StudyExperiment
Causal claimsCannot establish causation; confounders may lurk.Can establish causation when random assignment is used.
EthicsCan study harmful exposures (e.g., smoking) without forcing participation.Cannot assign harmful or dangerous treatments to subjects.
Cost & timeOften cheaper and faster; can use existing records.Usually more expensive and time-consuming to set up.
Real-world settingData reflects natural behavior; higher ecological validity.Controlled conditions may not reflect real life.
Sample sizeCan often include thousands or millions of subjects.Typically smaller due to logistical constraints.
Best used whenThe variable of interest cannot or should not be manipulated.The researcher wants to test a specific cause-and-effect hypothesis.
KEY TAKEAWAY
Think of it like a courtroom. An observational study is like circumstantial evidence — it can strongly suggest guilt (association), but it's not proof. An experiment with random assignment is more like DNA evidence — it directly connects the suspect (treatment) to the crime (outcome). Both types of evidence are useful, but only the direct evidence proves the case beyond reasonable doubt.

Connections to Inference & Advanced Statistics

Understanding the difference between observational studies and experiments is not just a vocabulary lesson — it forms the foundation for every statistical inference you'll encounter in later courses. When you learn about hypothesis tests and confidence intervals, the type of study determines what kind of conclusion you can draw from your data.

How this lesson's concepts connect to advanced topics
Concept in This LessonConnection to Advanced Statistics
Random assignment → causationIn AP Statistics, you'll learn to perform significance tests where the null hypothesis assumes no treatment effect. If results are statistically significant, causation can be claimed only if random assignment was used.
Random sampling → generalizationRandom sampling allows you to generalize results to the broader population. Confidence intervals rely on this assumption.
Confounding variablesIn regression analysis, statisticians use techniques like multiple regression to statistically control for confounders when experiments aren't possible.
Observational associationCorrelation coefficients (r) and scatterplots quantify the strength of associations found in observational data, but a strong r value never implies causation on its own.

A useful framework for remembering the scope of conclusions is the Scope of Inference table. It combines two questions: (1) Was random assignment used? and (2) Was random sampling used? Your answers determine what you can say about your results.

Scope of inference: what you can conclude depends on how the study was designed
Random Assignment: YesRandom Assignment: No
Random Sampling: YesCausation + Generalization (best case)Association + Generalization
Random Sampling: NoCausation, but only for subjects studiedAssociation only, for subjects studied (weakest)

As you move into AP Statistics or college-level courses, you'll learn formal methods — like randomization tests and propensity score matching — that attempt to draw stronger conclusions from observational data. However, the fundamental principle remains: without random assignment, you cannot rule out confounders with certainty.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher reads that people who drink green tea tend to live longer. She concludes that green tea causes longer life. What is the flaw in her reasoning, and what type of study likely produced this finding?
PROBLEM 2BASIC CALCULATION
A teacher randomly assigns 30 students to use a new study app and 30 students to study without the app. After a month, the app group averages 85% on a test and the no-app group averages 78%. (a) What type of study is this? (b) Can the teacher conclude the app caused the higher scores? Explain.
PROBLEM 3INTERMEDIATE
A hospital wants to know if a new drug reduces recovery time after surgery. Due to ethical concerns, doctors allow patients to choose whether they take the new drug or the standard medication. Patients who choose the new drug recover 2 days faster on average. A newspaper headline reads: 'New Drug Cuts Recovery Time by 2 Days.' (a) What type of study is this? (b) Identify at least one confounding variable. (c) Rewrite the headline to accurately reflect the evidence.
PROBLEM 4APPLIED
A tech company wants to test whether changing the color of a "Buy Now" button from blue to orange increases the percentage of customers who make a purchase. They randomly show half of website visitors the blue button and half the orange button, then track purchase rates. (a) Classify this study. (b) Suppose 4.2% of orange-button visitors purchase vs. 3.5% of blue-button visitors. Write an appropriate conclusion. (c) Name one feature the company should add to improve the study's design.
PROBLEM 5CRITICAL THINKING
A school board wants to know if banning cell phones during class improves student performance. They propose two study designs: (A) Compare test scores at schools that have already banned phones with schools that haven't. (B) Randomly select 20 classrooms in the district; flip a coin to determine which 10 classrooms ban phones for a semester and which 10 don't, then compare scores. For each design, state the study type, whether causation can be claimed, name a possible confounding variable (for design A), and discuss one ethical or practical challenge.

Lesson Summary

The central distinction in this lesson is between observational studies, where the researcher watches without intervening, and experiments, where the researcher deliberately imposes a treatment. The critical feature of a well-designed experiment is random assignment, which balances confounding variables across groups. Only experiments with random assignment can support claims of causation; observational studies can only reveal associations.

When evaluating any study, ask three questions: (1) Did the researcher impose a treatment? (2) Was random assignment used to form groups? (3) Are there confounding variables that could explain the results? Remember that random sampling (how subjects are selected) affects whether results generalize to a population, while random assignment (how subjects are divided into groups) determines whether causation can be claimed. Use association language for observational studies and causal language only for experiments.

Varsity Tutors • Math 3 • Observational Studies vs. Experiments