Statistics & Probability • Inferences & Conclusions

Surveys, Experiments & Observational Studies

Understanding how data is collected is the first step toward knowing what conclusions that data can actually support.

Why Do We Care How Data Is Collected?

Imagine you read a headline claiming "Students who eat breakfast earn higher GPAs." That sounds useful, but it leaves a critical question unanswered: does eating breakfast cause better grades, or do students who are already more organized simply tend to eat breakfast and study harder? The answer depends entirely on how the data behind the headline was gathered. Over the last two centuries, statisticians have developed distinct methods of data collection, each with its own power—and its own blind spots.

1786
John Sinclair of Scotland uses the word "statistics" for the first time in English, compiling parish-by-parish surveys of Scottish life. These early sample surveys aimed to describe a population without manipulating anything.
1882–1920s
Scientists like Louis Pasteur and later Sir Ronald A. Fisher pioneer controlled experiments—deliberately changing one factor while holding others constant—to prove cause-and-effect relationships in medicine and agriculture.
1926
Fisher publishes The Arrangement of Field Experiments, formally introducing randomization as the gold standard for assigning subjects to treatment groups, reducing bias in ways no other technique can.
1948
The landmark Framingham Heart Study begins—a long-running observational study tracking thousands of residents over decades without intervening. It reveals risk factors for heart disease through observation alone.
2010s–Present
Modern "big data" methods blend all three approaches. Understanding the differences matters more than ever, because the conclusions you can draw depend on the method used to collect the data.

The central question this lesson addresses is: given a real-world study, how do you identify what type it is, and what can (or can't) you legitimately conclude from its results?

Three Methods of Data Collection

Statistics recognizes three primary study designs. Each one answers a different kind of question and carries different rules about what conclusions are valid. Here they are, side by side.

1

Sample Survey

A study that selects a representative sample from a larger population and measures characteristics of that sample—usually through questions or measurements—without influencing the subjects in any way. The goal is to describe or estimate population parameters. Example: a political poll asking 1,200 registered voters which candidate they prefer.
2

Experiment

A study in which the researcher deliberately imposes a treatment on subjects (called experimental units) and then measures the response. Subjects are assigned to groups—at least one treatment group and typically a control group. The goal is to determine cause and effect. Example: giving one group of patients a new drug and another group a placebo to compare recovery rates.
3

Observational Study

A study that observes and records data on subjects without imposing any treatment. Unlike a survey, it often tracks variables over time or compares naturally occurring groups. The goal is to identify associations and patterns, but it cannot establish causation on its own. Example: comparing lung cancer rates between smokers and non-smokers by examining medical records.
4

Randomization (Cross-Cutting)

Randomization is not a study type—it is a technique used within study designs. In surveys it appears as random selection (choosing who is in the sample). In experiments it appears as random assignment (deciding who gets which treatment). Both forms reduce bias, but they serve different purposes, as we'll explore in Section 4.
✦ Key Takeaway
Think of a sample survey as taking a photograph of a crowd—you capture what already exists. An experiment is like a chemistry lab where you change the ingredients to see what happens. An observational study is like watching traffic from a bridge: you see patterns, but you didn't control the lights. The type of study determines the type of conclusion you're allowed to make.

Visual Guide: Identifying Study Types

The following flowchart walks you through a simple decision tree. When you encounter any study—in a textbook problem, a news article, or an AP exam question—follow these steps to classify it.

Figure 1 — Decision flowchart for classifying data collection methods

The key question is always: did somebody deliberately do something to the subjects? If yes, it's an experiment. If no, ask whether the study was designed to measure a whole population by looking at a representative sample (survey) or whether it simply records information about groups that already exist (observational study).

How Randomization Relates to Each Method

Randomization is the single most powerful tool statisticians have for reducing bias—systematic errors that push results in one direction. However, randomization shows up in different ways depending on the study type, and each form serves a distinct purpose.

Random Selection vs. Random Assignment

Random selection means using a chance process (like a random number generator) to choose which individuals from the population are included in the study. It is the hallmark of a well-designed sample survey. When every member of the population has a known, nonzero chance of being selected, the sample is likely to be representative, allowing us to generalize findings to the broader population.

Random assignment means using a chance process to decide which treatment each subject receives. It is the hallmark of a well-designed experiment. Random assignment tends to balance out all other variables—both the ones you know about and the ones you don't—across the treatment groups. This is what allows an experiment to establish causation.

In an observational study, neither random selection of a representative sample nor random assignment of treatments is typically present. Researchers simply observe subjects who have "self-selected" into different conditions (e.g., people who chose to smoke vs. those who didn't). Because of this, confounding variables—hidden factors that could explain the observed association—are always a concern, and causal claims cannot be made from observational studies alone.

Figure 2 — Random selection enables generalization; random assignment enables causal claims
✦ Key Takeaway
Think of it like jury selection in a courtroom. Random selection is like pulling jurors' names from a hat so the jury represents the whole community—that's generalizability. Random assignment is like flipping a coin to decide which jurors sit on the left side versus the right—it ensures neither side has a built-in advantage. Both are powerful, but they do different jobs.

Side-by-Side Comparison

The table below summarizes the critical distinctions. Study it carefully—questions on standardized tests often hinge on exactly these differences.

FeatureSample SurveyExperimentObservational Study
Treatment imposed?NoYesNo
Primary goalEstimate a population parameterDetermine cause and effectIdentify associations / patterns
Role of randomizationRandom selection of sample from populationRandom assignment of subjects to treatment groupsTypically none (though random sampling can appear)
Can generalize to population?Yes, if sample is randomOnly if subjects were also randomly selectedOnly if sample is representative
Can establish causation?NoYes, if properly randomizedNo — confounders may exist
Confounding variablesNot a major concern (not testing causal claims)Controlled by random assignmentMajor concern
ExampleGallup polls, census surveysClinical drug trials, A/B testingFramingham Heart Study, cohort studies

The Confounding Variable Problem

A confounding variable (or "lurking variable") is a factor that is related to both the explanatory variable and the response variable, making it impossible to tell which one is truly responsible for the observed effect. For example, in an observational study linking ice cream sales to drowning deaths, the confounding variable is temperature: hot weather increases both ice cream consumption and swimming, which increases drowning risk. Ice cream doesn't cause drowning—the study simply couldn't separate the variables because no treatment was imposed.

In a randomized experiment, confounders are handled automatically. Because subjects are split into groups by chance, every potential confounding factor—known or unknown—is likely to be distributed evenly across the groups. Any difference in outcomes can therefore be attributed to the treatment itself.

Worked Example

Read the scenario, then follow the step-by-step reasoning to classify the study and state what conclusions are valid.

Math Tutoring App Study
1
ScenarioA school district wants to know whether a new math tutoring app improves students' test scores. Researchers recruit 200 students from across the district by randomly selecting names from enrollment lists. Each selected student is then randomly assigned to either use the tutoring app for 8 weeks (treatment group: 100 students) or continue with their normal homework routine (control group: 100 students). At the end of the 8 weeks, both groups take the same standardized math test.
2
Step 1 — Was a treatment imposed?Yes. The researchers deliberately required one group to use the tutoring app while the other group did not. The researchers controlled who got the treatment. This rules out both a sample survey and an observational study.
Classification: Experiment.
3
Step 2 — Was random selection used?Yes. Students were selected from enrollment lists using a random process. This means the sample is likely representative of the district's student population.
4
Step 3 — Was random assignment used?Yes. Each student was randomly assigned to treatment or control. This means the groups should be balanced with respect to confounding variables such as prior math ability, motivation, and home environment.
5
Step 4 — What conclusions are valid?Because this study uses both random selection and random assignment, we can make the strongest possible conclusion. If the treatment group scores significantly higher:
✓ Causal claim: The tutoring app caused the improvement (random assignment). ✓ Generalization: This effect likely applies to all students in the district (random selection). If the study had used random assignment but not random selection (e.g., it only included volunteer students), we could still claim causation among the participants but could not generalize to the full district.

Strengths, Limitations & When to Use Each

No single method is "best" in all situations. Each has trade-offs, and real-world constraints often determine which approach researchers can actually use.

MethodStrengthsLimitations
Sample SurveyFast, cost-effective, can represent huge populations; allows generalization when sample is randomRelies on honest, accurate responses; nonresponse bias can skew results; cannot establish causation; question wording can introduce bias
ExperimentOnly method that can establish cause-and-effect; random assignment controls confounders; allows precise control of variablesCan be expensive and time-consuming; may involve ethical issues (e.g., you can't force people to smoke); subjects who know they're being studied may behave differently (Hawthorne effect); may lack generalizability if subjects aren't randomly selected
Observational StudyEthical when experiments aren't possible (e.g., effects of poverty, smoking); can study long-term, naturally occurring phenomena; often cheaper; can use existing dataCannot establish causation; confounding variables are always a concern; results can be misinterpreted by the public as causal
✦ Key Takeaway
When it's ethically and practically possible, a randomized experiment gives you the most powerful conclusions. But when you can't ethically impose a treatment—you can't randomly assign people to live in poverty or breathe polluted air—an observational study may be the best you can do. The key is to always match your conclusion to the strength of your study design. Think of it like a ladder: surveys describe, observational studies suggest, and experiments prove.

Connections to Advanced Theory

In a college-level statistics or AP Statistics course, these foundational ideas lead to more sophisticated topics. Here's a preview of where they connect.

This LessonAdvanced Extension
Random selection → generalizationSampling distributions & margin of error: Because the sample is random, statistical theory lets you calculate exactly how confident you should be in your estimate (e.g., "±3% with 95% confidence").
Random assignment → causationHypothesis testing & p-values: After running an experiment, you test whether the observed difference is statistically significant or could have occurred by random chance alone.
Confounding in observational studiesRegression & control variables: Advanced methods like multiple regression attempt to "statistically control" for confounders, though they can never fully replicate what random assignment achieves.
Bias in surveysStratified, cluster, & systematic sampling: Techniques for improving survey design beyond simple random sampling, reducing bias while managing cost.
Experimental design principlesBlocking, matched-pairs, & factorial designs: Ways to increase the sensitivity of experiments by grouping similar subjects before random assignment.

Understanding the three basic study types is not just an academic exercise. Every time you read a health study in the news, evaluate a product claim, or assess a social science finding, you need to ask: was this a survey, an experiment, or an observational study? Your ability to think critically about data starts here.

Practice Problems

PROBLEM 1CONCEPTUAL
A news article reports: "People who drink three or more cups of coffee per day tend to live longer than those who drink none." The study followed 400,000 adults over 14 years and tracked their coffee consumption and death rates. No one was told how much coffee to drink. Question: What type of study is this? Can the researchers conclude that coffee causes longer life? Explain your reasoning.
PROBLEM 2IDENTIFICATION
Classify each of the following as a sample survey, experiment, or observational study: (a) A cosmetics company randomly selects 500 customers from its database and emails them a questionnaire about product satisfaction. (b) A teacher gives one section of her chemistry class a new lab format and keeps the other section on the old format, then compares final exam scores. (c) A medical journal reviews hospital records to compare recovery times for patients who chose surgery versus those who chose physical therapy.
PROBLEM 3INTERMEDIATE
A researcher wants to test whether a new fertilizer increases tomato plant growth. She has 60 tomato seedlings of similar age and size. She randomly assigns 30 seedlings to receive the new fertilizer and 30 to receive the standard fertilizer. After 6 weeks, she measures the height of every plant. Question: (a) Identify the study type. (b) What is the explanatory variable? The response variable? (c) Why is random assignment important here? (d) Can the researcher generalize the results to all tomato plants everywhere? Why or why not?
PROBLEM 4APPLIED / MULTI-STEP
The principal of a high school wants to determine whether allowing students to listen to music during study hall improves their quiz scores. She is considering two plans: Plan A: Let students choose whether they want to listen to music. After a month, compare average quiz scores for listeners vs. non-listeners. Plan B: Randomly assign half the students to be allowed to listen to music and half to study in silence. After a month, compare quiz scores. Question: (a) Classify each plan. (b) Identify at least one confounding variable in Plan A. (c) Explain why Plan B would allow a stronger conclusion. (d) Even with Plan B, what additional step would be needed to generalize results to all high school students in the state?
PROBLEM 5SYNTHESIS / CRITICAL THINKING
Consider the following four study designs for investigating whether exercise reduces symptoms of anxiety in teenagers: Design I: Randomly select 500 teens from the state, survey them about their exercise habits and anxiety levels. Design II: Follow 200 volunteer teens for a year, recording their exercise and anxiety without intervening. Design III: Randomly assign 100 volunteer teens to a daily exercise program and 100 to a no-exercise control group; measure anxiety after 3 months. Design IV: Randomly select 200 teens from the state, then randomly assign them to exercise or control groups; measure anxiety after 3 months. Question: For each design, (a) classify it, (b) state whether it uses random selection, random assignment, both, or neither, and (c) state the strongest valid conclusion. Then (d) explain which design provides the most powerful evidence and why.

Lesson Summary

There are three fundamental methods of data collection in statistics. A sample survey selects a representative group from a population and asks questions or takes measurements, aiming to estimate population parameters—it describes what is, but does not explain why. An experiment deliberately imposes a treatment on subjects and uses a control group to isolate the effect of that treatment, making it the only method that can establish cause-and-effect relationships. An observational study watches and records data without any intervention, identifying associations and patterns, but always leaving open the possibility that confounding variables are driving the results.

Randomization is the thread that ties these methods together but plays different roles in each. Random selection—choosing who is in the study—enables generalization to the broader population and is the backbone of good surveys. Random assignment—choosing who gets which treatment—controls for confounders and enables causal conclusions, and is the backbone of good experiments. A study that uses both provides the most powerful evidence possible. The conclusion you are allowed to draw must always match the design of the study that produced the data.

Varsity Tutors • Statistics & Probability (Common Core) • Surveys, Experiments & Observational Studies