ACT MATH • PREPARING FOR HIGHER MATH

Data Collection Methods

Understand how surveys, experiments, and observational studies shape the data behind every statistical conclusion.

Historical Context & Motivation

Long before statisticians developed formal methods, people collected data to make decisions. Ancient civilizations conducted censuses to count populations, track grain harvests, and levy taxes. Over centuries, the question shifted from what to collect to how to collect it reliably. The way data is gathered determines whether the conclusions drawn from it are trustworthy, and that insight sits at the heart of modern statistics and the ACT's Preparing for Higher Math domain.

3800 BCE
Ancient Censuses
Babylonian and Egyptian civilizations conduct some of the earliest known population counts and agricultural inventories to guide taxation and resource allocation.
1747
First Clinical Trial
Scottish physician James Lind tests six treatments for scurvy aboard HMS Salisbury, establishing the controlled experiment as a powerful data collection method.
1936
The Literary Digest Fiasco
A flawed mail-in survey of 2.4 million people wrongly predicted the U.S. presidential election, demonstrating that sampling bias can destroy even massive datasets.
1948
Framingham Heart Study
This landmark observational study begins tracking thousands of residents for cardiovascular risk factors, running for decades and reshaping public health.
2000s–Present
Big Data & Digital Surveys
The internet enables rapid, large-scale data collection through online surveys, A/B testing on websites, and real-time sensor data, raising new questions about privacy and bias.

The central question that drives this entire topic is straightforward: How can we gather information so that the conclusions we draw actually reflect reality? On the ACT, you will encounter questions that ask you to identify which data collection method was used, evaluate whether the design supports a causal claim, or spot sources of bias. Mastering these ideas is essential for the statistics and probability strand of the test.

Core Principles & Definitions

Before diving into specific methods, you need a clear vocabulary. The ACT tests whether you understand the differences among the three major data collection strategies and the concepts that underpin each one. Every study begins with a population — the entire group you want to learn about — and typically examines a sample, a smaller subset chosen from that population. How you select that sample and what you do with it defines your method.

1

Observational Study

Researchers observe and record data without manipulating any variables. They watch what naturally happens. Because nothing is controlled, observational studies can show association but cannot establish causation.
2

Survey (Sample Survey)

Researchers ask questions of a selected group to collect self-reported data. A well-designed survey uses random sampling so results can be generalized to the broader population. Surveys are common but susceptible to response bias.
3

Experiment

Researchers deliberately impose a treatment on subjects and measure the response. By using random assignment, a control group, and controlled conditions, experiments can establish cause-and-effect relationships.
4

Random Sampling

Every member of the population has a known, nonzero chance of being selected. This minimizes selection bias and allows researchers to generalize results from the sample to the whole population.
5

Random Assignment

Once subjects are in an experiment, they are placed into treatment or control groups by chance. Random assignment helps ensure that confounding variables are balanced across groups, supporting causal conclusions.
KEY TAKEAWAY
Think of data collection like cooking. An observational study is like tasting dishes at a potluck — you notice what's good but can't tell why. A survey is like asking everyone at the potluck for their recipes — you gather lots of opinions, but people might exaggerate. An experiment is like entering your own kitchen, changing one ingredient at a time, and tasting the result — now you can figure out exactly what caused the difference in flavor.

Visual Explanation — The Data Collection Landscape

Follow the decision path from the top. If a treatment is imposed, the study is an experiment. If no treatment is imposed and subjects answer questions, it is a survey. If researchers simply watch and record, it is an observational study. This flowchart mirrors the logic ACT questions use when describing a research scenario.

When you encounter a statistics question on the ACT, your first job is to classify the study design. The flowchart above captures the two critical questions: (1) Was a treatment imposed? and (2) Were subjects questioned? Once you answer those, you can determine whether the study can claim causation or only association. This distinction is tested repeatedly on the ACT, so practice identifying it quickly.

How Each Method Works — Key Mechanics

Experiments: The Gold Standard for Causation

An experiment has several essential components. First, researchers identify an explanatory variable (also called the independent variable) that they will deliberately change, and a response variable (dependent variable) that they will measure. Subjects are divided into a treatment group and a control group using random assignment. The control group either receives no treatment or a placebo. Because random assignment balances out lurking variables across groups, any observed difference in the response variable can be attributed to the treatment.

Surveys: Collecting Opinions and Self-Reports

A well-designed survey selects participants through random sampling — for instance, using a simple random sample (SRS) where every individual in the population has an equal chance of being selected. The survey then poses carefully worded questions. Results can be generalized to the population, but they are subject to response bias (people may not answer truthfully) and nonresponse bias (certain groups may refuse to participate). Surveys do not manipulate variables, so they cannot establish causation.

Observational Studies: Watching Without Interfering

In an observational study, researchers record data on subjects as they naturally behave. No treatment is applied. A classic example: tracking whether students who eat breakfast earn higher GPAs. The researcher does not assign who eats breakfast; students self-select. This makes confounding variables — hidden factors like overall health habits or socioeconomic status — a serious concern. Observational studies can reveal correlations and associations but cannot prove that one variable causes another.

💡 ACT TIP
If an ACT question says the study shows that variable X causes variable Y, check whether the study used random assignment. If it did not, the correct answer will describe an association, not a causal relationship. The ACT loves to test this distinction.

Bias, Sampling Methods, and Study Design Details

Understanding data collection methods also means recognizing what can go wrong. Bias is any systematic error that causes your sample results to differ from the truth about the population. On the ACT, you may be asked to identify a source of bias or explain why a study's conclusions are limited.

The left column lists common forms of bias that weaken study conclusions. The right column shows sampling methods. Notice that a convenience sample carries a warning — it almost always introduces selection bias.

On the ACT, pay close attention to the wording of the study description. If the problem says participants were "volunteers" or "selected from the researcher's class," that is a convenience sample, which limits generalizability. If participants were "randomly selected from all students in the district," that is a random sample, and you can generalize results to the district. These phrases are the test-maker's clues — learn to spot them.

Worked Example — Classifying a Study and Drawing Conclusions

Let's walk through an ACT-style scenario step by step. This mirrors the kind of reasoning you need on test day.

📋 SCENARIO
A school nurse randomly selects 200 students from all 1,500 students at a high school. She asks each of the 200 students how many hours of sleep they got last night and records their most recent math test score. She finds that students who slept more than 7 hours scored an average of 12 points higher than students who slept fewer than 7 hours. The nurse concludes that sleeping more than 7 hours causes higher math scores.
Analyzing the Nurse's Study
1
Step 1 — Identify the Data Collection MethodAsk the key question: Did the nurse impose a treatment? No — she did not assign students to sleep a certain number of hours. She simply asked a question and recorded existing data. Since questions were involved but no treatment was applied, this is a survey combined with an observational study (she observed the test scores from school records).
Method: Observational study with survey-collected data
2
Step 2 — Evaluate the Sampling MethodThe nurse randomly selected 200 students from the entire school of 1,500. This is a simple random sample (SRS). Because the sample was randomly chosen, the results can be generalized to all 1,500 students at the school.
Random sample → Results generalize to the school population
3
Step 3 — Determine Whether Causation Can Be ClaimedSince no treatment was imposed and there was no random assignment to treatment/control groups, confounding variables may explain the association. Perhaps students who sleep more also have fewer extracurricular obligations, less stress, or better study habits. The study shows an association between sleep and math scores but cannot prove causation.
No causation — only association, because no random assignment was used
4
Step 4 — Evaluate the Nurse's ConclusionThe nurse's claim that sleep causes higher scores is not supported by this study design. A valid conclusion would be: "Among the 1,500 students at this school, there is an association between sleeping more than 7 hours and scoring higher on math tests."
Correct conclusion: Association between sleep and scores for the school population
🔑 REMEMBER THE TWO-PART TEST
On the ACT, always ask two questions about any study: (1) Was random sampling used? If yes, you can generalize to the population. (2) Was random assignment used? If yes, you can claim causation. Many ACT questions hinge on students confusing these two ideas.

Comparing Data Collection Methods — Strengths & Limitations

Comparison of the three major data collection methods
FeatureObservational StudySurveyExperiment
Treatment imposed?NoNoYes
Can establish causation?NoNoYes (with random assignment)
Can generalize to population?Only if random samplingYes, if random samplingOnly if random sampling
Main strengthEthical for sensitive topics; inexpensiveEfficient for large populationsControls confounding variables
Main limitationConfounding variablesResponse & nonresponse biasMay be unethical or impractical
ExampleTracking exercise habits vs. heart diseasePolling voters before an electionTesting a new drug vs. placebo
⚖️ CHOOSING THE RIGHT METHOD
Imagine you want to know whether a new energy drink improves test scores. An observational study just watches who drinks it and who doesn't — but maybe the motivated students are the ones buying it. A survey asks students if they drink it and how they scored — but they might not report honestly. An experiment randomly assigns half to drink it and half to drink a placebo, then compares scores. Only the experiment can tell you the drink truly made a difference.

Connecting to Statistical Inference

Data collection methods form the foundation for everything else in statistics. Once data is collected, statisticians use statistical inference — tools like confidence intervals and hypothesis tests — to draw conclusions. But those tools only work correctly if the data was collected properly. A beautifully calculated confidence interval is meaningless if the underlying sample was biased.

How today's concepts connect to advanced topics
Concept in This LessonWhere It Leads in Advanced Statistics
Random samplingEnables margin of error calculations and confidence intervals
Random assignmentJustifies using hypothesis tests to claim a treatment effect
Confounding variablesMotivates regression analysis, which statistically controls for confounders
Sample sizeLarger samples reduce variability and produce narrower confidence intervals
Bias identificationKey in AP Statistics and college research methods courses

While the ACT focuses on identifying study types and evaluating conclusions, these same ideas are central to AP Statistics and any college science course. Mastering them now gives you a strong advantage. Remember that the ACT is testing your reasoning about data — not your ability to calculate complex formulas. If you can correctly classify a study and identify its limitations, you can answer these questions quickly and confidently on test day.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher observes that people who own dogs tend to have lower blood pressure than people who do not own dogs. Can the researcher conclude that owning a dog causes lower blood pressure? Why or why not?
PROBLEM 2BASIC CALCULATION
A school wants to survey 50 of its 800 students about lunch preferences. They assign each student a number from 1 to 800, then use a random number generator to pick 50 numbers. What type of sampling method is this, and can the results be generalized to all 800 students?
PROBLEM 3INTERMEDIATE
A pharmaceutical company tests a new headache medication. They recruit 300 volunteers and randomly assign 150 to receive the medication and 150 to receive a sugar pill (placebo). Neither the participants nor the doctors know who received which pill. After two weeks, the medication group reported 40% fewer headaches. Identify: (a) the type of study, (b) the explanatory variable, (c) the response variable, and (d) whether a causal conclusion is justified.
PROBLEM 4APPLIED
A city council wants to know whether residents support building a new park. They post a poll on the city's website and receive 2,000 responses, with 85% in favor. A council member argues that since 85% is an overwhelming majority, they should proceed. Identify at least two problems with this data collection approach and explain how the study could be improved.
PROBLEM 5CRITICAL THINKING
A researcher randomly selects 500 high school students from across the state and randomly assigns half to use a new study app for one month while the other half studies without it. At the end of the month, the app group scores an average of 8 points higher on a standardized test. Describe the two types of conclusions this study supports and explain why each is valid.

Lesson Summary

Data collection methods fall into three main categories. An observational study records data without imposing any treatment and can only show association. A survey gathers self-reported data and can generalize to a population when random sampling is used. An experiment imposes a treatment and uses random assignment to establish cause-and-effect relationships.

For the ACT, remember the two-part framework: random sampling allows you to generalize results to the broader population, while random assignment allows you to claim causation. Watch for sources of bias — including selection bias, response bias, nonresponse bias, and confounding variables — that can weaken or invalidate conclusions. Classify the study, check for randomization, and identify potential bias to answer these questions confidently on test day.

Varsity Tutors • ACT Math • Data Collection Methods