MATH 3 • STATISTICS & PROBABILITY

Evaluating Study Conclusions — I can interpret survey/experiment descriptions and evaluate whether conclusions are justified.

Learn to separate valid conclusions from flawed reasoning when reading surveys and experiments.

Historical Context & Motivation

Every day, you encounter claims backed by data: a news headline says a new study "proves" that eating breakfast improves test scores, or an advertisement claims that 9 out of 10 dentists recommend a specific toothpaste. But how do you know whether those conclusions are actually supported by the data? The ability to evaluate study conclusions is one of the most practical skills in statistics, and its development has a rich history tied to the rise of modern science.

1747
First Clinical Trial
James Lind conducted one of the first controlled experiments on sailors with scurvy. He divided subjects into groups receiving different treatments and observed which group recovered, establishing the power of comparison groups.
1936
Literary Digest Poll Failure
The Literary Digest magazine surveyed 2.4 million people and incorrectly predicted that Alf Landon would defeat Franklin Roosevelt. The sampling bias — surveying mostly wealthy Americans — led to a famously wrong conclusion, proving that a large sample does not guarantee accuracy.
1948
Randomized Controlled Trials
The British Medical Research Council conducted the first randomized controlled trial (RCT) to test streptomycin for tuberculosis. Random assignment became the gold standard for establishing cause and effect.
2005–present
Replication Crisis
Researchers found that many published studies could not be replicated, sparking a worldwide conversation about statistical validity and the importance of critically evaluating conclusions rather than accepting them at face value.

These historical moments all point to the same fundamental question: Does the design of a study actually support the conclusion being drawn? In this lesson, you will learn the framework for answering that question with confidence.

Core Principles & Definitions

Before you can evaluate whether a conclusion is justified, you need to understand the building blocks of statistical studies. Every study has a design, and that design determines what kinds of conclusions are valid. Let's break down the key terms and ideas.

1

Observational Study vs. Experiment

An observational study watches and records data without interfering. An experiment deliberately applies a treatment to subjects. Only experiments can establish causation.
2

Random Sampling

Random sampling means every individual in the population has a known chance of being selected. This allows you to generalize findings to the larger population.
3

Random Assignment

Random assignment places subjects into treatment or control groups by chance. This controls for confounding variables and allows researchers to claim cause and effect.
4

Confounding Variable

A confounding variable is an outside factor that influences both the explanatory and response variables. If not controlled, it makes it impossible to determine true cause and effect.
5

Scope of Inference

The scope of inference describes who the findings apply to and whether a causal claim can be made. It depends entirely on how the study was designed — specifically, whether random sampling and random assignment were used.
KEY TAKEAWAY
Think of study design like a courtroom. Random sampling is like selecting a jury that represents the whole community — it lets you generalize. Random assignment is like a fair trial that eliminates bias — it lets you assign blame (causation). Without the right jury, your verdict only applies to the people in the room. Without a fair trial, you can't be sure the defendant is guilty.

Visual Explanation — The Scope of Inference Chart

The relationship between random sampling, random assignment, and the type of conclusion you can draw is best understood through a two-by-two chart. This visual is the single most important tool for evaluating study conclusions. Study it carefully — it will be your go-to reference.

The scope of inference chart shows four possible scenarios. The top-left cell (both random sampling and random assignment) is the strongest design. The bottom-right cell (neither) gives the weakest conclusions. Most real-world studies fall somewhere in between.

When you read a study description on a test or in the real world, your first move should be to identify which cell of this chart the study falls into. Ask yourself two questions: Was the sample randomly selected from a population? and Were subjects randomly assigned to treatment groups? The answers determine everything about what the study can and cannot conclude.

How Study Design Determines Valid Conclusions

The Two Key Questions Framework

Evaluating a study conclusion is a systematic process. You don't need formulas — you need a logical framework. Here is the step-by-step reasoning process you should follow every time you encounter a study and its conclusion.

QUESTION 1 — GENERALIZATION
Random Sampling → Can generalize to the population
If the study used a random sample from a defined population, the results can be extended beyond the study participants to that entire population. If the sample was a convenience sample (e.g., volunteers, one classroom), the results apply only to those specific subjects.
QUESTION 2 — CAUSATION
Random Assignment → Can claim cause and effect
If subjects were randomly assigned to treatment and control groups, confounding variables are balanced across groups. This lets the researcher attribute differences in outcomes to the treatment itself. Without random assignment, the conclusion must be limited to an association (correlation), not causation.

Common Conclusion Errors

The most frequent mistake in evaluating studies is claiming causation from an observational study. For example, if a survey finds that students who eat breakfast score higher on tests, you cannot conclude that eating breakfast causes higher scores. A confounding variable — perhaps family income — might explain both behaviors. Another common error is generalizing from a non-random sample. If you survey only students at one school, you cannot claim the results represent all teenagers nationwide.

💡 Remember This Phrase
"Correlation does not imply causation." This is the single most important sentence in introductory statistics. An observational study can show that two things are related, but only a well-designed experiment with random assignment can show that one causes the other.

Types of Studies & Their Valid Conclusions

Not all studies are created equal. Understanding the different types of study designs helps you quickly identify what conclusions are valid. The diagram below shows the decision tree you can use to classify any study you encounter.

This decision tree classifies any study. Start at the top: does the researcher impose a treatment? If yes, it is an experiment. If no, it is an observational study. Then check for random assignment and random sampling to determine the valid scope of inference.
Summary of study types and the conclusions each can support
Study TypeKey FeatureValid Conclusion
Randomized ExperimentTreatment imposed + random assignment to groupsCause-and-effect (and generalization if random sampling is also used)
Non-Randomized ExperimentTreatment imposed but groups are pre-existing (e.g., period 1 vs. period 3)Weak causal claim — confounders may not be balanced
Sample SurveyNo treatment; data collected from a sample of a populationAssociation only; can generalize if random sample
Prospective StudyFollows subjects forward in time, recording outcomesAssociation only — no causation
Retrospective StudyLooks backward at existing records or past dataAssociation only — no causation

Worked Example — Evaluating a Study Conclusion

Let's walk through a complete example of reading a study description and deciding whether its conclusion is justified.

📋 Study Description
A researcher at a large university wanted to know whether listening to classical music while studying improves exam performance. She posted flyers around campus asking for volunteers. Sixty students signed up. She randomly assigned 30 students to study with classical music and 30 to study in silence. After one week, both groups took the same exam. The music group scored an average of 82%, while the silent group scored 76%. The researcher concluded: "Classical music causes higher exam scores for college students."
Evaluating the Conclusion
1
Step 1 — Identify the study typeThe researcher imposed a treatment (listening to classical music vs. silence) on the subjects. Because a treatment was deliberately applied, this is an experiment, not an observational study.
Study type: Experiment
2
Step 2 — Check for random assignmentThe description says the researcher "randomly assigned" 30 students to each group. This means confounding variables (like prior GPA, sleep habits, or study skills) should be roughly balanced between the two groups. Random assignment is present, so a causal conclusion about the treatment is supported.
Random assignment: Yes → Causation is supported
3
Step 3 — Check for random samplingThe students were volunteers who responded to flyers — this is a convenience sample, not a random sample of all college students. The volunteers may differ from the general student population (perhaps they are more motivated or more interested in music). Therefore, the results cannot be generalized beyond the 60 participants.
Random sampling: No → Cannot generalize to all college students
4
Step 4 — Evaluate the conclusionThe researcher claimed: "Classical music causes higher exam scores for college students." The causal language ("causes") is supported because of random assignment. However, the phrase "for college students" implies generalization to the entire population of college students, which is not justified because the sample was not randomly selected.
Conclusion is PARTIALLY justified
5
Step 5 — Write a corrected conclusionA properly worded conclusion would narrow the scope: "For the 60 volunteer students in the study, listening to classical music while studying caused higher exam scores compared to studying in silence." This preserves the causal claim (from random assignment) while limiting the population (because there was no random sampling).
Corrected: causation claim limited to the study participants

Strengths & Limitations of Different Study Designs

Each study design has trade-offs. Randomized experiments provide the strongest evidence but are not always ethical or practical. For example, you cannot randomly assign people to smoke for 20 years to study lung cancer. In such cases, observational studies are the only option. Understanding these trade-offs helps you evaluate why certain designs were chosen and what that means for the conclusions.

Comparison of randomized experiments and observational studies
FeatureRandomized ExperimentObservational Study
Can establish causation?Yes — random assignment controls for confoundersNo — confounders remain uncontrolled
Can generalize?Only if random sampling was also usedOnly if random sampling was used
Ethical flexibilityLimited — cannot impose harmful treatmentsMore flexible — simply observes existing behaviors
Cost and timeOften expensive and time-consumingOften cheaper and faster
Main weaknessMay lack ecological validity (lab setting ≠ real world)Cannot rule out confounding variables
KEY TAKEAWAY
Think of study design like a lock and key. The conclusion is the lock, and the study design is the key. A causal conclusion requires the specific key of random assignment. A generalization conclusion requires the specific key of random sampling. If you try to open a lock with the wrong key, it simply doesn't work — and neither does a conclusion that overreaches its study design.

Connections to Advanced Statistical Reasoning

The principles you've learned in this lesson form the foundation for more advanced statistical concepts that you'll encounter in AP Statistics, college courses, and real-world data literacy. Understanding scope of inference now prepares you for the deeper reasoning ahead.

How this lesson connects to future topics
This LessonAdvanced Extension
Identifying confounding variablesMultiple regression analysis — statistically controlling for confounders
Random assignment → causationA/B testing in tech companies; clinical trials in medicine
Random sampling → generalizationMargin of error and confidence intervals for population parameters
Evaluating whether a conclusion is justifiedHypothesis testing (p-values, significance levels) — quantifying the strength of evidence
Correlation ≠ causationSimpson's Paradox — where aggregated data reverses a trend seen in subgroups

In AP Statistics, you will learn to quantify how strong the evidence is using tools like p-values and confidence intervals. But even the most sophisticated statistical test cannot fix a flawed study design. If a study lacks random assignment, no amount of math can turn an association into a causal claim. The concepts from this lesson remain your first line of defense throughout your statistical career.

Practice Problems

PROBLEM 1CONCEPTUAL
A study found that people who drink coffee daily tend to have lower rates of depression. A newspaper headline reads: "Coffee Prevents Depression." Is this headline justified? Explain your reasoning by identifying the study type and the valid scope of inference.
PROBLEM 2BASIC CALCULATION
A school principal randomly selects 200 students from the entire school population and surveys them about their sleep habits. She finds that students who sleep more than 8 hours per night have higher GPAs. She concludes: "Getting more sleep is associated with higher GPAs among students at this school." Is this conclusion justified? Identify whether random sampling and random assignment were used.
PROBLEM 3INTERMEDIATE
Researchers recruited 100 volunteer adults and randomly assigned them to either a meditation group or a control group for 8 weeks. At the end, the meditation group reported significantly lower stress levels. The researchers conclude: "Meditation reduces stress in adults." Identify two problems with this conclusion and rewrite it to be properly justified.
PROBLEM 4APPLIED
A tech company wants to know if changing the color of a "Buy Now" button from blue to green increases sales. They randomly select 10,000 users from their customer database and randomly assign half to see the blue button and half to see the green button. After one week, the green-button group had 15% more purchases. The company concludes: "Changing the button to green increases sales among our customers." Evaluate this conclusion.
PROBLEM 5CRITICAL THINKING
Consider two studies about exercise and heart health. Study A: Researchers randomly sample 5,000 adults from the U.S. population and survey them about their exercise habits and heart health, finding that more exercise is associated with better heart health. Study B: Researchers recruit 200 volunteers and randomly assign half to a 12-week exercise program. The exercise group shows improved heart health markers. Which study better supports the claim "Exercise improves heart health for American adults"? Explain why neither study alone fully supports this claim, and describe what an ideal study would look like.

Lesson Summary

Evaluating study conclusions requires a systematic approach. First, determine whether the study is an experiment (treatment imposed) or an observational study (no treatment). Then ask two critical questions: Was random sampling used? If yes, the results can be generalized to the population. Was random assignment used? If yes, a cause-and-effect conclusion is supported. Without random assignment, the conclusion must be limited to an association. Without random sampling, the conclusion cannot be extended beyond the study participants.

Always watch for confounding variables in observational studies and for convenience sampling that limits generalization. Remember: correlation does not imply causation. The scope of inference chart — with random sampling on one axis and random assignment on the other — is your essential tool for matching any study's design to the conclusions it can validly support.

Varsity Tutors • Math 3 • Evaluating Study Conclusions