MATH 3 • STATISTICS & PROBABILITY

Random Sampling & Assignment — I can describe random sampling and random assignment and explain their roles in valid conclusions.

Understanding how randomness in study design determines what conclusions you can legitimately draw.

Historical Context & Motivation

For centuries, people drew conclusions about large groups by studying just a handful of individuals. Doctors tested treatments on whichever patients happened to be available, and pollsters surveyed people who were easy to reach. The results were often wildly misleading because these convenience-based approaches introduced hidden biases that skewed the findings. The development of random sampling and random assignment transformed statistics from educated guesswork into a rigorous science capable of producing trustworthy evidence.

1747
Lind's Scurvy Trial
James Lind tested six treatments on sailors with scurvy. Although he didn't use true randomization, his comparative approach planted the seed for controlled experiments in medicine.
1936
The Literary Digest Poll Disaster
The magazine predicted Alf Landon would defeat FDR in a landslide, polling 2.4 million people. The prediction failed spectacularly because respondents were drawn from telephone and car-ownership lists — a biased, non-random sample that overrepresented wealthy voters.
1948
First Randomized Controlled Trial
The British Medical Research Council conducted the first modern randomized controlled trial (RCT), testing streptomycin for tuberculosis. Patients were randomly assigned to treatment or control groups, setting the gold standard for medical research.
1965
Federal Survey Standards
Governments worldwide adopted probability-based sampling for census and economic surveys, ensuring that conclusions about entire populations could be drawn with known margins of error.

These historical episodes highlight a central question in statistics: How do we design a study so that our conclusions are actually valid? The answer depends on two distinct uses of randomness — one for choosing who to study and another for deciding what treatment each participant receives. Understanding the difference is essential for evaluating any statistical claim you encounter in the news, in science class, or on social media.

Core Principles & Definitions

Before diving into examples, you need to grasp four foundational ideas that anchor every well-designed study. These concepts work together to determine which conclusions are valid and which are not. A population is the entire group you want to learn about — for example, all juniors at your school. A sample is the subset of the population you actually collect data from. How you select that sample and how you structure the study determine everything about the strength of your conclusions.

1

Random Sampling

Every member of the population has a known, non-zero chance of being selected for the study. This makes the sample representative of the population, allowing you to generalize your findings.
2

Random Assignment

Once participants are in the study, they are assigned to treatment or control groups using a random process (like a coin flip). This creates groups that are roughly equal in all characteristics, so any observed difference can be attributed to the treatment itself.
3

Generalization

The ability to extend conclusions from a sample to the broader population. This requires random sampling. Without it, your results may only describe the people you studied, not the population.
4

Causation

The ability to conclude that one variable actually caused a change in another. This requires random assignment. Without it, you can only identify associations, not cause-and-effect relationships.
KEY TAKEAWAY
Think of random sampling and random assignment as two separate doors in a study's design. Random sampling is the door that lets you say "this applies to everyone" (generalization). Random assignment is the door that lets you say "this treatment caused that effect" (causation). A study can open one door, both doors, or neither — and the doors it opens determine which conclusions are legitimate.

Visual Explanation

The diagram below illustrates the key difference between random sampling and random assignment. On the left, you see how subjects are chosen from a population. On the right, you see how those chosen subjects are placed into groups within an experiment. Notice that these are two completely separate steps that serve two completely different purposes.

The left column shows how random sampling selects a representative sample from the population, enabling generalization. The right column shows how random assignment distributes participants into treatment and control groups, enabling causal conclusions.

As the diagram makes clear, random sampling addresses the question "Who do we study?" while random assignment addresses the question "How do we structure the study once we have our participants?" A study might use one technique without the other, both, or neither. Each combination leads to a different set of valid conclusions, which we will explore in the next sections.

How It Works — The Four Study Designs

When researchers plan a study, the presence or absence of random sampling and random assignment creates a two-by-two grid of possible designs. Each design supports a different level of conclusion. Understanding this grid is one of the most powerful tools you can have for evaluating research claims.

The four study designs based on the presence or absence of random sampling and random assignment.
DesignRandom Sampling?Random Assignment?Valid Conclusions
Design AYesYesGeneralize AND establish causation
Design BYesNoGeneralize only (associations, not causes)
Design CNoYesEstablish causation only (for participants studied)
Design DNoNoNeither — only describes what was observed

Design A is the gold standard — it uses both forms of randomness, so you can say the results apply to the broader population and that the treatment caused the effect. Design B (like a well-conducted survey with random sampling but no experiment) lets you generalize but not claim causation. Design C (like many lab experiments using volunteer subjects) lets you claim causation for those specific participants but not generalize to the whole population. Design D (like an informal poll of your friends) supports neither conclusion.

💡 Why Does Random Assignment Enable Causation?
When you randomly assign participants to groups, you are distributing all potential confounding variables — age, health, motivation, prior knowledge — roughly equally across both groups. This means the only systematic difference between the groups is the treatment itself. If the treatment group performs differently, you can logically attribute that difference to the treatment rather than to some lurking variable.

Detailed Scenarios & Classification

The best way to master these concepts is to practice classifying real-world scenarios. The diagram below presents four study scenarios and maps each one to the appropriate design type from our grid. Pay close attention to whether participants were selected randomly from a larger population and whether they were randomly assigned to different conditions.

The 2 × 2 grid of study designs. Each quadrant shows whether random sampling and/or random assignment were used, along with a concrete scenario and the strongest valid conclusion.

Notice how Design C, which is extremely common in science classrooms and university labs, allows causal claims but only for the people actually in the study. For instance, if a psychology experiment uses college sophomores who volunteered, the researchers can say their treatment caused a particular result for those students, but they cannot confidently generalize that result to all adults everywhere. On the other hand, Design B (like a professionally conducted survey) gives you a trustworthy picture of an entire population but cannot tell you why variables are related.

Worked Example

Let's walk through a complete scenario to practice identifying the study design and determining which conclusions are valid.

Does Listening to Classical Music Improve Test Scores?
1
Step 1 — Read the ScenarioA school district wants to know whether playing classical music during study halls improves student performance on standardized tests. They randomly select 200 students from the district's enrollment list. Then, they randomly assign 100 of those students to study halls where classical music is played and 100 to study halls with no music. After eight weeks, they compare average test scores between the two groups.
2
Step 2 — Identify Random SamplingThe 200 students were randomly selected from the district's enrollment list. This means every student in the district had a chance of being chosen, making the sample representative of the entire district. Random sampling is present.
Random Sampling: ✓ Yes
3
Step 3 — Identify Random AssignmentThe 200 selected students were then randomly assigned to either the music group or the no-music group. This means the groups should be balanced in terms of prior ability, socioeconomic status, motivation, and every other characteristic. Any difference in outcomes can be attributed to the music. Random assignment is present.
Random Assignment: ✓ Yes
4
Step 4 — Classify the DesignSince both random sampling and random assignment are used, this is Design A — the gold standard.
Design Type: A (Both)
5
Step 5 — State Valid ConclusionsIf the music group scores significantly higher, the district can conclude that classical music caused improved test performance (because of random assignment) and that this effect likely applies to all students in the district (because of random sampling).
Valid conclusion: Classical music causes improved test scores for students across the district.
🔄 What If the Design Changed?
Suppose the district had used only volunteers instead of randomly selecting students. Then random sampling would be absent, making it Design C. You could still claim causation (the music caused the effect), but you could not generalize the result to all district students, because volunteers might be systematically different from the overall student body — perhaps more motivated or academically inclined.

Strengths & Limitations

Both random sampling and random assignment are powerful tools, but neither is perfect in practice. Understanding the trade-offs helps you evaluate research more critically and design better studies when the opportunity arises.

Comparing random sampling and random assignment across key features.
FeatureRandom SamplingRandom Assignment
PurposeMake the sample representative of the populationEqualize groups so differences can be attributed to treatment
EnablesGeneralizationCausal inference
When it occursBefore the study — during participant selectionDuring the study — when assigning to treatment groups
StrengthEliminates selection bias; results apply broadlyControls for confounding variables without having to identify them
LimitationExpensive and difficult for large or hard-to-reach populationsSometimes unethical or impractical (you can't randomly assign people to smoke cigarettes)
Common exampleNational surveys, census samplingClinical drug trials, A/B testing
KEY TAKEAWAY
Random sampling and random assignment are like the lenses on a pair of binoculars — each one sharpens a different aspect of your view. Random sampling sharpens your view of the population (who your results apply to). Random assignment sharpens your view of causation (whether the treatment actually worked). You get the clearest picture when both lenses are in focus, but even one lens is better than none.

Connection to Advanced Statistical Reasoning

The ideas of random sampling and random assignment are the foundation for more advanced topics you may encounter in AP Statistics, college-level courses, or professional research. The table below shows how these foundational concepts connect to more sophisticated statistical ideas.

How this lesson's concepts connect to more advanced statistical reasoning.
Concept in This LessonAdvanced Extension
Random sampling eliminates selection biasIn AP Statistics, you learn to calculate margin of error and confidence intervals, which quantify exactly how much uncertainty remains even after random sampling.
Random assignment controls confounding variablesAdvanced courses introduce hypothesis testing and p-values, which formalize whether an observed difference between groups is large enough to be considered statistically significant.
Design B (survey with random sampling)In college, you study stratified and cluster sampling — more sophisticated random sampling methods used to improve efficiency while maintaining representativeness.
Design C (experiment without random sampling)Researchers sometimes use techniques like replication across diverse populations to partially compensate for the lack of random sampling.

One particularly important idea for your future studies is the distinction between observational studies and experiments. An observational study measures variables without intervening — no treatment is applied, so no random assignment occurs. An experiment deliberately imposes a treatment on subjects. Because only experiments can use random assignment, only experiments can establish causation. This is why scientists say "correlation does not imply causation" — an observational study, no matter how large, cannot prove that one thing causes another.

🔮 Looking Ahead
In AP Statistics, you will learn to perform significance tests that calculate the probability of obtaining your observed results if the treatment had no real effect. These tests build directly on the logic of random assignment. Without random assignment, the mathematical assumptions behind these tests break down, which is why study design matters so much.

Practice Problems

PROBLEM 1CONCEPTUAL
In your own words, explain the difference between random sampling and random assignment. Which one allows a researcher to generalize results to a broader population, and which one allows a researcher to establish a cause-and-effect relationship?
PROBLEM 2BASIC CALCULATION
A university randomly selects 300 students from its enrollment database. It then surveys them about how many hours per week they exercise and what their GPA is. The study finds that students who exercise more tend to have higher GPAs. Identify whether random sampling and random assignment are present, classify the study design (A, B, C, or D), and state the strongest valid conclusion.
PROBLEM 3INTERMEDIATE
A pharmaceutical company wants to test a new allergy medication. They recruit 150 volunteers who respond to an online ad. They randomly assign 75 volunteers to take the new medication and 75 to take a placebo. After six weeks, they compare symptom severity between the two groups. (a) Classify the study design. (b) If the medication group shows significantly fewer symptoms, what can the researchers validly conclude? (c) What can they NOT conclude, and why?
PROBLEM 4APPLIED
Your school principal wants to determine whether a new tutoring program improves math test scores for all 800 students at the school. Design a study that would allow the principal to both generalize the results to all 800 students AND establish that the tutoring program caused any improvement. Describe specifically how you would use random sampling and random assignment.
PROBLEM 5CRITICAL THINKING
A news article reports: "A study of 10,000 people found that those who drink green tea daily have a 20% lower risk of heart disease." The article concludes that drinking green tea prevents heart disease. Critique this conclusion. What information about the study's design would you need to know before accepting this claim? Explain how random sampling and random assignment (or the lack thereof) affect the validity of the article's conclusion.

Lesson Summary

Random sampling is the process of selecting participants so that every member of the population has a known chance of being chosen. It produces a representative sample and enables generalization — the ability to extend conclusions from the sample to the entire population. Random assignment is the process of placing participants into treatment and control groups using a random mechanism. It balances confounding variables across groups and enables causal inference — the ability to conclude that the treatment caused the observed effect.

Together, these two tools create four possible study designs. Design A (both) supports generalization and causation. Design B (random sampling only) supports generalization. Design C (random assignment only) supports causation for participants studied. Design D (neither) supports only descriptive claims. When evaluating any statistical study, always ask two questions: Were participants randomly selected? Were they randomly assigned? The answers determine which conclusions are valid.

Varsity Tutors • Math 3 • Random Sampling & Assignment