Historical Context & Motivation
For centuries, people drew conclusions about large groups by studying just a handful of individuals. Doctors tested treatments on whichever patients happened to be available, and pollsters surveyed people who were easy to reach. The results were often wildly misleading because these convenience-based approaches introduced hidden biases that skewed the findings. The development of random sampling and random assignment transformed statistics from educated guesswork into a rigorous science capable of producing trustworthy evidence.
These historical episodes highlight a central question in statistics: How do we design a study so that our conclusions are actually valid? The answer depends on two distinct uses of randomness — one for choosing who to study and another for deciding what treatment each participant receives. Understanding the difference is essential for evaluating any statistical claim you encounter in the news, in science class, or on social media.
Core Principles & Definitions
Before diving into examples, you need to grasp four foundational ideas that anchor every well-designed study. These concepts work together to determine which conclusions are valid and which are not. A population is the entire group you want to learn about — for example, all juniors at your school. A sample is the subset of the population you actually collect data from. How you select that sample and how you structure the study determine everything about the strength of your conclusions.
Random Sampling
Random Assignment
Generalization
Causation
Visual Explanation
The diagram below illustrates the key difference between random sampling and random assignment. On the left, you see how subjects are chosen from a population. On the right, you see how those chosen subjects are placed into groups within an experiment. Notice that these are two completely separate steps that serve two completely different purposes.
As the diagram makes clear, random sampling addresses the question "Who do we study?" while random assignment addresses the question "How do we structure the study once we have our participants?" A study might use one technique without the other, both, or neither. Each combination leads to a different set of valid conclusions, which we will explore in the next sections.
How It Works — The Four Study Designs
When researchers plan a study, the presence or absence of random sampling and random assignment creates a two-by-two grid of possible designs. Each design supports a different level of conclusion. Understanding this grid is one of the most powerful tools you can have for evaluating research claims.
| Design | Random Sampling? | Random Assignment? | Valid Conclusions |
|---|---|---|---|
| Design A | Yes | Yes | Generalize AND establish causation |
| Design B | Yes | No | Generalize only (associations, not causes) |
| Design C | No | Yes | Establish causation only (for participants studied) |
| Design D | No | No | Neither — only describes what was observed |
Design A is the gold standard — it uses both forms of randomness, so you can say the results apply to the broader population and that the treatment caused the effect. Design B (like a well-conducted survey with random sampling but no experiment) lets you generalize but not claim causation. Design C (like many lab experiments using volunteer subjects) lets you claim causation for those specific participants but not generalize to the whole population. Design D (like an informal poll of your friends) supports neither conclusion.
Detailed Scenarios & Classification
The best way to master these concepts is to practice classifying real-world scenarios. The diagram below presents four study scenarios and maps each one to the appropriate design type from our grid. Pay close attention to whether participants were selected randomly from a larger population and whether they were randomly assigned to different conditions.
Notice how Design C, which is extremely common in science classrooms and university labs, allows causal claims but only for the people actually in the study. For instance, if a psychology experiment uses college sophomores who volunteered, the researchers can say their treatment caused a particular result for those students, but they cannot confidently generalize that result to all adults everywhere. On the other hand, Design B (like a professionally conducted survey) gives you a trustworthy picture of an entire population but cannot tell you why variables are related.
Worked Example
Let's walk through a complete scenario to practice identifying the study design and determining which conclusions are valid.
Strengths & Limitations
Both random sampling and random assignment are powerful tools, but neither is perfect in practice. Understanding the trade-offs helps you evaluate research more critically and design better studies when the opportunity arises.
| Feature | Random Sampling | Random Assignment |
|---|---|---|
| Purpose | Make the sample representative of the population | Equalize groups so differences can be attributed to treatment |
| Enables | Generalization | Causal inference |
| When it occurs | Before the study — during participant selection | During the study — when assigning to treatment groups |
| Strength | Eliminates selection bias; results apply broadly | Controls for confounding variables without having to identify them |
| Limitation | Expensive and difficult for large or hard-to-reach populations | Sometimes unethical or impractical (you can't randomly assign people to smoke cigarettes) |
| Common example | National surveys, census sampling | Clinical drug trials, A/B testing |
Connection to Advanced Statistical Reasoning
The ideas of random sampling and random assignment are the foundation for more advanced topics you may encounter in AP Statistics, college-level courses, or professional research. The table below shows how these foundational concepts connect to more sophisticated statistical ideas.
| Concept in This Lesson | Advanced Extension |
|---|---|
| Random sampling eliminates selection bias | In AP Statistics, you learn to calculate margin of error and confidence intervals, which quantify exactly how much uncertainty remains even after random sampling. |
| Random assignment controls confounding variables | Advanced courses introduce hypothesis testing and p-values, which formalize whether an observed difference between groups is large enough to be considered statistically significant. |
| Design B (survey with random sampling) | In college, you study stratified and cluster sampling — more sophisticated random sampling methods used to improve efficiency while maintaining representativeness. |
| Design C (experiment without random sampling) | Researchers sometimes use techniques like replication across diverse populations to partially compensate for the lack of random sampling. |
One particularly important idea for your future studies is the distinction between observational studies and experiments. An observational study measures variables without intervening — no treatment is applied, so no random assignment occurs. An experiment deliberately imposes a treatment on subjects. Because only experiments can use random assignment, only experiments can establish causation. This is why scientists say "correlation does not imply causation" — an observational study, no matter how large, cannot prove that one thing causes another.
Practice Problems
Lesson Summary
Random sampling is the process of selecting participants so that every member of the population has a known chance of being chosen. It produces a representative sample and enables generalization — the ability to extend conclusions from the sample to the entire population. Random assignment is the process of placing participants into treatment and control groups using a random mechanism. It balances confounding variables across groups and enables causal inference — the ability to conclude that the treatment caused the observed effect.
Together, these two tools create four possible study designs. Design A (both) supports generalization and causation. Design B (random sampling only) supports generalization. Design C (random assignment only) supports causation for participants studied. Design D (neither) supports only descriptive claims. When evaluating any statistical study, always ask two questions: Were participants randomly selected? Were they randomly assigned? The answers determine which conclusions are valid.