MATH 3 • STATISTICS & PROBABILITY

Sampling Bias — I can identify sources of bias in sampling and propose improvements to reduce bias at my level.

Learn why flawed samples lead to misleading conclusions and how to design studies that actually represent the population.

Historical Context & Motivation

Statistics has always been about making smart conclusions from limited data. When you can't ask every single person, measure every single item, or test every single case, you take a sample — a smaller group meant to represent the whole. But what happens when your sample doesn't actually represent the population? You get sampling bias, a systematic error that can lead to wildly incorrect conclusions. Some of the most famous failures in the history of polling and research stem from exactly this problem.

1936
The Literary Digest Debacle
The Literary Digest magazine polled 2.4 million people about the presidential election, yet predicted Alf Landon would crush Franklin D. Roosevelt. They were spectacularly wrong because their sample came from telephone directories and car registrations — sources that skewed toward wealthier voters during the Great Depression.
1948
Dewey Defeats Truman — or Does He?
Pollsters used quota sampling methods that systematically under-represented certain voter demographics. The Chicago Daily Tribune famously printed the headline "Dewey Defeats Truman" based on flawed poll data, only for Truman to win handily.
1976
Formalization of Sampling Theory
Statisticians like William Cochran published foundational works on sampling techniques, establishing rigorous frameworks for probability-based sampling that minimize bias. These methods became standard in government surveys and scientific research.
2016
Modern Polling Challenges
The 2016 U.S. presidential election shocked many pollsters who underestimated support for Donald Trump. Analysts later identified nonresponse bias and the under-sampling of voters without college degrees as key factors that skewed predictions.

These historic examples all raise the same core question: How do we select a sample that genuinely reflects the population we care about? Understanding the sources of bias in sampling — and knowing how to fix them — is one of the most practical skills in all of statistics.

Core Principles & Definitions

Before we can spot bias, we need to establish some key vocabulary. A population is the entire group you want to learn about — every student in your school, every fish in a lake, or every voter in a state. A sample is the subset you actually collect data from. When the sample consistently differs from the population in a way that distorts results, we call that sampling bias.

1

Voluntary Response Bias

Occurs when people choose whether to participate. Those with strong opinions (usually negative) are more likely to respond, making the sample unrepresentative of the broader population.
2

Convenience Sampling Bias

Arises when a researcher selects participants based on ease of access — like only surveying friends or people in one hallway. This ignores large portions of the population.
3

Undercoverage Bias

Happens when some groups in the population are left out of the sampling process entirely. For example, an online-only survey misses people without internet access.
4

Nonresponse Bias

Occurs when selected individuals refuse to participate or can't be reached. If nonrespondents differ systematically from respondents, the results become skewed.
5

Response Bias

Results from the way questions are worded, the presence of an interviewer, or social pressure. People may give answers they think are "correct" rather than truthful ones.
KEY TAKEAWAY
Think of sampling like taste-testing a pot of soup. If you only scoop from the top, you'll miss the flavors that sank to the bottom. A biased sample is like an unstirred pot — your "taste" doesn't represent the whole dish. To get a true picture, you need to stir the pot (randomize) and scoop from different depths (include all subgroups).

Visual Explanation — Biased vs. Unbiased Samples

On the left, a biased sample over-represents one subgroup (purple) and completely misses others. On the right, a representative sample includes all subgroups in roughly the same proportions as the population. The four colors represent different demographic subgroups within the population.

The diagram above illustrates the fundamental difference between a biased and an unbiased sample. Notice how the biased sample on the left is dominated by purple dots — it looks nothing like the actual population. The representative sample on the right, however, mirrors the population's proportions. In real research, these colored dots might represent different age groups, income levels, or geographic locations. When any of these groups is systematically excluded or over-included, the conclusions drawn from the data will be systematically off-target — not just slightly wrong, but wrong in a predictable direction.

Mathematical Framework — Quantifying Bias

While much of sampling bias is about study design rather than formulas, statistics does provide a mathematical way to think about how bias affects our estimates. Understanding these relationships helps you see why bias is such a big deal compared to random error.

TOTAL ERROR OF AN ESTIMATE
Total Error = Bias + Random Sampling Error
Total Error is the difference between the sample statistic and the true population parameter. Bias is the systematic component that does NOT decrease with larger samples. Random Sampling Error is the natural variation that DOES decrease as sample size increases.
BIAS OF A SAMPLE STATISTIC
Bias(x̄) = E(x̄) − μ
Where E(x̄) is the expected value (long-run average) of the sample mean across many samples, and μ is the true population mean. If this difference is zero, the sampling method is unbiased. If it is consistently positive or negative, the method has a systematic bias.
MARGIN OF ERROR (FOR RANDOM SAMPLES)
Margin of Error ≈ 1 / √n
Where n is the sample size. This formula shows that random error shrinks as your sample grows. However, increasing sample size does NOT reduce bias. A biased method with 10,000 people is still biased — just as the Literary Digest learned with 2.4 million respondents.
⚠️ Why Sample Size Can't Fix Bias
This is one of the most important ideas in the course. A bigger sample reduces random error but leaves systematic bias untouched. If your sampling method only reaches certain types of people, surveying more of those same types of people won't make your results more accurate. A well-designed small sample beats a poorly designed large one every time.

Types of Sampling Bias — A Closer Look

Now that you understand the general idea, let's examine each type of bias more carefully. Recognizing the specific type helps you propose the right fix. The diagram below maps out the five major types of sampling bias, their causes, and their remedies.

This flowchart maps each bias type (left, colored borders) through its cause (center) to its remedy (right, green borders). The overarching goal of all remedies is the same: ensure every member of the population has a known, non-zero chance of being selected.
Each bias type with a concrete example and the group that gets excluded or distorted.
Bias TypeReal-World ExampleWho Gets Left Out?
Voluntary ResponseA restaurant posts a QR code for a satisfaction survey on the receipt.Customers with moderate experiences — only those thrilled or furious bother scanning.
ConvenienceA student surveys people at the gym about exercise habits.People who don't go to the gym — the very group you'd need for a balanced view.
UndercoverageA phone survey uses landlines only to gauge political preferences.Younger adults who rely exclusively on cell phones.
NonresponseA mailed census form has a 40% return rate in certain neighborhoods.Residents who are busy, transient, or distrustful of government data collection.
ResponseA survey asks: "Don't you agree that our school needs more funding?"Opposing viewpoints are suppressed by the leading question's wording.

Worked Example — Identifying and Fixing Bias

Let's walk through a realistic scenario step by step. Suppose a student council wants to find out whether the school should switch to a four-day school week. They decide to post a poll on the school's Instagram story and let students vote. After 200 responses, 78% favor the four-day week. Should the council trust this result?

Analyzing the Four-Day School Week Poll
1
Step 1 — Identify the Population and SampleThe population is all students at the school (say, 1,200 students). The sample consists of the 200 students who responded to the Instagram story poll.
Population: 1,200 students. Sample: 200 Instagram respondents.
2
Step 2 — Check for Voluntary Response BiasThe poll was open for anyone to respond — no one was randomly selected. Students who feel strongly about a four-day week (excited about a day off) are much more likely to tap the poll than students who are indifferent. This is a classic case of voluntary response bias. The 78% figure almost certainly overestimates true support.
Bias identified: Voluntary response — strong opinions over-represented.
3
Step 3 — Check for Undercoverage BiasNot every student follows the school's Instagram account. Students without smartphones, students not on social media, and students who missed the 24-hour story window are all excluded. This is undercoverage bias — entire groups have zero probability of being included.
Bias identified: Undercoverage — non-Instagram users have no chance of selection.
4
Step 4 — Propose an Improved Sampling MethodInstead of an Instagram poll, the student council could use a simple random sample. Obtain the full student roster from the administration, assign each student a number, and use a random number generator to select 200 students. Then distribute the survey directly to those selected students via their school email accounts, with a follow-up reminder for nonrespondents.
Improvement: Use a random number generator on the full roster; follow up to reduce nonresponse.
5
Step 5 — Evaluate the ImprovementThe revised method eliminates voluntary response bias (students are selected, not self-selected) and undercoverage (every enrolled student is on the roster). Following up with nonrespondents reduces nonresponse bias. The 200-person sample is still manageable but now has a much better chance of reflecting the true opinions of all 1,200 students.
The improved study design addresses voluntary response, undercoverage, and nonresponse bias.

Sampling Methods — Strengths & Limitations

Knowing the types of bias is only half the battle. You also need to know which sampling methods reduce bias and which ones invite it. The table below compares common methods, noting their potential for bias and when each is appropriate.

Comparison of sampling methods by procedure and bias risk.
Sampling MethodHow It WorksBias Risk
Simple Random Sample (SRS)Every member of the population has an equal chance of being selected, typically using a random number generator.Low — the gold standard. Bias is minimal if the sampling frame is complete.
Stratified Random SampleThe population is divided into subgroups (strata) based on a key characteristic, and a random sample is drawn from each stratum.Low — guarantees subgroup representation. Excellent for diverse populations.
Systematic SampleSelect every kth individual from a list after a random starting point (e.g., every 10th name).Moderate — works well unless there's a hidden pattern in the list that matches the interval.
Convenience SampleResearcher selects whoever is easiest to reach (people nearby, friends, passersby).High — almost never representative. Results should not be generalized.
Voluntary Response SampleParticipants self-select by choosing to respond (online polls, call-in surveys).High — strongly favors people with extreme opinions. Never use for formal research.
KEY TAKEAWAY
Think of a sampling method like choosing players for a pickup basketball game. If you only invite your friends (convenience), you'll get a lopsided team. If you let anyone who shows up play (voluntary response), the most aggressive players dominate. But if you randomly draw names from the full roster (SRS) or make sure you pick from guards, forwards, and centers equally (stratified), you build a team that represents the full range of talent — and your game data is actually meaningful.

Connecting to Advanced Statistical Thinking

Understanding sampling bias at this level gives you a strong foundation, but the concept extends into more advanced territory. In college-level statistics, AP Statistics, and real-world data science, bias detection becomes even more nuanced. The table below shows how the ideas you've learned connect to concepts you may encounter next.

How current concepts connect to advanced topics in statistics.
What You Know NowWhat Comes Next
Identifying bias types (voluntary response, convenience, undercoverage, nonresponse, response)Measuring the magnitude of bias using confidence intervals, weighting adjustments, and propensity scores
Simple random samples give every individual an equal chanceCluster sampling, multi-stage sampling, and complex survey designs used by organizations like the U.S. Census Bureau
Bigger samples don't fix biasMean Squared Error (MSE) = Bias² + Variance — a formal equation that separates the two sources of error
Proposing design improvements to reduce biasExperimental design principles: randomization, blinding, control groups, and the distinction between observational studies and experiments

The key insight that carries forward is this: how you collect data matters just as much as how you analyze it. No amount of sophisticated analysis — regression, machine learning, or otherwise — can fully compensate for data that was collected in a biased way. This principle is so important that data scientists have a phrase for it: garbage in, garbage out. Learning to recognize bias now will serve you in every statistics course and data-driven career you might pursue.

Practice Problems

PROBLEM 1CONCEPTUAL
A school principal wants to know if students are satisfied with the cafeteria food. She announces over the loudspeaker that any student who wants to share their opinion can come to the office during lunch. What type of sampling bias is most likely present in this study, and why?
PROBLEM 2BASIC CALCULATION
A company surveys 500 customers by emailing everyone who purchased a product in the last month. Out of 500 emails sent, only 80 people respond. Of those 80, 70 say they are "very satisfied." The company reports that 87.5% of customers are very satisfied. Identify the type of bias and explain why the reported percentage is likely inaccurate.
PROBLEM 3INTERMEDIATE
A researcher wants to estimate how many hours per week high school students in a city spend on social media. She goes to the local library on a Wednesday afternoon and surveys 60 teenagers she finds there. Identify all sources of bias in this study and propose a specific improved sampling method.
PROBLEM 4APPLIED
A city council wants to decide whether to build a new skate park. They conduct a survey by calling landline phone numbers randomly selected from the city phone directory between 10 a.m. and 2 p.m. on weekdays. Of the 300 people who answer, only 15% support the skate park. A local youth group argues the survey is biased. Are they right? Explain, and design a better study.
PROBLEM 5CRITICAL THINKING
Consider the following claim: "We surveyed 10,000 people, so our results must be accurate." Using concepts from this lesson, write a well-reasoned argument explaining why this claim might be false. Include at least two specific types of bias in your argument and provide a concrete hypothetical scenario.

Lesson Summary

Sampling bias is a systematic error that occurs when a sample does not represent the population it is supposed to reflect. The five key types are voluntary response bias (participants self-select), convenience sampling bias (researcher picks the easiest-to-reach people), undercoverage bias (some groups are left out entirely), nonresponse bias (selected people don't respond), and response bias (the way questions are asked influences answers).

The most important principle is that increasing sample size does not fix bias — only improving the study design can do that. The best remedies involve probability-based sampling methods such as simple random samples and stratified random samples, where every member of the population has a known, non-zero chance of being selected. Identifying the specific type of bias in a study is the first step toward proposing a targeted, effective improvement.

Varsity Tutors • Math 3 • Sampling Bias