MATH 3 • STATISTICS & PROBABILITY

Designing Studies & Surveys — I can design a simple study or survey question set that measures a variable clearly at my level.

Learn to craft unbiased questions and choose study designs that produce trustworthy, meaningful data.

Historical Context & Motivation

Humans have been collecting data for thousands of years — ancient civilizations counted livestock, crops, and citizens to manage their resources. However, the idea of designing a systematic study to answer a specific question is surprisingly modern. For most of history, rulers simply ordered a census and accepted whatever numbers came back, with no thought about whether the process introduced errors or whether the questions themselves were fair. The shift toward carefully planned research designs transformed data collection from crude guesswork into a rigorous discipline that powers medicine, technology, and public policy today.

1085
The Domesday Book
William the Conqueror ordered a comprehensive survey of England's land, livestock, and population — one of the earliest large-scale data-collection efforts in Western history.
1747
First Clinical Trial
Scottish physician James Lind tested six different remedies on sailors with scurvy, comparing outcomes to identify that citrus fruit was the cure — a pioneering controlled experiment.
1936
Literary Digest Poll Disaster
A major magazine predicted Alf Landon would defeat Franklin Roosevelt in a landslide. Their survey sampled from telephone and car-ownership lists, missing millions of lower-income voters. The spectacular failure proved that how you sample matters as much as how many you sample.
1948
Randomized Controlled Trials Formalized
The British Medical Research Council conducted the first modern randomized controlled trial for streptomycin against tuberculosis, establishing the gold standard for experiments.
2020s
Digital Surveys & Big Data
Online survey platforms and data analytics allow researchers to collect millions of responses instantly, but new challenges in bias, privacy, and question design have emerged.

The central question running through all of this history is deceptively simple: How do we collect information so that the answers we get actually reflect reality? A poorly designed study can mislead doctors, waste billions of dollars, or sway elections. In this lesson, you will learn to design studies and survey questions that measure variables clearly and fairly — skills that make you a smarter consumer of data and a more effective researcher.

Core Principles & Definitions

Before you can design a study, you need a shared vocabulary. The building blocks of any research design are the population you care about, the variable you want to measure, the method you use to collect data, and the precautions you take to avoid bias — systematic errors that push results in one direction. Let's unpack these ideas with five foundational principles.

1

Population vs. Sample

The population is the entire group you want to learn about. A sample is the smaller subset you actually collect data from. Good design ensures the sample represents the population.
2

Variable of Interest

A variable is any characteristic that can change from one individual to another — height, opinion, test score, etc. Your study must define and measure this variable precisely.
3

Observational Study vs. Experiment

An observational study records data without changing conditions. An experiment deliberately imposes a treatment on subjects and measures the response.
4

Bias & Confounding

Bias is any systematic favoritism in how data is collected. A confounding variable is a hidden factor that influences both the explanatory variable and the response, making it hard to determine cause and effect.
5

Random Selection & Random Assignment

Random selection means every member of the population has an equal chance of being in the sample (fights sampling bias). Random assignment means subjects are placed into treatment groups by chance (fights confounding in experiments).
KEY TAKEAWAY
Think of designing a study like planning a taste test for a new pizza recipe. If you only ask your friends (biased sample), only test one topping at a time (controlled variable), and don't tell tasters which slice is which (blinding), your results will be much more trustworthy than just asking "Do you like my pizza?" at a party. Every design choice either adds or removes noise from your data.

Visual Explanation — The Study Design Flowchart

Choosing the right study design depends on two key decisions: first, whether you want to observe or intervene, and second, how you select your participants. The diagram below maps out this decision tree so you can quickly identify which design fits your research question.

Start at the top with your research question. If you cannot or should not impose a treatment, follow the left branch to observational designs. If you can impose a treatment, follow the right branch to experiments. The bottom boxes show what each design type can and cannot prove.

Notice the critical distinction at the bottom of the diagram. Observational studies — including surveys — can reveal associations between variables, but they cannot prove that one variable causes another. Only well-designed experiments with random assignment can establish causation, because random assignment balances out confounding variables across groups. When you design a survey, keep in mind that your conclusions will be limited to associations unless you build in an experimental component.

How Good Survey Questions Work

Even if you pick the perfect study design, your results are only as good as your questions. A confusing, leading, or vague survey question introduces response bias — a distortion caused by how people interpret or react to the wording. Writing clear survey items is part science, part craft, and there are concrete rules you can follow.

Anatomy of a Survey Question

Every survey question has three components. The stem is the question or statement itself. The response format is how the respondent answers — multiple choice, rating scale, open-ended text, etc. The instructions tell the respondent how to mark their answer. All three must align so the variable you intend to measure is the variable you actually capture.

Rules for Writing Unbiased Questions

  1. Use neutral language. Avoid loaded words. Instead of "Don't you agree that homework is excessive?", write "How would you rate the amount of homework you receive?"
  2. Ask one thing at a time. "Do you enjoy math and science?" is a double-barreled question. A student might love math but dislike science and not know how to answer.
  3. Provide exhaustive and mutually exclusive response options. If age ranges are 14–16 and 16–18, a 16-year-old doesn't know which box to check. Fix it: 14–15, 16–17, 18+.
  4. Define ambiguous terms. "Do you exercise regularly?" means different things to different people. Specify: "In a typical week, how many days do you exercise for at least 30 minutes?"
  5. Avoid leading or prestige wording. "Most experts recommend 8 hours of sleep — do you get enough?" pressures the respondent toward a socially desirable answer.
⚠️ COMMON TRAP
Voluntary response surveys (like online polls where anyone can click) attract people with strong opinions. If a school posts "Should we ban phones?" on its website, the students who feel most passionately — for or against — will respond, while the moderate majority stays silent. This creates voluntary response bias, making results unrepresentative.

Identifying and Avoiding Bias

Bias is the enemy of trustworthy data. It sneaks into a study at every stage — when you choose your sample, write your questions, or even decide where to collect responses. The diagram below illustrates the most common types of bias you need to guard against, along with the stage of the study where each one strikes.

The six most common bias types are organized by the stage at which they occur. Sampling stage biases (left) affect who gets into the study. Questioning stage biases (center) affect how people interpret the items. Response stage biases (right) affect who answers and how honestly they answer.

When reviewing your own study design, walk through all three stages. Ask: Is every subgroup of my population represented in the sample? Are my questions neutral, clear, and single-topic? Have I removed pressure for respondents to answer in a socially desirable way? Addressing each stage systematically is the best insurance against collecting misleading data.

Worked Example — Designing a School Survey

Imagine you are a member of the student council and you want to find out whether students at your high school support changing the start time from 7:45 AM to 8:30 AM. Your principal wants data, not just opinions from a few loud voices. Let's design a study from scratch.

Designing a Survey on School Start Times
1
Step 1 — Define the VariableThe variable of interest is student opinion on changing the school start time. This is a categorical variable because the responses will fall into distinct categories (e.g., strongly support, somewhat support, neutral, somewhat oppose, strongly oppose). Defining this clearly prevents us from accidentally measuring something else, like general school satisfaction.
Variable: Student opinion on a later start time (categorical, ordinal)
2
Step 2 — Identify the Population and Choose a Sampling MethodThe population is all currently enrolled students at the school (say, 1,200 students). Surveying everyone is impractical, so we need a sample. We decide on a stratified random sample: we split the population into strata by grade level (9th, 10th, 11th, 12th) and randomly select 30 students from each grade using a computer-generated random number list. This ensures every grade is proportionally represented and avoids the selection bias that would occur if we only surveyed students in one lunch period.
Sample: 120 students (30 per grade), selected by stratified random sampling
3
Step 3 — Write Unbiased Survey QuestionsWe need a clear, neutral stem with exhaustive, mutually exclusive response options. A bad question would be: "Research shows teens need more sleep. Don't you think we should start school later?" This is leading (cites authority) and uses a negative contraction that confuses. A good question: "Our school currently starts at 7:45 AM. A proposal would move the start time to 8:30 AM. How would you rate your support for this change?" with a 5-point scale from Strongly Oppose to Strongly Support.
Final question uses neutral framing and a balanced Likert scale
4
Step 4 — Plan Administration to Reduce Response BiasTo minimize social desirability bias, we make responses anonymous — no names on the form. To avoid voluntary response bias, we distribute the survey during a mandatory advisory period rather than posting it online for anyone to answer. We also include a brief instruction line: "Circle the one option that best reflects your personal opinion."
Administration: Anonymous paper survey during advisory, mandatory participation
5
Step 5 — Anticipate LimitationsEven with careful design, this is still an observational study — we are measuring opinions, not testing an intervention. Students absent on survey day will be missed (nonresponse). Also, opinions may change over time. We note these limitations honestly when reporting results. Despite this, the stratified random sample and neutral wording give us far more trustworthy data than an informal show of hands at a council meeting.
Limitations documented; design is strongest feasible observational approach

Comparing Study Types — Strengths & Limitations

Different study designs serve different purposes. A survey is great for measuring opinions at a single point in time, while an experiment is necessary to prove a treatment works. Understanding the trade-offs helps you choose the best design for your research question and communicate the limits of your conclusions honestly.

Comparison of common study designs
Design TypeStrengthsLimitations
Survey / CensusQuick and inexpensive; can reach large samples; measures opinions, behaviors, and demographicsCannot establish causation; subject to response and wording bias; self-reported data may be inaccurate
Observational Study (non-survey)Studies naturally occurring behaviors; ethical when experiments are not possible (e.g., smoking effects)Confounding variables are difficult to control; cannot prove causation; results may not generalize
Experiment (Completely Randomized)Random assignment controls for confounders; can establish cause-and-effect relationshipsCan be expensive and time-consuming; may raise ethical concerns; artificial settings may limit generalizability
Experiment (Block Design)Controls for known confounders by blocking; increases precision within subgroupsMore complex to set up and analyze; requires advance knowledge of blocking variables
KEY TAKEAWAY
Choosing a study design is like choosing a vehicle for a road trip. A bicycle (survey) is cheap and fast for short distances but won't cross a mountain range. A four-wheel-drive truck (randomized experiment) handles rough terrain and gets you to a causal conclusion, but it costs more time and resources. The right choice depends on where you need to go — that is, what question you need to answer.

Connection to Advanced Statistical Thinking

The skills you are building now — defining variables, selecting representative samples, and writing clear questions — form the foundation for every advanced statistics course you will encounter. In AP Statistics or college-level courses, you will layer on concepts like margin of error, confidence intervals, and hypothesis testing on top of the design framework you are learning here. Without a well-designed study, no amount of advanced math can rescue the results.

From basics to advanced: how today's concepts scale up
What You Learn NowWhat Comes Next
Defining a variable of interestChoosing between categorical and quantitative analysis; selecting appropriate test statistics
Random sampling to reduce biasComputing sampling distributions and standard error to quantify how much samples vary
Writing neutral survey questionsConstructing validated scales (Likert, semantic differential) and assessing reliability
Observational study vs. experimentDesigning double-blind, placebo-controlled randomized trials; using ANOVA for multi-group comparisons
Identifying confounding variablesMultiple regression and statistical control to adjust for confounders mathematically

Think of your current study-design skills as the blueprint for a building. Advanced statistics adds structural engineering, plumbing, and electrical systems, but none of those matter if the blueprint is flawed. Mastering these design fundamentals now will make every future statistics topic more intuitive and meaningful.

Practice Problems

PROBLEM 1CONCEPTUAL
A school newspaper wants to know how students feel about the cafeteria food. The editor asks 20 students who are sitting in the cafeteria at lunch whether they like the food. Identify two potential sources of bias in this approach and explain why each one is a problem.
PROBLEM 2BASIC CALCULATION
A high school has 800 students: 250 freshmen, 220 sophomores, 180 juniors, and 150 seniors. You want a stratified random sample of 80 students (10% of the school). How many students should be sampled from each grade to maintain proportional representation?
PROBLEM 3INTERMEDIATE
Rewrite each of the following biased survey questions to make them clear, neutral, and single-topic. (a) "Don't you think our school should obviously invest more money in the arts program?" (b) "How much do you enjoy your math and English classes?" (c) "Do you exercise a lot?"
PROBLEM 4APPLIED
You are the manager of a local movie theater and want to determine whether offering a discount on Tuesday evenings increases overall weekly attendance. Describe a study you would design to answer this question. In your answer, (a) state whether this is an observational study or an experiment, (b) identify the explanatory and response variables, (c) explain how you would collect data, and (d) describe one confounding variable and how you would try to control it.
PROBLEM 5CRITICAL THINKING
A researcher claims: "Our online poll of 50,000 people shows that 78% of Americans support Policy X. With such a huge sample size, the result is virtually certain to reflect the true opinion of the population." Evaluate this claim. Is a large sample size sufficient to guarantee accurate results? Explain your reasoning by referencing specific concepts from this lesson.

Lesson Summary

Designing a trustworthy study starts with clearly defining your variable of interest and identifying the population you want to learn about. You then choose between an observational study (which can reveal associations) and an experiment (which can establish causation through random assignment). For either design, use random selection from the population to ensure your sample is representative.

When writing survey questions, guard against bias at every stage: avoid leading questions, double-barreled items, and vague wording. Use neutral language with clearly defined, mutually exclusive response options. Finally, reduce voluntary response bias by sampling randomly rather than relying on opt-in polls, and reduce social desirability bias by making responses anonymous. These design principles ensure that the data you collect genuinely reflects the reality you set out to measure.

Varsity Tutors • Math 3 • Designing Studies & Surveys