AP Statistics Quiz: Introducing Statistics Learning From Data
18 questions · exam conditions
0:00
Introducing Statistics Learning From DataQuestion 1 of 18

A city transportation department selects a simple random sample of 120 intersections from all intersections in the city and records whether each has a functioning pedestrian "walk" signal. In the sample, 102 intersections have a functioning signal. Which conclusion is supported by these data?

About 85% of the city's intersections have a functioning pedestrian walk signal.
Exactly 85% of the city's intersections have a functioning pedestrian walk signal.
Installing walk signals causes intersections to be safer, so the city should install more.
About 85% of intersections in all cities have functioning pedestrian walk signals.
Because the sample is random, every intersection in the city has a functioning walk signal with probability 0.85.
← Back to quizzes

AP Statistics Quiz

AP Statistics Quiz: Introducing Statistics Learning From Data

Practice Introducing Statistics Learning From Data in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Introducing Statistics Learning From Data, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A city transportation department selects a simple random sample of 120 intersections from all intersections in the city and records whether each has a functioning pedestrian "walk" signal. In the sample, 102 intersections have a functioning signal. Which conclusion is supported by these data?

  1. About 85% of the city's intersections have a functioning pedestrian walk signal. (correct answer)
  2. Exactly 85% of the city's intersections have a functioning pedestrian walk signal.
  3. Installing walk signals causes intersections to be safer, so the city should install more.
  4. About 85% of intersections in all cities have functioning pedestrian walk signals.
  5. Because the sample is random, every intersection in the city has a functioning walk signal with probability 0.85.

Explanation: Learning from data in statistics involves making inferences from random samples to populations. The simple random sample of 120 intersections finds 85% functioning, allowing an approximate inference to all city intersections. Conclusions must match the data's scope, estimating for the city's intersections without exactness, causation, or extension to other cities. A lesson is that random sampling supports population estimates like 'about 85%,' but not precise probabilities per unit or unrelated recommendations. Avoid causal claims from descriptive data. Choice A correctly estimates for the city's intersections. Others add exactness, causation, or overreach.

Question 2

A researcher wants to study the relationship between weekly exercise time and resting heart rate among adults who belong to a particular gym. She records data from 60 gym members who volunteer to participate. The data show a negative association: members who report more weekly exercise tend to have lower resting heart rates. Which conclusion is supported by the data?

  1. Among the 60 volunteers from this gym, higher reported exercise time tended to be associated with lower resting heart rate. (correct answer)
  2. Increasing weekly exercise time will cause an adult's resting heart rate to decrease.
  3. Adults in general have lower resting heart rates when they exercise more, based on this study.
  4. The negative association proves that exercise time is the only factor affecting resting heart rate.
  5. Because 60 people were measured, the relationship must be strong and linear for all gym members.

Explanation: This observational study examines the relationship between exercise time and heart rate among gym volunteers. Option A correctly describes what the data shows: among these 60 volunteers, more exercise was associated with lower heart rates. We cannot claim causation (B) from observational data, generalize to all adults (C), or make absolute statements about the relationship (D, E). The key limitation is that participants were volunteers from one gym, not a random sample. When describing associations from observational studies, use tentative language and limit conclusions to the studied group.

Question 3

A streaming service wants to estimate the average number of hours per week its subscribers watch content. It takes a simple random sample of 500 subscribers from its current subscriber list and computes a sample mean of 9.4 hours per week. Which conclusion is supported by the data?

  1. The average viewing time for all people (subscribers and non-subscribers) is about 9.4 hours per week.
  2. For the 500 sampled subscribers, the average viewing time was 9.4 hours per week. (correct answer)
  3. Because the sample is random, every subscriber watches exactly 9.4 hours per week on average.
  4. Watching 9.4 hours per week causes people to subscribe to the streaming service.
  5. The sample mean proves that most subscribers watch between 9 and 10 hours per week.

Explanation: This question tests understanding of what sample statistics represent. The streaming service sampled 500 subscribers and found a mean of 9.4 hours. Option B correctly states what this means: the average for the 500 sampled subscribers was 9.4 hours. We cannot extend this to non-subscribers (A), claim every subscriber watches this exact amount (C), reverse causation (D), or make claims about the distribution (E). Sample statistics describe the sample itself; inference to the population requires additional statistical reasoning. Always distinguish between what you observed in your sample versus claims about the population.

Question 4

A company wants to estimate the average time its customer service phone calls last. It takes a random sample of 80 calls made last Tuesday and finds a mean duration of 6.4 minutes. Which conclusion is supported by the data?

  1. The mean call duration for all calls to the company is exactly 6.4 minutes.
  2. The mean call duration for all calls made last Tuesday was about 6.4 minutes. (correct answer)
  3. The random sample proves that call durations are normally distributed.
  4. The mean call duration for all calls made in the entire year is about 6.4 minutes.
  5. Shortening the call script will cause the mean call duration to be 6.4 minutes.

Explanation: In introductory statistics, this question evaluates estimating population parameters from sample data and scope of inference. The random sample of 80 calls from last Tuesday yields a mean of 6.4 minutes, supporting an estimate for all calls that day, not broader periods or distributions. The conclusion must connect to the sampled population without overextending to all calls ever or claiming causation. A key lesson is that random samples allow approximate inferences to the population from which they were drawn, like one day's calls, but not exact values or unrelated claims like normality. Avoid causal statements from descriptive data. Choice B properly estimates for last Tuesday's calls. Others generalize too far or misinterpret.

Question 5

A school district surveyed a simple random sample of 200 students from its 4 high schools about whether they usually eat breakfast on school days. In the sample, 118 students said "yes." Which conclusion is supported by these data?

  1. About 59% of students in the district's 4 high schools usually eat breakfast on school days. (correct answer)
  2. Exactly 59% of all students in the district's 4 high schools usually eat breakfast on school days.
  3. Eating breakfast causes higher grades for students in the district.
  4. About 59% of all teenagers in the state usually eat breakfast on school days.
  5. About 59% of the district's teachers usually eat breakfast on school days.

Explanation: This question assesses the skill of drawing appropriate conclusions from sample data in introductory statistics, focusing on the scope of inference. The data come from a simple random sample of 200 students from the district's 4 high schools, with 118 saying they eat breakfast, which is 59%. The supported conclusion must match the data's limitations, inferring to the population of students in those schools without overgeneralizing or claiming causation. A mini-lesson here is that when data are from a random sample, we can make approximate inferences to the sampled population, but not exact statements or extensions to broader groups like all teenagers or teachers. We also avoid causal claims unless there's experimental evidence. Thus, choice A correctly estimates about 59% for the district's high school students. Other choices overreach by claiming exactness, causation, or unrelated populations.

Question 6

A researcher wants to estimate the proportion of residents in a city who support a proposed public transit tax. She stands outside a downtown office building at 8:00 AM for one morning and surveys the first 150 people who enter. Of those surveyed, 72% say they support the tax. Which statement is best supported by these data?

  1. About 72% of all city residents support the tax because the sample size is large.
  2. About 72% of people entering that office building around 8:00 AM on that morning supported the tax. (correct answer)
  3. The tax proposal caused 72% of city residents to support public transit.
  4. About 72% of all downtown workers in the city support the tax.
  5. Because 72% is above 50%, a majority of all city residents definitely support the tax.

Explanation: This question in learning from data highlights recognizing sampling bias and limiting conclusions to the sampled group. The researcher used a convenience sample of 150 people entering a downtown office at 8:00 AM, with 72% supporting the tax, so inferences can't extend to all city residents due to potential bias. The data support conclusions only about the specific group surveyed, not broader populations or causal effects. A mini-lesson is that non-random samples, like convenience ones, may not represent the population, so match conclusions tightly to who was actually surveyed without assuming representativeness. We also avoid definitive majority claims without accounting for sampling error. Thus, choice B correctly limits to people entering that building that morning. Other options incorrectly generalize or imply causation.

Question 7

A marketing team tests two website layouts. For one week, the site uses Layout 1; the next week, it uses Layout 2. The purchase rate is 3.1% during the Layout 1 week and 3.8% during the Layout 2 week. No other data are collected. Which conclusion is supported by these data?

  1. Layout 2 caused the purchase rate to increase by 0.7 percentage points.
  2. The purchase rate was higher during the week when Layout 2 was used than during the week when Layout 1 was used. (correct answer)
  3. Layout 2 will increase purchase rates for all future weeks, regardless of seasonality or promotions.
  4. Because the purchase rate increased, customers preferred Layout 2 and reported higher satisfaction.
  5. The difference in purchase rates must be due to random assignment of visitors to layouts.

Explanation: This introductory statistics question examines conclusions from non-experimental data over time. The website used Layout 1 one week (3.1% purchases) and Layout 2 the next (3.8%), showing a descriptive difference but not causation due to potential time-based confounds. The supported conclusion describes the observed rates without implying cause or future effects. A mini-lesson is to recognize sequential designs aren't randomized, so stick to descriptive comparisons and avoid causal or preferential claims. Differences don't imply random assignment or ignore external factors. Choice B properly states the higher rate during Layout 2's week. Other choices introduce causation or unfounded predictions.

Question 8

A city health department randomly selects 250 adults from its list of registered residents and asks whether they have received a flu shot this year. Of those sampled, 148 say "yes." Which conclusion is supported by the data?

  1. About 148/250148/250 of the sampled registered residents reported getting a flu shot this year. (correct answer)
  2. About 148/250148/250 of all adults in the city got a flu shot this year.
  3. Getting a flu shot causes adults to register as residents in the city.
  4. Because the sample was random, exactly 148/250148/250 of all registered residents got a flu shot.
  5. Adults who did not respond to the survey must be less likely to have gotten a flu shot.

Explanation: This question examines understanding of what survey data can tell us. The health department surveyed 250 randomly selected registered residents and found 148 said they got flu shots. Option A correctly describes what the data shows: the proportion in the sample (148/250). We cannot extend this exact proportion to all adults in the city (B) because the survey only included registered residents. Options C, D, and E make unsupported claims about causation, exact population values, or non-respondents. When interpreting survey data, stick to describing what was actually observed in your sample.

Question 9

A wildlife biologist wants to estimate the proportion of trout in a lake that are infected with a certain parasite. She catches 80 trout using a net in a shallow cove near the shore and finds 12 infected. Which conclusion is supported by the data?

  1. About 12/8012/80 of the trout in the entire lake are infected with the parasite.
  2. The parasite infection rate in the shallow cove is exactly 12/8012/80 for all trout in that cove.
  3. Among the 80 trout caught in the shallow cove, 12/8012/80 were infected with the parasite. (correct answer)
  4. Catching trout in shallow water causes parasite infection.
  5. Because 80 trout were sampled, the sample must be representative of the whole lake regardless of how the trout were caught.

Explanation: This question highlights the importance of representative sampling. The biologist caught 80 trout from a shallow cove near shore - not a random sample from the entire lake. Option C correctly limits the conclusion to what was observed: 12/80 of the caught trout were infected. We cannot generalize to the whole lake (A) because trout in shallow coves might differ from those in deeper water. Options B, D, and E make unsupported claims about exact rates, causation, or representativeness. When samples aren't random or representative, conclusions must be restricted to the sampled group.

Question 10

A teacher wants to estimate the average number of hours of sleep per night for students at her school. She surveys the 32 students in her first-period class and finds a sample mean of 6.7 hours. Which conclusion is supported by the data?

  1. The average sleep for all students at the school is exactly 6.7 hours per night.
  2. The average sleep for students in the teacher's first-period class is 6.7 hours per night (based on the survey). (correct answer)
  3. Students at the school would sleep more if they were moved into a first-period class.
  4. Because 32 students were surveyed, the result can be generalized to all students at the school without concern.
  5. The survey proves that most students at the school sleep less than 7 hours per night.

Explanation: This question tests recognition of sampling limitations. The teacher surveyed only students in her first-period class - a convenience sample, not a random sample of all students. Option B correctly limits the conclusion to what the data actually represents: the average for students in that specific class. We cannot generalize to all students (A, D, E) because first-period students might differ systematically from others (perhaps early-risers get less sleep). Option C makes an unsupported causal claim. When data comes from a non-random sample, conclusions must be limited to that specific group.

Question 11

A principal wants to compare the average time it takes students to finish a standardized reading test when taken on paper versus on a computer. She randomly assigns 120 students to take the paper version and 120 students to take the computer version. The paper group's mean time is 54 minutes; the computer group's mean time is 49 minutes. Which conclusion is supported by the data?

  1. Students who take the test on a computer tend to finish faster than students who take it on paper, in this randomized experiment. (correct answer)
  2. Students at all schools will finish 5 minutes faster on a computer than on paper.
  3. Students who prefer computers were more likely to be placed in the computer group.
  4. Taking the test on a computer causes every student to finish faster than they would on paper.
  5. Because the computer group had a smaller mean time, the distribution of times must have been less variable for the computer group.

Explanation: This randomized experiment compared test completion times between paper and computer formats. With random assignment of 120 students to each group, we can make causal conclusions. Option A correctly states that computer-takers "tend to finish faster" in this experiment - appropriate language for experimental results. Options B and D overgeneralize (to all schools or every student). Option C incorrectly suggests assignment bias in a randomized study. Option E makes an unsupported claim about variability. Experimental conclusions should describe the observed effect without overgeneralizing beyond the experimental context.

Question 12

A student group wants to know whether students who participate in school sports report higher overall happiness. They survey 40 students who play a school sport and 40 students who do not, asking each to rate happiness on a 1–10 scale. The athletes' mean rating is 7.8; the non-athletes' mean rating is 6.9. Which conclusion is supported by the data?

  1. Playing a school sport causes students to be happier.
  2. At this school, students who play a school sport tended to report higher happiness ratings than students who do not (in this survey). (correct answer)
  3. Students at all schools who play sports tend to be happier than those who do not.
  4. The difference in means proves that sports participation is the only reason for differences in happiness.
  5. Because the group sizes are equal, the survey results are free of any possible confounding variables.

Explanation: This observational study compares happiness ratings between students who play sports and those who don't. Option B correctly describes the finding: at this school, sport players "tended to report higher happiness ratings" in this survey. This language appropriately describes an association without claiming causation (A) or generalizing beyond this school (C). Options D and E incorrectly claim the study proves causation or eliminates confounding. In observational studies comparing self-selected groups, we can only describe observed associations, not determine cause-and-effect relationships.

Question 13

A school counselor wants to learn whether an optional after-school tutoring program is associated with higher Algebra 1 exam scores. The counselor compares exam scores for all 86 ninth-graders who chose to attend tutoring at least once (mean =82=82) and all 190 ninth-graders who did not attend (mean =76=76) at one high school this semester. Which conclusion is supported by the data?

  1. Attending tutoring causes ninth-graders at this school to score higher on the Algebra 1 exam.
  2. Ninth-graders who attended tutoring at this school tended to have higher Algebra 1 exam scores than those who did not attend. (correct answer)
  3. Ninth-graders at all high schools would score about 6 points higher if they attended this tutoring program.
  4. The tutoring program is effective because the mean score difference proves the program works for every student who attends.
  5. Because the sample sizes are large, the difference in means must be due to the tutoring program rather than other factors.

Explanation: This question tests understanding of observational studies versus experiments. The key insight is that students chose to attend tutoring (not randomly assigned), making this an observational study. In observational studies, we can only describe associations, not claim causation. Option B correctly states that students who attended tutoring "tended to have higher" scores - this describes the observed pattern without claiming tutoring caused the difference. Options A, C, D, and E all make causal claims or overgeneralize beyond what the data supports.

Question 14

A student wants to estimate the average number of hours per week that students at her college spend studying. She posts an online poll on her social media and receives 310 responses. The mean reported study time is 12.1 hours per week. Which conclusion is supported by these data?

  1. About 12.1 hours per week is the average study time for all students at her college because the sample is large.
  2. About 12.1 hours per week is the average study time for all college students in the country.
  3. About 12.1 hours per week is the mean reported study time among the people who responded to her social media poll. (correct answer)
  4. Studying 12.1 hours per week causes higher GPAs for students at her college.
  5. Because the poll was online, the result is unbiased and representative of the entire college.

Explanation: This question tests recognizing voluntary response bias in sampling for statistical conclusions. The online social media poll got 310 responses with a mean of 12.1 hours, but self-selection means it only describes responders, not the college or beyond. Conclusions must limit to the data from respondents without assuming representativeness or causation. A key lesson is that voluntary polls are biased toward those motivated to respond, so match conclusions to the actual group surveyed, avoiding generalizations to populations like all students. Online doesn't ensure unbiasedness. Choice C accurately describes the mean for responders. Others overgeneralize or claim causation.

Question 15

A university wants to know whether a new tutoring program improves calculus exam scores. Students choose whether to participate in tutoring. At the end of the term, the mean exam score of participants is 84, and the mean score of nonparticipants is 78. Which conclusion is supported by these data?

  1. Participating in tutoring caused students to score 6 points higher on average.
  2. Among students in this university who chose to participate, the mean exam score was higher than among those who did not choose to participate. (correct answer)
  3. All calculus students everywhere would increase their scores by 6 points if they used this tutoring program.
  4. The tutoring program is the only reason participants scored higher.
  5. Because the mean is higher for participants, the program must have been randomly assigned.

Explanation: This question assesses distinguishing observational studies from experiments in statistics, focusing on causation. Students self-selected into tutoring, with participants averaging 84 and nonparticipants 78, showing an association but not causation due to lack of random assignment. Conclusions should describe the observed difference in this university's groups without implying cause or generalizing universally. A mini-lesson is that self-selection can introduce confounding, so limit to descriptive statements and avoid attributing differences solely to the program. Higher means don't prove causation or random assignment. Choice B accurately describes the means for choosers versus non-choosers at this university. Other options incorrectly claim causation or overgeneralize.

Question 16

A fitness app company wanted to compare two reminder messages. They randomly assigned 400 current users to receive either Message A or Message B for 2 weeks. At the end, 62% of users assigned to A met a weekly step goal, compared with 55% assigned to B. Which conclusion is supported by the data?

  1. Message A caused a higher step-goal success rate than Message B for the app's current users in this study. (correct answer)
  2. Message A will increase step-goal success for all adults, whether or not they use the app.
  3. Users who met the step goal were more likely to be assigned to Message A by choice.
  4. Message A caused a 7 percentage point increase in step-goal success for all people who download the app in the future.
  5. Message A and Message B have exactly the same effect because the percentages are close.

Explanation: In introductory statistics, this question tests understanding causal conclusions from experimental data versus observational data. The study randomly assigned 400 app users to two messages, finding 62% success with A versus 55% with B, allowing a causal inference within the study's scope. The conclusion should connect the data to the treatment effect for the participants without extrapolating to non-users or future scenarios. A key lesson is that random assignment enables causal claims about the studied group, but we must limit scope to the sample and conditions. Avoid assuming effects on broader populations or ignoring the difference due to closeness. Therefore, choice A appropriately states causation for the current users in this study. Choices like B or D overgeneralize beyond the data.

Question 17

A teacher wants to know whether students in her AP Statistics class sleep more on weekends than on school nights. She asks all 28 students in her class to report how many hours they slept last night (a school night) and how many hours they slept the most recent Saturday night. The average reported increase (Saturday minus school night) is 1.6 hours. Which conclusion is supported by these data?

  1. AP Statistics students nationwide sleep about 1.6 more hours on weekends than on school nights.
  2. For these 28 students, the average reported sleep was 1.6 hours higher on Saturday night than on the school night. (correct answer)
  3. Sleeping on Saturday night causes students to sleep less on school nights.
  4. All students at the school sleep about 1.6 more hours on weekends than on school nights.
  5. Because the teacher asked every student in the class, the result proves that teenagers generally sleep 1.6 more hours on weekends.

Explanation: Introductory statistics skills include distinguishing descriptive statistics from inferential ones and avoiding causation in observational data. The teacher collected data from all 28 students in her class, finding an average 1.6-hour increase on weekends, which describes only this group without sampling for broader inference. Conclusions must stick to the observed data for these students, not generalize to all AP students or imply causation. A lesson here is that when data are from a specific, non-random group like one class, conclusions are descriptive and limited to that group; don't extrapolate or confuse association with cause. Census of a small group doesn't prove nationwide trends. Choice B accurately describes the average for these 28 students. Others overgeneralize or claim causation.

Question 18

A university dining hall tests two different signs meant to reduce food waste. On 10 randomly chosen days, the hall displays Sign A; on 10 other randomly chosen days, it displays Sign B. The average pounds of food wasted per day is 320 with Sign A and 295 with Sign B. Which conclusion is supported by the data?

  1. Sign B caused lower average daily food waste than Sign A in this experiment on these 20 days. (correct answer)
  2. Sign B will reduce food waste at every dining hall by 25 pounds per day.
  3. Days with lower food waste were more likely to be assigned to Sign B, so the experiment is invalid.
  4. Sign B reduces food waste because students prefer it, as shown by the data.
  5. Because the mean waste was lower with Sign B, there must have been less variability in waste on Sign B days.

Explanation: This randomized experiment tested two signs' effects on food waste. With random assignment of days to each sign, we can make causal conclusions. Option A correctly states that Sign B "caused lower average daily food waste" in this specific experiment. This causal language is appropriate due to randomization. Options B and D overgeneralize or make unsupported claims about mechanisms. Option C incorrectly suggests the randomization failed. Option E makes an unfounded claim about variability. Experimental conclusions should state the causal effect observed in the specific experimental context.