What this quiz covers
This quiz focuses on How Collected Data Tells The Truth, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
A city health department investigates the research question: "Is the mean number of hours of sleep for adults in the city at least 7 hours?" Interviewers stood outside a downtown gym on weekday mornings and surveyed 85 adults leaving the gym; the sample mean was 7.4 hours. The department concluded that adults in the city average at least 7 hours of sleep. Which statement explains whether the conclusion is valid?
AP Statistics Quiz
Practice How Collected Data Tells The Truth in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.
This quiz focuses on How Collected Data Tells The Truth, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.
Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.
A city health department investigates the research question: "Is the mean number of hours of sleep for adults in the city at least 7 hours?" Interviewers stood outside a downtown gym on weekday mornings and surveyed 85 adults leaving the gym; the sample mean was 7.4 hours. The department concluded that adults in the city average at least 7 hours of sleep. Which statement explains whether the conclusion is valid?
Explanation: This question examines how convenience sampling undermines the validity of generalizations to a larger population. The health department surveyed only people leaving a downtown gym on weekday mornings, creating a convenience sample that likely overrepresents health-conscious individuals who exercise regularly and may have different sleep patterns than the general adult population. People who go to the gym in the morning might prioritize sleep more or have schedules that allow for adequate rest. The sample mean of 7.4 hours cannot reliably estimate the mean for all city adults because the sampling method systematically excludes many groups (non-exercisers, people with different work schedules, etc.). A valid conclusion would require random sampling from all adults in the city, not just gym-goers.
A county election office investigates the research question: "What proportion of registered voters in the county approve of the new voting center location?" The office used a random-digit dialing method that called only landline phone numbers in the county and completed interviews with 520 registered voters; 61% approved. The office concluded that about 61% of all registered voters in the county approve. Which statement explains whether the conclusion is valid?
Explanation: This question illustrates coverage bias, where the sampling method systematically excludes certain population segments. By using random-digit dialing limited to landline phones, the election office excludes voters who rely exclusively on cell phones, a group that tends to be younger and may have different political views. This creates undercoverage of an important demographic, potentially biasing the approval estimate. While random-digit dialing of landlines was once considered good practice, changing communication patterns have made it problematic for representing entire populations. The 61% approval estimate may not accurately reflect all registered voters' opinions if cell-phone-only users have different views about the voting center location. Modern polling must include cell phones or use other methods to ensure all voter segments have a chance of selection.
A school district investigates the research question: "What proportion of all district parents support requiring school uniforms?" The district randomly selected 8 schools from the 32 schools in the district, then surveyed all parents who attended a PTA meeting at each selected school; 64% of those surveyed supported uniforms. The district concluded that about 64% of all district parents support requiring uniforms. Which statement explains whether the conclusion is valid?
Explanation: This question illustrates how cluster sampling can introduce bias when combined with convenience sampling within clusters. While the district appropriately randomly selected 8 schools (clusters) from 32, they then surveyed only parents attending PTA meetings at those schools. PTA meeting attendees likely overrepresent parents who are more involved in school activities, have strong opinions about school policies, and have schedules permitting meeting attendance. These parents may have different views on uniforms than less-involved parents or those unable to attend meetings. The 64% support estimate probably doesn't represent all district parents accurately. For valid conclusions, the district should have randomly sampled parents within each selected school rather than relying on PTA meeting attendance, which creates a biased sub-sample within each cluster.
A wildlife agency investigates the research question: "What is the mean weight of adult trout in Lake Orion?" Biologists used nets at two easily accessible shoreline locations on a single afternoon and weighed 60 adult trout; the sample mean weight was 1.8 kg. They concluded that the mean weight of adult trout in Lake Orion is about 1.8 kg. Which statement explains whether the conclusion is valid?
Explanation: This question demonstrates how convenience sampling based on accessibility can create biased samples in ecological research. The biologists sampled trout from only two easily accessible shoreline locations on a single afternoon, which likely doesn't represent the full population of adult trout in Lake Orion. Fish at different depths, in different parts of the lake, or active at different times may have different weights due to varying food availability, water temperature, or habitat quality. Shoreline fish might be smaller or larger than those in deeper waters. Additionally, sampling on just one afternoon doesn't account for temporal variation. The 1.8 kg mean weight estimate is questionable because the sampling method systematically excludes trout from most of the lake. Valid conclusions would require sampling from multiple locations throughout the lake across different times.
A university investigates the research question: "Do first-year students at the university spend more time studying on weekends than on weekdays?" Researchers recruited 40 volunteers from an honors study-skills workshop and had them self-report hours studied on a typical weekday and a typical weekend day; the volunteers reported more hours on weekends. The researchers concluded that first-year students at the university generally study more on weekends than weekdays. Which statement explains whether the conclusion is valid?
Explanation: This question highlights how volunteer samples from specific subgroups cannot support generalizations to larger populations. The researchers recruited volunteers from an honors study-skills workshop, creating a sample that likely overrepresents academically motivated students who actively seek study improvement resources. These students probably have different study habits than typical first-year students, making any conclusions about weekend versus weekday study time invalid for the general first-year population. Additionally, self-selection bias occurs when students choose to participate, further skewing the sample toward those interested in the research topic. To validly answer their research question, researchers would need to randomly sample from all first-year students, not just workshop attendees, and address potential nonresponse issues.
A state transportation agency investigates the research question: "What proportion of all registered drivers in the state have received a speeding ticket in the past year?" The agency mailed surveys to 2000 randomly selected registered drivers; 620 returned the survey, and 18% of respondents reported a speeding ticket in the past year. The agency concluded that about 18% of all registered drivers received a speeding ticket in the past year. Which statement explains whether the conclusion is valid?
Explanation: This question demonstrates how nonresponse bias can specifically undermine conclusions when the characteristic being measured might influence response rates. The transportation agency randomly selected drivers, which is appropriate, but only 31% (620/2000) responded to the mailed survey. Critically, drivers who received speeding tickets might be less likely to respond due to embarrassment, legal concerns, or negative associations with government agencies. This differential nonresponse would cause the 18% estimate to underestimate the true proportion of drivers with speeding tickets. The low response rate combined with the sensitive nature of the question makes the conclusion questionable. A more valid approach might use official records rather than self-reported data, or employ methods to increase response rates and assess nonresponse bias.
A city transportation office wants to investigate the research question: "What proportion of all adult residents in the city support adding protected bike lanes on major streets?" The population is all adult residents of the city. The office posted a link to an online poll on the city's official social media accounts and received 2,450 responses; 71% of respondents said they support adding protected bike lanes. The office concludes that about 71% of all adult residents support adding protected bike lanes. Which statement explains whether the conclusion is valid?
Explanation: This question tests understanding of voluntary response bias in data collection. The city posted a poll on social media, which creates a voluntary response sample where people self-select to participate. This sampling method typically overrepresents people with strong opinions about the topic (bike lanes) and underrepresents those who are neutral or less engaged. The correct answer B identifies this fundamental flaw - voluntary response samples are not representative of the population. When data collection allows self-selection rather than using random sampling, the resulting estimates cannot be trusted to reflect the true population parameter, regardless of how large the sample size is.
A county wants to estimate the proportion of households that have access to high-speed internet. The county divided the county into 10 geographic regions and then randomly selected 2 regions to survey. Within each selected region, the county surveyed every household and found that 58% of surveyed households had high-speed internet. The county concluded that "about 58% of all households in the county have access to high-speed internet." Which statement explains whether the conclusion is valid?
Explanation: The skill tested in this AP Statistics question is recognizing limitations in cluster sampling with few clusters, and how it impacts the validity of population inferences. The county selected only 2 out of 10 regions as clusters and surveyed all households within them, but with so few clusters, the sample may not capture county-wide variability if the chosen regions differ from others. Choice C distracts by claiming cluster sampling is always representative for geographic areas, but representativeness depends on random selection and sufficient clusters. Trustworthy conclusions in cluster sampling require enough randomly selected clusters to approximate the population's diversity. This conclusion is invalid due to the small number of clusters, highlighting that inadequate cluster counts can lead to bias similar to non-random sampling.
A company has 4 departments with different numbers of employees. To estimate the proportion of all employees who prefer working from home at least 3 days per week, an analyst randomly selected 50 employees from each department (200 total) and found that 55% of those sampled preferred working from home at least 3 days per week. The analyst concluded that "about 55% of all employees prefer working from home at least 3 days per week." Which statement explains whether the conclusion is valid?
Explanation: This AP Statistics question focuses on evaluating stratified sampling and the impact of unequal strata sizes on data representativeness, relating to how collection methods influence truthful inferences. The problem arises because the analyst sampled equal numbers from departments of varying sizes, overrepresenting smaller departments and potentially skewing the overall proportion unless weighted adjustments are made. A distractor like choice B misleadingly suggests that equal sampling from strata ensures no bias, but this ignores the need for proportional allocation or weighting in stratified samples. Trustworthy conclusions require that stratified samples reflect the population's composition, either by sampling proportionally or weighting results accordingly to avoid misrepresentation. In this scenario, the conclusion may not be valid without such adjustments, underscoring the mini-lesson that sampling design must align with population structure for accurate generalizations.
A health clinic wants to estimate the proportion of adults in its county who received a flu shot this season. A nurse sampled 300 patients who came to the clinic during a two-day period and found that 64% reported receiving a flu shot. The clinic concluded that "about 64% of adults in the county received a flu shot this season." Which statement explains whether the conclusion is valid?
Explanation: Identifying selection bias in convenience samples is the focus of this AP Statistics question, exploring how data collection methods reveal or distort the truth. The clinic sampled only visitors, who may be more health-conscious and likely to get flu shots, creating a biased group that doesn't represent all county adults. Choice A distracts by claiming clinic patients are random, but they are self-selected and may differ systematically from the population. Trustworthy conclusions require a sampling method that includes non-clinic-goers, such as random digit dialing or address-based sampling. Here, the conclusion is invalid due to selection bias, teaching that samples from specific venues often fail to generalize unless the venue mirrors the population.
A streaming service wants to estimate the proportion of its subscribers who would recommend the service to a friend. It used a random number generator to select 1200 subscriber accounts from its full subscriber list and emailed those accounts a survey; 840 responded, and 72% of respondents said they would recommend the service. The company concluded that "about 72% of all subscribers would recommend the service." Which statement explains whether the conclusion is valid?
Explanation: This AP Statistics question evaluates nonresponse bias in random samples, stressing how incomplete responses affect the truthfulness of data conclusions. Although accounts were randomly selected, only 840 of 1200 responded, and if nonresponders differ (e.g., less satisfied subscribers ignore emails), the results could be biased. Choice A is a distractor suggesting randomness eliminates nonresponse bias, but bias can still occur if responders aren't representative. For trustworthy conclusions, high response rates or follow-ups are needed to mitigate nonresponse, ensuring the responding sample reflects the selected one. The conclusion may not be valid without addressing this, providing a lesson that even random starts require complete data to avoid distortion.
A state agency wants to estimate the proportion of registered voters in the state who approve of a proposed ballot measure. The agency used a list of landline telephone numbers and called numbers at random until 600 registered voters responded; 49% approved. The agency concluded that "about 49% of registered voters in the state approve of the measure." Which statement explains whether the conclusion is valid?
Explanation: This AP Statistics question examines undercoverage bias in sampling frames, particularly with outdated methods like landline surveys, and how it affects the truth of conclusions. The issue is that using only landline numbers excludes cell-phone-only households, which may include younger or mobile voters with different opinions, leading to a non-representative sample of registered voters. Choice C distracts by claiming random calling ensures equal chance, but this overlooks undercoverage if the list misses parts of the population. To ensure trustworthy conclusions, the sampling frame must cover the entire population adequately, using methods like mixed-mode surveys to reach all groups. The conclusion here is invalid due to potential undercoverage, teaching that biases from incomplete frames can distort estimates and must be addressed for accurate results.
A school principal wants to estimate the proportion of students at the school who eat breakfast every school day. She emailed a survey link to all 1800 students and used the first 300 responses received within 2 hours. Of those 300, 62% reported eating breakfast every school day. She concluded that "about 62% of all students at the school eat breakfast every school day." Which statement explains whether the conclusion is valid?
Explanation: In AP Statistics, this question assesses the ability to recognize voluntary response bias and nonresponse issues in data collection, emphasizing how the method affects the truthfulness of conclusions. The key issue is that the principal used the first 300 responses from an email survey, creating a voluntary response sample where quicker responders might differ from others, such as being more likely to eat breakfast or have strong opinions. Choice B is a distractor that wrongly claims a sample of 300 is always sufficient, disregarding the bias from self-selection and timing. For trustworthy conclusions, surveys should aim for high response rates and random selection to avoid nonresponse bias, where nonresponders might have different characteristics than responders. Here, the conclusion is invalid due to potential differences between early responders and the rest of the student body, teaching that response mechanisms must be considered for accurate population estimates.
A grocery chain wants to estimate the proportion of all its customers who use digital coupons. For one week, the chain recorded whether each customer used a digital coupon for every 20th customer who checked out between 5 p.m. and 8 p.m. at each store. In the resulting sample of 900 customers, 41% used a digital coupon. The chain concluded that "about 41% of all customers use digital coupons." Which statement explains whether the conclusion is valid?
Explanation: This question in AP Statistics tests understanding of systematic sampling and time-based biases in the sampling frame, relating to data collection's impact on truthful insights. The chain sampled every 20th customer only during evening hours, potentially missing daytime shoppers with different coupon habits, thus not representing all customers. Choice A is a distractor that falsely equates systematic sampling with simple random sampling, but systematic methods can be biased if patterns align with the interval. For trustworthy conclusions, systematic sampling should cover the full population variability, including different times if behaviors vary. The conclusion is invalid due to the limited time frame, emphasizing that sampling must encompass all relevant population segments to avoid bias.
A city's transportation office wants to know whether most adult residents support adding protected bike lanes downtown. Staff members stood outside a Saturday afternoon bike festival and asked 250 adults who walked by, "Do you support adding protected bike lanes downtown?" About 78% said yes. The office concluded that "about 78% of all adult residents in the city support adding protected bike lanes." Which statement explains whether the conclusion is valid?
Explanation: This question tests the skill of identifying sampling bias in data collection methods, specifically convenience sampling, as part of understanding how collected data can tell the truth in AP Statistics. The data collection issue here is that the survey was conducted at a bike festival using a convenience sample of passersby, which likely includes a higher proportion of bike enthusiasts who support bike lanes, leading to an overestimation of citywide support. A common distractor is choice A, which incorrectly assumes that a large sample size alone ensures representativeness, ignoring the bias introduced by the non-random selection method. To draw trustworthy conclusions from survey data, it's essential to use random sampling techniques that give every member of the population an equal chance of being selected, thereby minimizing bias and allowing generalization to the entire population. In this case, the conclusion is invalid because the sample does not represent all adult residents, highlighting the importance of evaluating the sampling frame and method before accepting results.
A restaurant chain investigates the research question: "What is the mean satisfaction rating among all customers who ate at our restaurants last month?" The chain printed a QR code on receipts inviting customers to rate their experience; 1,860 customers responded, with a mean rating of 4.6 out of 5. The chain concluded that the mean satisfaction rating for all customers last month was about 4.6. Which statement explains whether the conclusion is valid?
Explanation: This question examines voluntary response bias in customer satisfaction surveys, a common issue in business research. By printing QR codes on receipts and relying on customers to voluntarily scan and respond, the restaurant creates a sample biased toward customers with strong opinions (either very satisfied or very dissatisfied). Moderately satisfied customers are less motivated to take time to complete surveys, leading to overrepresentation of extreme views. The large sample size (1,860) doesn't correct this fundamental bias in who chooses to respond. The 4.6 average likely overestimates true satisfaction if happy customers are more motivated to share positive experiences. For valid conclusions, the restaurant would need to randomly select customers and actively solicit their feedback, perhaps offering incentives to increase response rates across all satisfaction levels.
A company investigates the research question: "What proportion of employees at the company are satisfied with their health insurance?" The HR department took an alphabetized list of all employees and surveyed every 10th person starting with the 7th name; 210 employees were selected and 196 responded, and 72% of respondents said they were satisfied. HR concluded that about 72% of all employees are satisfied with their health insurance. Which statement explains whether the conclusion is valid?
Explanation: This question tests understanding of systematic sampling and its potential validity. The HR department used systematic sampling (every 10th person from an alphabetized list), which can produce representative samples if the list ordering is unrelated to the characteristic being measured. Since alphabetical order of names is unlikely to correlate with health insurance satisfaction, this method should yield results similar to simple random sampling. The high response rate (196/210 = 93%) minimizes nonresponse bias concerns. The conclusion that 72% of employees are satisfied appears valid, assuming the employee list was complete and current. Systematic sampling is often more practical than simple random sampling while maintaining statistical validity, making it appropriate for this internal company survey.
A streaming service investigates the research question: "What percentage of all subscribers prefer watching with subtitles on?" The company randomly selected 1200 subscriber accounts from its full subscriber database and emailed one survey per selected account; 742 responded, and 55% of respondents reported preferring subtitles on. The company concluded that about 55% of all subscribers prefer subtitles on. Which statement explains whether the conclusion is valid?
Explanation: This question illustrates how nonresponse bias can affect conclusions even when the initial sample selection is random. The streaming service correctly started with a random sample of 1200 subscribers from their database, which is good practice. However, only 742 responded (about 62% response rate), introducing potential nonresponse bias if those who responded differ systematically from non-responders in their subtitle preferences. Despite this concern, the conclusion has some validity because the original selection was random and the response rate is reasonably high. The 55% estimate may be somewhat biased but is likely more accurate than samples from voluntary response or convenience sampling. To improve validity, the company could investigate whether responders and non-responders differ in relevant characteristics.
A school newspaper wants to investigate the research question: "What proportion of students at Central High support starting school at 9:00 a.m.?" The editors posted a poll link on the newspaper's Instagram story and received 312 responses; 68% of respondents selected "support." The editors concluded that about 68% of all Central High students support a 9:00 a.m. start. Which statement explains whether the conclusion is valid?
Explanation: This question tests understanding of how voluntary response samples affect the validity of statistical conclusions. The school newspaper posted a poll on Instagram, which creates a voluntary response sample where only followers who see the story and choose to respond are included. This sampling method likely overrepresents students who follow the newspaper (possibly more engaged students) and those with strong opinions about start times. While 312 responses seems substantial, sample size alone doesn't fix bias from non-representative sampling. To make valid conclusions about all Central High students, the newspaper would need a sampling method that gives every student an equal chance of being selected, such as a simple random sample from the student directory.
A health clinic investigates: What is the average systolic blood pressure of adults (ages 18+) in the county? The population is all adults in the county. The clinic sampled 150 adults by measuring blood pressure of the first 150 patients who came to the clinic in January. The sample mean was 132 mmHg, and the clinic concludes the county's adult mean systolic blood pressure is 132 mmHg. Which statement explains whether the conclusion is valid?
Explanation: This question illustrates classic convenience sampling bias in health research. The clinic sampled only patients who visited in January, creating a biased sample that likely overrepresents people with health concerns serious enough to seek medical care. Choice A correctly identifies that clinic patients are not representative of all county adults - they probably include more people with health issues, potentially including hypertension, compared to the general population. The sample size of 150 is reasonable for estimating a mean, and measuring blood pressure doesn't significantly affect the reading in a way that would invalidate results. Random distribution of patients geographically doesn't make them representative of all adults. To validly estimate county-wide blood pressure levels, researchers need a random sample from the entire adult population, not just those seeking medical care.