Math 3 Quiz: Communicating Statistical Conclusions
20 questions · exam conditions
0:00
Communicating Statistical ConclusionsQuestion 1 of 20

A sports psychologist studied confidence levels in 180 athletes before and after a mental training program. Pre-training confidence averaged 6.2 on a 10-point scale, while post-training averaged 7.4. The improvement was statistically significant (p = 0.001). However, no control group was used, and athletes knew they were being evaluated.

Which statement most appropriately communicates the results while acknowledging study limitations?

The lack of a control group completely invalidates these results, indicating the program likely has no real effect on confidence.
The program has been scientifically proven to increase athlete confidence and should be immediately implemented for all competitive sports teams.
The mental training program appears to be effective for improving athlete confidence, though the study design limits our certainty about the program's true impact.
Athletes probably experienced natural confidence growth over time, suggesting the training program contributed minimally to the observed improvements.
← Back to quizzes

Math 3 Quiz

Math 3 Quiz: Communicating Statistical Conclusions

Practice Communicating Statistical Conclusions in Math 3 with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Communicating Statistical Conclusions, giving you a quick way to practice the rules, question types, and explanations that matter most for Math 3.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A sports psychologist studied confidence levels in 180 athletes before and after a mental training program. Pre-training confidence averaged 6.2 on a 10-point scale, while post-training averaged 7.4. The improvement was statistically significant (p = 0.001). However, no control group was used, and athletes knew they were being evaluated.

Which statement most appropriately communicates the results while acknowledging study limitations?

  1. The lack of a control group completely invalidates these results, indicating the program likely has no real effect on confidence.
  2. The program has been scientifically proven to increase athlete confidence and should be immediately implemented for all competitive sports teams.
  3. The mental training program appears to be effective for improving athlete confidence, though the study design limits our certainty about the program's true impact. (correct answer)
  4. Athletes probably experienced natural confidence growth over time, suggesting the training program contributed minimally to the observed improvements.
Explanation: When you encounter research study questions, focus on balancing what the data shows with the study's methodological limitations. Strong conclusions require strong evidence, while weaker designs call for more cautious interpretations. This study shows a meaningful improvement in confidence scores (6.2 to 7.4) with strong statistical significance (p = 0.001), suggesting the change likely wasn't due to random chance. However, two major design flaws limit our confidence in attributing this improvement to the training program: no control group means we can't compare against athletes who didn't receive training, and the lack of blinding means participants' expectations could have influenced their responses. Option C correctly balances these factors by acknowledging the apparent effectiveness while recognizing the design limitations that prevent definitive conclusions. Option A goes too far by claiming the flaws "completely invalidate" the results - significant improvements still occurred, even if we can't be certain why. Option B makes the opposite error, treating preliminary evidence as definitive proof and recommending immediate widespread implementation without acknowledging any limitations. Option D assumes the improvements were primarily due to natural growth over time, but this dismisses the substantial observed changes without supporting evidence. Remember that in research interpretation, avoid both extremes: don't completely dismiss findings due to methodological issues, but also don't overstate conclusions beyond what the evidence supports. Look for answer choices that appropriately match the strength of language to the quality of the study design.

Question 2

A traffic safety study analyzed accident rates at intersections with three different traffic control methods over one year. Stop signs: 2.1 accidents per month (n=24 intersections), traffic lights: 1.3 accidents per month (n=28 intersections), roundabouts: 0.8 accidents per month (n=22 intersections). Statistical analysis showed p = 0.013. What conclusion is most appropriately stated?

  1. Roundabouts are the safest traffic control method and should replace all stop signs and traffic lights based on this definitive safety analysis.
  2. Installing roundabouts will reduce accidents to exactly 0.8 per month at any intersection according to this comprehensive traffic study.
  3. All traffic control methods are equally safe since accidents still occur with each method and the sample sizes are relatively small.
  4. The evidence suggests meaningful differences in accident rates between traffic control methods, with roundabouts showing relatively lower accident rates. (correct answer)
Explanation: When you encounter a statistical study comparing groups with a p-value, you're being tested on your ability to interpret research findings appropriately and avoid common misinterpretations of statistical evidence. The key information here is that three traffic control methods showed different accident rates (2.1, 1.3, and 0.8 per month) with a p-value of 0.013. Since p < 0.05, this indicates statistically significant differences between the groups, meaning the observed differences are unlikely due to chance alone. The data shows a clear pattern with roundabouts having the lowest rate, traffic lights intermediate, and stop signs highest. Answer D correctly interprets these findings by stating there are meaningful differences between methods and noting roundabouts show relatively lower rates. This appropriately acknowledges the statistical significance while using cautious language ("suggests," "relatively lower") that reflects proper scientific interpretation. Answer A overstates the conclusions by calling this "definitive" and recommending universal replacement—single studies rarely justify such sweeping policy changes. Answer B misinterprets the mean accident rate as a guarantee for future performance at any intersection, ignoring that averages don't predict individual outcomes. Answer C incorrectly dismisses significant findings by focusing on the fact that accidents still occur with all methods, missing the point that we're comparing relative safety, not absolute safety. Remember: when interpreting statistical studies, look for answers that acknowledge significant findings while avoiding overstatement. Proper scientific language uses terms like "suggests" and "relatively" rather than making absolute claims.

Question 3

A company analyzed customer satisfaction surveys from 500 customers across three service locations. Location A averaged 4.2/5 stars (n=180), Location B averaged 3.8/5 stars (n=165), and Location C averaged 4.0/5 stars (n=155). An ANOVA test produced F(2,497) = 8.34, p = 0.0003.

Which statement best communicates the statistical conclusion using appropriate uncertainty language?

  1. Location A provides the best customer service since it has the highest average rating and the ANOVA shows significant differences.
  2. The analysis indicates there are likely meaningful differences in customer satisfaction between locations, with Location A showing relatively higher ratings. (correct answer)
  3. All locations perform equally well in customer service since the differences in ratings are small and within normal business variation.
  4. The data proves that customers will definitely rate Location A higher than the other locations based on superior service quality.
Explanation: Choice B uses appropriate uncertainty language ('indicates', 'likely', 'relatively') while acknowledging the significant ANOVA result. Choice A makes definitive claims about service quality that go beyond what the data shows. Choice C ignores the significant statistical test results. Choice D overstates certainty and makes causal claims about service quality.

Question 4

A pharmaceutical study tested a new medication on 120 patients with high blood pressure. After 8 weeks, 78% of patients showed clinically significant blood pressure reduction compared to 45% in the placebo group (n=115). The chi-square test yielded p = 0.001. What conclusion is most appropriately stated?

  1. The medication definitely works for high blood pressure since 78% is much higher than 45% and the p-value proves effectiveness.
  2. The results strongly suggest the medication is associated with blood pressure improvement, though individual responses may vary considerably. (correct answer)
  3. The medication will work for approximately 78% of all patients with high blood pressure based on this clinical trial evidence.
  4. While the difference appears large, no firm conclusions can be drawn since this is only one study with a limited sample size.
Explanation: Choice B uses appropriate uncertainty language ('strongly suggest', 'associated with', 'may vary') while acknowledging the significant findings. Choice A makes definitive claims and says the p-value 'proves' effectiveness. Choice C inappropriately extrapolates the exact percentage to all patients. Choice D inappropriately dismisses strong significant results from an adequately sized study.

Question 5

A nutrition study tracked weight loss for 90 participants over 12 weeks using three different diet plans. Diet A participants lost an average of 8.2 pounds (SD = 3.1), Diet B participants lost 6.4 pounds (SD = 2.8), and Diet C participants lost 5.9 pounds (SD = 3.4). Each group had 30 participants. The ANOVA yielded p = 0.042.

Which statement represents the most appropriate conclusion using proper statistical language?

  1. Diet A is the most effective weight loss program since participants lost the most weight and the ANOVA proves significant differences exist.
  2. People using Diet A will lose approximately 8.2 pounds based on this controlled study demonstrating superior weight loss outcomes.
  3. The diets are essentially equivalent for weight loss since the differences are small and all participants lost some weight regardless of diet.
  4. The results suggest there may be differences in weight loss effectiveness between diets, with Diet A showing relatively greater average loss. (correct answer)
Explanation: When you encounter ANOVA results in statistical interpretation questions, focus on understanding what the analysis actually tells you versus what it doesn't prove definitively. The correct answer is D because it uses appropriately cautious statistical language. With p = 0.042 < 0.05, the ANOVA indicates statistically significant differences exist somewhere among the three groups, and Diet A did show the highest average weight loss (8.2 pounds). However, proper statistical interpretation requires acknowledging uncertainty—hence "suggests there may be differences" and "relatively greater average loss." Let's examine why the other options are problematic: Answer A makes overly strong causal claims ("Diet A is the most effective") and incorrectly states that ANOVA "proves" differences. Statistical tests suggest evidence, but don't provide absolute proof. Answer B commits a prediction fallacy by claiming "people using Diet A will lose approximately 8.2 pounds." This study shows what happened to these 90 participants, but you cannot guarantee future individual outcomes based on group averages. Answer C ignores the statistical significance entirely. While the numerical differences might seem small, the ANOVA p-value of 0.042 indicates these differences are unlikely due to chance alone, making the diets not "essentially equivalent." Remember this pattern for statistical interpretation questions: look for language that acknowledges uncertainty and avoids overgeneralization. Words like "suggests," "may indicate," and "appears to show" demonstrate proper statistical reasoning, while definitive claims about causation or future predictions usually signal incorrect answers.

Question 6

A medical study compared recovery times for two treatments. Treatment A had a mean recovery time of 12.3 days (SD = 2.1) for 45 patients. Treatment B had a mean recovery time of 14.7 days (SD = 2.8) for 38 patients. A t-test yielded p = 0.032. Which conclusion uses appropriate statistical language?

  1. Treatment A is definitely superior to Treatment B because the p-value is less than 0.05, proving statistical significance exists.
  2. The results suggest Treatment A may be associated with faster recovery times, though replication studies would strengthen this conclusion. (correct answer)
  3. There is insufficient evidence of any difference between treatments since the standard deviations overlap and sample sizes differ.
  4. Treatment A will reduce recovery time by exactly 2.4 days for any patient since this represents the observed difference in means.
Explanation: Choice B appropriately uses uncertainty language ('suggest', 'may be') and acknowledges the need for replication. Choice A overstates certainty by claiming the treatment is 'definitely superior' and that significance is 'proven'. Choice C incorrectly dismisses significant results based on irrelevant criteria. Choice D inappropriately guarantees individual outcomes based on group averages.

Question 7

A researcher studied the relationship between daily screen time and sleep quality scores (0-100 scale) for 200 high school students. The correlation coefficient was r = -0.73, and a regression analysis showed that students who used screens 2 hours more per day had sleep quality scores that were, on average, 15 points lower. The p-value for the correlation was 0.002.

Based on this study, which statement represents the most appropriate conclusion using proper uncertainty language?

  1. The evidence strongly suggests that increased screen time is associated with lower sleep quality, though causation cannot be established from this study alone. (correct answer)
  2. Screen time definitely causes poor sleep quality since the correlation is strong and the p-value shows statistical significance at the 0.05 level.
  3. The data proves that reducing screen time by 2 hours will improve any student's sleep quality score by exactly 15 points on average.
  4. There is likely no meaningful relationship between screen time and sleep quality since correlation doesn't equal causation in observational studies.
Explanation: Choice A correctly uses uncertainty language ('evidence strongly suggests') and acknowledges the limitation that correlation doesn't establish causation. Choice B incorrectly claims causation is proven. Choice C overstates certainty by saying the data 'proves' a causal effect. Choice D incorrectly dismisses a strong correlation (r = -0.73) with a significant p-value.

Question 8

A quality control study examined defect rates in manufacturing. Plant A had 2.3% defective items (n=2000), Plant B had 4.1% defective items (n=1800), and Plant C had 3.2% defective items (n=2200). A chi-square test comparing all three plants yielded p = 0.018. Which communication of results is most appropriate?

  1. Plant A has the best quality control system since it has the lowest defect rate and the statistical test confirms significant differences.
  2. The analysis suggests there are likely differences in defect rates between plants, with Plant A showing relatively lower defect rates. (correct answer)
  3. All plants have similar quality since the defect rates are all below 5% and the differences are practically insignificant for manufacturing.
  4. Plant A will always produce fewer defects than the other plants based on this comprehensive quality analysis of production data.
Explanation: Choice B uses appropriate uncertainty language ('suggests', 'likely', 'relatively') while acknowledging the significant statistical test. Choice A makes definitive claims about quality systems beyond what the data shows. Choice C ignores the significant statistical results by imposing an arbitrary 5% threshold. Choice D overstates certainty with 'always' and 'comprehensive'.

Question 9

A technology company studied employee productivity scores before and after implementing flexible work schedules. Before implementation: mean = 73.2, SD = 8.4 (n=95 employees). After implementation: mean = 79.8, SD = 7.9 (n=95 employees). A paired t-test showed t = 6.12, p < 0.001. The 95% confidence interval for the difference was 4.5 to 8.7 points.

What is the most appropriate way to communicate the results of this study?

  1. Flexible work schedules definitely improve employee productivity since the paired t-test proves a significant increase occurred after implementation.
  2. Implementing flexible schedules will increase any employee's productivity score by 6.6 points based on this definitive workplace study.
  3. The findings suggest that flexible work schedules may be associated with improved productivity, though other factors could influence these results. (correct answer)
  4. The productivity increase is likely due to random variation since the confidence interval shows the true effect could be anywhere from 4.5 to 8.7 points.
Explanation: When interpreting statistical research results, you must distinguish between what the data shows and what you can confidently claim. Statistical significance indicates a pattern in your sample, but proper scientific communication requires acknowledging limitations and avoiding overstated conclusions. The study found a statistically significant increase in productivity scores (mean difference of 6.6 points, t = 6.12, p < 0.001) with a 95% confidence interval of 4.5 to 8.7 points. This suggests a real association between flexible schedules and productivity improvements. However, good statistical practice means communicating results with appropriate caution about causation and generalizability. Choice C correctly uses tentative language ("suggest," "may be associated") and acknowledges that other factors could influence results. This reflects proper scientific communication that avoids overstating conclusions. Choice A is wrong because it uses definitive language ("definitely," "proves") that overstates what statistical tests can establish. Statistical tests show associations, not proof of causation. Choice B is incorrect because it makes absolute predictions ("will increase any employee's productivity") and claims the study is "definitive." No single study can make such broad, certain claims about all future cases. Choice D misinterprets the confidence interval, suggesting the results might be due to random variation. However, the significant p-value (p < 0.001) indicates the results are very unlikely due to chance alone. Remember: When interpreting research results, look for language that appropriately hedges conclusions. Words like "suggests," "may," and "associated with" indicate proper scientific caution, while "proves," "definitely," or "will" signal overstated claims.

Question 10

An educational researcher examined the relationship between homework time and GPA for 250 college students. Students were categorized into three groups: low homework time (0-5 hours/week), moderate (6-15 hours/week), and high (16+ hours/week). The mean GPAs were 2.4, 3.1, and 3.6 respectively, with p < 0.001 from ANOVA.

How should the researcher most appropriately communicate these findings?

  1. Students should spend more time on homework because increased homework time causes higher GPAs according to this definitive analysis.
  2. The evidence suggests a strong association between homework time and academic performance, though other factors likely influence this relationship. (correct answer)
  3. Homework time is the most important factor determining GPA since the p-value shows highly significant differences between all groups.
  4. The relationship is likely coincidental since correlation between homework and grades doesn't indicate any meaningful academic connection.
Explanation: Choice B appropriately uses uncertainty language ('suggests', 'likely influence') and acknowledges that correlation doesn't prove causation. Choice A incorrectly claims causation and calls the analysis 'definitive'. Choice C makes an unsupported claim about relative importance. Choice D inappropriately dismisses strong significant results as 'coincidental'.

Question 11

A psychological study investigated the effect of background music on concentration scores. Participants were randomly assigned to three conditions: no music (mean = 72, SD = 12, n = 35), classical music (mean = 78, SD = 11, n = 33), and pop music (mean = 68, SD = 14, n = 37). ANOVA results showed F(2,102) = 5.23, p = 0.007.

What is the most appropriate way to state the conclusion from this study?

  1. Classical music improves concentration while pop music impairs it, as proven by the significant ANOVA results and clear score differences.
  2. The findings suggest that background music type may be associated with concentration performance, though replication would strengthen these conclusions. (correct answer)
  3. Music has no real effect on concentration since the score differences are small and within one standard deviation of each other.
  4. Students should listen to classical music while studying because it will increase their concentration scores by approximately 6 points.
Explanation: Choice B appropriately uses uncertainty language ('suggest', 'may be associated') and acknowledges the need for replication. Choice A overstates certainty by claiming effects are 'proven' and music 'improves' or 'impairs' concentration. Choice C ignores significant statistical results. Choice D makes inappropriate causal recommendations and guarantees specific score improvements.

Question 12

A sleep study examined the relationship between caffeine consumption and sleep quality scores (1-10 scale) for 180 adults. Participants who consumed no caffeine after 2 PM averaged 7.8 on sleep quality, those who consumed 1-2 caffeinated drinks averaged 6.9, and those who consumed 3+ drinks averaged 5.4. Standard deviations were 1.2, 1.4, and 1.6 respectively. ANOVA results: F(2,177) = 42.1, p < 0.001.

How should researchers most appropriately communicate these findings?

  1. Caffeine consumption after 2 PM definitely impairs sleep quality, as demonstrated by the highly significant ANOVA and clear dose-response pattern.
  2. People should avoid all caffeine after 2 PM because it will reduce their sleep quality scores by 1-2 points on average.
  3. The evidence strongly suggests an association between afternoon caffeine intake and sleep quality, though individual responses likely vary considerably. (correct answer)
  4. Caffeine has no meaningful effect on sleep since people in all groups still achieved moderate sleep quality scores above the midpoint.
Explanation: When interpreting research findings, you need to distinguish between statistical significance and practical implications, while avoiding overstatement of causal claims from correlational data. The correct answer is C because it appropriately acknowledges the strong statistical evidence (F(2,177) = 42.1, p < 0.001) while maintaining scientific caution. The phrase "strongly suggests an association" correctly interprets the significant ANOVA without claiming causation. Additionally, noting that "individual responses likely vary considerably" is crucial given the standard deviations (1.2-1.6), which show substantial overlap between groups despite the mean differences. Answer A is wrong because it overstates the findings by claiming caffeine "definitely impairs" sleep quality. Observational studies show associations, not definitive causal relationships. Answer B makes an inappropriate prescriptive leap from research findings to specific recommendations, plus it mischaracterizes the effect sizes (the differences are 0.9 and 2.4 points, not uniformly "1-2 points"). Answer D incorrectly dismisses meaningful statistical differences just because all groups scored above the scale midpoint. Statistical significance combined with a clear dose-response pattern (7.8 → 6.9 → 5.4) indicates the relationship is meaningful, regardless of absolute score levels. Remember: Strong statistical evidence doesn't equal definitive proof of causation, especially in observational studies. Look for answer choices that acknowledge the strength of findings while maintaining appropriate scientific caution about causality and individual variation.

Question 13

An environmental study measured air quality index (AQI) values in three city districts over 6 months. District A averaged 45.2 AQI (n=180 daily measurements), District B averaged 52.7 AQI (n=175 measurements), and District C averaged 38.9 AQI (n=182 measurements). The 95% confidence intervals were: A (42.8-47.6), B (49.9-55.5), C (36.2-41.6). Which conclusion uses appropriate statistical communication?

  1. District C has the best air quality since 38.9 is the lowest AQI value and its confidence interval is completely below the others.
  2. All districts have similar air quality since the AQI values are all below 60 and within the acceptable range for urban areas.
  3. The data suggests there are likely meaningful differences in air quality between districts, with District C showing relatively cleaner air. (correct answer)
  4. District C will always have cleaner air than the other districts based on this comprehensive six-month environmental monitoring study.
Explanation: When interpreting statistical data with confidence intervals, you need to balance what the data shows while avoiding overconfident claims about causation or future predictions. The correct answer is C because it uses appropriately cautious statistical language. The phrase "data suggests there are likely meaningful differences" acknowledges uncertainty while noting that the non-overlapping confidence intervals indicate statistically significant differences between districts. District C's interval (36.2-41.6) doesn't overlap with A's (42.8-47.6) or B's (49.9-55.5), suggesting real differences exist beyond random variation. Option A makes an overly definitive claim by stating District C "has the best air quality" without acknowledging statistical uncertainty. While the conclusion about non-overlapping intervals is correct, the absolute language is inappropriate for statistical inference. Option B incorrectly focuses on absolute AQI values rather than the statistical comparison between districts. Just because all values fall below 60 doesn't mean the districts have "similar" air quality - the confidence intervals clearly show significant differences. Option D commits the classic error of overgeneralization, claiming District C "will always have cleaner air" based on six months of data. This inappropriately extends findings beyond the study period and ignores the inherent uncertainty in statistical inference. Study tip: When evaluating statistical conclusions, look for language that appropriately reflects uncertainty (words like "suggests," "likely," "indicates") rather than absolute claims. Confidence intervals show what's statistically significant, but proper interpretation requires acknowledging the limitations of any single study.

Question 14

A survey of 300 randomly selected adults found that 68% support a new environmental policy. The margin of error is ±4.2% at the 95% confidence level. A second survey of 150 adults from the same population found 71% support with a margin of error of ±5.8%.

What is the most appropriate way to communicate the conclusion from these two surveys?

  1. The evidence indicates that likely between 64% and 72% of adults support the policy, with consistent findings across both surveys. (correct answer)
  2. Exactly 68% of all adults support this policy since the first survey had a larger sample size and smaller margin of error.
  3. The surveys contradict each other since 68% and 71% are different values, making any conclusion about public opinion unreliable.
  4. Public support has increased from 68% to 71% between the first and second survey, showing a clear trend in opinion change.
Explanation: Choice A correctly uses uncertainty language ('likely') and recognizes that the confidence intervals overlap (63.8-72.2% and 65.2-76.8%), indicating consistent results. Choice B overstates certainty. Choice C incorrectly assumes different point estimates mean contradiction when confidence intervals overlap. Choice D assumes temporal change without evidence the surveys were conducted at different times.

Question 15

A randomized controlled trial tested a new teaching method with 80 students (40 in each group). The experimental group scored an average of 78.5 on the final exam compared to 75.2 for the control group. The 95% confidence interval for the difference is (0.1, 6.5). How should this result be communicated?

  1. The new method is definitely effective since the confidence interval doesn't include zero, proving statistical significance at p < 0.05.
  2. The evidence suggests the new method may improve test scores, though the effect size appears relatively modest and warrants further study. (correct answer)
  3. The difference of 3.3 points is too small to be educationally meaningful, regardless of statistical significance in this particular study.
  4. No conclusion can be drawn since the confidence interval is too wide, ranging from barely positive to moderately positive effects.
Explanation: Choice B appropriately uses uncertainty language ('suggests', 'may') and acknowledges both the positive finding and its limitations. Choice A overstates certainty by claiming the method is 'definitely effective' and that significance is 'proven'. Choice C makes an unsupported judgment about educational meaningfulness. Choice D incorrectly dismisses significant results due to interval width.

Question 16

A quality control inspector tested 500 randomly selected products from a manufacturing line. She found 23 defective items, giving a defect rate of 4.6%. The company's acceptable defect rate is 3%. How should she most appropriately report her findings to management?

  1. The sample size is too small to draw any reliable conclusions about the overall defect rate of the manufacturing process.
  2. The manufacturing line is definitely producing too many defective products and must be shut down immediately to prevent customer complaints.
  3. The defect rate of 4.6% is essentially equivalent to the 3% target since both values are relatively small percentages.
  4. The current defect rate appears to exceed acceptable limits, suggesting the manufacturing process likely requires adjustment and further monitoring. (correct answer)
Explanation: When analyzing quality control data, you need to balance statistical evidence with practical business judgment. The key is interpreting sample results appropriately without overreacting or underreacting to the findings. The correct approach is D because the data shows a meaningful difference between the observed defect rate (4.6%) and the target (3%). With 500 products tested, this sample size provides reasonable reliability for drawing conclusions about the manufacturing process. The 53% higher defect rate than acceptable suggests a genuine process issue that warrants investigation and corrective action, while acknowledging that further monitoring is needed to confirm the trend. A is incorrect because 500 is actually a substantial sample size for quality control purposes. This sample provides sufficient data to identify potential process issues, though continued monitoring remains important. B represents dangerous overreaction. While the defect rate exceeds targets, immediately shutting down production based on a single sample would be costly and premature. Quality control requires measured responses, not panic decisions. C demonstrates poor statistical judgment. A 4.6% defect rate is 53% higher than the 3% target—this isn't "essentially equivalent" but represents a significant deviation that could impact customer satisfaction and costs. Study tip: In quality control questions, look for balanced responses that acknowledge statistical evidence while avoiding extreme reactions. The best answers typically call for process investigation and continued monitoring rather than immediate shutdowns or dismissing concerning data.

Question 17

A researcher studied the relationship between daily exercise time and sleep quality scores among 200 college students. The correlation coefficient was r = 0.42, with a p-value of 0.003. Students who exercised more than 60 minutes daily had an average sleep quality score of 7.8 (on a 10-point scale), while those exercising less than 30 minutes daily averaged 6.1.

Which statement most appropriately communicates the conclusions from this study?

  1. The evidence suggests there is likely a positive association between exercise time and sleep quality, though we cannot conclude that exercise directly causes better sleep. (correct answer)
  2. The data proves that exercising more than 60 minutes daily will significantly improve sleep quality for all college students.
  3. There is likely no meaningful relationship between exercise and sleep quality since the correlation is less than 0.5.
  4. The evidence suggests that students should exercise less than 30 minutes daily to maintain adequate sleep quality scores.
Explanation: Choice A appropriately uses uncertainty language ('suggests', 'likely') and correctly distinguishes between association and causation. The correlation of 0.42 with p = 0.003 indicates a statistically significant positive relationship. Choice B overstates conclusions by claiming 'proof' and 'will improve' without acknowledging uncertainty. Choice C incorrectly dismisses a moderate correlation as meaningless. Choice D misinterprets the direction of the relationship.

Question 18

A meteorologist analyzed 20 years of weather data and found that when the barometric pressure drops below 29.8 inches of mercury, there is a 78% chance of rain within 24 hours. Last Tuesday, the pressure dropped to 29.6 inches.

Based on this analysis, how should the meteorologist most appropriately communicate the likelihood of rain for Wednesday?

  1. Rain is virtually certain for Wednesday since the pressure dropped below the critical threshold established by the historical data analysis.
  2. There will likely be rain on Wednesday, as historical patterns suggest a high probability when pressure drops to this level. (correct answer)
  3. Wednesday's weather remains unpredictable since the 78% figure represents only past data and cannot determine future weather events.
  4. Rain is unlikely on Wednesday because the pressure reading of 29.6 is still relatively close to the 29.8 threshold value.
Explanation: Choice B appropriately communicates high probability using uncertainty language ('likely', 'suggests') while acknowledging the basis in historical patterns. Choice A overstates certainty ('virtually certain') for a 78% probability. Choice C inappropriately dismisses the predictive value of well-established statistical patterns. Choice D incorrectly interprets the data by focusing on proximity to threshold rather than the actual probability associated with readings below 29.8.

Question 19

A school district analyzed standardized test scores from 1,200 students across 15 schools. They found that schools with smaller class sizes (under 20 students) had average scores 12 points higher than schools with larger classes (over 25 students). The difference was statistically significant (p = 0.02). However, the smaller schools also had newer facilities and more experienced teachers on average.

Which statement most appropriately communicates these findings while acknowledging potential confounding factors?

  1. Larger class sizes appear to be more effective since they likely provide students with greater peer interaction and collaborative learning opportunities.
  2. Reducing class sizes will definitely increase test scores by 12 points based on this comprehensive analysis of district data.
  3. The presence of confounding variables makes it impossible to determine any relationship between class size and academic performance.
  4. The evidence suggests smaller class sizes are likely associated with higher test scores, though other school factors may contribute to this relationship. (correct answer)
Explanation: When you encounter research findings with multiple variables, the key is distinguishing between correlation and causation while acknowledging confounding factors that could influence the results. The data shows a clear association: schools with smaller class sizes had higher test scores by 12 points, and this difference was statistically significant (p = 0.02). However, these same schools also had newer facilities and more experienced teachers. This creates confounding variables—factors that could explain the test score differences independently of class size. Answer D correctly acknowledges both the observed relationship and the limitation. It states that smaller class sizes are "likely associated" with higher scores while noting that "other school factors may contribute." This language appropriately reflects the strength of the evidence while maintaining scientific caution about causation. Answer A contradicts the actual findings by claiming larger classes are more effective, despite the data showing the opposite pattern. Answer B makes a causal claim ("will definitely increase") that the correlational data cannot support, especially with confounding variables present. Answer C goes too far in the opposite direction, claiming confounding variables make it "impossible to determine any relationship"—but we can identify associations even when causation is unclear. Remember that statistical significance indicates a real pattern exists, but correlation studies cannot prove causation, especially when confounding variables are present. Look for answer choices that acknowledge both the strength of the evidence and its limitations, using language like "associated with" or "suggests" rather than definitive causal claims.

Question 20

An environmental scientist measured air pollution levels in two neighborhoods over 6 months. Neighborhood A (near a major highway) showed average PM2.5 levels of 18.3 μg/m³, while Neighborhood B (residential area) averaged 12.1 μg/m³. The EPA standard is 15 μg/m³. Statistical analysis yielded p = 0.008 for the difference between neighborhoods.

How should the scientist most appropriately communicate these findings to the city council?

  1. Both neighborhoods have essentially similar air quality since the pollution measurements are relatively close to each other and the EPA standard.
  2. Neighborhood A is definitively unsafe for residents and should be evacuated immediately due to dangerous pollution levels exceeding federal limits.
  3. The evidence strongly suggests that air quality likely differs between the neighborhoods, with the highway area probably exceeding safe standards. (correct answer)
  4. The pollution data is inconclusive because environmental measurements are inherently unreliable and cannot guide policy decisions effectively.
Explanation: When interpreting scientific data for policy-makers, you need to balance statistical evidence with practical implications while avoiding both overstatement and understatement of findings. The correct approach is C because it appropriately reflects what the data shows. With p = 0.008 (well below the typical 0.05 threshold), there's strong statistical evidence that the neighborhoods genuinely differ in air quality. Neighborhood A's average of 18.3 μg/m³ exceeds the EPA standard of 15 μg/m³, while Neighborhood B's 12.1 μg/m³ stays within safe limits. The language "likely differs" and "probably exceeding" appropriately conveys statistical confidence while acknowledging inherent uncertainty in environmental measurements. A is wrong because it ignores both the statistical significance (p = 0.008) and the practical significance of one neighborhood exceeding EPA standards. A 6.2 μg/m³ difference isn't negligible when it represents crossing a safety threshold. B overstates the findings dramatically. While 18.3 μg/m³ exceeds the EPA standard, it doesn't indicate immediate danger requiring evacuation. Responsible science communication avoids catastrophizing moderate exceedances. D incorrectly dismisses valid environmental data. While all measurements have uncertainty, that doesn't make them "unreliable" or useless for policy. The statistical significance suggests the difference is real, not just measurement noise. Study tip: On science communication questions, look for answers that match the strength of evidence to the strength of language used. Strong statistical evidence (low p-values) supports confident but measured conclusions, not absolute certainty or complete dismissal.