All questions
Question 1
A researcher claims that a new study method improves test scores by an average of 15 points. The study included 8 students who used the new method for one week, with scores improving from 72 to 87 points on average. The standard deviation of the improvements was 18 points. Which statement best evaluates this claim?
- The claim is well-supported because the average improvement of 15 points matches the researcher's prediction exactly.
- The claim is questionable because the high variability relative to the effect size suggests the improvement may not be reliable. (correct answer)
- The claim is invalid because the study period of one week is too short to measure meaningful academic improvement.
- The claim is strong because all students showed positive improvement and the effect size exceeds typical measurement error.
Explanation: With a standard deviation of 18 points and a mean improvement of 15 points, the variability is larger than the claimed effect. Combined with the very small sample size (n=8), this high variability makes the results unreliable. Choice A ignores variability and sample size. Choice C focuses on study duration rather than statistical issues. Choice D incorrectly assumes all students improved and misinterprets the relationship between effect size and measurement error.
Question 2
A fitness tracker company claims their device motivates users to walk 2,000 more steps per day on average. They collected data from 150 users over 8 weeks, comparing their daily steps before and after getting the device. The mean increase was 2,100 steps with a standard deviation of 3,200 steps. What is the most concerning aspect of this claim?
- The study design lacks a control group of people who didn't receive fitness trackers to account for seasonal activity changes.
- The extremely high variability in step increases suggests many users had no benefit while others had unusually large increases. (correct answer)
- The 8-week study period is too brief to determine whether the step increase represents a sustainable long-term behavior change.
- The sample size of 150 users is inadequate for detecting reliable differences in daily step counts across diverse populations.
Explanation: A standard deviation of 3,200 steps when the mean increase is only 2,100 steps indicates enormous variability. This suggests the effect is highly inconsistent across users, with many likely showing no increase or even decreases. The high variability undermines confidence in the claimed average effect. Choice A identifies a valid design flaw but the variability issue is more fundamental. Choice C raises timing concerns but 8 weeks is reasonable for this type of study. Choice D incorrectly suggests 150 is too small.
Question 3
A consumer advocacy group tested battery life for two smartphone brands by having 25 users of each brand record their daily usage and battery duration over 2 weeks. Brand A averaged 11.2 hours of battery life (SD = 2.8 hours) while Brand B averaged 9.8 hours (SD = 1.9 hours).
The advocacy group concluded that Brand A has superior battery performance. What is the most significant methodological concern with this conclusion?
- The study period of 2 weeks is insufficient to account for battery degradation patterns that occur over months of typical usage.
- The sample size of 25 users per brand provides inadequate statistical power to detect meaningful differences in battery performance reliably.
- Users were not randomly assigned to phone brands, so differences in usage patterns between brand communities could explain the results. (correct answer)
- Self-reported usage data introduces measurement error that could systematically bias battery life estimates in either direction for both brands.
Explanation: When evaluating research conclusions, you need to distinguish between valid statistical findings and sound causal inferences. A study can show real differences between groups while still having serious flaws in its design that undermine the conclusions.
The correct answer is C because this study has a fundamental confounding variable problem. Users weren't randomly assigned to phone brands - they chose their own phones. This means Brand A and Brand B users likely differ in systematic ways beyond just their phone choice. Heavy users might gravitate toward phones marketed for long battery life, while casual users might prioritize other features. Different user communities could have vastly different usage patterns, charging habits, or even definitions of "battery duration." These pre-existing differences between the groups could entirely explain the 1.4-hour difference in battery life, making it impossible to attribute the results to the phones themselves.
Option A is wrong because while 2 weeks isn't ideal for long-term patterns, it's sufficient to detect basic performance differences if the study design were sound. Option B is incorrect - 25 users per group provides reasonable statistical power for detecting a 1.4-hour difference with these standard deviations. Option D identifies a real limitation (measurement error from self-reporting), but this would likely affect both brands equally and wouldn't systematically bias results toward one brand.
Remember this key principle: when you see comparative studies where participants weren't randomly assigned to groups, always ask whether the groups might differ in ways other than the treatment being studied. Self-selection bias is one of the most common threats to valid causal conclusions.
Question 4
A social media company reports that their new algorithm increased user engagement by 18% based on comparing average daily time spent on the platform before and after the change. The analysis used data from 2.3 million active users over a 6-week period. Despite the large sample size, what is the primary limitation in interpreting this engagement increase?
- Seasonal trends in social media usage could account for increased engagement rather than the algorithm change itself producing the effect.
- Individual user variations in engagement patterns could create misleading averages despite the large overall sample size of 2.3 million users.
- The 6-week observation period may be too short to distinguish between temporary novelty effects and sustained engagement improvements. (correct answer)
- The 18% increase might reflect changes in how engagement time is measured rather than actual changes in user behavior patterns.
Explanation: When evaluating study results that claim to show improvement, you need to distinguish between genuine long-term effects and temporary changes that might fade over time. This is especially critical when analyzing behavioral interventions like algorithm changes.
The correct answer is C because 6 weeks is insufficient to determine whether the 18% engagement increase represents lasting behavioral change or just a temporary "novelty effect." When platforms introduce new features or algorithms, users often initially engage more due to curiosity, changed content recommendations, or simply noticing something different. However, this heightened engagement frequently diminishes as users adapt to the changes and return to baseline behaviors. To claim genuine improvement, you'd need months of data to establish that the increase persists beyond the initial adjustment period.
Option A is incorrect because seasonal trends would affect engagement regardless of algorithm changes, and the comparison appears to be measuring before-versus-after changes during the same timeframe. Option B misunderstands statistical principles—with 2.3 million users, individual variations would actually average out effectively, making the sample size quite robust for detecting real effects. Option D suggests measurement issues, but there's no indication in the question that the company changed how they track engagement time.
Remember this pattern: when evaluating any intervention study, always consider the time dimension. Short observation periods can't distinguish between temporary novelty effects and sustainable improvements. Look for this issue especially in behavioral studies, where initial enthusiasm often masks the true long-term impact.
Question 5
A school principal claims that students who eat breakfast score 12 points higher on standardized tests than those who skip breakfast. This conclusion was based on comparing test scores of 45 breakfast-eating students (mean: 78) with 38 students who skipped breakfast (mean: 66). What additional information is most critical for evaluating this claim?
- The standard deviations of test scores within each group to assess whether the 12-point difference is statistically meaningful.
- The demographic characteristics of students in each group to determine if factors other than breakfast could explain the difference. (correct answer)
- The specific nutritional content of the breakfasts consumed to establish which nutrients might be responsible for improved performance.
- The time interval between eating breakfast and taking the test to verify that the nutritional effects would still be active.
Explanation: The most critical missing information is whether the groups differ in other important ways (socioeconomic status, sleep habits, family support, etc.) that could explain the score difference. This observational study cannot establish causation without controlling for confounding variables. Choice A is important but secondary to establishing whether breakfast is actually the cause. Choices C and D assume breakfast is the cause and focus on mechanism rather than establishing causation.
Question 6
An environmental group studied air quality improvements after a city implemented new emission standards. They measured pollution levels at 15 monitoring stations for 3 months before and 3 months after the standards took effect, finding an average reduction of 22% across all stations.
The group claims the emission standards successfully improved air quality. What information would be most valuable for evaluating this claim?
- Pollution measurements from similar cities without new emission standards during the same time period to control for seasonal and regional factors. (correct answer)
- Detailed breakdown of which specific pollutants decreased to determine if the standards targeted the most harmful emissions effectively.
- Analysis of weather patterns during both measurement periods since wind and precipitation significantly affect measured pollution levels.
- Information about industrial activity levels and traffic patterns that might have changed independent of the emission standards implementation.
Explanation: A control group of similar cities without emission standards would help determine if the 22% reduction was due to the standards or other factors (seasonal changes, regional weather patterns, economic conditions). This is the strongest way to establish causation. Choice B focuses on mechanism rather than establishing causation. Choice C addresses one potential confounding factor but control cities would handle this and other factors simultaneously. Choice D identifies some confounding variables but control cities would be more comprehensive.
Question 7
A university admissions officer claims that students who submit applications early are accepted at a 75% rate compared to 45% for regular applicants. This analysis included 320 early applications and 1,200 regular applications from last year. Which factor is most important for evaluating whether early submission actually improves acceptance chances?
- Whether early applicants systematically differ from regular applicants in academic qualifications, extracurricular activities, or other admission factors. (correct answer)
- Whether the university has explicit policies that give preference to early applications versus evaluating all applications using identical criteria.
- Whether the 30 percentage point difference in acceptance rates is statistically significant given the sample sizes of each applicant pool.
- Whether the university's acceptance rate has remained stable over multiple years or if last year's data represents an unusual pattern.
Explanation: The key question is whether early applicants are inherently different from regular applicants in ways that affect their acceptance chances. If early applicants tend to be more qualified, the higher acceptance rate might reflect their qualifications rather than the timing of their applications. Choice B addresses policy but not the fundamental confounding issue. Choice C focuses on statistical significance rather than causation. Choice D concerns generalizability rather than the causal relationship.
Question 8
A local newspaper reported that crime rates dropped 25% in neighborhoods with new street lighting based on police data from the past year. The reporter compared crime incidents in 12 neighborhoods that received new LED streetlights with 8 similar neighborhoods that kept old lighting systems.
Which factor most undermines the validity of concluding that new street lighting caused the crime reduction?
- The study period of one year may be insufficient to establish long-term trends in neighborhood crime patterns and seasonal variations.
- The unequal sample sizes of neighborhoods (12 vs 8) creates statistical imbalance that could skew the comparison results artificially.
- Neighborhoods selected for lighting upgrades may have had other crime prevention initiatives that could account for the observed reduction. (correct answer)
- The 25% reduction figure lacks context about baseline crime rates and whether this represents a meaningful absolute decrease in incidents.
Explanation: The biggest threat to causal inference is confounding variables. Neighborhoods chosen for lighting upgrades likely received them for specific reasons and may have had other concurrent crime prevention efforts. This makes it impossible to isolate the lighting effect. Choice A raises timing concerns but one year is reasonable for this analysis. Choice B wrongly suggests unequal sample sizes create bias. Choice D misses the point about causation vs. effect size.
Question 9
A city council member claims that the new bike lane reduced traffic accidents by 40% based on a comparison of accident data. Before the bike lane: 25 accidents over 6 months at the intersection. After the bike lane: 15 accidents over 6 months at the same intersection.
What is the most significant limitation in evaluating this accident reduction claim?
- The study fails to account for seasonal variations in traffic patterns and weather conditions that affect accident rates.
- The sample size of accidents is too small to detect meaningful differences between the two time periods reliably.
- The percentage calculation is incorrect because it should compare accidents per day rather than total accidents over each period.
- The study lacks a control location without bike lane changes to determine if the reduction occurred citywide or specifically here. (correct answer)
Explanation: The fundamental flaw is the lack of a control group. Without comparing to similar intersections that didn't get bike lanes, we can't determine if the reduction was due to the bike lane or other factors affecting the entire city. Choice A mentions valid concerns but they're secondary to the control group issue. Choice B overstates the sample size problem. Choice C is incorrect - the calculation method is appropriate for the claim being made.
Question 10
A health insurance company analyzed claims data and found that members who use their wellness app have 30% lower medical costs than non-users. The analysis included 12,000 app users and 45,000 non-users over an 18-month period. The company promotes this as evidence that their app reduces healthcare costs. What is the strongest criticism of this conclusion?
- The observational design cannot establish causation because app users likely differ from non-users in health consciousness and baseline health status. (correct answer)
- The 18-month timeframe may be insufficient to capture long-term healthcare cost patterns and the full impact of wellness interventions.
- The large difference in group sizes (12,000 vs 45,000) creates statistical imbalance that could artificially inflate the apparent cost difference.
- Self-selection bias affects the results since members chose whether to use the app rather than being randomly assigned to user groups.
Explanation: This is a classic selection bias problem. People who choose to use wellness apps are typically more health-conscious, may already be healthier, and likely engage in other healthy behaviors that reduce medical costs. The cost difference probably reflects these pre-existing differences rather than app effectiveness. Choice B underestimates the adequacy of 18 months. Choice C incorrectly suggests unequal sample sizes create bias. Choice D identifies self-selection but doesn't explain why it's problematic as clearly as Choice A.
Question 11
A supplement company surveyed 200 customers who had purchased their vitamin product and found that 78% reported feeling more energetic after 30 days of use. The company claims their vitamin increases energy levels in most users. Based on this survey design, what is the primary concern about this claim?
- The sample size of 200 customers is insufficient to make generalizations about the effectiveness of vitamin supplements.
- The 30-day time period may be too short to observe genuine physiological changes from vitamin supplementation.
- Selection bias likely occurred because only customers who purchased the product were surveyed, excluding potential users who chose not to buy.
- Response bias is probable since customers who felt no benefit may be less likely to respond to surveys from the company. (correct answer)
Explanation: The primary concern is response bias - customers who experienced positive effects are much more likely to respond to the company's survey than those who saw no benefit. This systematically inflates the percentage of positive responses. Choice A is wrong; 200 is adequate sample size. Choice B focuses on timing rather than bias. Choice C incorrectly identifies the bias - surveying customers who bought the product is appropriate, but getting biased responses from them is the real problem.
Question 12
A pharmaceutical company reports that their new medication reduced symptoms in 85% of trial participants compared to 45% in the placebo group. However, the trial had 40 participants in the treatment group and 35 in the placebo group. A medical reviewer questions whether this difference is meaningful. Which concern is most justified?
- The sample sizes are too small to reliably detect true differences between treatment and placebo effects in medical trials. (correct answer)
- The unequal group sizes (40 vs 35) create systematic bias that favors the treatment group in statistical comparisons.
- The 40 percentage point difference is too large to be credible and likely indicates measurement errors or data fabrication.
- The lack of a control group receiving no intervention makes it impossible to determine if either treatment provided real benefit.
Explanation: With only 40 and 35 participants, the sample sizes are indeed too small for reliable medical conclusions. Small samples have high variability and low statistical power, making chance differences appear significant. Choice B is wrong - slight size differences don't create systematic bias. Choice C incorrectly assumes large differences must be fabricated. Choice D misunderstands the design - the placebo group IS the control.
Question 13
A pharmaceutical company tests a new medication by giving it to 100 patients with a specific condition. After 3 months, 85 patients show improvement. The company claims their medication is highly effective.
Which factor most seriously undermines the validity of the company's effectiveness claim?
- The study period of 3 months may not be sufficient to observe long-term treatment effects
- The sample size of 100 patients is too small to establish medication effectiveness reliably
- The study should have included patients with different severity levels of the condition
- The absence of a placebo control group makes it impossible to isolate the medication's effect (correct answer)
Explanation: When evaluating scientific studies, you need to identify what factors could prevent researchers from drawing valid conclusions about cause and effect. The fundamental question here is whether we can confidently attribute the observed improvements to the medication itself.
The correct answer is D because without a placebo control group, there's no way to determine what would have happened to these patients without the medication. Many conditions improve naturally over time, and patients often experience psychological benefits just from receiving treatment (the placebo effect). The 85% improvement rate might seem impressive, but if 80% of patients typically improve without any treatment, then the medication would actually be harmful. A control group receiving a placebo would reveal the medication's true effect by showing the baseline improvement rate.
Choice A is wrong because while long-term effects matter for comprehensive evaluation, they don't undermine claims about short-term effectiveness. Choice B incorrectly assumes the sample size is inadequate—100 patients can provide meaningful results if the study is properly designed. Choice C addresses study comprehensiveness but doesn't fundamentally challenge the validity of measuring effectiveness within the studied population.
Study tip: In research design questions, always look for the absence of proper controls first. The most sophisticated study becomes meaningless without a way to isolate the variable being tested. Remember that correlation doesn't prove causation—you need controls to establish that the treatment, not other factors, caused the observed results.
Question 14
A fitness app company analyzes data from 10,000 users and finds that people who log workouts 5+ times per week lose an average of 2.3 pounds per month, while those who log fewer workouts lose 0.8 pounds per month. They conclude that frequent app use causes greater weight loss.
What represents the most serious flaw in concluding that app usage frequency causes the weight loss difference?
- People who frequently log workouts may have higher motivation levels that drive both app use and weight loss (correct answer)
- The study doesn't account for users' initial weight, age, or other demographic factors affecting weight loss
- Self-reported workout data may be inaccurate, leading to misclassification of user activity levels
- The 1.5-pound monthly difference may not be clinically significant for long-term health outcomes
Explanation: When you encounter studies claiming that one factor "causes" another, always examine whether the researchers have properly established causation versus correlation. The most fundamental threat to causal conclusions is the presence of confounding variables—third factors that influence both the supposed cause and the effect.
Choice A correctly identifies the most serious flaw: motivation level is likely a confounding variable that drives both frequent app usage AND weight loss success. Highly motivated people naturally log workouts more often and also stick to diet and exercise plans more consistently. The app usage itself may not be causing the weight loss—both behaviors might stem from the same underlying motivation.
Choice B describes a real limitation, but controlling for demographics wouldn't address the core causation problem. Even with matched groups by age and weight, the motivation confound would remain. Choice C points out data quality issues that could affect the study's accuracy, but this doesn't challenge the fundamental logic of the causal claim—it's more about measurement error. Choice D questions the practical significance of the results, but this doesn't address whether the relationship is truly causal.
The key insight is that correlation studies can never prove causation when alternative explanations exist. Here, the "third variable problem"—where motivation influences both app use and weight loss—provides a more plausible explanation than direct causation.
Study tip: When evaluating causal claims, always ask "What else could explain both variables?" The most serious flaws usually involve confounding variables, not just methodological limitations.
Question 15
A school district evaluates a reading intervention program by comparing reading scores of 60 students who received the intervention with 60 students who did not. Both groups started with similar average reading levels. After 6 months, the intervention group improved by an average of 1.2 grade levels while the control group improved by 0.7 grade levels.
Which aspect of this study design most strengthens confidence in attributing the 0.5 grade level difference to the intervention?
- The initial similarity in reading levels between groups, controlling for baseline differences (correct answer)
- The 6-month duration providing adequate time for reading improvements to develop and stabilize
- The use of equal sample sizes in both intervention and control groups for balanced comparison
- The measurement in grade levels rather than raw scores, providing educationally meaningful units
Explanation: When evaluating experimental studies, you need to identify what makes the results credible by eliminating alternative explanations for the observed differences. The key question is: could something other than the intervention explain why one group performed better?
The correct answer is A because initial similarity in reading levels is crucial for establishing causation. If the groups started at different baseline levels, any final difference might simply reflect those pre-existing differences rather than the intervention's effect. By ensuring both groups began with similar reading abilities, researchers can confidently attribute the 0.5 grade level difference to the intervention itself, not to one group being naturally stronger readers.
Let's examine why the other options don't strengthen causal inference as much: B suggests the 6-month duration was important, but timing doesn't address whether other factors caused the difference. C mentions equal sample sizes, which helps with statistical power but doesn't eliminate confounding variables—even with balanced groups, if one group started ahead, that advantage could explain the results. D focuses on measurement units, but using grade levels versus raw scores doesn't strengthen the causal argument; it's about interpretability, not validity.
Remember this pattern: in experimental design questions, look for what eliminates alternative explanations for the results. Controlling for baseline differences through matching or random assignment is often the strongest design feature because it ensures groups are truly comparable from the start.
Question 16
A university researcher claims that students who use study groups perform better on exams than those who study alone. The researcher surveyed 150 students and found that study group participants had a mean exam score of 84.2, while individual studiers had a mean score of 81.7.
Given that both groups had standard deviations of approximately 12 points, which evaluation of the researcher's claim is most appropriate?
- The claim is well-supported because study group participants consistently outperformed individual studiers
- The claim is questionable because the 2.5-point difference is small relative to the variability in scores (correct answer)
- The claim is invalid because the researcher should have used median scores instead of means
- The claim is unreliable because 150 students represents too small a sample for educational research
Explanation: With standard deviations of 12 points, a 2.5-point difference in means represents only about 0.2 standard deviations - a very small effect that could easily be due to random variation rather than a real difference. This difference is not practically meaningful given the variability in the data. Choice A ignores the statistical significance relative to variability. Choice C is incorrect because means are appropriate for this type of analysis. Choice D is incorrect because 150 students is generally sufficient for this type of study.
Question 17
A researcher claims that students who eat breakfast score higher on standardized tests. The study included 200 students: 120 who regularly eat breakfast (mean score: 78) and 80 who skip breakfast (mean score: 72). The standard deviation for both groups was 15 points. What additional information is most crucial for evaluating this claim?
- Whether the 6-point difference is statistically significant given the sample sizes and variability
- Whether the students were randomly selected from the same school district population
- Whether other factors like socioeconomic status were controlled for in the analysis (correct answer)
- Whether the standardized test used is a reliable measure of academic achievement
Explanation: While all factors matter, controlling for confounding variables like socioeconomic status is most crucial. Students who eat breakfast regularly may come from families with more resources, better home environments, or different cultural values around education - any of which could explain the score difference. Choice A addresses statistical significance but doesn't address causation. Choice B is important but less critical than controlling for confounders. Choice D addresses measurement validity but assumes the test is already standardized and reliable.
Question 18
A school district wants to determine if a new math curriculum improves student performance. They implement the new curriculum in 5 schools and compare test scores to the previous year's scores from the same schools. After one semester, they find that average test scores increased by 12 points across the 5 schools.
Which of the following represents the most significant limitation in concluding that the new curriculum caused the improvement in test scores?
- The sample size of 5 schools is too small to detect meaningful differences in educational outcomes
- The study lacks a control group using the old curriculum during the same time period (correct answer)
- The 12-point increase is not large enough to be considered educationally significant
- The study should have included students from multiple grade levels to ensure validity
Explanation: The most critical flaw is the lack of a proper control group. Comparing this year's scores to last year's scores introduces confounding variables (different students, different teachers, different conditions, potential natural improvement over time). A proper study would need a control group using the old curriculum during the same semester. Choice A is incorrect because 5 schools can provide meaningful data if properly designed. Choice C is incorrect because the significance of a 12-point increase depends on the scale and context. Choice D is incorrect because including multiple grade levels isn't necessary for validity if the study is properly focused.