All questions
Question 1
A research lab performs 8 independent hypothesis tests, investigating the effects of 8 different chemical compounds on crop yield. To control the family-wise error rate at 0.04, the researchers decide to use the Bonferroni correction. What is the adjusted significance level (alpha) they must use for each individual test?
- 0.005 (correct answer)
- 0.04
- 0.05
- 0.32
Explanation: The Bonferroni correction is a method used to counteract the multiple comparisons problem. To maintain a desired family-wise error rate (FWER), the significance level (alpha) for each individual test is adjusted. The formula is αadj=kFWER, where k is the number of tests. In this case, αadj=80.04=0.005. Each of the 8 tests must have a p-value less than 0.005 to be considered statistically significant. Question 2
A research team tests a learning intervention across 12 different schools. They find significant improvements in 4 schools (p<0.05) and conclude the intervention works in certain educational contexts. They then analyze characteristics of the 4 'successful' schools to develop implementation guidelines. What statistical reasoning error undermines their approach?
- They assume that statistical significance in individual schools translates to practical significance across different educational contexts and populations.
- They fail to account for school-level clustering effects that violate independence assumptions in their significance testing procedures.
- They treat chance occurrences as meaningful patterns by analyzing characteristics of schools that may have shown significance purely by random variation. (correct answer)
- They commit survivorship bias by only studying successful implementations while ignoring the factors that contributed to failure in other schools.
Explanation: With 12 schools tested at α = 0.05, we'd expect about 0.6 false positives by chance alone. Finding 4 significant results isn't dramatically different from chance expectation. The critical error is analyzing the 'successful' schools as if their success definitively indicates real effects, when some or all could be false positives. Building implementation guidelines based on characteristics of potentially random successes is fundamentally flawed. Choice A focuses on statistical vs. practical significance, which isn't the main issue. Choice B raises clustering concerns but doesn't address the core problem. Choice D mentions survivorship bias, but this scenario is more about multiple comparisons and post-hoc pattern seeking than selective survival.
Question 3
A marketing team conducts A/B tests on 15 different email subject lines throughout the month, testing each against their standard subject line. They find that 3 subject lines significantly outperform the standard (p<0.05) and immediately implement these as their new email strategy. Six months later, the performance gains have disappeared. Which combination of statistical pitfalls most likely explains this outcome?
- Multiple comparisons inflated false positive rates initially, and regression to the mean caused the apparent winners to perform closer to average over time. (correct answer)
- Survivorship bias in the initial testing phase and sampling error in the follow-up period created inconsistent results across different time frames.
- Data snooping during the testing phase and confirmation bias in the implementation phase led to overestimation of the true effect sizes.
- Correlation versus causation confusion initially and external validity issues during implementation caused the performance differences to disappear.
Explanation: This scenario involves two key pitfalls. First, testing 15 comparisons at α = 0.05 creates a high probability of false positives due to multiple comparisons (expected false positives ≈ 15 × 0.05 = 0.75). Second, the initially 'winning' subject lines were likely statistical outliers that regressed toward the mean over time. Choice B incorrectly identifies survivorship bias. Choice C mentions data snooping, which is related but not the primary issue here, and confirmation bias doesn't explain the disappearing performance. Choice D focuses on correlation/causation, which isn't the main problem in this A/B testing context.
Question 4
An analyst reports that a new website design increases conversion rates by 2.3% (p=0.048). When asked about their methodology, they reveal they tested 8 different design variants simultaneously and selected the best performer for this final comparison against the control. They also mention they stopped the test early when they first achieved significance. This result combines which statistical pitfalls?
- Multiple comparisons from testing 8 variants, data snooping from selecting the best performer, and optional stopping that inflates Type I error rates. (correct answer)
- Survivorship bias from only reporting the best variant, confirmation bias from stopping when significance was achieved, and correlation versus causation confusion.
- Selection bias from choosing among multiple variants, regression to the mean from early stopping, and misinterpretation of practical versus statistical significance.
- Cherry-picking the best result from multiple tests, publication bias from only reporting positive results, and external validity concerns from premature termination.
Explanation: This scenario demonstrates three distinct pitfalls: (1) Multiple comparisons - testing 8 variants inflates the family-wise error rate, (2) Data snooping - post-hoc selection of the 'best' variant for the final analysis compounds the multiple testing problem, and (3) Optional stopping - terminating the test upon reaching significance inflates Type I error because it doesn't account for the multiple interim analyses. Choice B incorrectly identifies survivorship bias and correlation/causation issues. Choice C mentions regression to the mean, which isn't the primary concern with optional stopping. Choice D mentions publication bias, which is about selective reporting across studies, not within a single analysis.
Question 5
A pharmaceutical researcher conducts 25 subgroup analyses on clinical trial data to identify which patient populations benefit most from a new drug. They find that patients aged 45-55 with BMI 25-30 show significant improvement (p=0.02), while the overall trial result was not significant (p=0.12). The researcher plans to seek FDA approval for this specific subgroup. Which statistical pitfall creates the strongest challenge to this strategy?
- The post-hoc nature of the subgroup analysis violates the intention-to-treat principle required for regulatory approval submissions.
- The overall non-significant result contradicts the subgroup finding, indicating a fundamental error in the statistical analysis methodology.
- The narrow age and BMI ranges represent selection bias that limits the external validity and generalizability of the findings.
- The multiple subgroup analyses created an inflated Type I error rate, making the significant result likely to be a false positive discovery. (correct answer)
Explanation: When you encounter questions about multiple hypothesis testing in clinical research, focus on how repeated statistical tests inflate the probability of false discoveries—this is a cornerstone concept in regulatory statistics.
The core issue here is multiple comparisons bias. When researchers conduct 25 subgroup analyses, each with a significance level of α=0.05, the probability of finding at least one false positive approaches 1−(0.95)25≈0.72, or 72%. This means there's a high likelihood that the significant result (p=0.02) is actually a false positive discovery rather than a true treatment effect. The FDA is acutely aware of this statistical pitfall and requires appropriate corrections for multiple testing.
Option A incorrectly conflates post-hoc analysis with intention-to-treat principles. While post-hoc subgroup analyses do raise concerns about data dredging, this doesn't violate intention-to-treat methodology, which relates to including all randomized participants regardless of compliance.
Option B misunderstands the relationship between overall and subgroup results. It's statistically possible—and not necessarily erroneous—for subgroups to show different effects than the overall population, especially if treatment effects vary across patient characteristics.
Option C addresses external validity concerns, which are important for clinical practice but don't represent the primary statistical challenge to regulatory approval. Narrow inclusion criteria are common in targeted therapies.
Remember: whenever you see multiple testing scenarios, immediately consider Type I error inflation. The more statistical tests performed, the higher the chance of spurious significant results—a critical concept for both exams and real-world research interpretation. Question 6
A pharmaceutical company tests 20 different drug compounds for effectiveness against a disease. They use a significance level of α=0.05 for each individual test and find that 2 compounds show statistically significant results (p<0.05). The research director concludes that they have identified 2 effective drugs. What is the primary statistical concern with this conclusion?
- The sample size is too small to detect meaningful differences between the drug compounds and requires larger clinical trials.
- The multiple comparisons problem inflates the family-wise error rate, making it likely that significant results occurred by chance alone. (correct answer)
- The significance level of 0.05 is too conservative for pharmaceutical research and should be increased to 0.10 for better sensitivity.
- The assumption of independence between drug compounds is violated since they may have similar chemical structures and mechanisms.
Explanation: This is a classic multiple comparisons problem. When conducting 20 tests at α = 0.05, the probability of finding at least one false positive is approximately 1 - (0.95)^20 ≈ 0.64, meaning there's about a 64% chance of finding spurious significant results. Finding 2 significant results out of 20 is actually close to what we'd expect by chance alone (20 × 0.05 = 1 expected false positive). Choice A focuses on sample size rather than the multiple testing issue. Choice C incorrectly suggests making the problem worse by using a higher α. Choice D mentions independence but misidentifies the core issue.
Question 7
In an exploratory data analysis project, a business analyst examines a dataset with 50 continuous variables. The analyst computes a correlation coefficient and a corresponding p-value for every possible pair of variables to identify any significant relationships. Using a significance level of α=0.05, approximately how many pairs of variables would be expected to show a statistically significant correlation purely by chance, even if no true relationships exist among any of the variables?
- 2.5
- 50
- 61 (correct answer)
- 1225
Explanation: This question requires calculating the number of comparisons and then the expected number of Type I errors. The number of unique pairs of variables from a set of 50 is given by the combination formula (250)=250×49=1225. The expected number of false positives (Type I errors) is the number of tests multiplied by the significance level: 1225×0.05=61.25. Therefore, one would expect to find about 61 'statistically significant' correlations due to random chance alone. Question 8
A marketing team A/B tests 20 different headlines for a new ad campaign, running a separate hypothesis test for each headline against the control. They find that one headline, 'Headline X,' yields a statistically significant increase in click-through rates with a p-value of 0.04. The team concludes that Headline X is genuinely superior. Which of the following statements presents the most significant statistical concern with this conclusion?
- The practical significance of the improvement may be too small to justify changing the headline.
- The reported p-value of 0.04 is not low enough to provide strong evidence against the null hypothesis.
- By conducting 20 tests, the probability of finding at least one significant result by random chance is substantially inflated. (correct answer)
- The study likely suffers from selection bias, as users who see the ads may not be representative of the target population.
Explanation: This scenario describes a classic multiple comparisons problem. When multiple hypothesis tests are conducted, the family-wise error rate (the probability of making at least one Type I error) increases. With 20 tests at an alpha of 0.05, the probability of getting at least one false positive is approximately 1−(0.95)20≈64. Therefore, the significant result is very likely to be spurious. Question 9
A study of several large companies reveals a strong, statistically significant positive correlation between the amount of money a company spends on its employee wellness programs and its annual profit margin. The CEO of a struggling company decides to immediately invest heavily in a new wellness program, stating that this will cause profits to rise. Why is the CEO's conclusion potentially flawed?
- The observed correlation does not prove causation; a third factor, like effective management or a positive company culture, may lead to both higher profits and more investment in wellness. (correct answer)
- The study may have used an incorrect statistical test to determine the correlation, thus invalidating the finding of significance.
- The CEO is confusing statistical significance with practical significance, as the cost of the program may exceed the profit increase.
- The relationship might be non-linear, meaning that beyond a certain point, increased wellness spending could lead to diminishing returns on profit.
Explanation: This is a classic case of confusing correlation with causation. While two variables may be strongly correlated, it does not mean one causes the other. An unobserved confounding variable is a common explanation. In this case, well-managed and profitable companies might simply have more resources and a greater inclination to invest in employee perks like wellness programs. The wellness program itself might not be the direct cause of the profitability.
Question 10
A university is investigating potential gender bias in its graduate school admissions. The overall data shows that 45% of male applicants are admitted, while only 35% of female applicants are admitted. However, when the data is broken down by the two largest colleges (the highly competitive College of Engineering and the less competitive College of Humanities), analysts find that within both colleges, a higher percentage of female applicants are admitted than male applicants.
Which statistical phenomenon is most likely responsible for this apparent contradiction?
- Ecological Fallacy
- Multiple Comparisons Problem
- Simpson's Paradox (correct answer)
- Data Snooping
Explanation: Simpson's Paradox occurs when a trend or relationship that appears in aggregate data reverses or disappears when the data is broken down into subgroups. In this case, the overall admission rate appears to favor males, but within each college, the rate favors females. This can happen if a much larger proportion of female applicants apply to the more competitive college (Engineering), where admission rates are low for everyone, which in turn drags down their overall admission rate.
Question 11
A national health study finds that states with a higher per-capita consumption of broccoli have lower rates of a certain type of cancer. Based on this aggregate data, a public health official launches a campaign encouraging all individuals to eat more broccoli to reduce their personal cancer risk. This conclusion is most at risk of committing which statistical error?
- The Base Rate Fallacy, by ignoring the low underlying rate of cancer in the population.
- Simpson's Paradox, as the trend might reverse if data were grouped by age or income.
- The Ecological Fallacy, by assuming a state-level correlation applies to individuals. (correct answer)
- The Multiple Comparisons Problem, by looking at too many types of vegetables and cancers.
Explanation: The Ecological Fallacy is an error in reasoning where an inference about an individual is deduced from aggregate data for the group to which that individual belongs. Just because states with higher broccoli consumption have lower cancer rates does not mean that it's the broccoli-eating individuals within those states who are avoiding cancer. It could be that states with higher broccoli consumption also have better healthcare, less pollution, or higher average incomes, and these are the real factors affecting individuals' health.
Question 12
A consulting firm tests 10 different management training programs to see if they improve team productivity. They run 10 separate hypothesis tests, one for each program against a control group, using a significance level of α=0.05 for each test. Assuming that, in reality, none of the programs have any effect on productivity, what is the approximate probability that the firm will incorrectly conclude that at least one program is effective?
- 5%
- 10%
- 40% (correct answer)
- 50%
Explanation: This is a multiple comparisons problem. The probability of correctly not finding a significant result (not making a Type I error) in one test is 1−0.05=0.95. If the tests are independent, the probability of not finding a significant result in any of the 10 tests is (0.95)10≈0.599. Therefore, the probability of finding at least one significant result just by chance (the family-wise error rate) is 1−0.599=0.401, or approximately 40%. Question 13
A company has a screening test for a rare manufacturing defect. The test is 99% accurate: it has a 99% probability of giving a positive result if the defect is present (sensitivity) and a 99% probability of giving a negative result if the defect is absent (specificity). The actual defect occurs in only 1 out of every 1,000 units (a base rate of 0.1%). If a randomly selected unit tests positive, what is the most likely conclusion?
- The unit almost certainly has the defect, since the test is 99% accurate.
- There is approximately a 50% chance that the unit has the defect.
- It is more likely that the positive test is an error than that the unit is actually defective. (correct answer)
- The accuracy of the test cannot be determined without knowing the total number of units produced.
Explanation: This problem illustrates the base rate fallacy. Even with a highly accurate test, a low base rate of the condition can lead to a high number of false positives. Consider 100,000 units. About 100 will have the defect, and the test will correctly identify 99 of them. About 99,900 will not have the defect. The test will incorrectly identify 1% of these as positive (false positives), which is 0.01×99,900≈999. So, out of 99+999=1098 positive tests, only 99 are true positives. The probability of a defect given a positive test is 99/1098≈9. Therefore, a positive result is far more likely to be a false alarm. Question 14
An analyst presents a simple linear regression model: Monthly Sales ()=40,000+3.2×TVAdBudget(). The model has a p-value of 0.01 for the coefficient of the ad budget. The analyst concludes, 'The model proves that every dollar we add to the TV ad budget will cause our sales to increase by $3.20.' What is the most significant misinterpretation in this conclusion?
- The model only describes a historical correlation; it does not prove that ad spending causes the increase in sales. (correct answer)
- The intercept of 40,000 is too high, suggesting the model is biased and therefore unreliable for forecasting.
- The conclusion is invalid because a simple linear regression is not sufficient to model a complex variable like sales.
- The analyst should have used a log transformation on the variables to correctly interpret the effect as a percentage change.
Explanation: The primary error is the leap from correlation to causation. A regression model, no matter how statistically significant, can only demonstrate the strength and direction of an association between variables in the data. It cannot, by itself, prove a causal link. It's possible that a third factor (e.g., seasonality, economic growth) drives both advertising spending and sales, or that the causal relationship is reversed (higher sales lead to bigger ad budgets). To claim causation, one would need a carefully designed experiment.
Question 15
A company pilots a new software tool to improve productivity. The initial study on 200 employees shows no statistically significant overall improvement. However, the researchers then analyze results for various subgroups and find that productivity increased significantly for employees in the R&D department who have been with the company for more than 5 years. In their press release, the company exclusively highlights this positive subgroup finding. This reporting practice is best characterized as:
- An effective market segmentation strategy to identify the ideal customer profile.
- An example of the ecological fallacy, where group results are applied to individuals.
- Data snooping and selective reporting ('cherry-picking') to present a misleadingly positive result. (correct answer)
- A robust analysis that correctly identifies an important interaction effect between department and tenure.
Explanation: The practice of conducting an analysis, failing to find a desired result, and then searching through various subgroups until a statistically significant result is found is known as data snooping or p-hacking. Selectively reporting only the favorable subgroup result while omitting the non-significant overall finding is a form of 'cherry-picking.' This is misleading because the subgroup finding is likely a result of chance due to the multiple comparisons problem and is not a reliable indicator of the software's true effect.
Question 16
Two A/B tests were conducted to improve website conversion rates:
- Test 1: Tested a minor headline change on 2,000,000 users. It resulted in a 0.2% lift in conversion, with a p-value of 0.005.
- Test 2: Tested a major website redesign on 2,000 users. It resulted in an 8% lift in conversion, with a p-value of 0.06.
A manager argues that the headline change from Test 1 is the more successful experiment and should be prioritized because it achieved 'high statistical significance.'
What is the primary flaw in the manager's reasoning?
- The manager is ignoring the much larger and more practically significant effect size of the website redesign in Test 2. (correct answer)
- The p-value for Test 2 is low enough to be considered significant, so both experiments are equally successful.
- The extremely large sample size in Test 1 makes its p-value artificially low and therefore untrustworthy.
- It is impossible to compare the two tests because their sample sizes and the nature of the changes are different.
Explanation: The manager is falling into the common trap of equating statistical significance with practical importance. Test 1's low p-value is a direct result of its enormous sample size, which gives it high power to detect even a minuscule effect (0.2% lift). Test 2, despite its higher p-value (due to a much smaller sample size), shows a dramatically larger effect size (8% lift). From a business perspective, the 8% lift is almost certainly more valuable, even if the statistical evidence for it is slightly less certain. The manager's focus on p-values alone is a misinterpretation of the results.
Question 17
A medical researcher is testing a new drug. The protocol calls for a study with 200 patients. However, the researcher decides to analyze the data and check the p-value after every 20 patients are enrolled. The researcher plans to stop the trial and declare success as soon as the p-value drops below 0.05. This 'optional stopping' strategy has what effect on the research findings?
- It increases the statistical power of the study, making it more likely to correctly identify a truly effective drug.
- It greatly inflates the Type I error rate, making it highly likely that a useless drug will be declared effective. (correct answer)
- It has no significant effect on the error rates if the final sample size is the same as the one specified in the original protocol.
- It leads to a biased, underestimated effect size because the trial is stopped before the full effect can be observed.
Explanation: This practice, known as p-hacking or optional stopping, severely inflates the Type I error rate. By repeatedly testing the data as it accumulates, the researcher gives random chance multiple opportunities to produce a 'significant' result. The true Type I error rate can be much higher than the nominal 5% level. A study conducted this way that finds a significant result is very likely to be a false positive.
Question 18
An auto insurance company's risk model finds a strong negative correlation between a driver having a 'classic car' license plate and their likelihood of filing a claim. The company proposes offering a large premium discount for drivers with classic car plates, believing this will attract lower-risk drivers and reduce payouts. What is the most significant flaw in this strategic reasoning?
- The model proves correlation, not causation; the license plate is a proxy for careful, conscientious owners, not the cause of their low risk. (correct answer)
- Offering the discount will lead to adverse selection, as only high-risk drivers will apply for it to lower their costs.
- The sample size of drivers with classic car plates is too small to build a reliable predictive model.
- The model should have included other variables like driver age and vehicle type to be considered valid.
Explanation: The company is mistaking a predictor variable (a proxy) for a causal factor. People who own, maintain, and register a classic car are likely to be hobbyists who drive carefully and infrequently. The license plate is an indicator of this underlying behavior; it does not cause the behavior. Simply getting a classic car plate (if that were possible without owning such a car) would not make a driver safer. The strategy might work to attract the right kind of customer, but the reasoning that the plate itself is linked to risk is flawed; it's a proxy for the driver's characteristics.
Question 19
A data analyst examines customer purchase data and notices that sales of ice cream and sunglasses are highly correlated (r=0.87). After finding this correlation, the analyst searches through the same dataset and discovers that both products also correlate strongly with temperature, outdoor events, and vacation bookings. The analyst then reports these findings as evidence of multiple interconnected market relationships. What statistical pitfall is most evident in this scenario?
- Correlation versus causation confusion, since the analyst assumes that ice cream sales directly influence sunglasses purchases without experimental evidence.
- Data snooping bias, since the analyst conducted post-hoc searches for additional correlations within the same dataset after finding the initial result. (correct answer)
- Survivorship bias, since the analysis only includes customers who made purchases and ignores those who visited but didn't buy anything.
- Confirmation bias, since the analyst only looked for positive correlations and ignored potential negative relationships between seasonal products.
Explanation: This scenario demonstrates data snooping (also called data mining or p-hacking). The analyst first found one correlation, then searched through the same dataset to find additional supporting correlations. This inflates the chance of finding spurious relationships because the analyst is conducting multiple unplanned comparisons on the same data. While choice A identifies a real issue with correlation vs. causation, the primary pitfall here is the post-hoc searching behavior. Choice C incorrectly identifies survivorship bias, which isn't relevant here. Choice D mischaracterizes the scenario - there's no evidence the analyst only sought positive correlations.
Question 20
An analyst at a beverage company examines hourly sales data across hundreds of products. He notices that sales for a specific brand of iced tea seem unusually high between 3:00 PM and 4:00 PM. He then formulates the hypothesis that 'the 3:00 PM hour is a peak sales period for this tea' and performs a t-test comparing sales in that hour to the daily average. The test yields a p-value of 0.02. What is the most critical flaw in the analyst's methodology?
- A t-test is not the correct statistical tool for analyzing time-series sales data; an ANOVA should have been used.
- The analyst generated the hypothesis after observing the pattern in the data, which invalidates the assumptions of the hypothesis test. (correct answer)
- The result lacks practical significance because a one-hour sales spike is unlikely to impact overall revenue.
- He failed to control for confounding variables, such as weather, which could be the true cause of the sales increase.
Explanation: This is an example of data snooping where a hypothesis is formulated after seeing a pattern in the data (post-hoc hypothesis generation). Standard hypothesis testing requires that the hypothesis be stated before data is examined. By selecting a pattern and then testing it with the same data that suggested it, the analyst has not conducted a valid confirmatory analysis. The low p-value is likely a result of capitalizing on chance, and the finding may not be reproducible with new data.