Biostatistics Quiz: P Values And Significance Levels
20 questions · exam conditions
0:00
P Values And Significance LevelsQuestion 1 of 20

A researcher conducts a two-tailed hypothesis test with α=0.05\alpha = 0.05 and obtains a p-value of 0.03. If the same data were analyzed using a one-tailed test in the direction of the observed effect, what can be concluded about the relationship between the new p-value and the significance level?

The one-tailed p-value would be 0.015, leading to rejection of the null hypothesis at α=0.05\alpha = 0.05
The one-tailed p-value would be 0.06, leading to failure to reject the null hypothesis at α=0.05\alpha = 0.05
The one-tailed p-value would be 0.015, but the conclusion would remain unchanged from the two-tailed test
The one-tailed p-value would be 0.03, leading to the same conclusion as the two-tailed test at α=0.05\alpha = 0.05
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: P Values And Significance Levels

Practice P Values And Significance Levels in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on P Values And Significance Levels, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher conducts a two-tailed hypothesis test with α=0.05\alpha = 0.05 and obtains a p-value of 0.03. If the same data were analyzed using a one-tailed test in the direction of the observed effect, what can be concluded about the relationship between the new p-value and the significance level?

  1. The one-tailed p-value would be 0.015, leading to rejection of the null hypothesis at α=0.05\alpha = 0.05 (correct answer)
  2. The one-tailed p-value would be 0.06, leading to failure to reject the null hypothesis at α=0.05\alpha = 0.05
  3. The one-tailed p-value would be 0.015, but the conclusion would remain unchanged from the two-tailed test
  4. The one-tailed p-value would be 0.03, leading to the same conclusion as the two-tailed test at α=0.05\alpha = 0.05
Explanation: For a two-tailed test with p = 0.03, the one-tailed p-value in the direction of the effect would be 0.03/2 = 0.015. Since 0.015 < 0.05, we reject the null hypothesis. Choice B incorrectly doubles instead of halving. Choice C correctly calculates the p-value but incorrectly states the conclusion would be unchanged (both reject, but for different reasons). Choice D incorrectly assumes the p-value stays the same.

Question 2

In a clinical trial, three independent tests are performed on the same dataset: Test A yields p = 0.02, Test B yields p = 0.04, and Test C yields p = 0.03. If no correction for multiple testing is applied and each test uses α=0.05\alpha = 0.05, what is the probability of making at least one Type I error across all three tests, assuming all null hypotheses are true?

  1. 0.05, because each individual test maintains the specified significance level
  2. 0.09, representing the sum of the three observed p-values
  3. 0.142625, calculated as 1(10.05)31 - (1 - 0.05)^3 (correct answer)
  4. 0.15, representing the cumulative significance level across three tests
Explanation: The probability of at least one Type I error is 1 - P(no Type I errors) = 1 - (0.95)³ = 1 - 0.857375 = 0.142625. The observed p-values are irrelevant when all null hypotheses are true. Choice A ignores multiple testing issues. Choice B incorrectly sums p-values. Choice D incorrectly multiplies the significance level by the number of tests.

Question 3

An investigator sets α=0.01\alpha = 0.01 for a study and obtains a p-value of 0.025. A colleague argues that if α\alpha had been set to 0.05, the results would be more convincing. What is the most appropriate response?

  1. The colleague is correct because a larger α\alpha makes the results more statistically significant
  2. The colleague is incorrect because the p-value provides the same evidence regardless of the chosen α\alpha level (correct answer)
  3. The colleague is correct because the results would change from non-significant to significant
  4. The colleague is incorrect because changing α\alpha after seeing the data invalidates the statistical inference
  5. The colleague is partially correct because α=0.05\alpha = 0.05 is the standard level used in most research
Explanation: When you encounter questions about statistical significance and p-values, remember that the p-value represents the strength of evidence against the null hypothesis, while α\alpha is simply the threshold you choose for decision-making. The p-value of 0.025 provides exactly the same evidence regardless of whether you compare it to α=0.01\alpha = 0.01 or α=0.05\alpha = 0.05. This value tells you there's a 2.5% probability of observing results this extreme or more extreme if the null hypothesis were true. That probability doesn't change based on your chosen significance level. Option A is incorrect because a larger α\alpha doesn't make results "more statistically significant" – it only makes them more likely to be declared significant by lowering the bar for rejection. The actual strength of evidence (the p-value) remains unchanged. Option C commits a common misconception by focusing on the significant/non-significant dichotomy rather than the underlying evidence. While the classification would change from non-significant to significant, this doesn't make the results more convincing. Option D raises an important point about post-hoc α\alpha adjustment, but the colleague's argument isn't necessarily about changing α\alpha after seeing data – they're discussing what would have happened with different initial criteria. The key insight is that statistical significance is a binary classification system, while the p-value provides continuous information about evidence strength. When evaluating research, focus on the p-value itself rather than arbitrary significance thresholds. A result with p = 0.025 carries the same evidential weight regardless of your α\alpha choice.

Question 4

A pharmaceutical company tests a new drug using α=0.05\alpha = 0.05 and obtains p = 0.048. The company claims this proves their drug is effective. A regulatory agency argues for using α=0.01\alpha = 0.01 instead. What is the key issue with the company's claim?

  1. The p-value is too close to the significance level to draw any meaningful conclusions
  2. Statistical significance does not prove effectiveness, and the choice of α\alpha should be predetermined (correct answer)
  3. The regulatory agency is correct because α=0.01\alpha = 0.01 is more stringent and reliable
  4. The company's claim is valid because p < 0.05 is the accepted standard in pharmaceutical research
  5. The p-value indicates only a 4.8% chance that the drug is effective, which is too low
Explanation: This question tests your understanding of hypothesis testing fundamentals and the distinction between statistical significance and practical significance. When you encounter questions about p-values and significance levels, focus on the proper interpretation and methodology rather than just the numerical comparison. The key issue here involves two critical principles: the predetermined nature of significance levels and the meaning of statistical significance. The significance level (α\alpha) must be established before data collection and analysis to avoid bias. When researchers choose α\alpha after seeing results, they're essentially moving the goalposts to favor their desired outcome. Additionally, statistical significance only indicates that an observed difference is unlikely due to chance alone—it doesn't prove that a treatment is clinically meaningful, safe, or "effective" in practical terms. Choice A is incorrect because while p = 0.048 is close to α=0.05\alpha = 0.05, proximity alone doesn't invalidate the statistical test—the real issues are methodological and interpretive. Choice C misses the main point by focusing on which α\alpha level is "better" rather than addressing the fundamental flaws in the company's approach. Choice D incorrectly assumes that meeting a conventional significance threshold automatically validates effectiveness claims, ignoring both the predetermined α\alpha requirement and the limitation of statistical significance. Remember this key principle: statistical significance ≠ clinical significance or proof of effectiveness. Always establish your significance level before collecting data, and interpret results within the broader context of clinical relevance, effect size, and practical importance.

Question 5

Two studies test the same research hypothesis. Study A (n = 100) reports p = 0.03, and Study B (n = 1000) reports p = 0.04. Both used α=0.05\alpha = 0.05. What is the most accurate interpretation?

  1. Study A provides stronger evidence because it has a smaller p-value despite the smaller sample size
  2. Study B provides stronger evidence because the larger sample size makes the result more reliable
  3. Both studies provide equivalent evidence since both p-values are less than 0.05 and statistically significant
  4. The studies cannot be compared because they have different sample sizes and different p-values
  5. Study A provides stronger evidence against the null hypothesis based solely on the p-value comparison (correct answer)
Explanation: When evaluating statistical evidence across studies, you need to consider both statistical significance and effect size magnitude, not just p-values alone. The p-value tells you the probability of observing your results if the null hypothesis were true, but it's heavily influenced by sample size. Study B provides stronger evidence despite the slightly larger p-value. With n=1000 versus Study A's n=100, Study B detected a significant effect with much greater precision. Large samples make it harder to achieve statistical significance by chance alone, so when you do find significance in a large study, it represents more robust evidence. The effect size in Study B is likely more reliable and generalizable. Choice A falls into the common trap of p-value worship—assuming smaller p-values always mean stronger evidence. This ignores how sample size affects p-values. Choice C incorrectly treats all statistically significant results as equivalent evidence, missing the crucial role of sample size in determining reliability. Choice D suggests the studies can't be compared, but comparing results across different sample sizes is exactly what meta-analyses and evidence synthesis do routinely. The key insight is that statistical significance is just the first hurdle. Once you've cleared α=0.05\alpha = 0.05, the strength of evidence depends on factors like sample size, effect size confidence intervals, and study design quality. A barely significant result in a tiny study often represents weaker evidence than a clearly significant result in a large, well-powered study. Remember: bigger samples generally provide more trustworthy evidence when both studies achieve statistical significance.

Question 6

A researcher obtains a p-value of 0.12 in a study with 50 participants. She argues that with a larger sample size, the p-value would definitely become significant at α=0.05\alpha = 0.05. What is wrong with this reasoning?

  1. The p-value would actually increase with a larger sample size, making significance less likely
  2. The effect size might be too small to achieve significance regardless of sample size increases (correct answer)
  3. P-values are independent of sample size, so increasing n would not change the result
  4. The reasoning assumes the effect is real, but p = 0.12 suggests the null hypothesis is likely true
  5. The significance level should be adjusted downward when sample size increases to maintain validity
Explanation: When you encounter questions about p-values and sample size, remember that statistical significance depends on both effect size and sample size working together. The relationship isn't automatic. The researcher's reasoning contains a critical flaw: she assumes that simply increasing sample size will guarantee significance, but this ignores whether there's actually a meaningful effect to detect. While larger samples do increase statistical power (the ability to detect true effects), they can only amplify effects that actually exist. If the true effect size in the population is very small or zero, even massive sample increases won't reliably push the p-value below 0.05. You could end up with p-values hovering around the significance threshold regardless of sample size. Looking at the wrong answers: (A) incorrectly suggests p-values increase with larger samples - actually, if there's a real effect, larger samples typically decrease p-values by providing more precise estimates. (C) is fundamentally wrong because p-values are definitely influenced by sample size through the standard error calculations. (D) misinterprets what p = 0.12 means - this p-value doesn't suggest the null hypothesis is "likely true," it simply means the evidence isn't strong enough to reject it at the 0.05 level. Study tip: Remember that statistical significance requires both adequate sample size AND meaningful effect size. When you see questions about "guaranteed significance with larger samples," always consider whether the effect itself might be too small to matter, regardless of sample size increases.

Question 7

In a clinical trial, the primary endpoint shows p = 0.06, while a secondary endpoint shows p = 0.02. Both tests used α=0.05\alpha = 0.05. The researchers conclude that the treatment is effective based on the secondary endpoint. What is the main concern with this approach?

  1. The secondary endpoint should use a more stringent α\alpha level to account for multiple testing
  2. The primary endpoint result contradicts the secondary endpoint, making interpretation impossible
  3. Secondary endpoints are less reliable than primary endpoints and should not be used for efficacy claims
  4. The researchers are cherry-picking significant results while ignoring the primary outcome measure (correct answer)
  5. Both p-values should be averaged to determine the overall treatment effect significance
Explanation: When you encounter clinical trial results with multiple endpoints, the key principle is that conclusions should be based on pre-specified primary outcomes, not on hunting for favorable results among secondary measures. The correct answer is D because this scenario represents a classic case of cherry-picking or "p-hacking." The researchers designed their study with a specific primary endpoint to test their main hypothesis, but when that endpoint failed to reach significance (p = 0.06), they shifted focus to a secondary endpoint that happened to be significant (p = 0.02). This invalidates the study's inferential framework because they're essentially changing their hypothesis after seeing the data. Let's examine why the other options miss the mark. Option A suggests using stricter alpha levels for secondary endpoints, but the real issue isn't the alpha level—it's that secondary endpoints weren't meant to be the basis for primary efficacy claims. Option B incorrectly states that interpretation is impossible; while the results are mixed, the primary endpoint should take precedence in interpretation. Option C makes an overly broad claim that secondary endpoints are inherently less reliable, when actually they can be quite valid for their intended exploratory purposes. The fundamental problem is prioritizing convenience over scientific rigor. Secondary endpoints are valuable for generating hypotheses and understanding mechanisms, but using them to rescue a failed primary analysis undermines the entire statistical framework of hypothesis testing. Study tip: Remember that in clinical trials, the primary endpoint is king. When you see researchers celebrating secondary results while downplaying primary failures, that's a major red flag for biased interpretation.

Question 8

A researcher conducts 20 independent hypothesis tests, each with α=0.05\alpha = 0.05, and finds that exactly one test yields p < 0.05. She concludes this represents a significant finding. What is the expected number of false positives in this scenario?

  1. Exactly 1, confirming that the significant result is likely a false positive discovery
  2. Approximately 1, suggesting the significant result might be due to chance alone (correct answer)
  3. Zero, because the tests are independent and each has equal probability of significance
  4. Cannot be determined without knowing the true effect sizes in each test
  5. Exactly 0.05, representing the probability of Type I error for the significant test
Explanation: When you encounter multiple testing scenarios, you need to understand the concept of expected false positives due to Type I error accumulation across independent tests. With 20 independent tests each using α=0.05\alpha = 0.05, the expected number of false positives is simply 20×0.05=120 \times 0.05 = 1. This means that even when all null hypotheses are true, you'd expect about one test to yield p < 0.05 purely by chance. Since the researcher found exactly one significant result, this aligns perfectly with what we'd expect from random variation alone. Answer A is incorrect because while the expected number is indeed 1, this doesn't confirm that the significant result is "likely" a false positive—it's possible but not certain. The wording overstates the conclusion we can draw. Answer B correctly identifies that approximately 1 false positive is expected and appropriately suggests (rather than definitively concludes) that the significant result might be due to chance alone. This measured interpretation is statistically sound. Answer C is wrong because independence doesn't eliminate false positives—it actually makes them predictable. Each test has a 5% chance of false significance regardless of the others, so false positives are expected, not prevented. Answer D is incorrect because the expected number of false positives depends only on α\alpha and the number of tests, not on true effect sizes. Even with unknown effect sizes, we can calculate the expected false positive rate. Remember: in multiple testing, always calculate expected false positives as (number of tests × α\alpha) to put your significant findings in proper context.

Question 9

A study reports: 'The treatment showed a statistically significant improvement (p = 0.045) with a 2% increase in success rate.' A critic argues this result, while statistically significant, may not be practically meaningful. What does this illustrate?

  1. The critic is incorrect because statistical significance always implies practical importance in medical research
  2. The p-value is too close to 0.05 to be considered reliable evidence of any effect
  3. Statistical significance and practical significance are distinct concepts that should both be considered (correct answer)
  4. The sample size was likely too large, making trivial differences appear statistically significant
  5. The 2% improvement contradicts the significant p-value, indicating an error in the analysis
Explanation: This question tests your understanding of a critical distinction in biostatistics: the difference between statistical significance and practical (clinical) significance. When you encounter results that are statistically significant but involve small effect sizes, you should always consider whether the finding is meaningful in real-world terms. The correct answer is C because statistical significance and practical significance are indeed separate concepts that researchers must evaluate independently. A p-value of 0.045 indicates that the observed 2% improvement is unlikely due to chance alone, satisfying the criterion for statistical significance. However, whether a 2% increase in success rate represents a clinically meaningful improvement depends on the specific medical context, costs, risks, and patient outcomes involved. Option A is wrong because statistical significance never automatically guarantees practical importance. Even tiny, clinically irrelevant differences can achieve statistical significance with large enough sample sizes. Option B incorrectly suggests that p-values near 0.05 are unreliable—while p-hacking is a concern, a p-value of 0.045 is still below the conventional 0.05 threshold. Option D makes an assumption about sample size that, while plausible, isn't definitively supported by the information given and misses the main point about the distinction between types of significance. Remember this key principle: always ask two questions when evaluating research results—"Is this difference statistically significant?" and "Is this difference large enough to matter in practice?" Both questions are essential for proper interpretation of biostatistical findings.

Question 10

A researcher reports p = 0.08 and states: 'While not statistically significant at α=0.05\alpha = 0.05, this p-value indicates a trend toward significance.' What is the most appropriate evaluation of this interpretation?

  1. The interpretation is correct because p-values between 0.05 and 0.10 represent meaningful trends
  2. The interpretation is misleading because it treats p-values as if they have inherent meaning beyond the chosen α\alpha (correct answer)
  3. The interpretation is acceptable because it acknowledges the limitation while noting the proximity to significance
  4. The interpretation is incorrect because any p-value above 0.05 indicates no effect whatsoever
  5. The interpretation depends on whether this was a one-tailed or two-tailed test design
Explanation: When you encounter questions about p-value interpretation, focus on understanding what p-values actually measure versus how researchers sometimes misinterpret them. P-values tell us the probability of observing our data (or more extreme) if the null hypothesis were true, but they don't have inherent meaning beyond our predetermined significance threshold. The correct interpretation recognizes that p = 0.08 is simply a non-significant result at α=0.05\alpha = 0.05. There's no statistical basis for calling this a "trend toward significance" – this phrase treats p-values as if they exist on a meaningful continuum where values close to 0.05 have special importance. Option B correctly identifies this as misleading because it suggests the p-value has inherent meaning beyond our chosen alpha level. Looking at the wrong answers: Option A incorrectly validates the "trend" concept by suggesting p-values between 0.05 and 0.10 represent meaningful trends – this is a common misconception with no statistical foundation. Option C seems reasonable but actually perpetuates the problematic thinking by accepting that "proximity to significance" is meaningful. Option D goes too far in the opposite direction, incorrectly stating that p > 0.05 means "no effect whatsoever" – we simply fail to reject the null hypothesis, which doesn't prove no effect exists. Study tip: Remember that statistical significance is binary at your chosen alpha level – results are either significant or not. Phrases like "trending toward significance," "marginally significant," or "approaching significance" are red flags indicating misinterpretation of p-values.

Question 11

Two identical studies of the same intervention are conducted. Study 1 reports p = 0.049, Study 2 reports p = 0.051. Using α=0.05\alpha = 0.05, how should these results be interpreted?

  1. Study 1 proves the intervention works, while Study 2 proves it doesn't work, showing contradictory evidence
  2. The results are essentially equivalent, and the difference in conclusions highlights the arbitrariness of significance thresholds (correct answer)
  3. Study 1 provides stronger evidence because its p-value is below the significance threshold
  4. Study 2 is more reliable because p-values just above 0.05 are less likely to be false positives
  5. The studies should be combined using meta-analysis since individual results are inconclusive
Explanation: When you encounter questions about p-values that fall just above or below significance thresholds, you're dealing with one of the most important conceptual issues in statistical interpretation: the arbitrary nature of significance cutoffs. The correct answer is B because these two studies show nearly identical evidence strength. A p-value of 0.049 versus 0.051 represents virtually no meaningful difference in the strength of evidence against the null hypothesis. The 0.05 threshold is a human convention, not a natural boundary where evidence quality suddenly changes. Both studies suggest similar, modest evidence against the null hypothesis. Option A is wrong because neither study "proves" anything—statistical significance doesn't equal proof, and the tiny difference in p-values doesn't justify claiming contradictory evidence. Option C misunderstands what "stronger evidence" means; while Study 1 technically meets the significance criterion, the difference of 0.002 in p-values is practically meaningless. Option D incorrectly suggests that p-values just above 0.05 have some special reliability property regarding false positives—this isn't true and misrepresents how Type I error rates work. The key insight here is that statistical significance is a binary decision rule applied to continuous evidence. When p-values hover around your α\alpha level, focus on the practical similarity of the evidence rather than the artificial significance/non-significance distinction. Remember: p-values measure evidence strength on a continuum, but significance testing forces an artificial yes/no decision that can be misleading when results fall near the threshold.

Question 12

A researcher conducts a study with n=25n = 25 and obtains p = 0.30. She argues that the lack of significance is due to insufficient power and plans to increase the sample size to n=100n = 100. What assumption is she making?

  1. That the effect size will increase proportionally with the larger sample size
  2. That there is a true effect to detect, which may or may not be a valid assumption (correct answer)
  3. That the significance level will automatically adjust to accommodate the larger sample size
  4. That increasing sample size will reduce the Type I error rate in future studies
  5. That the p-value of 0.30 indicates the power was exactly 70% in the original study
Explanation: When you encounter questions about statistical power and sample size decisions, focus on the underlying assumptions researchers make about effect existence and detectability. The researcher's logic reveals a critical assumption: she believes a true effect exists in the population that her initial study failed to detect due to insufficient power. By increasing sample size from 25 to 100, she's operating under the premise that with greater power, she'll be able to detect this presumed true effect. This assumption may or may not be valid—the non-significant result could reflect either insufficient power to detect a real effect or the genuine absence of an effect. Let's examine why the other options miss the mark. Choice A incorrectly suggests effect size grows with sample size, but effect size is a population parameter that remains constant regardless of your sample size—only your ability to detect it changes. Choice C misunderstands significance levels, which researchers set deliberately (typically at 0.05) and don't automatically adjust based on sample size. Choice D confuses the relationship between sample size and error types—larger samples increase power (reduce Type II error) but don't change the Type I error rate, which is determined by your chosen significance level. The correct answer is B because it captures the fundamental assumption underlying the researcher's decision: that there's a true effect worth detecting. Study tip: Remember that increasing sample size only helps if there's actually an effect to find. When you see power analysis questions, always consider whether the researcher is assuming effect existence versus effect absence.

Question 13

A clinical trial uses α=0.025\alpha = 0.025 for the primary analysis to account for an interim analysis. The final result shows p = 0.04. The researchers state they cannot conclude efficacy. A reviewer argues they should use α=0.05\alpha = 0.05 since this is the final analysis. Who is correct?

  1. The reviewer is correct because interim analyses shouldn't affect the final analysis interpretation
  2. The researchers are correct because the α level must account for all planned analyses to control Type I error (correct answer)
  3. Both are wrong because the p-value should be adjusted rather than the significance level
  4. The reviewer is correct because 0.025 is too conservative for clinical trial research
  5. Neither can be determined correct without knowing the results of the interim analysis
Explanation: When you encounter questions about multiple comparisons in clinical trials, think about the fundamental principle of Type I error control. The alpha level must be set to maintain the overall probability of falsely rejecting the null hypothesis across all planned analyses, not just individual tests. The researchers are correct because they properly applied alpha spending to control the familywise error rate. When conducting an interim analysis followed by a final analysis, you're essentially performing multiple hypothesis tests on the same data. Without adjustment, each test at α=0.05\alpha = 0.05 would give you a cumulative Type I error rate exceeding 5%. By using α=0.025\alpha = 0.025 for the final analysis, they allocated part of their "alpha budget" to the interim look while maintaining overall Type I error control at 0.05. Option A is wrong because interim analyses absolutely affect interpretation—they create a multiple comparisons problem that demands alpha adjustment. Option C misunderstands the approach: adjusting the significance level (alpha spending) is the standard, appropriate method rather than post-hoc p-value adjustment. Option D incorrectly suggests that 0.025 is arbitrarily conservative, when it's actually a calculated adjustment based on the study's sequential design. The p-value of 0.04 exceeds their adjusted threshold of 0.025, so claiming efficacy would inflate Type I error risk beyond acceptable levels. Study tip: Remember that in sequential trial designs, alpha spending is like a budget—you must allocate it across all planned analyses. The total can't exceed your target Type I error rate, regardless of how many looks you take at the data.

Question 14

A researcher plans to test 10 related hypotheses and wants to maintain an overall Type I error rate of 0.05. Using the Bonferroni correction, she sets each individual test at α=0.005\alpha = 0.005. One test yields p = 0.02. What should she conclude?

  1. The result is significant because p = 0.02 < 0.05, regardless of the Bonferroni correction
  2. The result is not significant because p = 0.02 > 0.005, and the correction must be maintained (correct answer)
  3. The result is marginally significant and should be interpreted as a trend requiring replication
  4. The Bonferroni correction is too conservative and should be replaced with a less stringent method
  5. The result significance depends on whether the other 9 tests were also significant
Explanation: When you encounter multiple hypothesis testing questions, focus on understanding why we need correction methods and how to apply the chosen correction consistently. The Bonferroni correction addresses the multiple comparisons problem: when testing many hypotheses simultaneously, your chance of finding at least one false positive increases dramatically. With 10 tests at α=0.05\alpha = 0.05 each, you'd have roughly a 40% chance of making a Type I error somewhere. The Bonferroni method controls this by dividing your desired overall error rate by the number of tests: αadjusted=0.05/10=0.005\alpha_{adjusted} = 0.05/10 = 0.005. Once you've chosen this correction, you must apply it consistently. Since p=0.02>0.005p = 0.02 > 0.005, this result is not statistically significant under the corrected threshold, making option B correct. Option A ignores the multiple testing problem entirely—you can't selectively apply corrections based on convenience. This defeats the purpose of controlling family-wise error rate. Option C introduces the problematic concept of "marginal significance," which isn't a valid statistical category and encourages p-hacking. While the Bonferroni correction is indeed conservative, option D doesn't address what the researcher should conclude about this specific test result. Key strategy: When researchers choose a multiple comparison correction method, they must stick with it for all tests in that family. You can't cherry-pick which correction to apply based on whether results meet your preferred threshold. Always compare the p-value to the adjusted alpha level, not the original one.

Question 15

A researcher obtains p = 0.15 in a pilot study with n = 20 and calculates that n = 80 would provide 80% power to detect the observed effect size. She conducts the larger study and obtains p = 0.06. How should this sequence of results be interpreted?

  1. The pilot study was underpowered, and the larger study confirms the effect exists
  2. The larger study contradicts the pilot study since one is significant and the other is not
  3. Both studies suggest weak evidence, and the effect may not be practically important (correct answer)
  4. The pilot study was misleading, and only the larger study should be considered reliable
  5. The results are consistent with the power analysis predictions and support the presence of an effect
Explanation: When interpreting a sequence of studies, you need to look beyond simple significance thresholds and consider the broader pattern of evidence. Both p-values here (0.15 and 0.06) fall in the "gray zone" where results are neither clearly significant nor clearly null. The correct interpretation is C because both studies actually tell a consistent story of weak evidence. A p-value of 0.15 suggests some signal but insufficient evidence for significance, while 0.06 indicates a slightly stronger signal that still falls short of the conventional 0.05 threshold. The fact that even with quadrupled sample size (n=20 to n=80) the effect barely approached significance suggests the true effect size may be smaller than initially estimated, raising questions about practical importance. A is wrong because while the pilot study was indeed underpowered, p=0.06 doesn't "confirm" the effect exists—it's still not statistically significant. B incorrectly frames this as contradictory results when both are actually on the same side of the significance threshold and tell a coherent story of marginal evidence. D is wrong because pilot studies aren't "misleading" just because they're underpowered—they provide valuable preliminary information that should be interpreted in context with follow-up studies. Study tip: Don't fall into the "significant vs. not significant" binary thinking trap. P-values near 0.05 (whether slightly above or below) often indicate weak evidence that warrants careful interpretation. Always consider the pattern across studies and whether effect sizes justify practical significance, not just statistical significance.

Question 16

A journal editor receives a manuscript reporting p = 0.052 and requests that the authors 'discuss the implications of this near-significant result.' The authors argue that p-values should be reported as exact values without reference to arbitrary thresholds. Who has the better argument?

  1. The editor is correct because p = 0.052 represents a meaningful category distinct from p < 0.05
  2. The authors are correct because the exact p-value provides more information than threshold-based interpretations (correct answer)
  3. The editor is correct because readers need guidance on how to interpret borderline results
  4. The authors are wrong because statistical conventions exist for good reasons and should be maintained
  5. Both are partially correct, but the p-value should be supplemented with confidence intervals and effect sizes
Explanation: This question tests your understanding of p-value interpretation and the ongoing debate about statistical significance thresholds in biostatistics. When you encounter questions about p-values near 0.05, focus on what information is actually being conveyed versus arbitrary categorizations. The authors have the stronger argument because exact p-values provide readers with the complete statistical information needed to make their own interpretations. A p-value of 0.052 tells you the precise probability of observing your data (or more extreme) under the null hypothesis. This allows readers to assess the strength of evidence on a continuous scale rather than forcing a binary significant/not-significant decision. The exact value also enables proper meta-analyses and helps readers understand how close the result was to conventional thresholds. Option A is incorrect because p = 0.052 versus p < 0.05 represents an artificial categorical distinction. The difference between 0.049 and 0.052 is trivial and doesn't represent meaningfully different evidence against the null hypothesis. Option C seems reasonable but is flawed because "guidance" shouldn't come from arbitrary threshold enforcement. Better guidance comes from teaching readers to interpret the continuous nature of evidence that p-values represent. Option D incorrectly assumes that maintaining conventions is more important than providing complete information. While conventions have their place, the scientific community increasingly recognizes that rigid adherence to p < 0.05 can be misleading. Remember: p-values are continuous measures of evidence, not binary indicators of truth. Always report exact values and avoid treating 0.05 as a magical boundary that transforms evidence quality.

Question 17

A randomized trial reports the primary outcome as 'not significant (p = 0.18)' but emphasizes a subgroup analysis showing p = 0.01. The authors conclude the treatment is effective in the subgroup. What are the main concerns with this interpretation?

  1. Subgroup analyses require different statistical methods than primary analyses, making comparison invalid
  2. The p-value of 0.18 contradicts the subgroup finding, indicating an error in the analysis
  3. Subgroup analyses are exploratory and prone to false positives, especially when the overall result is negative (correct answer)
  4. The sample size in the subgroup was likely too small to provide reliable statistical inference
  5. The significance level should be adjusted for the subgroup analysis to account for the failed primary endpoint
Explanation: When you encounter a study where the primary outcome is non-significant but the authors emphasize a significant subgroup result, you're looking at a classic statistical interpretation pitfall known as "subgroup hunting" or multiple testing bias. The correct answer is C because subgroup analyses are inherently exploratory and carry high risk of false positives, especially when the overall trial is negative. Here's why: when you divide your sample into subgroups, you're essentially performing multiple statistical tests. Each test has a 5% chance of showing significance by random chance alone. With multiple subgroups, the probability of finding at least one "significant" result increases dramatically, even when no true effect exists. When the primary analysis is already negative (p = 0.18), emphasizing a positive subgroup finding is particularly suspect. A is incorrect because subgroup analyses use the same statistical methods as primary analyses - the issue isn't methodological incompatibility. B misunderstands the relationship between overall and subgroup results. A non-significant overall result doesn't contradict a significant subgroup finding - it's actually expected when subgroup hunting occurs. D, while potentially true, misses the main concern. Small sample sizes in subgroups are problematic, but the primary issue is the multiple testing bias, not just power. Remember this key principle: when the primary endpoint fails, be extremely skeptical of post-hoc subgroup claims. True subgroup effects should be pre-specified, biologically plausible, and ideally replicated in independent studies. This is a high-yield concept for biostatistics exams.

Question 18

A meta-analysis combines 5 studies, each with p-values of 0.08, 0.12, 0.09, 0.11, and 0.07. None were individually significant at α=0.05\alpha = 0.05. The combined analysis yields p = 0.02. What does this suggest?

  1. The individual studies were all flawed since they failed to detect what the meta-analysis found
  2. The meta-analysis is incorrect because you cannot combine non-significant results to get significance
  3. Each study may have had insufficient power to detect a real but small effect (correct answer)
  4. The significance level should have been set higher in the original studies to match the meta-analysis
  5. The meta-analysis p-value should be the average of the individual p-values, which is 0.094
Explanation: When you encounter meta-analysis questions in biostatistics, focus on the fundamental concept of statistical power—the ability to detect a true effect when it exists. Meta-analysis works by pooling data to increase overall sample size and power. The correct answer is C because each individual study likely had insufficient power to detect a real but small effect. Notice that all five p-values (0.08, 0.12, 0.09, 0.11, 0.07) cluster just above the α=0.05\alpha = 0.05 threshold, suggesting a consistent but weak signal across studies. When meta-analysis combines these studies, it increases the effective sample size, boosting statistical power and allowing detection of the true underlying effect (p = 0.02). Answer A is wrong because studies aren't "flawed" simply for lacking power—they may have been appropriately designed for their available resources, but individually underpowered for small effects. Answer B reflects a common misconception; you absolutely can combine non-significant results to achieve significance if there's a consistent underlying effect that individual studies were underpowered to detect. This is precisely why meta-analysis exists. Answer D is incorrect because raising the significance level (making α\alpha larger) would make it easier to find significance in individual studies but wouldn't address the fundamental power issue. Remember this pattern: when you see multiple studies with p-values hovering just above 0.05 that become significant when combined, think "power problem solved by larger sample size," not methodological flaws or statistical errors.

Question 19

A medical journal requires p < 0.001 for publication of clinical trials, while another journal uses p < 0.05. A researcher obtains p = 0.02 in a well-designed study. What does this situation illustrate about p-values and significance levels?

  1. The first journal's requirement is scientifically invalid because 0.05 is the universally accepted standard
  2. The researcher's evidence is inadequate because it fails to meet the more stringent requirement
  3. Significance levels are arbitrary thresholds, and the strength of evidence should be evaluated independently (correct answer)
  4. The researcher should conduct a larger study to reduce the p-value below 0.001 for publication
  5. The p-value of 0.02 indicates the study has a 98% probability of being correct regardless of journal requirements
Explanation: This question tests your understanding of p-values as measures of evidence strength rather than absolute determinants of truth. When you encounter questions about significance thresholds, focus on the underlying statistical reasoning rather than rigid rules. The correct answer is C because significance levels (α) are indeed arbitrary cutoff points that researchers and journals establish by convention. A p-value of 0.02 represents the same strength of evidence regardless of which journal evaluates it. The evidence suggests that if the null hypothesis were true, you'd see results this extreme or more only 2% of the time. This statistical evidence doesn't change based on publication requirements. Option A is wrong because there's no universal scientific law mandating p < 0.05. Different fields and contexts may appropriately use different thresholds based on the consequences of Type I errors. Option B incorrectly suggests the evidence itself is inadequate when the study may be perfectly valid and informative. The evidence strength remains constant at p = 0.02. Option D reflects a common misconception about p-hacking. Simply increasing sample size to achieve a predetermined p-value threshold is methodologically problematic and doesn't improve the quality of evidence—it manipulates statistics rather than strengthening the underlying research. Key takeaway: Remember that p-values measure evidence strength on a continuum, not binary significance. When you see biostatistics questions about publication standards or significance thresholds, focus on the actual evidence presented rather than arbitrary cutoff points. The strength of your statistical evidence doesn't depend on editorial policies.

Question 20

A pharmaceutical company conducts a large trial (n = 5000) and reports: 'Treatment significantly reduced symptoms (p = 0.03) with a mean improvement of 0.2 points on a 50-point scale.' What is the most important consideration for interpreting this result?

  1. The p-value indicates strong evidence, so the 0.2-point improvement should be considered clinically meaningful
  2. The large sample size makes any significant result highly reliable regardless of effect magnitude
  3. The clinical significance of a 0.2-point improvement should be evaluated independently of statistical significance (correct answer)
  4. The p-value is too close to 0.05 to be trustworthy given the large sample size
  5. The result should be replicated in a smaller, more controlled study to confirm the findings
Explanation: When interpreting research results, you must distinguish between statistical significance (whether an effect exists) and clinical significance (whether that effect matters in practice). This distinction becomes especially critical with large sample sizes, which can detect tiny differences that may be statistically significant but clinically meaningless. In this study, the 0.2-point improvement on a 50-point scale represents just a 0.4% change. While the large sample size (n = 5000) provided enough statistical power to detect this small difference as "significant" (p = 0.03), this doesn't automatically make it clinically meaningful. Clinical significance depends on factors like what constitutes a noticeable improvement for patients, the scale's properties, and the intervention's costs and risks. Option A incorrectly assumes that statistical significance guarantees clinical importance. A p-value only tells you the likelihood of seeing such results by chance—it says nothing about whether the effect size matters to patients. Option B makes the dangerous assumption that large samples validate any significant finding regardless of effect magnitude. While large samples do increase reliability and reduce random error, they also make it easier to find trivial differences. Option D misunderstands p-value interpretation; p = 0.03 is statistically significant regardless of sample size, and being "close to 0.05" doesn't invalidate results. Study tip: Remember that with very large samples, almost any difference can become statistically significant. Always ask yourself: "Even if this effect is real, does the magnitude actually matter?" Look for effect sizes and confidence intervals, not just p-values, when evaluating research.