All questions
Question 1
A research team publishes a study showing a significant treatment effect (p = 0.02). Later, another team attempts to replicate the study with a larger sample size and finds a non-significant result (p = 0.18) with a similar effect size. The original authors claim the replication 'failed' because it didn't achieve significance. What is the most accurate assessment of this situation?
- The replication failed because it didn't reproduce the significant finding
- The replication succeeded because it found a similar effect size (correct answer)
- The original study was likely a false positive finding
- The replication used inappropriate methodology with larger sample size
- Both studies are equally valid and contradictory results are expected
Explanation: When evaluating research replications, you need to distinguish between statistical significance and effect size - two fundamentally different concepts that students often conflate.
The key insight here is that a successful replication reproduces the underlying effect, not necessarily the p-value. Both studies found similar effect sizes, meaning they detected the same magnitude of treatment benefit. The difference in statistical significance simply reflects the inherent variability in p-values across studies, even when the true effect remains constant.
Answer B is correct because finding a similar effect size indicates the replication successfully detected the same underlying phenomenon. Statistical significance is just one piece of evidence, and p-values naturally vary between studies due to sampling variation.
Answer A falls into the common trap of equating replication success with achieving significance. This misconception ignores that p-values fluctuate around the significance threshold even when effects are real. Answer C jumps to an unwarranted conclusion - while false positives are possible, the similar effect sizes actually support the original finding's validity. Answer D incorrectly suggests that larger sample sizes are problematic, when they actually provide more precise estimates and are methodologically superior.
The larger sample in the replication likely provided a more precise effect estimate, even though random variation pushed the p-value above 0.05. Remember: focus on effect sizes and confidence intervals when evaluating replications, not just whether studies cross the arbitrary p = 0.05 threshold. Successful science replicates effects, not p-values.
Question 2
An epidemiological study examines 25 potential risk factors for cardiovascular disease. The researchers report that 'several risk factors showed significant associations' and present only the 4 factors with p < 0.05 in their main results table. The remaining 21 factors are mentioned briefly as 'non-significant' without specific p-values. What practice does this represent?
- Efficient presentation of only clinically relevant findings
- Standard practice for exploratory epidemiological research
- Selective reporting that may mislead readers about effect prevalence (correct answer)
- Appropriate focus on statistically significant relationships
- Necessary space-saving measure for journal publication limits
Explanation: When you encounter questions about research reporting practices, focus on the principles of transparency and complete disclosure that underpin scientific integrity. This scenario describes a classic case of selective reporting bias.
The researchers' approach represents selective reporting that misleads readers about effect prevalence (C). By presenting only the 4 "significant" factors while burying the 21 non-significant ones, they create a distorted picture. Readers seeing "several risk factors showed significant associations" might assume most factors tested were significant, when actually only 16% (4/25) showed associations. This selective presentation inflates the apparent prevalence of meaningful risk factors and violates principles of complete scientific reporting.
Option A is wrong because clinical relevance isn't determined solely by statistical significance—some non-significant findings may have important clinical implications or help rule out suspected risk factors. Option B incorrectly suggests this is standard practice; while exploratory studies do test multiple factors, transparent reporting requires disclosing all results, not just favorable ones. Option D reflects a common misconception that statistical significance alone determines what merits reporting—this thinking perpetuates publication bias and incomplete scientific records.
The selective reporting also raises concerns about multiple comparisons. With 25 tests, you'd expect 1-2 significant results by chance alone at α = 0.05, making the 4 significant findings less impressive than presented.
Study tip: Watch for reporting scenarios where researchers emphasize positive findings while downplaying negative ones. Complete transparency—reporting all tested variables with their results—is always the gold standard in research ethics.
Question 3
A researcher conducts 15 independent statistical tests on the same dataset, each testing a different hypothesis at α = 0.05. Two tests yield p-values of 0.03 and 0.04, while the remaining 13 tests have p-values > 0.05. The researcher reports only the two significant results in the publication. What is the primary statistical concern with this approach?
- The sample size is too small for reliable inference
- The alpha level should be adjusted for multiple comparisons (correct answer)
- The effect sizes are likely too small to be meaningful
- The confidence intervals should be reported instead of p-values
- The assumptions of independence have been violated
Explanation: When you encounter multiple hypothesis testing scenarios, the key concern is maintaining the overall error rate across all tests performed. This is a classic multiple comparisons problem.
The researcher conducted 15 independent tests, each at α = 0.05. Without adjustment, the probability of finding at least one "significant" result by chance alone is approximately 1−(0.95)15=0.54, meaning there's a 54% chance of a Type I error across all tests. When only the "significant" results are reported, this creates a misleading picture of the evidence. Answer B correctly identifies that the alpha level should be adjusted using methods like Bonferroni correction (α/number of tests) or false discovery rate controls to maintain the intended error rate.
Answer A misses the point—sample size isn't the issue here; it's about controlling error rates across multiple tests. The sample size for each individual test may be perfectly adequate. Answer C about effect sizes is irrelevant to the statistical validity concern. Small effect sizes can still be scientifically meaningful, and the p-values alone don't tell us about effect magnitude. Answer D suggests reporting confidence intervals instead, but this doesn't address the fundamental multiple comparisons problem. You'd still have the same inflated Type I error rate whether using p-values or confidence intervals.
Study tip: Whenever you see multiple tests on the same dataset, immediately think "multiple comparisons correction." The more tests performed, the higher the chance of false positives, regardless of what gets reported in the final publication. Question 4
A clinical trial protocol specifies that the primary endpoint will be analyzed when 200 patients complete the study. After enrolling 150 patients, the investigators notice the treatment effect is not as strong as expected. They decide to change the primary endpoint to a different outcome measure that shows more promise in the current data. What type of research practice does this represent?
- Appropriate adaptive trial design methodology
- Necessary protocol amendment for patient safety
- Outcome switching that compromises study validity (correct answer)
- Standard interim analysis with stopping rules
- Required modification due to insufficient power
Explanation: When you encounter questions about changes made to study protocols after data collection has begun, focus on the principle of scientific integrity and the potential for bias introduction.
Why C is correct: This scenario represents outcome switching, a serious threat to study validity. The investigators changed their primary endpoint mid-study specifically because the original endpoint wasn't showing favorable results. This constitutes "cherry-picking" outcomes based on observed data trends, which can dramatically inflate the apparent treatment effect and compromise the study's credibility. The timing is crucial here—they made this change after seeing disappointing results in 150 patients, not based on legitimate scientific or safety concerns.
Why the other answers are wrong: A) Adaptive trial designs involve pre-specified rules for modifications that are built into the original protocol, not post-hoc changes based on unfavorable interim results. B) Protocol amendments for patient safety would involve safety endpoints or adverse events, not switching to a "more promising" efficacy outcome. D) Standard interim analyses follow predetermined stopping rules and statistical boundaries—they don't involve changing the primary endpoint because results aren't as strong as hoped.
Key study tip: Remember that legitimate protocol modifications must be pre-specified or based on safety concerns, never on the desire to find more favorable results. When you see scenarios describing endpoint changes after interim data review, particularly when motivated by disappointing efficacy results, this almost always represents problematic outcome switching that undermines the study's validity and introduces bias.
Question 5
In a meta-analysis of 20 studies examining a new drug's effectiveness, researchers find that studies with significant results (p < 0.05) are more likely to be published in high-impact journals, while studies with non-significant results are often unpublished or appear in lower-tier journals. This pattern most directly threatens which aspect of the meta-analysis?
- The statistical power of individual studies included
- The external validity of the combined results
- The internal validity of each component study
- The representativeness of the pooled effect estimate (correct answer)
- The heterogeneity assessment between studies
Explanation: When you encounter questions about systematic bias in research synthesis, focus on how different types of bias affect the validity and representativeness of meta-analytic results.
The scenario describes classic publication bias - the tendency for studies with statistically significant, positive results to be published more frequently and in higher-impact venues than studies with null or negative findings. This creates a fundamental sampling problem for meta-analyses.
Answer D is correct because publication bias directly compromises the representativeness of the pooled effect estimate. When a meta-analysis predominantly includes published studies with significant results while missing unpublished null studies, the combined effect size becomes artificially inflated. The pooled estimate no longer represents the true population effect - it represents a biased sample skewed toward positive findings.
Answer A is incorrect because publication bias doesn't affect the statistical power of individual studies that are already completed. Power is determined during study design, not publication decisions.
Answer B misses the mark because external validity refers to generalizability across populations, settings, or conditions. While publication bias affects what studies contribute to the meta-analysis, it doesn't directly threaten whether results apply to different contexts.
Answer C is wrong because internal validity concerns whether individual studies accurately measure causal relationships within their own design. Publication bias affects which studies get included in the synthesis, not the quality of the studies themselves.
Study tip: Remember that publication bias creates a "file drawer problem" - significant results get published while null results stay hidden. This sampling bias distorts meta-analytic effect estimates, making them appear larger than the true population effect.
Question 6
A researcher tests the association between a biomarker and disease outcome using three different statistical approaches: logistic regression, Cox proportional hazards model, and Mann-Whitney U test. The p-values are 0.08, 0.04, and 0.12, respectively. The researcher reports only the Cox model results, stating it is the 'most appropriate' method without justification. What is the main concern with this reporting strategy?
- The Cox model assumptions have not been verified
- Multiple testing correction should be applied to all three tests
- The choice of statistical method appears to be result-driven (correct answer)
- The sample size may be insufficient for reliable Cox modeling
- The biomarker may not meet proportional hazards assumptions
Explanation: This question tests your understanding of research integrity and proper statistical reporting practices. When researchers apply multiple statistical methods to the same data, the method selection should be predetermined based on study design and data characteristics, not chosen after seeing which produces the most favorable results.
The correct answer is C because the researcher's behavior suggests "p-hacking" or result-driven analysis. Running three different tests and then selecting only the one that yielded a significant p-value (0.04) while dismissing the non-significant results (0.08 and 0.12) without methodological justification represents questionable research practice. The claim that Cox modeling is "most appropriate" appears to be reverse-engineered from the results rather than based on a priori statistical reasoning.
Option A is incorrect because while Cox model assumptions should indeed be verified, the primary concern here isn't about assumption checking—it's about the selective reporting pattern. Option B is wrong because multiple testing correction only applies when you're genuinely testing multiple hypotheses simultaneously by design, not when you're choosing between fundamentally different analytical approaches for the same research question. Option D misses the point entirely—sample size adequacy, while important for Cox modeling, doesn't address the core ethical issue of cherry-picking significant results.
Remember this key principle: statistical method selection should always be justified by your research question, data structure, and study design—never by which approach gives you the p-value you want. This type of selective reporting undermines scientific integrity and inflates Type I error rates.
Question 7
A pharmaceutical company conducts 8 Phase II trials for a new medication across different populations. Three trials show statistically significant benefits (p < 0.05), while five show no significant effect. In their FDA submission, they emphasize the three positive trials and briefly mention the negative ones in an appendix. From a regulatory perspective, what is the most serious issue with this approach?
- The statistical power was likely inadequate across all trials
- The populations studied were too heterogeneous for comparison
- The selective emphasis may misrepresent the overall evidence (correct answer)
- The trials should have been combined in a single large study
- The significance threshold should be lowered for exploratory trials
Explanation: When you encounter questions about research reporting and regulatory submissions, focus on the fundamental principle of scientific integrity: presenting a complete and balanced view of all available evidence.
The most serious regulatory concern here is selective emphasis that misrepresents the overall evidence (C). When a company highlights 3 positive trials while downplaying 5 negative ones, they're creating a misleading impression of the drug's effectiveness. Regulatory agencies like the FDA need to see the complete picture—both successes and failures—to make informed decisions about public safety. This selective presentation violates the principle that all conducted studies should be given appropriate weight in regulatory submissions, regardless of their outcomes.
Looking at the other options: (A) assumes inadequate statistical power without evidence—some trials may have been properly powered, and this doesn't address the core reporting issue. (B) suggests population heterogeneity is the main problem, but diverse populations are often desirable in Phase II trials to understand drug performance across different groups. (D) implies the study design was flawed, but conducting multiple smaller trials instead of one large trial is often a valid Phase II approach that allows for adaptive learning.
Remember this key principle: in biostatistics and regulatory affairs, questions about study reporting often test whether you recognize threats to scientific integrity. The most serious issues typically involve selective presentation, publication bias, or incomplete disclosure—not just statistical methodology. Always consider whether the approach gives regulators and clinicians an honest, complete view of the evidence.
Question 8
A research team pre-registers a study protocol specifying that they will analyze the relationship between exercise frequency and depression scores using linear regression. After data collection, they discover the relationship is non-linear and decide to use polynomial regression instead, finding a significant result (p = 0.02). How should they report this finding?
- Report the polynomial regression as planned, since it's more appropriate
- Report both linear and polynomial results with transparency about the protocol deviation (correct answer)
- Report only the linear regression to maintain protocol adherence
- Update the pre-registration retroactively to reflect the polynomial approach
- Report the polynomial regression without mentioning the original protocol
Explanation: When you encounter questions about pre-registered study protocols and post-hoc analysis changes, you're dealing with research transparency and scientific integrity principles. Pre-registration is designed to prevent p-hacking and selective reporting, but researchers must balance protocol adherence with appropriate statistical methods.
The correct approach here is Option B - reporting both the pre-registered linear regression and the post-hoc polynomial regression with full transparency about the protocol deviation. This maintains scientific integrity by showing readers both what was planned and what was actually done, allowing them to evaluate potential bias. The polynomial regression may indeed be more appropriate for the data, but the deviation from protocol must be acknowledged.
Option A fails because simply reporting the polynomial regression without mentioning the protocol deviation conceals important methodological information that could affect interpretation. Option C is problematic because rigidly adhering to an inappropriate statistical method (linear regression for non-linear data) compromises the validity of your findings and wastes valuable data. Option D represents research misconduct - retroactively changing pre-registrations defeats their entire purpose and constitutes falsification of the research record.
Study tip: Remember that pre-registration doesn't lock you into inappropriate methods forever, but any deviations must be transparently reported. On biostatistics exams, questions about research integrity often test whether you understand that transparency trumps convenience - always choose the option that maintains scientific honesty even when it's more work to report.
Question 9
A researcher analyzes survey data and finds that income is not significantly associated with health satisfaction (p = 0.12). They then decide to examine the relationship separately for men and women, finding significant associations in women (p = 0.03) but not men (p = 0.67). They report: 'Income significantly predicts health satisfaction in women.' What practice does this exemplify?
- Appropriate subgroup analysis based on biological rationale
- Data dredging through post-hoc stratification (correct answer)
- Necessary adjustment for effect modification
- Standard exploratory data analysis technique
- Required analysis to address confounding by gender
Explanation: When researchers conduct statistical analyses, the timing and justification for subgroup analyses is crucial for maintaining scientific integrity. This question tests your ability to distinguish between appropriate analytical strategies and problematic data mining practices.
The scenario describes a classic case of data dredging (also called p-hacking or fishing expeditions). The researcher initially found no significant overall association between income and health satisfaction (p = 0.12). Rather than accepting this null result, they then searched through subgroups until finding a significant result in women. This post-hoc exploration without prior hypothesis or theoretical justification exemplifies data dredging, making B correct.
A is wrong because there's no mention of biological rationale driving the subgroup analysis. The decision appears purely motivated by the initial non-significant result, not by theoretical considerations about sex-based differences.
C is incorrect because effect modification (interaction) should be planned and tested formally using interaction terms in statistical models, not discovered through post-hoc subgroup hunting after finding null overall results.
D is wrong because while exploratory analysis has its place, reporting a post-hoc subgroup finding as a definitive conclusion ("Income significantly predicts health satisfaction in women") without acknowledging the exploratory nature crosses into inappropriate territory.
Study tip: Watch for scenarios where researchers find non-significant results, then slice their data various ways until something becomes significant. Legitimate subgroup analyses should be pre-specified with clear scientific rationale, not fishing expeditions triggered by disappointing overall results.
Question 10
A clinical trial compares three treatment groups (A, B, and placebo) with the primary outcome measured at 6 months. The protocol specified analysis at 6 months only. However, measurements were also taken at 3 months, and the researchers notice that treatment A shows significance at 3 months (p = 0.04) but not at 6 months (p = 0.08). What is the most appropriate interpretation of these results?
- Treatment A is effective based on the 3-month significant result
- Treatment A shows early promise but fails to meet the primary endpoint (correct answer)
- The 6-month result is invalid due to patient dropout
- Both time points should be considered equally important
- Treatment A demonstrates sustained effectiveness over time
Explanation: When you encounter clinical trial questions involving multiple time points, remember that the primary endpoint specified in the protocol takes precedence over exploratory or secondary analyses. This principle protects against data dredging and maintains statistical integrity.
The key issue here is distinguishing between pre-specified primary outcomes and post-hoc observations. Since the protocol specified 6-month analysis as the primary endpoint, this is where the definitive conclusion must be drawn. At 6 months, treatment A failed to reach significance (p = 0.08), meaning it did not meet the study's primary objective.
Answer B correctly captures this nuance: treatment A showed "early promise" at 3 months but "failed to meet the primary endpoint" at 6 months. This interpretation acknowledges the 3-month finding while properly prioritizing the pre-specified primary analysis.
Answer A is wrong because declaring effectiveness based solely on a secondary time point violates the study design and increases Type I error risk through multiple testing without adjustment.
Answer C incorrectly assumes patient dropout invalidates results. Dropout is a common issue in trials, but there's no evidence presented that dropout rates compromise the 6-month analysis validity.
Answer D is incorrect because not all time points are equal in clinical trials. The primary endpoint, defined prospectively in the protocol, carries the most weight for regulatory and clinical decision-making.
Study tip: In biostatistics questions about clinical trials, always identify what was pre-specified versus post-hoc. Pre-specified primary endpoints trump exploratory findings, even when the exploratory results seem more favorable.
Question 11
A researcher plans to conduct pairwise comparisons between 5 treatment groups using t-tests. Without any adjustment for multiple comparisons, how many individual tests will be performed, and what is the family-wise error rate if each test uses α = 0.05?
- 5 tests; family-wise error rate = 0.25
- 10 tests; family-wise error rate = 0.40 (correct answer)
- 10 tests; family-wise error rate = 0.50
- 15 tests; family-wise error rate = 0.54
- 20 tests; family-wise error rate = 0.64
Explanation: When you encounter questions about multiple comparisons, you need to consider two key calculations: the number of pairwise tests and the cumulative probability of making at least one Type I error across all tests.
For 5 treatment groups, the number of unique pairwise comparisons follows the combination formula: (25)=2!(5−2)!5!=25×4=10 tests. Each group must be compared with every other group exactly once.
The family-wise error rate (FWER) represents the probability of making at least one Type I error across all tests. With independent tests, this is calculated as: FWER=1−(1−α)k where α = 0.05 and k = 10 tests. So: FWER=1−(0.95)10=1−0.599=0.401≈0.40
Answer A incorrectly calculates only 5 tests, likely confusing the number of groups with the number of comparisons needed. Answer C uses the correct number of tests but miscalculates the FWER, possibly using a simpler but incorrect formula like 10 × 0.05. Answer D shows 15 tests, which would be correct for 6 groups, not 5, though the FWER calculation appears roughly consistent with that incorrect number.
Remember this pattern: for n groups, you need (2n) comparisons, and the FWER grows exponentially with the number of tests. This is why multiple comparison corrections like Bonferroni are essential in practice to control Type I error inflation. Question 12
A medical journal implements a policy requiring authors to provide access to raw datasets upon reasonable request. A submitted manuscript reports significant results but the authors refuse to share data, citing 'proprietary concerns.' From a transparency and reproducibility perspective, what should the journal's response be?
- Accept the paper since proprietary concerns are legitimate
- Require data sharing or consider rejection based on transparency standards (correct answer)
- Accept the paper but publish it with a data availability disclaimer
- Request only the statistical analysis code instead of raw data
- Accept the paper if the results can be verified through peer review
Explanation: When you encounter questions about research transparency and data sharing policies, you're dealing with core principles of scientific integrity and reproducibility. These issues have become increasingly important as the scientific community grapples with the replication crisis.
The correct approach here is option B - requiring data sharing or considering rejection based on transparency standards. When a journal establishes a clear data sharing policy, consistent enforcement is essential for maintaining scientific integrity. Transparency allows other researchers to verify results, identify potential errors, and build upon findings. Without access to raw data, the scientific community cannot adequately evaluate the validity of reported results, especially when those results claim statistical significance.
Option A is problematic because while proprietary concerns may sometimes be legitimate, they cannot override a journal's established transparency requirements. If authors cannot meet these standards, the venue may not be appropriate for their work. Option C represents a weak compromise that undermines the journal's policy - a disclaimer doesn't provide the transparency needed for reproducibility and essentially creates a two-tier system where some papers meet standards and others don't. Option D fails because statistical code alone, without the underlying data, still prevents full verification and replication of the analysis.
Remember that data sharing policies in academic publishing are designed to strengthen scientific rigor, not create barriers. When you see transparency-related questions, prioritize approaches that maintain consistent standards and enable reproducible research over accommodations that compromise scientific integrity.
Question 13
A researcher conducts a study with 200 participants and finds a correlation between variables X and Y of r = 0.18 (p = 0.045). To increase the sample size, they collect data from 100 additional participants. In the combined dataset (n = 300), the correlation becomes r = 0.15 (p = 0.08). Which result should they report?
- Only the original significant result with n = 200
- Only the final result with the complete dataset n = 300
- Both results, explaining the data collection process transparently (correct answer)
- The result that shows the stronger effect size (r = 0.18)
- A meta-analysis combining both samples as separate studies
Explanation: When you encounter questions about reporting research results, especially when data collection occurs in phases, you're dealing with scientific transparency and ethical reporting standards. The key principle is that researchers must report their complete methodology and all relevant findings.
The correct approach is to report both results with full transparency about the data collection process (C). This allows readers to understand how the study evolved and make informed judgments about the findings. The original analysis with n=200 showed r=0.18 (p=0.045), suggesting a statistically significant but weak correlation. After adding 100 participants, the correlation weakened to r=0.15 (p=0.08), losing statistical significance. Both pieces of information are scientifically valuable and should be disclosed.
Option A is problematic because reporting only the significant result while hiding the additional data constitutes selective reporting bias - a serious ethical violation that can mislead the scientific community. Option B fails because it ignores the sequential nature of data collection and doesn't explain why the sample size changed, potentially confusing readers about the study design. Option D represents effect size cherry-picking, which violates principles of honest reporting by selecting results based on strength rather than methodological completeness.
The change from significant to non-significant results with increased sample size actually provides important information about the robustness (or lack thereof) of the original finding. This scenario illustrates why effect sizes near significance thresholds require careful interpretation.
Study tip: Remember that transparency in reporting methodology and results is always preferable to selective reporting, even when results become less favorable with additional data.
Question 14
A researcher tests 20 different biomarkers for association with disease risk. They find that 3 biomarkers have p-values less than 0.05. Using the Bonferroni correction, what adjusted significance level should be used, and how many of the findings would remain significant at this level?
- Adjusted α = 0.0025; all 3 findings remain significant
- Adjusted α = 0.0025; number significant depends on actual p-values (correct answer)
- Adjusted α = 0.01; all 3 findings remain significant
- Adjusted α = 0.05; no adjustment needed for biomarker studies
- Adjusted α = 0.002; approximately 1 finding remains significant
Explanation: When you encounter multiple comparisons in biostatistics, you need to understand the multiple testing problem. Testing many hypotheses simultaneously inflates your chance of finding false positives, so corrections like Bonferroni are essential to maintain the overall Type I error rate.
The Bonferroni correction divides your original significance level by the number of tests performed. Here, with 20 biomarkers tested at α=0.05, the adjusted significance level becomes αadjusted=0.05/20=0.0025. This is the new threshold each individual p-value must meet to be considered significant.
The key insight is that knowing 3 biomarkers had p-values less than 0.05 doesn't tell you whether they'll remain significant at the stricter 0.0025 level. For example, p-values of 0.001, 0.01, and 0.04 would all be less than 0.05, but only the first would survive Bonferroni correction.
Option A incorrectly assumes all findings automatically remain significant after correction. Option C miscalculates the adjusted alpha as 0.01 (this would be correct for 5 tests, not 20). Option D wrongly suggests biomarker studies don't need multiple testing corrections – they absolutely do when testing many markers simultaneously.
Remember this pattern: Bonferroni correction always makes your significance threshold more stringent (smaller), and whether findings survive depends on their exact p-values, not just whether they initially met the uncorrected threshold. Always calculate the adjusted alpha first, then evaluate each finding against this stricter standard. Question 15
A pharmaceutical company conducts 4 clinical trials for a new drug. Trial 1 shows p = 0.02, Trial 2 shows p = 0.15, Trial 3 shows p = 0.08, and Trial 4 shows p = 0.03. In their marketing materials, they state: 'Clinical trials demonstrate significant efficacy (p < 0.05).' What aspect of this statement is most problematic from a transparency perspective?
- The statement should report exact p-values for all trials
- The statement implies consistent evidence while ignoring mixed results (correct answer)
- The significance threshold should be adjusted for multiple trials
- Marketing materials should not include statistical information
- The statement should include confidence intervals instead of p-values
Explanation: When evaluating research transparency and ethical reporting, you need to consider whether the presented information gives readers a complete and honest picture of the evidence. This question tests your understanding of selective reporting and cherry-picking issues in scientific communication.
The pharmaceutical company's statement is problematic because it creates a misleading impression by emphasizing only the favorable results. While technically true that some trials showed significance (p < 0.05), the statement "clinical trials demonstrate significant efficacy" suggests consistent, strong evidence across studies. In reality, the results are mixed: two trials were significant (p = 0.02, p = 0.03) while two were not (p = 0.15, p = 0.08). This selective emphasis while ignoring contradictory findings violates transparency principles and can mislead healthcare providers and patients about the drug's true efficacy profile.
Looking at the other options: (A) is incorrect because while reporting exact p-values would be more informative, the core transparency problem isn't the level of detail but the selective presentation. (C) addresses multiple testing corrections, which is a valid statistical concern but not the primary transparency issue highlighted here. (D) is wrong because statistical information in marketing can be appropriate when presented honestly and completely.
Remember that transparency in biostatistics isn't just about statistical accuracy—it's about presenting the complete picture. When you see research claims, always ask: "What information might be missing?" Look for patterns where only favorable results are highlighted while unfavorable or mixed findings are downplayed or omitted entirely.
Question 16
A researcher analyzing survey data discovers that their primary hypothesis yields p = 0.07. They then examine the data using different variable transformations: log transformation gives p = 0.03, square root transformation gives p = 0.09, and standardization gives p = 0.04. They report the log transformation result as their primary finding. What best describes this practice?
- Appropriate model selection based on data characteristics
- Standard exploratory data analysis methodology
- Analytical flexibility that inflates Type I error (correct answer)
- Necessary adjustment for non-normal data distribution
- Proper validation of statistical assumptions
Explanation: This question tests your understanding of p-hacking and multiple testing problems in statistical analysis. When researchers test multiple analytical approaches and selectively report the most favorable result, they're engaging in questionable research practices that compromise statistical validity.
The scenario describes a classic case of analytical flexibility inflating Type I error. The researcher started with a non-significant result (p = 0.07) and tried various transformations until finding one that achieved significance (p = 0.03 with log transformation). By testing multiple approaches and cherry-picking the significant result, they've essentially conducted multiple statistical tests without adjusting for this multiplicity. This inflates the probability of finding a false positive result well above the nominal α = 0.05 level, making answer C correct.
Answer A is wrong because appropriate model selection should be driven by theoretical considerations and data characteristics determined before analysis, not by which method yields significant results. Answer B incorrectly frames this as standard practice—while exploratory analysis is legitimate, reporting exploratory results as confirmatory findings without acknowledging the multiple testing is problematic. Answer D misses the point entirely; even if the data were non-normal and transformations were justified, the issue here is the selective reporting of results based on statistical significance rather than pre-planned analytical strategy.
Remember: whenever you see researchers trying multiple analytical approaches and reporting only the significant result, think "p-hacking." Legitimate data transformation should be justified by the data's distributional properties and planned in advance, not selected post-hoc based on p-values.
Question 17
A meta-analysis protocol specifies inclusion of all randomized controlled trials published between 2010-2020. During the search process, researchers identify 15 relevant studies but exclude 3 studies with non-significant results, stating they had 'methodological limitations' not mentioned in their exclusion criteria. The final meta-analysis of 12 studies shows a significant treatment effect. What is the primary methodological concern?
- The search period was too restrictive for adequate evidence
- Post-hoc exclusions based on results introduce selection bias (correct answer)
- The sample of 12 studies is too small for reliable meta-analysis
- Methodological quality should always be assessed before inclusion
- The protocol should have specified stricter inclusion criteria
Explanation: When you encounter meta-analysis questions, focus on whether researchers followed their pre-specified protocol and avoided bias that could distort results.
The critical issue here is that researchers deviated from their original inclusion criteria by excluding studies after seeing their results. They initially planned to include all relevant RCTs from 2010-2020, but then removed three studies with non-significant findings, claiming "methodological limitations" that weren't part of their original exclusion criteria. This is classic selection bias—systematically excluding studies based on their outcomes rather than predetermined quality standards.
Option B correctly identifies this post-hoc exclusion problem. When you exclude studies after seeing they don't support your hypothesis, you artificially inflate the treatment effect and compromise the meta-analysis's validity.
Option A is wrong because a 10-year search window is typically adequate for evidence synthesis. The timeframe itself isn't the methodological flaw. Option C misses the point—12 studies can provide reliable meta-analysis results depending on their quality and sample sizes. The number alone isn't problematic. Option D sounds reasonable but misses the core issue. While methodological assessment is important, the problem isn't when quality was assessed, but that quality concerns were invented post-hoc to justify excluding inconvenient results.
Remember this pattern: In systematic review and meta-analysis questions, always check whether researchers stuck to their original protocol. Post-hoc changes to inclusion criteria, especially ones that seem to favor positive results, are major red flags for selection bias.
Question 18
A systematic review includes 12 studies investigating a treatment effect. A funnel plot shows that smaller studies with non-significant results appear to be missing from the literature, while larger studies show a range of results. What statistical approach would be most appropriate to quantitatively assess this observation?
- Cochran's Q test for heterogeneity
- Egger's test for publication bias (correct answer)
- Forest plot with confidence intervals
- I² statistic for inconsistency
- Begg's rank correlation test
Explanation: When you encounter questions about missing studies in systematic reviews, you're dealing with publication bias detection. The key clue here is "smaller studies with non-significant results appear to be missing" — this classic pattern suggests that negative results from small studies weren't published, creating an asymmetric funnel plot.
Egger's test (B) is specifically designed to quantitatively detect this asymmetry in funnel plots. It uses linear regression to test whether smaller studies (with larger standard errors) show systematically different effect sizes than larger studies. When publication bias exists, you'll see a significant relationship between study precision and effect size, indicating missing studies.
Cochran's Q test (A) measures whether studies show more variation in results than expected by chance alone — it detects heterogeneity between studies, not missing studies. The I² statistic (D) quantifies the same heterogeneity as a percentage, telling you what proportion of variation comes from true differences rather than sampling error. Both are about existing studies disagreeing with each other, not about studies being absent. Forest plots (C) are visualization tools that display individual study results and overall effects, but they don't quantitatively assess publication bias.
Remember this pattern: when you see "missing studies," "asymmetric funnel plot," or "small studies with negative results absent," think publication bias detection. Egger's test is the gold standard quantitative approach, while Begg's test is an alternative. Don't confuse publication bias (missing studies) with heterogeneity (existing studies showing different results).
Question 19
In a genome-wide association study (GWAS), researchers test 500,000 SNPs for association with a disease phenotype. Using a standard significance threshold of p < 0.05, they would expect approximately how many false positive associations by chance alone?
- 25 false positives
- 2,500 false positives
- 25,000 false positives (correct answer)
- 250,000 false positives
- 500 false positives
Explanation: When you encounter GWAS multiple testing problems, you're dealing with the fundamental concept that testing many hypotheses simultaneously inflates your chance of false discoveries. With a p < 0.05 threshold, you expect 5% of tests to show "significant" results purely by random chance, even when no true associations exist.
The calculation is straightforward: multiply the total number of tests by your significance threshold. With 500,000 SNPs tested at p < 0.05, you'd expect 500,000×0.05=25,000 false positive associations by chance alone.
This confirms answer C (25,000 false positives) is correct.
Answer A (25 false positives) reflects a major miscalculation—perhaps confusing the number of tests or the significance threshold. This would only be correct if testing 500 SNPs, not 500,000. Answer B (2,500 false positives) suggests using either 50,000 tests instead of 500,000, or a 0.005 significance level instead of 0.05. Answer D (250,000 false positives) implies a 50% false positive rate rather than 5%, which would mean p < 0.5—an absurdly lenient threshold.
Remember this key principle: in any multiple testing scenario, multiply the number of tests by your alpha level to estimate false positives under the null hypothesis. This is why GWAS studies use stringent corrections like Bonferroni (p < 0.05/500,000 ≈ 1×10⁻⁷) or false discovery rate methods—without them, you'd be drowning in false discoveries. Question 20
In a randomized controlled trial, researchers pre-specify that they will analyze outcomes using intention-to-treat (ITT) analysis. After data collection, they find non-significant results with ITT (p = 0.12) but significant results with per-protocol analysis (p = 0.03). They report both analyses but emphasize the per-protocol results in their conclusion. What is the primary concern with this approach?
- Per-protocol analysis is never appropriate in randomized trials
- The emphasis contradicts the pre-specified primary analysis method (correct answer)
- Both analyses should use the same statistical test
- The sample size was likely inadequate for per-protocol analysis
- ITT analysis should only be used for non-inferiority trials
Explanation: When you encounter questions about analysis choices in clinical trials, focus on the principle of scientific integrity and pre-specification. The key issue here is deviation from the predetermined statistical analysis plan without proper justification.
The researchers violated a fundamental principle by emphasizing per-protocol results when they had pre-specified intention-to-treat (ITT) as their primary analysis. This represents a form of selective reporting that undermines the trial's credibility. When you pre-specify your analysis method, you're making a commitment to follow that approach regardless of whether it yields the results you hope for. Switching emphasis to a secondary analysis because it shows significant results constitutes a form of data mining or "p-hacking."
Let's examine why the other options miss the mark. Choice A is incorrect because per-protocol analysis can be appropriate as a secondary or sensitivity analysis in randomized trials - it's just not the gold standard for primary analysis. Choice C misses the point entirely; different analytical approaches can legitimately use different statistical tests depending on their assumptions and data handling. Choice D incorrectly assumes the issue is statistical power when the real problem is methodological integrity.
The ITT analysis preserves the benefits of randomization and provides a more conservative, clinically relevant estimate of treatment effectiveness in real-world conditions, while per-protocol analysis is more susceptible to bias from differential dropout patterns.
Study tip: Remember that pre-specification is sacred in clinical research. When you see questions about conflicting analysis results, ask yourself whether the researchers stuck to their original plan or went "fishing" for significant results.