Biostatistics Quiz: Writing Statistical Results
20 questions · exam conditions
0:00
Writing Statistical ResultsQuestion 1 of 20

A confidence interval for the difference in mean cholesterol levels between treatment and control groups is reported as (2.4, 18.7) mg/dL. How should this result be interpreted in a research report?

The treatment group had significantly higher cholesterol levels with 95% certainty that the true difference lies between 2.4 and 18.7 mg/dL.
We are 95% confident that the true difference in mean cholesterol levels between groups is between 2.4 and 18.7 mg/dL.
There is a 95% probability that future samples will show differences between 2.4 and 18.7 mg/dL in cholesterol levels.
The observed difference in cholesterol levels falls within the expected range of 2.4 to 18.7 mg/dL with 95% reliability.
95% of individuals in the treatment group will have cholesterol differences between 2.4 and 18.7 mg/dL compared to controls.
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: Writing Statistical Results

Practice Writing Statistical Results in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Writing Statistical Results, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A confidence interval for the difference in mean cholesterol levels between treatment and control groups is reported as (2.4, 18.7) mg/dL. How should this result be interpreted in a research report?

  1. The treatment group had significantly higher cholesterol levels with 95% certainty that the true difference lies between 2.4 and 18.7 mg/dL.
  2. We are 95% confident that the true difference in mean cholesterol levels between groups is between 2.4 and 18.7 mg/dL. (correct answer)
  3. There is a 95% probability that future samples will show differences between 2.4 and 18.7 mg/dL in cholesterol levels.
  4. The observed difference in cholesterol levels falls within the expected range of 2.4 to 18.7 mg/dL with 95% reliability.
  5. 95% of individuals in the treatment group will have cholesterol differences between 2.4 and 18.7 mg/dL compared to controls.
Explanation: When you encounter confidence interval interpretation questions, remember that these intervals make statements about population parameters, not individual samples or future observations. The confidence interval (2.4, 18.7) mg/dL represents our best estimate of where the true population difference in mean cholesterol levels lies. Since both bounds are positive, we can infer that the treatment group likely has higher cholesterol levels than the control group, with the true difference falling somewhere in this range. Answer B correctly interprets this as a statement about our confidence in where the true population parameter (the difference in means) lies. The 95% confidence level means that if we repeated this study many times with the same methodology, 95% of the resulting confidence intervals would contain the true population difference. Answer A incorrectly uses the phrase "95% certainty," which implies probability about the parameter itself. The parameter is fixed; our interval either contains it or doesn't. The 95% refers to our confidence in the method, not certainty about this specific interval. Answer C misinterprets the interval as predicting future sample results. Confidence intervals describe population parameters, not future sample outcomes, which would require prediction intervals instead. Answer D confuses the confidence interval with an expected range for observed differences. The interval estimates the true population difference, not the range where we expect future observed differences to fall. Study tip: Remember that confidence intervals always make statements about population parameters, and the confidence level refers to the long-run performance of the method, not the probability that this specific interval contains the parameter.

Question 2

A clinical trial reports: 'The mean reduction in systolic blood pressure was 12.3 mmHg (SD = 4.8) in the treatment group compared to 3.1 mmHg (SD = 4.2) in the placebo group (p < 0.001).' What additional information is most critical for complete reporting?

  1. The correlation coefficient between baseline blood pressure and treatment response in both groups.
  2. The confidence interval for the difference between group means and the sample sizes. (correct answer)
  3. The median blood pressure values and interquartile ranges for both treatment groups.
  4. The effect size calculation and power analysis results from the original study design.
  5. The individual patient responses and range of blood pressure changes observed.
Explanation: When evaluating statistical reporting in clinical trials, you should always ask: "Do I have enough information to understand the magnitude of the effect and how confident I can be in these results?" The reported means, standard deviations, and p-value tell you there's a statistically significant difference, but they don't give you the complete picture needed for clinical interpretation. Answer B is correct because confidence intervals and sample sizes are essential for complete reporting. The confidence interval for the difference between means (12.3 - 3.1 = 9.2 mmHg difference) would show you the range of plausible values for the true treatment effect, giving you crucial information about both the precision of the estimate and its clinical significance. Sample sizes are equally critical because they affect the reliability of your results and help readers assess whether the study was adequately powered. Answer A is incorrect because correlation between baseline values and response, while potentially interesting, isn't required for basic statistical reporting and doesn't address the completeness of the main comparison. Answer C focuses on median and IQR, which would be more relevant if you suspected non-normal distributions, but the study already reports means and standard deviations appropriately. Answer D mentions effect size and power analysis, which are valuable but secondary to the fundamental need for confidence intervals and sample sizes that directly relate to the reported comparison. Remember: whenever you see means compared between groups, immediately look for confidence intervals around the difference and sample sizes. These transform a simple "statistically significant" finding into clinically interpretable evidence.

Question 3

A researcher writes: 'The treatment was effective (p = 0.04) with a mean improvement of 2.1 points on the depression scale.' What is the primary issue with this statistical reporting?

  1. The p-value is too close to 0.05 to definitively conclude effectiveness of the treatment intervention.
  2. Statistical significance does not necessarily indicate clinical effectiveness without context about meaningful change. (correct answer)
  3. The direction of improvement is unclear since depression scale scoring methods were not specified.
  4. The sample size and confidence intervals should be reported before drawing conclusions about effectiveness.
  5. A one-tailed test should have been used instead of a two-tailed test for effectiveness studies.
Explanation: When you encounter statistical reporting in research, you need to distinguish between statistical significance and clinical significance—two fundamentally different concepts that are often confused. The correct answer is B because statistical significance (p = 0.04) only tells you that the observed difference is unlikely due to chance. However, it says nothing about whether that difference is meaningful in real-world clinical practice. A 2.1-point improvement might be statistically significant but clinically trivial if, for example, the depression scale ranges from 0-100 and typical meaningful changes require 10+ points. Without knowing what constitutes a clinically meaningful difference on this particular scale, you cannot conclude the treatment is "effective." Answer A is wrong because p = 0.04 is definitively below the conventional 0.05 threshold—there's no ambiguity about statistical significance here. Answer C misses the main issue; while scale direction matters for interpretation, the primary problem is the leap from statistical to clinical significance. Answer D identifies important missing information, but sample size and confidence intervals wouldn't resolve the fundamental confusion between statistical and clinical significance. Remember this key principle: statistical significance answers "Is this difference real?" while clinical significance answers "Does this difference matter to patients?" A p-value can never tell you about clinical effectiveness—you need additional context about what size of change is meaningful in practice. Always be skeptical when researchers jump from p-values to effectiveness claims without discussing clinical significance thresholds.

Question 4

A researcher reports a 95% confidence interval for mean weight loss as (1.2, 4.8) kg. A colleague suggests this should be written as: 'Participants lost 3.0 ± 1.8 kg (95% CI).' What is wrong with the colleague's suggestion?

  1. The confidence interval should be reported as (1.2, 4.8) kg rather than using the plus-minus notation format.
  2. The calculation is incorrect because 3.0 ± 1.8 gives (1.2, 4.8), but this assumes a symmetric distribution.
  3. The plus-minus notation suggests a standard error rather than a confidence interval for the population mean.
  4. The colleague's format incorrectly implies that 3.0 kg is the true population mean rather than the sample mean.
  5. The notation 3.0 ± 1.8 represents margin of error, not the confidence interval bounds for this estimate. (correct answer)
Explanation: When interpreting confidence intervals, you need to understand what different notation formats communicate about statistical uncertainty and what parameter is being estimated. The colleague's suggestion of "3.0 ± 1.8 kg (95% CI)" is mathematically correct—it does yield the interval (1.2, 4.8). However, this plus-minus format creates confusion about what type of uncertainty is being reported. In statistical writing, the ± notation is typically reserved for standard errors or standard deviations, not confidence intervals. When readers see ±1.8, they naturally interpret this as one standard error, suggesting they could multiply by 1.96 to get the 95% confidence interval. This misrepresents the actual statistical calculation that produced the interval. Looking at the wrong answers: Choice A incorrectly suggests that confidence intervals should never use plus-minus notation—while not preferred, it's not inherently wrong. Choice B misunderstands the issue by focusing on distribution symmetry, but the calculation itself is mathematically sound regardless of the underlying distribution shape. Choice D misses the mark by suggesting confusion about sample versus population means, but both formats are clearly reporting sample statistics as estimates of population parameters. The core problem is that the plus-minus notation misleadingly suggests standard error reporting rather than a confidence interval, potentially confusing readers about the type of uncertainty measure being presented. Study tip: Always report confidence intervals in interval notation (a, b) to avoid confusion with standard errors. Reserve ± notation for standard deviations and standard errors only.

Question 5

A study comparing three pain medications reports: 'ANOVA revealed significant differences among groups (F(2,87) = 4.32, p = 0.017).' What important limitation should be noted in interpreting this result?

  1. The F-statistic is too small to indicate practically meaningful differences among the three medication groups.
  2. The degrees of freedom suggest the sample size is too small for reliable comparisons among groups.
  3. Post-hoc testing is needed to determine which specific pairs of medications differ significantly from each other. (correct answer)
  4. The p-value indicates only a weak level of statistical significance for the overall comparison.
  5. ANOVA assumes equal variances, but this assumption likely cannot be met with medication effectiveness data.
Explanation: When you encounter ANOVA results in biostatistics, remember that ANOVA is an omnibus test - it tells you whether there are any significant differences among your groups, but not which specific groups differ from each other. The reported ANOVA result (F(2,87) = 4.32, p = 0.017) indicates that at least one of the three pain medications differs significantly from the others. However, this leaves several possibilities: maybe medication A differs from B but not C, maybe all three differ from each other, or maybe only one medication differs from the other two. Without post-hoc testing (like Tukey's HSD or Bonferroni corrections), you cannot determine which specific pairwise comparisons are driving the significant omnibus result. This is why answer C correctly identifies the key limitation. Answer A is incorrect because the F-statistic's magnitude doesn't directly indicate practical significance - that depends on effect sizes and clinical relevance, not the F-value itself. Answer B misinterprets the degrees of freedom: with df = (2,87), this represents 90 total participants across three groups, which is generally adequate for ANOVA. Answer D incorrectly characterizes p = 0.017 as "weak" significance - this p-value is clearly below the conventional α = 0.05 threshold and represents solid statistical significance. Study tip: Whenever you see significant ANOVA results with more than two groups, immediately think "post-hoc testing needed." ANOVA answers the question "Are there differences?" but post-hoc tests answer "Where are the differences?"

Question 6

A regression analysis yields: 'For each additional year of age, blood pressure increases by 0.8 mmHg (95% CI: 0.3, 1.3; p = 0.003).' How should the confidence interval be interpreted?

  1. We can be 95% certain that individual patients will experience blood pressure increases between 0.3 and 1.3 mmHg per year.
  2. There is a 95% probability that the true population slope coefficient lies between 0.3 and 1.3 mmHg per year.
  3. We are 95% confident that the true relationship between age and blood pressure is between 0.3 and 1.3 mmHg per year. (correct answer)
  4. 95% of the observed data points fall within 0.3 to 1.3 mmHg of the predicted regression line.
  5. Future studies will find age-related blood pressure increases between 0.3 and 1.3 mmHg with 95% probability.
Explanation: When you encounter confidence intervals for regression coefficients, you're dealing with statistical inference about the true population relationship between variables, not predictions about individual cases or data distribution. The confidence interval (0.3, 1.3) represents our uncertainty about the true slope coefficient in the population. Since we calculated this from a sample, our estimate of 0.8 mmHg per year has sampling variability. The 95% confidence interval tells us the plausible range for the true population parameter. Answer C correctly captures this - we're 95% confident the true relationship (slope coefficient) between age and blood pressure lies between 0.3 and 1.3 mmHg per year. Answer A is wrong because confidence intervals don't predict individual patient outcomes - they estimate population parameters. Individual predictions would require prediction intervals, which account for both parameter uncertainty and individual variability around the regression line. Answer B uses problematic language by saying "probability that the true slope lies between..." The true slope either is or isn't in this range - we can't assign probability to a fixed parameter. The confidence level refers to our method's long-run performance, not the probability for this specific interval. Answer D confuses confidence intervals with the distribution of residuals around the regression line. This describes how data points scatter around the fitted line, which is unrelated to our uncertainty about the slope coefficient. Study tip: Remember that confidence intervals for regression coefficients always estimate uncertainty about the true population relationship, never individual predictions or data scatter. Focus on the parameter being estimated when interpreting any confidence interval.

Question 7

A researcher writes: 'The correlation between vitamin D levels and bone density was statistically significant (r = 0.34, p < 0.05).' What additional information would strengthen this statistical reporting?

  1. The coefficient of determination and confidence interval for the correlation coefficient should be included.
  2. The exact p-value rather than 'p < 0.05' and the direction of the correlation should be specified.
  3. The sample size, confidence interval for r, and exact p-value should be provided for complete reporting. (correct answer)
  4. A scatter plot and the linear regression equation should accompany all correlation coefficient reports.
  5. The statistical test used to calculate significance and assumptions that were verified should be documented.
Explanation: When evaluating statistical reporting of correlation coefficients, you need to assess whether the researcher has provided enough information for readers to fully interpret and replicate the findings. Complete statistical reporting requires transparency about sample characteristics, effect size precision, and exact statistical values. Answer C is correct because it identifies the three most critical missing elements. Sample size (n) is essential because correlation coefficients are heavily influenced by sample size - a correlation of 0.34 might be meaningful with n=200 but questionable with n=15. The confidence interval for r shows the precision of the estimate and helps readers understand the range of plausible true correlation values. The exact p-value (like p=0.032) provides more information than just "p < 0.05" and allows for more nuanced interpretation. Answer A focuses on coefficient of determination (r²) and confidence intervals, but r² is often unnecessary for correlation reporting since readers can easily calculate it from r. Answer B mentions the exact p-value and direction, but the direction is already clear from the positive r value (0.34), making this less comprehensive than C. Answer D suggests scatter plots and regression equations, which are helpful for visualization but not required for complete statistical reporting of correlations. Remember this reporting hierarchy: sample size is always critical for interpreting any correlation, confidence intervals show precision, and exact p-values provide more information than threshold reporting. These three elements together allow readers to properly evaluate the strength and reliability of the reported relationship.

Question 8

A study reports: 'Mean survival time was 14.2 months (SD = 6.8) in the experimental group versus 11.7 months (SD = 5.9) in the control group.' What is the most significant reporting deficiency?

  1. Survival data should be reported using median values and interquartile ranges rather than means and standard deviations. (correct answer)
  2. The difference in survival times and its statistical significance testing results are not provided.
  3. Confidence intervals for both group means and the sample sizes should be included in the report.
  4. The hazard ratio and results from Cox proportional hazards analysis should replace descriptive statistics.
  5. Kaplan-Meier survival curves and log-rank test results are the appropriate statistics for survival outcomes.
Explanation: When analyzing survival data, you're dealing with measurements that are typically right-skewed and may include censored observations (patients lost to follow-up or still alive at study end). This fundamental characteristic of survival data determines how it should be summarized and reported. Answer A is correct because survival times rarely follow a normal distribution. They're usually right-skewed, with most patients surviving for shorter periods and a "tail" of long-term survivors. In skewed distributions, the mean is pulled toward the extreme values and doesn't represent the typical patient experience well. The median survival time tells you when 50% of patients are still alive, which is far more clinically meaningful. Interquartile ranges are also more appropriate than standard deviations for non-normal data. Answer B is wrong because while statistical testing results would be valuable, they don't address the fundamental problem of using inappropriate descriptive statistics for survival data. Answer C is wrong because confidence intervals and sample sizes, though important for complete reporting, still don't solve the core issue of using means and standard deviations for skewed survival data. Answer D is wrong because while Cox regression and hazard ratios are important for survival analysis, descriptive statistics are still needed to summarize the data. The issue isn't replacing descriptive statistics entirely, but using the right ones. Key takeaway: Whenever you see survival data reported with means and standard deviations, immediately flag this as inappropriate. Survival data should virtually always use medians and interquartile ranges due to its characteristic right-skewed distribution.

Question 9

A chi-square test yields χ2=8.94,df=3,p=0.030\chi^2 = 8.94, df = 3, p = 0.030. The researcher writes: 'The variables are significantly associated.' What additional reporting would be most valuable?

  1. The exact cell counts and expected frequencies for each category in the contingency table analysis.
  2. A measure of effect size such as Cramér's V and a description of the pattern of association. (correct answer)
  3. Post-hoc comparisons between individual cells to identify which categories drive the significant result.
  4. Verification that assumptions of independence and minimum expected cell counts were satisfied.
  5. The odds ratios for each category comparison and their corresponding confidence intervals.
Explanation: When you encounter a significant chi-square test result, you're only halfway through proper statistical reporting. A significant p-value tells you that an association exists, but it doesn't tell you how strong that association is or what pattern it follows. Answer B is correct because it addresses both critical gaps in the researcher's statement. Cramér's V quantifies the strength of association on a 0-1 scale, giving readers meaningful context about whether this significant result represents a weak, moderate, or strong relationship. With χ2=8.94\chi^2 = 8.94 and df=3df = 3, you can calculate Cramér's V to determine practical significance beyond statistical significance. Equally important is describing the pattern—which categories are over- or under-represented relative to what you'd expect if variables were independent. Answer A provides transparency but doesn't enhance interpretation of the significant finding. Raw data alone doesn't help readers understand the strength or nature of the association. Answer C suggests post-hoc cell comparisons, but chi-square tests don't require multiple comparison corrections like ANOVA does—the overall test already tells you about association patterns. Answer D addresses methodological rigor, which is important for validity but doesn't improve interpretation of an already-conducted analysis. Remember this hierarchy for chi-square reporting: statistical significance first (already established), then effect size and pattern description (missing here), then methodological details if space permits. On biostatistics exams, questions about "most valuable additional reporting" typically prioritize interpretive information that helps readers understand both the magnitude and nature of significant findings.

Question 10

A study reports: 'Treatment A showed superior outcomes with an effect size of d = 0.8.' A reviewer questions this interpretation. What is the most likely concern?

  1. Effect size should be reported as Cohen's d with confidence intervals and sample sizes for proper interpretation.
  2. An effect size of 0.8 is not large enough to support claims of superior outcomes in clinical research.
  3. The term 'superior outcomes' implies clinical significance that cannot be determined from effect size alone. (correct answer)
  4. Cohen's d assumes equal variances between groups, which may not be valid for this comparison.
  5. Effect sizes should be interpreted relative to the specific clinical context rather than general guidelines.
Explanation: When you encounter questions about effect sizes and clinical interpretations, focus on the crucial distinction between statistical significance and clinical significance. These are fundamentally different concepts that researchers and clinicians must carefully separate. The reviewer's most likely concern is that claiming "superior outcomes" implies clinical significance, which cannot be determined from Cohen's d alone. While d = 0.8 indicates a large statistical effect, clinical significance depends on whether this translates to meaningful improvements in patient outcomes, quality of life, or other clinically relevant measures. A statistically large effect might represent a small, clinically irrelevant difference in the actual measured outcome. Let's examine why the other options miss the mark: Option A incorrectly suggests the primary issue is reporting format—while confidence intervals and sample sizes improve interpretation, they don't address the clinical significance problem. Option B is factually wrong since d = 0.8 represents a large effect size by Cohen's conventions (small = 0.2, medium = 0.5, large = 0.8). Option D focuses on a statistical assumption that, while potentially important, isn't the most likely concern a reviewer would raise about claims of "superior outcomes." Remember this key principle: effect size tells you the magnitude of difference between groups, but only clinical expertise and context can determine if that difference matters to patients. When you see researchers making clinical claims based solely on statistical measures, always question whether they've bridged the gap between statistical and clinical significance.

Question 11

A researcher reports: 'The odds ratio for smoking and lung cancer was 12.4 (p < 0.001), confirming that smoking causes lung cancer.' What is the primary statistical reporting error?

  1. The confidence interval for the odds ratio should be provided along with the point estimate.
  2. Odds ratios cannot establish causation; they only quantify the strength of association between variables. (correct answer)
  3. The p-value should be exact rather than reported as an inequality for proper interpretation.
  4. An odds ratio of 12.4 is too large to be credible and suggests methodological problems.
  5. Case-control studies require relative risk calculations rather than odds ratios for cancer research.
Explanation: When you encounter questions about statistical associations and causation, remember that statistical measures can only quantify relationships between variables—they cannot prove that one variable causes another, no matter how strong the association. The researcher's primary error is claiming that the odds ratio establishes causation. An odds ratio of 12.4 means that people who smoke have 12.4 times the odds of developing lung cancer compared to non-smokers. This is a strong association, but association is not causation. Even with a highly significant p-value, this statistical measure alone cannot prove that smoking causes lung cancer. Establishing causation requires additional evidence including temporal relationships, dose-response patterns, biological plausibility, and consideration of confounding variables. Looking at the other options: Choice A is a valid concern since confidence intervals provide important information about precision and statistical significance, but it's not the primary error here. Choice C misunderstands p-value reporting—using "p < 0.001" is perfectly acceptable and often preferred when the exact value is extremely small. Choice D incorrectly suggests that large odds ratios are inherently problematic. While unusually large effect sizes warrant scrutiny, an odds ratio of 12.4 is not implausible for well-established risk factors like smoking and lung cancer. Study tip: Remember the fundamental rule: correlation ≠ causation. No matter how strong an association appears statistically (high odds ratios, low p-values), observational studies can only demonstrate association. Watch for questions that try to trick you into accepting causal language based solely on statistical significance.

Question 12

A multiple regression reports: 'Age was a significant predictor (β = 0.23, p = 0.01) while gender was not (β = 0.08, p = 0.34).' What key information should be added for complete reporting?

  1. The standardized beta coefficients should be reported instead of the unstandardized regression coefficients.
  2. The confidence intervals for both regression coefficients and the overall model R-squared value. (correct answer)
  3. The correlation matrix showing relationships among all predictor variables included in the model.
  4. The variance inflation factors to assess multicollinearity among the predictor variables used.
  5. The residual analysis results and tests of regression assumptions for the fitted model.
Explanation: When evaluating multiple regression reporting, you need to assess whether the results provide enough information for proper interpretation and replication. Complete statistical reporting requires both effect size measures and uncertainty estimates. The correct answer is B because confidence intervals and R-squared are essential missing components. Confidence intervals show the range of plausible values for each coefficient, giving you crucial information about precision and clinical significance that p-values alone cannot provide. The overall model R-squared tells you how much variance in the outcome is explained by all predictors combined, which is fundamental for understanding the model's practical utility. Option A is incorrect because both standardized and unstandardized coefficients have value - the choice depends on your research question. Unstandardized coefficients show real-world effect sizes in original units, while standardized coefficients allow comparison across different scales. Option C is wrong because correlation matrices, while sometimes helpful, aren't essential for basic regression reporting. They're more relevant during exploratory analysis or when examining multicollinearity. Option D is incorrect because variance inflation factors (VIFs) are diagnostic tools for multicollinearity assessment, not standard reporting requirements. You'd typically check VIFs during analysis but wouldn't necessarily report them unless multicollinearity was a specific concern. Remember this reporting hierarchy: significance tests (p-values) tell you if an effect exists, confidence intervals tell you the precision and magnitude, and R-squared tells you overall model performance. All three components together provide the complete statistical story that readers need for proper interpretation.

Question 13

A t-test comparing two groups yields: t(48) = 2.34, p = 0.023. The researcher writes: 'Group A performed significantly better than Group B.' What reporting element is most critically missing?

  1. The direction of the difference cannot be determined from the t-statistic and p-value alone.
  2. The actual group means and standard deviations are needed to interpret the magnitude of difference. (correct answer)
  3. The confidence interval for the difference between means should accompany the significance test.
  4. The effect size should be calculated and reported along with the statistical significance results.
  5. The assumption of equal variances should be tested and reported before interpreting t-test results.
Explanation: When interpreting statistical test results, you must be able to determine what the numbers actually tell you about the real-world difference between groups. The researcher's conclusion about "Group A performing significantly better" cannot be validated from the given information. The t-statistic and p-value tell us there's a statistically significant difference, but they don't reveal which group had higher values. A positive t-statistic (2.34) could mean Group A > Group B or Group B > Group A, depending on how the researcher coded the groups in the analysis. Without seeing the actual means, you cannot determine the direction of the difference, making the researcher's directional claim unjustified. Looking at the wrong answers: (A) is incorrect because the direction actually cannot be determined from a t-statistic alone - this supports why (B) is correct rather than contradicting it. (C) identifies confidence intervals as missing, which would be valuable but isn't the most critical omission when you can't even verify the basic claim about which group performed better. (D) points to missing effect size, which is important for interpreting practical significance, but again, you first need to know the direction of the difference before worrying about its magnitude. For biostatistics questions involving test interpretation, always ask yourself: "Can I verify the researcher's specific claims from the given statistics?" When you see directional conclusions ("A is better than B"), immediately check whether the reported statistics actually support that direction. Means and standard deviations are your foundation for interpreting any between-group comparison.

Question 14

A study found a correlation coefficient of r = 0.72 between hours of exercise per week and cardiovascular fitness scores. Which statement best describes how this finding should be reported?

  1. Exercise causes a 72% improvement in cardiovascular fitness scores among study participants.
  2. There is a strong positive association between weekly exercise hours and fitness scores (r = 0.72). (correct answer)
  3. Weekly exercise hours explain 72% of the variation observed in cardiovascular fitness scores.
  4. Cardiovascular fitness increases by 0.72 units for each additional hour of weekly exercise performed.
  5. The relationship between exercise and fitness demonstrates a 72% probability of clinical significance.
Explanation: When you encounter correlation coefficients in biostatistics, remember that they measure the strength and direction of linear relationships between variables, not causation or specific predictive values. The correlation coefficient r = 0.72 indicates a strong positive association between exercise hours and fitness scores. Since correlations range from -1 to +1, a value of 0.72 represents a strong positive relationship—as one variable increases, the other tends to increase as well. Answer B correctly describes this as "a strong positive association" and properly includes the correlation value for transparency. Let's examine why the other options are incorrect. Answer A claims "exercise causes a 72% improvement," which confuses correlation with causation—correlation studies cannot establish causal relationships, and the 0.72 value doesn't represent a percentage change. Answer C states that "exercise hours explain 72% of the variation," but this misinterprets the correlation coefficient. To find explained variation, you'd calculate r², which would be 0.72² = 0.52 or 52%. Answer D suggests "fitness increases by 0.72 units for each additional exercise hour," which confuses correlation coefficients with regression slopes—these are entirely different statistical measures. Study tip: Remember that correlation coefficients only describe association strength and direction. For causation, you need experimental designs. For percentage of explained variance, square the correlation coefficient (r²). For predicting specific unit changes, you need regression analysis, not correlation.

Question 15

A study comparing medication effectiveness reports: 'Response rates were 45% (18/40) in Group A and 65% (26/40) in Group B, with a significant difference (χ2=3.2,p=0.04\chi^2 = 3.2, p = 0.04).' What should be questioned about this reporting?

  1. The chi-square statistic appears too small given the observed difference in response rates between groups.
  2. Response rates should be compared using Fisher's exact test rather than chi-square with these sample sizes.
  3. The 20% difference in response rates may not represent a clinically meaningful improvement.
  4. Confidence intervals for the difference in proportions and effect size measures should be included. (correct answer)
  5. The sample sizes are equal between groups, which suggests potential bias in the randomization process.
Explanation: When evaluating statistical reporting in biomedical research, complete interpretation requires both statistical significance testing and practical effect assessment. While this study correctly reports the chi-square test results, it's missing crucial information for proper clinical interpretation. Answer D is correct because comprehensive statistical reporting should include confidence intervals and effect size measures. The confidence interval for the 20% difference in proportions would show the range of plausible true differences, helping readers assess both statistical precision and clinical relevance. Effect size measures (like odds ratio or relative risk) provide standardized ways to interpret the magnitude of treatment effects that complement p-values. Option A is incorrect—the chi-square value of 3.2 is actually appropriate for this comparison. With these sample sizes and proportions, this statistic correctly reflects the observed difference. Option B is wrong because chi-square is perfectly acceptable here; Fisher's exact test is typically preferred only when expected cell counts fall below 5, which doesn't occur with these sample sizes (both groups have 40 participants). Option C, while raising a valid clinical consideration, isn't the primary statistical reporting issue—the 20% absolute difference could indeed be clinically meaningful depending on the condition and intervention context. Remember that p-values only tell you if a difference is statistically significant, not how large or clinically important that difference is. Always look for confidence intervals and effect sizes in research reports—their absence is a red flag for incomplete statistical communication, even when significance testing is performed correctly.

Question 16

A study reports: 'The relative risk of developing diabetes was 2.3 (95% CI: 1.4, 3.8) for obese compared to normal-weight individuals.' Which interpretation is most precise?

  1. Obese individuals are 2.3 times more likely to develop diabetes than normal-weight individuals.
  2. The risk of diabetes is 130% higher among obese individuals compared to those with normal weight.
  3. Obese individuals have 2.3 times the risk of developing diabetes compared to normal-weight individuals. (correct answer)
  4. The probability of diabetes increases by a factor of 2.3 when individuals become obese.
  5. For every 2.3 obese individuals with diabetes, 1 normal-weight individual develops the condition.
Explanation: When you encounter relative risk (RR) questions, focus on precise mathematical interpretation rather than colloquial language that can introduce ambiguity. Relative risk represents a ratio: the risk in the exposed group divided by the risk in the unexposed group. An RR of 2.3 means the exposed group (obese individuals) has exactly 2.3 times the risk of the unexposed group (normal-weight individuals). This is a multiplicative relationship, not an additive one. Answer C correctly states this mathematical relationship: obese individuals have 2.3 times the risk compared to normal-weight individuals. This phrasing captures the precise ratio without additional interpretation. Answer A uses "more likely," which can be ambiguous—does this mean 2.3 times as likely or 2.3 times more likely? In statistics, we avoid this ambiguous phrasing. Answer B incorrectly converts to percentage increase. A 130% increase would mean the risk is 2.3 times higher (230% of baseline), but this adds unnecessary conversion and potential confusion. The original RR value is more precise. Answer D suggests causation ("when individuals become obese") and implies the RR measures probability change over time, but relative risk compares risks between groups at a point in time, not probability changes within individuals. Remember: When interpreting relative risk, stick to the mathematical ratio language ("X times the risk") rather than converting to percentages or using potentially ambiguous phrases like "more likely." This maintains statistical precision and avoids common misinterpretations.

Question 17

A study reports: 'No significant difference was found between groups (p = 0.12).' Which follow-up statement would be most appropriate?

  1. This proves that the two treatment approaches are equally effective for all patients in the population.
  2. The study failed to detect a difference, but this does not prove that no difference exists. (correct answer)
  3. There is an 88% probability that the null hypothesis of no difference is correct.
  4. The lack of significance indicates that any observed differences were due to random error.
  5. Additional studies are unnecessary since no treatment effect has been established at this time.
Explanation: When you encounter a non-significant p-value in biostatistics, you're dealing with one of the most commonly misunderstood concepts in statistical inference. The key principle here is that "absence of evidence is not evidence of absence." A p-value of 0.12 means that if there truly were no difference between groups, you would observe results this extreme or more extreme about 12% of the time by random chance alone. Since this exceeds the conventional significance threshold of 0.05, we fail to reject the null hypothesis. However, this doesn't prove the null hypothesis is true. Answer B correctly captures this nuance: failing to detect a difference doesn't prove no difference exists. The study may have been underpowered, the effect size may be small, or there may be other methodological limitations preventing detection of a real difference. Answer A is wrong because statistical tests never "prove" anything about entire populations with certainty. Answer C misinterprets the p-value—it's not the probability that the null hypothesis is correct. The p-value tells us about the probability of observing our data given the null hypothesis, not the probability of the null hypothesis being true. Answer D is incorrect because we cannot definitively attribute observed differences to random error when we fail to find significance; there could still be a real effect we simply didn't detect. Remember: non-significant results indicate insufficient evidence for a difference, not proof of no difference. Always consider study power and effect size when interpreting negative results.

Question 18

A researcher conducted a study comparing mean blood pressure between two groups and obtained p = 0.03. Which of the following statements most appropriately describes this result?

  1. There is a 3% probability that the null hypothesis is true given the observed data.
  2. There is a 97% probability that the alternative hypothesis is correct based on this sample.
  3. The probability of observing a difference this large or larger, assuming no true difference exists, is 0.03. (correct answer)
  4. There is a 3% chance that the study results occurred due to random sampling variation alone.
  5. The clinical significance of the blood pressure difference has been established at the 3% level.
Explanation: When you encounter p-values in biostatistics, you're dealing with one of the most misunderstood concepts in research. The p-value specifically measures the probability of observing your data (or more extreme data) under the assumption that the null hypothesis is true. Answer C correctly defines this concept. With p = 0.03, there's a 3% probability of observing a blood pressure difference this large or larger if there truly is no difference between the groups. This is exactly what p-values measure—the likelihood of your observed results under the null hypothesis. Answer A commits the classic "prosecutor's fallacy" by confusing P(data|null hypothesis) with P(null hypothesis|data). The p-value doesn't tell you the probability that the null hypothesis is true given your data. Answer B makes a similar error—p-values cannot directly tell you the probability that the alternative hypothesis is correct, as this would require prior probabilities and Bayesian reasoning. Answer D mischaracterizes what "random sampling variation" means in this context. While sampling variation is involved, the p-value specifically assumes the null hypothesis is true when calculating this probability. Remember this key distinction: p-values tell you about the compatibility of your data with the null hypothesis, not about the truth of hypotheses themselves. When you see p-value questions, always ask yourself "probability of what, given what?" The answer is always "probability of the observed data or more extreme, given the null hypothesis is true."

Question 19

A study reports: '15 of 60 patients in the treatment group experienced side effects compared to 8 of 55 in the control group (p = 0.08).' How should this result be most appropriately described?

  1. The treatment and control groups showed no significant difference in side effect rates at the 0.05 level.
  2. There was a trend toward higher side effect rates in the treatment group, but this did not reach significance.
  3. The side effect rate was 25% in the treatment group compared to 14.5% in the control group (p = 0.08). (correct answer)
  4. The study was underpowered to detect a significant difference in side effect rates between groups.
  5. No conclusion about side effect differences can be drawn due to the non-significant p-value obtained.
Explanation: When interpreting statistical results, you should focus on presenting the actual data clearly and accurately, including both the descriptive statistics and the p-value without over-interpreting the statistical significance. Answer C correctly presents the essential information: it calculates and reports the actual percentages (15/60 = 25% for treatment, 8/55 = 14.5% for control) and includes the p-value of 0.08. This approach provides readers with both the clinical effect size and the statistical evidence, allowing them to draw their own conclusions about the meaningfulness of the difference. Answer A is incorrect because it makes a definitive claim about "no significant difference" based solely on the p-value exceeding 0.05. While statistically accurate, this black-and-white interpretation ignores the actual magnitude of difference and doesn't present the data itself. Answer B is problematic because it uses vague language like "trend toward higher side effect rates" which is imprecise and doesn't quantify the actual difference observed. The phrase "trend" is often misused in statistics and doesn't add meaningful information. Answer D makes an assumption about statistical power that cannot be determined from the given information. Without knowing the study design, effect size the study was powered to detect, or conducting a formal power analysis, you cannot conclude the study was "underpowered." Remember: The best approach to reporting results is to present the actual data (percentages, rates, means) along with the statistical test results. Let the numbers speak for themselves rather than over-interpreting statistical significance thresholds.

Question 20

A case-control study investigating risk factors for lung cancer included 380 cases and 760 controls. Smoking history showed an odds ratio of 4.2 (95% CI: 2.8-6.3, p<0.001), while family history of cancer had an OR of 1.8 (95% CI: 0.9-3.6, p=0.09). How should these findings be reported in the results section?

  1. Smoking history was significantly associated with lung cancer risk (OR=4.2, 95% CI: 2.8-6.3, p<0.001), while family cancer history showed a trend toward increased risk (OR=1.8, 95% CI: 0.9-3.6, p=0.09).
  2. Current or former smoking was associated with substantially increased lung cancer odds (OR=4.2, 95% CI: 2.8-6.3, p<0.001), whereas family history was not significantly associated with cancer risk (OR=1.8, 95% CI: 0.9-3.6, p=0.09).
  3. Smoking was strongly associated with lung cancer (OR=4.2, 95% CI: 2.8-6.3, p<0.001). Family cancer history was associated with increased odds (OR=1.8, 95% CI: 0.9-3.6) but this did not reach statistical significance (p=0.09). (correct answer)
  4. Participants with smoking history had 4.2 times higher odds of lung cancer (95% CI: 2.8-6.3, p<0.001). Family cancer history showed elevated but non-significant odds (OR=1.8, 95% CI: 0.9-3.6, p=0.09).
Explanation: Choice C appropriately describes associations in case-control studies, provides complete statistical information, and correctly interprets non-significant findings without dismissing them. Choice A inappropriately uses 'trend' language without justification. Choice B dismisses the non-significant finding inappropriately. Choice D uses acceptable language but is less precise in describing the association and uses awkward phrasing for the non-significant result.