Biostatistics Quiz: Tables And Figures For Reporting
20 questions · exam conditions
0:00
Tables And Figures For ReportingQuestion 1 of 20

A survival analysis comparing two cancer treatments shows median survival times of 24 months (Treatment A) and 18 months (Treatment B). When presenting these results graphically, what approach provides the most appropriate information for clinical interpretation?

Bar chart showing median survival times with error bars representing standard errors of the medians
Kaplan-Meier curves with confidence bands, number at risk tables, and log-rank test results prominently displayed
Kaplan-Meier curves with confidence bands, number at risk tables below the plot, and hazard ratio with confidence interval in the figure caption
Box plots showing the distribution of survival times for each treatment group with overlaid individual data points
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: Tables And Figures For Reporting

Practice Tables And Figures For Reporting in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Tables And Figures For Reporting, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A survival analysis comparing two cancer treatments shows median survival times of 24 months (Treatment A) and 18 months (Treatment B). When presenting these results graphically, what approach provides the most appropriate information for clinical interpretation?

  1. Bar chart showing median survival times with error bars representing standard errors of the medians
  2. Kaplan-Meier curves with confidence bands, number at risk tables, and log-rank test results prominently displayed
  3. Kaplan-Meier curves with confidence bands, number at risk tables below the plot, and hazard ratio with confidence interval in the figure caption (correct answer)
  4. Box plots showing the distribution of survival times for each treatment group with overlaid individual data points
Explanation: Survival data requires Kaplan-Meier curves with confidence bands and number at risk tables. The hazard ratio should be in the caption, not prominently on the plot to avoid overemphasizing statistical significance. Choice A loses time-to-event information. Choice B overemphasizes p-values. Choice D doesn't account for censoring in survival data.

Question 2

When reporting genetic association results from a genome-wide association study (GWAS), a researcher wants to present findings that meet genome-wide significance (p < 5 × 10⁻⁸). The study identified 12 such associations. What tabular presentation would be MOST informative for the scientific community?

  1. A simple table listing SNP identifiers, chromosomal locations, p-values, and associated genes only
  2. A detailed table including SNP information, effect sizes, allele frequencies, and functional annotations (correct answer)
  3. A summary table grouping SNPs by chromosome with counts of significant associations per chromosome
  4. A table focusing on p-values and Bonferroni-corrected significance levels for conservative interpretation
  5. A table showing only novel associations not previously reported in the literature
Explanation: When evaluating GWAS results for scientific publication, you need to consider what information readers require to interpret, replicate, and build upon your findings. Genome-wide significant associations represent potentially important biological discoveries that warrant comprehensive characterization. Option B provides the most complete scientific value because it includes all essential components for meaningful interpretation. SNP information and chromosomal locations enable replication studies. Effect sizes (odds ratios or beta coefficients) quantify the magnitude of association, which is crucial for understanding biological relevance and clinical significance. Allele frequencies help readers assess generalizability across populations and understand why certain associations might not replicate in different ethnic groups. Functional annotations connect statistical findings to biological mechanisms, helping prioritize variants for follow-up studies. Option A falls short because it lacks effect sizes and functional context, making it difficult for readers to assess biological importance or plan subsequent research. Option C provides only summary statistics without the granular detail needed for scientific evaluation or replication - chromosome-level grouping doesn't advance understanding of individual associations. Option D focuses unnecessarily on statistical corrections when genome-wide significance threshold (p<5×108p < 5 \times 10^{-8}) already accounts for multiple testing more appropriately than Bonferroni correction in GWAS contexts. For GWAS questions, remember that the scientific community values reproducibility and biological insight. Tables should provide enough detail for other researchers to understand not just what you found, but how significant the effects are and where to look for functional mechanisms.

Question 3

A researcher wants to present baseline characteristics of study participants stratified by treatment group in a clinical trial. The study includes 200 participants with variables including age, gender, BMI, and smoking status. Which approach would be MOST appropriate for organizing this information?

  1. A table with variables in rows and treatment groups in columns, including p-values for between-group comparisons
  2. A table with variables in rows and treatment groups in columns, without statistical comparisons (correct answer)
  3. A series of separate figures showing distributions for each variable by treatment group
  4. A table with treatment groups in rows and variables in columns, including confidence intervals
  5. A combined table and figure showing both summary statistics and visual distributions
Explanation: When presenting baseline characteristics in clinical trials, you're creating what's called a "Table 1" - a foundational summary that shows how participant characteristics are distributed across treatment groups. The primary purpose is descriptive: to help readers understand your study population and assess whether randomization created balanced groups. The most appropriate format uses variables in rows (age, gender, BMI, smoking status) and treatment groups in columns, presenting summary statistics without statistical comparisons. This creates a clean, readable table that efficiently displays means, standard deviations, counts, and percentages for each characteristic by group. Option A includes p-values for between-group comparisons, which is problematic because baseline characteristics should be balanced by design through proper randomization. Including p-values suggests you're testing whether randomization "worked," but statistical significance at baseline often reflects random variation rather than meaningful imbalance. This can mislead readers and is discouraged by major medical journals. Option C using separate figures wastes space and makes comparisons difficult. Baseline characteristics are efficiently presented in tabular format, reserving figures for outcome data or complex distributions. Option D reverses the conventional layout with treatment groups in rows and variables in columns, creating an awkward, hard-to-read format. Confidence intervals aren't standard for baseline characteristics since you're describing your sample, not making inferences about populations. Study tip: Remember that Table 1 is purely descriptive. Save statistical testing and confidence intervals for your outcome analyses - baseline tables should let the data speak for itself about group balance.

Question 4

When reporting survival data from a clinical trial comparing two treatments, a researcher creates a Kaplan-Meier curve. What additional element is MOST critical to include to ensure proper interpretation of the results?

  1. Individual patient data points plotted on the curve for complete transparency
  2. Number at risk table showing participants remaining at specific time intervals (correct answer)
  3. Confidence bands around each survival curve for uncertainty quantification
  4. Hazard ratio values displayed directly on the figure for immediate comparison
  5. Median survival times annotated with arrows pointing to curve intersections
Explanation: When interpreting Kaplan-Meier survival curves, you need to understand not just the survival probabilities, but also how reliable those estimates are at different time points. The key insight is that survival estimates become less reliable as time progresses because fewer patients remain in the study. Answer B is correct because a number at risk table is essential for proper interpretation. This table shows how many participants remain under observation at each time interval, allowing readers to assess the reliability of the survival estimates. When only a few patients remain at later time points, the survival curve becomes much less precise and potentially misleading. Without this information, readers cannot distinguish between a curve based on 100 patients versus one based on 5 patients at a given time point. Answer A is wrong because individual patient data points would create visual clutter without adding meaningful interpretive value. The stepped nature of Kaplan-Meier curves already incorporates all relevant individual events. Answer C, while statistically valuable, is less critical than the number at risk table. Confidence bands show uncertainty but don't explain why that uncertainty exists or varies over time. Answer D is incorrect because hazard ratios are summary statistics typically reported in text or tables, not displayed on the survival curves themselves. They don't address the fundamental issue of time-varying reliability of the estimates. Remember: In survival analysis, always look for information about sample size over time. The most beautiful survival curve is meaningless if it's based on only a handful of patients at later time points.

Question 5

A study reports odds ratios for multiple risk factors in a logistic regression model. The researcher wants to create a visual display that effectively communicates both the point estimates and their precision. Which graphical approach would be MOST appropriate?

  1. A bar chart showing odds ratio values with error bars representing standard errors
  2. A forest plot displaying odds ratios with horizontal lines representing confidence intervals (correct answer)
  3. A scatter plot with odds ratios on the y-axis and p-values on the x-axis
  4. A table with odds ratios, confidence intervals, and p-values in separate columns
  5. A box plot showing the distribution of odds ratios across different risk factors
Explanation: When you encounter questions about visualizing results from logistic regression models, think about how to best communicate both the magnitude of effects and their statistical uncertainty to your audience. A forest plot (Option B) is the gold standard for displaying odds ratios from multiple risk factors. This specialized visualization places odds ratios as points on a logarithmic scale, with horizontal lines extending to show confidence intervals. The vertical reference line at OR = 1.0 makes it immediately clear which factors increase risk (points to the right) versus decrease risk (points to the left). The width of each confidence interval instantly communicates precision – narrow intervals indicate more precise estimates. Option A creates problems because odds ratios have a skewed distribution, making standard errors misleading when displayed as symmetric error bars. Plus, bar charts imply that the "area" of each bar has meaning, which isn't true for odds ratios. Option C (scatter plot with p-values) focuses too heavily on statistical significance rather than clinical significance. P-values don't directly tell you about effect size or precision, and this format makes it difficult to compare the magnitude of different risk factors. Option D, while containing all necessary information, isn't a graphical display at all – it's just a table. Tables require readers to mentally process numbers rather than providing the immediate visual comparison that graphs offer. Study tip: Remember that forest plots are specifically designed for displaying effect estimates with confidence intervals. Whenever you see odds ratios, relative risks, or hazard ratios from multiple comparisons, forest plots should come to mind as the preferred visualization method.

Question 6

In a table presenting results from a randomized controlled trial, a researcher includes 95% confidence intervals for the primary outcome. The confidence interval for the treatment effect is (0.15, 0.85) with p = 0.023. How should this information be interpreted and potentially modified for clearer communication?

  1. The result is statistically significant and no modification is needed since the confidence interval excludes 1.0
  2. The confidence interval should be changed to 99% to provide more conservative estimates of precision
  3. The p-value should be removed since the confidence interval provides sufficient information about significance
  4. Both the confidence interval and p-value should be reported, but the clinical significance should be discussed (correct answer)
  5. The confidence interval should be replaced with a standard error to reduce reader confusion
Explanation: When interpreting clinical trial results, you need to consider both statistical significance and clinical meaningfulness. A confidence interval tells you the range of plausible values for the true treatment effect, while the p-value indicates whether the result could reasonably be due to chance. The correct answer is D because comprehensive reporting requires both statistical measures AND discussion of clinical relevance. The confidence interval (0.15, 0.85) suggests the true treatment effect likely falls within this range, and since it doesn't include 1.0 (assuming this is a relative risk or odds ratio), it indicates statistical significance. However, statistical significance doesn't automatically mean clinical importance. A relative risk of 0.15 represents a much more clinically meaningful effect than 0.85, even though both are statistically significant. Option A is incomplete because excluding the null value only confirms statistical significance, not clinical importance. Option B incorrectly suggests changing to 99% confidence intervals - while these would be more conservative, 95% is the standard and changing the interval doesn't improve interpretation without justification. Option C wrongly recommends removing the p-value; both confidence intervals and p-values provide complementary information and are typically reported together in clinical trials. Remember that in biostatistics, statistical significance is just the first step. Always ask: "Is this difference large enough to matter clinically?" Look for questions that test whether you can distinguish between statistical and clinical significance - this is a key concept in evidence-based medicine.

Question 7

A researcher analyzing longitudinal data wants to show how a biomarker changes over time in two treatment groups. The study has measurements at baseline, 3 months, 6 months, and 12 months with substantial dropout over time. What approach would BEST communicate both the temporal pattern and the limitation due to missing data?

  1. A line graph with mean values connected by lines, excluding participants with missing data
  2. A line graph showing individual patient trajectories for all participants with complete data only
  3. A line graph with means and error bars, plus a table showing sample sizes at each timepoint (correct answer)
  4. Separate scatter plots for each timepoint showing all available data without connecting lines
  5. A table presenting means and standard deviations at each timepoint without graphical display
Explanation: When analyzing longitudinal data with dropout, you need to balance clear communication of temporal trends with transparent reporting of data limitations. This is crucial in biostatistics because missing data can bias results and mislead readers about the reliability of findings. Option C provides the optimal approach because it accomplishes both goals simultaneously. The line graph with means and error bars clearly shows the temporal pattern of biomarker changes in both treatment groups, allowing easy comparison of trends over time. The accompanying table with sample sizes at each timepoint transparently communicates how much data underlies each mean estimate, letting readers judge the reliability of later timepoints where dropout has occurred. Option A fails because excluding participants with any missing data wastes valuable information and can introduce bias—you'd only analyze the "completers" who may differ systematically from those who dropped out. Option B has the same fundamental flaw of complete-case analysis, plus individual trajectories become visually cluttered and harder to interpret when showing group-level trends. Option D avoids the missing data problem but fails to show the longitudinal nature of the data—separate scatter plots don't communicate how the biomarker changes over time within individuals. Remember this principle for longitudinal studies: use all available data at each timepoint while being transparent about sample sizes. This approach maximizes statistical power while maintaining scientific integrity. Look for visualization strategies that both communicate the main finding and acknowledge data limitations—this dual transparency is essential in clinical research.

Question 8

When creating a table to report adverse events in a clinical trial, a researcher must decide how to handle events that occurred in fewer than 5% of participants in either group. Which approach represents the BEST practice for transparent reporting?

  1. Exclude rare events entirely to focus on clinically relevant findings and reduce table complexity
  2. Group all rare events together under a single 'Other adverse events' category with total counts
  3. Report all individual events but use footnotes to identify those occurring in <5% of participants
  4. Create a separate supplementary table for rare events while summarizing them in the main table (correct answer)
  5. Report individual rare events only if they are considered medically important by the investigators
Explanation: When reporting adverse events in clinical trials, you're balancing transparency with readability. Regulatory guidelines and good scientific practice require complete disclosure of safety data, but presenting it in an accessible way is equally important. Option D represents best practice because it achieves both goals: complete transparency and clear communication. By creating a supplementary table for rare events while summarizing them in the main table, you provide all the detailed information researchers and regulators need while keeping your primary results table focused and digestible. This approach is standard in high-quality journals and meets regulatory requirements for comprehensive safety reporting. Option A is problematic because excluding rare events entirely violates transparency principles. Even uncommon adverse events can be clinically significant or may signal safety concerns that become apparent when data from multiple studies are combined in meta-analyses. Option B loses important specificity by lumping different rare events together. A reader can't distinguish between different types of adverse events, which is crucial for safety assessment. Individual events might have different clinical implications even if they're all rare. Option C attempts transparency but creates a cluttered main table that's difficult to interpret. While footnotes help identify rare events, having too much detail in the primary table undermines its effectiveness for communicating key findings to readers. Remember this pattern: in biostatistics reporting, the gold standard is often a tiered approach—detailed data available for those who need it, but primary tables optimized for clear communication of main findings.

Question 9

A meta-analysis includes 12 studies examining the effect of a dietary intervention on cholesterol levels. The studies vary considerably in sample size (ranging from 50 to 1,200 participants) and follow-up duration (3 months to 2 years). What visualization approach would BEST represent both the individual study results and the overall pooled estimate?

  1. A forest plot with studies ordered chronologically, showing individual effect sizes and the pooled estimate
  2. A forest plot with studies ordered by sample size, with point sizes proportional to study weight (correct answer)
  3. A scatter plot showing effect size versus sample size with a regression line through the data points
  4. Separate forest plots for different follow-up durations to account for temporal heterogeneity
  5. A funnel plot displaying effect sizes versus standard errors to assess publication bias
Explanation: When evaluating meta-analysis visualizations, you need to consider how well the approach displays individual study contributions, the overall pooled result, and the relative importance of each study in the analysis. A forest plot ordered by sample size with point sizes proportional to study weight (B) is the optimal choice because it immediately shows you which studies carry more influence in the pooled estimate. In meta-analysis, larger studies and those with more precise estimates receive greater weight in calculating the overall effect. When point sizes reflect these weights and studies are ordered by sample size, you can quickly assess whether the pooled result is driven by a few large studies or represents a consensus across multiple smaller ones. This ordering also helps identify potential small-study effects or publication bias patterns. Option A fails because chronological ordering doesn't provide meaningful insight into study quality or influence—publication date rarely correlates with methodological rigor or statistical power. Option C is problematic because a scatter plot with regression line suggests you're trying to model a relationship between effect size and sample size, rather than pooling estimates. This approach also loses the confidence intervals that are crucial for interpreting individual study precision. Option D unnecessarily fragments the analysis—while temporal heterogeneity might warrant subgroup analysis, creating separate plots makes it harder to see the overall pattern and may not be justified without evidence of time-dependent effects. Study tip: In meta-analysis questions, always prioritize visualizations that emphasize study weights and precision. The goal is showing both individual contributions and how they combine into the pooled estimate.

Question 10

A case-control study examining risk factors for lung cancer includes 500 cases and 500 controls. The analysis produces odds ratios for smoking, asbestos exposure, family history, and age. When presenting these results in a table, what approach would BEST address the different nature of these risk factors?

  1. Present all variables in a single table with identical formatting regardless of variable type
  2. Separate tables for categorical variables (smoking, asbestos, family history) and continuous variables (age) (correct answer)
  3. Group variables by strength of association, with strongest predictors presented first
  4. Present crude odds ratios in one table and adjusted odds ratios in a separate table
  5. Use different reference categories for each variable to optimize the clinical interpretation
Explanation: When presenting results from epidemiological studies, the key consideration is how different variable types require different presentation approaches to maximize clarity and interpretability for readers. Option B is correct because categorical and continuous variables fundamentally differ in how their associations are measured and interpreted. Categorical variables like smoking status (yes/no), asbestos exposure (exposed/unexposed), and family history (present/absent) are naturally presented with odds ratios comparing defined groups. Continuous variables like age require different handling—they might be categorized into age groups, presented per unit increase, or per standard deviation increase. Separating these variable types allows you to use appropriate formatting, confidence intervals, and reference categories for each type without forcing awkward compromises. Option A fails because identical formatting ignores the fundamental differences between variable types, making the table harder to interpret and potentially misleading. Option C is problematic because grouping by association strength obscures the logical organization by variable type and doesn't address the presentation challenges posed by mixing categorical and continuous variables. Option D misses the point entirely—while crude versus adjusted analyses is an important distinction, it doesn't address the core issue of how different variable types should be presented. Study tip: When reviewing epidemiological tables, always ask yourself whether the presentation method matches the variable type. Categorical variables need clear reference groups, while continuous variables need clear units of measurement. Good table design should make these distinctions obvious to readers.

Question 11

A researcher is creating a figure to show the relationship between two continuous biomarkers measured in 150 patients, with the goal of demonstrating their correlation and identifying potential outliers. The correlation coefficient is r = 0.72. What additional element would MOST enhance the figure's scientific value?

  1. A fitted regression line with the equation displayed to show the predictive relationship
  2. Different colors or symbols to distinguish patient subgroups based on relevant clinical characteristics (correct answer)
  3. Marginal histograms showing the distribution of each biomarker along the axes
  4. Confidence ellipses around the data points to show the uncertainty in the correlation
  5. A logarithmic transformation of both axes to improve the linearity of the relationship
Explanation: When evaluating scientific figures, especially scatterplots showing biomarker relationships, you should prioritize what adds the most interpretive value and clinical relevance to the data visualization. Option B is correct because distinguishing patient subgroups based on clinical characteristics (like disease stage, treatment response, or demographic factors) transforms a simple correlation display into a clinically meaningful analysis. This approach allows you to see whether the correlation holds across different patient populations, reveals potential confounding variables, and provides actionable insights for clinical practice. With r = 0.72 showing a strong correlation, understanding how this relationship varies by clinically relevant subgroups is crucial for interpretation and application. Option A, while statistically informative, adds limited value since the correlation coefficient already quantifies the linear relationship strength. The regression equation would be redundant given the stated purpose of demonstrating correlation rather than prediction. Option C (marginal histograms) provides useful distributional information but doesn't enhance understanding of the relationship between the biomarkers or add clinical context. This is more of a descriptive enhancement than a scientifically valuable addition. Option D misrepresents statistical concepts—confidence ellipses don't show "uncertainty in correlation" but rather data distribution patterns. True correlation uncertainty would be shown through confidence intervals for the correlation coefficient, not ellipses around data points. Study tip: In biostatistics figures, always prioritize elements that add clinical context and interpretive depth over purely decorative statistical additions. Ask yourself: "What would help a clinician better understand and apply these results?"

Question 12

In a table reporting laboratory values before and after treatment, a researcher notes that some values are below the limit of detection (LOD). The LOD varies by analyte: cytokine A (LOD = 0.5 pg/mL), cytokine B (LOD = 2.0 pg/mL). What approach would be MOST appropriate for handling and reporting these values?

  1. Replace all values below LOD with zero and note this approach in the table footnote
  2. Replace values below LOD with LOD/2 and clearly indicate this substitution method
  3. Exclude all participants with any values below LOD from the analysis and table
  4. Report the number of values below LOD for each analyte and use appropriate statistical methods (correct answer)
  5. Set all values below LOD to the LOD value and treat them as exact measurements
Explanation: When analyzing laboratory data with values below the limit of detection (LOD), you're dealing with left-censored data – a common challenge in biostatistics where you know a value exists but can't measure it precisely. The most appropriate approach is to transparently report the extent of the problem and use statistical methods designed for censored data. Answer D correctly recognizes that you should document how many values fell below the LOD for each analyte and then apply appropriate analytical techniques like survival analysis methods, maximum likelihood estimation, or non-parametric tests that can handle censored observations. Here's why the other approaches are problematic: Answer A (replacing with zero) artificially creates false precision and introduces bias by assuming undetectable means absent, which isn't necessarily true. Answer B (using LOD/2) is a common but flawed substitution method that creates pseudo-data and can lead to incorrect statistical inferences, especially when a large proportion of values are below the LOD. Answer C (excluding participants) reduces your sample size and statistical power while potentially introducing selection bias, since the pattern of missing values might not be random. Different analytes having different LODs (0.5 vs 2.0 pg/mL in this case) makes substitution methods particularly problematic because you'd be making different assumptions about the underlying data distribution for each analyte. Study tip: When you encounter LOD questions on biostatistics exams, remember that transparency and appropriate statistical methodology trump artificial data manipulation. Look for answers that acknowledge the limitation rather than trying to "fix" it with substitution.

Question 13

A systematic review includes studies that report outcomes using different scales: some use continuous scores (0-100), others use categorical outcomes (improved/not improved), and others use time-to-event measures. What presentation strategy would BEST communicate these diverse findings?

  1. Convert all outcomes to standardized mean differences for uniform presentation in a single forest plot
  2. Present separate forest plots for each outcome type with appropriate effect measures for each (correct answer)
  3. Focus only on studies with continuous outcomes since these provide the most precise estimates
  4. Create a narrative summary table describing the direction of effects without quantitative synthesis
  5. Convert all outcomes to p-values and present a forest plot of significance levels
Explanation: When you encounter a systematic review with diverse outcome measures, you're dealing with one of the fundamental challenges in meta-analysis: how to appropriately synthesize different types of data while preserving their unique characteristics and interpretability. The best approach is to present separate forest plots for each outcome type with their most appropriate effect measures (Answer B). Continuous outcomes should use mean differences or standardized mean differences, categorical outcomes should use odds ratios or risk ratios, and time-to-event data should use hazard ratios. This respects the nature of each data type and allows readers to interpret effects in clinically meaningful units. Answer A is problematic because you cannot convert categorical or time-to-event outcomes to standardized mean differences – these are fundamentally different types of measurements that require different statistical approaches. Answer C wastes valuable information by excluding studies with different but valid outcome measures, reducing both the comprehensiveness and generalizability of your review. Answer D abandons the quantitative power of meta-analysis entirely, when appropriate statistical synthesis is actually possible with the right approach. The key insight is that different outcome types aren't a barrier to systematic review – they just require thoughtful organization and appropriate statistical methods for each type. Study tip: Remember that in systematic reviews, preservation of data integrity trumps visual uniformity. Multiple forest plots with appropriate effect measures provide more valid and interpretable results than forcing disparate data types into a single analytical framework.

Question 14

A clinical trial reports time-to-event outcomes using both Kaplan-Meier curves and a Cox proportional hazards model. The hazard ratio is 0.68 (95% CI: 0.52-0.89) with p = 0.005. In presenting these results, what potential issue should be addressed to ensure appropriate interpretation?

  1. The confidence interval should be widened to account for multiple interim analyses during the trial
  2. Verification that the proportional hazards assumption is satisfied should be reported or tested (correct answer)
  3. The hazard ratio should be converted to a relative risk for easier clinical interpretation
  4. Median survival times should replace hazard ratios since they are more clinically meaningful
  5. The p-value should be adjusted for the correlation between the log-rank test and Cox regression
Explanation: When you encounter survival analysis results using Cox proportional hazards models, always consider whether the underlying statistical assumptions have been met. The Cox model relies on a critical assumption called proportional hazards, which means the hazard ratio between groups remains constant over time. Option B is correct because the proportional hazards assumption must be verified for valid interpretation of the hazard ratio. If this assumption is violated, the reported hazard ratio of 0.68 may not accurately represent the treatment effect throughout the study period. Common methods to test this include examining Schoenfeld residuals, testing for interaction with time, or visual inspection of log-minus-log plots. Without this verification, readers cannot confidently interpret the results. Option A is incorrect because confidence interval adjustment for interim analyses would have been handled during the trial design phase through methods like alpha spending functions, not post-hoc widening. Option C misunderstands the difference between hazard ratios and relative risks—they measure different quantities and conversion isn't straightforward in survival analysis. Hazard ratios reflect instantaneous risk at any given time, while relative risks compare cumulative probabilities. Option D incorrectly suggests median survival times are always more meaningful. Hazard ratios are actually preferred because they utilize the entire survival curve and handle censored data better than simple median comparisons. Study tip: For biostatistics questions involving regression models (Cox, logistic, linear), always consider whether key assumptions have been tested. Assumption checking is a fundamental part of proper statistical reporting that examiners frequently test.

Question 15

A researcher analyzing health survey data wants to present demographic characteristics of 5,000 respondents. The data include age, income, education level, employment status, and health insurance type. Some variables have substantial missing data (up to 15% for income). What table design would MOST effectively communicate this information?

  1. Present complete case analysis results only, noting the final sample size in the table title
  2. Include a column showing the number of valid responses for each variable alongside descriptive statistics (correct answer)
  3. Create separate tables for variables with and without substantial missing data
  4. Use multiple imputation for missing values and present results without mentioning missing data
  5. Present percentages based on total sample size, treating missing values as a separate category
Explanation: When presenting demographic data from large-scale surveys, transparency about data quality is essential for proper interpretation. Missing data is a common reality in survey research, and how you present it affects whether readers can assess the reliability and generalizability of your findings. Option B is correct because including a column with the number of valid responses serves multiple critical functions. It allows readers to immediately see which variables have complete versus incomplete data, assess whether missing data patterns might introduce bias, and determine if sample sizes are adequate for drawing conclusions about specific characteristics. This approach maintains transparency while keeping all information in one comprehensive table. Option A is problematic because complete case analysis would dramatically reduce your sample size when multiple variables have missing data. With 15% missing income data alone, you'd lose substantial power and potentially introduce selection bias by excluding respondents systematically. Option C creates unnecessary fragmentation that makes it harder to see relationships between variables and compare sample sizes across characteristics. It also implies that variables with missing data are somehow less important or reliable. Option D is misleading and potentially unethical. While multiple imputation can be appropriate for analysis, presenting imputed results without acknowledging the imputation process violates principles of scientific transparency and prevents readers from making informed judgments about data quality. Remember: In biostatistics, transparency about data limitations builds credibility rather than undermining it. Readers need to know both what you found and how much confidence they should place in those findings based on data completeness.

Question 16

In a cost-effectiveness analysis comparing three medical interventions, a researcher wants to present both costs and health outcomes (quality-adjusted life years, QALYs) in a way that facilitates decision-making. The analysis includes uncertainty estimates for both cost and effectiveness measures. Which presentation approach would be MOST useful for healthcare decision-makers?

  1. A table showing costs and QALYs separately with confidence intervals for each measure
  2. A cost-effectiveness plane scatter plot showing the joint uncertainty in costs and effects
  3. A table with incremental cost-effectiveness ratios (ICERs) and cost-effectiveness acceptability curves
  4. A combination of summary table with ICERs and a cost-effectiveness plane showing uncertainty (correct answer)
  5. A decision tree diagram showing all possible cost and outcome combinations
Explanation: When analyzing cost-effectiveness of medical interventions, decision-makers need both summary measures for quick comparison and detailed uncertainty information to assess confidence in their choices. This requires multiple complementary presentation formats. Option D provides the most comprehensive approach by combining summary tables with ICERs (incremental cost-effectiveness ratios) for easy comparison alongside cost-effectiveness planes that visualize the joint uncertainty in both costs and effects. The summary table gives decision-makers the key numbers they need—how much extra it costs per additional QALY gained—while the cost-effectiveness plane shows whether that estimate is reliable by displaying the scatter of possible cost-effect combinations from uncertainty analysis. Option A fails because presenting costs and QALYs separately doesn't show their relationship or provide the cost-per-QALY ratios that decision-makers actually use. Even with confidence intervals, you can't easily determine cost-effectiveness. Option B provides excellent visualization of uncertainty through the cost-effectiveness plane but lacks the summary statistics (ICERs) that decision-makers need for practical comparison between interventions. Option C gives you the key summary measures and acceptability curves showing the probability each intervention is cost-effective at different willingness-to-pay thresholds, but misses the intuitive visual representation of uncertainty that planes provide. For biostatistics questions about presenting complex analyses, remember that healthcare decision-making typically requires both summary statistics for quick comparison and uncertainty visualization for confidence assessment. Look for answers that combine multiple complementary presentation methods rather than relying on a single approach.

Question 17

In reporting results from a dose-response study with four dose levels (0, 10, 20, 40 mg) of an investigational drug, a researcher wants to present both the raw response data and evidence for a dose-response relationship. Which combination of table and figure would be MOST informative?

  1. A table with means and standard deviations by dose group, paired with a bar chart showing the same means
  2. A table with detailed descriptive statistics by dose, paired with a line graph showing means connected by lines
  3. A correlation table showing pairwise comparisons between doses, paired with a scatter plot of individual responses
  4. A table with summary statistics and trend test results, paired with a box plot showing distributions by dose (correct answer)
  5. A frequency table showing response categories by dose, paired with a stacked bar chart of proportions
Explanation: When evaluating dose-response relationships, you need to accomplish two key goals: present the raw data clearly and provide statistical evidence that a true dose-response trend exists. This requires both descriptive summaries and formal trend analysis. Option D is correct because it combines comprehensive summary statistics with trend test results in the table, paired with box plots that show the complete distribution of responses at each dose level. The trend test (like a test for linear trend) provides crucial statistical evidence that the observed pattern isn't due to chance, while box plots reveal the full data distribution, including medians, quartiles, outliers, and variability at each dose. This combination gives readers both the statistical rigor and visual clarity needed to evaluate dose-response relationships. Option A fails because bar charts and means alone don't show the underlying data distributions or provide statistical evidence of a true dose-response relationship. Option B includes helpful trend visualization with connected means, but line graphs can be misleading since they imply continuous dosing between discrete dose levels, and this option lacks formal statistical testing for trends. Option C focuses on pairwise comparisons rather than the overall dose-response trend, and scatter plots of individual responses don't clearly organize data by dose groups. For biostatistics exam success, remember that dose-response studies require both descriptive presentation of data distributions and formal statistical testing for trends. Look for combinations that include trend analysis (not just group comparisons) and visualizations that preserve information about data variability and distribution shape.

Question 18

A multi-center study reports results from 8 different sites with varying sample sizes (ranging from 45 to 320 participants per site). When presenting overall results, what approach would BEST acknowledge the multi-center design while providing clinically relevant summary estimates?

  1. Report results from each center separately in individual tables without overall summary
  2. Present pooled results with random effects meta-analysis to account for between-center heterogeneity (correct answer)
  3. Weight all centers equally regardless of sample size to avoid bias toward larger centers
  4. Report only results from the three largest centers to ensure adequate statistical power
  5. Use fixed effects meta-analysis assuming all centers represent the same underlying population
Explanation: When you encounter multi-center studies in biostatistics, you're dealing with hierarchical data where participants are nested within centers. This creates potential correlation within centers and heterogeneity between centers that must be properly addressed in your analysis. The best approach is random effects meta-analysis (option B) because it treats each center as contributing separate but related estimates, then combines them while accounting for both within-center sampling error and between-center variation. This method acknowledges that centers may have genuine differences due to population characteristics, local practices, or other factors, while still providing a meaningful overall estimate with appropriate confidence intervals. Option A fails because reporting only separate results wastes the opportunity to provide clinically useful summary estimates and offers no overall conclusions. Option C is statistically flawed because equal weighting ignores the reality that larger samples provide more precise estimates - this approach would actually increase uncertainty in your pooled estimate unnecessarily. Option D throws away valuable data from smaller centers and introduces selection bias, potentially making results less generalizable to diverse clinical settings. The random effects approach properly weights centers by their precision (larger centers get more weight) while incorporating between-center heterogeneity into the uncertainty estimates. This produces summary results that are both statistically valid and clinically interpretable. Study tip: In multi-center study questions, look for approaches that acknowledge the hierarchical structure. Random effects models are usually preferred over fixed effects when you want results generalizable beyond the specific centers studied.

Question 19

A pharmaceutical company is preparing to submit results from a Phase III randomized controlled trial comparing a new antihypertensive drug to placebo. The primary endpoint is change in systolic blood pressure from baseline to 12 weeks. Secondary endpoints include change in diastolic blood pressure, proportion of patients achieving target blood pressure (<140/90 mmHg), and adverse events. The study enrolled 400 participants (200 per group).

For the primary efficacy results table, which elements should be included to meet regulatory reporting standards while ensuring clinical interpretability?

  1. Baseline and endpoint means with standard deviations, mean change, and p-values for within-group changes
  2. Mean change from baseline with confidence intervals, between-group difference with confidence interval, and p-value (correct answer)
  3. Baseline and endpoint medians with interquartile ranges, given the potential for non-normal distributions
  4. Individual patient data presented in a detailed appendix with summary statistics in the main table
  5. Effect size calculations with Cohen's d values to facilitate comparison with other antihypertensive studies
Explanation: When you encounter questions about reporting clinical trial results for regulatory submission, focus on what provides both statistical rigor and clinical interpretability while meeting standard reporting requirements. The correct approach (B) includes the essential elements for primary efficacy reporting: mean change from baseline with confidence intervals shows the treatment effect within each group, the between-group difference with its confidence interval demonstrates the comparative effectiveness, and the p-value provides statistical significance testing. This combination allows clinicians to assess both the magnitude of effect (clinically meaningful?) and statistical certainty, while meeting FDA and other regulatory expectations. Option A falls short because it emphasizes within-group changes rather than the crucial between-group comparison. While baseline and endpoint data are useful, within-group p-values don't address the key question of whether the drug outperforms placebo. The focus should be on comparative effectiveness. Option C suggests using medians and interquartile ranges, but blood pressure changes typically follow approximately normal distributions, making means more appropriate and expected by regulatory standards. Additionally, this approach doesn't specify the critical between-group statistical comparisons. Option D proposes including individual patient data in the main results, which would be overwhelming and inappropriate for a primary results table. While individual data might belong in appendices, summary statistics should dominate the main presentation for clarity and regulatory compliance. Remember: For Phase III trial reporting, always prioritize between-group comparisons over within-group changes, and include both effect size estimates (with confidence intervals) and significance testing to meet regulatory standards.

Question 20

When presenting longitudinal data from a 12-month clinical trial measuring depression scores at baseline, 3, 6, 9, and 12 months for two treatment groups, which graphical approach best handles missing data and dropout patterns?

  1. Line graph with mean ± SEM at each timepoint, noting sample sizes at each visit in the figure legend, and complete case analysis
  2. Line graph showing individual patient trajectories with different line colors for completers vs. dropouts, overlaid with group mean trajectories
  3. Line graph with means and confidence intervals, sample sizes displayed below x-axis at each timepoint, with last-observation-carried-forward imputation
  4. Spaghetti plot showing all individual trajectories, with group means highlighted and dropout patterns indicated by line termination points (correct answer)
Explanation: Spaghetti plots reveal individual variation and dropout patterns while showing group trends, providing transparency about missing data. Option A hides dropout information. Option B may be cluttered and LOCF (option C) can bias results. Option D best shows the complexity of longitudinal data including missingness patterns.