Historical Context & Motivation
Every statistical model rests on a scaffold of assumptions—conditions that must hold for the analysis to produce valid inferences. Throughout the history of biostatistics, failures to document those assumptions have led to retracted papers, ineffective therapies, and eroded public trust. The practice of documenting assumptions and limitations evolved as the field recognized that transparent reporting is not an optional courtesy but a scientific obligation. Understanding how this norm developed helps explain why modern journals, regulatory agencies, and funding bodies now require explicit statements of analytic assumptions, data limitations, and potential sources of bias.
Early biostatistical work in the nineteenth and early twentieth centuries often presented results with little discussion of the conditions under which they held. Ronald A. Fisher's foundational work in the 1920s formalized assumptions like normality and independence, yet even Fisher seldom devoted space to discussing when those assumptions might fail. As biostatistics matured through the mid-twentieth century—spurred by landmark clinical trials and growing regulatory oversight—researchers began to appreciate that unstated assumptions could silently undermine entire studies.
This historical arc reveals a clear trajectory: from an era in which assumptions were implicit and limitations unstated, to the modern expectation that every biostatistical analysis must include a forthright account of what was assumed, what could go wrong, and how bias may have shaped the results. The central question this lesson addresses is: How do we systematically identify, categorize, and communicate the assumptions, limitations, and biases inherent in any biostatistical analysis?
Core Principles & Definitions
Before you can document assumptions and limitations effectively, you need a precise vocabulary. Three concepts sit at the core of this practice: assumptions, limitations, and sources of bias. Although they are related, each term captures a distinct dimension of analytic uncertainty. Understanding the differences—and the relationships—between them is essential for producing credible biostatistical reports.
Assumptions
Limitations
Sources of Bias
Transparency Principle
Sensitivity Analysis
Visual Explanation — The Assumption–Bias–Limitation Framework
The diagram below illustrates how assumptions, limitations, and sources of bias interact within a biostatistical analysis pipeline. Data collection, analytic modeling, and interpretation each introduce distinct categories of assumptions that, when violated, produce different forms of bias or limitation. Tracing these connections visually helps you construct a comprehensive documentation checklist for any study.
Notice that the framework is organized around the three stages where assumptions enter an analysis. At the data collection stage, the analyst assumes the sample is representative, measurements are accurate, and follow-up is complete. Violations produce selection bias or information bias. At the analytic modeling stage, assumptions about distributional form, independence, and linearity determine whether parameter estimates and their standard errors are valid. Finally, at the interpretation stage, assumptions about causality and generalizability shape whether conclusions are warranted. Effective documentation addresses all three stages systematically.
Quantifying the Impact of Assumption Violations
While documenting assumptions is fundamentally a qualitative reporting practice, biostatisticians frequently quantify the potential impact of assumption violations through mathematical sensitivity analysis. Understanding the formulaic underpinnings helps you appreciate why certain assumptions matter more than others—and it equips you to provide numerical evidence in your documentation rather than vague disclaimers.
Bias Due to Unmeasured Confounding
One of the most common assumption violations in observational biostatistics is the presence of unmeasured confounding. The bias formula for confounding provides a way to express how far an observed association might deviate from the true causal effect when a confounder is not controlled.
E-Value for Sensitivity to Unmeasured Confounding
The E-value, introduced by VanderWeele and Ding in 2017, provides a more elegant approach. It quantifies the minimum strength of association (on the risk ratio scale) that an unmeasured confounder would need to have with both the treatment and the outcome to fully explain away an observed association.
Quantifying Misclassification Bias
Taxonomy of Assumptions, Limitations & Biases
To document assumptions and limitations comprehensively, it helps to work from a structured taxonomy rather than relying on memory or intuition alone. The table below classifies the most common assumptions, their associated biases or consequences when violated, and the documentation strategies recommended for each. This classification covers the three pipeline stages identified in Section 3—data collection, analytic modeling, and interpretation—and adds a fourth dimension for reporting and communication.
| Category | Specific Assumption | Consequence of Violation | Documentation Strategy |
|---|---|---|---|
| Sampling | Random or probability-based sampling from target population | Selection bias; limited external validity | Describe sampling frame, inclusion/exclusion criteria, and population to which results apply |
| Measurement | Exposure and outcome measured without error (non-differential or differential) | Information bias (misclassification); attenuation or exaggeration of effect | Report measurement validity (sensitivity, specificity); quantify expected bias direction |
| Missing Data | Data are missing completely at random (MCAR) or missing at random (MAR) | Biased estimates if data are missing not at random (MNAR); loss of power | Report missingness patterns; justify chosen imputation method; conduct complete-case sensitivity analysis |
| Distributional | Outcome or residuals follow a specified distribution (e.g., normal, Poisson) | Invalid standard errors, p-values, and confidence intervals | Present diagnostic plots (Q-Q, residual plots); use robust methods if violated |
| Structural | Linear (or specified functional) relationship between predictors and outcome | Model misspecification; biased coefficient estimates | Test linearity with splines or polynomial terms; report residual-vs-fitted plots |
| Independence | Observations are independent of one another | Underestimated standard errors; inflated type I error | Acknowledge clustering (e.g., patients within hospitals); use mixed-effects or GEE models |
| No Unmeasured Confounding | All relevant confounders are measured and adjusted for | Residual confounding; spurious or masked associations | List potential unmeasured confounders; compute E-value; conduct negative-control analyses |
| Causal Inference | Study design supports causal interpretation (e.g., exchangeability, positivity, consistency) | Overstatement of causal claims; policy errors | Specify causal framework (DAG, potential outcomes); distinguish association from causation |
Worked Example — Drafting an Assumptions & Limitations Section
Consider a hypothetical observational cohort study examining whether daily aspirin use is associated with reduced risk of colorectal cancer among adults aged 50–75. The study uses electronic health records from a single hospital system, follows patients for five years, and applies a multivariable Cox proportional hazards model to estimate the hazard ratio. Below, we walk through the systematic construction of an assumptions and limitations section using the four-layer framework.
Best Practices vs. Common Pitfalls
Not all assumptions-and-limitations sections are created equal. Many published papers include perfunctory disclaimers—generic sentences that acknowledge limitations without specificity—while others provide richly detailed, actionable documentation that enhances the paper's credibility. The table below contrasts best practices with the most frequently encountered pitfalls, drawn from systematic reviews of published biostatistical work.
| Best Practice | Common Pitfall |
|---|---|
| Name each assumption explicitly and cite the statistical test or diagnostic used to assess it (e.g., 'We tested normality with the Shapiro-Wilk test, p = 0.41'). | Vaguely state 'standard assumptions were met' without specifying which assumptions or how they were evaluated. |
| Describe the expected direction and magnitude of bias when an assumption is likely violated (e.g., 'Non-differential misclassification would bias the OR toward the null'). | Acknowledge a limitation exists but omit any discussion of its likely impact on results (e.g., 'Measurement error may be present'). |
| Report quantitative sensitivity analyses such as E-values, bias-adjusted estimates, or results under alternative assumptions. | Provide no quantitative evidence about robustness; rely entirely on qualitative statements. |
| Distinguish between limitations that affect internal validity (bias, confounding) and those that affect external validity (generalizability). | Lump all limitations together without distinguishing their implications for different aspects of validity. |
| Discuss how limitations inform the interpretation of results—qualifying claims where uncertainty is highest. | Isolate the limitations section from the discussion; conclusions are stated without referencing acknowledged weaknesses. |
| Acknowledge strengths alongside limitations to provide balanced context for the reader. | Either omit strengths entirely or use them to dismiss limitations ('Despite these limitations, our large sample size ensures validity'). |
Connection to Advanced Methods in Bias Analysis
Documenting assumptions and limitations is the qualitative foundation upon which more sophisticated quantitative bias analysis (QBA) methods are built. As you advance in biostatistics, you will encounter formal frameworks that transform qualitative limitations into probabilistic adjustments—allowing you to present bias-corrected estimates alongside conventional results. Understanding the progression from basic documentation to advanced QBA methods clarifies why rigorous documentation matters at every level of expertise.
| Feature | Basic Documentation (This Lesson) | Advanced QBA Methods |
|---|---|---|
| Goal | Transparently identify and describe assumptions, limitations, and bias sources | Quantify the magnitude and direction of bias; produce bias-adjusted effect estimates |
| Output | Narrative limitations section with E-values and qualitative assessment | Bias-adjusted point estimates, confidence intervals incorporating systematic error, probabilistic distributions of bias parameters |
| Techniques | Assumption enumeration, diagnostic tests (Schoenfeld, Q-Q), E-value computation | Monte Carlo sensitivity analysis, Bayesian bias adjustment, multiple bias modeling, external validation |
| Prerequisite Knowledge | Foundational biostatistics: regression, study design, basic epidemiology | Bayesian inference, simulation methods, causal inference frameworks (DAGs, potential outcomes) |
| Key Reference | STROBE, CONSORT checklists; VanderWeele & Ding (2017) for E-values | Lash, Fox, & Fink, Applying Quantitative Bias Analysis to Epidemiologic Data (2009, 2nd ed. 2021) |
The key insight is that basic documentation and advanced QBA are not competing approaches but successive rungs on the same ladder. Every quantitative bias analysis begins with the exact skills taught in this lesson—identifying the relevant assumptions, assessing their plausibility, and specifying the direction of potential bias. As you encounter courses in advanced epidemiologic methods or Bayesian statistics, you will find that the quality of your quantitative bias analysis is only as good as the quality of your qualitative documentation. A Monte Carlo simulation that draws bias parameters from poorly justified prior distributions is no better than a vague limitation statement. Rigorous documentation, therefore, is not merely a beginner's exercise but the enduring foundation of credible statistical communication.
Practice Problems
Lesson Summary
Documenting assumptions, limitations, and sources of bias is a foundational skill in biostatistical communication that transforms raw analytic output into credible, actionable science. This lesson established that assumptions are the conditions required for valid inference (such as normality, independence, and linearity), limitations are inherent design constraints (such as small sample size or restricted generalizability), and biases are systematic errors including selection bias, information bias, and confounding. The four-layer framework (data collection → analytic model → interpretation → reporting) provides a systematic checklist for comprehensive documentation.
Quantitative tools such as the E-value and misclassification bias formulas convert qualitative concerns into numerical evidence about robustness. Best practices call for naming each assumption explicitly, describing the likely direction and magnitude of any bias, reporting sensitivity analyses, and distinguishing between threats to internal validity and external validity. Reporting guidelines like CONSORT and STROBE codify these expectations. Ultimately, transparent documentation is not an admission of weakness but a hallmark of scientific integrity—it builds the trust upon which evidence-based medicine depends.