BIOSTATISTICS • STUDY DESIGN & DATA

Population, Sample, Parameter & Statistic — Define population, sample, parameter, and statistic in health contexts

Understanding how inference bridges finite observations to the broader health reality they represent.

Historical Context & Motivation

Long before modern clinical trials, physicians and public health officials grappled with a fundamental tension: they needed to make decisions about entire communities—whether a water source was contaminated, whether a treatment was efficacious—yet they could never observe every individual in those communities. The intellectual machinery for navigating that gap evolved over centuries, drawing on political arithmetic, probability theory, and epidemiological fieldwork. The concepts of population, sample, parameter, and statistic crystallized as formalized ideas only in the late nineteenth and early twentieth centuries, yet their roots are far older.

1662
John Graunt's Bills of Mortality
Graunt compiled London's weekly death records into systematic tables, pioneering the idea that a finite set of parish records could reveal regularities about the broader population's mortality patterns—an early instance of using sample data to infer population-level truths.
1854
John Snow & the Broad Street Pump
Snow mapped cholera deaths in a London neighborhood, treating a geographically defined group as a population and interviewing a subset of households as a sample. His work demonstrated how targeted observations could support causal conclusions about disease transmission.
1908
Student's t-Distribution
William Sealy Gosset, publishing as 'Student,' showed that the sampling distribution of the mean differs from the normal curve when samples are small. His work formalized the relationship between sample statistics and population parameters under uncertainty.
1948
The Framingham Heart Study
This landmark cohort study enrolled 5,209 residents of Framingham, Massachusetts, as a sample of the U.S. adult population. Researchers measured sample statistics—blood pressure means, cholesterol levels—and used them to estimate population parameters for cardiovascular risk.
1996
CONSORT Guidelines Published
The Consolidated Standards of Reporting Trials required explicit reporting of target populations, sampling methods, and whether sample statistics could generalize. This codified the population–sample distinction into clinical research standards worldwide.

The recurring challenge across these milestones is the same question biostatistics still asks today: how do we move from what we can observe (a sample and its statistics) to what we want to know (a population and its parameters)? Answering this requires precise definitions of all four terms and a clear understanding of the logical relationship between them.

Core Definitions & Principles

At its core, biostatistical inference rests on a clean conceptual architecture: two types of groups (populations and samples) and two types of numerical summaries (parameters and statistics). Understanding these four pillars is prerequisite to every subsequent topic in biostatistics—from hypothesis testing to regression—because each technique is, in essence, a method for relating sample statistics back to population parameters. The four definitions that follow are complementary; each gains meaning in contrast to the others.

1

Population

The complete set of individuals, observations, or units about which you wish to draw conclusions. In health research, a population might be 'all adults in the United States with type 2 diabetes' or 'every surgical procedure performed at Hospital X in 2024.' Populations can be finite or conceptually infinite (e.g., all future patients receiving a drug).
2

Sample

A subset of the population that is actually observed or measured. Good study design ensures the sample is representative—often through random selection—so that findings can be generalized. In the Framingham Heart Study, the 5,209 enrolled residents constituted a sample of the broader U.S. adult population.
3

Parameter

A fixed numerical summary that describes some characteristic of the population. Parameters are typically unknown and are the targets of inference. The true mean systolic blood pressure of all American adults (μ) is a parameter. Greek letters (μ, σ, π) conventionally denote parameters.
4

Statistic

A numerical summary computed from sample data. Because the sample varies from study to study, a statistic is a random variable. The sample mean blood pressure (x̄) calculated from 500 participants is a statistic. Roman letters (x̄, s, p̂) conventionally denote statistics.
KEY TAKEAWAY
Think of a population parameter as the temperature of an entire lake. You cannot measure every water molecule, so you dip a thermometer into several spots (your sample) and compute an average reading (your statistic). The thermometer readings are not the lake's true temperature—they are estimates of it. If you sampled different spots, you would get slightly different readings, but a well-chosen sampling strategy keeps those readings close to the truth. The relationship between statistic and parameter is exactly the same in health research: sample statistics are our best available estimates of population parameters, and the quality of those estimates depends on how the sample was drawn.

Visual Explanation — From Population to Statistic

The diagram below illustrates the logical flow from the population level to the sample level, showing how parameters and statistics occupy parallel roles. Arrows represent the processes of sampling (moving from population to sample) and inference (moving from sample back to population). This bidirectional relationship is the engine of biostatistical reasoning.

The upper box represents the population—all individuals of interest—which is described by fixed parameters (μ, σ, π). Through sampling (yellow arrow), a subset is drawn. Computed summaries of this subset are statistics (x̄, s, p̂). Through inference (pink arrow), we use those statistics to estimate the unknown parameters.

Notice that the diagram encodes a critical asymmetry: we can always compute a statistic from our sample data, but we can never directly observe the parameter we are trying to estimate. The entire apparatus of confidence intervals, p-values, and Bayesian posteriors exists to quantify the uncertainty introduced by this asymmetry. The clearer you keep the population–sample and parameter–statistic distinctions in your mind, the more naturally the rest of biostatistics will follow.

Mathematical Framework — Notation & Relationships

Biostatistics uses a disciplined notational convention that maps directly onto the conceptual distinction between parameters and statistics. Greek letters are reserved for population parameters; Roman (Latin) letters designate sample statistics. This is not mere typographic convention—it carries logical meaning. Whenever you see μ in an equation, you know the author is referring to a fixed but usually unknown population quantity. Whenever you see x̄, you know it was computed from observed data and could change if a new sample were drawn. The following equations formalize the most common parameter–statistic pairs encountered in health research.

POPULATION MEAN (PARAMETER)
μ = (1/N) × Σᵢ₌₁ᴺ Xᵢ
Where μ is the population mean, N is the total number of individuals in the population, and Xᵢ is the value for the i-th individual. In practice, μ is almost never computable because N is too large or infinite.
SAMPLE MEAN (STATISTIC)
x̄ = (1/n) × Σᵢ₌₁ⁿ xᵢ
Where x̄ is the sample mean, n is the number of sampled observations, and xᵢ is the value for the i-th sampled individual. The sample mean x̄ serves as an estimator of the population parameter μ.
POPULATION PROPORTION (PARAMETER)
π = (number with attribute) / N
Where π (or sometimes P) is the true proportion of the population possessing a characteristic, such as the proportion of all ICU patients who develop ventilator-associated pneumonia.
SAMPLE PROPORTION (STATISTIC)
p̂ = x / n
Where p̂ is the sample proportion, x is the count of sampled individuals with the attribute, and n is the sample size. The hat symbol (^) is a standard indicator that the quantity is an estimate of a parameter.
💡 Notation Tip
A helpful mnemonic: Population → Parameter (both start with P). Sample → Statistic (both start with S). Parameters use Greek letters; statistics use Roman letters. Keep this pairing automatic and you will never confuse μ with x̄ or π with p̂.

Classification — Health Research Examples

The abstract definitions of population, sample, parameter, and statistic take on vivid specificity when embedded in real health research scenarios. The table below maps several common study designs to their corresponding populations, samples, parameters, and statistics. Studying these examples side by side reveals a recurring pattern: the researcher defines a population, draws a sample from it, computes statistics from the sample, and then uses those statistics to make inferences about the population parameters. This pattern holds regardless of whether the study is observational or experimental, cross-sectional or longitudinal.

Examples of population, sample, parameter, and statistic across health study designs
Study ContextPopulationSampleParameterStatistic
Vaccine efficacy RCTAll adults aged 18–65 eligible for vaccination30,000 randomized participantsTrue relative risk reduction (π)Observed efficacy rate (p̂ = 0.95)
Hospital readmission surveyAll patients discharged from Hospital X in 2024400 randomly selected discharge recordsTrue 30-day readmission rate (π)Sample readmission rate (p̂ = 0.12)
National blood pressure studyAll U.S. adults aged ≥ 205,000 participants in NHANES cyclePopulation mean SBP (μ)Sample mean SBP (x̄ = 126 mmHg)
Birth weight cohort studyAll singleton live births in state Y1,200 births at three hospitalsPopulation SD of birth weight (σ)Sample SD (s = 480 g)
Each row pairs a population parameter (left, in Greek) with its corresponding sample statistic (right, in Roman). Dashed lines emphasize the one-to-one mapping. Note the n − 1 denominator in the sample variance, which is Bessel's correction for unbiased estimation.

A subtlety worth noting is the denominator difference between population variance (dividing by N) and sample variance (dividing by n − 1). This adjustment—known as Bessel's correction—ensures that s² is an unbiased estimator of σ². Without it, the sample variance would systematically underestimate population variability because sample points tend to cluster around their own mean rather than the true population mean. This small correction exemplifies how the parameter–statistic distinction has direct computational consequences.

Worked Example — Blood Glucose Screening

A county health department wants to estimate the mean fasting blood glucose level (in mg/dL) of all adults aged 40–60 in the county. The department cannot test every resident, so it recruits a random sample of 200 adults from local clinics. The recorded fasting blood glucose values yield a sample mean of 104 mg/dL and a sample standard deviation of 18 mg/dL. Let us walk through identifying the population, sample, parameter, and statistic, and then compute a 95% confidence interval for the population mean.

Estimating Population Mean Fasting Blood Glucose
1
Step 1 — Define the PopulationThe population is all adults aged 40–60 residing in the county. This is a finite but large group whose exact size (N) may be estimated from census data but whose individual glucose values are not all known.
Population = all county adults aged 40–60
2
Step 2 — Identify the SampleThe sample consists of the 200 adults who were randomly recruited and tested. Random selection is crucial: it supports the generalizability of findings from this subset to the broader population.
Sample: n = 200 randomly selected adults
3
Step 3 — Specify the Parameter of InterestThe parameter is μ, the true mean fasting blood glucose of the entire population. This value is fixed but unknown—it is the target of our investigation.
Parameter: μ (unknown population mean FBG)
4
Step 4 — Compute the Sample StatisticFrom the 200 observed values, we compute x̄ = 104 mg/dL and s = 18 mg/dL. These are statistics—numerical summaries calculated from sample data. If we had drawn a different random sample of 200 adults, x̄ would likely differ slightly.
Statistics: x̄ = 104 mg/dL, s = 18 mg/dL
5
Step 5 — Construct a 95% Confidence Interval for μUsing the formula CI = x̄ ± z* × (s / √n), with z* = 1.96 for 95% confidence: Standard error = 18 / √200 = 18 / 14.14 ≈ 1.273 mg/dL. Margin of error = 1.96 × 1.273 ≈ 2.50 mg/dL. The 95% confidence interval is 104 ± 2.50, or (101.50, 106.50) mg/dL. We are 95% confident that the true population mean fasting blood glucose μ falls within this range.
95% CI for μ: (101.50, 106.50) mg/dL
📌 Interpretation Note
The confidence interval is a statement about the parameter μ, not the statistic x̄. We already know x̄ = 104—it is computed, not uncertain. The uncertainty lies in how well 104 approximates the unknown μ. The CI quantifies that uncertainty by saying: if we repeated this sampling process many times, about 95% of the resulting intervals would contain the true μ.

Strengths, Limitations & Common Pitfalls

The population–sample and parameter–statistic framework is the backbone of inferential statistics, but it is only as reliable as the sampling process that links them. Understanding where the framework shines—and where it can mislead—helps researchers design better studies and interpret results more carefully.

Strengths and limitations of the population–sample–parameter–statistic framework
FeatureStrengthLimitation / Pitfall
GeneralizabilityWhen sampling is random and representative, statistics generalize to population parameters with quantifiable precision.Convenience or volunteer samples may produce biased statistics that do not approximate the true parameter—no amount of statistical technique can fix a fundamentally non-representative sample.
Quantified uncertaintyConfidence intervals and standard errors provide formal measures of how much a statistic might differ from its target parameter.These measures assume correct model assumptions (e.g., normality, independence). Violated assumptions produce misleading uncertainty estimates.
EfficiencySampling is far cheaper and faster than a census. NHANES uses ~5,000 participants to characterize the health of ~260 million U.S. adults.Small samples yield wide confidence intervals and low statistical power, limiting the practical utility of estimates.
Population definitionPrecisely defining the target population focuses the research question and clarifies to whom conclusions apply.Vaguely defined populations (e.g., 'patients') make generalization ambiguous. The target and accessible populations may differ.
Sampling variabilityThe sampling distribution framework allows us to predict how much statistics will vary across repeated samples.Researchers sometimes confuse the standard deviation (variability of individuals) with the standard error (variability of the statistic), leading to incorrect inference.
⚠️ COMMON CONFUSION
One of the most frequent errors in published health research is conflating the target population with the study population. The target population is the group you want to learn about (e.g., all adults with asthma); the study or accessible population is the group from which you can realistically recruit (e.g., asthma patients at three urban clinics). Your sample is drawn from the study population, and your statistics estimate parameters of that study population. Whether those parameters also characterize the target population depends on how similar the study population is to the target population—a judgment that requires domain expertise, not just statistical methods.

Connection to Sampling Distributions & Inferential Theory

The conceptual framework introduced in this lesson is not merely definitional; it is the foundation on which the entire superstructure of inferential statistics rests. Once you recognize that a statistic is a random variable—because it changes with every new sample—you naturally arrive at the concept of the sampling distribution, which is the probability distribution of a statistic over all possible samples of size n. The Central Limit Theorem guarantees that, under mild conditions, the sampling distribution of the sample mean x̄ is approximately normal with mean μ and standard deviation σ/√n, regardless of the shape of the original population distribution. This remarkable result is what makes confidence intervals and hypothesis tests possible.

How this lesson's concepts connect to advanced biostatistical theory
Concept in This LessonAdvanced Extension
Sample statistic x̄ varies across samplesSampling distribution of x̄; standard error = σ/√n
Parameter μ is unknown and fixedBayesian framework treats μ as a random variable with a prior distribution, updated by sample data
Sample is a subset of the populationSampling designs (stratified, cluster, multistage) control how subsets are drawn to optimize precision
Statistic estimates parameterProperties of estimators: unbiasedness, consistency, efficiency, sufficiency
Confidence interval for μHypothesis testing (H₀: μ = μ₀) and p-values as formalized decisions about parameters

As you progress through biostatistics, you will encounter increasingly sophisticated methods—logistic regression, survival analysis, mixed-effects models—but every one of these techniques is performing the same fundamental operation: using sample-level statistics to estimate or test population-level parameters. The vocabulary and logic you are building now will remain the conceptual scaffolding for every method you learn hereafter.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher reports: 'The average hemoglobin A1c of 350 enrolled patients was 7.2%.' Identify the population, the sample, the parameter, and the statistic in this statement. Explain why 7.2% is a statistic rather than a parameter.
PROBLEM 2BASIC CALCULATION
In a sample of 150 emergency department patients, 42 tested positive for influenza. Calculate the sample proportion (p̂) who tested positive. Then identify the corresponding population parameter this estimate targets.
PROBLEM 3INTERMEDIATE
A public health agency surveys 800 randomly selected adults and finds a sample mean body mass index (BMI) of 27.4 kg/m² with a sample standard deviation of 5.1 kg/m². Compute the standard error of the mean and construct a 95% confidence interval for the population mean BMI (μ). Interpret the interval in context.
PROBLEM 4APPLIED
A clinical trial enrolls 500 volunteers from three urban hospitals to test a new antihypertensive drug, with the goal of generalizing results to 'all U.S. adults with stage 1 hypertension.' The trial reports a mean reduction in systolic blood pressure of 8.3 mmHg (sample statistic). Discuss at least two reasons why this statistic might not accurately estimate the corresponding population parameter, and suggest one design modification that could improve generalizability.
PROBLEM 5CRITICAL THINKING
Consider a census scenario in which a hospital audits every one of its 12,000 discharge records from the past year to compute a 30-day readmission rate of 11.4%. Is 11.4% a parameter or a statistic? Now suppose a policy analyst uses this same 11.4% figure to argue about the readmission performance of 'all similar-sized community hospitals in the state.' In this second context, is 11.4% a parameter or a statistic? Justify your reasoning and discuss the implications for inference.

Lesson Summary

This lesson established the four foundational concepts of biostatistical inference. A population is the complete group about which we wish to draw conclusions—for example, all adults in the United States with a given condition. A sample is the subset of that population that is actually observed, ideally through random selection to ensure representativeness. A parameter is a fixed numerical characteristic of the population (denoted by Greek letters such as μ, σ, and π), and a statistic is the corresponding numerical summary computed from sample data (denoted by Roman letters such as x̄, s, and p̂).

The core logic of biostatistics flows in two directions: sampling moves from population to sample, while inference moves from sample back to population. Every confidence interval, hypothesis test, and regression model you will encounter is a formal method for using a statistic to estimate or test a parameter. The quality of that inference depends critically on how the sample was drawn: random and representative samples support valid generalization, while biased samples do not—regardless of sample size or analytical sophistication. Mastering these four terms and their relationships is the essential first step toward rigorous reasoning in health research.

Varsity Tutors • Biostatistics • Population, Sample, Parameter & Statistic