BIOSTATISTICS • DATA METHODS & STATISTICAL COMMUNICATION

Human Subjects & Data Privacy — Human subjects basics and data privacy concepts (overview)

Understanding the ethical and legal foundations that govern how researchers collect, store, and share data involving human participants.

Historical Context & Motivation

The modern framework for protecting human subjects in research did not emerge from abstract philosophical debate; it was forged in response to deeply troubling episodes of exploitation and harm. Throughout the twentieth century, a series of revelations about unethical experimentation forced governments, academic institutions, and professional organizations to develop formal codes of conduct and regulatory structures. Understanding this history is essential for any biostatistician or data scientist, because the statistical methods you apply, the datasets you analyze, and the conclusions you communicate are all shaped by the ethical constraints that arose from these events.

At its core, human subjects research refers to any systematic investigation involving living individuals from whom a researcher obtains data through intervention, interaction, or access to identifiable private information. The concept of data privacy extends this concern into the digital age, addressing how personal health information and research data are collected, stored, shared, and potentially re-identified. Together, these twin pillars — ethical treatment of participants and protection of their data — define the landscape within which modern biostatistical research operates.

1947
The Nuremberg Code
Following the Doctors' Trial at Nuremberg, the international community established ten principles for ethical human experimentation, centering on voluntary informed consent as an absolute requirement. This was the first formal articulation that subjects must consent to participation without coercion.
1964
Declaration of Helsinki
The World Medical Association adopted comprehensive ethical principles for medical research involving human subjects. This declaration introduced the concept that the well-being of the individual research participant must take precedence over all other interests, including those of science and society.
1974
National Research Act & the Belmont Report
Prompted by revelations of the Tuskegee Syphilis Study (1932–1972), the U.S. Congress passed the National Research Act, creating the National Commission for the Protection of Human Subjects. The resulting Belmont Report (1979) codified three core ethical principles: Respect for Persons, Beneficence, and Justice.
1991
The Common Rule (45 CFR 46)
The U.S. federal government adopted a uniform policy across multiple agencies — known as the Common Rule — that standardized the requirements for Institutional Review Boards (IRBs), informed consent procedures, and additional protections for vulnerable populations.
1996–2018
HIPAA and GDPR
The Health Insurance Portability and Accountability Act (HIPAA, 1996) established U.S. standards for protecting health information. The European Union's General Data Protection Regulation (GDPR, 2018) set a global benchmark for data privacy rights, including the right to erasure and data portability.

Each of these milestones responded to a critical gap: how do we ensure that the pursuit of knowledge does not come at the cost of individual dignity, autonomy, and safety? For biostatisticians, this question is not merely philosophical — it directly shapes study design, sampling strategies, data management protocols, and the permissible scope of secondary data analysis. The regulatory infrastructure that emerged from these events provides both a moral compass and a practical framework for responsible research.

Core Principles & Definitions

The ethical framework governing human subjects research rests upon a set of interconnected principles that translate moral philosophy into actionable requirements. These principles apply whether a researcher is conducting a randomized controlled trial, analyzing an existing dataset, or designing a survey instrument. The Belmont Report's three principles — Respect for Persons, Beneficence, and Justice — serve as the foundational pillars, while additional data privacy concepts extend these principles into the realm of information management.

1

Respect for Persons

Individuals must be treated as autonomous agents capable of making their own decisions about participation. This principle mandates informed consent — a process (not merely a document) through which participants receive adequate information, comprehend it, and voluntarily agree to participate. Persons with diminished autonomy (e.g., prisoners, children) are entitled to additional protections.
2

Beneficence

Researchers have a dual obligation: to maximize possible benefits and to minimize possible harms. This translates into rigorous risk–benefit assessment — a systematic evaluation of whether the anticipated knowledge gain justifies the risks to participants. The concept applies to physical, psychological, social, economic, and informational harms.
3

Justice

The burdens and benefits of research must be distributed fairly across populations. Historically marginalized communities should not bear disproportionate research risk, nor should privileged groups monopolize access to research benefits. In biostatistics, this principle influences sampling design and the generalizability of findings.
4

Privacy & Confidentiality

Privacy refers to an individual's right to control access to information about themselves. Confidentiality refers to the researcher's obligation to protect data that has been shared within the research relationship. These are distinct but complementary concepts that underpin all data management protocols.
5

Data Minimization & Purpose Limitation

Data minimization means collecting only the data elements necessary for the stated research objectives. Purpose limitation restricts the use of collected data to the purposes disclosed at the time of consent. Both principles reduce risk exposure and are codified in regulations such as GDPR and the HIPAA Minimum Necessary Standard.
KEY TAKEAWAY
Think of human subjects protections like the safety engineering in a laboratory. Just as a chemistry lab has fume hoods, safety goggles, and waste disposal protocols that constrain how you work but ultimately protect everyone in the room, the Belmont principles and data privacy regulations constrain how you design studies and handle data — but they protect both the participants and the integrity of your research. Skipping them is not a shortcut; it is a structural failure that can collapse the entire enterprise.

Visual Explanation — The Human Subjects Research Ecosystem

The relationship between ethical principles, regulatory bodies, institutional oversight, and data protection mechanisms can be understood as a layered ecosystem. At the outermost layer sit the guiding ethical principles. These are translated into federal and international regulations, which are then implemented locally through Institutional Review Boards (IRBs). At the innermost layer, individual researchers apply data privacy safeguards in their day-to-day work. The following diagram illustrates this nested architecture.

The outermost ellipse represents the Belmont Report's ethical principles. The next ring shows the regulatory layer (Common Rule, HIPAA, GDPR). The IRB translates regulations into actionable protocol requirements. At the center, researcher practice implements de-identification, secure storage, and access controls.

Notice that each layer depends on those surrounding it. A researcher's data management practices are ultimately accountable to the ethical principles at the outermost boundary. If an IRB determines that a protocol's data handling procedures do not adequately minimize risk, the protocol cannot proceed — regardless of how scientifically compelling the study may be. Conversely, a regulation like HIPAA provides the specific technical and procedural standards (such as the 18 HIPAA identifiers) that give concrete meaning to abstract principles like beneficence and respect for persons.

How It Works — The IRB Review Process & Consent Framework

In the United States, federally funded research involving human subjects must undergo review by an Institutional Review Board (IRB) before any data collection begins. The IRB is a committee composed of at least five members with diverse backgrounds, including at least one non-scientist and one member unaffiliated with the institution. Their mandate is to evaluate proposed research protocols against the Belmont principles and the Common Rule, ensuring that risks are minimized, consent is adequate, and vulnerable populations receive appropriate safeguards.

Categories of IRB Review

Not all research carries the same level of risk, and the IRB review process is calibrated accordingly. There are three primary categories of review. Exempt review applies to research that poses minimal risk and falls into one of several predefined categories — for example, the analysis of existing de-identified datasets or the use of anonymous educational surveys. Although the term "exempt" suggests no oversight, the determination itself must be made by the IRB or an authorized designee; researchers cannot self-exempt.

Expedited review is available when the research involves no more than minimal risk — defined as risks no greater than those encountered in daily life or during routine medical examinations — and fits within specific regulatory categories such as the collection of blood samples by venipuncture or the use of voice recordings for research purposes. In an expedited review, the IRB chairperson or a designated experienced reviewer evaluates the protocol without convening the full board.

Full board review is required for research that exceeds minimal risk. This involves a convened meeting of a majority of IRB members, with deliberation and a formal vote. Clinical trials involving investigational drugs, studies with vulnerable populations such as prisoners or children, and protocols that involve more than minimal psychological stress typically require full board review.

Informed Consent: Elements and Process

The Common Rule specifies eight required and six additional elements of informed consent. Among the required elements are: a description of the research purpose and procedures; a disclosure of foreseeable risks and discomforts; a description of expected benefits; a statement regarding confidentiality; and an explanation that participation is voluntary and can be withdrawn at any time without penalty. For biostatisticians, the confidentiality element is particularly significant, as it requires a clear explanation of how data will be de-identified, stored, and potentially shared.

⚖️ Waiver of Informed Consent
An IRB may waive the requirement for informed consent under specific conditions codified in 45 CFR 46.116(f): the research involves no more than minimal risk, the waiver will not adversely affect participants' rights and welfare, the research could not practicably be carried out without the waiver, and (whenever appropriate) participants will be provided with additional pertinent information after participation. This is particularly relevant for retrospective analyses of large existing datasets where re-contacting individuals is infeasible.

Risk–Benefit Assessment Framework

RISK–BENEFIT RATIO (CONCEPTUAL)
Acceptability = f(Σ Benefits to participants and society) / f(Σ Risks to participants)
This is a conceptual representation, not a strict quantitative formula. Benefits include direct benefits to participants and the expected value of knowledge gained. Risks encompass physical, psychological, social, economic, and informational harms. The IRB qualitatively evaluates whether the ratio is acceptable — there is no numerical threshold, but the assessment must be systematic and documented.

Data Privacy in Depth — Identifiers, De-identification, and Regulatory Standards

Data privacy in human subjects research is not a single binary state but rather a spectrum of identifiability. At one extreme, a dataset contains direct identifiers — names, social security numbers, medical record numbers — that unambiguously link records to specific individuals. At the other extreme, data has been aggregated or transformed to the point where re-identification is effectively impossible. Between these poles lie several intermediate states that biostatisticians must understand in order to select appropriate analytical and reporting strategies.

Spectrum of Data Identifiability
Directly Identifiable
Coded / Pseudonymized
Limited Dataset
De-identified (Safe Harbor)
Fully Anonymous
PHI present
Key exists
No re-identification path
High RiskLow Risk
HIPAA provides two pathways to achieve de-identification. The Safe Harbor method (left) requires removal of all 18 specified identifiers and any residual knowledge that could identify an individual. The Expert Determination method (right) relies on a qualified statistical expert to certify that the risk of re-identification is very small.

The distinction between de-identification and anonymization is critical. De-identification under HIPAA means removing specified identifiers such that there is no reasonable basis to believe the information can be used to identify an individual. However, a de-identified dataset under HIPAA may still retain a key that links back to identifiable records — this is known as pseudonymization or coding. True anonymization, as referenced in GDPR, implies that re-identification is irreversible — no key exists and the data cannot be traced back to any individual, even by the original data holder. The choice between these approaches has profound implications for what statistical analyses are permissible, what regulatory frameworks apply, and what consent requirements must be met.

Key Concept: k-Anonymity

K-ANONYMITY CRITERION
For every combination of quasi-identifiers in dataset D, there exist at least k records sharing that combination.
Quasi-identifiers are attributes that are not direct identifiers but can be combined to identify individuals (e.g., zip code + date of birth + gender). A dataset satisfies k-anonymity if each record is indistinguishable from at least k − 1 other records with respect to its quasi-identifiers. For example, if k = 5, any combination of quasi-identifier values appears in at least 5 rows.

Worked Example — Evaluating a Research Protocol

Suppose you are a biostatistician on a research team that wants to study the association between a new dietary supplement and blood pressure reduction. The team proposes a randomized controlled trial with 200 adult participants recruited from a university hospital. You are asked to evaluate whether the protocol meets human subjects and data privacy requirements. Let us walk through this evaluation systematically.

Protocol Evaluation: Dietary Supplement and Blood Pressure Study
1
Step 1 — Determine if the Study Involves Human SubjectsUnder the Common Rule, a human subject is a living individual about whom a researcher obtains data through intervention or interaction, or identifiable private information. In this study, researchers will administer a supplement (intervention), measure blood pressure (interaction), and collect health records (identifiable private information). All three criteria are met.
Yes — this study involves human subjects under 45 CFR 46.
2
Step 2 — Assess the Level of IRB Review RequiredThe study involves an intervention (ingesting a supplement) that could have adverse effects, placing it above the minimal-risk threshold. Participants will also have blood drawn for lab tests. Because the research involves more than minimal risk — the risks exceed those ordinarily encountered in daily life — and involves a direct intervention, this cannot qualify for exempt or expedited review.
Full board IRB review is required.
3
Step 3 — Evaluate the Informed Consent DocumentThe consent form must include: (1) a clear description of the research purpose and the randomization procedure; (2) an explanation that participants may receive a placebo; (3) known risks of the dietary supplement and of blood draws; (4) potential benefits (blood pressure reduction, contribution to knowledge); (5) a confidentiality statement detailing how blood pressure readings, lab results, and health records will be protected; (6) a statement that participation is voluntary; (7) contact information for questions; and (8) information about alternatives to participation. The biostatistician should verify that the confidentiality section specifies the de-identification strategy.
All eight required elements of informed consent must be present.
4
Step 4 — Develop the Data Privacy PlanAs the biostatistician, you recommend: (a) assigning each participant a unique study ID at enrollment and storing the ID-to-name key in a separate, encrypted file accessible only to the PI; (b) removing all 18 HIPAA identifiers from the analytical dataset; (c) storing data on a HIPAA-compliant server with role-based access controls; (d) specifying data retention and destruction timelines in the protocol; and (e) ensuring that any published results present data in aggregate form such that no cell in a cross-tabulation contains fewer than 5 individuals (to prevent inferential disclosure).
The data management plan implements pseudonymization with Safe Harbor de-identification and a minimum cell-size policy.
5
Step 5 — Apply the Risk–Benefit AssessmentRisks include potential adverse reactions to the supplement, discomfort from blood draws, and the possibility of a confidentiality breach. Benefits include direct health benefit if the supplement is effective, advancement of knowledge regarding non-pharmacological blood pressure management, and the contribution to evidence-based dietary guidelines. The risk–benefit assessment must document that the study design minimizes risks (e.g., exclusion criteria for individuals with known allergies, adverse event monitoring, data security measures) and that the remaining risks are reasonable in relation to the anticipated benefits.
The protocol passes the risk–benefit assessment when appropriate safeguards are in place.

Comparing Regulatory Frameworks — HIPAA vs. GDPR vs. the Common Rule

Biostatisticians working with international data or multi-site studies must navigate multiple, sometimes overlapping, regulatory frameworks. The three most influential frameworks for human subjects and data privacy are the Common Rule (governing federally funded research), HIPAA (governing protected health information in the U.S.), and the GDPR (governing personal data in the European Union). While they share a commitment to protecting individuals, they differ in scope, definitions, and enforcement mechanisms.

Comparison of the three major regulatory frameworks governing human subjects research and data privacy.
FeatureCommon Rule (45 CFR 46)HIPAA Privacy RuleGDPR
ScopeFederally funded human subjects research in the U.S.Protected health information held by covered entities and business associatesAll personal data of EU residents, regardless of where processing occurs
Key SubjectHuman subjects (living individuals)Patients and health plan membersData subjects (any identified or identifiable natural person)
Consent BasisInformed consent with specific required elements; waiver possible under defined criteriaAuthorization for use/disclosure of PHI; research exception with IRB/privacy board waiverLawful basis required (consent, legitimate interest, public interest, etc.); explicit consent for sensitive data
De-identification StandardNot specified in detail; focuses on coded vs. identifiable dataSafe Harbor (remove 18 identifiers) or Expert DeterminationAnonymization (irreversible) or pseudonymization (reversible with key); GDPR still applies to pseudonymized data
Right to Withdraw / ErasureParticipants may withdraw at any time without penaltyIndividuals may revoke authorization; previously disclosed data may be retainedRight to erasure ("right to be forgotten"); exceptions for research in the public interest
EnforcementOHRP; sanctions include suspension of fundingHHS Office for Civil Rights; civil and criminal penaltiesNational Data Protection Authorities; fines up to €20M or 4% of global revenue
KEY TAKEAWAY
Think of these three frameworks as overlapping jurisdictional maps. If your research involves U.S. federally funded data collection from human subjects, you operate within the Common Rule. If that data includes protected health information from a covered entity, HIPAA also applies. If any of your participants are EU residents, GDPR enters the picture as well. In multi-site international studies, all three frameworks may apply simultaneously, and you must satisfy the most restrictive requirement in each domain. This is analogous to an engineer designing a bridge that must meet building codes in two different municipalities — you design to whichever standard is stricter for each structural element.

Connection to Advanced Theory — Emerging Challenges in Data Privacy

The foundational concepts covered in this lesson form the bedrock upon which more advanced data privacy techniques are built. As biostatistical datasets grow in size, complexity, and interconnectedness, traditional de-identification methods face new threats. The rise of linkage attacks — where an adversary combines a de-identified research dataset with publicly available auxiliary data to re-identify individuals — has exposed the limitations of simple identifier removal. Latanya Sweeney's landmark demonstration that 87% of the U.S. population could be uniquely identified by the combination of zip code, date of birth, and sex underscored the inadequacy of naive de-identification approaches.

How foundational concepts in this lesson connect to advanced data privacy and research ethics topics.
ConceptBasic (This Lesson)Advanced Extension
De-identificationSafe Harbor removal of 18 HIPAA identifiersDifferential privacy — mathematical guarantee that individual records do not significantly affect query outputs
k-AnonymityEach quasi-identifier combination appears ≥ k timesl-Diversity (sensitive attribute diversity within equivalence classes) and t-Closeness (distribution similarity)
ConsentInformed consent for specific studyBroad consent, dynamic consent platforms, and tiered consent models for biobanks and longitudinal studies
Data SharingSharing de-identified datasets with collaboratorsFederated learning — training models across institutions without centralizing raw data
IRB OversightSingle-site IRB reviewSingle IRB of record for multi-site trials (2018 Common Rule revision); reliance agreements

The concept of differential privacy represents one of the most significant advances in data privacy theory. Formally, a randomized algorithm M satisfies ε-differential privacy if, for all datasets D₁ and D₂ differing in a single record, and for all subsets S of possible outputs, Pr[M(D₁) ∈ S] ≤ eε × Pr[M(D₂) ∈ S]. The parameter ε (epsilon) quantifies the privacy loss: smaller values of ε provide stronger privacy guarantees but introduce more noise into the output, creating a fundamental tension between privacy and statistical utility. The U.S. Census Bureau adopted differential privacy for the 2020 Decennial Census, illustrating its practical relevance at scale.

ε-DIFFERENTIAL PRIVACY
Pr[M(D₁) ∈ S] ≤ e^ε × Pr[M(D₂) ∈ S]
M is a randomized mechanism, D₁ and D₂ are neighboring datasets differing by one record, S is any subset of outputs, and ε (epsilon) is the privacy budget. A smaller ε means stronger privacy but more noise in the output.

As you progress through your biostatistics training, you will encounter these advanced methods in the context of genomic data sharing, electronic health record research, and precision medicine initiatives. The foundational understanding of human subjects ethics and data privacy concepts developed in this lesson will provide the conceptual vocabulary needed to engage meaningfully with these challenges.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher at a university wants to analyze an existing hospital database of patient blood pressure readings from the past 10 years. The dataset has been stripped of names and social security numbers but still contains full dates of birth, zip codes, and gender. Does this study involve human subjects under the Common Rule? Is the dataset de-identified under HIPAA Safe Harbor? Explain your reasoning for both questions.
PROBLEM 2BASIC CALCULATION
A dataset contains records with three quasi-identifiers: age group (6 categories), gender (3 categories), and region (4 categories). What is the maximum number of unique quasi-identifier combinations? If the dataset contains 1,200 records and we want to achieve k-anonymity with k = 5, what is the minimum number of records that must share the most common quasi-identifier combination, assuming a perfectly uniform distribution?
PROBLEM 3INTERMEDIATE
A multi-site clinical trial is being conducted at three hospitals: one in Texas (USA), one in London (UK), and one in Berlin (Germany). The trial is funded by the U.S. National Institutes of Health (NIH). For each site, identify which regulatory frameworks apply (Common Rule, HIPAA, GDPR). Then explain why the consent form at the Berlin site might need to include elements not required at the Texas site.
PROBLEM 4APPLIED
You are a biostatistician asked to design the data management plan for a study examining the relationship between genetic markers and treatment response in 500 cancer patients. The PI wants to share the dataset with collaborators at three other universities for secondary analyses. Outline a comprehensive data privacy strategy that addresses: (a) the type of consent needed, (b) the de-identification approach, (c) data sharing safeguards, and (d) how you would handle a participant who withdraws consent after their data has already been analyzed and results published.
PROBLEM 5CRITICAL THINKING
A technology company offers a free health-tracking app that collects heart rate, sleep patterns, GPS location, and self-reported mood data from 10 million users. Users agree to terms of service that permit the company to share 'de-identified aggregate data' with researchers. A university biostatistics team receives a dataset with no names or emails, but it includes hourly GPS coordinates, minute-level heart rate data, age, and gender. Critically evaluate: (1) whether this dataset is truly de-identified, (2) what ethical concerns arise from using terms-of-service consent rather than informed consent for research, (3) whether the Belmont principles are applicable even though this is not federally funded research, and (4) what responsibility the biostatistician has when the regulatory framework may be ambiguous.

Summary — Human Subjects & Data Privacy

The protection of human subjects in research is grounded in three core ethical principles from the Belmont Report: Respect for Persons (mandating informed consent), Beneficence (requiring systematic risk–benefit assessment), and Justice (ensuring fair distribution of research burdens and benefits). These principles are implemented through three overlapping regulatory frameworks: the Common Rule for federally funded research, HIPAA for protected health information, and the GDPR for personal data of EU residents. Institutional Review Boards (IRBs) serve as the local enforcement mechanism, reviewing protocols through exempt, expedited, or full board review depending on the level of risk.

Data privacy exists on a spectrum of identifiability, ranging from directly identifiable data to fully anonymous data. HIPAA provides two pathways to de-identification: the Safe Harbor method (removing all 18 HIPAA identifiers) and the Expert Determination method (statistical certification of low re-identification risk). Advanced techniques such as k-anonymity and differential privacy offer stronger formal guarantees. For biostatisticians, every stage of the research pipeline — from study design and sampling through analysis and publication — must be conducted within this ethical and legal framework. Principles of data minimization and purpose limitation should guide every data management decision.

Varsity Tutors • Biostatistics • Human Subjects & Data Privacy