EPPP: PART 1, KNOWLEDGE • DOMAIN 7: RESEARCH METHODS AND STATISTICS

Sampling Methods — Differentiate probability and non-probability sampling strategies

Understanding how participant selection shapes the validity and generalizability of behavioral health research findings.

Historical Context & Motivation

The history of sampling in research is, in many respects, a history of learning from spectacular failures. Before formalized sampling methods existed, researchers and pollsters frequently drew conclusions from whichever participants happened to be available, often producing results that were misleading or flatly wrong. The evolution of sampling methodology was driven by a fundamental insight: the way you select participants from a population determines whether your findings can be trusted and generalized. In the behavioral health sciences, where we study diverse human experiences—depression, addiction, trauma, neurodevelopmental conditions—the stakes of poor sampling are especially high because biased samples can lead to treatments that work for some groups but fail others entirely.

1895
Bowley's Random Sampling Advocacy
Arthur Bowley, a British statistician, was among the first to advocate for random sampling as a principled alternative to complete enumeration (census). His work in poverty surveys demonstrated that carefully selected subsets could represent entire populations with quantifiable accuracy.
1936
The Literary Digest Debacle
The Literary Digest predicted Alf Landon would defeat Franklin Roosevelt in the presidential election based on a massive but deeply biased sample drawn from telephone directories and automobile registrations—sources that overrepresented wealthier Americans. The poll's spectacular failure became the canonical cautionary tale about sampling bias and demonstrated that sample quality matters far more than sample size.
1934–1950
Neyman and the Formalization of Probability Sampling
Jerzy Neyman developed the mathematical theory underpinning probability sampling, establishing that every member of a population must have a known, nonzero probability of selection for results to be statistically generalizable. This framework became the gold standard for survey research and epidemiological studies.
1967
Glaser & Strauss and Theoretical Sampling
With the publication of The Discovery of Grounded Theory, Barney Glaser and Anselm Strauss formalized theoretical sampling as a legitimate non-probability strategy in qualitative research, arguing that purposeful participant selection was essential for theory development in behavioral and social sciences.
2000s–Present
Mixed-Methods and Adaptive Designs
Contemporary behavioral health research increasingly employs mixed-methods designs that intentionally combine probability and non-probability approaches. Community-based participatory research and respondent-driven sampling have expanded access to hard-to-reach populations such as individuals experiencing homelessness or those with severe mental illness.

The central question that these historical developments converge upon is deceptively simple: How should a researcher select participants so that findings accurately reflect the broader population of interest? For the EPPP, you must understand not only the mechanics of each sampling strategy but also the trade-offs each introduces with respect to external validity, feasibility, and ethical constraints—considerations that are particularly salient in behavioral health research where vulnerable populations are frequently the focus of study.

Core Principles & Definitions

Before distinguishing between probability and non-probability sampling, it is essential to ground the discussion in several foundational concepts. A population is the entire set of individuals (or units) about whom the researcher wishes to draw conclusions—for instance, all adults in the United States diagnosed with major depressive disorder. A sample is the subset of that population actually included in the study. The sampling frame is the operational list from which participants are drawn, and discrepancies between the sampling frame and the true population are a major source of bias. External validity—the degree to which findings generalize beyond the study sample—depends heavily on how well the sampling strategy mirrors the target population.

1

Probability Sampling

Every member of the population has a known, nonzero probability of being selected. This enables calculation of sampling error and supports statistical generalization. Examples include simple random, stratified, cluster, and systematic sampling.
2

Non-Probability Sampling

Selection is based on researcher judgment, convenience, or participant self-selection rather than random processes. Sampling error cannot be formally estimated. Includes convenience, purposive, quota, and snowball sampling.
3

Sampling Bias

A systematic error that occurs when certain members of the population are more or less likely to be included than others. Bias cannot be reduced by increasing sample size—it is a structural problem with the selection process itself.
4

Representativeness

A representative sample mirrors the population on key characteristics (e.g., age, gender, diagnosis severity). Probability methods maximize the likelihood of representativeness, while non-probability methods require additional justification.
5

Sampling Error vs. Non-Sampling Error

Sampling error is the natural discrepancy between a sample statistic and the population parameter; it decreases with larger samples. Non-sampling error (measurement error, nonresponse bias) arises from flaws in data collection, not selection.
KEY TAKEAWAY
Think of probability sampling like drawing names from a hat that contains every person in your target population—everyone has a fair chance of being selected, and you can mathematically describe how likely any particular draw is. Non-probability sampling is more like recruiting people who walk past your clinic's front door: you may get useful information, but you cannot know whether the people walking by differ in important ways from those who do not. In behavioral health research, this distinction is critical because the populations we study—individuals with serious mental illness, substance use disorders, or trauma histories—often have characteristics that correlate with accessibility, meaning convenience samples may systematically exclude the very individuals whose experiences matter most.

Visual Explanation — Probability vs. Non-Probability Sampling

This taxonomy diagram illustrates the two major branches of sampling methodology. Probability methods (left, blue) rely on randomization and known selection probabilities, while non-probability methods (right, pink) rely on researcher judgment or participant availability. The bottom panels compare key features, showing the trade-off between generalizability and practicality.

The diagram above captures the fundamental bifurcation in sampling methodology. On the left side, all four probability methods share a common feature: randomization ensures that every individual in the target population has a calculable chance of being selected, which in turn permits the researcher to estimate sampling error and construct confidence intervals. On the right side, non-probability methods sacrifice this mathematical guarantee in exchange for practical advantages—lower cost, faster recruitment, or access to populations for which no sampling frame exists. For EPPP preparation, it is essential to recognize that this distinction directly maps onto the concept of external validity: probability sampling supports generalization to defined populations, whereas non-probability sampling limits conclusions to the sample itself unless additional analytic steps (e.g., weighting, sensitivity analyses) are employed.

How Each Sampling Method Works

Probability Sampling Methods in Detail

Simple random sampling (SRS) is the most straightforward probability method: every member of the population has an equal probability of selection. The researcher assigns each individual a unique identifier and uses a random number generator (or random number table) to select participants. If the population contains N individuals and the desired sample size is n, each individual has a selection probability of n/N. SRS is conceptually elegant but requires a complete and accurate sampling frame—something that can be difficult to obtain in behavioral health research, where clinical registries may be incomplete or stigmatized conditions may lead to underreporting.

SELECTION PROBABILITY IN SRS
P(selection) = n / N
Where n = desired sample size and N = total population size. In SRS, this probability is equal for every member.

Stratified random sampling divides the population into mutually exclusive subgroups (strata) based on a characteristic relevant to the research question—such as age group, severity of diagnosis, or treatment modality—and then conducts a separate random sample within each stratum. This method guarantees representation of key subgroups and typically produces more precise estimates than SRS because within-stratum variability is reduced. In proportional stratified sampling, the number drawn from each stratum reflects its proportion in the population; in disproportional stratified sampling, smaller strata may be oversampled to ensure adequate statistical power for subgroup analyses.

Cluster sampling is used when a complete list of individuals is unavailable but a list of naturally occurring groups (clusters) is accessible—for example, community mental health centers, schools, or hospital units. The researcher randomly selects entire clusters and then either includes all members of the selected clusters (single-stage) or takes a random sample within each selected cluster (two-stage or multi-stage). Cluster sampling is more cost-effective for geographically dispersed populations but introduces greater sampling error because individuals within clusters tend to be more similar to one another than to the population at large.

Systematic sampling selects every kth individual from an ordered list after a random start. The sampling interval k is calculated as N/n. This method is operationally simpler than SRS and functions well when the list is randomly ordered; however, if the list contains a periodic pattern that aligns with the sampling interval, systematic bias can result.

SYSTEMATIC SAMPLING INTERVAL
k = N / n
Every kth individual is selected after choosing a random starting point between 1 and k. If the list is randomly ordered, this approximates simple random sampling.

Non-Probability Sampling Methods in Detail

Convenience sampling selects participants who are readily available to the researcher—patients in a particular clinic, students enrolled in a psychology course, or individuals who respond to a posted flyer. It is the most common sampling method in behavioral health research because of practical constraints, but it carries the highest risk of bias. There is no mechanism to ensure that the convenience sample mirrors the broader population on relevant variables.

Purposive (judgment) sampling involves the deliberate selection of participants based on specific criteria established by the researcher. A clinician-researcher studying treatment-resistant schizophrenia might hand-select patients who meet strict diagnostic and treatment-history criteria. While this approach is indispensable for qualitative research and studies of rare conditions, the reliance on researcher judgment means that selection probabilities are unknown and results are not statistically generalizable.

Snowball (chain-referral) sampling begins with a small number of initial participants who then recruit additional participants from their personal networks. This method is particularly valuable for accessing hidden or stigmatized populations—such as individuals engaged in illicit substance use, undocumented immigrants seeking mental health services, or people living with HIV who are not connected to formal care systems. The limitation is that the resulting sample tends to be socially clustered, reflecting the networks of the initial participants rather than the broader population.

Quota sampling resembles stratified sampling in appearance but lacks the randomization that makes stratified sampling a probability method. The researcher establishes quotas for key demographic or clinical categories (e.g., 40% female, 30% with comorbid anxiety) and recruits until each quota is filled. The selection of specific individuals within each category is non-random, which means that quota sampling cannot support formal statistical inference despite its structured appearance.

Detailed Classification and Comparison

This scatter plot conceptualizes the trade-off between generalizability (y-axis) and feasibility (x-axis). Stratified random sampling offers the highest generalizability but demands more resources, while convenience sampling is the easiest to implement but yields the weakest external validity. The dashed boundaries separate probability methods (upper region) from non-probability methods (lower region).
Comprehensive comparison of eight major sampling methods with behavioral health examples
MethodTypeProcedureBehavioral Health Example
Simple RandomProbabilityRandom number generator selects n individuals from complete list of NSelecting 200 patients from a state registry of 10,000 adults with PTSD
Stratified RandomProbabilityPopulation divided into strata; random sample drawn from each stratumStratifying by depression severity (mild, moderate, severe) before random selection
ClusterProbabilityRandomly select groups (clusters); include all or subsample within clustersRandomly selecting 15 community mental health centers, then sampling patients within each
SystematicProbabilitySelect every kth individual from ordered list after random startSelecting every 5th intake at an addiction treatment facility over a 6-month period
ConvenienceNon-ProbabilityRecruit whoever is available and willingRecruiting undergraduate psychology students for a study on anxiety
PurposiveNon-ProbabilityResearcher selects participants based on specific criteria or expertiseInterviewing therapists with 10+ years of DBT experience for a qualitative study
SnowballNon-ProbabilityInitial participants refer others from their networksStudying opioid misuse among rural populations through peer referral chains
QuotaNon-ProbabilitySet demographic/clinical quotas; fill non-randomlyEnsuring 50% male and 50% female in a trauma exposure study, but selecting participants by convenience
💡 EPPP EXAM TIP
A common EPPP question format presents a research scenario and asks you to identify the sampling method being used. Pay close attention to whether the scenario describes a random mechanism (probability) or a judgment-based or availability-based selection process (non-probability). Also distinguish between quota and stratified sampling: both create subgroups, but only stratified sampling uses randomization within those subgroups.

Worked Example — Selecting a Sampling Strategy

Consider the following scenario: Dr. Alvarez is conducting a study on the effectiveness of a new group therapy intervention for generalized anxiety disorder (GAD). She has access to a statewide electronic health records database that lists 8,000 adults diagnosed with GAD across 40 outpatient clinics. She wants a sample of 400 participants and wants to ensure that participants represent the full range of anxiety severity (mild, moderate, severe) found in the population. Her budget is limited, and she cannot travel to all 40 clinics. Let us walk through the decision-making process for selecting an appropriate sampling strategy.

Choosing and Implementing a Sampling Strategy for a GAD Treatment Study
1
Step 1 — Define the Target Population and Sampling FrameThe target population is all adults diagnosed with GAD in the state. The sampling frame is the electronic health records database of 8,000 patients. Dr. Alvarez should evaluate whether this frame is complete—are there individuals with GAD who are not in the system (e.g., uninsured, untreated, or diagnosed in private practice)? If so, the sampling frame underrepresents certain subgroups, introducing coverage error.
Population: all state adults with GAD. Frame: N = 8,000 patients in EHR database.
2
Step 2 — Evaluate Feasibility ConstraintsDr. Alvarez cannot visit all 40 clinics. Simple random sampling from the full database would likely scatter participants across many sites, making in-person group therapy impractical. Stratified random sampling would ensure severity representation but also spread participants geographically. A multi-stage cluster approach is most appropriate: first randomly select a feasible number of clinics, then stratify patients within those clinics by severity, and finally randomly sample within each stratum.
Multi-stage design selected: cluster sampling (clinics) combined with stratified random sampling (severity).
3
Step 3 — Execute the Sampling PlanStage 1: Randomly select 10 clinics from the 40 available (cluster sampling). Stage 2: Within each selected clinic, classify patients as mild, moderate, or severe GAD based on standardized assessment scores. Stage 3: Conduct proportional stratified random sampling within each clinic to select approximately 40 patients per clinic (total n = 400), preserving the population proportions of each severity level. For instance, if the population is 30% mild, 50% moderate, and 20% severe, then each clinic's sample of 40 would include approximately 12 mild, 20 moderate, and 8 severe patients, selected randomly within each stratum.
Per clinic: ~12 mild + ~20 moderate + ~8 severe = 40 participants × 10 clinics = 400 total.
4
Step 4 — Assess Threats to ValidityThis multi-stage design introduces some loss of precision compared to pure stratified random sampling (because patients within clinics may share characteristics), but it is far more feasible. Dr. Alvarez should report the design effect—the ratio of the variance of her cluster-based estimate to the variance that would have been obtained from SRS—and adjust confidence intervals accordingly. She should also document any nonresponse and assess whether nonrespondents differ systematically from respondents.
The final sample is a probability sample with known selection probabilities, supporting generalization to the EHR population with appropriate statistical adjustments.

Strengths, Limitations, and Contextual Considerations

Comparative analysis of probability and non-probability sampling across seven evaluative criteria
CriterionProbability SamplingNon-Probability Sampling
External ValidityHigh — supports generalization to the defined populationLimited — findings pertain to the sample; generalization requires additional justification
Sampling ErrorCalculable — enables confidence intervals and hypothesis testingNot calculable — cannot formally estimate the margin of error
Cost and TimeHigher — requires complete sampling frame, more complex logisticsLower — faster recruitment, fewer resources needed
Access to Hidden PopulationsDifficult — requires an enumerable populationSuperior — snowball and purposive methods reach populations without sampling frames
Bias RiskLower — randomization minimizes systematic selection biasHigher — researcher judgment, self-selection, and network homogeneity introduce bias
Qualitative ResearchRarely used — statistical generalization is not the goal of qualitative inquiryPreferred — purposive and theoretical sampling align with qualitative aims of depth and meaning
Ethical FeasibilityMay raise concerns if random assignment to study conditions involves withholding treatmentMay be more ethically acceptable when studying vulnerable populations who cannot be systematically enumerated
KEY TAKEAWAY
Neither probability nor non-probability sampling is inherently superior—the optimal choice depends on the research question, available resources, population characteristics, and the type of inference desired. In behavioral health, researchers frequently face a tension between the ideal (probability sampling for maximum generalizability) and the reality (many populations of clinical interest are hidden, stigmatized, or not captured in existing databases). The EPPP expects you to recognize this tension and to evaluate sampling decisions in context rather than applying a blanket rule. A well-designed convenience study with transparent reporting of limitations may contribute more to clinical knowledge than a poorly executed probability study with massive nonresponse.

Connections to Validity, Power, and Advanced Designs

Sampling decisions do not exist in isolation—they ripple through every aspect of a study's design, analysis, and interpretation. The EPPP frequently tests your ability to connect sampling methodology to broader concepts in research design, and three connections are particularly important: the relationship between sampling and external validity, the impact of sampling on statistical power, and the role of sampling in evidence-based practice guidelines.

Connecting sampling to broader research methodology concepts tested on the EPPP
ConceptBasic UnderstandingAdvanced Connection to Sampling
External ValidityGeneralizability of findings beyond the study sample to the target populationProbability sampling directly supports external validity; non-probability sampling requires analytical generalization, replication, or meta-analytic synthesis to build evidence for generalizability
Statistical PowerThe probability of detecting a true effect (avoiding Type II error)Stratified sampling increases power by reducing within-group variance; cluster sampling decreases effective sample size due to intraclass correlation, requiring larger total n for equivalent power
Internal ValidityThe degree to which the study establishes a causal relationshipSampling method is distinct from random assignment. Random sampling determines who is in the study; random assignment determines who gets which condition. A study can use convenience sampling but still employ random assignment, maintaining internal validity while limiting external validity
Respondent-Driven Sampling (RDS)An advanced variant of snowball samplingRDS uses mathematical modeling of the referral process to derive population-level estimates from network-based recruitment, partially bridging the gap between non-probability and probability methods. Increasingly used in HIV and substance use research
Mixed-Methods DesignsStudies combining quantitative and qualitative approachesThe quantitative strand may use probability sampling while the qualitative strand uses purposive sampling from the same population. This requires careful integration and transparent reporting of each strand's sampling logic
⚠️ CRITICAL DISTINCTION
Do not confuse random sampling with random assignment. Random sampling determines who enters the study (addressing external validity), while random assignment determines which condition each participant receives (addressing internal validity). A study can employ one without the other. This distinction is a perennial EPPP favorite.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher conducts a study on burnout among licensed clinical psychologists. She obtains a list of all licensed clinical psychologists in her state (N = 5,200), assigns each a number, and uses a random number generator to select 300 participants. What type of sampling method is she using, and what is the primary advantage of this approach?
PROBLEM 2BASIC CALCULATION
A substance abuse treatment center has 1,500 patient records. A researcher wants to conduct a systematic sample of 100 patients. What is the sampling interval (k), and if the random starting point is patient #7, which are the first five patients selected?
PROBLEM 3INTERMEDIATE
Dr. Chen wants to study the lived experiences of transgender adolescents accessing mental health services. He knows that no comprehensive registry of transgender adolescents exists, and many in this population are not connected to formal health services. He begins by recruiting three participants from a local LGBTQ+ youth center, each of whom then refers two peers, and so on. (a) Identify the sampling method. (b) Explain why this is classified as non-probability sampling. (c) Identify two specific threats to the validity of Dr. Chen's findings.
PROBLEM 4APPLIED
A state health department wants to estimate the prevalence of co-occurring PTSD and substance use disorder among veterans across 120 VA medical centers. Budget constraints limit data collection to 20 sites. The department wants to ensure that findings are representative of veterans across different geographic regions (urban, suburban, rural). Design a multi-stage sampling strategy, specifying the sampling method used at each stage and explaining how the design addresses both feasibility and representativeness.
PROBLEM 5CRITICAL THINKING
A clinical psychology researcher publishes a study demonstrating that a new cognitive-behavioral intervention significantly reduces symptoms of social anxiety disorder. The study used a convenience sample of 60 undergraduate students who scored above the clinical cutoff on a social anxiety measure and were recruited through a departmental participant pool. A reviewer criticizes the study's external validity. The researcher responds that random assignment to treatment vs. control conditions ensures the validity of the findings. Evaluate both the reviewer's criticism and the researcher's defense. Under what circumstances might the use of a convenience sample be acceptable for this type of study?

Summary — Sampling Methods in Behavioral Health Research

Sampling methods divide into two fundamental categories. Probability sampling—including simple random, stratified random, cluster, and systematic sampling—ensures that every member of the population has a known, nonzero chance of selection, enabling formal estimation of sampling error and supporting strong external validity. Non-probability sampling—including convenience, purposive, snowball, and quota sampling—relies on researcher judgment, availability, or referral networks and does not permit calculation of sampling error, limiting generalizability but offering practical access to populations that probability methods cannot reach.

For the EPPP, remember three critical distinctions. First, random sampling (who is in the study) is distinct from random assignment (who receives which treatment). Second, stratified sampling and quota sampling both create subgroups, but only stratified sampling uses randomization within those subgroups. Third, sampling bias is a systematic error that cannot be corrected by increasing sample size—only by improving the selection process. The choice between probability and non-probability sampling should be guided by the research question, population accessibility, resource constraints, and the type of inference desired.

Varsity Tutors • EPPP: Part 1, Knowledge • Sampling Methods — Differentiate probability and non-probability sampling strategies