EPPP: PART 1, KNOWLEDGE • DOMAIN 7: RESEARCH METHODS AND STATISTICS

Data Collection Methods — Evaluate strengths and weaknesses of survey, observational, and archival data methods

Understanding when and why to use surveys, observations, or archival records to generate valid behavioral health evidence.

Historical Context & Motivation

The question of how best to gather information about human behavior has preoccupied researchers since the emergence of psychology as a formal discipline. Early investigators relied almost exclusively on clinical case notes and introspective reports—approaches that were deeply intertwined with the observer's theoretical commitments and therefore difficult to replicate. As the field matured, the demand for systematic, reproducible methods gave rise to three broad families of data collection: surveys and self-report instruments, direct observational procedures, and the use of pre-existing archival records. Each method crystallized in response to specific limitations of the others, and understanding their historical trajectory illuminates why no single approach dominates modern behavioral health research.

1890s
Early Survey Instruments
G. Stanley Hall and other pioneers disseminated questionnaires to study child development and educational attitudes, marking one of the first large-scale uses of self-report in psychology.
1920s–1930s
Systematic Observation Emerges
Researchers such as Jean Piaget and the behaviorist school formalized direct observation protocols, introducing time-sampling and event-recording methods to study behavior in naturalistic and laboratory settings.
1950s–1960s
Archival Data Gains Scientific Credibility
Émile Durkheim's sociological tradition was extended into psychology; epidemiological researchers began mining hospital records, census data, and vital statistics to study mental health trends at the population level.
1970s–1980s
Psychometric Rigor and Standardized Scales
The development of classical test theory and item response theory led to an explosion of validated self-report measures—Beck Depression Inventory, MMPI-2, SCL-90-R—transforming survey-based research in clinical psychology.
2000s–Present
Big Data and Digital Methods
Electronic health records, smartphone-based ecological momentary assessment, and social media analytics blur traditional boundaries, creating hybrid methods that combine survey, observational, and archival elements at unprecedented scale.

This historical progression reveals a central tension in behavioral health research: how do we collect data that are simultaneously rich, accurate, efficient, and ethically defensible? Each method represents a different resolution of that tension, trading off internal validity, external validity, feasibility, and participant burden in characteristic ways. For the EPPP, you need to evaluate these trade-offs fluently, matching data collection strategies to research questions and identifying the threats to validity that each method introduces.

Core Principles & Definitions

Before comparing the three major data collection methods, it is essential to anchor the discussion in several foundational concepts that cut across all research designs in behavioral health. These principles provide the evaluative framework through which one judges whether a given method is appropriate for a particular research question.

1

Reactivity

The degree to which the act of measurement changes the behavior being measured. Methods differ dramatically in their susceptibility to reactivity; surveys may trigger socially desirable responding, while archival records are typically nonreactive.
2

Ecological Validity

The extent to which findings generalize to real-world settings. Ecological validity is often highest for naturalistic observation and archival data drawn from everyday contexts, and lowest for laboratory-based paradigms.
3

Response Bias

Systematic distortions in self-reported data, including social desirability bias, acquiescence bias, and recall bias. These are especially relevant to surveys and structured interviews.
4

Observer Bias

In observational research, the observer's expectations, training quality, and fatigue can distort data recording. Establishing high inter-rater reliability is the primary safeguard.
5

Nonreactive (Unobtrusive) Measurement

Data collection that does not alert participants to the fact they are being studied. Archival methods are the prototypical unobtrusive measure, though covert observation also qualifies—albeit with significant ethical caveats.
KEY TAKEAWAY
Think of the three data collection methods as three different cameras aimed at the same scene. A survey is like a selfie—participants control the angle, which is convenient but filtered. Observation is like a documentary crew—detailed and vivid, but the presence of the camera may change the action. Archival data is like reviewing security footage after the fact—completely unfiltered, but you can only see what the camera happened to capture, and you cannot reposition it.

Visual Explanation — The Three Methods at a Glance

A side-by-side comparison of the three primary data collection families: survey (self-report), observational (direct behavior coding), and archival (pre-existing records). Note how strengths in one column often correspond to weaknesses in another, which is why multi-method designs are considered best practice.

The diagram above highlights a complementary pattern: where one method is strong, another tends to be weak. Survey methods excel at accessing internal states (thoughts, feelings, attitudes) but are vulnerable to participants' willingness and ability to report accurately. Observational methods capture behavior as it actually unfolds, yet the observer's presence can alter the very behavior under study—a phenomenon known as the Hawthorne effect. Archival methods sidestep reactivity entirely because the data were generated before the research question existed, but the investigator sacrifices control over what was recorded and how consistently the records were maintained. This trade-off structure is the conceptual backbone of the EPPP's data-methods questions.

How Each Method Works — Deep Dive

Survey Methods: From Questionnaires to Structured Interviews

Survey research encompasses any procedure in which participants provide verbal or written responses to a predetermined set of questions. The spectrum ranges from mailed paper questionnaires and online forms to face-to-face structured clinical interviews such as the SCID-5. A key distinction exists between closed-ended items (Likert scales, multiple-choice), which yield quantitative data amenable to statistical analysis, and open-ended items, which produce richer qualitative data but require content analysis or thematic coding. In behavioral health contexts, validated self-report instruments—such as the PHQ-9 for depression screening or the GAD-7 for generalized anxiety—are widely used because their psychometric properties (reliability, sensitivity, specificity) have been extensively documented.

The primary threats to survey validity include social desirability bias (respondents present themselves favorably), acquiescence bias (tendency to agree with statements regardless of content), recall bias (inaccurate memory for past events), and nonresponse bias (systematic differences between responders and non-responders). Each of these threats can be partially mitigated through careful instrument design—for example, embedding validity scales (as in the MMPI-2) or using forced-choice formats.

Observational Methods: Naturalistic, Participant, and Structured

Observational methods involve trained researchers directly watching and coding behavior. Three major subtypes exist. Naturalistic observation occurs in the participant's everyday environment without intervention—for example, coding parent-child interactions at a playground. Participant observation involves the researcher embedding within the group being studied, a method with deep roots in anthropology and qualitative psychology. Structured (or systematic) observation uses a controlled setting and a detailed behavioral coding scheme to maximize reliability. Sampling strategies—time sampling, event sampling, and interval recording—determine which behaviors are recorded and when, and the choice among them can substantially affect the resulting data.

Central to high-quality observational research is inter-rater reliability, typically quantified using Cohen's kappa (κ) for categorical codes or intraclass correlation coefficients (ICC) for continuous ratings. A κ value above .80 is generally considered excellent agreement, while values between .60 and .79 represent substantial agreement. Without demonstrating adequate inter-rater reliability, observational findings are viewed as insufficiently objective.

COHEN'S KAPPA
κ = (P₀ − Pₑ) / (1 − Pₑ)
where P₀ = observed proportion of agreement between two raters, and Pₑ = proportion of agreement expected by chance alone. A κ of 1.0 indicates perfect agreement; 0 indicates chance-level agreement.

Archival Methods: Clinical Records, Databases, and Historical Documents

Archival research analyzes data that already exist, having been collected for purposes other than the current study. In behavioral health, common archival sources include electronic health records (EHRs), insurance claims databases, court records, school disciplinary logs, and national datasets such as the National Comorbidity Survey Replication (NCS-R) or the Behavioral Risk Factor Surveillance System (BRFSS). The researcher's task is to extract, code, and analyze relevant variables from these pre-existing repositories.

Two archival-specific validity threats deserve attention. Selective deposit refers to the fact that not all events are recorded—certain populations or diagnoses may be systematically underrepresented in the records. Selective survival refers to the differential preservation of records over time; older or less-valued documents may be lost or destroyed, biasing the surviving sample. Additionally, coding conventions and diagnostic criteria change over time (e.g., the transition from DSM-IV to DSM-5), which can introduce measurement non-equivalence in longitudinal archival analyses.

Detailed Classification — Threats, Safeguards, and Decision Criteria

A decision flowchart for selecting a data collection method based on the nature of the research question. Begin at the top and follow branching criteria. Each terminal node also lists recommended safeguards for that method.

The flowchart above provides a simplified decision heuristic, but real-world research often demands a more nuanced evaluation. Consider, for example, that a clinical researcher interested in substance use frequency might use a self-report screening tool (survey), verify consumption through urinalysis records (archival), and observe social drinking behavior in a controlled bar-lab setting (observation). This multi-method or triangulation approach is widely regarded as the gold standard because the weaknesses of one method are offset by the strengths of another. On the EPPP, questions frequently present a scenario and ask which method—or combination of methods—would maximize validity for that particular research aim.

💡 EPPP TIP
When an EPPP item describes a study of sensitive behaviors (e.g., sexual behavior, drug use, violence), be alert for the correct answer involving archival or unobtrusive methods. Participants are most likely to underreport socially stigmatized behaviors on surveys, making archival verification or indirect observation particularly valuable.

Worked Example — Evaluating Data Collection in a Behavioral Health Study

Imagine a research team at a community mental health center wants to study the relationship between therapist alliance and treatment dropout among clients diagnosed with borderline personality disorder (BPD). The following worked example walks through how to evaluate each data collection method for this research question.

Selecting and Evaluating Data Collection Methods for a BPD Treatment Study
1
Step 1 — Define the ConstructsThe two key constructs are therapeutic alliance (a subjective experience of the working relationship) and treatment dropout (a behavioral event). Alliance is inherently subjective—it reflects the client's felt sense of collaboration with the therapist—which makes it a prime candidate for self-report measurement. Treatment dropout, by contrast, is an observable, documentable event.
Alliance → subjective → survey; Dropout → behavioral event → archival or observational
2
Step 2 — Evaluate Survey Method for AllianceThe Working Alliance Inventory (WAI) is a well-validated self-report measure with demonstrated internal consistency (Cronbach's α ≈ .93) and predictive validity for therapy outcomes. Strengths: standardized, quantifiable, captures the client's private experience. Weaknesses: clients with BPD may exhibit unstable self-perceptions, fluctuating alliance ratings session-to-session, and may provide socially desirable responses to avoid conflict with their therapist.
Survey (WAI) is appropriate for alliance but requires awareness of BPD-specific response patterns.
3
Step 3 — Evaluate Archival Method for DropoutClinic records (EHR data) provide a nonreactive measure of dropout: session attendance logs, discharge summaries, and no-show records. Strengths: no risk of reactivity, large sample sizes possible, exact dates available. Weaknesses: the records may not distinguish client-initiated dropout from clinician-initiated discharge, insurance-mandated termination, or relocation. Selective deposit is a concern if some clinicians document termination reasons more thoroughly than others.
Archival (EHR) is strong for dropout timing but requires supplemental data to classify dropout type.
4
Step 4 — Evaluate Observational Method as a ComplementThe research team could video-record therapy sessions and code alliance-related behaviors (e.g., therapist empathy, client engagement) using an established coding system such as the System for Observing Family Therapy Alliances (SOFTA). This offers a non-self-report perspective on alliance. Strengths: captures actual interpersonal dynamics, triangulates with self-report. Weaknesses: extremely time-intensive, raises ethical concerns about recording clients with BPD (potential emotional dysregulation), the Hawthorne effect may alter therapist behavior, and requires extensive coder training to achieve adequate κ.
Observation is valuable for triangulation but introduces feasibility and ethical challenges specific to the BPD population.
5
Step 5 — Synthesize a Multi-Method DesignThe optimal design combines all three methods. Self-report (WAI) measures perceived alliance, archival data (EHR) provides an objective dropout indicator, and observational coding adds behavioral data on alliance processes. This triangulation approach maximizes construct validity by converging multiple operationalizations of the same construct and reduces reliance on any single method's limitations.
Multi-method triangulation is the strongest design: survey for subjective alliance, archival for dropout, observation for behavioral process.

Comparative Strengths, Weaknesses, and Mitigation Strategies

Comparative evaluation of survey, observational, and archival methods across seven key criteria.
CriterionSurveyObservationalArchival
ReactivityHigh — participants know they are being studied and may alter responsesModerate — Hawthorne effect can alter behavior, mitigated by habituationNone — data were collected before study began
Ecological ValidityVariable — depends on whether items reflect real-world constructsHigh for naturalistic; lower for lab-based observationHigh — data reflect real-world events as originally documented
Access to Internal StatesExcellent — the only method that directly taps thoughts, feelings, beliefsPoor — inferences about internal states from behavior are indirectPoor — limited to what was documented; internal states rarely recorded
Cost & EfficiencyLow cost per participant; scalable via online platformsHigh cost — training coders, recording equipment, time-intensive codingLow marginal cost once data access is secured; initial access may be difficult
Sample SizeLarge samples feasible (thousands via online survey)Typically small due to labor demandsVery large — national databases may contain millions of records
Researcher ControlHigh — researcher designs items, format, and samplingModerate — can control setting (structured) or not (naturalistic)None — researcher cannot alter what was recorded
Primary Bias ThreatSocial desirability, recall bias, acquiescenceObserver expectancy, drift, and Hawthorne effectSelective deposit, selective survival, coding changes over time
KEY TAKEAWAY
No single data collection method is universally superior; each represents a strategic trade-off. The EPPP tests your ability to identify the most appropriate method for a given scenario. A useful heuristic: if the question asks about internal experiences, lean toward surveys; if it asks about behavioral processes, lean toward observation; if it asks about population-level trends or historical patterns, lean toward archival data. When the item asks for the strongest design, the answer is often a multi-method approach.

Connection to Advanced Methodological Concepts

An understanding of data collection methods connects directly to several advanced topics that appear elsewhere on the EPPP. The concept of construct validity is fundamentally a question about whether the chosen data collection method actually captures the latent construct of interest. When multiple methods converge on the same finding—a phenomenon studied under the framework of convergent validity within the multitrait-multimethod (MTMM) matrix proposed by Campbell and Fiske (1959)—confidence in construct validity increases. Conversely, when a measure correlates more strongly with method-similar but construct-dissimilar measures than with method-dissimilar but construct-similar measures, method variance is said to be inflating the findings—a problem particularly acute when a study relies on a single data collection method.

How foundational data collection concepts connect to advanced EPPP topics.
Basic ConceptAdvanced Extension
Reactivity in surveysDemand characteristics and experimenter expectancy effects (Orne, 1962; Rosenthal, 1966)
Inter-rater reliability (κ)Generalizability theory (G-theory): partitioning variance across observers, occasions, and items simultaneously
Triangulation across methodsMultitrait-multimethod (MTMM) matrix for evaluating convergent and discriminant validity
Archival trend analysisInterrupted time-series designs and cohort-sequential analysis for causal inference from archival data
Ecological momentary assessment (EMA)Hybrid survey-observation method using multilevel modeling to analyze within-person fluctuations

As you advance in your preparation, recognize that modern behavioral health research increasingly relies on hybrid and digital methods—ecological momentary assessment (EMA), wearable biosensors, and natural language processing of clinical notes—that blur traditional boundaries between survey, observational, and archival methods. Nevertheless, the core evaluative framework you have learned here remains the foundation for appraising any data collection strategy, no matter how technologically novel.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher studies attitudes toward involuntary psychiatric hospitalization by mailing a Likert-scale questionnaire to 2,000 mental health professionals. Which type of bias is MOST likely to threaten the validity of the results?
PROBLEM 2BASIC APPLICATION
Two observers independently code aggressive behaviors in a sample of children during recess. Observer A records 40 aggressive incidents and Observer B records 35 aggressive incidents. They agree on 30 of those incidents. If chance agreement (Pₑ) is estimated at 0.25, what is the approximate Cohen's kappa (κ)?
PROBLEM 3INTERMEDIATE
A health psychologist wants to examine whether childhood adversity predicts adult cardiovascular disease over a 30-year period. She decides to use archival data from a longitudinal birth cohort study (e.g., the Dunedin Multidisciplinary Health and Development Study). Identify TWO specific validity threats inherent to this archival approach and suggest one safeguard for each.
PROBLEM 4APPLIED
A forensic psychologist is hired to evaluate a diversion program for justice-involved individuals with serious mental illness. The program claims to reduce recidivism and improve psychiatric symptom severity. Design a multi-method evaluation plan, specifying which data collection method you would use for each outcome and explaining your rationale.
PROBLEM 5CRITICAL THINKING
A research team publishes a study finding that self-reported mindfulness (measured by the MAAS questionnaire) is positively associated with self-reported emotional regulation (measured by the DERS). Critics argue that the association is inflated by shared method variance. Explain this critique using the concepts of convergent validity and the multitrait-multimethod matrix, and propose a redesign that would address the concern.

Summary — Data Collection Methods in Behavioral Health Research

Three major families of data collection—survey (self-report), observational (direct behavior coding), and archival (pre-existing records)—each offer distinct advantages and introduce characteristic validity threats. Surveys provide unmatched access to internal states (attitudes, beliefs, subjective experiences) but are vulnerable to social desirability, recall bias, and nonresponse bias. Observational methods capture actual behavior in context with high ecological validity but require extensive training for adequate inter-rater reliability (κ) and are susceptible to the Hawthorne effect. Archival methods are uniquely nonreactive and can span large populations and long time periods, but face threats of selective deposit, selective survival, and measurement non-equivalence.

For the EPPP, remember that the optimal design frequently involves methodological triangulation—using two or more methods to measure the same construct—which strengthens construct validity by separating true construct variance from shared method variance. Match the data collection method to the nature of the construct: subjective states call for surveys, overt behaviors call for observation, and population-level trends call for archival analysis. When an exam item describes a study using only one method to measure multiple constructs, look for answer choices that identify common method bias as a limitation.

Varsity Tutors • EPPP: Part 1, Knowledge • Data Collection Methods — Evaluate strengths and weaknesses of survey, observational, and archival data methods