Historical Context & Motivation
The question of how best to gather information about human behavior has preoccupied researchers since the emergence of psychology as a formal discipline. Early investigators relied almost exclusively on clinical case notes and introspective reports—approaches that were deeply intertwined with the observer's theoretical commitments and therefore difficult to replicate. As the field matured, the demand for systematic, reproducible methods gave rise to three broad families of data collection: surveys and self-report instruments, direct observational procedures, and the use of pre-existing archival records. Each method crystallized in response to specific limitations of the others, and understanding their historical trajectory illuminates why no single approach dominates modern behavioral health research.
This historical progression reveals a central tension in behavioral health research: how do we collect data that are simultaneously rich, accurate, efficient, and ethically defensible? Each method represents a different resolution of that tension, trading off internal validity, external validity, feasibility, and participant burden in characteristic ways. For the EPPP, you need to evaluate these trade-offs fluently, matching data collection strategies to research questions and identifying the threats to validity that each method introduces.
Core Principles & Definitions
Before comparing the three major data collection methods, it is essential to anchor the discussion in several foundational concepts that cut across all research designs in behavioral health. These principles provide the evaluative framework through which one judges whether a given method is appropriate for a particular research question.
Reactivity
Ecological Validity
Response Bias
Observer Bias
Nonreactive (Unobtrusive) Measurement
Visual Explanation — The Three Methods at a Glance
The diagram above highlights a complementary pattern: where one method is strong, another tends to be weak. Survey methods excel at accessing internal states (thoughts, feelings, attitudes) but are vulnerable to participants' willingness and ability to report accurately. Observational methods capture behavior as it actually unfolds, yet the observer's presence can alter the very behavior under study—a phenomenon known as the Hawthorne effect. Archival methods sidestep reactivity entirely because the data were generated before the research question existed, but the investigator sacrifices control over what was recorded and how consistently the records were maintained. This trade-off structure is the conceptual backbone of the EPPP's data-methods questions.
How Each Method Works — Deep Dive
Survey Methods: From Questionnaires to Structured Interviews
Survey research encompasses any procedure in which participants provide verbal or written responses to a predetermined set of questions. The spectrum ranges from mailed paper questionnaires and online forms to face-to-face structured clinical interviews such as the SCID-5. A key distinction exists between closed-ended items (Likert scales, multiple-choice), which yield quantitative data amenable to statistical analysis, and open-ended items, which produce richer qualitative data but require content analysis or thematic coding. In behavioral health contexts, validated self-report instruments—such as the PHQ-9 for depression screening or the GAD-7 for generalized anxiety—are widely used because their psychometric properties (reliability, sensitivity, specificity) have been extensively documented.
The primary threats to survey validity include social desirability bias (respondents present themselves favorably), acquiescence bias (tendency to agree with statements regardless of content), recall bias (inaccurate memory for past events), and nonresponse bias (systematic differences between responders and non-responders). Each of these threats can be partially mitigated through careful instrument design—for example, embedding validity scales (as in the MMPI-2) or using forced-choice formats.
Observational Methods: Naturalistic, Participant, and Structured
Observational methods involve trained researchers directly watching and coding behavior. Three major subtypes exist. Naturalistic observation occurs in the participant's everyday environment without intervention—for example, coding parent-child interactions at a playground. Participant observation involves the researcher embedding within the group being studied, a method with deep roots in anthropology and qualitative psychology. Structured (or systematic) observation uses a controlled setting and a detailed behavioral coding scheme to maximize reliability. Sampling strategies—time sampling, event sampling, and interval recording—determine which behaviors are recorded and when, and the choice among them can substantially affect the resulting data.
Central to high-quality observational research is inter-rater reliability, typically quantified using Cohen's kappa (κ) for categorical codes or intraclass correlation coefficients (ICC) for continuous ratings. A κ value above .80 is generally considered excellent agreement, while values between .60 and .79 represent substantial agreement. Without demonstrating adequate inter-rater reliability, observational findings are viewed as insufficiently objective.
Archival Methods: Clinical Records, Databases, and Historical Documents
Archival research analyzes data that already exist, having been collected for purposes other than the current study. In behavioral health, common archival sources include electronic health records (EHRs), insurance claims databases, court records, school disciplinary logs, and national datasets such as the National Comorbidity Survey Replication (NCS-R) or the Behavioral Risk Factor Surveillance System (BRFSS). The researcher's task is to extract, code, and analyze relevant variables from these pre-existing repositories.
Two archival-specific validity threats deserve attention. Selective deposit refers to the fact that not all events are recorded—certain populations or diagnoses may be systematically underrepresented in the records. Selective survival refers to the differential preservation of records over time; older or less-valued documents may be lost or destroyed, biasing the surviving sample. Additionally, coding conventions and diagnostic criteria change over time (e.g., the transition from DSM-IV to DSM-5), which can introduce measurement non-equivalence in longitudinal archival analyses.
Detailed Classification — Threats, Safeguards, and Decision Criteria
The flowchart above provides a simplified decision heuristic, but real-world research often demands a more nuanced evaluation. Consider, for example, that a clinical researcher interested in substance use frequency might use a self-report screening tool (survey), verify consumption through urinalysis records (archival), and observe social drinking behavior in a controlled bar-lab setting (observation). This multi-method or triangulation approach is widely regarded as the gold standard because the weaknesses of one method are offset by the strengths of another. On the EPPP, questions frequently present a scenario and ask which method—or combination of methods—would maximize validity for that particular research aim.
Worked Example — Evaluating Data Collection in a Behavioral Health Study
Imagine a research team at a community mental health center wants to study the relationship between therapist alliance and treatment dropout among clients diagnosed with borderline personality disorder (BPD). The following worked example walks through how to evaluate each data collection method for this research question.
Comparative Strengths, Weaknesses, and Mitigation Strategies
| Criterion | Survey | Observational | Archival |
|---|---|---|---|
| Reactivity | High — participants know they are being studied and may alter responses | Moderate — Hawthorne effect can alter behavior, mitigated by habituation | None — data were collected before study began |
| Ecological Validity | Variable — depends on whether items reflect real-world constructs | High for naturalistic; lower for lab-based observation | High — data reflect real-world events as originally documented |
| Access to Internal States | Excellent — the only method that directly taps thoughts, feelings, beliefs | Poor — inferences about internal states from behavior are indirect | Poor — limited to what was documented; internal states rarely recorded |
| Cost & Efficiency | Low cost per participant; scalable via online platforms | High cost — training coders, recording equipment, time-intensive coding | Low marginal cost once data access is secured; initial access may be difficult |
| Sample Size | Large samples feasible (thousands via online survey) | Typically small due to labor demands | Very large — national databases may contain millions of records |
| Researcher Control | High — researcher designs items, format, and sampling | Moderate — can control setting (structured) or not (naturalistic) | None — researcher cannot alter what was recorded |
| Primary Bias Threat | Social desirability, recall bias, acquiescence | Observer expectancy, drift, and Hawthorne effect | Selective deposit, selective survival, coding changes over time |
Connection to Advanced Methodological Concepts
An understanding of data collection methods connects directly to several advanced topics that appear elsewhere on the EPPP. The concept of construct validity is fundamentally a question about whether the chosen data collection method actually captures the latent construct of interest. When multiple methods converge on the same finding—a phenomenon studied under the framework of convergent validity within the multitrait-multimethod (MTMM) matrix proposed by Campbell and Fiske (1959)—confidence in construct validity increases. Conversely, when a measure correlates more strongly with method-similar but construct-dissimilar measures than with method-dissimilar but construct-similar measures, method variance is said to be inflating the findings—a problem particularly acute when a study relies on a single data collection method.
| Basic Concept | Advanced Extension |
|---|---|
| Reactivity in surveys | Demand characteristics and experimenter expectancy effects (Orne, 1962; Rosenthal, 1966) |
| Inter-rater reliability (κ) | Generalizability theory (G-theory): partitioning variance across observers, occasions, and items simultaneously |
| Triangulation across methods | Multitrait-multimethod (MTMM) matrix for evaluating convergent and discriminant validity |
| Archival trend analysis | Interrupted time-series designs and cohort-sequential analysis for causal inference from archival data |
| Ecological momentary assessment (EMA) | Hybrid survey-observation method using multilevel modeling to analyze within-person fluctuations |
As you advance in your preparation, recognize that modern behavioral health research increasingly relies on hybrid and digital methods—ecological momentary assessment (EMA), wearable biosensors, and natural language processing of clinical notes—that blur traditional boundaries between survey, observational, and archival methods. Nevertheless, the core evaluative framework you have learned here remains the foundation for appraising any data collection strategy, no matter how technologically novel.
Practice Problems
Summary — Data Collection Methods in Behavioral Health Research
Three major families of data collection—survey (self-report), observational (direct behavior coding), and archival (pre-existing records)—each offer distinct advantages and introduce characteristic validity threats. Surveys provide unmatched access to internal states (attitudes, beliefs, subjective experiences) but are vulnerable to social desirability, recall bias, and nonresponse bias. Observational methods capture actual behavior in context with high ecological validity but require extensive training for adequate inter-rater reliability (κ) and are susceptible to the Hawthorne effect. Archival methods are uniquely nonreactive and can span large populations and long time periods, but face threats of selective deposit, selective survival, and measurement non-equivalence.
For the EPPP, remember that the optimal design frequently involves methodological triangulation—using two or more methods to measure the same construct—which strengthens construct validity by separating true construct variance from shared method variance. Match the data collection method to the nature of the construct: subjective states call for surveys, overt behaviors call for observation, and population-level trends call for archival analysis. When an exam item describes a study using only one method to measure multiple constructs, look for answer choices that identify common method bias as a limitation.