EPPP: PART 1, KNOWLEDGE • DOMAIN 5: ASSESSMENT AND DIAGNOSIS

Assessment Methods — Evaluate strengths and limitations of interviews, observation, self-report, and multi-informant data

Understanding how clinical data sources converge and diverge to shape accurate psychological assessment.

Historical Context & Motivation

The history of psychological assessment is, in many respects, the history of the discipline itself. Before the advent of standardized testing, clinicians relied almost exclusively on the clinical interview — an unstructured conversation designed to elicit diagnostic impressions. While this method offered flexibility and rapport, it also introduced considerable subjectivity, and the reliability of diagnoses across practitioners remained poor. The twentieth century witnessed a systematic effort to supplement interviews with more rigorous methods, each aimed at capturing different facets of human behavior and experience.

The development of formal behavioral observation methods in the 1920s and 1930s grew out of the behaviorist tradition, emphasizing what could be directly seen and measured. Concurrently, self-report inventories such as the Woodworth Personal Data Sheet (1917) introduced the idea that individuals could systematically report on their own internal states. By the late twentieth century, researchers increasingly recognized that no single informant or method could capture the full complexity of psychological functioning, giving rise to multi-informant assessment paradigms.

1917
Woodworth Personal Data Sheet
Robert Woodworth develops the first standardized self-report personality inventory to screen military recruits for shell shock susceptibility during World War I.
1943
MMPI Published
Hathaway and McKinley publish the Minnesota Multiphasic Personality Inventory, establishing empirical keying as the gold standard for self-report personality assessment and introducing validity scales to detect response bias.
1965
Structured Interview Development
The development of structured diagnostic interviews such as the PSE (Present State Examination) dramatically improves inter-rater reliability in psychiatric diagnosis, laying groundwork for the SCID and DISC.
1991
Achenbach Multi-Informant System
Thomas Achenbach formalizes the multi-informant approach with parallel forms of the CBCL, TRF, and YSR, demonstrating that cross-informant correlations are typically low to moderate and proposing methods to integrate discrepant data.
2013
DSM-5 and Dimensional Assessment
The DSM-5 introduces cross-cutting symptom measures and Level 2 assessments incorporating self-report and clinician ratings, reflecting a growing consensus that multi-method, multi-informant assessment yields the most valid diagnostic picture.

The central question driving the evolution of these methods is both simple and profound: How do we gather the most valid and reliable information about a person's psychological functioning, given that every method introduces its own sources of error? Understanding each method's strengths and limitations is essential for competent clinical practice and is a core competency tested on the EPPP.

Core Principles of Assessment Methods

Before evaluating specific assessment methods, it is important to understand the overarching principles that govern their use. Every assessment method can be evaluated along several dimensions, including its reliability (consistency of measurement), validity (accuracy in measuring what it purports to measure), clinical utility (practical value in guiding treatment decisions), and susceptibility to various forms of bias. These dimensions interact in complex ways: a method that maximizes ecological validity through naturalistic observation may sacrifice the standardization necessary for high inter-rater reliability.

1

Reliability

The degree to which an assessment yields consistent, reproducible results. Key types include inter-rater reliability (agreement among evaluators), test-retest reliability (stability over time), and internal consistency (coherence among items).
2

Validity

The extent to which a method measures the construct it claims to measure. Construct validity, criterion validity (predictive and concurrent), and content validity each reflect different facets of measurement accuracy.
3

Reactivity

The phenomenon by which the act of measurement itself alters the behavior being measured. Observation is particularly susceptible to reactivity, though self-report can also be affected when respondents adjust answers to perceived expectations.
4

Incremental Validity

The degree to which adding a new source of information (e.g., a teacher rating to a parent rating) improves prediction or classification accuracy beyond what existing data already provide. Central to the rationale for multi-informant assessment.
5

Method Variance

Variance in scores attributable to the assessment method itself rather than to the construct being measured. Same-method correlations tend to be inflated relative to cross-method correlations, a pattern described in Campbell and Fiske's multitrait-multimethod matrix.
KEY TAKEAWAY
Think of assessment methods like different camera angles in a film. A single camera captures one perspective — it may miss crucial action happening off-screen. Using multiple cameras (methods and informants) provides a composite view, but each angle has its own blind spots and distortions. The clinician's task is analogous to the film editor: integrating footage into a coherent narrative while recognizing artifacts introduced by each lens.

Visual Overview of Assessment Methods

This diagram illustrates how each assessment method feeds into an integrated clinical formulation. The top row shows the four primary methods; the middle section highlights their respective strengths; the bottom section notes their limitations. All streams converge at the bottom into a unified formulation where convergent data strengthen diagnostic confidence.

The diagram above captures a fundamental principle of clinical assessment: each method occupies a distinct vantage point. Interviews access subjective experience through dynamic conversation. Observation captures overt behavior in situ. Self-report instruments standardize the collection of internal experiences. And multi-informant data triangulates across perspectives. When data from multiple sources converge, clinicians can place greater confidence in their diagnostic conclusions; when they diverge, the discrepancy itself becomes clinically informative.

How Each Method Works — Mechanisms and Decision Points

The Clinical Interview

Clinical interviews exist along a continuum from fully unstructured to fully structured, with semi-structured formats occupying the middle ground. Unstructured interviews allow clinicians to pursue idiosyncratic leads and build therapeutic alliance, but they yield poor inter-rater reliability — studies suggest kappa values as low as 0.30 for some diagnostic categories. Structured interviews such as the SCID-5 or the MINI prescribe exact questions and decision trees, achieving kappa values frequently exceeding 0.70. Semi-structured interviews (e.g., the SCID-5-CV) provide standardized prompts with latitude for follow-up, balancing reliability with clinical flexibility.

Behavioral Observation

Behavioral observation can be naturalistic (conducted in the client's everyday environment) or analog (conducted in a controlled setting designed to simulate natural conditions). Observation systems typically employ coding schemes that specify target behaviors, recording methods (e.g., event recording, interval recording, time sampling), and operational definitions. Inter-observer agreement is the primary reliability metric, usually expressed as percentage agreement or Cohen's kappa. A critical threat to validity is reactivity: individuals who know they are being observed often modify their behavior, a phenomenon sometimes called the Hawthorne effect. Observer drift — the gradual, unintentional shift in how an observer applies coding criteria — is another systematic source of error.

Self-Report Measures

Self-report instruments ask individuals to endorse items describing their thoughts, feelings, or behaviors. They are among the most efficient and widely normed assessment tools in clinical psychology. However, their validity depends on several assumptions: that the respondent possesses adequate insight into their own functioning, sufficient reading comprehension to interpret items accurately, and willingness to respond honestly. Common threats include social desirability bias (presenting oneself favorably), acquiescence (tendency to agree regardless of content), extreme responding, and malingering (deliberate fabrication or exaggeration of symptoms). Modern inventories such as the MMPI-3 include validity scales (e.g., L, F, K) specifically designed to detect these response styles.

Multi-Informant Assessment

Multi-informant assessment gathers data from two or more informants who observe the individual in different contexts. In child and adolescent assessment, this commonly involves parallel reports from parents, teachers, and the child. Achenbach's seminal meta-analysis found that the mean cross-informant correlation was approximately r = 0.28 for informants occupying different roles, underscoring that discrepancy is the norm rather than the exception. The Operations Triad Model (OTM) proposed by De Los Reyes and Kazdin offers a framework for understanding when and why informants disagree, distinguishing between discrepancies attributable to true contextual variation in behavior versus those reflecting informant bias. Understanding these patterns is critical because clinicians must decide how to weight conflicting information — a process that is more art than algorithm.

EPPP ALERT
The EPPP frequently tests the distinction between structured and unstructured interviews, particularly regarding their relative inter-rater reliability. Remember: structure increases reliability but may reduce the clinician's ability to explore idiosyncratic presentations. Semi-structured interviews represent the most commonly recommended compromise for research and clinical practice.

Detailed Classification of Methods and Error Sources

Bar chart illustrating Achenbach's meta-analytic findings on cross-informant correlations. Same-informant stability (test-retest) shows the highest correlation (r ≈ .64), while agreement between informants in different roles (e.g., parent vs. teacher) averages only r ≈ .28. The dashed red line marks this mean cross-informant correlation.

The bar chart above visually encodes one of the most important empirical findings in clinical assessment: informants who observe the same individual in similar contexts (e.g., two parents) tend to agree reasonably well (r ≈ .60), but those who observe the individual across different settings — such as a parent at home and a teacher at school — agree only modestly (r ≈ .28). This pattern is not merely a psychometric inconvenience; it reflects the genuine situational specificity of behavior. A child who is compliant in a structured classroom may be oppositional at home, and both reports are accurate within their respective contexts.

Common Error Sources Across Assessment Methods
Error TypeMethods AffectedDescription & Mitigation
Social DesirabilitySelf-Report, InterviewTendency to present oneself favorably. Mitigated by validity scales (e.g., MMPI L, K scales), forced-choice formats, and building rapport before sensitive questions.
ReactivityObservationBehavior changes because the individual knows they are being observed. Mitigated by habituation periods, unobtrusive recording, or participant observation.
Observer DriftObservationGradual, unintentional shift in how observers apply operational definitions over time. Mitigated by periodic recalibration, random reliability checks, and anchored coding manuals.
Halo EffectInterview, Multi-InformantGlobal impressions (positive or negative) contaminate ratings of specific attributes. Mitigated by using behaviorally anchored rating scales and assessing domains independently.
Confirmatory BiasInterview, ObservationTendency to seek or interpret information in ways that confirm preexisting hypotheses. Mitigated by structured protocols, considering alternative diagnoses, and using actuarial decision rules.
Informant BiasMulti-InformantEach informant filters observations through their own personality, psychopathology, and relationship with the client. A depressed mother may overreport child problems. Mitigated by assessing informant characteristics and using aggregation models.

Worked Example: Designing a Multi-Method Assessment Battery

Consider a referral question: A 10-year-old boy, Marcus, is referred by his teacher for disruptive behavior in the classroom. His mother reports no behavioral concerns at home. The goal is to determine whether Marcus meets criteria for ADHD and/or ODD, and to develop a treatment plan. Let us work through how a clinician might design and evaluate a multi-method, multi-informant assessment battery.

Multi-Method Assessment Battery for Marcus
1
Step 1 — Identify the Referral Question and Select MethodsThe referral question involves diagnostic clarification (ADHD vs. ODD vs. comorbid) and treatment planning. This calls for methods that assess both internalizing and externalizing symptoms across settings. The clinician selects: (a) a semi-structured diagnostic interview (KIDDIE-SADS) with the mother and Marcus; (b) teacher and parent rating scales (CBCL, TRF, Conners-3); (c) classroom observation using the BOSS (Behavioral Observation of Students in Schools); and (d) Marcus's self-report (YSR).
Four methods selected: semi-structured interview, rating scales (parent and teacher), direct observation, self-report
2
Step 2 — Evaluate the Strengths Each Method ContributesThe semi-structured interview provides diagnostic specificity and allows the clinician to probe for onset, duration, and impairment — critical DSM-5 criteria. The rating scales are normed, efficient, and provide T-scores for cross-setting comparison. The classroom observation yields direct behavioral data (e.g., on-task vs. off-task intervals) free from informant filtering. Marcus's self-report captures his subjective experience and any internalizing symptoms the adults might miss.
Each method addresses a different facet: diagnostic precision, norm-referenced comparison, direct behavior, and internal states
3
Step 3 — Anticipate Limitations and Error SourcesThe mother may underreport due to limited exposure to structured settings or due to a depressive attribution bias that normalizes disruptive behavior. The teacher may overreport due to the halo effect (Marcus is struggling academically, which colors her perception of his behavior). Classroom observation captures only a single time window and is subject to reactivity. Marcus, at age 10, may lack sufficient metacognitive skill to accurately report attention difficulties on the YSR.
Anticipated errors: informant bias (mother, teacher), reactivity (observation), limited insight (self-report)
4
Step 4 — Integrate Data and Interpret DiscrepanciesResults show: TRF Attention Problems T = 72 (clinical range), CBCL Attention Problems T = 58 (normative), BOSS on-task rate = 42% (well below classroom mean of 78%), YSR internalizing T = 66. The parent-teacher discrepancy is consistent with situational specificity — ADHD symptoms are more pronounced in structured environments demanding sustained attention. Marcus's elevated internalizing self-report suggests possible comorbid anxiety that neither adult informant detected.
Convergent evidence for school-based attention deficits; divergent data reveal possible internalizing comorbidity missed by adult informants
5
Step 5 — Formulate Diagnostic Impression and Treatment RecommendationsIntegrating across methods, the clinician concludes that Marcus presents with ADHD, Predominantly Inattentive Presentation, with possible comorbid generalized anxiety. The multi-informant discrepancy is consistent with the expected cross-setting pattern rather than an indication that one informant is wrong. Treatment recommendations include classroom-based behavioral interventions, psychoeducation for the family, and further evaluation of anxiety symptoms with targeted self-report measures (e.g., SCARED).
Final formulation: ADHD-PI with possible GAD comorbidity; discrepancy interpreted as clinically informative rather than erroneous

Comparative Strengths and Limitations

Comparative Strengths and Limitations of Assessment Methods
MethodKey StrengthsKey Limitations
Unstructured InterviewMaximal flexibility; builds rapport; excellent for exploring unique presentations; allows observation of nonverbal cues and affect regulation in session.Lowest inter-rater reliability; susceptible to confirmatory bias and primacy/recency effects; systematic coverage of diagnostic criteria not guaranteed.
Structured InterviewHigh inter-rater reliability (κ > .70 for many diagnoses); systematic criterion coverage; replicable across clinicians and settings.Time-intensive; may feel impersonal; limited room for follow-up; requires training; may miss presentations not covered by the instrument.
Semi-Structured InterviewBalances reliability with flexibility; standardized probes with room for clinical follow-up; widely accepted in both research and practice.Still time-intensive; inter-rater reliability depends on training quality; reliability lower than fully structured formats.
Naturalistic ObservationHigh ecological validity; captures actual behavior in context; does not require verbal ability or insight from the client.Reactivity; observer drift; labor-intensive; limited to overt behaviors; may capture atypical samples due to time constraints.
Self-Report InventoryEfficient; standardized norms; accesses internal states (cognitions, emotions, subjective distress); large-scale administration possible.Social desirability; response sets (acquiescence, extreme responding); requires literacy and insight; malingering and faking.
Multi-Informant DataCross-context coverage; incremental validity; discrepancies are clinically informative; reduces method variance through triangulation.Low mean cross-informant correlations (r ≈ .28); integration is complex; no gold standard for resolving discrepancies; increased cost and logistical burden.
KEY TAKEAWAY
No single assessment method is inherently superior to the others — each is optimized for a different measurement niche. Just as a research team benefits from methodological triangulation (using surveys, experiments, and qualitative interviews to converge on a finding), a clinician benefits from multi-method assessment. The EPPP rewards candidates who can articulate not only what each method does well, but also where each one falls short and how combining methods addresses those gaps.

Connecting to Advanced Psychometric Theory

The principles underlying multi-method, multi-informant assessment connect directly to two foundational psychometric frameworks: Campbell and Fiske's Multitrait-Multimethod (MTMM) Matrix and Generalizability Theory (G-Theory). The MTMM matrix, introduced in 1959, provided a systematic way to evaluate convergent and discriminant validity by examining correlations across traits and methods simultaneously. When same-trait, different-method correlations (convergent validity) exceed different-trait, same-method correlations (discriminant validity), the constructs being measured are well-defined and method effects are minimal.

Advanced Psychometric Frameworks Related to Assessment Methods
FrameworkKey ContributionRelation to Assessment Methods
MTMM MatrixSeparates trait variance from method variance; establishes criteria for convergent and discriminant validity.Explains why same-method correlations (e.g., two self-report scales) tend to be inflated: shared method variance inflates the apparent relationship between constructs.
Generalizability TheoryModels multiple sources of measurement error (facets) simultaneously; produces generalizability coefficients (G-coefficients).Allows clinicians and researchers to estimate how much variance in scores is attributable to persons, items, occasions, raters, and their interactions — directly informing decisions about how many informants or observations are needed.
Operations Triad ModelProvides a theoretical framework for predicting when and why informants will agree or disagree.Distinguishes measurement artifact from genuine contextual variation; guides clinicians in interpreting discrepant multi-informant data without defaulting to a 'one informant is right' heuristic.

Looking forward, advances in ecological momentary assessment (EMA) are beginning to blur traditional method boundaries. EMA uses smartphone-based self-report collected multiple times per day in naturalistic settings, combining the standardization of self-report with the ecological validity of observation. Similarly, passive sensing (e.g., GPS, accelerometry, voice analysis) offers the promise of continuous, unobtrusive behavioral measurement. These innovations do not eliminate the fundamental challenges of method variance and informant bias, but they expand the clinician's toolkit and create new opportunities for triangulation.

Practice Problems

PROBLEM 1CONCEPTUAL
A psychologist wants to maximize inter-rater reliability for a diagnostic evaluation of major depressive disorder. Which type of interview format would best achieve this goal, and why?
PROBLEM 2BASIC APPLICATION
A teacher rating on the CBCL-TRF yields a T-score of 70 on the Aggressive Behavior scale for an 8-year-old girl, while her mother's CBCL report yields a T-score of 55 on the same scale. Using Achenbach's meta-analytic findings, explain whether this discrepancy is surprising or expected.
PROBLEM 3INTERMEDIATE
A clinician is conducting a classroom observation of a student suspected of having ADHD. After three sessions, the student's on-task behavior appears to improve markedly. The clinician wonders if the student is genuinely improving or if a measurement artifact is responsible. Identify the most likely artifact and describe two strategies to mitigate it.
PROBLEM 4APPLIED
You are evaluating a 35-year-old man for possible antisocial personality disorder (ASPD) in a forensic setting. He has a strong motivation to deny symptoms. Design a multi-method assessment strategy that addresses the specific challenges of this referral context, identifying at least three methods and explaining what each uniquely contributes.
PROBLEM 5CRITICAL THINKING
A researcher proposes that because cross-informant correlations are typically low (r ≈ .28), clinicians should simply average scores across informants to arrive at the 'true' level of a child's symptoms. Critically evaluate this proposal using concepts from the lesson, including the Operations Triad Model, situational specificity, and incremental validity.

Summary

Psychological assessment draws on four primary data-gathering methods, each with distinct strengths and vulnerabilities. Clinical interviews range from unstructured (maximum flexibility, lowest reliability) to structured (maximum reliability, reduced flexibility), with semi-structured interviews representing the recommended balance for most clinical and research contexts. Behavioral observation provides direct, ecologically valid data on overt behavior but is vulnerable to reactivity and observer drift. Self-report inventories efficiently access internal states with normative comparisons but are susceptible to social desirability bias, acquiescence, and require adequate literacy and insight.

Multi-informant assessment provides cross-context data and incremental validity, though mean cross-informant correlations are low (r ≈ .28), reflecting the situational specificity of behavior rather than error alone. The MTMM matrix and Generalizability Theory provide the psychometric foundation for understanding method variance, while the Operations Triad Model guides interpretation of informant discrepancies. Competent clinical assessment integrates data across methods and informants, treating convergence as evidence of validity and divergence as clinically informative — not dismissible.

Varsity Tutors • EPPP: Part 1, Knowledge • Assessment Methods — Evaluate strengths and limitations of interviews, observation, self-report, and multi-informant data