EPPP: PART 1, KNOWLEDGE • DOMAIN 5: ASSESSMENT AND DIAGNOSIS

Assessment Technology — Evaluate validity and ethical considerations in technology-based testing

Examining how digital innovations transform psychological assessment while raising critical validity and ethical challenges.

Historical Context & Motivation

The integration of technology into psychological assessment did not emerge overnight; rather, it evolved across several decades as computing power expanded and the behavioral health field increasingly demanded efficiency, standardization, and accessibility. Early psychological tests were administered exclusively through paper-and-pencil formats, with clinicians manually scoring protocols and interpreting results. While these methods established foundational psychometric principles, they were time-intensive, susceptible to clerical errors, and constrained by geographic proximity between examiner and examinee. The advent of computer-based testing (CBT) and later internet-based assessment offered solutions to many of these limitations, but simultaneously introduced novel concerns about test validity, data security, and equitable access that continue to shape the field today.

1960s
Mainframe-Based Scoring
Psychologists began using mainframe computers at universities and medical centers to score tests such as the MMPI, reducing clerical errors and enabling large-scale data aggregation for normative studies.
1980s
Computerized Adaptive Testing Emerges
Computerized adaptive testing (CAT) algorithms, grounded in Item Response Theory (IRT), allowed tests to dynamically select items based on examinee performance, shortening test length while preserving measurement precision.
1999
APA Guidelines on Computer-Based Testing
The American Psychological Association published guidelines addressing the equivalence of computerized and traditional administrations, establishing formal standards for validity evidence in technology-based formats.
2007
Telehealth Assessment Growth
With the expansion of telehealth, remote psychological assessment via videoconferencing and web-based platforms became more common, raising new ethical questions around informed consent, proctoring, and cross-jurisdictional licensure.
2020–Present
AI-Augmented and Remote Assessment
The COVID-19 pandemic accelerated adoption of remote assessment tools. Artificial intelligence began influencing automated scoring of narrative responses, behavioral coding from video, and predictive risk algorithms, sparking intense ethical debate about algorithmic bias and transparency.

This historical trajectory reveals a central question that persists in contemporary behavioral health practice: When we change the medium through which a psychological test is delivered, do we change what the test measures? Answering this question requires a rigorous understanding of validity evidence, psychometric equivalence, and the ethical frameworks that govern responsible technology use in clinical and forensic contexts.

Core Principles & Definitions

Before evaluating any technology-based assessment, clinicians must understand the foundational concepts that undergird sound psychometric practice. The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) conceptualizes validity not as a binary property of a test but as an ongoing, evidence-based argument about the appropriateness of score interpretations for a given use and population. Technology-based testing introduces additional sources of variance—such as hardware differences, internet connectivity, and user familiarity with digital interfaces—that must be accounted for within this validity framework. The following principles serve as conceptual anchors throughout this lesson.

1

Construct Validity

The degree to which a test measures the theoretical construct it claims to measure. When tests move online, extraneous variables—such as computer anxiety or digital literacy—may introduce construct-irrelevant variance, threatening the meaning of test scores.
2

Measurement Equivalence

Also called mode equivalence, this refers to whether scores obtained via technology-based administration are comparable to those obtained via traditional methods. Equivalence must be demonstrated empirically, not assumed.
3

Informed Consent in Digital Contexts

Ethical guidelines (APA Ethics Code Standard 9.03) require that examinees understand the nature, purpose, and limitations of assessment. In technology-based testing, this extends to disclosing data storage, potential breaches, and algorithmic decision-making processes.
4

Equity and Access

The digital divide refers to disparities in technology access across socioeconomic, racial, geographic, and age groups. Ethical practice demands that technology-based testing does not systematically disadvantage certain populations.
5

Data Security and Confidentiality

Technology-based testing generates digital records that must comply with HIPAA, FERPA, and state-specific regulations. Encryption, secure servers, and access controls are ethical necessities, not optional features.
KEY TAKEAWAY
Think of a psychological test as a recipe that was originally designed for a specific kitchen. Moving that recipe to a different kitchen—perhaps one with different ovens, utensils, and altitude—may change the final product in subtle but meaningful ways. Mode equivalence research is the process of carefully verifying that the dish tastes the same regardless of the kitchen. Without this verification, you cannot be confident that you are measuring the same construct.

Visual Explanation — Validity Threats in Technology-Based Assessment

The following diagram maps the relationship between traditional validity evidence categories—as outlined in the Standards (2014)—and the specific threats introduced by technology-based testing modalities. Each arrow indicates a pathway through which a technological factor can compromise a particular form of validity evidence. Understanding these pathways is essential for EPPP examinees, who must evaluate whether a given technology-based assessment meets professional standards.

This diagram illustrates how four categories of validity evidence (top row) are each vulnerable to specific technology-related threats (middle rows). Dashed arrows connect each evidence type to its most relevant threat sources. The bottom panel emphasizes the clinical implication: validity evidence is modality-specific and must be reestablished when the administration format changes.

As the diagram makes clear, technological mediation does not affect a single dimension of validity in isolation. For example, computer anxiety primarily threatens internal structure evidence by introducing a secondary dimension to what should be a unidimensional measure, but it also compromises response process evidence because the examinee's cognitive strategy shifts from engaging with item content to managing frustration with the interface. Clinicians preparing for the EPPP must recognize these interconnections and be prepared to evaluate whether a test publisher has provided sufficient evidence that technology-specific threats have been mitigated for the population and setting in question.

How Technology Affects Measurement — Mechanisms and Frameworks

Construct-Irrelevant Variance (CIV) in Digital Formats

One of the most critical psychometric concepts for evaluating technology-based assessments is construct-irrelevant variance (CIV). CIV occurs when extraneous factors systematically influence test scores in ways unrelated to the target construct. In traditional assessment, CIV might stem from poor lighting or ambient noise. In technology-based assessment, CIV can arise from factors such as typing speed on performance-based measures, differential screen sizes altering item presentation, or internet latency causing response timing artifacts. The formal decomposition of observed score variance in technology-based testing can be expressed as follows.

OBSERVED SCORE DECOMPOSITION
X = T + E_random + E_tech
Where X = observed score, T = true score on the target construct, E_random = random measurement error, and E_tech = systematic error attributable to the technology medium. The goal of mode equivalence research is to demonstrate that E_tech ≈ 0.

Mode Equivalence Testing

Establishing mode equivalence requires more than simply correlating scores across formats. Researchers employ a hierarchy of equivalence criteria. At the most basic level, mean score equivalence ensures that group-level averages do not differ meaningfully across administration modes. More stringently, rank-order equivalence requires that individuals maintain their relative standing across formats, typically evaluated through high correlations (r ≥ .90). The most demanding standard, measurement invariance, uses confirmatory factor analysis (CFA) to demonstrate that factor loadings, intercepts, and residual variances are equivalent across modes. The International Test Commission (ITC) recommends that publishers provide evidence at the measurement invariance level before claiming mode equivalence.

MEASUREMENT INVARIANCE — SCALAR LEVEL
Y_ij = τ_j + λ_j × η_i + ε_ij
Where Y_ij = observed score for person i on item j, τ_j = item intercept, λ_j = factor loading, η_i = latent trait score, and ε_ij = residual. Scalar invariance requires that both λ_j and τ_j remain constant across administration modes.

Computerized Adaptive Testing (CAT)

Computerized adaptive testing (CAT) represents a distinct technology-driven assessment paradigm grounded in Item Response Theory (IRT). Unlike fixed-form tests, CAT algorithms select subsequent items based on the examinee's responses to prior items, converging on a precise ability estimate with fewer items. The probability that an examinee with ability θ endorses item j correctly is governed by the item characteristic curve, often modeled using the two-parameter logistic function.

TWO-PARAMETER IRT MODEL
P(X_j = 1 | θ) = 1 / (1 + e^(−a_j(θ − b_j)))
Where a_j = discrimination parameter for item j, b_j = difficulty parameter, and θ = examinee ability estimate. Higher discrimination values indicate items that better differentiate between adjacent ability levels.
Clinical Implication
Because CAT tests administer different items to different examinees, traditional internal consistency reliability estimates (e.g., Cronbach's α) are not appropriate. Instead, CAT precision is reported as conditional standard error of measurement (CSEM) at each ability level, and clinicians should evaluate whether the CSEM is sufficiently small at clinical decision points (e.g., diagnostic cutoff scores).

Ethical Considerations in Technology-Based Testing

Ethical practice in technology-based assessment is governed by multiple overlapping frameworks, including the APA Ethical Principles of Psychologists and Code of Conduct (particularly Standards 9.01–9.11), the International Test Commission (ITC) Guidelines on Computer-Based and Internet-Delivered Testing, and HIPAA regulations governing protected health information. These sources converge on several key ethical domains that clinicians must address when employing technology-based assessments in behavioral health settings.

The ethical decision framework illustrates six interconnected domains that converge on ethical practice (center). Each domain connects via dashed lines to the central hub, emphasizing that neglecting any single domain compromises the ethical integrity of the entire assessment process.

Key Ethical Standards in Detail

Key APA ethical standards and their technology-specific applications
Ethical DomainAPA StandardTechnology-Specific Requirement
Informed Consent9.03Disclose that test data will be stored digitally, explain automated scoring procedures, notify of recording/proctoring tools, and describe data breach protocols.
Competence2.01, 2.04Clinician must understand how the technology may affect test validity. Using a computer-administered test without knowledge of mode equivalence data constitutes practicing outside one's competence.
Test Security9.11Prevent unauthorized copying, screen capture, or redistribution of test items. Use encrypted platforms and monitor item exposure rates.
Fairness9.02, 9.06Ensure technology does not introduce bias; provide alternative formats for individuals with disabilities or limited technology access; review algorithms for differential impact.
Data Protection4.01, 6.02Comply with HIPAA for PHI; use end-to-end encryption; establish data retention and destruction policies; conduct regular security audits.
EPPP Alert: Algorithmic Bias
A growing area of concern is the use of machine learning algorithms in automated test interpretation. If training data over-represents certain demographic groups, the algorithm may produce systematically biased results for underrepresented populations. Psychologists have an ethical obligation to evaluate such tools for differential validity before deploying them in clinical decision-making.

Worked Example — Evaluating a Technology-Based Assessment

Consider the following clinical scenario: A psychologist in a rural behavioral health clinic is considering adopting a web-based version of a well-validated depression screening instrument (originally developed for paper-and-pencil administration) to serve patients who cannot travel to the clinic. The psychologist must evaluate both the validity evidence and the ethical considerations before implementing this technology-based assessment.

Evaluating a Web-Based Depression Screener
1
Step 1 — Review Mode Equivalence EvidenceThe psychologist examines the test publisher's technical manual and published literature for studies comparing scores on the web-based version with the paper-and-pencil version. She finds a study (N = 420) reporting a correlation of r = .92 between modes, no significant mean difference (d = 0.08), and confirmatory factor analysis supporting configural and metric invariance but not scalar invariance. This means factor loadings are equivalent, but item intercepts differ slightly across modes.
Partial mode equivalence established—adequate for screening but not for precise score comparisons across formats.
2
Step 2 — Assess Construct-Irrelevant VarianceThe psychologist considers her patient population. Many are older adults with limited computer experience. Research suggests that computer anxiety can inflate scores on negative affect measures by 0.3–0.5 standard deviations in older adults. She notes that the publisher's validation study used college undergraduates (mean age 21), which does not match her clinical population.
Significant CIV risk identified—validation sample does not match target population.
3
Step 3 — Evaluate Informed Consent RequirementsThe psychologist drafts a technology-specific informed consent addendum that discloses: (a) the test will be administered via a third-party platform, (b) responses will be encrypted and stored on HIPAA-compliant servers, (c) the platform uses automated scoring but the clinician will interpret all results, and (d) the patient may request a paper-and-pencil alternative at no disadvantage.
Informed consent requirements met—technology-specific disclosures included.
4
Step 4 — Address Equity and Access ConcernsSeveral patients lack reliable internet access. The psychologist establishes a protocol providing clinic-based tablets for patients who cannot complete the assessment at home, along with a brief orientation session to familiarize patients with the digital interface. She also ensures the platform meets ADA accessibility standards for patients with visual or motor impairments.
Equity concerns addressed through alternative access protocols and ADA compliance.
5
Step 5 — Formulate Clinical DecisionIntegrating all evidence, the psychologist concludes that the web-based screener may be used as a first-level screening tool with appropriate caveats. She will not use web-based scores interchangeably with paper-and-pencil norms for diagnostic decisions. She will document in each patient's record: the administration mode, any observed technology-related difficulties, and the rationale for score interpretation. Patients who screen positive will receive a comprehensive in-person diagnostic evaluation.
Decision: Adopt for screening with documented limitations; do not use for diagnostic determination without corroborating in-person assessment.

Strengths and Limitations of Technology-Based Assessment

Technology-based assessment offers substantial advantages to behavioral health practice, but these advantages must be weighed against genuine limitations. The EPPP requires candidates to evaluate both sides of this equation with clinical sophistication, recognizing that the appropriateness of any technology-based tool depends on the specific use case, population, and setting. The following comparison synthesizes the current literature.

Comparative strengths and limitations of technology-based assessment in behavioral health
DimensionStrengthsLimitations
EfficiencyCAT reduces test length by 30–60% while maintaining precision; automated scoring eliminates clerical errors and returns results in seconds.Technology failures (server crashes, connectivity loss) can invalidate entire sessions; troubleshooting requires technical support infrastructure.
AccessibilityRemote administration reaches underserved rural populations and homebound patients; multilingual interfaces can be updated easily.Digital divide disadvantages older adults, low-SES individuals, and those with disabilities; hardware requirements may be prohibitive.
StandardizationComputer administration ensures uniform item presentation, precise timing, and consistent instructions across all examinees.Environmental variability in unsupervised settings (noise, distractions, assistance from others) undermines standardization benefits.
Data QualityReal-time data capture enables response time analysis, pattern detection, and immediate flagging of invalid protocols.Data security risks (breaches, hacking) create ethical and legal liabilities; metadata (IP address, geolocation) raises privacy concerns.
ValidityLarge-scale digital data collection facilitates rapid norming studies and ongoing validity monitoring.Mode equivalence cannot be assumed; validation studies often use convenience samples that may not represent clinical populations.
KEY TAKEAWAY
Consider technology-based assessment tools as you would a new medical device: they may offer improved precision and convenience, but they require independent validation for each intended use, population, and clinical context. Just as a surgical robot validated for cardiac surgery cannot be assumed safe for neurosurgery, a depression screener validated online with college students cannot be assumed equivalent when administered to elderly patients in a rural clinic. The burden of proof for validity and ethical compliance always rests with the clinician deploying the tool.

Connection to Advanced Theory — AI, Telehealth, and Future Directions

The current landscape of technology-based assessment is rapidly evolving beyond computerized versions of traditional tests toward fundamentally new assessment paradigms. Artificial intelligence (AI) is being applied to analyze speech patterns for depression markers, code facial expressions for emotional regulation assessment, and generate automated narrative interpretations of personality inventories. Ecological momentary assessment (EMA) uses smartphone-based prompts to collect real-time symptom data in natural environments, offering temporal granularity impossible with traditional single-session testing. Virtual reality (VR) environments are being developed to assess PTSD responses, social anxiety, and executive functioning in immersive simulated contexts. Each of these innovations raises questions that extend well beyond classical mode equivalence research.

Contrasting current and emerging technology-based assessment paradigms
Current Assessment TechnologyEmerging Assessment Technology
Computerized versions of existing paper tests; mode equivalence is the central validity question.AI-driven novel measures (e.g., digital phenotyping); construct definition itself is the central validity question—what exactly is being measured?
Clinician interprets all scores; automated scoring is limited to objective rules.Machine learning generates interpretive narratives and risk predictions; clinician's role shifts to evaluating algorithmic output.
Single-session assessment in controlled environment.Continuous passive data collection via smartphones and wearables; assessment becomes ongoing rather than episodic.
Ethical concerns focus on informed consent, test security, and data storage.Ethical concerns expand to include algorithmic transparency, consent for passive data collection, ownership of behavioral data, and liability for AI-generated errors.

For EPPP preparation, it is essential to recognize that while these advanced technologies hold considerable promise, they remain in early stages of psychometric validation. The fundamental principles of this lesson—that validity evidence must be established for each specific use and that ethical obligations intensify rather than diminish as technology becomes more complex—apply with even greater force to these emerging tools. The clinician's responsibility to maintain interpretive authority, evaluate evidence critically, and protect client welfare remains the bedrock of ethical assessment practice regardless of how sophisticated the technology becomes.

Practice Problems

PROBLEM 1CONCEPTUAL
A test publisher releases an online version of a well-established anxiety inventory and states that the test 'measures the same construct regardless of administration format.' What type of evidence would be most critical to support this claim, and why is the publisher's statement, standing alone, insufficient?
PROBLEM 2BASIC APPLICATION
A psychologist plans to use a computerized adaptive test (CAT) for cognitive screening. She calculates Cronbach's alpha from one administration session and obtains α = .68. Should she conclude that the test has inadequate reliability? Explain your reasoning.
PROBLEM 3INTERMEDIATE
A mode equivalence study for a personality inventory finds that the correlation between paper and online versions is r = .93, but the online version yields mean scores that are 0.4 SD higher on the Neuroticism scale. The study establishes metric invariance but not scalar invariance. How should a clinician interpret these findings when deciding whether to use online-derived norms?
PROBLEM 4APPLIED
A forensic psychologist is asked to conduct a court-ordered competency evaluation using a videoconference-administered neuropsychological battery. The test publisher has validated the instrument for in-person administration only. Identify at least three ethical concerns the psychologist must address and describe how each should be managed.
PROBLEM 5CRITICAL THINKING
A behavioral health system implements an AI-driven predictive algorithm that analyzes patients' electronic health records, self-report questionnaire responses, and smartphone usage patterns to generate suicide risk scores. The algorithm demonstrates strong predictive validity (AUC = .87) in the development sample. Critically evaluate this tool from both validity and ethical perspectives, addressing at least four distinct concerns.

Lesson Summary

Technology-based assessment has transformed behavioral health practice by enabling computerized adaptive testing, remote administration, and automated scoring and interpretation. However, clinicians must recognize that changing the administration medium does not preserve validity automatically. Mode equivalence must be empirically demonstrated through studies examining mean score equivalence, rank-order consistency, and—ideally—measurement invariance at the scalar level via confirmatory factor analysis. Construct-irrelevant variance from factors such as computer anxiety, digital literacy, and unsupervised testing environments can systematically distort scores if not addressed.

Ethical practice in technology-based assessment requires attention to six interconnected domains: informed consent (including technology-specific disclosures), clinician competence in understanding psychometric implications, test security, equity and access across the digital divide, data protection under HIPAA and related regulations, and ongoing clinical supervision of automated systems. As AI-augmented tools, ecological momentary assessment, and virtual reality platforms emerge, these foundational principles become even more critical. The clinician's ultimate obligation is to ensure that technology enhances—rather than undermines—the accuracy, fairness, and ethical integrity of psychological assessment.

Varsity Tutors • EPPP: Part 1, Knowledge • Assessment Technology — Evaluate validity and ethical considerations in technology-based testing