EPPP: PART 2, SKILLS • DOMAIN 6: COLLABORATION, CONSULTATION, AND SUPERVISION

Program Evaluation — Evaluate service and program effectiveness

Systematic methods for determining whether behavioral health programs achieve their intended outcomes and warrant continued investment.

Historical Context & Motivation

The formal practice of program evaluation in behavioral health emerged from a broader movement in social science and public policy that sought to apply empirical reasoning to questions about whether government-funded programs were actually working. Before systematic evaluation methods existed, decisions about program continuation or termination were largely political, guided by anecdotal reports and stakeholder impressions rather than data. The rise of large-scale social programs in the mid-twentieth century—coupled with growing demand for accountability—created an urgent need for structured approaches to assessing program impact. In behavioral health specifically, the proliferation of community mental health centers, substance abuse treatment programs, and crisis intervention services made evaluation not merely desirable but ethically essential: clinicians and administrators needed to know whether the services they provided genuinely helped their clients or inadvertently caused harm.

1965
Great Society Programs & Accountability
President Johnson's Great Society legislation, including the Elementary and Secondary Education Act, mandated evaluation of federally funded programs. This created unprecedented demand for systematic program evaluation methodology across social services, including early behavioral health interventions.
1972
Scriven's Formative vs. Summative Distinction
Michael Scriven formalized the distinction between formative evaluation (conducted during program implementation to improve it) and summative evaluation (conducted after implementation to judge overall worth), establishing a conceptual framework that remains foundational to the field.
1986
Rossi & Freeman's Comprehensive Framework
Publication of Rossi and Freeman's influential textbook synthesized decades of evaluation theory into a hierarchical model encompassing needs assessment, process evaluation, outcome evaluation, and efficiency analysis—a framework widely adopted in behavioral health settings.
2001
Evidence-Based Practice Movement
The APA Presidential Task Force on Evidence-Based Practice underscored the integration of research evidence with clinical expertise. Program evaluation became a core competency for psychologists, linking service delivery to empirically supported outcomes.
2020s
Health Equity & Culturally Responsive Evaluation
Contemporary program evaluation in behavioral health increasingly incorporates culturally responsive evaluation frameworks, participatory methods, and health equity metrics, reflecting the field's growing commitment to addressing disparities in service access and outcomes.

This historical trajectory reveals a central, persistent question: How do we know whether a behavioral health program is truly effective, and for whom? Program evaluation provides the systematic tools to answer this question, moving beyond clinical intuition toward evidence-grounded decision-making about resource allocation, service design, and client welfare. For psychologists preparing for the EPPP, competency in program evaluation means understanding not only the technical methods of data collection and analysis, but also the ethical, cultural, and organizational contexts in which evaluation occurs.

Core Principles & Definitions

Program evaluation in behavioral health rests on several foundational principles that distinguish it from basic research. While both endeavors employ empirical methods, program evaluation is fundamentally applied—its purpose is to inform real-world decisions about service delivery, funding, and policy. Understanding these core principles equips the evaluator to design assessments that are not only methodologically sound but also practically useful and ethically responsible within the complex ecosystem of behavioral health services.

1

Logic Model Thinking

Every program evaluation begins with a logic model—a visual representation of the causal chain linking program inputs, activities, outputs, short-term outcomes, and long-term impact. The logic model makes the program's theory of change explicit and testable.
2

Formative vs. Summative Evaluation

Formative evaluation is conducted during program implementation to guide improvement. Summative evaluation occurs at or after program completion to determine overall effectiveness and inform continuation or termination decisions.
3

Stakeholder Engagement

Effective evaluation requires engaging diverse stakeholders—including clients, clinicians, administrators, funders, and community members—whose perspectives shape evaluation questions, methods, and the interpretation of findings.
4

Utilization-Focused Approach

Drawing on Michael Quinn Patton's framework, utilization-focused evaluation prioritizes the actual use of findings. An evaluation is only valuable if its results inform decision-making and lead to actionable improvements in service delivery.
5

Cultural Responsiveness

Evaluations must account for the cultural context of the populations served. Culturally responsive evaluation examines whether programs are equitable, accessible, and appropriate for diverse communities, attending to systemic barriers and cultural strengths.
KEY TAKEAWAY
Think of program evaluation like a GPS navigation system for behavioral health services. A GPS doesn't just tell you where you are—it tells you whether you're on the right route (formative evaluation), whether you've arrived at your destination (summative evaluation), how efficiently you traveled (cost-effectiveness analysis), and whether the route works equally well for all passengers (equity analysis). Without it, you might drive for hours believing you're making progress while actually heading in the wrong direction. Program evaluation provides the empirical feedback loop that keeps services on course toward meaningful client outcomes.

Visual Explanation — The Logic Model

The logic model is the conceptual backbone of program evaluation, providing a visual map of how a program is expected to produce change. Understanding this model is essential for designing evaluations because each component of the logic model corresponds to a specific type of evaluation question. The diagram below illustrates a generic logic model for a behavioral health program, showing the causal chain from resource inputs through to long-term community impact.

The logic model maps the causal chain from inputs (resources) through activities and outputs to outcomes and impact. Each component aligns with a specific evaluation type shown below, and the feedback loop illustrates how evaluation results inform program improvement.

Notice how each evaluation type corresponds to a specific segment of the logic model. A needs assessment occurs before the program begins, determining whether the target population actually requires the proposed services. Process evaluation examines whether activities are being implemented as planned and whether outputs meet expected benchmarks—this is the formative dimension. Outcome evaluation asks the summative question: did the program produce the intended changes in client functioning, symptom reduction, or quality of life? Finally, efficiency analysis compares the costs of inputs against the value of outcomes achieved, informing decisions about resource allocation and program sustainability. The feedback loop at the bottom represents the iterative nature of evaluation: findings from any stage should cycle back to refine the program's design and delivery.

How Program Evaluation Works — Design & Measurement

Designing a rigorous program evaluation requires decisions across multiple domains: selecting an appropriate evaluation design, choosing valid and reliable measures, determining data collection procedures, and planning for the analysis of results. In behavioral health settings, these decisions must also balance methodological rigor with practical constraints—ethical concerns about withholding treatment, limited budgets, small sample sizes, and the complex, multidimensional nature of behavioral health outcomes.

Evaluation Designs

The choice of evaluation design depends on the questions being asked and the level of causal inference required. Randomized controlled trials (RCTs) represent the gold standard for establishing causal relationships between program participation and outcomes, but they are often impractical or unethical in behavioral health contexts where denying services to a control group raises serious concerns. Quasi-experimental designs—such as non-equivalent control group designs and interrupted time series—offer a middle ground, providing stronger evidence than purely descriptive approaches while respecting practical limitations. Pre-experimental designs (e.g., one-group pretest-posttest) are common in practice but vulnerable to threats to internal validity such as maturation, history, and regression to the mean.

Quantitative Metrics in Outcome Evaluation

While program evaluation in behavioral health is not primarily a mathematical discipline, certain quantitative frameworks are essential for interpreting outcomes. Two metrics commonly used to assess program effectiveness are effect size and the Reliable Change Index (RCI). Effect sizes quantify the magnitude of change attributable to the program, while the RCI determines whether an individual client's change exceeds what would be expected from measurement error alone.

COHEN'S d EFFECT SIZE
d = (M₁ − M₂) / SD_pooled
Where M₁ is the mean of the treatment group, M₂ is the mean of the comparison group (or pretest mean), and SDpooled is the pooled standard deviation. Cohen's conventions: d = 0.2 (small), d = 0.5 (medium), d = 0.8 (large).
RELIABLE CHANGE INDEX (RCI)
RCI = (X₂ − X₁) / S_diff
Where X₁ is the pretest score, X₂ is the posttest score, and Sdiff = SD₁ × √(2(1 − rxx)). An RCI > 1.96 indicates statistically reliable change (p < .05), meaning the observed change is unlikely due to measurement error.
COST-EFFECTIVENESS RATIO
CER = Total Program Cost / Number of Successful Outcomes
This ratio expresses the cost per unit of outcome achieved (e.g., cost per client who achieves clinically significant improvement). Lower CER values indicate greater efficiency. This metric allows comparison across programs with similar outcome definitions.
📊 Mixed-Methods Designs
Behavioral health program evaluations increasingly employ mixed-methods designs that combine quantitative outcome data with qualitative data from client interviews, focus groups, and case studies. Qualitative data can illuminate why a program succeeded or failed, capture client experiences that standardized measures miss, and identify implementation barriers—enriching the quantitative findings with contextual depth.

Detailed Breakdown — Types of Program Evaluation

Program evaluation encompasses several distinct but interconnected types, each addressing different questions about a program's value and functioning. Understanding the relationships among these types—and knowing when to deploy each one—is a critical competency assessed on the EPPP. The following diagram and table provide a comprehensive classification of evaluation types commonly encountered in behavioral health settings.

This hierarchy shows how evaluation types build on one another. A needs assessment establishes the foundation, followed by theory evaluation and process evaluation during implementation, culminating in outcome evaluation and efficiency analysis that inform program decisions.
Overview of evaluation types, their guiding questions, methods, and behavioral health applications
Evaluation TypeKey QuestionCommon MethodsBehavioral Health Example
Needs AssessmentWhat is the nature and extent of the problem?Epidemiological surveys, key informant interviews, community forumsAssessing rates of opioid use disorder in a rural county to justify a new treatment program
Theory EvaluationIs the program's logic model evidence-based and plausible?Literature review, expert panel review, evaluability assessmentReviewing whether a CBT-based anxiety program's theory of change is supported by research
Process EvaluationIs the program being implemented as designed?Fidelity checklists, attendance logs, staff surveys, observationMonitoring whether therapists adhere to a manualized DBT protocol
Outcome EvaluationDid the program produce the intended changes?Pre-post measures, comparison groups, standardized assessments (PHQ-9, GAD-7)Measuring reduction in depression scores after a 12-week group therapy program
Efficiency AnalysisAre outcomes justified relative to costs?Cost-effectiveness analysis (CEA), cost-benefit analysis (CBA), cost-utility analysisComparing cost per QALY gained between telehealth and in-person therapy models

Worked Example — Evaluating a Substance Use Treatment Program

Consider a community behavioral health center that has operated an intensive outpatient program (IOP) for substance use disorders for two years. The program's funder requests an outcome evaluation to determine whether the program should receive continued support. The evaluator must design and conduct a systematic evaluation of program effectiveness. Let us walk through this process step by step.

Evaluating the Effectiveness of a Substance Use IOP
1
Step 1 — Clarify Evaluation Questions with StakeholdersThe evaluator meets with the program director, clinical staff, client advisory board, and the funder to identify evaluation questions. Through this stakeholder engagement process, three primary questions emerge: (1) Do clients show significant reductions in substance use frequency at 3-month and 6-month follow-up? (2) Do clients demonstrate improvements in psychosocial functioning? (3) Is the program cost-effective relative to comparable services in the region?
Three evaluation questions defined through stakeholder collaboration
2
Step 2 — Review the Logic Model and Select MeasuresThe evaluator reviews the program's logic model: inputs include licensed counselors, group rooms, and curricula; activities include 3× weekly group therapy, individual counseling, and drug screening; expected outcomes include reduced substance use, improved functioning, and fewer ER visits. The evaluator selects validated measures: the Addiction Severity Index (ASI) for substance use, the World Health Organization Disability Assessment Schedule (WHODAS 2.0) for functioning, and program financial records for cost data. The evaluation design is a one-group pretest-posttest with 6-month follow-up, supplemented by qualitative exit interviews.
Validated instruments selected: ASI, WHODAS 2.0; pre-post design with follow-up
3
Step 3 — Collect and Analyze DataOver the evaluation period, 84 clients enrolled in the IOP. Pretest ASI composite scores averaged 0.42 (SD = 0.15), while 3-month posttest scores averaged 0.28 (SD = 0.14). Using Cohen's d: d = (0.42 − 0.28) / √((0.15² + 0.14²) / 2) = 0.14 / 0.145 ≈ 0.97. This represents a large effect size. The evaluator also calculates the RCI for individual clients to identify the proportion showing clinically reliable improvement, finding that 62% (52/84) exceeded the RCI threshold of 1.96.
Cohen's d ≈ 0.97 (large effect); 62% of clients showed reliable change
4
Step 4 — Assess Cost-EffectivenessThe total program cost for the evaluation period was $378,000 (including staff salaries, facility costs, and materials). With 52 clients achieving clinically significant improvement, the cost-effectiveness ratio is CER = $378,000 / 52 = $7,269 per successful outcome. Compared to inpatient treatment programs in the region averaging $15,000–$25,000 per client episode, the IOP demonstrates favorable cost-effectiveness for clients who respond positively to outpatient-level care.
CER = $7,269 per successful outcome — favorable vs. inpatient alternatives
5
Step 5 — Interpret, Report, and RecommendThe evaluator prepares a comprehensive report. Strengths include the large effect size and favorable cost-effectiveness. Limitations include the absence of a comparison group (threatening internal validity), significant attrition (24 clients did not complete post-testing), and the short follow-up period. Qualitative interviews revealed that clients valued the peer support component but found scheduling inflexible. Recommendations include continued funding, adding evening sessions to improve retention, implementing a waitlist control design for the next evaluation cycle, and tracking 12-month outcomes.
Report delivered with evidence-based recommendations for program continuation and improvement

Strengths, Limitations, & Ethical Considerations

Program evaluation in behavioral health occupies a unique position at the intersection of research methodology, clinical practice, and organizational management. Each evaluation approach carries distinct advantages and vulnerabilities that the competent evaluator must weigh carefully. Furthermore, ethical considerations—many of which parallel those in clinical research—permeate every phase of the evaluation process.

Comparative strengths and limitations of evaluation designs commonly used in behavioral health
DimensionStrengthsLimitations / Challenges
RCT DesignsHighest internal validity; strong causal inference; controls for confounds through randomizationOften unethical in clinical settings (withholding treatment); expensive; limited ecological validity; selection bias if clients refuse randomization
Quasi-ExperimentalPractical for real-world settings; preserves treatment access; can use existing comparison groupsVulnerable to selection bias; groups may differ on unmeasured variables; requires statistical controls (e.g., propensity score matching)
Pre-Post (No Control)Simple; low cost; feasible with small programs; no ethical issues regarding withholding treatmentCannot rule out maturation, history, regression to the mean; weak causal claims; most common but least rigorous design
Qualitative MethodsRich contextual understanding; captures client voice; identifies implementation barriers; flexible and emergentSubjective interpretation; limited generalizability; time-intensive analysis; may not satisfy funders who require quantitative outcome data
Cost-Effectiveness AnalysisDirectly informs resource allocation; compelling to funders and policymakers; allows cross-program comparisonDifficulty monetizing behavioral health outcomes; may reduce complex human changes to simplistic ratios; risk of privileging efficiency over equity
⚖️ ETHICAL IMPERATIVES IN PROGRAM EVALUATION
Program evaluators in behavioral health must navigate several ethical tensions. The APA Ethics Code and the American Evaluation Association's Guiding Principles both emphasize respect for participants, informed consent, confidentiality, and attention to cultural context. A particularly salient tension arises when evaluation findings suggest a program is ineffective: the evaluator has an obligation to report accurately, even when findings threaten the program's funding or organizational relationships. Conversely, evaluators must guard against evaluator bias—the temptation to design or interpret evaluations in ways that favor a predetermined conclusion, whether positive or negative. Maintaining independence, transparency, and methodological rigor is essential for the evaluation's credibility and ultimately for the welfare of the clients the program serves.

Connection to Advanced Theory & Contemporary Practice

Program evaluation has evolved significantly from its early emphasis on simple outcome measurement. Contemporary approaches in behavioral health integrate sophisticated theoretical frameworks, advanced methodologies, and a heightened awareness of social justice concerns. Understanding these advanced connections positions the developing psychologist not only for success on the EPPP but also for leadership in shaping effective, equitable behavioral health services.

Evolution from traditional to contemporary program evaluation approaches
Traditional ApproachContemporary / Advanced ApproachKey Advancement
Expert-driven evaluation; evaluator determines all questions and methodsParticipatory evaluation; stakeholders co-design evaluation with evaluatorIncreased relevance, ownership, and utilization of findings
One-size-fits-all outcome measurementCulturally responsive evaluation; equity-focused outcome analysisExamines differential outcomes across racial, ethnic, and socioeconomic groups
Summative focus; evaluation occurs at program endDevelopmental evaluation (Patton); continuous feedback in complex adaptive systemsSupports innovation and adaptation in rapidly changing environments
Single-site, program-level analysisImplementation science; studying how EBPs are adopted across contextsBridges the research-practice gap; focuses on scalability and sustainability
Paper-based data collection at discrete time pointsRoutine outcome monitoring (ROM); real-time digital data dashboardsEnables session-by-session tracking; supports measurement-based care

A particularly important contemporary development is the integration of program evaluation with implementation science. While traditional program evaluation asks "Did this program work?" implementation science asks "How and why does this program work (or fail) when transported to new settings?" This shift reflects the field's recognition that evidence-based practices often lose effectiveness during real-world dissemination due to contextual factors, organizational barriers, and adaptations that compromise fidelity. For psychologists in behavioral health, competency in both program evaluation and implementation science is increasingly essential for ensuring that effective interventions reach the populations who need them most.

📋 EPPP Application Note
On the EPPP Part 2 (Skills), you may be presented with vignettes requiring you to select appropriate evaluation designs, identify threats to validity in proposed evaluations, interpret outcome data, or recommend program modifications based on evaluation findings. Focus on demonstrating your ability to integrate methodological knowledge with practical constraints and ethical reasoning. The examiners value nuanced responses that acknowledge trade-offs between rigor and feasibility.

Practice Problems

PROBLEM 1CONCEPTUAL
A community mental health center is planning to launch a new peer support program for individuals with serious mental illness. The program director asks you, as the consulting psychologist, to help plan an evaluation. She states: "We want to know if the program works." Explain why this question is insufficient for designing a rigorous evaluation and describe at least three specific, evaluable questions you would recommend instead.
PROBLEM 2BASIC CALCULATION
A depression treatment program reports the following data: Pretest PHQ-9 mean = 17.4 (SD = 4.8), Posttest PHQ-9 mean = 10.2 (SD = 5.1). Calculate the within-group effect size (Cohen's d) for this pre-post comparison and interpret the result using Cohen's conventions.
PROBLEM 3INTERMEDIATE
You are evaluating an anxiety reduction program using a one-group pretest-posttest design. Your data show a statistically significant reduction in GAD-7 scores (p < .01). However, during the evaluation period, a major employer in the community announced it would not be conducting layoffs, reducing economic uncertainty for many program participants. A colleague argues this invalidates your findings. Identify the specific threat to internal validity at play, explain how it could account for the observed improvement, and propose a design modification for the next evaluation cycle that would partially address this threat.
PROBLEM 4APPLIED
You are hired to evaluate a school-based suicide prevention program that has been operating for three years. The program serves 800 students annually across four high schools. The program director wants a rigorous outcome evaluation but faces several constraints: (a) the school board will not allow randomized assignment of students to a no-treatment control group, (b) the budget for evaluation is $25,000, and (c) the program uses a proprietary curriculum with no published psychometric data on its assessment tools. Design a feasible evaluation plan that addresses these constraints while maximizing methodological rigor.
PROBLEM 5CRITICAL THINKING
A state behavioral health authority commissions you to evaluate whether its system of community-based services is reducing racial disparities in treatment outcomes. The system includes 42 agencies serving approximately 15,000 clients annually across diverse geographic regions. Discuss the theoretical frameworks you would draw upon, the evaluation design considerations unique to equity-focused evaluation, the types of data you would need, and at least two ethical tensions you would anticipate in conducting this evaluation.

Lesson Summary

Program evaluation is the systematic process of assessing the design, implementation, and outcomes of behavioral health services. Its foundations include the logic model (linking inputs → activities → outputs → outcomes → impact), the distinction between formative evaluation (improving a program during implementation) and summative evaluation (judging overall effectiveness), and the hierarchy of evaluation types: needs assessment, process evaluation, outcome evaluation, and efficiency analysis. Key quantitative tools include Cohen's d effect size, the Reliable Change Index, and cost-effectiveness ratios.

Competent evaluators engage stakeholders throughout the evaluation process, select designs that balance internal validity with practical and ethical constraints, use validated measures, and attend to cultural responsiveness and health equity. Contemporary advances—including implementation science, participatory evaluation, developmental evaluation, and routine outcome monitoring—extend these foundations into a dynamic, equity-oriented practice that positions psychologists as leaders in ensuring that behavioral health services genuinely serve the populations they are designed to help.

Varsity Tutors • EPPP: Part 2, Skills • Program Evaluation — Evaluate service and program effectiveness