Historical Context & Motivation
The systematic evaluation of social and behavioral health programs has a relatively recent but impactful history. Before the mid-twentieth century, programs in mental health, education, and social services were often launched with noble intentions but minimal attention to whether they actually achieved their goals. Program evaluation emerged as a formal discipline in response to the growing need for accountability — particularly as government-funded initiatives expanded dramatically during the 1960s and 1970s. Legislators, administrators, and clinicians alike began asking a fundamental question: Are the programs we invest in actually making a difference?
These historical developments collectively shaped the modern landscape of program evaluation. Today, behavioral health professionals are expected to understand not just whether a program works, but when, how, and at what cost it achieves its effects. The EPPP tests your ability to differentiate among evaluation approaches — formative, summative, outcome, and cost-benefit — each of which addresses a distinct evaluative question at a different stage of program development.
Core Principles & Definitions
Program evaluation is the systematic collection and analysis of information about the activities, characteristics, and outcomes of programs to make judgments about the program, improve program effectiveness, and inform decisions about future programming. While there are numerous evaluation models, the EPPP focuses on four primary types that differ in their timing, purpose, and the questions they seek to answer. Understanding these distinctions is essential for both clinical practice and the licensing examination.
Formative Evaluation
Summative Evaluation
Outcome Evaluation
Cost-Benefit Evaluation
Visual Explanation — The Evaluation Lifecycle
The diagram above highlights a critical concept for the EPPP: these evaluation types are not mutually exclusive, but they serve different functions and address different stakeholder questions. A well-designed behavioral health program might employ all four types across its lifespan. Formative evaluation guides mid-course corrections during a pilot phase, outcome evaluation documents the changes experienced by clients, summative evaluation informs the decision about whether to continue or expand the program, and cost-benefit analysis helps administrators compare the program's value to alternative uses of the same resources.
How Each Evaluation Method Works
Formative Evaluation in Depth
Formative evaluation is process-oriented and typically occurs during the development or early implementation of a program. Its primary purpose is quality improvement rather than judgment. In behavioral health contexts, this might involve collecting feedback from therapists delivering a new group intervention, observing session fidelity to a treatment manual, or conducting focus groups with participants midway through a psychoeducational program. The data gathered inform real-time modifications — perhaps the session length is too long, the language is inaccessible to the target population, or a critical component is being inconsistently delivered.
Summative Evaluation in Depth
Summative evaluation occurs at the conclusion of a program or at a predetermined milestone and renders an overall judgment about the program's merit, worth, or value. The audience for summative evaluation is typically external — funders, administrators, policymakers, or regulatory bodies — who need to decide whether to continue, expand, replicate, or terminate a program. Whereas formative evaluation feeds information back to the program developers, summative evaluation feeds information outward to decision-makers. A summative evaluation of a substance abuse treatment program, for example, might compare completion rates, client satisfaction scores, and relapse rates against established benchmarks or against comparison programs.
Outcome Evaluation in Depth
Outcome evaluation (sometimes called impact evaluation when focused on broader community-level effects) specifically examines changes in the target variables that the program was designed to influence. This type of evaluation often employs quasi-experimental or experimental designs — including pre-test/post-test comparisons, waitlist control groups, or randomized controlled trials — to establish whether observed changes can be attributed to the program rather than to maturation, history, or other confounding factors. In behavioral health, typical outcome variables include symptom severity scores (e.g., PHQ-9 for depression), level of functioning (e.g., GAF or WHODAS scores), hospitalization rates, or quality of life measures.
Cost-Benefit Evaluation in Depth
Cost-benefit analysis (CBA) translates all program outcomes into monetary terms so that costs and benefits can be directly compared. If a community mental health center spends $500,000 annually on an early intervention program for psychosis, a cost-benefit analysis would attempt to monetize the benefits — such as reduced emergency room visits, decreased inpatient days, increased employment, and averted criminal justice costs — and compare these to the program expenditure. The result is typically expressed as a benefit-cost ratio (BCR) or as net benefits. A related but distinct approach is cost-effectiveness analysis (CEA), which does not require monetizing outcomes; instead, it compares the cost per unit of outcome (e.g., cost per symptom-free day, cost per quality-adjusted life year or QALY).
Detailed Classification & Comparison
One common source of confusion on the EPPP involves distinguishing summative evaluation from outcome evaluation. While both typically occur at or near the conclusion of a program, they differ in scope and emphasis. Outcome evaluation is narrowly focused on measuring whether the target variables changed — it is essentially an empirical question. Summative evaluation, by contrast, renders a holistic judgment that may incorporate outcome data alongside considerations such as participant satisfaction, implementation fidelity, cost, and alignment with organizational priorities. In this sense, outcome data often serve as one input to a broader summative evaluation, but the two are not synonymous.
Worked Example — Evaluating a Behavioral Health Program
Consider the following scenario: A community mental health center launches a new 12-week cognitive-behavioral group therapy (CBT-G) program for adults with generalized anxiety disorder. The program serves 60 clients per year and costs $120,000 annually to operate. The program director wants to conduct a comprehensive evaluation. Let us walk through how each evaluation type would be applied.
Strengths & Limitations of Each Method
| Method | Strengths | Limitations |
|---|---|---|
| Formative | Enables real-time improvement; increases likelihood of program success; responsive to stakeholder needs; identifies implementation problems early | Does not establish effectiveness; may lack rigor if informal; findings may not generalize; can be biased by evaluator proximity to program staff |
| Summative | Provides accountability; supports high-stakes decisions (funding, continuation); integrates multiple data sources for a comprehensive judgment | Occurs too late to fix problems; may oversimplify complex outcomes; political pressures may influence findings; does not always clarify what caused success or failure |
| Outcome | Directly measures client change; can establish causal attribution (with experimental designs); uses standardized instruments for comparability | Requires adequate sample sizes and control conditions; internal validity threats (history, maturation, attrition); does not explain why outcomes occurred; expensive if RCT is used |
| Cost-Benefit | Provides economic justification; allows comparison across different types of programs; highly persuasive to policymakers and funders; objective metric (BCR or net benefit) | Difficult to monetize intangible benefits (e.g., dignity, reduced suffering); may undervalue outcomes that resist dollar conversion; requires complex assumptions about future savings; ethically contentious when applied to human wellbeing |
Connection to Advanced Evaluation Theory
The four evaluation types covered in this lesson represent the foundational categories tested on the EPPP, but contemporary evaluation theory has expanded considerably beyond these classical distinctions. Understanding how these basics connect to more advanced frameworks will both deepen your conceptual grasp and help you navigate challenging exam questions that reference overlapping models.
| Basic Concept | Advanced Extension | Key Difference |
|---|---|---|
| Formative Evaluation | Developmental Evaluation (Patton, 2011) | DE is used in highly complex, emergent, or innovative programs where the intervention itself is still being designed. The evaluator is embedded in the team and evaluation is continuous — not just periodic check-ins. |
| Summative Evaluation | Goal-Free Evaluation (Scriven, 1991) | Instead of judging whether stated goals were met, the evaluator deliberately avoids learning about program goals and instead examines all effects — intended and unintended — to reduce confirmation bias. |
| Outcome Evaluation | Theory-Driven Evaluation (Chen, 1990) | Goes beyond simply measuring outcomes to specifying and testing the causal mechanisms (the program theory or logic model) that link program activities to outcomes — asking not just 'what' but 'how' and 'why.' |
| Cost-Benefit Analysis | Social Return on Investment (SROI) | SROI extends CBA by incorporating social, environmental, and community-level impacts that traditional CBA may overlook — particularly relevant in behavioral health where benefits include family stability, community safety, and social inclusion. |
Another important advanced concept is the logic model (also called a program theory), which visually maps the assumed causal pathway from program inputs and activities to outputs, short-term outcomes, and long-term impact. Logic models are frequently used to guide all four types of evaluation: formative evaluation checks whether activities are being delivered as planned (process), outcome evaluation tests the expected changes, summative evaluation assesses the overall chain from input to impact, and cost-benefit analysis monetizes the outcomes at the end of the chain. On the EPPP, you may encounter questions that present a logic model and ask you to identify which component is being evaluated.
Practice Problems
Summary & Review
Program evaluation in behavioral health encompasses four primary methods, each serving a distinct purpose along the program lifecycle. Formative evaluation occurs during program implementation and focuses on process improvement — think of the cook tasting the soup. Summative evaluation occurs at the conclusion and renders an overall judgment about program merit — the dinner guest tasting the final dish. Outcome evaluation empirically measures whether target variables changed as a result of the program, using designs ranging from simple pre-post comparisons to randomized controlled trials. Cost-benefit analysis converts all program outcomes to monetary values to calculate the benefit-cost ratio (BCR) and net benefit, while the related cost-effectiveness analysis (CEA) compares costs to outcomes in their natural clinical units.
For the EPPP, focus on three key differentiators: timing (during vs. after the program), purpose (improvement vs. judgment vs. measurement vs. fiscal analysis), and audience (program staff vs. funders vs. researchers vs. policymakers). Remember that these evaluation types are complementary, not mutually exclusive, and that a comprehensive program evaluation in behavioral health often integrates elements of all four approaches across the program lifecycle.