EPPP: PART 1, KNOWLEDGE • DOMAIN 7: RESEARCH METHODS AND STATISTICS

Program Evaluation — Differentiate formative, summative, outcome, and cost-benefit evaluation methods

Understanding how behavioral health programs are assessed for improvement, effectiveness, results, and fiscal justification.

Historical Context & Motivation

The systematic evaluation of social and behavioral health programs has a relatively recent but impactful history. Before the mid-twentieth century, programs in mental health, education, and social services were often launched with noble intentions but minimal attention to whether they actually achieved their goals. Program evaluation emerged as a formal discipline in response to the growing need for accountability — particularly as government-funded initiatives expanded dramatically during the 1960s and 1970s. Legislators, administrators, and clinicians alike began asking a fundamental question: Are the programs we invest in actually making a difference?

1932
Tyler's Evaluation Framework
Ralph Tyler developed one of the first systematic approaches to program evaluation in the context of educational assessment, emphasizing objective-based evaluation and measurable learning outcomes.
1967
Scriven's Formative–Summative Distinction
Michael Scriven introduced the critical distinction between formative evaluation (improving a program while it runs) and summative evaluation (judging a program's overall merit after completion), terms that remain foundational today.
1969
Stufflebeam's CIPP Model
Daniel Stufflebeam proposed the Context, Input, Process, Product (CIPP) model, offering a comprehensive framework that incorporated multiple evaluation types and emphasized decision-oriented evaluation.
1975
Cost-Benefit Analysis in Health
The application of cost-benefit and cost-effectiveness analysis expanded into public health and behavioral health, driven by the need to justify expenditures on prevention, treatment, and community mental health programs.
1993
Government Performance and Results Act (GPRA)
U.S. federal legislation mandated that all federally funded programs establish measurable outcome goals and report results, cementing outcome evaluation as a requirement for public behavioral health services.

These historical developments collectively shaped the modern landscape of program evaluation. Today, behavioral health professionals are expected to understand not just whether a program works, but when, how, and at what cost it achieves its effects. The EPPP tests your ability to differentiate among evaluation approaches — formative, summative, outcome, and cost-benefit — each of which addresses a distinct evaluative question at a different stage of program development.

Core Principles & Definitions

Program evaluation is the systematic collection and analysis of information about the activities, characteristics, and outcomes of programs to make judgments about the program, improve program effectiveness, and inform decisions about future programming. While there are numerous evaluation models, the EPPP focuses on four primary types that differ in their timing, purpose, and the questions they seek to answer. Understanding these distinctions is essential for both clinical practice and the licensing examination.

1

Formative Evaluation

Conducted during program implementation to identify strengths and weaknesses so that improvements can be made in real time. Focuses on process, delivery, and fidelity. Asks: "How can we make this program better while it's running?"
2

Summative Evaluation

Conducted after program completion (or at a designated endpoint) to determine overall effectiveness, merit, and worth. Addresses accountability and continuation decisions. Asks: "Did this program work as intended?"
3

Outcome Evaluation

Examines the results or effects of a program on participants — both intended and unintended. Often involves pre-post comparisons or control groups to measure change in target variables (e.g., symptom reduction, functioning). Asks: "What changed because of this program?"
4

Cost-Benefit Evaluation

Compares the monetary value of program benefits to the monetary costs of running the program. Related methods include cost-effectiveness analysis (CEA), which compares costs to non-monetary outcomes. Asks: "Is this program worth the investment?"
KEY TAKEAWAY
Think of program evaluation like the lifecycle of a new treatment protocol at a community mental health center. Formative evaluation is like monitoring patients during the first few weeks to adjust dosage and delivery. Summative evaluation is the final review committee that decides whether the protocol should be adopted permanently. Outcome evaluation measures whether patients actually improved. Cost-benefit evaluation asks whether the gains justified the expense relative to alternative treatments.

Visual Explanation — The Evaluation Lifecycle

This diagram illustrates how the four evaluation types map onto the program lifecycle. Formative evaluation occurs during implementation, outcome evaluation spans data collection before and after, summative evaluation occurs at the end, and cost-benefit analysis often follows the summative phase to inform fiscal decisions.

The diagram above highlights a critical concept for the EPPP: these evaluation types are not mutually exclusive, but they serve different functions and address different stakeholder questions. A well-designed behavioral health program might employ all four types across its lifespan. Formative evaluation guides mid-course corrections during a pilot phase, outcome evaluation documents the changes experienced by clients, summative evaluation informs the decision about whether to continue or expand the program, and cost-benefit analysis helps administrators compare the program's value to alternative uses of the same resources.

How Each Evaluation Method Works

Formative Evaluation in Depth

Formative evaluation is process-oriented and typically occurs during the development or early implementation of a program. Its primary purpose is quality improvement rather than judgment. In behavioral health contexts, this might involve collecting feedback from therapists delivering a new group intervention, observing session fidelity to a treatment manual, or conducting focus groups with participants midway through a psychoeducational program. The data gathered inform real-time modifications — perhaps the session length is too long, the language is inaccessible to the target population, or a critical component is being inconsistently delivered.

💡 EPPP TIP
A common EPPP question stem describes a scenario where an evaluator is observing a program in progress and making recommendations to improve delivery. This is almost always formative evaluation. The key differentiator is that the program is still ongoing and the goal is improvement, not final judgment.

Summative Evaluation in Depth

Summative evaluation occurs at the conclusion of a program or at a predetermined milestone and renders an overall judgment about the program's merit, worth, or value. The audience for summative evaluation is typically external — funders, administrators, policymakers, or regulatory bodies — who need to decide whether to continue, expand, replicate, or terminate a program. Whereas formative evaluation feeds information back to the program developers, summative evaluation feeds information outward to decision-makers. A summative evaluation of a substance abuse treatment program, for example, might compare completion rates, client satisfaction scores, and relapse rates against established benchmarks or against comparison programs.

Outcome Evaluation in Depth

Outcome evaluation (sometimes called impact evaluation when focused on broader community-level effects) specifically examines changes in the target variables that the program was designed to influence. This type of evaluation often employs quasi-experimental or experimental designs — including pre-test/post-test comparisons, waitlist control groups, or randomized controlled trials — to establish whether observed changes can be attributed to the program rather than to maturation, history, or other confounding factors. In behavioral health, typical outcome variables include symptom severity scores (e.g., PHQ-9 for depression), level of functioning (e.g., GAF or WHODAS scores), hospitalization rates, or quality of life measures.

Cost-Benefit Evaluation in Depth

Cost-benefit analysis (CBA) translates all program outcomes into monetary terms so that costs and benefits can be directly compared. If a community mental health center spends $500,000 annually on an early intervention program for psychosis, a cost-benefit analysis would attempt to monetize the benefits — such as reduced emergency room visits, decreased inpatient days, increased employment, and averted criminal justice costs — and compare these to the program expenditure. The result is typically expressed as a benefit-cost ratio (BCR) or as net benefits. A related but distinct approach is cost-effectiveness analysis (CEA), which does not require monetizing outcomes; instead, it compares the cost per unit of outcome (e.g., cost per symptom-free day, cost per quality-adjusted life year or QALY).

BENEFIT-COST RATIO
BCR = Total Monetized Benefits ÷ Total Program Costs
A BCR > 1.0 indicates that benefits exceed costs; a BCR < 1.0 indicates costs exceed benefits. For example, a BCR of 3.2 means every $1 invested returns $3.20 in benefits.
NET BENEFIT
Net Benefit = Total Monetized Benefits − Total Program Costs
A positive net benefit means the program generated more value than it consumed. This metric is useful for comparing programs of different sizes.
COST-EFFECTIVENESS RATIO
CER = Total Program Costs ÷ Total Units of Outcome
Unlike CBA, cost-effectiveness analysis does not require outcomes to be monetized. The CER might be expressed as cost per QALY gained, cost per depression remission, or cost per successful treatment completion.

Detailed Classification & Comparison

This comprehensive comparison chart maps each evaluation type across six critical dimensions: timing, purpose, primary audience, key question, typical methods, and a behavioral health example. Note Scriven's classic mnemonic at the bottom — a frequently cited analogy on the EPPP.

One common source of confusion on the EPPP involves distinguishing summative evaluation from outcome evaluation. While both typically occur at or near the conclusion of a program, they differ in scope and emphasis. Outcome evaluation is narrowly focused on measuring whether the target variables changed — it is essentially an empirical question. Summative evaluation, by contrast, renders a holistic judgment that may incorporate outcome data alongside considerations such as participant satisfaction, implementation fidelity, cost, and alignment with organizational priorities. In this sense, outcome data often serve as one input to a broader summative evaluation, but the two are not synonymous.

⚠️ CBA vs. CEA — Know the Difference
Cost-benefit analysis (CBA) requires all outcomes to be converted to monetary values, allowing a direct comparison of dollars spent versus dollars saved. Cost-effectiveness analysis (CEA) compares costs to outcomes measured in their natural units (e.g., symptom-free days, QALYs). On the EPPP, if a question asks about comparing the monetary value of outcomes to costs, the answer is CBA. If it asks about cost per unit of clinical improvement, the answer is CEA.

Worked Example — Evaluating a Behavioral Health Program

Consider the following scenario: A community mental health center launches a new 12-week cognitive-behavioral group therapy (CBT-G) program for adults with generalized anxiety disorder. The program serves 60 clients per year and costs $120,000 annually to operate. The program director wants to conduct a comprehensive evaluation. Let us walk through how each evaluation type would be applied.

Comprehensive Evaluation of a CBT Group Program for Anxiety
1
Step 1 — Formative Evaluation (During Program)At weeks 4 and 8, the evaluator observes group sessions and rates therapist adherence to the CBT manual using a fidelity checklist. Client satisfaction surveys are administered at the midpoint. Focus groups with facilitators reveal that the homework assignments are too complex for clients with lower literacy levels. Result: The homework materials are simplified and visual aids are added for the remaining sessions.
Formative evaluation led to a mid-course improvement in program materials.
2
Step 2 — Outcome Evaluation (Pre-Post Measurement)All 60 clients complete the GAD-7 (Generalized Anxiety Disorder 7-item scale) at intake and again at week 12. Mean pre-treatment GAD-7 score = 16.4 (severe anxiety); mean post-treatment GAD-7 score = 8.2 (mild anxiety). A paired-samples t-test yields t(59) = 7.83, p < .001, Cohen's d = 1.02, indicating a large effect. Result: Statistically and clinically significant improvement in anxiety symptoms.
Outcome evaluation documented a clinically meaningful reduction in anxiety (d = 1.02).
3
Step 3 — Summative Evaluation (End-of-Year Review)At the end of the fiscal year, the advisory board reviews the outcome data alongside additional information: 85% of clients completed all 12 sessions (above the 75% benchmark), client satisfaction averaged 4.6/5.0, and therapist fidelity ratings averaged 92%. However, only 40 of 60 slots were filled due to low referral volume. The board weighs all of this evidence to determine whether the program merits continued funding.
Summative evaluation: Program was effective but underutilized; board recommends continuation with enhanced outreach.
4
Step 4 — Cost-Benefit AnalysisThe evaluator calculates total program cost: $120,000. Next, benefits are monetized. Clients in the program had an average of 1.8 fewer ER visits per year (ER visit cost ≈ $1,500 each), for a total savings of 60 × 1.8 × $1,500 = $162,000. Additionally, improved functioning led to an estimated 15 clients returning to part-time employment, generating approximately $90,000 in economic productivity. Total monetized benefits = $162,000 + $90,000 = $252,000.
BCR = $252,000 ÷ $120,000 = 2.10. For every $1 invested, $2.10 in benefits were generated. Net benefit = $252,000 − $120,000 = $132,000.

Strengths & Limitations of Each Method

Strengths and limitations of the four primary evaluation methods in behavioral health
MethodStrengthsLimitations
FormativeEnables real-time improvement; increases likelihood of program success; responsive to stakeholder needs; identifies implementation problems earlyDoes not establish effectiveness; may lack rigor if informal; findings may not generalize; can be biased by evaluator proximity to program staff
SummativeProvides accountability; supports high-stakes decisions (funding, continuation); integrates multiple data sources for a comprehensive judgmentOccurs too late to fix problems; may oversimplify complex outcomes; political pressures may influence findings; does not always clarify what caused success or failure
OutcomeDirectly measures client change; can establish causal attribution (with experimental designs); uses standardized instruments for comparabilityRequires adequate sample sizes and control conditions; internal validity threats (history, maturation, attrition); does not explain why outcomes occurred; expensive if RCT is used
Cost-BenefitProvides economic justification; allows comparison across different types of programs; highly persuasive to policymakers and funders; objective metric (BCR or net benefit)Difficult to monetize intangible benefits (e.g., dignity, reduced suffering); may undervalue outcomes that resist dollar conversion; requires complex assumptions about future savings; ethically contentious when applied to human wellbeing
KEY TAKEAWAY
No single evaluation method is sufficient on its own. Just as a clinical psychologist uses multiple assessment methods (interview, testing, observation) to form a comprehensive case conceptualization, a thorough program evaluation typically integrates formative, summative, outcome, and economic data to provide a complete picture. On the EPPP, recognize that question stems often test whether you can match the correct evaluation type to the scenario described — pay close attention to when the evaluation occurs and what question it aims to answer.

Connection to Advanced Evaluation Theory

The four evaluation types covered in this lesson represent the foundational categories tested on the EPPP, but contemporary evaluation theory has expanded considerably beyond these classical distinctions. Understanding how these basics connect to more advanced frameworks will both deepen your conceptual grasp and help you navigate challenging exam questions that reference overlapping models.

How foundational evaluation types connect to advanced evaluation frameworks
Basic ConceptAdvanced ExtensionKey Difference
Formative EvaluationDevelopmental Evaluation (Patton, 2011)DE is used in highly complex, emergent, or innovative programs where the intervention itself is still being designed. The evaluator is embedded in the team and evaluation is continuous — not just periodic check-ins.
Summative EvaluationGoal-Free Evaluation (Scriven, 1991)Instead of judging whether stated goals were met, the evaluator deliberately avoids learning about program goals and instead examines all effects — intended and unintended — to reduce confirmation bias.
Outcome EvaluationTheory-Driven Evaluation (Chen, 1990)Goes beyond simply measuring outcomes to specifying and testing the causal mechanisms (the program theory or logic model) that link program activities to outcomes — asking not just 'what' but 'how' and 'why.'
Cost-Benefit AnalysisSocial Return on Investment (SROI)SROI extends CBA by incorporating social, environmental, and community-level impacts that traditional CBA may overlook — particularly relevant in behavioral health where benefits include family stability, community safety, and social inclusion.

Another important advanced concept is the logic model (also called a program theory), which visually maps the assumed causal pathway from program inputs and activities to outputs, short-term outcomes, and long-term impact. Logic models are frequently used to guide all four types of evaluation: formative evaluation checks whether activities are being delivered as planned (process), outcome evaluation tests the expected changes, summative evaluation assesses the overall chain from input to impact, and cost-benefit analysis monetizes the outcomes at the end of the chain. On the EPPP, you may encounter questions that present a logic model and ask you to identify which component is being evaluated.

Practice Problems

PROBLEM 1CONCEPTUAL
A program director collects feedback from both facilitators and participants at the midpoint of a 10-session anger management group in order to make adjustments for the remaining sessions. What type of evaluation is the director conducting?
PROBLEM 2BASIC CALCULATION
A substance abuse treatment program costs $200,000 per year. An evaluation determines that the program produces monetized benefits of $350,000 annually (including reduced healthcare utilization, decreased criminal justice involvement, and increased employment). Calculate the benefit-cost ratio (BCR) and the net benefit.
PROBLEM 3INTERMEDIATE
A state mental health authority funds two depression treatment programs: Program A costs $80,000 per year and achieves depression remission in 40 of 100 patients (40%); Program B costs $150,000 per year and achieves remission in 90 of 120 patients (75%). Using cost-effectiveness analysis, which program achieves a lower cost per remission? How might a cost-benefit analysis yield a different conclusion?
PROBLEM 4APPLIED
You are hired as an external evaluator for a school-based suicide prevention program that has been running for three years. The school board wants to know whether to continue funding the program next year. They ask you to compare the program to two alternative prevention programs. Which type of evaluation is most appropriate for the school board's question, and what additional evaluation type would strengthen your report? Explain your reasoning.
PROBLEM 5CRITICAL THINKING
A colleague argues that formative and summative evaluations are fundamentally incompatible — that the same evaluator cannot serve both improvement and judgment functions without a conflict of interest. Another colleague disagrees, citing Stufflebeam's CIPP model as evidence that a single evaluation framework can integrate both functions. Critically evaluate both positions and describe under what conditions each position might be correct.

Summary & Review

Program evaluation in behavioral health encompasses four primary methods, each serving a distinct purpose along the program lifecycle. Formative evaluation occurs during program implementation and focuses on process improvement — think of the cook tasting the soup. Summative evaluation occurs at the conclusion and renders an overall judgment about program merit — the dinner guest tasting the final dish. Outcome evaluation empirically measures whether target variables changed as a result of the program, using designs ranging from simple pre-post comparisons to randomized controlled trials. Cost-benefit analysis converts all program outcomes to monetary values to calculate the benefit-cost ratio (BCR) and net benefit, while the related cost-effectiveness analysis (CEA) compares costs to outcomes in their natural clinical units.

For the EPPP, focus on three key differentiators: timing (during vs. after the program), purpose (improvement vs. judgment vs. measurement vs. fiscal analysis), and audience (program staff vs. funders vs. researchers vs. policymakers). Remember that these evaluation types are complementary, not mutually exclusive, and that a comprehensive program evaluation in behavioral health often integrates elements of all four approaches across the program lifecycle.

Varsity Tutors • EPPP: Part 1, Knowledge • Program Evaluation — Differentiate formative, summative, outcome, and cost-benefit evaluation methods