COLLEGE POLITICAL SCIENCE • PUBLIC POLICY AND ADMINISTRATION

Policy Evaluation Design — Design basic policy evaluation questions and outcome metrics

Learn to craft rigorous evaluation questions and measurable outcome metrics that determine whether public policies achieve their intended goals.

Historical Context & Motivation

Governments have always sought to understand whether their actions produce the desired results, but the systematic study of policy evaluation is a relatively modern development, emerging from the convergence of social science methodology and the expansion of the administrative state. In the early twentieth century, public programs were largely assessed through anecdotal evidence, political rhetoric, and budgetary audits that tracked spending rather than results. The idea that government interventions could be scrutinized with the same rigor applied to scientific experiments represented a paradigm shift—one that fundamentally altered the relationship between policymakers, analysts, and the public they serve.

The growth of policy evaluation as a discipline was driven by a central problem: how do we know if a policy is actually working? This question, deceptively simple on its surface, requires evaluators to define what "working" means, establish baseline conditions, identify causal mechanisms, and measure outcomes in ways that are both valid and reliable. Without a structured approach to evaluation design, governments risk perpetuating ineffective programs, misallocating scarce resources, or abandoning initiatives that are, in fact, producing important benefits.

1932
New Deal and Early Program Assessment
Franklin Roosevelt's New Deal programs generated unprecedented demand for understanding whether large-scale federal interventions produced their intended economic effects, prompting early social surveys and output tracking.
1965
Great Society and the Evaluation Boom
Lyndon Johnson's Great Society legislation—including the Elementary and Secondary Education Act—mandated formal program evaluations, institutionalizing the practice within federal agencies and spurring the growth of evaluation as a professional field.
1970s
Rise of Experimental Methods
Landmark studies such as the RAND Health Insurance Experiment and the Negative Income Tax Experiments demonstrated that randomized controlled trials could be applied to social policy, raising the methodological bar for credible evaluation.
1993
GPRA and Performance Measurement
The Government Performance and Results Act required all federal agencies to develop strategic plans with measurable performance goals, embedding outcome metrics into the fabric of American public administration.
2010s
Evidence-Based Policymaking Movement
The creation of the Commission on Evidence-Based Policymaking (2016) and the Foundations for Evidence-Based Policymaking Act (2018) signaled a bipartisan commitment to using rigorous evaluation questions and outcome data to guide legislative and administrative decisions.

This historical trajectory reveals a persistent gap that policy evaluation design seeks to close: the distance between a policy's stated intentions and its measurable effects. Designing effective evaluation questions and outcome metrics is the foundational skill that makes credible assessment possible—without it, even the most sophisticated analytical methods lack direction and purpose.

Core Principles of Policy Evaluation Design

Before constructing specific evaluation questions or selecting metrics, evaluators must ground their work in a set of foundational principles that ensure rigor, relevance, and utility. These principles distinguish systematic evaluation from casual judgment, and they provide the intellectual scaffolding upon which all subsequent design decisions rest. A well-designed evaluation is not merely a post-hoc accounting exercise; it is a structured inquiry that begins with clear conceptual commitments about what matters, why it matters, and how we can know whether change has occurred.

1

Theory of Change

Every evaluation begins with an explicit theory of change—a causal narrative that maps how inputs and activities are expected to produce outputs, outcomes, and ultimately impacts. This theory provides the logical foundation for choosing what to measure.
2

Evaluability

Not every policy question is evaluable. Evaluability assessment examines whether a program is sufficiently well-defined, whether data are accessible, and whether stakeholders are willing to use findings—ensuring that evaluation resources are invested wisely.
3

Validity and Reliability

Evaluation questions must be answerable using metrics that are valid (they measure what they claim to measure) and reliable (they produce consistent results under similar conditions). Without these properties, findings cannot support credible conclusions.
4

Stakeholder Relevance

Evaluation questions should address the information needs of key stakeholders—legislators, program managers, beneficiaries, and the public. A technically rigorous evaluation that answers questions nobody is asking fails the utilization test.
5

Counterfactual Reasoning

Causal evaluation questions require a counterfactual—an estimate of what would have happened in the absence of the policy. Designing for this from the outset ensures that outcome metrics can support causal claims rather than mere correlation.
KEY TAKEAWAY
Think of evaluation design like an architect's blueprint: before you pour concrete or hang drywall, you need a plan that specifies what the structure is supposed to achieve, what materials you will use, and how you will verify that the building meets code. A policy evaluation question is your blueprint's purpose statement, and outcome metrics are the engineering specifications—together, they determine whether the finished product stands up to scrutiny.

The Logic Model — Mapping Policy to Outcomes

The logic model is the most widely used visual framework for organizing evaluation design. It traces the causal chain from the resources invested in a policy (inputs) through the work performed (activities), to the direct products of that work (outputs), the short- and medium-term changes observed (outcomes), and the ultimate societal changes the policy aims to achieve (impacts). The following diagram illustrates this chain using a hypothetical job training program as an example.

The logic model traces the causal chain from inputs to impacts. Evaluation questions are mapped to specific stages: process evaluation questions address the left side (inputs, activities, outputs), while outcome evaluation questions target the right side (outcomes and impacts).

Notice that the logic model makes the theory of change explicit and testable. Each arrow in the diagram represents a causal assumption that can be interrogated through a specific evaluation question. For instance, the arrow between activities and outputs invites a process question such as "Are training sessions being delivered as designed?" The arrow between outputs and outcomes invites an outcome question such as "Did program completers experience higher employment rates than comparable non-participants?" By mapping evaluation questions to the logic model, evaluators ensure comprehensive coverage and avoid the common pitfall of measuring only what is easy rather than what is important.

Designing Evaluation Questions — A Structured Approach

Evaluation questions are the engine of any policy evaluation; they determine what data to collect, what methods to employ, and what conclusions can be drawn. A well-crafted evaluation question exhibits several structural properties: it is specific enough to be answerable, measurable through available or collectible data, aligned with the policy's theory of change, and useful to stakeholders who must act on the findings.

Three Types of Evaluation Questions

Evaluation scholars commonly distinguish among three types of evaluation questions, each corresponding to a different layer of the logic model. Descriptive questions ask "what is happening?" and are appropriate for documenting implementation fidelity, describing the population served, or cataloging program outputs. Normative questions ask "is what is happening consistent with what should be happening?" and require a benchmark, standard, or target against which observed performance is compared. Causal questions ask "did the policy cause the observed change?" and demand the most rigorous research designs—typically experimental or strong quasi-experimental methods that establish a credible counterfactual.

Three types of evaluation questions with templates and examples
Question TypeTemplateExample (Job Training Program)
DescriptiveWhat is the [characteristic] of [population/process]?What are the demographic characteristics of program enrollees?
NormativeTo what extent does [observed metric] meet [benchmark/standard]?Does the program's completion rate meet the agency target of 70%?
CausalDid [intervention] cause a change in [outcome] compared to [counterfactual]?Did participation in the program increase employment rates relative to non-participants?

Crafting Strong Evaluation Questions: The SMART-E Framework

A useful heuristic for crafting evaluation questions is the SMART-E framework, adapted from management science for evaluation contexts. Each question should be Specific (identifies the population, intervention, and comparison), Measurable (linked to quantifiable or systematically observable indicators), Attributable (the design permits causal inference or clearly disclaims it), Relevant (addresses stakeholder priorities), Time-bound (specifies the observation window), and Ethical (respects the rights and dignity of those studied). Applying SMART-E at the question design stage prevents many downstream methodological problems.

💡 Weak vs. Strong Evaluation Questions
Weak: "Is the program successful?" — This question is vague, unmeasurable, and provides no guidance for data collection. Strong: "Among adults aged 18–45 who completed the six-month training program in 2024, did the program increase the probability of full-time employment within 12 months of completion compared to eligible non-participants?" — This question specifies the population, intervention, outcome metric, time frame, and comparison group.

Designing Outcome Metrics — From Concepts to Indicators

An evaluation question without a corresponding outcome metric is a question without an answer key. Outcome metrics operationalize abstract policy goals—transforming concepts like "improved health," "reduced crime," or "increased economic opportunity" into concrete, observable, and measurable indicators. The process of operationalization involves moving through three levels of abstraction: the construct (the theoretical concept), the indicator (the observable dimension of the construct), and the measure (the specific data point collected).

The operationalization pyramid shows how the abstract construct of "Economic Opportunity" decomposes into observable indicators (employment status, earnings level, job stability) and then into specific, measurable data points with defined variable types and time frames.

Criteria for Strong Outcome Metrics

  • Construct validity: The metric genuinely captures the underlying concept it is intended to represent. Standardized test scores may or may not be a valid indicator of "learning," depending on the context.
  • Sensitivity: The metric must be capable of detecting a plausible change. If a policy is expected to produce a modest effect, the metric should have sufficient precision and variability to register it.
  • Feasibility: Data must be collectible within the evaluation's budget, timeline, and ethical constraints. Administrative data (e.g., unemployment insurance records) are often more feasible than primary surveys.
  • Minimal gaming risk: Metrics that can be easily manipulated by program operators (e.g., counting only favorable cases) undermine credibility. Evaluators should consider how incentives might distort measurement.
  • Disaggregability: Strong metrics can be broken down by subgroup (race, gender, geography, income level), enabling equity-focused analysis and identification of differential program effects.

Output Metrics vs. Outcome Metrics

A common error in evaluation design is conflating output metrics with outcome metrics. Output metrics count what the program produces (number of training sessions held, participants enrolled, brochures distributed), while outcome metrics capture changes in the condition the policy aims to affect (employment rates, health status, recidivism rates). A program can generate impressive outputs without producing any meaningful outcomes—an important distinction that evaluation questions must explicitly address.

Worked Example — Designing an Evaluation for a Municipal Anti-Recidivism Program

Suppose a mid-sized city has launched a Reentry Support Program (RSP) providing case management, housing assistance, and vocational training to formerly incarcerated individuals within 90 days of release. The city council has asked for an evaluation. Walk through the process of designing evaluation questions and outcome metrics step by step.

Designing an Evaluation for the Reentry Support Program
1
Step 1 — Articulate the Theory of ChangeBegin by drafting the program's theory of change. The RSP assumes that formerly incarcerated individuals reoffend partly because they lack stable housing, employment skills, and social support. By providing these resources (inputs/activities), the program expects to produce service connections (outputs) that lead to stable housing and employment (outcomes), which in turn reduce recidivism (impact). Writing this narrative explicitly reveals each causal assumption that the evaluation can test.
Theory of change: Case management + housing + training → service connections → stable employment & housing → reduced recidivism
2
Step 2 — Identify Stakeholder Information NeedsConsult with key stakeholders to understand their priority questions. The city council wants to know whether the program reduces re-arrest rates. The program director wants to know whether the case management model is being implemented consistently. Community advocates want to know whether the program serves participants equitably across racial groups. These distinct needs generate different types of evaluation questions.
Three stakeholder groups → three categories of evaluation questions (causal, process, equity)
3
Step 3 — Draft Evaluation Questions Using SMART-EApply the SMART-E framework. Causal question: "Among adults released from the county jail between January and December 2024 who enrolled in the RSP within 90 days, did the program reduce the probability of re-arrest within 24 months compared to eligible individuals who did not enroll?" Process question: "To what extent did RSP participants receive the full package of case management contacts, housing referrals, and vocational training as specified in the program manual?" Equity question: "Do recidivism reduction outcomes differ by race, gender, or prior conviction history?" Each question specifies population, intervention, comparison, outcome, and time frame.
Three SMART-E-compliant evaluation questions drafted, covering causal impact, process fidelity, and equity.
4
Step 4 — Operationalize Outcome MetricsFor each evaluation question, select concrete outcome metrics. The primary outcome metric for the causal question is the re-arrest rate within 24 months of release, operationalized as a binary variable (re-arrested or not) using county criminal justice records. Secondary metrics include time-to-first-re-arrest (a survival variable), re-conviction rate, and days of incarceration. For the process question, metrics include the percentage of participants receiving ≥ 12 case management contacts, the percentage placed in housing within 30 days, and the percentage enrolled in vocational training within 60 days. For the equity question, all outcome metrics are disaggregated by race, gender, and conviction history.
Primary metric: 24-month re-arrest rate (binary). Secondary metrics: time-to-re-arrest, re-conviction rate, days incarcerated. Process metrics: service dosage rates. All disaggregated by demographics.
5
Step 5 — Construct an Evaluation MatrixCompile all elements into an evaluation matrix—a table that links each evaluation question to its corresponding metrics, data sources, data collection methods, and analysis plan. This matrix serves as the master planning document for the entire evaluation. For example, the causal question maps to the re-arrest rate metric, sourced from county jail records, collected through administrative data extraction, and analyzed using propensity score matching to approximate a counterfactual comparison group.
Evaluation matrix complete: Questions → Metrics → Data Sources → Methods → Analysis Plan

Strengths and Limitations of Common Outcome Metrics

No single outcome metric perfectly captures a policy's effect; each involves trade-offs between validity, feasibility, and interpretability. Understanding these trade-offs is essential for designing evaluations that produce credible and actionable findings. The following table compares several common metric types across key evaluation design criteria.

Comparison of common outcome metric types
Metric TypeStrengthsLimitations
Administrative records (e.g., arrest data, tax filings)Low cost; large sample sizes; objective; longitudinal tracking possible; no respondent burdenLimited to what agencies collect; may not capture constructs of interest (e.g., well-being); data quality varies across jurisdictions; privacy constraints
Self-report surveysCapture subjective experiences (satisfaction, perceived safety); customizable to evaluation questions; can collect data not available elsewhereSocial desirability bias; recall errors; non-response bias; costly to administer at scale; attrition in longitudinal designs
Standardized indices (e.g., poverty rate, Gini coefficient)Widely understood; comparable across jurisdictions and time; pre-validated measurement instruments availableMay mask variation within subgroups; can be insensitive to small-scale interventions; threshold definitions may be arbitrary
Behavioral measures (e.g., observed compliance, attendance)Objective; directly observable; difficult to fake; high face validityObserver bias; Hawthorne effects (behavior changes because of observation); may not capture intent or motivation; resource-intensive to collect
Composite metrics (e.g., Human Development Index)Capture multi-dimensional constructs; useful for ranking and benchmarking; reduce data overload into single summary statisticWeighting choices are subjective; difficult to interpret changes in component parts; can obscure trade-offs between dimensions
KEY TAKEAWAY
Choosing outcome metrics is analogous to selecting instruments for a scientific laboratory: no single thermometer measures temperature at every scale, and no single metric captures every dimension of a policy's effect. Effective evaluation design typically employs a portfolio of complementary metrics—combining, for instance, administrative data on employment with survey data on job satisfaction—to triangulate findings and compensate for the limitations of any single measure.

Connecting Basic Design to Advanced Evaluation Methods

The evaluation questions and outcome metrics you design at the outset directly constrain and enable the advanced methods available downstream. A causal evaluation question, for instance, demands a research design that establishes a credible counterfactual—this might involve a randomized controlled trial (RCT), a regression discontinuity design (RDD), difference-in-differences (DiD), or propensity score matching. Each of these methods has distinct data requirements that must be anticipated at the design stage.

How basic evaluation design concepts connect to advanced methods
Concept in This LessonAdvanced ExtensionWhy Design Decisions Matter Early
Evaluation question types (descriptive, normative, causal)Research design selection (RCT, quasi-experimental, mixed-methods)A causal question asked after implementation without baseline data cannot be answered with an RCT; design must precede implementation.
Theory of change / Logic modelProcess tracing, contribution analysis, realist evaluationComplex policies with multiple causal pathways require explicit theories of change to guide data collection on mediating variables.
Outcome metric operationalizationStatistical power analysis, minimum detectable effect sizeMetric variability and expected effect size determine the sample size needed; this must be calculated before data collection begins.
Counterfactual reasoningPotential outcomes framework (Rubin causal model)The formal notation of causal inference—Y₁ − Y₀ = treatment effect—makes explicit what the evaluation question implies informally.
Evaluation matrixPre-analysis plans and evaluation registriesFormalizing the evaluation matrix into a pre-analysis plan prevents post-hoc data mining and enhances the credibility of findings.

The overarching lesson is that evaluation design is not merely a preliminary step to be rushed through before "real" analysis begins—it is the intellectual foundation upon which the entire edifice of credible evidence rests. Students who master the art of crafting precise evaluation questions and valid outcome metrics will find that advanced methods become tools in service of well-defined purposes, rather than techniques in search of a problem.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain the difference between an output metric and an outcome metric. Using a hypothetical public library literacy program as your example, provide one output metric and one outcome metric, and explain why the distinction matters for policy evaluation.
PROBLEM 2BASIC APPLICATION
A state government implements a program offering subsidized health insurance to uninsured residents earning below 200% of the federal poverty level. Using the SMART-E framework, draft one causal evaluation question for this program. Identify which element of your question corresponds to each letter in SMART-E.
PROBLEM 3INTERMEDIATE
You are designing an evaluation for a city's new police body camera program. The mayor wants to know if body cameras reduce use-of-force incidents and improve public trust. Draft a logic model (listing inputs, activities, outputs, outcomes, and impacts) for this program, and then write one normative evaluation question and one causal evaluation question, each linked to a specific outcome metric.
PROBLEM 4APPLIED
A federal agency is evaluating a nationwide school nutrition program that provides free breakfast and lunch to students in low-income schools. The agency has administrative data on student enrollment, meal counts, and standardized test scores. Propose a set of three complementary outcome metrics (one primary, two secondary) that address both health and educational dimensions of the program. For each metric, identify the data source, variable type, and one potential validity concern.
PROBLEM 5CRITICAL THINKING
A state legislature mandates that all publicly funded social programs must demonstrate a statistically significant improvement in at least one pre-specified outcome metric within three years or face automatic defunding. Critically analyze this mandate from the perspective of evaluation design principles covered in this lesson. Identify at least three problems with this approach and propose a more nuanced alternative framework for linking evaluation findings to funding decisions.

Summary — Policy Evaluation Design

Policy evaluation design begins with a theory of change that maps the causal chain from inputs through activities and outputs to outcomes and impacts, typically visualized through a logic model. Evaluation questions fall into three categories—descriptive, normative, and causal—each requiring different levels of methodological rigor and each mapped to specific stages of the logic model. The SMART-E framework ensures that questions are specific, measurable, attributable, relevant, time-bound, and ethical.

Outcome metrics operationalize abstract policy goals by moving from constructs to indicators to specific measures, evaluated for construct validity, sensitivity, feasibility, gaming risk, and disaggregability. All of these elements are compiled into an evaluation matrix that links questions to metrics, data sources, and analytical methods—serving as the master planning document for credible, useful policy evaluation.

Varsity Tutors • College Political Science • Policy Evaluation Design — Design basic policy evaluation questions and outcome metrics