All questions
Question 1
A philanthropic foundation aims to fund programs that increase 'civic engagement' in young adults. The foundation must choose a reliable outcome metric to evaluate the effectiveness of its grants. Which of the following metrics would be the most problematic and least valid proxy for the abstract concept of civic engagement?
- Changes in voter turnout rates for individuals aged 18-25 in local and national elections.
- The number of hours per capita that young adults report volunteering for community organizations.
- The quantity of political lawn signs and bumper stickers displayed by young adults during an election season. (correct answer)
- The percentage of young adults who report contacting a public official in the last year.
Explanation: The correct answer is C. Displaying lawn signs or bumper stickers is a superficial form of expression that requires minimal effort and may not reflect genuine, sustained civic engagement. Its measurement is also difficult and highly susceptible to confounding factors like campaign distribution efforts. Distractors A, B, and D, while all imperfect proxies, are more established and direct indicators of active participation in civic life (voting, volunteering, contacting officials) and are commonly used in political science research.
Question 2
A state agency launches a new vocational training program for recently unemployed individuals. To rigorously evaluate the program's causal effect on employment, which evaluation question provides the strongest research design?
- How do the employment rates of program graduates compare to the statewide average unemployment rate?
- Among program participants, what is the difference in income between those who completed the training and those who dropped out?
- What is the difference in subsequent employment and earnings between individuals who entered the program and a randomly assigned control group of similar individuals who did not? (correct answer)
- According to surveys, what percentage of program graduates believe the training was instrumental in their finding a new job?
Explanation: The correct answer is C because it describes a randomized controlled trial (RCT). By comparing the treatment group (program participants) to a randomly assigned control group, this design is best able to isolate the causal effect of the program from other factors that could influence employment. Distractor A uses a poor comparison group (the statewide average), which is not comparable to the specific population of unemployed individuals. Distractor B suffers from selection bias, as people who complete the program may be inherently more motivated than those who drop out. Distractor D relies on subjective self-reporting, which is less reliable than objective employment data.
Question 3
A police department adopts a 'broken windows' strategy, significantly increasing enforcement against minor quality-of-life offenses in specific neighborhoods with the stated goal of deterring more serious crime. An analyst plans to evaluate the policy's effectiveness. Which proposed evaluation design is most likely to commit an ecological fallacy?
- Comparing precinct-level rates of serious crime before and after the strategy was implemented in those same precincts.
- Correlating the city-wide total number of minor infraction citations with the city-wide total number of violent crimes. (correct answer)
- Interviewing residents of the targeted neighborhoods about their perceptions of safety and police presence.
- Tracking a cohort of individuals cited for minor infractions to measure their subsequent involvement in serious crime.
Explanation: The correct answer is B. The ecological fallacy is the error of making inferences about individuals based on aggregate data for a group. Correlating two city-wide totals (citations and violent crimes) and then drawing a conclusion about whether the specific strategy worked in specific neighborhoods is a classic example. The city-wide crime rate could be driven by trends in areas where the policy was not even implemented. The other options are more valid: A is a localized pre-post analysis, C gathers individual perceptions, and D uses individual-level longitudinal data.
Question 4
A city's parks department establishes a new strategic goal to 'enhance the quality of public green spaces.' To track progress, the department needs to operationalize this abstract goal. Which of the following sets of metrics offers the most valid and comprehensive measurement of 'quality'?
- The total acreage of parkland managed by the department and the annual departmental budget.
- The number of new trees planted and the linear feet of new walking paths installed each year.
- The total number of unique visitors to city parks as measured by gate counters or other tracking technology.
- A composite index based on user satisfaction surveys, independent safety audits, and standardized measures of cleanliness. (correct answer)
Explanation: The correct answer is D. 'Quality' is a multi-dimensional concept. This option provides a comprehensive operationalization by combining user perception (surveys), objective safety (audits), and maintenance standards (cleanliness). Distractor A consists of input metrics (budget, acreage), not quality outcomes. Distractor B measures specific outputs which may or may not improve overall quality. Distractor C measures usage, which could be high even if quality is low (e.g., due to lack of other options), making it an unreliable proxy for quality.
Question 5
A non-governmental organization (NGO) provides microloans to female entrepreneurs in a rural region of a developing country. The program's central goal is to foster their economic empowerment. Which of the following provides the most robust and meaningful outcome metric for this specific goal?
- The repayment rate on the loans and the number of new loans disbursed per year.
- The growth in annual business profits and personal savings for program participants, compared to a matched control group of non-participants. (correct answer)
- The change in the gross domestic product (GDP) of the country where the program operates.
- The number of women who report feeling 'more confident' or 'empowered' in qualitative, post-program interviews.
Explanation: The correct answer is B. This metric uses concrete economic indicators (profits, savings) that are central to economic empowerment and, crucially, uses a control group to help establish a causal link between the program and the outcomes. Distractor A measures program outputs and financial health, not the impact on the entrepreneurs. Distractor C uses a metric that is far too broad; a small microloan program cannot have a measurable effect on national GDP. Distractor D relies on subjective self-reporting, which can be valuable but is less robust and more prone to bias than objective economic data.
Question 6
The Food and Drug Administration (FDA) mandates a new, more prominent 'black box' warning on a common over-the-counter painkiller, highlighting the risk of severe liver damage. The immediate goal of this regulatory policy is to increase consumer awareness of this specific risk. Which of the following is the most direct metric for evaluating the success of this specific goal?
- The pre/post change in the percentage of the medication's users who can correctly identify its primary risk in a survey. (correct answer)
- The pre/post change in the number of poison control center calls related to overdoses of the medication.
- The change in the medication's total unit sales in the quarter following the label change.
- The average cost per unit for the manufacturer to produce the newly compliant packaging.
Explanation: When evaluating regulatory policy effectiveness, you need to match your evaluation metric directly to the stated policy goal. The FDA's goal here is specifically to "increase consumer awareness" of liver damage risk, so you need a metric that directly measures awareness levels.
Option A is correct because it directly measures the policy's stated objective. A survey asking users to identify the medication's primary risk before and after the label change would show whether the black box warning successfully increased awareness of the specific liver damage risk. This metric has a clear causal link to the regulatory intervention and measures exactly what the policy aimed to achieve.
Option B measures behavioral outcomes (poison control calls) rather than awareness. While fewer overdose calls might be a desirable long-term effect, it doesn't directly measure whether people are more aware of the risks—they could be aware but still take risky actions, or calls could decrease for unrelated reasons.
Option C focuses on sales volume, which measures market response rather than consumer knowledge. Sales could drop due to price changes, competition, or seasonal factors completely unrelated to awareness of health risks.
Option D measures implementation costs for manufacturers, which tells you nothing about whether consumers gained awareness. This is an input metric related to compliance costs, not an outcome metric related to policy effectiveness.
Strategy tip: On political science questions about policy evaluation, always identify the specific policy goal first, then look for the metric that most directly measures that exact outcome. Avoid metrics that measure related but indirect effects.
Question 7
A city offers a significant property tax abatement to developers who include a certain percentage of affordable housing units in new construction projects. The goal is to increase the net supply of affordable housing. Which outcome metric best captures the policy's true effect, accounting for potential market dynamics?
- The gross number of new, subsidized affordable housing units created by developers receiving the tax abatement.
- The total monetary value of the property tax abatements granted to developers under the program.
- The change in the city's total stock of affordable housing units, compared to the projected trend without the policy. (correct answer)
- The percentage of residential developers in the city who report that the tax abatement was a key factor in their decision-making.
Explanation: The correct answer is C because it measures the net change in affordable housing stock against a counterfactual baseline (the projected trend). This approach accounts for the possibility that some units would have been built anyway or that the new construction displaced other affordable units. Distractor A is a gross measure that suffers from this attribution problem. Distractor B is a cost/input metric. Distractor D is a subjective measure of influence, not a measure of the actual housing supply outcome.
Question 8
A large city enacts a ban on single-use plastic bags to reduce plastic pollution in local waterways. To evaluate the policy, the city's sanitation department proposes four potential outcome metrics. Which metric most directly measures the policy's intended environmental impact?
- The total number of fines issued to businesses for non-compliance with the bag ban.
- The reported sales figures for reusable shopping bags from major retailers within the city.
- The tonnage of plastic bags collected from storm drain filters and waterway cleanup events, per quarter. (correct answer)
- A city-wide survey measuring changes in public attitudes toward environmental conservation.
Explanation: The correct answer is C because it is a direct measurement of the targeted problem: plastic bag pollution in the local environment. This metric captures the actual reduction of the pollutant. Distractor A measures policy enforcement (an output), not its environmental effect. Distractor B is a proxy metric for behavior change, but it doesn't confirm that the use of reusable bags is replacing plastic bags or reducing overall waste. Distractor D measures public opinion, which may or may not correlate with the actual environmental outcome.
Question 9
A state legislature decriminalizes the possession of small amounts of marijuana. The policy's primary goals are to reduce the burden on the criminal justice system and to mitigate racial disparities in drug-related arrests. Which of the following evaluation questions best addresses these specific goals?
- What has been the subsequent trend in tax revenue from legal cannabis sales in the state?
- How has the rate of marijuana-related hospitalizations changed since the policy was enacted?
- Have arrest rates for marijuana possession decreased, and has the racial gap in those arrest rates narrowed? (correct answer)
- What is the level of public support for the decriminalization policy one year after implementation?
Explanation: The correct answer is C because it directly frames the evaluation around the two explicit goals of the policy: reducing the justice system's burden (measured by overall arrest rates) and addressing racial disparities (measured by the racial gap in arrest rates). Distractor A is irrelevant as decriminalization is distinct from legal sales and taxation. Distractor B focuses on a potential public health consequence, not the stated criminal justice goals. Distractor D measures public opinion, not the policy's tangible effects on the justice system.
Question 10
A city government imposes a tax on sugar-sweetened beverages with the long-term goal of reducing obesity and diabetes. An evaluation is planned to assess the policy's progress after 18 months. Which of the following represents the most appropriate intermediate outcome metric to measure at this stage?
- The change in the prevalence of type 2 diabetes among the city's residents.
- The total annual revenue generated for the city by the beverage tax.
- The per capita volume of taxed beverages sold, compared to pre-tax levels and sales in a neighboring city. (correct answer)
- The number of beverage retailers that have successfully implemented the tax collection system.
Explanation: The correct answer is C. The policy's causal logic is that the tax will reduce consumption, which will then lead to better health outcomes. Measuring the change in beverage sales is a crucial intermediate step to see if the policy's primary mechanism is working. Distractor A is a long-term impact, which is unlikely to be measurable in just 18 months and is affected by many other factors. Distractor B is a fiscal byproduct, not a health outcome. Distractor D is a process metric related to implementation fidelity.
Question 11
A school district implements a new anti-bullying curriculum across all middle schools. After the first year, administrators want to determine if the program is being delivered as intended by the curriculum designers. Which of the following is a process evaluation question, as opposed to an outcome evaluation question?
- Has the average number of disciplinary actions related to bullying decreased since the curriculum was introduced?
- Do students report feeling a greater sense of safety and belonging in school compared to the previous year?
- Are teachers dedicating the prescribed number of classroom hours to the curriculum and using the official materials? (correct answer)
- Is there a statistically significant difference in bullying incidents between schools that adopted the curriculum and those that did not?
Explanation: The correct answer is C. A process evaluation (or implementation evaluation) focuses on how a program is being operated and whether it is being delivered with fidelity to its original design. This question addresses the 'how' and 'what' of program delivery. Distractors A, B, and D are all examples of outcome or impact evaluation questions because they focus on the effects and results of the program on its target population.
Question 12
A municipal government installs an extensive network of high-visibility surveillance cameras in public squares and commercial districts to deter street crime. While the primary evaluation will focus on crime statistics, an administrator wants to also assess potential negative externalities. Which evaluation question best targets a widely discussed 'chilling effect' of such surveillance?
- What was the total cost of installing and maintaining the camera network over its first five years of operation?
- Did the installation of cameras displace criminal activity to adjacent, non-surveilled neighborhoods?
- How did public perception of safety in the commercial districts change after the cameras were installed?
- Was there a measurable change in the frequency and size of lawful public demonstrations and political rallies in the surveilled areas? (correct answer)
Explanation: When evaluating public surveillance systems, you need to distinguish between direct policy outcomes and broader societal effects known as "externalities." The "chilling effect" is a specific concept in political science referring to when government surveillance or oversight discourages citizens from exercising their constitutional rights, particularly free speech and assembly, even when those activities are perfectly legal.
Option D correctly identifies this chilling effect by measuring whether surveillance cameras reduced lawful political demonstrations and rallies. If fewer people are willing to participate in legal protests or political gatherings because they feel watched and potentially monitored by authorities, this represents the classic chilling effect on democratic participation.
Option A focuses purely on financial costs, which are important for budget analysis but don't address negative externalities affecting civil liberties. Option B examines crime displacement, which is certainly a relevant externality for crime policy, but it's not the "chilling effect" specifically mentioned in the question. Option C measures public safety perceptions, which would actually be considered a positive outcome of the surveillance program rather than a negative externality.
The key distinction is that the chilling effect specifically concerns how government surveillance might suppress constitutionally protected activities. Citizens might avoid participating in protests or political events not because these activities are illegal, but because they fear being identified, recorded, or potentially targeted later.
Remember: when you see questions about surveillance and civil liberties, look for answers that address impacts on constitutional rights and democratic participation, not just crime or economic measures.
Question 13
A transportation authority is conducting a cost-benefit analysis for a proposed new bridge. The primary purpose of the bridge is to alleviate traffic congestion and reduce travel times. In quantifying the project's benefits, which metric should be given the most weight?
- The number of temporary construction jobs that will be created during the building phase of the project.
- The projected annual toll revenue the bridge is expected to generate for the authority.
- The increase in property values in the neighborhoods located closest to the bridge access points.
- The calculated monetary value of the cumulative time saved by commuters and commercial vehicles. (correct answer)
Explanation: Cost-benefit analysis is a cornerstone of public policy evaluation that requires weighing a project's total social benefits against its costs. When analyzing public infrastructure like bridges, you need to focus on the primary purpose and measure benefits that directly address the problem the project aims to solve.
The correct answer is D because the bridge's stated primary purpose is alleviating traffic congestion and reducing travel times. The monetary value of cumulative time saved represents the core benefit that justifies the project's existence. This metric captures the economic value of productivity gains for both commuters and commercial vehicles, translating the bridge's main function into quantifiable terms that can be compared against construction and maintenance costs.
Option A is incorrect because temporary construction jobs are short-term economic effects, not lasting benefits of the completed bridge. These jobs would exist for any major construction project regardless of its ultimate value. Option B focuses on revenue generation rather than social benefits – toll revenue is simply cost recovery, not a measure of public benefit. A bridge could generate substantial tolls while providing minimal traffic relief. Option C addresses property value increases, which are secondary spillover effects rather than the primary benefit the bridge is designed to deliver.
When approaching cost-benefit analysis questions, always identify the project's stated primary objective first, then look for the metric that most directly measures success in achieving that goal. Don't be distracted by secondary effects or revenue considerations – focus on the core public benefit the policy is designed to create.
Question 14
The Centers for Disease Control and Prevention (CDC) launches a national public service announcement (PSA) campaign to increase the rate of annual influenza vaccinations. Which evaluation design would be most effective at isolating the specific impact of the media campaign itself?
- A comparison of the national vaccination rate in the year of the campaign to the rate from the previous year.
- A survey administered to a random sample of vaccinated individuals, asking if the PSA campaign influenced their decision.
- An analysis tracking the correlation between the number of times the PSAs were aired on national television and the weekly national vaccination data.
- A comparison of vaccination rates between a set of randomly chosen media markets where the PSAs were aired and a similar set of markets where they were not. (correct answer)
Explanation: This question tests your understanding of causal inference and experimental design in policy evaluation. When evaluating whether a specific intervention (like a PSA campaign) caused an observed outcome, you need to isolate that intervention's effect from other factors that might influence the result.
Option D provides the strongest causal evidence because it uses a randomized controlled design. By randomly selecting some media markets to receive the PSAs while withholding them from similar markets, you create a natural experiment. Any difference in vaccination rates between these groups can be attributed to the campaign itself, since random assignment controls for other variables that might affect vaccination behavior.
Option A fails because it can't separate the campaign's effect from countless other factors that changed between years—seasonal patterns, news events, policy changes, or shifting public attitudes. A simple before-and-after comparison proves correlation, not causation.
Option B introduces serious selection bias. People who got vaccinated may not accurately remember or report what influenced them, and those who didn't vaccinate aren't included at all. Self-reported influence is notoriously unreliable for measuring actual behavioral impact.
Option C shows correlation between PSA frequency and vaccination rates, but correlation doesn't establish causation. Higher vaccination rates might drive more PSA airings (reverse causation), or both might respond to the same underlying factors like seasonal flu outbreaks.
When you see policy evaluation questions, look for designs that create comparison groups and control for confounding variables—randomized experiments provide the gold standard for establishing causal relationships.
Question 15
A state agency responsible for issuing professional licenses overhauls its system, moving from a paper-based process to a fully online portal. A key goal of the new policy is to 'increase agency efficiency.' Which metric would best evaluate the achievement of this specific goal?
- The reduction in the average staff hours and material costs required to process one complete application. (correct answer)
- The average applicant-reported satisfaction score with the new online portal.
- The change in the total number of licenses the agency issues per fiscal year.
- The year-over-year percentage change in the number of applications received by the agency.
Explanation: When evaluating public policy effectiveness, you need to match your measurement metric directly to the stated policy goal. This question tests whether you can distinguish between measuring efficiency versus other policy outcomes like satisfaction or volume.
The agency's specific goal is to "increase efficiency" through digitization. Efficiency fundamentally means accomplishing the same task with fewer resources—less time, money, or effort per unit of output. Answer A captures this perfectly by measuring the reduction in staff hours and material costs required to process each application. If the new system reduces these inputs while maintaining the same output quality, efficiency has demonstrably improved.
Answer B measures user satisfaction, which relates to service quality or user experience rather than operational efficiency. While satisfied users are valuable, high satisfaction doesn't necessarily mean the agency is operating more efficiently. Answer C tracks total licenses issued annually, but this measures overall productivity or capacity, not efficiency. The agency could issue more licenses simply due to increased demand while actually becoming less efficient per application. Answer D examines application volume received, which reflects external demand factors completely outside the agency's control and unrelated to internal operational efficiency.
Remember this key distinction for policy analysis questions: efficiency specifically measures the input-to-output ratio, while other common policy goals include effectiveness (achieving desired outcomes), equity (fair distribution), or responsiveness (meeting citizen needs). Always match your evaluation metric to the precise policy objective stated in the question.
Question 16
To evaluate a 'College Promise' program that provides free community college tuition, a researcher interviews only students who enrolled in the program and successfully graduated. This research design, which suffers from selection bias, would be least able to provide a valid answer to which of the following evaluation questions?
- What were the primary motivations for successful graduates to enroll in the program?
- What is the program's overall impact on the college completion rate for all eligible students? (correct answer)
- What academic and support services did the program's successful graduates find most helpful?
- How do employers in the local community perceive the job-readiness of the program's graduates?
Explanation: The correct answer is B. The research design exclusively examines successes (graduates) and completely ignores failures (students who enrolled but did not complete). Therefore, it is impossible to calculate an overall completion rate or determine the program's net impact on this metric. The design can, however, provide valid (though limited) insights for the other questions, which are specifically about the experiences and perceptions of the successful graduates themselves.
Question 17
A city government makes a major investment to create a downtown arts district, with the broad goal of 'urban revitalization.' The evaluation plan includes several outcome metrics. Which of the following metrics provides the LEAST direct evidence of whether this specific goal has been achieved?
- The number of individuals who identify as 'artists' taking up primary residence in the city as a whole. (correct answer)
- The change in commercial property values and retail vacancy rates inside the district compared to similar commercial areas elsewhere in the city.
- The year-over-year change in sales tax revenue from businesses located within the arts district's designated boundaries.
- Pre- and post-investment data on pedestrian foot traffic and visitor spending within the district.
Explanation: When evaluating public policy outcomes, you need to distinguish between metrics that directly measure your stated goal versus those that capture related but broader effects. Urban revitalization specifically refers to breathing new life into a particular area—making it more economically vibrant, attractive, and functionally successful as a destination.
Option A measures artist residency across the entire city, which is too geographically broad and conceptually indirect. While artists moving to the city might be loosely connected to the arts district's existence, this doesn't tell you whether the downtown district itself is being revitalized. Artists could settle anywhere in the city for reasons unrelated to your investment, making this metric poorly targeted to your specific goal.
Options B, C, and D all provide direct evidence of district-level revitalization. Option B measures economic vitality through property values and vacancy rates specifically within the target area, comparing it to control areas. Option C tracks actual economic activity (sales tax revenue) within the district's boundaries, showing whether businesses are thriving there. Option D captures whether the area is successfully attracting people and generating economic activity through foot traffic and spending data.
The key distinction is geographic specificity and causal proximity—B, C, and D measure changes happening in the exact location where you invested, while A measures a city-wide phenomenon that could occur independently of your district's success.
Remember: Strong policy evaluation requires metrics that are geographically and causally proximate to your intervention. Always ask whether the metric directly captures change in your target area.
Question 18
A large city enacts a ban on single-use plastic bags to reduce plastic pollution in local waterways. To evaluate the policy, the city's sanitation department proposes four potential outcome metrics. Which metric most directly measures the policy's intended environmental impact?
- The total number of fines issued to businesses for non-compliance with the bag ban.
- The reported sales figures for reusable shopping bags from major retailers within the city.
- The tonnage of plastic bags collected from storm drain filters and waterway cleanup events, per quarter. (correct answer)
- A city-wide survey measuring changes in public attitudes toward environmental conservation.
Explanation: The correct answer is C because it is a direct measurement of the targeted problem: plastic bag pollution in the local environment. This metric captures the actual reduction of the pollutant. Distractor A measures policy enforcement (an output), not its environmental effect. Distractor B is a proxy metric for behavior change, but it doesn't confirm that the use of reusable bags is replacing plastic bags or reducing overall waste. Distractor D measures public opinion, which may or may not correlate with the actual environmental outcome.
Question 19
A philanthropic foundation aims to fund programs that increase 'civic engagement' in young adults. The foundation must choose a reliable outcome metric to evaluate the effectiveness of its grants. Which of the following metrics would be the most problematic and least valid proxy for the abstract concept of civic engagement?
- Changes in voter turnout rates for individuals aged 18-25 in local and national elections.
- The number of hours per capita that young adults report volunteering for community organizations.
- The quantity of political lawn signs and bumper stickers displayed by young adults during an election season. (correct answer)
- The percentage of young adults who report contacting a public official in the last year.
Explanation: The correct answer is C. Displaying lawn signs or bumper stickers is a superficial form of expression that requires minimal effort and may not reflect genuine, sustained civic engagement. Its measurement is also difficult and highly susceptible to confounding factors like campaign distribution efforts. Distractors A, B, and D, while all imperfect proxies, are more established and direct indicators of active participation in civic life (voting, volunteering, contacting officials) and are commonly used in political science research.
Question 20
A school district implements a new anti-bullying curriculum across all middle schools. After the first year, administrators want to determine if the program is being delivered as intended by the curriculum designers. Which of the following is a process evaluation question, as opposed to an outcome evaluation question?
- Has the average number of disciplinary actions related to bullying decreased since the curriculum was introduced?
- Do students report feeling a greater sense of safety and belonging in school compared to the previous year?
- Are teachers dedicating the prescribed number of classroom hours to the curriculum and using the official materials? (correct answer)
- Is there a statistically significant difference in bullying incidents between schools that adopted the curriculum and those that did not?
Explanation: The correct answer is C. A process evaluation (or implementation evaluation) focuses on how a program is being operated and whether it is being delivered with fidelity to its original design. This question addresses the 'how' and 'what' of program delivery. Distractors A, B, and D are all examples of outcome or impact evaluation questions because they focus on the effects and results of the program on its target population.