All questions
Question 1
A town newsletter reports that 68 percent of residents support converting an unused municipal building into an arts center. The figure comes from an online questionnaire linked on the town's social-media page. Four hundred twelve people completed the questionnaire. Participation was voluntary, respondents were not asked whether they lived in town, and no information was collected about age, neighborhood, or other characteristics. The newsletter describes the result as "clear evidence that a large majority of all town residents favor the project."
What is the most important gap in the evidence supporting the newsletter's claim?
- The questionnaire did not ask respondents to explain the artistic programs they would prefer the new center to offer.
- The town cannot determine whether the voluntary respondents were residents or represented the broader population whose views are being described. (correct answer)
- The newsletter reports only the percentage favoring the project and does not convert that percentage into a whole-number count.
- The questionnaire received fewer responses than the town's total population, so no conclusion can be drawn from the results.
Explanation: Whenever you evaluate a claim based on survey or poll data, your first instinct should be to ask: Who actually responded, and do they represent the group being described? This question tests your ability to spot flaws in statistical reasoning — specifically, the problem of an unrepresentative sample.
The newsletter makes a sweeping claim about "all town residents," but the data comes from a voluntary online questionnaire shared only on the town's social-media page. No one verified whether respondents actually lived in town, and nothing was collected about their demographics. That means the 412 responses could include non-residents, people with unusually strong opinions (who were motivated to seek out and complete the survey), or a slice of the population that simply doesn't reflect the town as a whole. Because of these gaps, there is no valid basis for generalizing the result to all residents — making B the correct answer and the most critical flaw in the evidence.
A is irrelevant to the claim being evaluated. Whether respondents shared program preferences has nothing to do with whether the 68% figure is representative. C describes a non-issue: percentages are a standard, legitimate way to report survey results, and converting them to whole numbers wouldn't change the underlying problem. D reflects a common misconception — a sample doesn't need to equal the total population to be valid. What matters is whether the sample is representative, not just large enough.
Your strategy tip: when a question asks about a gap in evidence, focus on the match between the sample and the population being described. Voluntary, self-selected surveys are a classic red flag on this exam.
Question 2
After a county library extended its weekday hours until 9 p.m., the director announced that the change had improved completion rates in the library's adult education courses. As evidence, she quoted six graduates who said the later hours allowed them to attend after work. Enrollment records also showed that most classes held after 6 p.m. were nearly full. The library did not report completion rates from before the schedule change, results from branches that retained their previous hours, or information about students who enrolled but left the courses.
Which additional evidence would most directly test the director's claim that extended hours improved course completion?
- A list of the evening courses that attracted the greatest number of registrations after the hours were extended
- Interviews with additional graduates about whether they preferred evening classes to classes held during the day
- Records showing how much the library spent on staffing, security, and utilities during the extended hours
- Completion rates before and after the change, compared with similar branches that did not extend their hours (correct answer)
Explanation: When evaluating whether a policy change caused an improvement, you need comparison data — specifically, what happened before versus after the change, and what happened to similar groups that didn't experience the change. This is the logic behind a controlled comparison, and it's exactly what the director's evidence lacks.
The director claims extended hours improved completion rates, but she never actually reports completion rates at all — only full classrooms and positive testimonials. To directly test her claim, you'd need to know whether completion rates genuinely rose, and whether the schedule change (not some other factor) was responsible. D provides both: completion rates before and after the change, compared against branches that kept their original hours. That comparison isolates the effect of extended hours and rules out coincidental improvements happening across all branches regardless of schedule.
A is wrong because registration numbers don't tell you whether students finished their courses — high enrollment and high completion are different things. B is wrong because graduates' preferences about class timing don't reveal whether more students completed courses overall; it also samples only people who already succeeded. C is wrong because operational costs are relevant to budgeting decisions, not to whether the policy improved educational outcomes.
A common trap on reasoning questions like this is mistaking related evidence for relevant evidence. Full classrooms (A), happy graduates (B), and cost data (C) all relate to the extended hours in some way, but none of them actually measure what the director claimed improved. Always ask: does this evidence directly measure the outcome being claimed?
Question 3
A property-management firm claims that digital notices are more effective than mailed notices for reaching tenants. In three buildings, the firm emailed a maintenance notice to tenants who had provided email addresses and sent certified letters to the remaining tenants. Its software recorded that 78 percent of the emails were opened, while postal records showed delivery receipts for 61 percent of the certified letters. The firm did not determine whether an opened email was read, whether another household member saw a letter, or how many tenants had not provided an email address because they lacked dependable internet access.
What is the strongest reason the evidence does not fully support the firm's claim?
- The two percentages measure different actions and come from tenant groups that may differ in ways related to receiving digital communication. (correct answer)
- Email open-tracking software cannot record whether a message was read, so the 78 percent figure is too unreliable to use in any comparison with postal records.
- A valid comparison requires that both methods achieve delivery to every tenant in the buildings, and the postal percentage falls well short of that standard.
- The study would need to report the cost per contact for email and certified mail before any conclusion about effectiveness in reaching tenants can be drawn.
Explanation: When evaluating a claim that one method outperforms another, you need to ask: are these two things actually being compared fairly? A valid comparison requires that both groups being measured are similar and that both measurements capture the same thing. If either condition fails, the comparison breaks down.
That's exactly the flaw here. The firm compared email open rates among tenants who chose to provide email addresses against delivery rates for tenants who did not. The passage even hints at why those groups differ — some tenants lacked dependable internet access, which is precisely why they didn't submit an email address. So the "email group" was already pre-selected for internet reliability. Comparing their open rate to the postal group's delivery rate is like comparing apples to oranges: different actions (opening vs. receiving), different populations (digitally connected vs. not). Answer A correctly identifies both of these problems together, making it the strongest and most complete objection.
Answer B is too absolute. Saying the open-rate figure is too unreliable to use in any comparison overstates the problem. Imperfect data can still be useful — the real issue is what's being compared, not the metric alone.
Answer C sets up an impossible standard. No delivery method reaches 100% of tenants, so demanding that threshold would disqualify virtually every real-world study.
Answer D introduces cost as a factor, but the firm's claim is about reach, not cost-efficiency. Cost data is simply irrelevant here.
When you see comparison claims, always ask: are the two groups equivalent, and are the two measurements capturing the same thing? If not, the comparison is flawed at its foundation.
Question 4
A developer's transportation study predicts that a proposed apartment complex will not increase congestion at a nearby intersection. The developer funded the study, but the report publishes its traffic counts, formulas, and modeling assumptions. A neighborhood group challenges the prediction using residents' descriptions of existing backups. Later, an independent engineer checks the published traffic counts and finds them accurate but does not evaluate the model's assumptions about how many future residents will use public transit or travel outside peak hours.
Which evaluation gives the most appropriate weight to the available evidence?
- The prediction should be rejected because evidence funded by an interested party is automatically less reliable than residents' firsthand observations.
- The prediction should be accepted because the independent engineer confirmed one part of the study and therefore validated the entire traffic model.
- The residents' descriptions settle the issue because present-day backups prove that any additional housing will increase future congestion.
- The verified counts strengthen the study's starting data, but its prediction remains uncertain until the key assumptions about future travel are examined. (correct answer)
Explanation: When evaluating a prediction that relies on both data and assumptions, you need to ask two separate questions: What has actually been verified? and What remains unexamined? Treating a partial confirmation as a full one — or dismissing everything because of one flaw — are the two most common traps in evidence-evaluation questions.
Here, the independent engineer confirmed that the traffic counts are accurate. That matters — it means the study's raw input data is trustworthy. However, the study's prediction about future congestion also depends on assumptions about how many residents will use public transit and whether they'll travel outside peak hours. Those assumptions were never independently checked. So the honest conclusion is that the foundation is solid, but the structure built on it is still unverified. That's exactly what D captures: verified counts strengthen the starting data, but the prediction stays uncertain until the key assumptions are examined.
A is wrong because it dismisses the study entirely based on who funded it. Funding source can raise a red flag worth investigating, but the study published its methodology — which is precisely why independent review became possible. Bias isn't automatic just because there's a financial interest. B commits the opposite error, treating confirmation of one component as validation of the whole model — a logical leap the passage doesn't support. C mistakes current conditions for future proof. Existing backups show present congestion, but they don't tell you how a new complex's residents will actually behave.
Your strategy: when a question involves partial verification, always ask what was checked and what wasn't — both matter equally for judging a prediction's reliability.
Question 5
A city report credits its free shade-tree program with lowering household summer electricity bills. In neighborhoods that received the most trees, average bills fell by 9 percent over two years; elsewhere, they fell by 2 percent. Trees were distributed first to neighborhoods identified as having the least existing shade. During the same period, the electric utility offered substantial insulation rebates in several of those same neighborhoods. The report does not state how many households used the rebates or whether the tree recipients' homes differed in size, occupancy, or air-conditioning equipment.
Which statement best identifies how the evidence should be evaluated?
- The larger decline proves the trees reduced electricity use because neighborhoods with fewer trees had a smaller decline.
- The evidence is irrelevant because electricity bills cannot provide any information about household energy consumption.
- The difference is consistent with a benefit from trees, but rebate use and household differences could also explain part of the decline. (correct answer)
- The trees probably raised electricity use because neighborhoods selected for the program initially had less shade than other neighborhoods.
Explanation: Whenever a question asks you to evaluate evidence rather than simply accept a conclusion, your job is to think like a skeptic: Does the data actually prove what the report claims, or could something else explain the results?
Here, the city report points to a 9% bill reduction in tree-recipient neighborhoods versus 2% elsewhere and concludes the trees deserve credit. That comparison is suggestive, but a careful reader notices two red flags buried in the passage: (1) insulation rebates were offered in many of those same neighborhoods during the same period, and (2) the homes may have differed in size, equipment, or occupancy. Either factor could independently reduce electricity bills, meaning the trees may deserve only partial credit — or possibly less. Answer C captures this balance perfectly: the difference is consistent with a tree benefit, but the uncontrolled variables mean you can't isolate the cause. That's honest, accurate evaluation.
A commits a classic logical error called "correlation implies causation." Just because trees were present and bills fell more doesn't prove the trees caused the decline — especially when competing explanations exist right in the passage. B goes to the opposite extreme, claiming bills tell us nothing about energy use, which is false — bills directly reflect consumption costs and are a reasonable proxy. D misreads the selection process entirely. The fact that those neighborhoods had less initial shade is the reason they were chosen for the program — it doesn't suggest trees raised electricity use.
Your strategy tip: when a passage mentions a program and other simultaneous interventions, always ask yourself, "Could those other factors explain the result?" That's the heart of evaluating causation on this exam.
Question 6
A regional insurance company tested flexible start times in one claims-processing unit for three months. According to a management report, the number of employees recorded as late fell from 14 in the previous three-month period to 6 during the pilot. The report concludes, "Flexible scheduling cut lateness by more than half." However, two other changes occurred when the pilot began: road construction near the office ended, and the new unit supervisor recorded an employee as late only after a ten-minute delay rather than the five-minute delay used previously.
Which evaluation of the report's conclusion is most accurate?
- The conclusion is well supported because the number recorded as late fell by more than half after flexible scheduling began.
- The conclusion is unsupported because workplace records can never provide reliable evidence about changes in employee attendance.
- The records show a decline in recorded lateness, but the timing and measurement changes prevent attributing that decline solely to flexible scheduling. (correct answer)
- The records show flexible scheduling was ineffective because the company did not eliminate lateness among every employee in the unit.
Explanation: When a report claims that one specific change caused a result, your job is to ask: could something else explain this? This is the heart of evaluating evidence — identifying confounding factors, meaning other variables that changed at the same time and could have produced the same outcome.
Here, the report credits flexible scheduling for cutting recorded lateness in half. But two other changes happened simultaneously: road construction near the office ended (so commutes got easier for everyone), and the new supervisor redefined "late" by adding five extra minutes of grace. Either of these changes alone could explain the drop in recorded lateness — regardless of scheduling flexibility. That's why C is the most accurate evaluation. It acknowledges the real decline in the data while correctly identifying that the timing and measurement changes make it impossible to isolate flexible scheduling as the cause.
Answer A fails because it confuses correlation with causation. Yes, lateness fell after flexible scheduling began, but "after" doesn't mean "because of." A treats the raw numbers as proof without accounting for what else changed.
Answer B goes too far in the opposite direction. Workplace records can be useful evidence — the problem here isn't that records are inherently unreliable, it's that this specific situation has uncontrolled variables. Dismissing all workplace records is an overgeneralization.
Answer D applies an unreasonable standard. No policy is expected to achieve a 100% elimination of a problem. Judging effectiveness by a perfect outcome is a logical trap.
Strategy tip: Whenever a passage credits one cause for a change, scan for other factors that shifted at the same time — those are your confounding variables, and questions about them are extremely common on critical-reasoning exams.
Question 7
A human-resources memo argues that low morale is the principal cause of employee absence. It compares eight departments and reports that departments with lower average scores on an annual engagement survey also had more absence days per employee. The memo does not examine whether the employees who reported low engagement were the same employees who were frequently absent. It also does not account for differences in physical demands, shift schedules, exposure to illness, or access to paid sick leave across departments.
What does the evidence most reasonably support?
- Low morale must cause absence because the two measures move in opposite directions across all eight departments.
- Department-level morale and absence are associated, but the evidence does not show that low morale is the principal cause of individual absences. (correct answer)
- Absence must cause low morale rather than the reverse because absent employees have fewer opportunities to engage with coworkers.
- The engagement survey is useless because employees' opinions cannot be compared with numerical attendance records.
Explanation: When a passage presents data showing two things moving together — like morale scores and absence rates — your job is to ask: does this evidence actually prove a causal relationship, or does it only show a pattern? This question tests your ability to distinguish correlation from causation and recognize what evidence can and cannot support.
The passage reports a department-level pattern: lower engagement scores tend to accompany higher absence rates. That's a real finding, and B correctly names it — an association between morale and absence at the department level. But the memo has two critical gaps. First, it never confirms that the same individuals with low engagement are the ones going absent. Second, it ignores confounding factors like physical job demands, shift schedules, and sick-leave access, any of which could independently drive absence rates. Because these gaps exist, calling low morale the principal cause goes far beyond what the data supports. B is the only answer that accurately reflects both what the evidence shows and what it fails to prove.
Choice A is wrong because two measures moving together across departments establishes correlation, not causation — "must cause" is far too strong a conclusion. Choice C makes the opposite causal claim (absence causes low morale), which is equally unsupported and introduces a speculative mechanism the passage never tests. Choice D is wrong because comparing survey scores with attendance records is a legitimate and common research method; a limitation in interpretation doesn't make the survey "useless."
Strategy tip: When evidence is correlational, correct answers almost always include hedging language like "associated with" or "suggests." Watch for answer choices using absolute words like "must" or "proves" — those are usually traps.
Question 8
A company's annual report states, "Our mentorship program causes new employees to remain with the company longer." Of the new employees assigned mentors last year, 88 percent were still employed after twelve months, compared with 70 percent of employees without mentors. Supervisors selected as mentees those employees they believed showed strong potential, and participation was optional. No adjustments were made for job type, prior experience, starting salary, or supervisor ratings.
Which revision of the annual report's statement is most justified by the evidence?
- The mentorship program retains exactly 18 percent of new employees who would otherwise leave within their first year.
- The mentorship program causes higher retention mainly among employees whom supervisors identify as having strong potential.
- Participation in mentorship was associated with higher one-year retention, although selection differences may account for some of the gap. (correct answer)
- Mentorship has no relationship to retention because employees were not randomly required to participate in the company program.
Explanation: When a passage makes a causal claim ("causes retention"), your job is to ask: does the evidence actually support that cause, or could something else explain the pattern? This question tests your ability to distinguish correlation from causation and identify confounding variables.
The passage contains two major red flags: mentees were selected by supervisors for showing strong potential, and participation was optional. This means the mentored and unmentored groups were already different before the program began — higher-potential employees who chose to participate may have stayed longer for reasons unrelated to mentorship itself. These are called confounding variables, and they undermine any clean causal conclusion. Answer C captures this perfectly: it acknowledges the real finding (mentored employees had higher one-year retention) while honestly noting that "selection differences may account for some of the gap." That's a justified, appropriately cautious revision of the original overclaim.
Answer A is wrong because the 18-percentage-point gap (88% − 70%) doesn't tell you how many employees who would have left were retained because of the program — that calculation requires controlled conditions the study didn't have. Answer B is tempting but still makes a causal claim ("causes higher retention"), just a narrower one. The evidence doesn't justify any causal language because the groups weren't comparable to begin with. Answer D goes too far in the other direction — the lack of randomization doesn't mean there's no relationship, only that the relationship may not be causal.
Study tip: When evaluating conclusions from data, watch for selection bias and missing controls. If groups differ before an intervention begins, correlation language ("associated with") is safer than causal language ("causes").
Question 9
A manufacturer advertises a new exterior sealant as providing "ten years of protection in every climate." An independent laboratory applied the sealant to 12 identical wooden panels and exposed them for six weeks to repeated cycles of intense ultraviolet light and sprayed salt water. None of the treated panels cracked or peeled, while several untreated panels did. The laboratory report warns that its accelerated test did not reproduce freezing temperatures, variations in application technique, or long-term expansion and contraction of different building materials.
Which conclusion is best supported by the laboratory evidence?
- The sealant resisted damage under the tested conditions, but the study does not establish ten-year performance across all climates and uses. (correct answer)
- The sealant will probably last ten years in coastal climates because salt water and ultraviolet light are the main causes of exterior damage.
- The sealant provides no dependable protection because the laboratory did not test every climate or every type of building material.
- The sealant is superior to competing commercial sealants because the untreated panels showed more damage during the laboratory test.
Explanation: When a question asks you to draw a conclusion from evidence that includes both findings and stated limitations, your job is to match the conclusion precisely to what the data actually shows — no more, no less.
The laboratory found that treated panels survived UV light and salt water exposure without cracking or peeling. That's a real, meaningful result. However, the researchers themselves flagged critical gaps: no freezing temperatures, no variation in application technique, and no testing across different building materials. A sound conclusion honors both the positive finding and those acknowledged limits. That's exactly what A does — it confirms the sealant resisted damage under the tested conditions while noting the study cannot establish ten-year performance across all climates and uses. This makes A the best supported conclusion.
B oversteps the evidence by claiming the sealant "will probably last ten years in coastal climates." The study ran for six weeks under artificial conditions, not ten years in real coastal environments. Assuming UV light and salt water are the main causes of all exterior damage is also an unsupported leap.
C goes too far in the opposite direction. Saying the sealant offers "no dependable protection" ignores the clear positive result the study did produce. Incomplete evidence doesn't mean zero evidence.
D introduces a comparison that was never made. The untreated panels were controls, not competing commercial sealants, so you cannot use this study to claim superiority over rivals.
Your strategy: watch for conclusions that either overreach the evidence (B, D) or dismiss it entirely (C). The best conclusion stays within the boundaries the data and its limitations define.
Question 10
Six months after launching a curbside composting program, a county announced that the program had reduced the amount of waste sent to its landfill by 18 percent. Scale records confirm that total landfill deliveries declined by that amount compared with the same six months of the previous year. During the period, however, a large food-processing plant closed, and the county's population decreased after a major employer relocated. The county has records of how much material the composting trucks collected, but its announcement does not cite those records or separate household food waste from commercial and industrial waste.
Which evidence would be most useful for determining how much of the landfill decline was attributable to the composting program?
- Residents' opinions about whether curbside composting is more convenient than taking food waste to a community collection site
- The amount of household organic waste collected by the program, together with changes in other waste sources and the number of households served (correct answer)
- The total cost of purchasing composting trucks, containers, and educational materials during the program's first six months
- A comparison of the county's current landfill total with its highest annual landfill total from the previous decade
Explanation: When a question asks you to identify the most useful evidence for determining causation, your job is to think like an investigator: what specific data would let you isolate one cause from several competing explanations? Here, the landfill declined 18%, but three things changed simultaneously — the composting program launched, a food-processing plant closed, and the population dropped. You need evidence that separates the program's contribution from those other factors.
Answer B does exactly that. Knowing how much organic waste the composting trucks actually collected tells you the program's direct output. Pairing that with changes in other waste sources (like the closed plant) and the number of households served lets you subtract out unrelated causes and calculate what portion of the 18% decline the program genuinely produced. This is the targeted, multi-variable evidence an investigator would need.
Answer A is a trap — residents' opinions about convenience measure satisfaction, not actual waste diversion. How people feel about the program tells you nothing about how much landfill reduction it caused. Answer C describes startup costs, which are relevant to budgeting but irrelevant to measuring environmental impact. Knowing what the trucks cost doesn't tell you what they collected. Answer D compares the current landfill total to a historical peak, which is far too broad — it spans years of unrelated changes and can't isolate the composting program's six-month effect.
The key strategy: when a question involves multiple possible causes for one outcome, the best evidence is always the data that lets you isolate and measure each cause separately — not general opinions, costs, or broad historical comparisons.