All questions
Question 1
A library tested two reminder methods for returning books on time. They randomly assigned 50 patrons to Treatment A (text reminder) and 50 patrons to Treatment B (email reminder). The outcome was whether the book was returned on time. Treatment A had 41 on-time returns (0.82) and Treatment B had 33 on-time returns (0.66). The observed difference in proportions was 0.82−0.66=0.16 (A − B). A randomization test was performed by shuffling the on-time/late outcomes among all 100 patrons 4000 times under “no treatment effect.” In the 4000 shuffles, 12 produced a difference in proportions of at least +0.16 (A − B).
Which conclusion is most reasonable about the treatment effect (A − B)?
- The difference must be due to patrons in the text group being more responsible, since confounding always explains group differences.
- Because 12 out of 4000 shuffles were at least +0.16, the observed difference is common under no effect; there is not enough evidence of a difference.
- Text reminders are guaranteed to make exactly 16% more patrons return books on time in any library.
- Because only 12 out of 4000 shuffles were at least +0.16, the observed difference would be unusual under no effect; there is evidence text reminders increase the on-time return proportion compared with email reminders. (correct answer)
Explanation: This question covers comparing reminder methods in a randomized experiment and simulation to see if texts improve on-time book returns over emails. Random assignment makes groups similar, allowing attribution of differences to treatments for cause-and-effect inferences. The observed difference is in return proportions, 0.16 higher for Treatment A than B. The 'no effect' simulation shuffles outcomes across patrons multiple times, generating a distribution of possible differences by chance alone under no effect. Only 12 out of 4000 shuffles were at least +0.16, meaning the observed is unusual and provides evidence for Treatment A's effectiveness. Remember, random assignment supports causality in the study unlike random sampling for generalization, and rarity offers evidence without absolute proof. To use this approach, evaluate the count of simulations as or more extreme than observed.
Question 2
In a randomized experiment, 72 volunteers were randomly assigned to two different puzzle-solving strategies: Treatment A (work backward) and Treatment B (trial-and-error), 36 per group. The outcome was the number of hints used (lower is better). Treatment A had a mean of 2.1 hints and Treatment B had a mean of 2.9 hints, so the observed difference in means was 2.1−2.9=−0.8 (A − B). A randomization test shuffled the treatment labels 2500 times under “no treatment effect.” In the 2500 shuffles, 5 produced a difference in means of at most −0.8 (A − B).
Which conclusion is most reasonable about the treatment effect (A − B), where more negative values indicate fewer hints for Treatment A?
- The observed difference must be due to different volunteer backgrounds, so random assignment cannot support a causal conclusion.
- Because only 5 out of 2500 shuffles were at most −0.8, the observed difference would be unusual under no effect; there is evidence Treatment A reduces the mean number of hints used compared with Treatment B. (correct answer)
- Treatment A is proven to reduce hints for every person because the mean difference is −0.8.
- Because 5 out of 2500 shuffles were at most −0.8, the observed difference is common under no effect; there is not enough evidence of a treatment effect.
Explanation: This question examines comparing puzzle strategies in a randomized experiment and simulation to check if working backward reduces hints needed. Random assignment minimizes differences between groups, supporting cause-and-effect conclusions from treatments. The observed difference is in mean hints, -0.8 for Treatment A minus B, showing fewer for A. The 'no effect' simulation shuffles labels on fixed hint counts many times, building a distribution of chance differences assuming no treatment effect. Only 5 out of 2500 shuffles were at most -0.8, indicating the observed is unusual and evidencing Treatment A's benefit. A misconception is equating random assignment to random sampling; it facilitates causality in this context, with rarity providing evidence, not certainty. For transfer, focus on how often simulations are as extreme or more than observed.
Question 3
In a randomized experiment, 90 participants were randomly assigned (45 each) to two different instruction videos: Treatment A (interactive) and Treatment B (standard). After viewing, each participant answered 10 comprehension questions. Treatment A had a mean score of 7.1 and Treatment B had a mean score of 6.2, so the observed difference in means was 7.1−6.2=0.9 (A − B). A randomization test shuffled the A/B labels 1000 times under “no treatment effect.” In the 1000 shuffles, 8 produced a difference in means of at least +0.9 (A − B).
Based on the randomization test, is the observed difference surprising under no effect?
- No. The randomization test is invalid because it uses shuffling instead of collecting new samples.
- Yes. Any observed difference must be caused by the interactive video because random assignment eliminates all variability.
- No. Since 8 out of 1000 shuffles were at least +0.9, the observed difference is not surprising under no effect, so there is no evidence of an effect.
- Yes. Since only 8 out of 1000 shuffles were at least +0.9, the observed difference would be unusual under no effect; there is evidence the interactive video increases the mean score. (correct answer)
Explanation: This question deals with comparing instructional videos in a randomized experiment and simulation to check if an interactive one boosts comprehension scores. Random assignment equalizes groups on other factors, permitting cause-and-effect conclusions from treatment differences. The observed difference is in mean scores, 0.9 points higher for Treatment A than B. The 'no effect' simulation shuffles labels on fixed scores many times, forming a distribution of chance-based differences assuming no treatment impact. With just 8 out of 1000 shuffles at least +0.9, the observed difference is surprising under no effect, indicating evidence for Treatment A's benefit. A key misconception is confusing random assignment with random sampling; the first aids causality in the experiment, and rarity supports but doesn't prove the effect universally. Apply this by seeing the proportion of extreme simulated differences to determine rarity.
Question 4
A student organization compared two ways to encourage event attendance. They randomly assigned 40 members to Treatment A (personalized message) and 40 members to Treatment B (generic message). The outcome was whether the member attended the event. Treatment A had 22 attendees (0.55) and Treatment B had 21 attendees (0.525). The observed difference in proportions was 0.55−0.525=0.025 (A − B). A randomization test shuffled the attendance outcomes among the 80 members 3000 times under “no treatment effect.” In the 3000 shuffles, 1459 produced a difference in proportions of at least +0.025 (A − B).
Based on the randomization test, is the observed difference surprising under no effect?
- Yes. Because Treatment A’s proportion (0.55) is larger, it guarantees the personalized message causes higher attendance.
- Yes. Since only 1459 out of 3000 shuffles were at least +0.025, the result is very rare under no effect and suggests a real advantage for Treatment A.
- No. Since 1459 out of 3000 shuffles were at least +0.025, a difference like +0.025 is fairly common under no effect; there is not strong evidence that the personalized message increases attendance. (correct answer)
- No. The study cannot be used because random assignment only matters if the members were randomly sampled from the whole campus.
Explanation: This question involves comparing messaging strategies in a randomized experiment with simulation to see if personalized messages boost event attendance. Random assignment ensures comparability, allowing cause-and-effect attributions to treatments. The observed difference is in attendance proportions, 0.025 higher for Treatment A than B. The 'no effect' simulation shuffles outcomes among members repeatedly, creating a distribution of differences from random chance if treatments didn't differ. Since 1459 out of 3000 shuffles were at least +0.025, such a difference is common under no effect, suggesting it's not surprising and lacks strong evidence. Note that random assignment supports causality here but differs from random sampling for broader applicability, and rarity indicates evidence without proof. To apply, count the proportion of simulations equaling or surpassing the observed extremity.
Question 5
A coach compared two warm-up routines. Using a random number generator, 34 athletes were randomly assigned to Treatment A (dynamic warm-up) and 34 to Treatment B (light jogging). Each athlete then attempted a target drill; success was “hit the target at least 7 times out of 10.” In Treatment A, 20 of 34 succeeded (0.588). In Treatment B, 19 of 34 succeeded (0.559). The observed difference in proportions was 0.588−0.559=0.029 (A − B). A randomization test shuffled the success/failure outcomes among athletes 2000 times under “no treatment effect.” In the 2000 shuffles, 836 produced a difference in proportions of at least +0.029 (A − B).
Which conclusion is most reasonable about whether Treatment A is more effective than Treatment B?
- Because 836 out of 2000 shuffles were at least +0.029, a difference this large is fairly common under no effect; there is not strong evidence that Treatment A is more effective. (correct answer)
- No conclusion can be drawn because the athletes were not randomly assigned; the coach chose who did which warm-up.
- Treatment A is proven to be better because its success proportion (0.588) is higher than Treatment B’s (0.559).
- Because 836 out of 2000 shuffles were at least +0.029, the result is extremely rare under no effect; there is strong evidence Treatment A is more effective.
Explanation: This question investigates comparing warm-up routines in a randomized experiment with simulation to assess effectiveness in a target drill. Random assignment helps isolate treatment effects by balancing groups, supporting cause-and-effect claims. The observed difference is in success proportions, 0.029 higher for Treatment A than B. The 'no effect' simulation shuffles outcomes among athletes repeatedly, creating a distribution of differences from random chance if treatments were identical. Since 836 out of 2000 shuffles were at least +0.029, such a difference is common under no effect, suggesting insufficient evidence for Treatment A's superiority. Misconceptions include thinking random assignment equals random sampling—it enables causality here, and rarity evidences but doesn't prove effects. For other cases, look at how frequently simulations match or exceed the observed extremity.
Question 6
In a randomized experiment, 60 students were randomly assigned by shuffling identical cards (30 labeled A and 30 labeled B) to either Treatment A (a new study app) or Treatment B (a standard app). After one week, each student took the same 20-question quiz. The mean score for Treatment A was 16.4 and for Treatment B was 14.9, so the observed difference in means was 16.4−14.9=1.5 points (A − B). To test whether this difference could be due to chance, the researcher performed a randomization test by keeping all 60 quiz scores fixed and repeatedly shuffling the A/B labels 2000 times under “no treatment effect.” In the 2000 shuffles, 18 produced a difference in means of at least +1.5 points (A − B).
Which conclusion is most reasonable about the treatment effect (A − B)?
- Because the students were randomly assigned, Treatment A is proven to increase quiz scores by exactly 1.5 points.
- The result cannot be used to compare treatments because the students were not randomly sampled from all students.
- Because only 18 out of 2000 shuffles produced a difference at least as large as +1.5, the observed difference would be unusual under no effect, so there is evidence Treatment A increases the mean quiz score compared with Treatment B. (correct answer)
- Because 18 out of 2000 is not zero, the observed difference is common under no effect, so there is not enough evidence of a treatment effect.
Explanation: This question explores comparing treatments using a randomized experiment and simulation to assess if a new study app improves quiz scores. Random assignment of students to treatments helps ensure that any observed differences are likely due to the treatments themselves, allowing for cause-and-effect conclusions by balancing out other factors across groups. The observed difference is the difference in mean quiz scores, here 1.5 points higher for Treatment A than B. The 'no effect' simulation shuffles the treatment labels many times while keeping scores fixed, creating a distribution of possible differences that could occur just by chance if the treatments had no real impact. By comparing the observed +1.5 to this distribution, we see that only 18 out of 2000 shuffles were at least as extreme, indicating the result is rare under no effect and provides evidence of a treatment effect. A common misconception is that random assignment is the same as random sampling from a population, but it actually supports causal claims within the study group, and rarity offers evidence but not absolute proof. To apply this strategy elsewhere, check how often simulated differences are as extreme or more than the observed one to gauge surprise under no effect.
Question 7
A school randomly assigned 60 students to use either Treatment A (a spaced-practice study plan) or Treatment B (a cramming study plan) for one week before a quiz. Treatment A had n=30 with a mean quiz score of 82 points; Treatment B had n=30 with a mean quiz score of 74 points. The observed difference in means was 82−74=8 points (A − B). To test whether this difference could be due to chance under “no treatment effect,” the teacher performed a randomization test by shuffling the treatment labels 1,000 times and recalculating (A − B) each time. In the simulation, 9 of the 1,000 shuffled differences were at least as large as 8 points.
Which conclusion is most reasonable about the treatment effect?
- Because only 9 of 1,000 shuffled differences were 8, the observed difference would be rare if there were no treatment effect, so the results provide evidence that Treatment A increases mean quiz score compared with Treatment B. (correct answer)
- Because 9 of 1,000 is small, the observed difference is common under no effect, so there is not enough evidence that Treatment A is more effective.
- The randomization test is not relevant because simulations are not based on real data, so no conclusion about treatment effectiveness can be made.
- Because the students were randomly assigned, Treatment A is guaranteed to be better than Treatment B by exactly 8 points.
Explanation: In comparing treatments using a randomized experiment and simulation, we investigate if one treatment performs better than another by assigning participants randomly to groups. Random assignment helps balance out other factors, allowing us to attribute differences in outcomes to the treatments themselves rather than chance or biases, supporting cause-and-effect conclusions. The observed difference here is the difference in mean quiz scores between Treatment A and B, which was 8 points higher for A. The 'no effect' simulation shuffles the treatment labels many times, assuming treatments make no difference, and creates a distribution of possible differences that could occur just by chance. By comparing the observed 8-point difference to this simulated distribution, we see that only 9 out of 1,000 shuffled differences were as large as or larger than 8, indicating the observed result is rare under no effect. A common misconception is that random assignment is the same as random sampling from a population, but it actually helps with causal inference within the study group, and rarity provides evidence but not absolute proof of an effect. To apply this strategy elsewhere, count how often simulated differences are at least as extreme as the observed one to assess if the result is surprising.
Question 8
A game designer randomly assigned 90 players to try either Treatment A (a new tutorial) or Treatment B (the old tutorial), 45 players per group. The outcome was whether the player completed the first level without hints.
Results: Treatment A had 33/45 successes (73.3%), Treatment B had 21/45 successes (46.7%). Observed difference in proportions (A − B) = 0.733−0.467=0.266.
A randomization test shuffled the success/failure outcomes across groups 5,000 times under “no treatment effect.” In the simulation, 18 of the 5,000 shuffled differences were at least as large as 0.266.
Based on the randomization test, is the observed difference surprising under no effect?
- No. Since 18 of 5,000 is not zero, the observed difference is expected under no effect and does not suggest a treatment effect.
- Yes. Only 18 of 5,000 shuffled differences were 0.266, so the observed difference would be rare under no effect, suggesting the new tutorial increases the completion rate. (correct answer)
- Yes. Random assignment means confounders must explain the difference, so the simulation is unnecessary.
- No. The randomization test should count shuffled differences with absolute value 0.266 even though the question is about A being higher, so the given count cannot be used.
Explanation: To compare treatments, randomized experiments use simulation to test if differences are real or chance-based. Random assignment balances groups, allowing cause-and-effect conclusions by minimizing biases. The observed difference here is 0.266 higher proportion of successes for Treatment A versus B. Under 'no effect,' the simulation shuffles outcomes many times, creating a distribution of chance differences. Only 18 out of 5,000 simulations had differences as large as or larger than 0.266, marking the observed as rare without an effect. A misconception is equating random assignment with random sampling; assignment enables causality, and rarity evidences but doesn't prove the effect. Apply this by counting simulated differences at least as extreme as observed in other scenarios.
Question 9
A community center randomly assigned 100 adults to two reminder systems for attending a weekly class: Treatment A (text reminder) and Treatment B (email reminder), 50 adults per group. The outcome was whether the adult attended the class that week.
Results: Treatment A attendance = 31/50, Treatment B attendance = 29/50. Observed difference in proportions (A − B) = 0.62−0.58=0.04.
A randomization test shuffled the attendance outcomes across groups 10,000 times under “no treatment effect.” In the simulation, 6,820 of the 10,000 shuffled differences were at least as large as 0.04.
Which conclusion is most reasonable about the treatment effect?
- There is evidence that text reminders increase attendance, because 6,820 of 10,000 is a small number of simulations.
- There is not enough evidence that text reminders increase attendance, because differences of 0.04 or larger occurred frequently (6,820/10,000) under no effect. (correct answer)
- No conclusion is possible because random assignment does not balance confounders, so any difference must be due to confounding.
- There is evidence that email reminders increase attendance, because the observed difference is positive (A − B = 0.04).
Explanation: Comparing treatments via randomized experiments and simulations involves testing if differences exceed chance variation. Random assignment balances factors, enabling cause-and-effect claims. The observed difference is 0.04 higher attendance proportion for Treatment A. In 'no effect' simulations, outcomes are shuffled to create a distribution of random differences. 6,820 out of 10,000 simulations had differences as large as or larger than 0.04, showing it's frequent by chance. Remember, random assignment differs from sampling; it helps causality, and rarity evidences but doesn't prove. In applications, count simulations with differences at least as extreme.
Question 10
A fitness coach randomly assigned 50 clients to Treatment A (interval training) and 50 clients to Treatment B (steady-state training) for 4 weeks. The outcome was the change in resting heart rate (in beats per minute, bpm), where a more negative value indicates a larger improvement.
Summaries: Treatment A mean change =−6.2 bpm, Treatment B mean change =−3.1 bpm. The observed difference in means (A − B) was −6.2−(−3.1)=−3.1 bpm.
To test for a treatment effect, the coach performed a randomization test by shuffling the treatment labels 1,000 times and recalculating (A − B). In the simulation under “no effect,” 6 of the 1,000 shuffled differences were at most −3.1 bpm (i.e., as negative as the observed difference).
Do the results provide evidence that Treatment A leads to a larger average decrease in resting heart rate than Treatment B?
- No. Only 6 of 1,000 is too many to be considered unusual, so the difference is likely due to chance.
- Yes. Because the sample size is 100, any difference in means must be statistically significant and caused by the training program.
- No. Since the observed difference is negative, it must mean Treatment B worked better than Treatment A.
- Yes. Only 6 of 1,000 shuffled differences were -3.1, so such an extreme negative difference would be rare under no effect, providing evidence that Treatment A causes a larger average decrease. (correct answer)
Explanation: We use randomized experiments and simulations to compare treatments by checking if differences in outcomes are likely caused by the treatments. Through random assignment, groups start balanced, so observed differences can be causally tied to the treatments rather than other variables. The observed difference in this case is -3.1 bpm in mean heart rate changes, meaning Treatment A showed a larger decrease. In the 'no effect' simulation, we shuffle labels repeatedly to mimic what differences might occur purely by chance, forming a distribution of those chance differences. Only 6 out of 1,000 simulations produced differences as negative as or more than -3.1, showing the observed result is rare if there's no real effect. It's a misconception that random assignment equals random sampling; assignment aids causal claims, and rarity suggests evidence of an effect without proving it universally. In similar analyses, focus on how frequently the simulated differences match or exceed the observed extremity.
Question 11
A cafeteria randomly assigned 40 students to receive either Treatment A (fruit placed at eye level) or Treatment B (fruit placed on a lower shelf), 20 students per group. The outcome was whether the student chose fruit.
Results: Treatment A fruit choice = 14/20, Treatment B fruit choice = 6/20. Observed difference in proportions (A − B) = 0.70−0.30=0.40.
A randomization test shuffled the 40 yes/no outcomes across the two groups 1,000 times under “no treatment effect,” recalculating (A − B) each time. In the simulation, 3 of the 1,000 shuffled differences were at least as large as 0.40.
Which conclusion is most reasonable about the treatment effect?
- The difference must be due to pre-existing differences between the groups, because random assignment cannot balance confounders.
- There is not enough evidence of an effect because 3 of 1,000 indicates the observed difference happens frequently by chance.
- The results prove that every student would choose fruit if it is placed at eye level, since the study used random assignment.
- There is evidence that placing fruit at eye level increases the probability of choosing fruit, because only 3 of 1,000 shuffled differences were 0.40, making the observed difference rare under no effect. (correct answer)
Explanation: To compare treatments, randomized experiments with simulations assess if differences are surprising under chance. Random assignment balances groups for causal inference. Observed difference: 0.40 higher fruit choice proportion for A. 'No effect' simulation shuffles outcomes for a chance distribution. 3 out of 1,000 simulations >=0.40, showing rarity. Note: assignment ≠ sampling; rarity evidences, not proves. Apply by counting extreme simulations.
Question 12
A company randomly assigned 70 employees to two email-subject-line styles for an internal newsletter: Treatment A (question-style subject lines) and Treatment B (statement-style subject lines), 35 employees per group. The outcome was the time (in seconds) to open the email after it was sent.
Results: Treatment A mean open time = 52 s, Treatment B mean open time = 55 s. Observed difference in means (A − B) = −3 s.
To test whether Treatment A leads to faster opening times, the company conducted a randomization test by shuffling the treatment labels 2,000 times and recomputing (A − B). In the simulation, 1,540 of the 2,000 shuffled differences were at most −3 s (i.e., as negative as observed).
Which conclusion is most reasonable about the treatment effect?
- There is not enough evidence that Treatment A makes open times faster, because differences at least as negative as −3 seconds happened often (1,540/2,000) under no effect. (correct answer)
- There is evidence that Treatment B makes employees open emails faster, because the observed difference is negative.
- There is evidence that Treatment A makes employees open emails faster, because 1,540 of 2,000 is a very small proportion.
- There is not enough evidence because random assignment only works if the employees were randomly sampled from all possible employees.
Explanation: Randomized experiments and simulations help compare treatments by assessing if outcomes differ more than expected by chance. Random assignment ensures fair groups, supporting that treatments cause observed effects. The observed difference is -3 seconds in mean open times, with A faster than B. The 'no effect' simulation shuffles labels repeatedly, producing a distribution of possible differences from randomness alone. 1,540 out of 2,000 simulations showed differences as negative as or more than -3, meaning it's common under no effect. Don't confuse random assignment with random sampling; it aids causal inference, and rarity suggests evidence, not proof. For other cases, look at how often simulations yield differences as extreme as observed.
Question 13
A teacher randomly assigned 60 students to try two study plans for a vocabulary quiz: Treatment A (spaced practice) and Treatment B (single long review). Thirty students were randomly assigned to each treatment. The outcome was the quiz score (out of 100). Treatment A had a mean score of 84.2 and Treatment B had a mean score of 78.5, for an observed difference in means of 84.2−78.5=5.7 points (A − B). To test whether this difference could be due to chance under “no treatment effect,” the teacher performed a randomization test by shuffling the treatment labels 2000 times and recalculating the difference in means each time. In the simulations, 18 of the 2000 shuffled differences were at least as large as 5.7 (A − B).
Which conclusion is most reasonable about the treatment effect?
- Since only 18 out of 2000 shuffled differences were at least as large as 5.7, the observed difference would be rare under no effect, so there is evidence that Treatment A tends to produce higher mean scores than Treatment B. (correct answer)
- Because the teacher did not randomly sample students from all students everywhere, no conclusion can be made about whether Treatment A caused higher scores in this class.
- Because 18 simulated differences were at least as large as 5.7, the observed difference is common under no effect, so there is not enough evidence that Treatment A is better.
- Because the students were randomly assigned, Treatment A is guaranteed to raise scores by exactly 5.7 points for every student.
Explanation: This question tests understanding of randomized experiments and simulation-based inference. Random assignment allows us to attribute differences to the treatment rather than confounding factors. The observed difference of 5.7 points (Treatment A minus Treatment B) represents how much higher the mean score was for the spaced practice group. The randomization test simulates what differences we'd see if there were no treatment effect by shuffling labels 2000 times. Finding only 18 out of 2000 shuffled differences at least as large as 5.7 means the observed difference would be rare (less than 1% chance) under no effect. This rarity provides evidence that Treatment A tends to produce higher scores, though it doesn't guarantee every student benefits by exactly 5.7 points. Random assignment differs from random sampling—we can make causal conclusions about this class even without sampling from all students everywhere.
Question 14
A coach randomly assigned 48 runners to two warm-up routines before a timed 1-mile run: Treatment A (dynamic stretching) and Treatment B (light jogging). There were 24 runners in each group. The outcome was time in seconds (lower is better). Treatment A had a mean time of 412 seconds and Treatment B had a mean time of 420 seconds, giving an observed difference of 412−420=−8 seconds (A − B). To test for a treatment effect, the coach conducted a randomization test by shuffling treatment labels 1000 times and computing (A − B) each time. Under these shuffles, 410 of the 1000 simulated differences were ≤−8 seconds (at least as extreme in the direction of faster times for A).
Which conclusion is most reasonable about the treatment effect?
- Because runners were randomly assigned, any observed difference must be caused by the warm-up routine, so Treatment A definitely makes runners 8 seconds faster on average.
- Because 410 out of 1000 is a small number, the observed difference is rare under no effect, so Treatment A is clearly better.
- Because 410 out of 1000 shuffled differences were ≤−8, the observed result is not rare under no effect, so there is not strong evidence that Treatment A produces faster mean times than Treatment B. (correct answer)
- The randomization test is irrelevant because simulations are not real data, so no conclusion can be made about whether the warm-up matters.
Explanation: This question examines interpreting randomization test results when many simulated differences are as extreme as observed. Random assignment of runners allows us to attribute differences to the warm-up routine rather than runner characteristics. The observed difference of -8 seconds (Treatment A minus B) indicates dynamic stretching led to faster times on average. The randomization test simulates differences under no treatment effect by shuffling labels 1000 times. Finding 410 out of 1000 shuffled differences at least as extreme (≤ -8 seconds) means the observed difference is not rare under no effect—it would happen about 41% of the time by chance alone. This lack of rarity means we don't have strong evidence that Treatment A produces faster times. A key misconception is thinking 410 is a "small number"—we must consider it as a proportion (410/1000 = 41%). Random assignment doesn't guarantee any specific effect size, and simulations help us judge whether observed differences are surprising.
Question 15
A researcher randomly assigned 100 houseplants to two fertilizers: Treatment A and Treatment B (50 plants each). After 4 weeks, the outcome was the increase in height (cm). Treatment A had a mean increase of 6.4 cm and Treatment B had a mean increase of 6.0 cm, so the observed difference in means was 6.4−6.0=0.4 cm (A − B). The researcher performed a randomization test by shuffling the fertilizer labels 5000 times and recalculating the difference in means each time. In the simulations, 2400 of the 5000 shuffled differences were at least as large as 0.4 cm (A − B).
Do the results provide evidence that Treatment A is more effective than Treatment B?
- No. Because many shuffled differences (2400 out of 5000) were at least as large as 0.4 cm, the observed difference is not rare under no effect, so there is not strong evidence that Treatment A increases mean growth compared with Treatment B. (correct answer)
- No. The correct effect measure is the ratio 6.4/6.0, so the difference in means cannot be used to judge effectiveness.
- Yes. Since 2400 is a large number, that means the observed difference is extremely rare under no effect, so Treatment A is more effective.
- Yes. Random assignment makes the two groups identical, so any observed difference must be due to the fertilizer.
Explanation: This question tests recognizing when many simulated differences indicate lack of evidence. Random assignment of plants to fertilizers allows us to attribute differences to the treatment rather than other factors. The observed difference of 0.4 cm (6.4 minus 6.0) shows Treatment A had slightly more growth on average. The randomization test simulates differences under no treatment effect by shuffling labels 5000 times. Finding 2400 out of 5000 shuffled differences at least as large as 0.4 means the observed difference is common—it would occur about 48% of the time by chance alone. This commonness under no effect means we lack strong evidence that Treatment A increases growth. A misconception is thinking 2400 being "large" makes the difference rare; we must consider the proportion (2400/5000 ≈ 48%). Random assignment doesn't make every difference meaningful, and using ratios instead of differences doesn't change the conclusion about treatment effectiveness.
Question 16
A robotics club randomly assigned 50 students to learn programming with two methods: Treatment A (video tutorials) and Treatment B (written tutorials). Twenty-five students were assigned to each method. After one week, students completed a 20-question quiz. Treatment A had a mean of 15.1 correct and Treatment B had a mean of 14.8 correct, so the observed difference in means was 15.1−14.8=0.3 (A − B). To check if this could happen by chance under no treatment effect, the club shuffled the A/B labels 2000 times. In the shuffled results, 980 of the 2000 differences were at least as large as 0.3 (A − B).
Based on the randomization test, is the observed difference surprising under no effect?
- No. Differences at least as large as 0.3 occurred 980 out of 2000 times under shuffling, so the observed difference is not rare under no effect and does not provide strong evidence of a real advantage for Treatment A. (correct answer)
- Yes. Because students were randomly assigned, any nonzero observed difference must be caused by the treatment method.
- No. The correct comparison is to count shuffled differences ≤−0.3 because that is the only direction that matters when checking if A is better than B.
- Yes. Since 980 is close to 2000, that means 0.3 is extremely rare under no effect, so video tutorials clearly work better.
Explanation: This question tests recognizing when randomization results indicate no evidence of treatment effect. Random assignment allows us to attribute any real differences to the teaching method rather than student characteristics. The observed difference of 0.3 correct answers (15.1 minus 14.8) shows video tutorials had a slightly higher mean score. The randomization test simulates differences under no treatment effect by shuffling labels 2000 times. Finding 980 out of 2000 shuffled differences at least as large as 0.3 means the observed difference is not rare—it would occur about 49% of the time by chance alone. This commonness under no effect means we lack evidence that video tutorials work better. A misconception is thinking 980 being "close to 2000" makes it rare; actually, 980/2000 ≈ 50% indicates the difference is completely typical under no effect. Random assignment doesn't make every nonzero difference meaningful—we must check if it's surprising compared to chance variation.
Question 17
A community center randomly assigned 72 participants to two ways of learning a new card game: Treatment A (in-person demonstration) and Treatment B (instruction booklet). There were 36 participants in each group. The outcome was whether the participant could correctly play a full round without help (success/failure). In Treatment A, 27 of 36 succeeded (p^A=0.75). In Treatment B, 18 of 36 succeeded (p^B=0.50). The observed difference in proportions was 0.75−0.50=0.25 (A − B). A randomization test shuffled the success/failure outcomes across the two groups 4000 times under “no treatment effect.” In the simulations, 40 of the 4000 shuffled differences were at least as large as 0.25 (A − B).
Which conclusion is most reasonable about the treatment effect?
- The result proves the demonstration causes exactly a 25 percentage-point increase for every participant, since participants were randomly assigned.
- There is evidence that the in-person demonstration increases the probability of success, because only 40 out of 4000 shuffled differences were at least 0.25, which is rare under no effect. (correct answer)
- There is evidence that the booklet is better, because the observed difference is 0.25 and we should look for simulated differences at most 0.25 (A − B).
- There is no evidence of a treatment effect because randomization tests cannot be used with proportions, only with means.
Explanation: This problem involves comparing treatments using proportions in a randomized experiment. Random assignment of participants allows causal conclusions about which learning method increases success probability. The observed difference of 0.25 (75% minus 50%) shows the in-person demonstration had a 25 percentage point higher success rate. The randomization test checks if this could occur by chance under no effect by shuffling outcomes 4000 times. Finding only 40 out of 4000 shuffled differences at least as large as 0.25 means the observed difference would be rare (1%) under no effect. This rarity provides evidence that the demonstration increases success probability, though it doesn't guarantee a 25-point increase for every individual. Randomization tests work perfectly well with proportions, not just means. When checking if A is better than B, we look for simulated differences at least as large as observed in the same direction (positive when A−B is calculated).
Question 18
A language app randomly assigned 88 users to two reminder settings: Treatment A (daily reminder) and Treatment B (no reminder), with 44 users in each group. The outcome was the number of practice sessions completed in a week. Treatment A users completed an average of 5.9 sessions and Treatment B users completed an average of 4.7 sessions, so the observed difference in means was 5.9−4.7=1.2 sessions (A − B). A randomization test shuffled the reminder labels 3000 times. In the shuffled results, 9 of the 3000 simulated differences were at least as large as 1.2 sessions (A − B).
Based on the randomization test, is the observed difference surprising under no effect?
- No. To check if A is better, we should count simulated differences ≤−1.2 because negative differences are more extreme in favor of A.
- Yes. Only 9 out of 3000 shuffled differences were at least 1.2, so the observed difference would be rare under no effect, suggesting the daily reminder increases the mean number of sessions. (correct answer)
- Yes. But only because users were randomly sampled from all app users worldwide, not because of random assignment.
- No. Since 9 is not zero, the observed difference is expected under no effect, so there is no evidence the reminder helps.
Explanation: This problem examines interpreting rare occurrences in randomization tests as evidence of treatment effect. Random assignment of app users allows causal conclusions about whether reminders increase practice sessions. The observed difference of 1.2 sessions (5.9 minus 4.7) shows the reminder group practiced more on average. The randomization test checks if this could occur by chance under no effect by shuffling labels 3000 times. Finding only 9 out of 3000 shuffled differences at least as large as 1.2 means the observed difference would be rare (0.3%) under no effect. This rarity provides strong evidence that daily reminders increase mean practice sessions. The fact that 9 isn't zero doesn't mean the difference is expected—9/3000 indicates extreme rarity. Random sampling from all users isn't necessary for causal conclusions; random assignment is sufficient. When A−B is positive and we're checking if A is better, we correctly count differences ≥ observed value.
Question 19
In a randomized experiment, 48 students were randomly assigned to two different methods for learning keyboard shortcuts: Treatment A (game-based practice) and Treatment B (worksheet practice). There were 24 students per group. After one week, the outcome was the number of shortcuts correctly recalled (out of 30).
Treatment A mean = 21.5, Treatment B mean = 18.9, so the observed difference (A − B) is 2.6.
A randomization test was conducted by shuffling the 48 recall scores between groups 2,000 times under “no treatment effect.” In the simulations, 28 of the 2,000 shuffled differences were at least as large as 2.6.
Which conclusion is most reasonable about the treatment effect (one-sided, A − B)?
- Because 28 out of 2,000 simulated differences were at least as large as 2.6, the result would be rare under no effect, so the data suggest Treatment A increases the mean number of shortcuts recalled. (correct answer)
- Because 28 out of 2,000 simulated differences were at least as large as 2.6, the observed difference is common under no effect, so there is strong evidence that Treatment B is better than Treatment A.
- Because the students were randomly assigned, the difference of 2.6 proves Treatment A improves every student by exactly 2.6 shortcuts.
- Because the students were not randomly sampled from all students, the random assignment does not allow any conclusion about whether the treatments caused different outcomes in this experiment.
Explanation: In comparing treatments using a randomized experiment and simulation, we evaluate if game-based practice (Treatment A) boosts keyboard shortcut recall more than worksheet practice (Treatment B). Random assignment creates comparable groups, so differences can be linked causally to the methods. The observed difference is A's mean recall minus B's, equaling 2.6 here. The 'no effect' simulation reassigns scores 2,000 times, building a distribution of differences from random variation alone. Only 28 out of 2,000 were at least as large as 2.6, showing rarity under no effect and evidence for A's advantage. Remember, random assignment differs from random sampling; it allows causal conclusions here but not population-wide, and rarity provides supportive evidence, not proof. In other cases, look at the count of simulated extremes to judge if the observed difference is rare.
Question 20
A school randomly assigned 60 students to try two different study plans for a vocabulary quiz: Treatment A (spaced practice) and Treatment B (single long review). Each group had 30 students. The quiz was scored out of 50 points. Treatment A had a mean score of 41.2 and Treatment B had a mean score of 38.6, for an observed difference (A − B) of 2.6 points. To test whether this difference could be due to chance under “no treatment effect,” the teacher performed a randomization test by shuffling the treatment labels 1,000 times and recalculating (A − B) each time. In the simulations, 12 of the 1,000 shuffled differences were at least as large as 2.6.
Which conclusion is most reasonable about the treatment effect (using the one-sided direction A − B)?
- Because the students were randomly assigned, Treatment A is guaranteed to be better than Treatment B by exactly 2.6 points.
- Because 12 out of 1,000 simulated differences were at least as large as 2.6, the observed difference would be fairly common under no effect, so there is not much evidence that Treatment A is better than Treatment B.
- Because only 12 out of 1,000 simulated differences were at least as large as 2.6, the result would be rare under no effect, so the data provide evidence that Treatment A leads to higher mean scores than Treatment B. (correct answer)
- The difference must be due to pre-existing differences between groups, since random assignment does not help with confounding.
Explanation: In comparing treatments using a randomized experiment and simulation, we investigate if spaced practice (Treatment A) leads to higher vocabulary quiz scores than a single long review (Treatment B). Random assignment of students to treatments helps ensure that any observed difference in scores is likely due to the treatments themselves, not other factors, allowing us to draw cause-and-effect conclusions. The observed difference is the mean score for A minus the mean for B, which here is 2.6 points. The 'no effect' simulation shuffles the treatment labels many times, assuming the treatments make no difference, and creates a distribution of possible differences that could occur just by chance. We compare the observed difference of 2.6 to this simulated distribution and find that only 12 out of 1,000 simulated differences were at least as large, making the observed result rare under no effect and providing evidence for a treatment effect. A common misconception is that random assignment is the same as random sampling, but random assignment balances groups in the experiment while random sampling would generalize to a larger population; also, rarity provides evidence but doesn't prove the effect beyond doubt. To apply this strategy elsewhere, count how often simulated differences are as extreme as or more extreme than the observed one to assess if the result is surprising.