Home

Tutoring

Subjects

Live Classes

Study Coach

Essay Review

On-Demand Courses

Colleges

Games


Sign up

Log in

Opening subject page...

Loading your content

Practice

  • All Subjects
  • Algebra Flashcards
  • SAT Math Practice Tests
  • Math Question of the Day
  • Live Classes
  • On-Demand Courses

Varsity Tutors

  • Find a Tutor
  • Test Prep
  • Online Classes
  • K-12 Learning
  • College Search
  • VarsityTutors.com

© 2026 Varsity Tutors. All rights reserved.

← Back to quizzes

Statistics Quiz

Statistics Quiz: Evaluating Models With Data And Simulation

Practice Evaluating Models With Data And Simulation in Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

Question 1 / 20

0 of 20 answered

A survey company claims that P(supports the proposal)=0.40P(\text{supports the proposal})=0.40P(supports the proposal)=0.40 for the population. They survey 80 people and find 28 support the proposal. You simulate 100 samples of 80 people each using P=0.40P=0.40P=0.40. In the simulations, 41 out of 100 samples had 28 or fewer supporters ("as extreme or more extreme" means ≤28\le 28≤28). Which conclusion is most reasonable?

Select an answer to continue

What this quiz covers

This quiz focuses on Evaluating Models With Data And Simulation, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A survey company claims that P(supports the proposal)=0.40P(\text{supports the proposal})=0.40P(supports the proposal)=0.40 for the population. They survey 80 people and find 28 support the proposal. You simulate 100 samples of 80 people each using P=0.40P=0.40P=0.40. In the simulations, 41 out of 100 samples had 28 or fewer supporters ("as extreme or more extreme" means ≤28\le 28≤28). Which conclusion is most reasonable?

  1. The result is unusual because you should instead count samples with 28 or more supporters, and that would be close to 0 out of 100.
  2. The result proves the model is correct because many simulations were at least as extreme.
  3. The result strongly contradicts the model because 28 is not equal to 0.40×80=320.40\times 80=320.40×80=32, so the model must be wrong.
  4. The result is not especially unusual under the model because 41 out of 100 simulated samples were as extreme, so it does not provide strong evidence against P=0.40P=0.40P=0.40. (correct answer)

Explanation: Evaluating fit to a probability model means assessing if 28 supporters in 80 surveys match P(support)=0.40 or seem inconsistent. Chance allows for variation, but only very rare outcomes strongly question the model. Through simulation, we generate 100 samples and count those as extreme as or more extreme—here, 28 or fewer supporters. With 41 out of 100 showing that, the result is not unusual, providing no strong evidence against the model. This leads to concluding the outcome is plausible under P=0.40 without notable doubt. Many think models require exact matches to the expected, like precisely 32, but randomness means samples vary, and not alternating support is normal. To apply broadly, define 'as extreme' from your data, simulate many times, and high frequencies like 41% indicate good model-data consistency.

Question 2

A factory claims 5% of its lightbulbs are defective, so P(defective)=0.05P(\text{defective})=0.05P(defective)=0.05. A quality inspector checks 60 bulbs and finds 9 defective. You decide to use simulation: repeatedly generate 60 bulbs with a 0.05 chance of being defective and count how often you get 9 or more defectives ("as extreme or more extreme" means ≥9\ge 9≥9). If your simulation shows that 0 out of 200 simulated runs had 9 or more defectives, would this result cause you to question the model? Why?

  1. No, because 60 is a small sample, and in small samples you should expect exactly 0.05×60=30.05\times 60=30.05×60=3 defectives.
  2. Yes, because finding 9 defectives proves the true defect rate is exactly 9/60=0.159/60=0.159/60=0.15.
  3. No, because simulation cannot be used to evaluate probability models; only exact calculations can.
  4. Yes, because outcomes this extreme did not occur in 200 simulated runs, so 9 defectives in 60 is very unlikely under P(defective)=0.05P(\text{defective})=0.05P(defective)=0.05. (correct answer)

Explanation: Checking data against a probability model means seeing if 9 defectives in 60 bulbs fits with P(defective)=0.05, or if it's inconsistent. Chance allows for some unusual results, but those that are exceedingly rare under the model can lead us to question its validity. Through simulation, we recreate the process 200 times, counting how often we get as extreme as or more extreme—here, 9 or more defectives. With 0 out of 200 simulations reaching that, the observed 9 appears very unlikely, supporting doubt about the 5% claim. This makes it reasonable to question the model because such a result is so rare under it. It's a misconception that small samples should exactly match the expected value, like precisely 3 defectives, but variation is normal and randomness doesn't mean alternation. To transfer this, define your extreme measure, run many simulations, and low frequencies like 0% suggest the observed data may not fit the model.

Question 3

A jar is said to contain red marbles with probability P(red)=0.50P(\text{red})=0.50P(red)=0.50 on each draw (with replacement). You draw 12 times and get 10 reds. You run 60 simulated sets of 12 draws under P(red)=0.50P(\text{red})=0.50P(red)=0.50. In those simulations, 11 out of 60 runs produced 10 or more reds ("as extreme or more extreme" means ≥10\ge 10≥10 reds). Which conclusion is most reasonable based on the simulations?

  1. Yes, you should doubt the model because 11 out of 60 is extremely rare, so the jar cannot have P(red)=0.50P(\text{red})=0.50P(red)=0.50.
  2. No, because 11 out of 60 runs were as extreme, so getting 10 or more reds in 12 draws is not especially unusual under the model. (correct answer)
  3. Yes, because you should instead count simulations with 10 or fewer reds, and that would be near 60 out of 60.
  4. No, because a probability model is confirmed whenever the observed result appears at least once in simulation.

Explanation: To verify a probability model, we examine if data like 10 reds in 12 draws match P(red)=0.50 or seem off. Outcomes can be unusual due to chance, but only the very rare ones truly challenge the model's credibility. Simulation recreates the draws many times—60 sets here—and measures how often we get as extreme as or more extreme, specifically 10 or more reds. Since 11 out of 60 simulations did, this frequency shows it's not especially unusual, so no strong reason to doubt the model. The reasonable conclusion is that the result is plausible under P=0.50 without raising significant questions. People often confuse randomness with alternation, but runs of the same color are possible, and small samples don't always reflect the probability closely. Generally, outline 'as extreme,' simulate repeatedly, and if the proportion is moderate like 18%, the data likely support the model.

Question 4

A spinner is advertised to land on Blue with probability P(Blue)=0.30P(\text{Blue})=0.30P(Blue)=0.30 each spin. A student spins it 30 times and gets Blue 7 times. To evaluate the claim, the student simulates 80 runs of 30 spins each using P(Blue)=0.30P(\text{Blue})=0.30P(Blue)=0.30. In the simulations, 29 out of 80 runs produced 7 or fewer Blues ("as extreme or more extreme" means ≤7\le 7≤7 Blues). What does the simulation suggest about how unusual the observed result is?

  1. It is very unusual because 29 out of 80 is close to 0, so results this low almost never happen under the model.
  2. It proves P(Blue)=0.30P(\text{Blue})=0.30P(Blue)=0.30 is correct because the observed count is not exactly 0.30×30=90.30\times 30=90.30×30=9.
  3. It is not especially unusual because 29 out of 80 runs were as extreme or more extreme, so the result is plausible under the model. (correct answer)
  4. It is unusual because you should count runs with 7 or more Blues, not 7 or fewer, and that would be rare.

Explanation: To check consistency with a probability model, like a spinner with P(Blue)=0.30 yielding 7 Blues in 30 spins, we assess if the data seem typical or odd under that probability. Outcomes can deviate from the expected by chance, but extremely rare deviations make us skeptical of the model. Simulation replicates the scenario many times—80 runs of 30 spins here—to find how often we see results as extreme as or more extreme, meaning 7 or fewer Blues. Since 29 out of 80 simulations showed that, the observed 7 is fairly common, suggesting it's plausible and not especially unusual. Thus, the simulation indicates the result aligns reasonably with the model without strong reason to doubt it. A misconception is that random outcomes must alternate colors or match the probability exactly in every sample, but small samples can vary a lot without being non-random. To use this approach generally, specify what counts as 'as extreme,' perform numerous simulations, and if the frequency is high, like over 20-30%, the data likely fit the model well.

Question 5

A delivery service claims that a package arrives on time with probability P(on time)=0.90P(\text{on time})=0.90P(on time)=0.90. Over 50 deliveries, a customer reports only 38 arrived on time. A simulation of 100 runs of 50 deliveries each (using P(on time)=0.90P(\text{on time})=0.90P(on time)=0.90) found that 0 out of 100 runs had 38 or fewer on-time deliveries ("as extreme or more extreme" means ≤38\le 38≤38 on time). What does the simulation suggest?

  1. The outcome proves the model is false, because simulation showed it cannot happen at all.
  2. The outcome is plausible because 38 is still more than half of 50, so it should occur often under P(on time)=0.90P(\text{on time})=0.90P(on time)=0.90.
  3. The outcome is not unusual because you should count runs with 38 or more on-time deliveries, not 38 or fewer.
  4. The outcome is very unusual under the model because none of the 100 simulated runs were as low as 38 on time, so the data raise doubt about P(on time)=0.90P(\text{on time})=0.90P(on time)=0.90. (correct answer)

Explanation: Checking data-model consistency involves seeing if 38 on-time deliveries in 50 fit P(on time)=0.90 or appear mismatched. Chance can cause deviations, but extremely rare ones under the model prompt skepticism. We use simulation to repeat the process 100 times, counting occurrences of results as extreme as or more extreme—here, 38 or fewer on time. None of the 100 runs showed that, indicating the outcome is very unusual and raises doubt about the 90% claim. This suggests the simulation points to the result being atypical under the model, challenging its accuracy. A misconception is that random should avoid long strings of delays or always hover near the average, but in reality, small samples can vary, though extremes like this are telling. To use this method, define your extreme criterion, run ample simulations, and low frequencies signal potential issues with the model.

Question 6

A card trick performer claims they can correctly guess a hidden card's color (red/black) with probability P(correct)=0.50P(\text{correct})=0.50P(correct)=0.50 per attempt if they are just guessing. In 30 attempts, they are correct 22 times. You simulate 90 runs of 30 random guesses with P(correct)=0.50P(\text{correct})=0.50P(correct)=0.50. In the simulations, 4 out of 90 runs had 22 or more correct ("as extreme or more extreme" means ≥22\ge 22≥22 correct). Would this result cause you to question the guessing model? Why?

  1. No, because 22 correct is close to half of 30, so it matches the guessing model.
  2. Yes, because only 4 out of 90 simulated runs were as extreme, so 22 or more correct is fairly rare under guessing and suggests the performer may be better than chance. (correct answer)
  3. Yes, because after several correct guesses, the next guess should be wrong to keep the overall rate near 50%.
  4. No, because rare results happen only when the sample size is 1; with 30 attempts, the result must match 15 correct exactly.

Explanation: We test a model's validity by checking if data, like 22 correct in 30 attempts, align with P(correct)=0.50 for guessing. Unusual results can occur by chance, but rarity under the model can lead us to question it. Simulation means running many trials—90 here—and finding the proportion with outcomes as extreme as or more extreme, meaning 22 or more correct. Only 4 out of 90 did, showing it's fairly rare, which suggests the performer might exceed chance and challenges the guessing model. Thus, yes, it causes doubt because such success is uncommon under pure guessing. It's wrong to think that after successes, failures must follow to balance, as randomness doesn't compensate, and small samples can have imbalances. For other cases, specify extremity, simulate extensively, and evaluate based on how often extremes appear in simulations.

Question 7

A game app claims its bonus wheel lands on Gold with probability P(Gold)=0.20P(\text{Gold})=0.20P(Gold)=0.20 each spin. A player spins the wheel 40 times and gets Gold 19 times. To check the claim, you simulate 100 runs of 40 spins each using P(Gold)=0.20P(\text{Gold})=0.20P(Gold)=0.20. In the simulations, 0 out of 100 runs produced 19 or more Golds ("as extreme or more extreme" means ≥19\ge 19≥19 Golds in 40 spins). Which conclusion is most reasonable based on the simulations?

  1. Because 19 is close to half of 40, it is not surprising; the simulations show outcomes like this are common under P(Gold)=0.20P(\text{Gold})=0.20P(Gold)=0.20.
  2. The result proves the wheel is rigged, because getting 19 Golds cannot happen if P(Gold)=0.20P(\text{Gold})=0.20P(Gold)=0.20.
  3. The relevant comparison is getting 19 or fewer Golds, so the simulations do not suggest the result is unusual.
  4. Because 0 out of 100 simulated runs were as extreme, the result is very unusual under the model and raises doubt about P(Gold)=0.20P(\text{Gold})=0.20P(Gold)=0.20. (correct answer)

Explanation: When we check if data fit a probability model, like a wheel landing on Gold with P(Gold)=0.20, we're seeing if the observed outcome—19 Golds in 40 spins—is consistent with that chance. Unusual outcomes can happen by chance alone, but if something is extremely rare under the model, it makes us doubt whether the model is accurate. Simulation helps by repeating the process many times under the assumed model to count how often we get results as extreme as or more extreme than observed, here meaning 19 or more Golds. In this case, 0 out of 100 simulations hit that mark, suggesting the observed 19 is very rare under P=0.20. This rarity supports concluding that the result raises doubt about the model's claim, as it's not what we'd typically expect. A common misconception is that random means outcomes should be evenly spread or alternate, but actually, clusters or streaks can occur, and small samples might not match the long-run probability closely. To apply this elsewhere, define 'as extreme' based on your observation, run lots of simulations, and if the frequency is very low, like under 5%, it might indicate the data don't fit the model well.

Question 8

A website claims that 60% of visitors click a certain button, so P(click)=0.60P(\text{click})=0.60P(click)=0.60. In a sample of 25 visitors, only 9 clicked. You simulate 120 samples of 25 visitors each using P(click)=0.60P(\text{click})=0.60P(click)=0.60. In the simulations, 3 out of 120 samples had 9 or fewer clicks ("as extreme or more extreme" means ≤9\le 9≤9 clicks). Would this result cause you to question the model? Why?

  1. No, because if the first few visitors did not click, later visitors are more likely to click to balance it out.
  2. Yes, because the result proves the true click rate is exactly 9/25=0.369/25=0.369/25=0.36.
  3. No, because 9 is close to 60% of 25, so it matches the model well.
  4. Yes, because only 3 out of 120 simulated samples were as extreme, so 9 or fewer clicks is rare under P(click)=0.60P(\text{click})=0.60P(click)=0.60 and raises doubt. (correct answer)

Explanation: Assessing a probability model requires checking if observations, such as 9 clicks out of 25 with P(click)=0.60, seem consistent or suspicious. Chance variation means not every outcome matches expectations perfectly, but very rare ones can make us doubt the model. We simulate the process multiple times—120 samples here—to see the frequency of results as extreme as or more extreme, meaning 9 or fewer clicks. With just 3 out of 120 showing that, the outcome is rare, justifying questions about the 60% claim. This supports saying yes, it raises doubt because such lows are uncommon under the model. A common error is thinking that after few clicks, more must follow to 'balance' it, but randomness doesn't work that way, and small samples can fluctuate. To apply this strategy, define extremity relative to your data, conduct many simulations, and use the resulting frequency to judge model fit.

Question 9

A basketball player claims they make free throws with probability P(make)=0.70P(\text{make})=0.70P(make)=0.70. In one practice session, they take 20 free throws and make 8. You decide to use simulation: repeatedly simulate 20 shots with P(make)=0.70P(\text{make})=0.70P(make)=0.70 and record how often you get 8 or fewer makes. In 100 simulated sessions, 0 produced 8 or fewer makes. Would this result cause you to question the model? Why?

  1. Yes; but only because the next free throw is now more likely to be a make after so many misses.
  2. No; a random process should have the same number of makes and misses in a short run, so 8 makes supports the model.
  3. Yes; 8 or fewer makes appears extremely rare under P(make)=0.70P(\text{make})=0.70P(make)=0.70, since it occurred 0 out of 100 simulated sessions. (correct answer)
  4. No; because 8 makes out of 20 is close to 70% when rounded.

Explanation: Checking data against a model, like P(make)=0.70 for free throws, involves seeing if 8 makes in 20 shots fits. Unusual results occur by chance, but very rare ones can cast doubt on the model. We simulate many sessions under the model to find the frequency of extreme outcomes, defined as 8 or fewer makes. With 0 out of 100 simulations showing that, it's extremely rare, so yes, it causes us to question the 70% claim. This rarity implies the low performance is unlikely if the model holds true. Misconceptions include thinking random means equal makes and misses in short runs, but variability is normal in small samples. To apply this strategy, specify 'as extreme' (e.g., 8 or less), simulate a bunch, and evaluate based on how often it appears.

Question 10

A bag is said to contain 40% blue marbles, so P(blue)=0.40P(\text{blue})=0.40P(blue)=0.40 on each draw with replacement. You draw 25 times and get 6 blue marbles. You simulate 100 runs of 25 draws under P(blue)=0.40P(\text{blue})=0.40P(blue)=0.40. In the simulations, 22 out of 100 runs produced 6 or fewer blues. Which conclusion is most reasonable?

  1. The result raises doubt because we should count 6 or more blues as extreme, not 6 or fewer.
  2. The result raises strong doubt because 22 out of 100 is extremely rare.
  3. The result does not provide strong evidence against the model because 6 or fewer blues happens fairly often (22 out of 100) in the simulations. (correct answer)
  4. The result proves the bag does have 40% blue marbles because 22 out of 100 is not 0.

Explanation: The idea is to verify if observed data, such as 6 blue marbles in 25 draws, match a model like P(blue)=0.40. Random chance can lead to odd outcomes sometimes, but only very rare ones typically prompt serious doubt. Simulation recreates the draws repeatedly under the model to measure how often we see results as extreme, here 6 or fewer blues. Since 22 out of 100 simulations showed that, it's fairly common (22%), so it doesn't provide strong evidence against the model. This means the result is plausible and doesn't raise much doubt. A common mix-up is expecting randomness to alternate colors evenly, but small samples naturally fluctuate. For broader use, define extremes (e.g., 6 or less), run simulations, and check if the frequency indicates rarity or not.

Question 11

A box of candies is said to contain 25% strawberry candies, so P(strawberry)=0.25P(\text{strawberry})=0.25P(strawberry)=0.25. You randomly select 16 candies (with replacement) and get 7 strawberry candies. You simulate 200 runs of 16 selections under P(strawberry)=0.25P(\text{strawberry})=0.25P(strawberry)=0.25. In the simulations, 6 out of 200 runs produced 7 or more strawberry candies. What does the simulation suggest about how unusual the observed result is, and whether you should doubt the model?

  1. It is not unusual because the correct comparison is 7 or fewer strawberries, not 7 or more.
  2. It is somewhat rare (6 out of 200), so it may raise doubt about the 25% model, though it does not prove the model is wrong. (correct answer)
  3. It proves the model is wrong because the observed count is not exactly 0.25×16=40.25\times 16=40.25×16=4.
  4. It is fairly common (6 out of 200), so it strongly supports the 25% model.

Explanation: We're checking if data like 7 strawberry candies in 16 selections fit P(strawberry)=0.25. Unusual outcomes happen by chance, but very rare ones can prompt doubt. Simulation mimics selections many times under the model to frequency-count extremes, such as 7 or more strawberries. With 6 out of 200 simulations (3%) showing that, it's somewhat rare, possibly raising doubt without proving the model wrong. This suggests mild evidence against the 25% claim. Misconception: randomness isn't alternating flavors; small samples fluctuate. Broadly, define 'extreme,' simulate extensively, and interpret the frequency for consistency.

Question 12

A website claims that 60% of visitors click an ad, so P(click)=0.60P(\text{click})=0.60P(click)=0.60. In a sample of 30 visitors, only 9 clicked. To evaluate this, you simulate many samples of 30 visitors using P(click)=0.60P(\text{click})=0.60P(click)=0.60 and look for results 9 or fewer clicks. In 100 simulated samples, 2 had 9 or fewer clicks. What does the simulation suggest about how unusual the observed result is?

  1. It proves the model is false because the simulated results did not match the observed result exactly.
  2. It is somewhat rare under the model (about 2%), so it provides evidence that may raise doubt about P(click)=0.60P(\text{click})=0.60P(click)=0.60. (correct answer)
  3. It is common under the model because 2 out of 100 is about half the time.
  4. It is not unusual because randomness should alternate, and 9 clicks is close to 15 clicks.

Explanation: To check consistency with a probability model, such as P(click)=0.60 for ad clicks, we assess if the data—like only 9 clicks in 30 visitors—seem plausible under that model. Outcomes can be unusual due to random chance, but very rare ones might indicate the model is off. Simulation involves running the scenario many times under the model to see how frequently we get results as extreme, defined here as 9 or fewer clicks. The 2 out of 100 simulations with such low clicks suggest it's somewhat rare, about 2%, providing evidence that could raise doubt about the 60% claim. Thus, the simulation implies the observed result is unusual enough to question the model without proving it false. A misconception is that random sequences should alternate evenly, but small samples often show clumps or lows naturally. In general, define 'as extreme' clearly, perform many simulations, and use the frequency to decide if doubt is warranted.

Question 13

A factory claims 5% of its light bulbs are defective, so P(defective)=0.05P(\text{defective})=0.05P(defective)=0.05. A quality inspector checks 60 bulbs and finds 7 defective. They simulate 100 runs of checking 60 bulbs under P(defective)=0.05P(\text{defective})=0.05P(defective)=0.05. In the simulations, 3 out of 100 runs had 7 or more defectives. Which conclusion is most reasonable based on the simulations?

  1. The result raises some doubt about the 5% model because 7 or more defectives happens rarely (3 out of 100) in the simulations. (correct answer)
  2. The result does not raise doubt because 3 out of 100 means it will happen in nearly every run.
  3. The result proves the defect rate is higher than 5% because 7 defectives is above the expected number.
  4. The result does not raise doubt because we should count 7 or fewer defectives as extreme, not 7 or more.

Explanation: We evaluate model consistency by checking if data like 7 defectives in 60 bulbs align with P(defective)=0.05. Chance can produce outliers, but rare extremes suggest the model might not fit. Simulation runs the process multiple times under the model to count extreme results, here 7 or more defectives. Occurring in 3 out of 100 simulations (3%) indicates it's somewhat rare, raising some doubt about the 5% claim. This supports a conclusion of mild evidence against the model without definitive proof. People often mistakenly believe randomness equals alternation, but small samples can deviate. Generally, define your extreme measure, simulate repeatedly, and use the frequency to judge plausibility.

Question 14

A coin is claimed to be fair: P(heads)=0.5P(\text{heads})=0.5P(heads)=0.5. In 20 flips, you observe 18 heads. You decide to evaluate the claim by simulation: repeat 20-flip experiments many times using a random number generator with P(heads)=0.5P(\text{heads})=0.5P(heads)=0.5 and count how often you get 18 or more heads. Would this result cause you to question the model? Why? (Here “as extreme or more extreme” means ≥18\ge 18≥18 heads out of 20.)

  1. Yes; it proves the coin must have P(heads)=0.9P(\text{heads})=0.9P(heads)=0.9 since 18/20 is about 0.9.
  2. No; because the next flips are more likely to be tails after so many heads, the model will balance out.
  3. Yes; ≥18\ge 18≥18 heads out of 20 would be very rare for a fair coin, so it would raise doubt about the model. (correct answer)
  4. No; a fair coin should alternate often, and 18 heads shows it is random so the model is fine.

Explanation: This question examines whether getting 18 heads in 20 coin flips is consistent with a fair coin model where P(heads) = 0.5. With a fair coin, we'd expect about 10 heads, but random results vary—the question is whether 18 heads is unusually extreme. To evaluate this, we'd simulate many sets of 20 flips using P(heads) = 0.5 and count how often we get 18 or more heads. For a fair coin, getting 18+ heads out of 20 would be extremely rare (probability ≈ 0.0002), so observing this outcome would strongly suggest the coin isn't fair. A key misconception is believing that random outcomes should alternate regularly—true randomness can produce streaks. Another error is thinking past results affect future flips (gambler's fallacy). The evaluation strategy remains: simulate under the model, check how often results as extreme occur, and if very rarely, doubt the model.

Question 15

A cafeteria claims that 40% of students choose pizza on Fridays: P(pizza)=0.40P(\text{pizza})=0.40P(pizza)=0.40. You survey 25 students and 5 choose pizza. A student simulates 200 surveys of 25 students under P(pizza)=0.40P(\text{pizza})=0.40P(pizza)=0.40; 3 of the 200 simulated surveys had 5 or fewer pizza-choosers. Would this result cause you to question the model? Why? (Treat “as extreme or more extreme” as ≤5\le 5≤5 out of 25 choosing pizza.)

  1. Yes; only 3 out of 200 simulations had 5 or fewer, so the observed result is very unusual under the model and raises doubt. (correct answer)
  2. No; 3 out of 200 shows it is common, so there is no reason to doubt the model.
  3. Yes; it proves the true pizza rate is exactly 5/25=0.205/25=0.205/25=0.20 because that’s what you observed.
  4. No; because 5 is exactly 40% of 25, the result matches the model.

Explanation: This question asks whether 5 students choosing pizza out of 25 is consistent with a model where P(pizza) = 0.40. Under this model, we'd expect about 10 students to choose pizza (40% of 25), so 5 is notably low—but is it unusually low? The simulation shows that only 3 out of 200 simulated surveys had 5 or fewer pizza-choosers, giving a frequency of 1.5% (3/200 = 0.015). When an observed result would occur less than 2% of the time under the assumed model, it's quite rare and raises significant doubt about whether the true pizza preference rate is really 40%. A common error is thinking that because 5/25 = 20% is still a reasonable percentage, the result supports the model, but the key is how rarely this outcome occurs under the 40% assumption. The strategy is clear: simulate under the claimed model, and when results as extreme as observed happen very rarely (1.5%), question the model's accuracy.

Question 16

A dartboard game claims a player hits the bullseye with probability P(bullseye)=0.20P(\text{bullseye})=0.20P(bullseye)=0.20. In 30 throws, the player hits 2 bullseyes. A student simulates 100 sets of 30 throws under P(bullseye)=0.20P(\text{bullseye})=0.20P(bullseye)=0.20; 6 of the 100 simulated sets had 2 or fewer bullseyes. Would this result cause you to question the model? Why? (Treat “as extreme or more extreme” as getting ≤2\le 2≤2 bullseyes out of 30.)

  1. Yes; 6 out of 100 simulations suggests ≤2\le 2≤2 bullseyes is pretty common, so the model is likely correct.
  2. No; after many non-bullseyes, the next throw is more likely to be a bullseye, so the low count is not surprising.
  3. Yes; only 6 out of 100 simulations were as low as 2 or fewer, so it seems fairly rare and raises some doubt about the 0.20 claim. (correct answer)
  4. No; 2 bullseyes is exactly what you’d expect because 0.20×30=20.20\times 30=20.20×30=2.

Explanation: This question asks whether hitting 2 bullseyes in 30 throws is consistent with a model where P(bullseye) = 0.20. Under this model, we'd expect about 6 bullseyes (20% of 30), so 2 is notably below expectation—but is it unusually low? The simulation results show that 6 out of 100 simulated sets had 2 or fewer bullseyes, giving a frequency of 6% (6/100 = 0.06). While not extremely rare, a 6% occurrence rate is fairly uncommon and would raise some doubt about whether the true probability really is 0.20. A misconception is believing in the "hot hand" fallacy—that after many misses, a bullseye becomes more likely, when in fact each throw is independent. The strategy remains consistent: define what's extreme (≤2 bullseyes), run simulations, and when the observed result happens infrequently (like 6% here), it suggests the model might not be accurate.

Question 17

A card app claims it deals a heart with probability P(♡)=0.25P(\heartsuit)=0.25P(♡)=0.25 on each deal (with replacement). In 12 deals, you get 6 hearts. A student simulates 300 sets of 12 deals under P(♡)=0.25P(\heartsuit)=0.25P(♡)=0.25; 4 of the 300 simulated sets had 6 or more hearts. Would this result cause you to question the model? Why? (Treat “as extreme or more extreme” as ≥6\ge 6≥6 hearts out of 12.)

  1. No; 4 out of 300 means it happens most of the time, so the model looks consistent.
  2. Yes; only 4 out of 300 simulations had ≥6\ge 6≥6 hearts, so the observed result is quite rare under the model and raises doubt. (correct answer)
  3. Yes; it proves the app is rigged to give hearts every other deal, since randomness should alternate suits.
  4. No; because 6 hearts is exactly half, and randomness should give about half hearts in any short run.

Explanation: This question evaluates whether getting 6 hearts in 12 card deals is consistent with a model where P(♥) = 0.25. Under this model, we'd expect about 3 hearts (25% of 12), so 6 hearts is double the expectation—but could it happen by chance? The simulation shows that only 4 out of 300 simulated sets had 6 or more hearts, giving a frequency of about 1.3% (4/300 ≈ 0.013). When an observed result would occur less than 2% of the time under the assumed model, it's quite rare and raises significant doubt about whether the true probability of dealing a heart is really 0.25. A misconception is thinking randomness means perfect alternation between suits—true randomness can produce clusters. Another error is misreading 4/300 as meaning it happens most of the time. The evaluation method remains: simulate under the model, and when extreme results occur very rarely (1.3%), doubt the model's validity.

Question 18

A video game claims that a rare item drops with probability P(drop)=0.15P(\text{drop})=0.15P(drop)=0.15 each time a player defeats a boss. A player defeats the boss 30 times and gets 2 drops. A student simulates 100 sets of 30 defeats under P(drop)=0.15P(\text{drop})=0.15P(drop)=0.15; 18 of the 100 simulated sets had 2 or fewer drops. Would this result cause you to question the model? Why? (Treat “as extreme or more extreme” as getting ≤2\le 2≤2 drops in 30 tries.)

  1. No; 2 or fewer drops occurred in 18 out of 100 simulations, so the observed result is not especially rare under the model. (correct answer)
  2. Yes; 18 out of 100 means the result almost never happens, so the model should be doubted.
  3. No; the only relevant comparison is getting ≥2\ge 2≥2 drops, and that will happen in almost every simulation.
  4. Yes; because 2 is less than the expected 4.5, it proves the true drop rate is below 0.15.

Explanation: This question evaluates whether getting 2 drops in 30 boss defeats is consistent with a model where P(drop) = 0.15. Under this model, we'd expect about 4.5 drops (15% of 30), so 2 drops is below expectation—but is it unusually low? The simulation results show that 18 out of 100 simulated sets had 2 or fewer drops, meaning this outcome occurred in 18% of simulations. Since getting 2 or fewer drops happens in nearly one-fifth of all simulations, this result is fairly common under the assumed model and provides no strong reason to doubt it. A misconception is misinterpreting what 18/100 means—it doesn't mean the result almost never happens; rather, it happens reasonably often. The evaluation process remains: define what's extreme (≤2 drops), simulate many times, and when the observed result occurs with reasonable frequency (18%), the model remains plausible.

Question 19

A game booth claims its prize wheel lands on Blue with probability P(Blue)=0.25P(\text{Blue})=0.25P(Blue)=0.25. You watch 40 spins and see Blue come up 22 times. A student runs 100 simulated sets of 40 spins using the model P(Blue)=0.25P(\text{Blue})=0.25P(Blue)=0.25; in those simulations, 0 out of 100 sets had 22 or more Blues. Would this result cause you to question the model? Why? (Treat “as extreme or more extreme” as getting ≥22\ge 22≥22 Blues in 40 spins.)

  1. Yes; getting ≥22\ge 22≥22 Blues seems extremely rare under the simulations (0 out of 100), so it raises doubt about P(Blue)=0.25P(\text{Blue})=0.25P(Blue)=0.25. (correct answer)
  2. No; 22 Blues out of 40 is close to 25% because random results should vary a lot, so it supports the model.
  3. Yes; it proves the wheel cannot have P(Blue)=0.25P(\text{Blue})=0.25P(Blue)=0.25 because the observed result is different from exactly 10 Blues.
  4. No; since 0 out of 100 simulations matched exactly 22 Blues, that means 22 is the most likely outcome under the model.

Explanation: This question asks whether observing 22 Blues in 40 spins is consistent with a probability model where P(Blue) = 0.25. If the model is correct, we'd expect about 10 Blues (25% of 40), but random outcomes naturally vary—the key is determining whether 22 Blues is unusually high. The student ran 100 simulations under the claimed model, and none of them produced 22 or more Blues, meaning this outcome appears extremely rare (0/100 = 0%). When simulations show that an observed result would almost never happen under the assumed model, it raises serious doubt about whether that model is correct. A common misconception is thinking that any deviation from the expected value disproves the model, but here the simulation tells us that 22 Blues is far beyond normal variation. The strategy is: define what counts as extreme (≥22 Blues), run many simulations, and check how often such extreme results occur—if very rarely, question the model.

Question 20

A factory claims 5% of its lightbulbs are defective, so P(defective)=0.05P(\text{defective})=0.05P(defective)=0.05. An inspector checks 80 bulbs and finds 10 defective. To evaluate the claim, 50 simulated inspections of 80 bulbs are generated under the model. In the 50 simulations, 0 inspections had 10 or more defectives ("as extreme or more extreme" means ≥10\ge 10≥10 defectives). Would this result cause you to question the model? Why?​​

  1. Yes; because none of the simulations reached 10 or more defectives, the observed result appears extremely rare under P=0.05P=0.05P=0.05 and raises doubt. (correct answer)
  2. No; because 10 out of 80 is only 12.5%, and small samples often differ from 5% by that much.
  3. No; the result proves the 5% model is correct because defectives are random and anything can happen.
  4. No; because the correct tail is "10 or fewer" defectives, and that would happen in nearly every simulation.

Explanation: To check a factory's 5% defective rate model, we assess if 10 defectives in 80 bulbs fits or raises questions. Variability happens by chance, but extremely high counts under the model can lead to doubt. Simulation runs many inspections of 80 bulbs with P(defective)=0.05 to see the frequency of >=10. With 0 out of 50 simulations at >=10, this is extremely rare. Therefore, yes, it causes questioning, as such a high defective count is improbable under the claim. Some think after non-defectives, one's 'due,' but independence rules that out, and small inspections can differ from 5% naturally. For transfer, set 'as extreme' to >=10, simulate repeatedly, and if zero occurrences, doubt the model.