AP Statistics Quiz: Difference Of Two Means Test
20 questions · exam conditions
0:00
Difference Of Two Means TestQuestion 1 of 20

A school district wants to know whether a new reading program changes mean reading-comprehension scores. A random sample of 40 students used the new program (Group N) and a separate random sample of 38 students used the old program (Group O). A two-sample tt test for μNμO\mu_N-\mu_O was conducted with hypotheses H0:μNμO=0H_0: \mu_N-\mu_O=0 and Ha:μNμO>0H_a: \mu_N-\mu_O>0. The test produced a p-value of 0.018. Using α=0.05\alpha=0.05, what conclusion is appropriate?

Fail to reject H0H_0; there is not convincing evidence that the new program increases the mean score.
Reject H0H_0; there is convincing evidence that the new program increases the population mean score compared with the old program.
Reject H0H_0; there is convincing evidence that the old program increases the population mean score compared with the new program.
Fail to reject H0H_0; the samples prove the two sample means are equal.
Reject H0H_0; the new program causes higher scores for every student in the district.
← Back to quizzes

AP Statistics Quiz

AP Statistics Quiz: Difference Of Two Means Test

Practice Difference Of Two Means Test in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Difference Of Two Means Test, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A school district wants to know whether a new reading program changes mean reading-comprehension scores. A random sample of 40 students used the new program (Group N) and a separate random sample of 38 students used the old program (Group O). A two-sample tt test for μNμO\mu_N-\mu_O was conducted with hypotheses H0:μNμO=0H_0: \mu_N-\mu_O=0 and Ha:μNμO>0H_a: \mu_N-\mu_O>0. The test produced a p-value of 0.018. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence that the new program increases the mean score.
  2. Reject H0H_0; there is convincing evidence that the new program increases the population mean score compared with the old program. (correct answer)
  3. Reject H0H_0; there is convincing evidence that the old program increases the population mean score compared with the new program.
  4. Fail to reject H0H_0; the samples prove the two sample means are equal.
  5. Reject H0H_0; the new program causes higher scores for every student in the district.

Explanation: This question tests understanding of hypothesis test conclusions for a difference of two means. Since the p-value (0.018) is less than α (0.05), we reject the null hypothesis. The alternative hypothesis states μN - μO > 0, which means we're testing if the new program has a higher mean than the old program. When we reject H0, we have convincing evidence supporting the alternative hypothesis - that the new program increases the population mean score. Choice D incorrectly states we "prove" equality, but hypothesis tests never prove anything. Choice E overstates the conclusion by claiming the program works for "every student," when we can only make claims about population means.

Question 2

A city compares mean response time (minutes) for two ambulance dispatch systems. Independent random samples of calls are taken from System 1 (S1) and System 2 (S2). A two-sample tt test for μS1μS2\mu_{S1}-\mu_{S2} is performed with H0:μS1μS2=0H_0: \mu_{S1}-\mu_{S2}=0 and Ha:μS1μS2>0H_a: \mu_{S1}-\mu_{S2}>0. The p-value is 0.11. Using α=0.10\alpha=0.10, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that System 1 has a larger population mean response time than System 2.
  2. Fail to reject H0H_0; there is not convincing evidence that System 1 has a larger population mean response time than System 2. (correct answer)
  3. Reject H0H_0; there is convincing evidence that System 1 has a smaller population mean response time than System 2.
  4. Fail to reject H0H_0; therefore the two systems have exactly the same mean response time in the population.
  5. Reject H0H_0; using System 1 causes slower response times for all calls.

Explanation: This question involves comparing ambulance response times with a right-tailed test. The p-value (0.11) is greater than α (0.10), so we fail to reject the null hypothesis. The alternative Ha: μS1 - μS2 > 0 tests whether System 1 has a larger (worse) mean response time than System 2. Since we fail to reject H0, we don't have convincing evidence that System 1 has a larger population mean response time. Choice D incorrectly interprets failing to reject as proof of equality - we simply lack evidence of a difference. Choice E makes an inappropriate causal claim. When p-value > α, we always fail to reject H0 and conclude there's insufficient evidence for the alternative hypothesis.

Question 3

A school district wants to know whether a new online homework system changes mean weekly math quiz scores. A random sample of students used the new system (group N) and another random sample used the old system (group O). A two-sample tt test was performed for H0:μNμO=0H_0: \mu_N-\mu_O=0 versus Ha:μNμO>0H_a: \mu_N-\mu_O>0. The test produced a p-value of 0.03. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence that the new system increases the mean quiz score.
  2. Reject H0H_0; there is convincing evidence that the new system increases the mean quiz score. (correct answer)
  3. Reject H0H_0; there is convincing evidence that the old system increases the mean quiz score.
  4. Fail to reject H0H_0; the two samples have the same mean quiz score, so the population means are equal.
  5. Reject H0H_0; the new system caused individual students' quiz scores to increase.

Explanation: This question tests your ability to interpret a two-sample t-test for the difference of means. Since the p-value (0.03) is less than the significance level α = 0.05, we reject the null hypothesis. The alternative hypothesis states μ_N - μ_O > 0, which means we're testing if the new system has a higher mean than the old system. By rejecting H₀, we conclude there is convincing evidence that the new system increases the mean quiz score. Choice D incorrectly claims we can conclude the population means are equal from sample data, and Choice E incorrectly makes a causal claim about individual students. When conducting hypothesis tests for difference of means, we compare the p-value to α and make conclusions about population means, not individual values.

Question 4

A psychologist studies whether a mindfulness app affects mean stress score (higher = more stress). Participants were randomly assigned to use the app (group A) or a placebo app (group P). A two-sample tt test was conducted for H0:μAμP=0H_0: \mu_A-\mu_P=0 versus Ha:μAμP<0H_a: \mu_A-\mu_P<0. The p-value was 0.018. At α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence that the mindfulness app lowers mean stress score.
  2. Reject H0H_0; there is convincing evidence that the mindfulness app lowers mean stress score. (correct answer)
  3. Reject H0H_0; there is convincing evidence that the mindfulness app raises mean stress score.
  4. Reject H0H_0; the mindfulness app is proven to reduce stress for every individual.
  5. Reject H0H_0; the two sample means are different, so the app group mean must be exactly 0.018 lower.

Explanation: This randomized experiment tests whether a mindfulness app lowers mean stress score, with H_a: μ_A - μ_P < 0. The p-value of 0.018 is less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the mindfulness app lowers mean stress score. Choice D incorrectly claims the app reduces stress for every individual, while Choice E wrongly states that we can determine the exact difference in population means. Even in randomized experiments, hypothesis tests provide evidence about population parameters (means), not guarantees about individual outcomes. The p-value tells us about statistical significance, not the size of the effect.

Question 5

A school compares mean time (in minutes) to complete a standardized math test for students using a paper booklet versus an online version. Two independent random samples were taken (paper: n=40n=40; online: n=38n=38). A two-sample tt test for μpaperμonline\mu_{paper}-\mu_{online} used H0:μpaperμonline=0H_0:\mu_{paper}-\mu_{online}=0 and Ha:μpaperμonline<0H_a:\mu_{paper}-\mu_{online}<0. The p-value was 0.041. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is sufficient evidence that the population mean completion time is lower for paper than for online. (correct answer)
  2. Fail to reject H0H_0; there is sufficient evidence that the population mean completion time is lower for paper than for online.
  3. Reject H0H_0; there is sufficient evidence that the population mean completion time is higher for paper than for online.
  4. Fail to reject H0H_0; the sample data show paper was faster, so paper is faster in the population.
  5. Reject H0H_0; using paper causes students to finish faster.

Explanation: This question tests the application of a two-sample t-test for the difference in mean completion times between paper and online math tests. The one-sided alternative hypothesizes that paper is faster, and since the p-value of 0.041 is less than α=0.05, we reject H0, supporting that the population mean time is lower for paper. A frequent distractor is choice B, which mistakenly says 'fail to reject' while claiming sufficient evidence, confusing the decision rule. When concluding two-mean tests, specify the direction if it's a one-sided test and evidence supports it, but don't generalize to causation like in choice E. Focus on population inferences, as sample differences alone (as in choice D) don't guarantee population differences without statistical significance.

Question 6

An environmental scientist tests whether the mean nitrate concentration (mg/L) differs between two nearby lakes, Lake 1 and Lake 2. Independent random water samples were collected from each lake. A two-sample tt test was performed for H0:μ1μ2=0H_0: \mu_1-\mu_2=0 versus Ha:μ1μ20H_a: \mu_1-\mu_2\ne 0. The p-value was 0.049. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence of a difference in mean nitrate concentration between the lakes.
  2. Reject H0H_0; there is convincing evidence that the mean nitrate concentration differs between the two lakes. (correct answer)
  3. Reject H0H_0; there is convincing evidence that Lake 1 has a higher mean nitrate concentration than Lake 2.
  4. Reject H0H_0; the sampling shows the lakes' nitrate means are different, so one lake caused the other to change.
  5. Fail to reject H0H_0; because 0.049 is close to 0.05, the result is inconclusive and no decision can be made.

Explanation: This two-sample t-test uses a two-tailed alternative hypothesis (μ₁ - μ₂ ≠ 0) to test for any difference in mean nitrate concentration between the lakes. The p-value of 0.049 is just barely less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the mean nitrate concentration differs between the two lakes. Choice C incorrectly specifies a direction (Lake 1 higher) when the two-tailed test doesn't indicate which lake has higher concentration. Choice E wrongly suggests that a p-value close to α makes the result inconclusive. In hypothesis testing, we use a clear decision rule: if p-value < α, we reject H₀, regardless of how close the values are.

Question 7

A sports scientist compares mean resting heart rate for two populations: endurance athletes (A) and nonathletes (N). Independent random samples are taken, and a two-sample tt test for μAμN\mu_A-\mu_N is performed with H0:μAμN=0H_0: \mu_A-\mu_N=0 and Ha:μAμN<0H_a: \mu_A-\mu_N<0. The p-value is 0.27. At the α=0.05\alpha=0.05 level, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that athletes have a lower population mean resting heart rate than nonathletes.
  2. Fail to reject H0H_0; there is not convincing evidence that athletes have a lower population mean resting heart rate than nonathletes. (correct answer)
  3. Reject H0H_0; there is convincing evidence that athletes have a higher population mean resting heart rate than nonathletes.
  4. Fail to reject H0H_0; therefore the two population means are equal.
  5. Reject H0H_0; being an endurance athlete causes a lower resting heart rate in the population.

Explanation: This problem involves a left-tailed test for the difference between athlete and non-athlete mean heart rates. The p-value (0.27) is much larger than α (0.05), so we fail to reject the null hypothesis. The alternative hypothesis Ha: μA - μN < 0 tests whether athletes have a lower mean heart rate than non-athletes. Since we fail to reject H0, we do not have convincing evidence that athletes have a lower population mean resting heart rate. Choice D incorrectly claims that failing to reject H0 means the means are equal - we simply don't have evidence of a difference. Remember that failing to reject H0 never proves the null hypothesis is true; it only means we lack sufficient evidence against it.

Question 8

A teacher compares mean quiz scores for students using two different study apps. Two independent random samples are taken: App A and App B. A two-sample tt test for μAμB\mu_A-\mu_B is conducted with H0:μAμB=0H_0: \mu_A-\mu_B=0 and Ha:μAμB0H_a: \mu_A-\mu_B\neq 0. The p-value is 0.52. At α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that the population mean quiz scores differ between the two apps.
  2. Fail to reject H0H_0; there is not convincing evidence of a difference in population mean quiz scores between the two apps. (correct answer)
  3. Reject H0H_0; there is convincing evidence that App A has a higher population mean quiz score than App B.
  4. Fail to reject H0H_0; this proves both apps produce the same mean score for all students.
  5. Reject H0H_0; using App A causes higher quiz scores than using App B.

Explanation: This problem presents a two-tailed test comparing study apps. The p-value (0.52) is much larger than α (0.05), so we fail to reject the null hypothesis. With a two-tailed alternative (Ha: μA - μB ≠ 0), failing to reject means we don't have convincing evidence of any difference between the population mean quiz scores. Choice A incorrectly suggests rejecting H0 when the large p-value clearly indicates we should fail to reject. Choice D wrongly claims this "proves" equality - hypothesis tests never prove the null hypothesis. A large p-value simply means our sample data are consistent with the null hypothesis of no difference between population means.

Question 9

A school district compares mean math test scores for students taught with Method 1 versus Method 2. Two independent random samples of students were selected (Method 1: n=28n=28; Method 2: n=30n=30). A two-sample tt test for μ1μ2\mu_1-\mu_2 was performed with H0:μ1=μ2H_0:\mu_1=\mu_2 and Ha:μ1>μ2H_a:\mu_1>\mu_2. The p-value was 0.004. Using α=0.01\alpha=0.01, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence that Method 1 leads to a higher mean score than Method 2.
  2. Reject H0H_0; there is convincing evidence that Method 1 has a higher population mean score than Method 2. (correct answer)
  3. Reject H0H_0; Method 1 causes higher scores for every student.
  4. Reject H0H_0; there is convincing evidence that Method 2 has a higher population mean score than Method 1.
  5. Fail to reject H0H_0; the sample means are equal, so the population means are equal.

Explanation: This question tests a one-tailed hypothesis where we're examining if Method 1 produces higher mean scores than Method 2 (H_a: μ₁ > μ₂). The p-value (0.004) is less than α = 0.01, so we reject the null hypothesis. This provides convincing evidence that Method 1 has a higher population mean score than Method 2. Choice C incorrectly claims the effect applies to "every student," which is an overstatement—hypothesis tests make conclusions about population means, not individual outcomes. Choice D reverses the direction of the conclusion. When conducting hypothesis tests, we make inferences about population parameters based on sample data, not claims about every individual in the population.

Question 10

A public health researcher compares mean systolic blood pressure for adults who exercise regularly versus those who do not. Two independent random samples were taken (Exercise: n=50n=50; No exercise: n=48n=48). A two-sample tt test for μexμno\mu_{\text{ex}}-\mu_{\text{no}} was conducted with H0:μex=μnoH_0:\mu_{\text{ex}}=\mu_{\text{no}} and Ha:μex<μnoH_a:\mu_{\text{ex}}<\mu_{\text{no}}. The p-value was 0.031. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that the population mean systolic blood pressure is lower for adults who exercise regularly than for those who do not. (correct answer)
  2. Fail to reject H0H_0; there is not convincing evidence that exercisers have lower mean systolic blood pressure.
  3. Reject H0H_0; exercising causes lower systolic blood pressure for every adult.
  4. Reject H0H_0; there is convincing evidence that exercisers have higher mean systolic blood pressure than non-exercisers.
  5. Fail to reject H0H_0; the sample means show the two groups have the same mean blood pressure.

Explanation: This question tests understanding of a one-tailed test where we're examining if exercisers have lower mean systolic blood pressure (H_a: μ_ex < μ_no). The p-value (0.031) is less than α = 0.05, so we reject the null hypothesis. This provides convincing evidence that the population mean systolic blood pressure is lower for adults who exercise regularly than for those who do not. Choice C incorrectly claims causation for "every adult," which overstates the conclusion. Choice E misinterprets what failing to reject would mean. When we reject H₀ in a one-tailed test with "less than" alternative, we conclude there's evidence supporting the specific directional claim in the alternative hypothesis.

Question 11

A researcher compares mean daily screen time (hours) for teenagers in Urban areas versus Rural areas. Independent random samples were taken (Urban: n=45n=45; Rural: n=43n=43). A two-sample tt test for μUμR\mu_U-\mu_R used hypotheses H0:μU=μRH_0:\mu_U=\mu_R and Ha:μUμRH_a:\mu_U\ne\mu_R. The p-value was 0.051. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that mean screen time differs between Urban and Rural teenagers.
  2. Fail to reject H0H_0; there is not convincing evidence at the 0.05 level that population mean screen time differs between Urban and Rural teenagers. (correct answer)
  3. Reject H0H_0; Urban living causes higher screen time than Rural living.
  4. Fail to reject H0H_0; therefore the population means are equal.
  5. Reject H0H_0; there is convincing evidence that Urban teenagers have lower mean screen time than Rural teenagers.

Explanation: This problem involves a two-tailed test comparing mean screen time between Urban and Rural teenagers. The p-value (0.051) is just slightly greater than α = 0.05, so we fail to reject the null hypothesis. This means there is not convincing evidence at the 0.05 level that population mean screen time differs between Urban and Rural teenagers. Choice D incorrectly concludes that the population means are equal—failing to reject H₀ doesn't prove equality. Choice C makes an inappropriate causal claim. When the p-value is very close to α but slightly exceeds it, we still fail to reject H₀, though the evidence is nearly significant at the chosen level.

Question 12

An engineer compares mean battery life (hours) for two brands, Brand X and Brand Y. Independent random samples were tested (Brand X: n=18n=18; Brand Y: n=20n=20). A two-sample tt test for μXμY\mu_X-\mu_Y used H0:μX=μYH_0:\mu_X=\mu_Y and Ha:μXμYH_a:\mu_X\ne\mu_Y. The p-value was 0.62. At α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that the mean battery lives differ between Brand X and Brand Y.
  2. Fail to reject H0H_0; there is not convincing evidence of a difference in population mean battery life between Brand X and Brand Y. (correct answer)
  3. Fail to reject H0H_0; the two brands have exactly the same mean battery life.
  4. Reject H0H_0; there is convincing evidence that Brand X has a higher mean battery life than Brand Y.
  5. Reject H0H_0; Brand Y causes longer battery life than Brand X.

Explanation: This problem involves a two-tailed test comparing mean battery life between two brands. With a p-value of 0.62, which greatly exceeds α = 0.05, we fail to reject the null hypothesis. This means there is not convincing evidence of a difference in population mean battery life between Brand X and Brand Y. Choice C incorrectly states that the brands have "exactly the same mean"—failing to reject H₀ doesn't prove equality, it only indicates insufficient evidence for a difference. Choice E makes an inappropriate causal claim. In hypothesis testing, a large p-value suggests the observed difference could easily occur by chance if the null hypothesis were true, so we maintain our assumption of no difference.

Question 13

A company tests whether a new training program reduces the mean time (in minutes) to complete a standard task. A random sample of employees used the new program (group T) and a separate random sample used the old program (group C). A two-sample tt test was performed for H0:μTμC=0H_0: \mu_T-\mu_C=0 versus Ha:μTμC<0H_a: \mu_T-\mu_C<0. The p-value was 0.08. Using α=0.10\alpha=0.10, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not convincing evidence that the new program reduces the mean completion time.
  2. Reject H0H_0; there is convincing evidence that the new program reduces the mean completion time. (correct answer)
  3. Reject H0H_0; there is convincing evidence that the new program increases the mean completion time.
  4. Reject H0H_0; the new program caused every employee to complete the task faster.
  5. Fail to reject H0H_0; the sample data prove there is no difference in population mean completion times.

Explanation: In this two-sample t-test, we're testing whether the new training program reduces mean completion time, with H_a: μ_T - μ_C < 0 (new minus old is negative). The p-value of 0.08 is less than the significance level α = 0.10, so we reject the null hypothesis. This provides convincing evidence that the new program reduces the mean completion time. Choice D incorrectly makes a claim about every individual employee rather than the population mean. Choice E wrongly states that failing to reject H₀ would prove no difference exists. When interpreting hypothesis tests, focus on whether the p-value is less than α, and remember that conclusions apply to population parameters, not individual observations.

Question 14

A bookstore owner compares mean amount spent per customer on weekdays versus weekends. Independent random samples were taken (weekday: n=25n=25; weekend: n=27n=27). A two-sample tt test for μweekendμweekday\mu_{weekend}-\mu_{weekday} used H0:μweekendμweekday=0H_0:\mu_{weekend}-\mu_{weekday}=0 and Ha:μweekendμweekday>0H_a:\mu_{weekend}-\mu_{weekday}>0. The p-value was 0.073. At α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is sufficient evidence that the population mean amount spent is higher on weekends than on weekdays.
  2. Fail to reject H0H_0; there is not sufficient evidence that the population mean amount spent is higher on weekends than on weekdays. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that the population mean amount spent is lower on weekends than on weekdays.
  4. Fail to reject H0H_0; therefore weekend and weekday customers spend the same amount on average.
  5. Fail to reject H0H_0; weekend shopping causes customers to spend more, but the sample was too small to show it.

Explanation: This question tests a two-sample t-test for comparing mean spending between weekdays and weekends. The one-sided alternative posits higher weekend spending, but p=0.073 > α=0.05 leads to failing to reject H0, with insufficient evidence for the claim. Choice D distracts by claiming fail to reject proves equal means, but it doesn't—it just lacks evidence of difference. Two-mean test conclusions should carefully state what the evidence supports regarding population means, avoiding causal or definitive equality statements like in choice E. Always compare p to α explicitly in your reasoning.

Question 15

A city compares mean commute time (minutes) for residents who take the bus versus residents who drive. Two independent random samples were selected (bus: n=55n=55; drive: n=50n=50). A two-sample tt test for μbusμdrive\mu_{bus}-\mu_{drive} used hypotheses H0:μbusμdrive=0H_0:\mu_{bus}-\mu_{drive}=0 and Ha:μbusμdrive0H_a:\mu_{bus}-\mu_{drive}\neq 0. The p-value was 0.004. Using α=0.01\alpha=0.01, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not sufficient evidence that the population mean commute times differ for bus riders and drivers.
  2. Reject H0H_0; there is sufficient evidence that the population mean commute times differ for bus riders and drivers. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that bus riders have a lower population mean commute time than drivers.
  4. Reject H0H_0; the sample means are different, so the population means must be different.
  5. Reject H0H_0; taking the bus causes a different commute time than driving.

Explanation: This question examines a two-sample t-test for differences in mean commute times between bus riders and drivers. The two-sided alternative allows for any difference, and with p=0.004 less than α=0.01, we reject H0, providing evidence of a population mean difference. Choice D is a distractor, incorrectly assuming different sample means automatically mean different population means without considering significance. For two-mean tests, conclusions should avoid causal claims like in choice E and focus on evidence for or against the null in the population context. Note that even with rejection, we don't specify direction unless the test is one-sided or further analysis is done.

Question 16

An engineer compares mean battery life (hours) for two brands of rechargeable batteries. Independent random samples were tested (Brand X: n=15n=15; Brand Y: n=17n=17). A two-sample tt test for μXμY\mu_X-\mu_Y was conducted with H0:μXμY=0H_0:\mu_X-\mu_Y=0 and Ha:μXμY>0H_a:\mu_X-\mu_Y>0. The p-value was 0.11. At α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Reject H0H_0; there is sufficient evidence that Brand X has a greater population mean battery life than Brand Y.
  2. Fail to reject H0H_0; there is not sufficient evidence that Brand X has a greater population mean battery life than Brand Y. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that Brand X has a smaller population mean battery life than Brand Y.
  4. Fail to reject H0H_0; this shows the two brands have equal population mean battery life.
  5. Fail to reject H0H_0; therefore Brand Y causes longer battery life.

Explanation: This question involves interpreting a two-sample t-test for comparing mean battery life between two brands. With a one-sided alternative for Brand X being greater and p-value 0.11 exceeding α=0.05, we fail to reject H0, meaning not enough evidence that X has longer life. Choice D distracts by saying fail to reject proves equal means, but it only indicates insufficient evidence against equality. In two-mean test conclusions, remember that failing to reject doesn't affirm the null or imply causation, as choice E wrongly suggests for Brand Y. Always reference the population means and the specific alternative hypothesis in your statement.

Question 17

A nutrition researcher compares mean systolic blood pressure for adults who follow Diet A versus Diet B. Two independent random samples were taken (Diet A: n=22n=22; Diet B: n=24n=24). A two-sample tt test for μAμB\mu_A-\mu_B was performed with H0:μAμB=0H_0:\mu_A-\mu_B=0 and Ha:μAμB0H_a:\mu_A-\mu_B\neq 0. The p-value was 0.62. At the 0.05 significance level, what conclusion is appropriate?

  1. Reject H0H_0; there is sufficient evidence that Diet A and Diet B have different population mean systolic blood pressures.
  2. Fail to reject H0H_0; there is not sufficient evidence of a difference in population mean systolic blood pressure between Diet A and Diet B. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that Diet A results in a lower population mean systolic blood pressure than Diet B.
  4. Fail to reject H0H_0; therefore the two diets have exactly the same mean systolic blood pressure in the population.
  5. Fail to reject H0H_0; this proves Diet B causes higher blood pressure.

Explanation: This question evaluates understanding of a two-sample t-test for comparing means of systolic blood pressure between two diets. The two-sided alternative tests for any difference, and with a p-value of 0.62 greater than α=0.05, we fail to reject the null hypothesis, indicating insufficient evidence of a population mean difference. Choice D is a distractor because it claims failing to reject proves exact equality in population means, but it only means we lack evidence to say they differ. In conclusions for two-mean tests, emphasize that 'fail to reject' doesn't confirm the null—there might still be a difference undetected by the sample. Avoid implying causation, as in choice E, since the test doesn't establish cause-and-effect relationships. Always tie the conclusion back to the population parameters, not just the samples.

Question 18

A sports scientist investigates whether mean vertical jump height differs between athletes who follow Program 1 and those who follow Program 2. Two independent random samples were taken (Program 1: n=12n=12; Program 2: n=14n=14). A two-sample tt test for μ1μ2\mu_1-\mu_2 was conducted with H0:μ1μ2=0H_0:\mu_1-\mu_2=0 and Ha:μ1μ20H_a:\mu_1-\mu_2\neq 0. The p-value was 0.049. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not sufficient evidence of a difference in population mean jump height between the two programs.
  2. Reject H0H_0; there is sufficient evidence of a difference in population mean jump height between the two programs. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that Program 1 produces a lower population mean jump height than Program 2.
  4. Reject H0H_0; therefore the two sample means are different.
  5. Reject H0H_0; Program 1 causes higher jump heights than Program 2 for all athletes.

Explanation: This question assesses interpreting a two-sample t-test for mean jump height differences between two programs. The two-sided test yields p=0.049 just under α=0.05, so we reject H0, indicating sufficient evidence of a population mean difference. A common distractor is choice D, which confuses sample mean differences with the hypothesis test outcome. In concluding two-mean tests, emphasize population-level inferences and avoid absolutes like 'for all athletes' in choice E, as results are probabilistic. Rejection supports the alternative but doesn't prove causation or direction without additional context.

Question 19

A gardener compares mean plant height after 8 weeks for plants given Fertilizer 1 versus Fertilizer 2. Two independent random samples were used (Fertilizer 1: n=20n=20; Fertilizer 2: n=20n=20). A two-sample tt test for μ1μ2\mu_1-\mu_2 was performed with H0:μ1μ2=0H_0:\mu_1-\mu_2=0 and Ha:μ1μ20H_a:\mu_1-\mu_2\neq 0. The p-value was 0.032. Using α=0.05\alpha=0.05, what conclusion is appropriate?

  1. Fail to reject H0H_0; there is not sufficient evidence that the population mean heights differ between the two fertilizers.
  2. Reject H0H_0; there is sufficient evidence that the population mean heights differ between the two fertilizers. (correct answer)
  3. Reject H0H_0; there is sufficient evidence that Fertilizer 1 produces a greater population mean height than Fertilizer 2.
  4. Reject H0H_0; the sample means are different, so the fertilizers must have different population means.
  5. Reject H0H_0; Fertilizer 1 causes plants to grow taller than Fertilizer 2.

Explanation: This question tests interpreting a two-sample t-test for mean plant heights with different fertilizers. The two-sided alternative and p=0.032 < α=0.05 lead to rejecting H0, showing evidence of a population mean difference. Choice D is a distractor, wrongly equating sample differences to automatic population differences without significance. In two-mean conclusions, avoid causal language like choice E and focus on evidence for the alternative in the population. Rejection indicates a difference but not its direction in a two-sided test unless specified.

Question 20

A nutritionist compares mean sodium intake (mg/day) for adults who follow Diet A versus Diet B. Independent random samples were taken from each diet group. A two-sample tt test was conducted for H0:μAμB=0H_0: \mu_A-\mu_B=0 versus Ha:μAμB0H_a: \mu_A-\mu_B\ne 0. The p-value was 0.41. At the 0.05 significance level, what conclusion is appropriate?

  1. Reject H0H_0; there is convincing evidence that mean sodium intake differs between Diet A and Diet B.
  2. Fail to reject H0H_0; there is not convincing evidence that mean sodium intake differs between Diet A and Diet B. (correct answer)
  3. Reject H0H_0; there is convincing evidence that Diet A has a lower mean sodium intake than Diet B.
  4. Fail to reject H0H_0; the sample means are close, so the population means must be equal.
  5. Fail to reject H0H_0; therefore Diet A and Diet B have exactly the same mean sodium intake in the populations.

Explanation: This problem involves a two-sample t-test with a two-tailed alternative hypothesis (μ_A - μ_B ≠ 0). The p-value of 0.41 is much larger than the significance level of 0.05, so we fail to reject the null hypothesis. This means there is not convincing evidence that mean sodium intake differs between Diet A and Diet B. Choice D incorrectly suggests that similar sample means guarantee equal population means, while Choice E wrongly claims that failing to reject H₀ proves the means are exactly equal. When we fail to reject H₀, we simply lack sufficient evidence to conclude the means differ; we cannot prove they are equal. Remember that hypothesis tests can only provide evidence against the null hypothesis, never proof that it's true.