College Statistics Quiz: T Test For Difference Of Means
20 questions · exam conditions
0:00
T Test For Difference Of MeansQuestion 1 of 20

A researcher conducts a two-tailed test for the difference in mean scores between Group 1 and Group 2. The hypotheses are H0:μ1μ2=0H_0: \mu_1 - \mu_2 = 0 and Ha:μ1μ20H_a: \mu_1 - \mu_2 \neq 0. The test yields a statistic of t=2.10t = 2.10 and a p-value of 0.04. If the researcher had instead defined the difference as μ2μ1\mu_2 - \mu_1 and tested the hypotheses H0:μ2μ1=0H_0: \mu_2 - \mu_1 = 0 and Ha:μ2μ10H_a: \mu_2 - \mu_1 \neq 0, what would the new test statistic and p-value be?

t = 2.10, p-value = 0.04
t = –2.10, p-value = 0.96
t = –2.10, p-value = 0.04
The results cannot be determined without the original data.
← Back to quizzes

College Statistics Quiz

College Statistics Quiz: T Test For Difference Of Means

Practice T Test For Difference Of Means in College Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on T Test For Difference Of Means, giving you a quick way to practice the rules, question types, and explanations that matter most for College Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher conducts a two-tailed test for the difference in mean scores between Group 1 and Group 2. The hypotheses are H0:μ1μ2=0H_0: \mu_1 - \mu_2 = 0 and Ha:μ1μ20H_a: \mu_1 - \mu_2 \neq 0. The test yields a statistic of t=2.10t = 2.10 and a p-value of 0.04. If the researcher had instead defined the difference as μ2μ1\mu_2 - \mu_1 and tested the hypotheses H0:μ2μ1=0H_0: \mu_2 - \mu_1 = 0 and Ha:μ2μ10H_a: \mu_2 - \mu_1 \neq 0, what would the new test statistic and p-value be?

  1. t = 2.10, p-value = 0.04
  2. t = –2.10, p-value = 0.96
  3. t = –2.10, p-value = 0.04 (correct answer)
  4. The results cannot be determined without the original data.
Explanation: The test statistic is calculated as t=(sample difference)(hypothesized difference)SEt = \frac{(\text{sample difference}) - (\text{hypothesized difference})}{SE}. When the order of subtraction is reversed from (Group 1 - Group 2) to (Group 2 - Group 1), the sample difference (xˉ1xˉ2\bar{x}_1 - \bar{x}_2) becomes (xˉ2xˉ1\bar{x}_2 - \bar{x}_1), which simply negates its value. Therefore, the new test statistic will be the negative of the original: t = –2.10. For a two-tailed test, the p-value is the probability of observing a test statistic as extreme or more extreme than the one calculated, in either direction (i.e., P(Tt)P(|T| \ge |t|)). Since 2.10=2.10|-2.10| = |2.10|, the p-value remains exactly the same.

Question 2

A pharmaceutical company develops a new drug to reduce blood pressure. A clinical trial is conducted where one group receives the new drug and a control group receives a placebo. Let μdrug\mu_{drug} be the true mean reduction in systolic blood pressure for the drug group and μplacebo\mu_{placebo} be the true mean reduction for the placebo group. The company will only market the drug if it is shown to be more effective than the placebo. The research team conducts a hypothesis test and obtains a p-value of 0.03.

  1. H0:μdrugμplacebo=0H_0: \mu_{drug} - \mu_{placebo} = 0; Ha:μdrugμplacebo0H_a: \mu_{drug} - \mu_{placebo} \neq 0. The result is significant, so the drug is more effective.
  2. H0:μdrugμplacebo=0H_0: \mu_{drug} - \mu_{placebo} = 0; Ha:μdrugμplacebo<0H_a: \mu_{drug} - \mu_{placebo} < 0. The result is significant, providing evidence the drug is less effective.
  3. H0:μdrugμplacebo=0H_0: \mu_{drug} - \mu_{placebo} = 0; Ha:μdrugμplacebo>0H_a: \mu_{drug} - \mu_{placebo} > 0. The result is not significant at α=0.05\alpha = 0.05.
  4. H0:μdrugμplacebo=0H_0: \mu_{drug} - \mu_{placebo} = 0; Ha:μdrugμplacebo>0H_a: \mu_{drug} - \mu_{placebo} > 0. The result is significant at α=0.05\alpha = 0.05, providing evidence the drug is more effective. (correct answer)
Explanation: The research question is whether the drug is more effective than the placebo, which translates to a one-tailed alternative hypothesis: Ha:μdrug>μplaceboH_a: \mu_{drug} > \mu_{placebo}, or Ha:μdrugμplacebo>0H_a: \mu_{drug} - \mu_{placebo} > 0. The null hypothesis is that there is no difference: H0:μdrugμplacebo=0H_0: \mu_{drug} - \mu_{placebo} = 0. The obtained p-value is 0.03. Since 0.03 is less than the standard significance level of α=0.05\alpha = 0.05, the null hypothesis is rejected in favor of the alternative. This provides statistically significant evidence that the drug is more effective than the placebo.

Question 3

A researcher conducts a two-tailed test for the difference in mean scores between Group 1 and Group 2. The hypotheses are H0:μ1μ2=0H_0: \mu_1 - \mu_2 = 0 and Ha:μ1μ20H_a: \mu_1 - \mu_2 \neq 0. The test yields a statistic of t=2.10t = 2.10 and a p-value of 0.04. If the researcher had instead defined the difference as μ2μ1\mu_2 - \mu_1 and tested the hypotheses H0:μ2μ1=0H_0: \mu_2 - \mu_1 = 0 and Ha:μ2μ10H_a: \mu_2 - \mu_1 \neq 0, what would the new test statistic and p-value be?

  1. t = 2.10, p-value = 0.04
  2. t = –2.10, p-value = 0.96
  3. t = –2.10, p-value = 0.04 (correct answer)
  4. The results cannot be determined without the original data.
Explanation: The test statistic is calculated as t=(sample difference)(hypothesized difference)SEt = \frac{(\text{sample difference}) - (\text{hypothesized difference})}{SE}. When the order of subtraction is reversed from (Group 1 - Group 2) to (Group 2 - Group 1), the sample difference (xˉ1xˉ2\bar{x}_1 - \bar{x}_2) becomes (xˉ2xˉ1\bar{x}_2 - \bar{x}_1), which simply negates its value. Therefore, the new test statistic will be the negative of the original: t = –2.10. For a two-tailed test, the p-value is the probability of observing a test statistic as extreme or more extreme than the one calculated, in either direction (i.e., P(Tt)P(|T| \ge |t|)). Since 2.10=2.10|-2.10| = |2.10|, the p-value remains exactly the same.

Question 4

Two different studies are conducted to compare the mean blood pressure reduction from two drugs, Drug A and Drug B. Summary statistics are as follows:

  • Study 1: nA=25,nB=25n_A = 25, n_B = 25; sample variances sA2=100,sB2=120s_A^2 = 100, s_B^2 = 120.
  • Study 2: nA=50,nB=50n_A = 50, n_B = 50; sample variances sA2=50,sB2=60s_A^2 = 50, s_B^2 = 60.

Assuming the difference in sample means (xˉAxˉB\bar{x}_A - \bar{x}_B) is the same for both studies, how would the width of a 95% confidence interval for μAμB\mu_A - \mu_B from Study 2 compare to the one from Study 1?

  1. The interval from Study 2 would be wider.
  2. The interval from Study 2 would be narrower. (correct answer)
  3. The interval widths would be approximately the same.
  4. The relationship cannot be determined without the sample means.
Explanation: The width of a confidence interval for the difference of two means is proportional to the standard error of the difference, SE=sA2nA+sB2nBSE = \sqrt{\frac{s_A^2}{n_A} + \frac{s_B^2}{n_B}}. For Study 1, SE1=10025+12025=4+4.8=8.82.97SE_1 = \sqrt{\frac{100}{25} + \frac{120}{25}} = \sqrt{4+4.8} = \sqrt{8.8} \approx 2.97. For Study 2, SE2=5050+6050=1+1.2=2.21.48SE_2 = \sqrt{\frac{50}{50} + \frac{60}{50}} = \sqrt{1+1.2} = \sqrt{2.2} \approx 1.48. Study 2 has both larger sample sizes and smaller sample variances, both of which contribute to a smaller standard error. A smaller standard error results in a narrower confidence interval.

Question 5

A researcher is comparing two groups and is deciding whether to use a pooled or unpooled (Welch's) two-sample t-test. The sample sizes are n1=15n_1 = 15 and n2=18n_2 = 18. The sample standard deviations are s1=10.5s_1 = 10.5 and s2=11.2s_2 = 11.2. A common rule of thumb suggests that if the larger sample standard deviation is not more than twice the smaller one, one can assume equal variances.

  1. The pooled t-test must be used because the rule of thumb is satisfied and it is always more powerful.
  2. The unpooled (Welch's) t-test must be used because the sample standard deviations are not identical.
  3. The unpooled (Welch's) t-test is generally preferred as it is more robust to violations of the equal variance assumption. (correct answer)
  4. Neither test is appropriate because the sample sizes are small.
Explanation: While the rule of thumb (11.2 / 10.5 ≈ 1.07 < 2) is satisfied, modern statistical practice recommends using the unpooled (Welch's) t-test as the default choice. It performs nearly as well as the pooled test when the population variances are equal and performs much better (i.e., maintains the correct Type I error rate) when they are not. Since one rarely knows if the population variances are truly equal, the safer and more robust option is the unpooled test. Thus, it is generally preferred.

Question 6

A researcher compares the effectiveness of two different fertilizers on crop yield. They test Fertilizer A on 10 plots and Fertilizer B on 12 plots. Boxplots of the yields for each group show significant skewness and several outliers. The sample standard deviations are 25.4 and 28.1, respectively.

  1. The sample sizes are unequal (10 versus 12).
  2. The sample standard deviations are not identical.
  3. The presence of significant skewness and outliers with small sample sizes. (correct answer)
  4. The experimental plots were not randomly assigned to the fertilizers.
Explanation: The validity of a two-sample t-test relies on three main assumptions: (1) independent samples, (2) normal populations or large enough sample sizes for the Central Limit Theorem to apply, and (3) equal population variances (for the pooled test). The most significant issue here is the violation of the normality assumption. The presence of significant skewness and outliers in small samples (n < 30) means the sampling distribution of the difference in means may not be well-approximated by a t-distribution, making the calculated p-value unreliable.

Question 7

A 95% confidence interval for the difference in mean test scores between two teaching methods (Method A minus Method B) is calculated to be (–0.5, 4.5). A researcher wants to use this information to conduct a two-tailed hypothesis test of H0:μAμB=0H_0: \mu_A - \mu_B = 0 versus Ha:μAμB0H_a: \mu_A - \mu_B \neq 0.

  1. The null hypothesis would be rejected because the interval contains positive values.
  2. The null hypothesis would not be rejected, and the p-value for the test would be greater than 0.05. (correct answer)
  3. The outcome of the test cannot be determined from the confidence interval alone.
  4. The null hypothesis would be rejected because the interval is not symmetric around 0.
Explanation: There is a direct correspondence between a two-sided hypothesis test and a confidence interval. The null hypothesis H0:μAμB=0H_0: \mu_A - \mu_B = 0 is rejected at significance level α\alpha if and only if the value 0 is not contained within the (1α)100%(1-\alpha)100\% confidence interval for the difference. Since the 95% confidence interval (–0.5, 4.5) contains 0, the null hypothesis would not be rejected at the α=0.05\alpha = 0.05 level. This implies that the corresponding p-value must be greater than 0.05.

Question 8

Two independent-sample t-tests are performed at the same alpha level (α=0.05\alpha = 0.05, two-tailed).

  • Test 1: Compares a group of size 10 to a group of size 12.
  • Test 2: Compares a group of size 50 to a group of size 60.

Let t1t_1^* be the critical value for Test 1 and t2t_2^* be the critical value for Test 2. Which of the following statements is true?

  1. t1>t2t_1^* > t_2^* (correct answer)
  2. t1=t2t_1^* = t_2^*
  3. t1<t2t_1^* < t_2^*
  4. The relationship depends on the sample variances.
Explanation: When you encounter questions about critical values in t-tests, focus on how degrees of freedom affect the t-distribution. The key insight is that as sample sizes increase, the t-distribution approaches the standard normal distribution, making critical values smaller. For independent-sample t-tests, the degrees of freedom equal df=n1+n22df = n_1 + n_2 - 2. Test 1 has df=10+122=20df = 10 + 12 - 2 = 20, while Test 2 has df=50+602=108df = 50 + 60 - 2 = 108. Since both tests use α=0.05\alpha = 0.05 (two-tailed), you're looking for the critical values that cut off 2.5% in each tail. The t-distribution has heavier tails than the normal distribution, but these tails get lighter as degrees of freedom increase. With fewer degrees of freedom (Test 1), you need a larger critical value to capture the same area in the tails. With more degrees of freedom (Test 2), the distribution is closer to normal, so the critical value is smaller. Therefore, t1>t2t_1^* > t_2^*, making choice A correct. Choice B is wrong because critical values definitely change with different degrees of freedom. Choice C reverses the relationship—smaller sample sizes actually require larger critical values, not smaller ones. Choice D incorrectly suggests that sample variances affect critical values, but critical values depend only on the significance level and degrees of freedom, not on the actual data variances. Remember: smaller samples mean fewer degrees of freedom, which means larger critical values needed for significance. This makes it harder to reject the null hypothesis with small samples.

Question 9

A study compared the mean recovery times for two surgical procedures. The resulting p-value for a two-sample t-test was 0.25. The lead surgeon concluded, 'This p-value proves that the two procedures have the same mean recovery time.'

  1. The conclusion is flawed because a p-value of 0.25 is small and suggests a difference.
  2. The conclusion is flawed because failing to find evidence of a difference is not proof of no difference. (correct answer)
  3. The conclusion is correct because a large p-value means the null hypothesis is true.
  4. The conclusion is flawed because the surgeon should have used a one-tailed test.
Explanation: This is a common misinterpretation of hypothesis testing. A large p-value (typically > 0.05) leads to a failure to reject the null hypothesis. This does not mean the null hypothesis is true or has been proven. It simply means that the sample data did not provide sufficient evidence to conclude that a difference exists. This is analogous to a 'not guilty' verdict in a trial: it doesn't prove innocence, only that there was not enough evidence for a conviction. The phrase 'absence of evidence is not evidence of absence' applies here.

Question 10

A researcher wants to determine if a new fuel additive improves gas mileage. Which of the following experimental designs is most appropriate for analysis using an independent-samples t-test for the difference in means?

  1. Select 25 cars. Measure their mileage, add the additive, and then measure their mileage again.
  2. Randomly select 50 cars, assign 25 to use the additive and 25 to not use it, then compare the two groups' mean mileage. (correct answer)
  3. Select 25 pairs of identical twin car models. For each pair, randomly assign one to use the additive and one to not.
  4. Select 50 cars. Have each car drive one week with the additive and one week without, in a random order.
Explanation: An independent-samples t-test is appropriate when the two groups being compared are independent of each other. In choice B, the 50 cars are randomly assigned to two separate groups (treatment and control), and these groups are independent. Choices A and D describe repeated measures (or crossover) designs, where the same subjects (cars) are measured under both conditions. Choice C describes a matched-pairs design. These three designs (A, C, D) create dependent samples and must be analyzed with a paired-samples t-test, not an independent-samples t-test.

Question 11

A researcher wants to determine if a new fuel additive improves gas mileage. Which of the following experimental designs is most appropriate for analysis using an independent-samples t-test for the difference in means?

  1. Select 25 cars. Measure their mileage, add the additive, and then measure their mileage again.
  2. Randomly select 50 cars, assign 25 to use the additive and 25 to not use it, then compare the two groups' mean mileage. (correct answer)
  3. Select 25 pairs of identical twin car models. For each pair, randomly assign one to use the additive and one to not.
  4. Select 50 cars. Have each car drive one week with the additive and one week without, in a random order.
Explanation: An independent-samples t-test is appropriate when the two groups being compared are independent of each other. In choice B, the 50 cars are randomly assigned to two separate groups (treatment and control), and these groups are independent. Choices A and D describe repeated measures (or crossover) designs, where the same subjects (cars) are measured under both conditions. Choice C describes a matched-pairs design. These three designs (A, C, D) create dependent samples and must be analyzed with a paired-samples t-test, not an independent-samples t-test.

Question 12

A sociologist is studying the average number of hours of television watched per week by adults in two cities. Summary statistics are collected:

  • City A: nA=40,xˉA=15.5,sA=4.0n_A = 40, \bar{x}_A = 15.5, s_A = 4.0
  • City B: nB=50,xˉB=17.5,sB=5.0n_B = 50, \bar{x}_B = 17.5, s_B = 5.0

Assuming the conditions for inference are met, what is the value of the unpooled test statistic for testing H0:μA=μBH_0: \mu_A = \mu_B?

  1. –4.47
  2. –2.06
  3. –1.49
  4. –2.11 (correct answer)
Explanation: The unpooled (Welch's) t-test statistic is calculated as t=(xˉAxˉB)0sA2nA+sB2nBt = \frac{(\bar{x}_A - \bar{x}_B) - 0}{\sqrt{\frac{s_A^2}{n_A} + \frac{s_B^2}{n_B}}}. Plugging in the values: t=15.517.54.0240+5.0250=21640+2550=20.4+0.5=20.92.11t = \frac{15.5 - 17.5}{\sqrt{\frac{4.0^2}{40} + \frac{5.0^2}{50}}} = \frac{-2}{\sqrt{\frac{16}{40} + \frac{25}{50}}} = \frac{-2}{\sqrt{0.4 + 0.5}} = \frac{-2}{\sqrt{0.9}} \approx -2.11. Distractor A results from forgetting to square the standard deviations. Distractor B is the result of using the pooled standard error. Distractor C results from incorrectly adding the standard errors of the mean.

Question 13

A researcher conducts a two-sample t-test to compare the mean effectiveness of two therapies and obtains a p-value of 0.08. Since the result is not significant at α=0.05\alpha = 0.05, they fail to reject the null hypothesis of no difference. The researcher suspects that a true difference exists, but the study lacked sufficient power. Which of the following actions, if taken in a new study, would be most likely to increase the power of the test to detect a true difference of the same magnitude?

  1. Decrease the significance level α\alpha to 0.01.
  2. Conduct a two-tailed test instead of a one-tailed test.
  3. Increase the sample sizes for both groups. (correct answer)
  4. Ensure the sample sizes are exactly equal in both groups, keeping the total sample size the same.
Explanation: Statistical power is the probability of correctly rejecting a false null hypothesis. Power is increased by four main factors: increasing the sample size, increasing the significance level (α\alpha), increasing the true effect size, and decreasing the population variance. Of the choices given, increasing the sample sizes is a standard and effective method for a researcher to increase the power of a new study. Decreasing α\alpha (A) would decrease power. Using a two-tailed test instead of a correctly specified one-tailed test (B) would decrease power. Equalizing sample sizes (D) can be beneficial, but increasing the total sample size has a much larger and more certain impact on power.

Question 14

A study compares the salaries of 20 randomly selected employees from a tech company and 20 from a manufacturing company. A two-sample t-test is performed. Later, it is discovered that the highest salary in the tech company sample, which was an extreme outlier, was recorded incorrectly and should be removed. The sample mean for the tech company was much higher than for the manufacturing company.

  1. The t-statistic will increase because the removal of the outlier improves the normality of the data.
  2. The sample mean for the tech company will decrease, but the sample standard deviation will increase.
  3. The degrees of freedom will decrease, causing the p-value to decrease significantly.
  4. The sample mean and standard deviation for the tech company will both decrease, likely leading to a smaller t-statistic. (correct answer)
Explanation: Removing an extreme high outlier from a sample will have two primary effects: it will decrease the sample mean (xˉ\bar{x}) and it will decrease the sample standard deviation (ss), as the data will be less spread out. The t-statistic is calculated as (xˉ1xˉ2)/SE(\bar{x}_1 - \bar{x}_2) / SE. Both the numerator (the difference in means) and the denominator (which depends on the standard deviations) will likely decrease. However, the outlier's effect on the mean is typically more pronounced. The reduction in the difference of means (numerator) will generally be larger than the reduction in the standard error (denominator), leading to a smaller t-statistic and, consequently, a larger p-value.

Question 15

A researcher tests if a new fertilizer (F) increases plant height compared to an old fertilizer (O). The hypotheses are H0:μFμO=0H_0: \mu_F - \mu_O = 0 and Ha:μFμO>0H_a: \mu_F - \mu_O > 0. The data yields nF=30,xˉF=55n_F = 30, \bar{x}_F = 55 cm and nO=30,xˉO=57n_O = 30, \bar{x}_O = 57 cm. The resulting test statistic is t=2.50t = -2.50.

  1. Since |–2.50| is large, the null hypothesis should be rejected in favor of the alternative hypothesis.
  2. An error was likely made in the calculation, as the test statistic cannot be negative when testing for an increase.
  3. The p-value is the area to the left of –2.50, which is very small, providing strong evidence for the alternative hypothesis.
  4. The sample data are in the opposite direction of the alternative hypothesis, so the p-value will be large. (correct answer)
Explanation: The alternative hypothesis Ha:μFμO>0H_a: \mu_F - \mu_O > 0 suggests we are looking for evidence that the new fertilizer produces taller plants. However, the sample data show the opposite: the mean height for the new fertilizer (55 cm) is less than the mean height for the old fertilizer (57 cm). This leads to a negative test statistic (t = -2.50). For a right-tailed test, a negative t-statistic provides no evidence in favor of the alternative hypothesis. The p-value is the probability of observing a result as extreme or more extreme in the direction of the alternative hypothesis. Thus, the p-value would be the area to the right of t = -2.50, which is very large (close to 1).

Question 16

In a study comparing two groups with n1=10n_1 = 10 and n2=10n_2 = 10, the sample standard deviations were s1=5.0s_1 = 5.0 and s2=7.0s_2 = 7.0. Assuming the population variances are equal, the researcher calculates a pooled standard deviation, sps_p. Which of the following must be true about the value of sps_p?

  1. Its value will be between 5.0 and 7.0. (correct answer)
  2. It will be larger than 7.0 because variability is combined.
  3. It will be equal to 6.0, the average of the two standard deviations.
  4. It will be closer to 7.0 because the larger standard deviation has more influence.
Explanation: When you encounter problems involving pooled standard deviation, you're dealing with a way to combine variability measures from two groups under the assumption of equal population variances. The pooled standard deviation creates a weighted average that accounts for both sample sizes and variances. The pooled standard deviation formula is: sp=(n11)s12+(n21)s22n1+n22s_p = \sqrt{\frac{(n_1-1)s_1^2 + (n_2-1)s_2^2}{n_1+n_2-2}} Let's calculate: sp=(101)(5.0)2+(101)(7.0)210+102=9(25)+9(49)18=225+44118=37=6.08s_p = \sqrt{\frac{(10-1)(5.0)^2 + (10-1)(7.0)^2}{10+10-2}} = \sqrt{\frac{9(25) + 9(49)}{18}} = \sqrt{\frac{225 + 441}{18}} = \sqrt{37} = 6.08 Answer A is correct because the pooled standard deviation will always fall between the individual sample standard deviations when sample sizes are equal. This makes intuitive sense—you're combining information from both groups, creating a compromise value. Answer B is wrong because pooling doesn't simply add variability; it creates a weighted average that falls between the original values. Answer C incorrectly assumes the pooled standard deviation equals the simple arithmetic mean of the two standard deviations (6.0). The pooling formula weights by degrees of freedom, not just averages. Answer D is incorrect because with equal sample sizes, neither standard deviation has more influence—they're weighted equally. Remember: pooled standard deviation with equal sample sizes always falls between the individual standard deviations. The exact value depends on the specific numbers, but it won't exceed the bounds set by the original values.

Question 17

A company claims its new machine can process widgets at a rate that is, on average, more than 5 widgets per hour faster than the old machine. Let μnew\mu_{new} and μold\mu_{old} be the mean processing rates. Summary data from trials are:

  • New Machine: n1=20,xˉ1=58,s1=4n_1 = 20, \bar{x}_1 = 58, s_1 = 4
  • Old Machine: n2=20,xˉ2=51,s2=3n_2 = 20, \bar{x}_2 = 51, s_2 = 3

To test the company's claim, what are the appropriate null and alternative hypotheses and the value of the test statistic?

  1. H0:μnewμold=0H_0: \mu_{new} - \mu_{old} = 0; Ha:μnewμold>0H_a: \mu_{new} - \mu_{old} > 0; t ≈ 6.26
  2. H0:μnewμold=5H_0: \mu_{new} - \mu_{old} = 5; Ha:μnewμold>5H_a: \mu_{new} - \mu_{old} > 5; t ≈ 1.79 (correct answer)
  3. H0:μnewμold=5H_0: \mu_{new} - \mu_{old} = 5; Ha:μnewμold>5H_a: \mu_{new} - \mu_{old} > 5; t ≈ –1.79
  4. H0:μnewμold=0H_0: \mu_{new} - \mu_{old} = 0; Ha:μnewμold>5H_a: \mu_{new} - \mu_{old} > 5; t ≈ 1.79
Explanation: The company's claim is that the new machine is more than 5 widgets per hour faster, so the alternative hypothesis is Ha:μnewμold>5H_a: \mu_{new} - \mu_{old} > 5. The corresponding null hypothesis is H0:μnewμold=5H_0: \mu_{new} - \mu_{old} = 5. The test statistic for a non-zero hypothesized difference is t=(xˉ1xˉ2)(μ1μ2)0s12n1+s22n2t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)_0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}. Plugging in the values: t=(5851)54220+3220=751620+920=22520=21.251.79t = \frac{(58 - 51) - 5}{\sqrt{\frac{4^2}{20} + \frac{3^2}{20}}} = \frac{7-5}{\sqrt{\frac{16}{20} + \frac{9}{20}}} = \frac{2}{\sqrt{\frac{25}{20}}} = \frac{2}{\sqrt{1.25}} \approx 1.79.

Question 18

An ecologist measures the shell diameter of a specific snail species in two different coastal areas, Area A (sandy) and Area B (rocky). They collect 15 snails from Area A and 40 snails from Area B. The data from Area A show a roughly symmetric, mound-shaped distribution. The data from Area B are moderately right-skewed. The ecologist is concerned about the different sample sizes and the skewness in the Area B data.

  1. Proceed with an unpooled (Welch's) t-test, as the test is robust due to the large sample size in the skewed group. (correct answer)
  2. Do not perform a t-test because the normality assumption is violated for Area B.
  3. Use a pooled two-sample t-test because the conditions are approximately met.
  4. The data cannot be analyzed without equal sample sizes.
Explanation: When you encounter questions about comparing means between two groups with different sample sizes and distributions, focus on the conditions required for two-sample t-tests and how robust these tests are to violations. The key insight here is understanding when t-tests remain valid despite imperfect conditions. With Area B having 40 snails, this large sample size makes the t-test robust to moderate skewness due to the Central Limit Theorem. The sampling distribution of the mean becomes approximately normal even when the population isn't perfectly normal. Since the sample sizes are unequal (15 vs 40), an unpooled (Welch's) t-test is appropriate because it doesn't assume equal variances. Answer A correctly identifies this approach. Answer B is too restrictive—it assumes any deviation from normality invalidates the t-test. While normality is an assumption, the test is robust to violations when sample sizes are reasonably large (typically n ≥ 30). Answer C suggests a pooled t-test, but this assumes equal variances between groups, which is questionable given the different environments and distributions. The unequal sample sizes also make pooling less appropriate. Answer D reflects a common misconception—t-tests don't require equal sample sizes. Welch's t-test was specifically designed to handle unequal sample sizes and variances. Remember this pattern: For two-sample comparisons, large sample sizes (especially n ≥ 30) allow t-tests to work well even with moderate skewness. Always choose unpooled methods when sample sizes or variances differ significantly between groups.

Question 19

A sociologist is studying the average number of hours of television watched per week by adults in two cities. Summary statistics are collected:

  • City A: nA=40,xˉA=15.5,sA=4.0n_A = 40, \bar{x}_A = 15.5, s_A = 4.0
  • City B: nB=50,xˉB=17.5,sB=5.0n_B = 50, \bar{x}_B = 17.5, s_B = 5.0

Assuming the conditions for inference are met, what is the value of the unpooled test statistic for testing H0:μA=μBH_0: \mu_A = \mu_B?

  1. –4.47
  2. –2.06
  3. –1.49
  4. –2.11 (correct answer)
Explanation: The unpooled (Welch's) t-test statistic is calculated as t=(xˉAxˉB)0sA2nA+sB2nBt = \frac{(\bar{x}_A - \bar{x}_B) - 0}{\sqrt{\frac{s_A^2}{n_A} + \frac{s_B^2}{n_B}}}. Plugging in the values: t=15.517.54.0240+5.0250=21640+2550=20.4+0.5=20.92.11t = \frac{15.5 - 17.5}{\sqrt{\frac{4.0^2}{40} + \frac{5.0^2}{50}}} = \frac{-2}{\sqrt{\frac{16}{40} + \frac{25}{50}}} = \frac{-2}{\sqrt{0.4 + 0.5}} = \frac{-2}{\sqrt{0.9}} \approx -2.11. Distractor A results from forgetting to square the standard deviations. Distractor B is the result of using the pooled standard error. Distractor C results from incorrectly adding the standard errors of the mean.

Question 20

A large-scale educational study involving 10,000 students in Group A and 10,000 students in Group B compared two online learning platforms. The study found that the mean final exam score for Group A was 85.2 and for Group B was 85.6. A two-sample t-test for the difference in means yielded a p-value of 0.001.

  1. Since the p-value is very small, the platform for Group B is substantially better and should be adopted.
  2. The result is unreliable because the large sample size artificially deflated the p-value.
  3. The difference in mean scores is statistically significant, but it may not be practically significant. (correct answer)
  4. There is no meaningful difference between the two platforms because the means are very close.
Explanation: This question tests the distinction between statistical significance and practical significance. A very small p-value (like 0.001) indicates that the observed difference is unlikely to be due to random chance, so the result is statistically significant. However, the magnitude of the difference in means (85.6 - 85.2 = 0.4 points) is very small. In a practical context, a 0.4-point difference on an exam may be considered negligible. Large sample sizes provide high power to detect even very small true differences, which may not be meaningful in a real-world context.