College Statistics Quiz: Power
20 questions · exam conditions
0:00
PowerQuestion 1 of 20

A materials scientist is testing two new alloys. Alloy A is expected to have a dramatically higher tensile strength than the current standard (a large effect size). Alloy B is expected to have only a marginal improvement (a small effect size). If two separate hypothesis tests are conducted with the same sample size and significance level, how will the power of the test for Alloy A compare to the power of the test for Alloy B?

The power will be lower for Alloy A because a larger effect is more likely to be due to random chance.
The power will be higher for Alloy A because a larger true difference between the null and alternative hypotheses is easier to detect.
The power will be identical for both tests because the sample size and significance level are the same.
It is impossible to compare the power without knowing the p-values from the tests.
← Back to quizzes

College Statistics Quiz

College Statistics Quiz: Power

Practice Power in College Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Power, giving you a quick way to practice the rules, question types, and explanations that matter most for College Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A materials scientist is testing two new alloys. Alloy A is expected to have a dramatically higher tensile strength than the current standard (a large effect size). Alloy B is expected to have only a marginal improvement (a small effect size). If two separate hypothesis tests are conducted with the same sample size and significance level, how will the power of the test for Alloy A compare to the power of the test for Alloy B?

  1. The power will be lower for Alloy A because a larger effect is more likely to be due to random chance.
  2. The power will be higher for Alloy A because a larger true difference between the null and alternative hypotheses is easier to detect. (correct answer)
  3. The power will be identical for both tests because the sample size and significance level are the same.
  4. It is impossible to compare the power without knowing the p-values from the tests.
Explanation: Power is the probability of detecting a true effect. The 'effect size' is the magnitude of this effect. A larger effect size means there is a greater true difference between the parameter value under the null hypothesis and the true parameter value. It is intuitively and mathematically easier to detect a large difference than a small one. Therefore, holding all else constant (n, α\alpha, σ\sigma), the test for Alloy A (large effect size) will have substantially higher power than the test for Alloy B (small effect size). Distractor A is a misconception. Distractor C is incorrect because it ignores the crucial role of effect size in determining power. Distractor D is incorrect because power is a pre-study characteristic of a test, calculated based on assumed parameters, not a post-study result like a p-value.

Question 2

A researcher simultaneously decreases the desired significance level α\alpha from 0.05 to 0.01 and increases the sample size nn fourfold. What is the net effect of these two changes on the power of the hypothesis test?

  1. The power will definitely increase. (correct answer)
  2. The power will definitely decrease.
  3. The power will remain exactly the same.
  4. The net effect cannot be determined without more information on the effect size and variance.
Explanation: This question involves assessing the impact of two competing changes. Decreasing α\alpha from 0.05 to 0.01 makes the test more conservative and, by itself, decreases power. However, increasing the sample size nn fourfold makes the test much more sensitive and, by itself, increases power. The effect of sample size on power is generally much more substantial than the effect of changing α\alpha in this range. A fourfold increase in nn halves the standard error (since SE 1/n\propto 1/\sqrt{n}), which typically leads to a very large increase in power that will more than compensate for the modest decrease in power from the change in α\alpha. Thus, the power will definitely increase. While the exact amount of increase depends on other factors (Choice D), the direction of the net change is clear.

Question 3

An education researcher is investigating the effect of a new teaching method on test scores. The budget for the study is fixed, so the sample size cannot be increased. Which of the following adjustments to the study protocol is most likely to increase the statistical power of the test?

  1. Lowering the significance level (α\alpha) to reduce the chance of a Type I error.
  2. Implementing a more standardized test administration procedure to reduce random measurement error. (correct answer)
  3. Switching from a one-tailed to a two-tailed hypothesis test to capture effects in either direction.
  4. Testing for a smaller effect size, as smaller effects are more common and thus easier to find.
Explanation: This is a multi-step reasoning question. Statistical power is increased by reducing the variability (σ\sigma) of the outcome measure. Random measurement error is a component of this overall variability. By implementing a more standardized test administration procedure, the researcher reduces the 'noise' or random error in the collected test scores. This effectively lowers the standard deviation of the measured scores, which in turn increases the power of the test, even if the sample size is fixed. Distractor A is incorrect; lowering α\alpha decreases power. Distractor C is incorrect; a one-tailed test is more powerful than a two-tailed test for detecting an effect in the specified direction. Distractor D is incorrect; tests have less power to detect smaller effect sizes.

Question 4

A researcher achieves a power of 0.95 in a hypothesis test with a sample of size n=100n=100 and a significance level of α=0.05\alpha=0.05. Which of the following changes, if made before conducting the study, would have resulted in a decrease in statistical power?

  1. Increasing the sample size to n=200n=200.
  2. Changing the significance level to α=0.10\alpha=0.10.
  3. Decreasing the sample size to n=50n=50. (correct answer)
  4. Investigating a larger effect size while keeping nn and α\alpha the same.
Explanation: The question asks which change would decrease power. Let's analyze the options: A) Increasing the sample size increases power. B) Increasing the significance level α\alpha (from 0.05 to 0.10) makes it easier to reject H0H_0, thus increasing power. D) Investigating a larger effect size makes the effect easier to detect, thus increasing power. C) Decreasing the sample size (from 100 to 50) increases the standard error, making the test less sensitive and thus decreasing power. Therefore, decreasing the sample size is the change that would have led to lower power.

Question 5

A study fails to reject the null hypothesis. A subsequent analysis reveals that the statistical power of the study was only 0.20 for the minimum effect size considered practically important. What is the most valid conclusion to draw from these findings?

  1. The null hypothesis is very likely to be true.
  2. There is strong evidence to support the null hypothesis.
  3. The study was inconclusive because it had an inadequate chance of detecting a meaningful effect if one existed. (correct answer)
  4. A calculation error must have occurred, because power must be at least 0.50 for a study to be valid.
Explanation: Failing to reject the null hypothesis in a low-power study is an ambiguous result. A power of 0.20 means that even if a meaningful effect truly existed, the study only had a 20% chance of detecting it. Therefore, the non-significant result could be due to the absence of an effect OR the study's inability to detect the effect. This makes the study inconclusive. Distractors A and B are too strong; 'absence of evidence is not evidence of absence', especially in an underpowered study. Distractor D is incorrect; while a conventional target for power is 0.80, there is no rule that a study with power less than 0.50 is invalid, only that it is likely to be uninformative.

Question 6

A cognitive scientist has a strong theoretical reason to hypothesize that a specific memory-enhancing drug will improve recall scores, and has no reason to believe it could harm them. In designing an experiment to test this hypothesis against a placebo, which choice regarding the hypothesis test structure and its corresponding rationale is most appropriate for maximizing the ability to detect the anticipated effect?

  1. A two-tailed test should be used because it is more conservative and reduces the risk of a Type I error, which increases power.
  2. A one-tailed test should be used because it allocates the entire rejection region (α\alpha) to one side of the distribution, making the test more sensitive to an effect in the hypothesized direction. (correct answer)
  3. A two-tailed test should be used because it has greater power than a one-tailed test to detect any deviation from the null hypothesis, regardless of direction.
  4. A one-tailed test should be used, but the significance level α\alpha must be halved to maintain the same power as a two-tailed test.
Explanation: When there is a strong directional hypothesis (e.g., improvement only), a one-tailed test is more powerful than a two-tailed test for detecting an effect in that specific direction. This is because the entire significance level α\alpha is placed in one tail. For a given α\alpha, this results in a critical value that is less extreme than the critical values for a two-tailed test, making it easier to reject the null hypothesis if the effect is indeed in the predicted direction. Answer A is incorrect because being more conservative against Type I error (using a two-tailed test when a one-tailed is justified) decreases power. Answer C is incorrect; a one-tailed test is more powerful for the specified direction. Answer D is a misunderstanding of the relationship between test type and power.

Question 7

Before launching a large-scale study, a sociologist performs a power analysis and determines that the planned experiment has a power of 0.80 to detect the effect of interest. Which of the following is the correct interpretation of this result?

  1. There is an 80% probability that the alternative hypothesis is true.
  2. There is an 80% probability that the study's conclusions will be correct.
  3. If the study yields a statistically significant result, there is an 80% probability that the alternative hypothesis is true.
  4. If the alternative hypothesis is true, there is an 80% probability that the study will yield a statistically significant result. (correct answer)
Explanation: When you encounter power analysis questions, focus on understanding what statistical power actually measures: the probability of correctly rejecting a false null hypothesis. Power of 0.80 means that if the alternative hypothesis is truly correct in the population, there's an 80% chance your study will detect this effect and yield a statistically significant result. Think of it as the test's sensitivity - how good it is at finding real effects when they exist. Answer D captures this definition perfectly: "If the alternative hypothesis is true, there is an 80% probability that the study will yield a statistically significant result." This conditional probability statement (given Ha is true) is exactly what power measures. Answer A incorrectly suggests power tells us the probability that Ha is true. Power doesn't determine whether hypotheses are true - it assumes Ha is true and calculates detection probability. Answer B is too broad, claiming 80% probability of correct conclusions overall. Power only addresses one type of correct decision (detecting true effects), not general accuracy including correctly accepting true null hypotheses. Answer C confuses power with the complement of Type I error or relates to positive predictive value. This describes the probability Ha is true given a significant result, which depends on prior probabilities and isn't what power analysis calculates. Remember: Power always assumes the alternative hypothesis is true as a starting point. It answers "If there really is an effect, how likely am I to find it?" This conditional thinking is crucial for interpreting power analysis results correctly.

Question 8

A pharmaceutical company is conducting a clinical trial for a new drug. The U.S. Food and Drug Administration (FDA) requires the company to demonstrate the drug's efficacy with a significance level of α=0.01\alpha = 0.01 instead of the more common α=0.05\alpha = 0.05. Assuming the sample size and the true effect size of the drug remain fixed, what is the primary consequence of the FDA's requirement on the statistical power of the trial?

  1. The power of the trial will increase because the stricter significance level makes the test more rigorous.
  2. The power of the trial will decrease because the critical value for the test statistic becomes more extreme, making it harder to reject the null hypothesis. (correct answer)
  3. The power of the trial will remain unchanged because power is only dependent on sample size and effect size, not the significance level.
  4. The power of the trial will be directly proportional to the change in α\alpha, resulting in a five-fold decrease in power.
Explanation: Power is the probability of correctly rejecting a false null hypothesis (H0H_0). Lowering the significance level α\alpha from 0.05 to 0.01 means the threshold for statistical significance becomes more stringent. This makes the rejection region for H0H_0 smaller. Consequently, for any given true effect, it is more difficult to obtain a test statistic that falls within this smaller rejection region. This increases the probability of a Type II error (β\beta), which is the failure to reject a false H0H_0. Since power is 1β1 - \beta, an increase in β\beta leads to a decrease in power. Distractor A is incorrect because 'rigor' in the sense of reducing Type I error comes at the cost of power. Distractor C is incorrect because power is fundamentally linked to α\alpha, nn, and effect size. Distractor D incorrectly suggests a simple proportional relationship, which is not the case; the relationship is more complex.

Question 9

A team of psychologists conducted a pilot study with n=30n=30 participants to test a new intervention and found a non-significant result (p = 0.25). They suspect the intervention has a small but meaningful effect. They are now designing a larger, definitive study and want to maximize their chances of correctly identifying this effect if it truly exists. Which of the following strategies provides the most direct and substantial increase in statistical power, assuming all other factors are held constant?

  1. Increasing the sample size from n=30n=30 to n=300n=300. (correct answer)
  2. Changing the significance level from α=0.05\alpha = 0.05 to α=0.01\alpha = 0.01.
  3. Recruiting a more diverse and heterogeneous sample of participants.
  4. Using a two-tailed test instead of the one-tailed test used in the pilot study.
Explanation: Statistical power is heavily dependent on sample size. Increasing the sample size (n) reduces the standard error of the test statistic. This results in a narrower sampling distribution, which makes it easier to distinguish the alternative hypothesis distribution from the null hypothesis distribution, thus increasing the probability of correctly rejecting a false null hypothesis. Answer B is incorrect because decreasing α\alpha would decrease power. Answer C is incorrect because a more heterogeneous sample would likely increase population variance (σ2\sigma^2), which would decrease power. Answer D is incorrect because, for a fixed α\alpha and an effect in the expected direction, a one-tailed test is more powerful than a two-tailed test; switching to a two-tailed test would decrease power.

Question 10

A researcher is planning an experiment and can choose to draw samples from one of two populations. Population 1 has a known standard deviation of σ1=10\sigma_1 = 10. Population 2 is more homogeneous and has a known standard deviation of σ2=5\sigma_2 = 5. To test a hypothesis about the population mean, the researcher plans to use the same sample size nn and significance level α\alpha for either population. To maximize the power of the test, which population should be chosen and why?

  1. Population 1, because higher variability provides a greater chance of observing an extreme and significant result.
  2. Population 2, because lower variability reduces the standard error of the mean, making it easier to detect a true effect. (correct answer)
  3. Either population, as the standard deviation does not affect the power of a test concerning the mean.
  4. Population 1, because a larger standard deviation implies that the effect size will also be larger.
Explanation: The power of a hypothesis test is inversely related to the population variability (standard deviation, σ\sigma). A smaller σ\sigma leads to a smaller standard error of the sample mean (SE = σ/n\sigma / \sqrt{n}). A smaller standard error means the sampling distribution of the mean is narrower and more precise. This reduces the overlap between the sampling distributions under the null and alternative hypotheses, thus making it easier to detect a true difference from the null value. Therefore, choosing Population 2 with the lower standard deviation will result in a more powerful test. Distractor A is incorrect; high variability adds 'noise' and makes it harder to detect the 'signal' (the true effect). Distractor C is incorrect as σ\sigma is a key component in the test statistic and power calculation. Distractor D incorrectly conflates population variability with effect size.

Question 11

A medical research team is designing a trial for a new cancer therapy. They determine that a Type I error (concluding the therapy is effective when it is not) is dangerous, but a Type II error (failing to recognize an effective therapy) is also a major concern as it would shelve a potentially life-saving treatment. How does the desire to avoid a Type II error relate to the concept of statistical power?

  1. Avoiding a Type II error is unrelated to power, which only concerns the probability of a Type I error.
  2. To avoid a Type II error, the researchers should aim for a test with low power, as high power increases the chance of error.
  3. Minimizing the probability of a Type II error is statistically equivalent to maximizing the power of the test. (correct answer)
  4. The probability of a Type II error is equal to the power of the test, so they are the same concept.
Explanation: A Type II error is the failure to reject a null hypothesis when it is in fact false. The probability of a Type II error is denoted by β\beta. Statistical power is defined as the probability of correctly rejecting a null hypothesis when it is false. By definition, power is equal to 1β1 - \beta. Therefore, minimizing the probability of a Type II error (minimizing β\beta) is the same as maximizing the statistical power (maximizing 1β1 - \beta). These are two sides of the same coin. Distractor A is incorrect as power is directly related to Type II error. Distractor B is the opposite of the correct relationship. Distractor D is incorrect; power is 1β1 - \beta, not β\beta.

Question 12

Before launching a large-scale study, a sociologist performs a power analysis and determines that the planned experiment has a power of 0.80 to detect the effect of interest. Which of the following is the correct interpretation of this result?

  1. There is an 80% probability that the alternative hypothesis is true.
  2. There is an 80% probability that the study's conclusions will be correct.
  3. If the study yields a statistically significant result, there is an 80% probability that the alternative hypothesis is true.
  4. If the alternative hypothesis is true, there is an 80% probability that the study will yield a statistically significant result. (correct answer)
Explanation: When you encounter power analysis questions, focus on understanding what statistical power actually measures: the probability of correctly rejecting a false null hypothesis. Power of 0.80 means that if the alternative hypothesis is truly correct in the population, there's an 80% chance your study will detect this effect and yield a statistically significant result. Think of it as the test's sensitivity - how good it is at finding real effects when they exist. Answer D captures this definition perfectly: "If the alternative hypothesis is true, there is an 80% probability that the study will yield a statistically significant result." This conditional probability statement (given Ha is true) is exactly what power measures. Answer A incorrectly suggests power tells us the probability that Ha is true. Power doesn't determine whether hypotheses are true - it assumes Ha is true and calculates detection probability. Answer B is too broad, claiming 80% probability of correct conclusions overall. Power only addresses one type of correct decision (detecting true effects), not general accuracy including correctly accepting true null hypotheses. Answer C confuses power with the complement of Type I error or relates to positive predictive value. This describes the probability Ha is true given a significant result, which depends on prior probabilities and isn't what power analysis calculates. Remember: Power always assumes the alternative hypothesis is true as a starting point. It answers "If there really is an effect, how likely am I to find it?" This conditional thinking is crucial for interpreting power analysis results correctly.

Question 13

Two research groups, Group 1 and Group 2, are planning studies to achieve a statistical power of 0.80. Group 1 is investigating a phenomenon with a large expected effect size. Group 2 is investigating a different phenomenon with a small expected effect size. Assuming both groups use the same significance level and have similar population variability, what can be concluded about the sample sizes required for their respective studies?

  1. Group 1 will need a larger sample size than Group 2.
  2. The required sample sizes cannot be compared without knowing the specific test being used.
  3. Both groups will require approximately the same sample size to achieve the same power.
  4. Group 2 will need a larger sample size than Group 1. (correct answer)
Explanation: When you encounter questions about statistical power and sample size, remember that these four factors are interconnected: sample size, effect size, significance level (α), and power. When three are held constant, the fourth must adjust accordingly. Statistical power represents your ability to detect a true effect when it exists. To achieve the same power level (0.80 in this case), you need enough data to reliably distinguish your effect from random noise. Here's the key relationship: larger effect sizes are easier to detect and require smaller sample sizes, while smaller effect sizes are harder to detect and require larger sample sizes. Think of it like trying to hear a whisper versus a shout in a noisy room. A large effect size is like a shout—you don't need many observations to confidently say it's real. A small effect size is like a whisper—you need many more observations to distinguish it from background noise with the same level of confidence. Since Group 2 is studying a phenomenon with a small expected effect size while Group 1 has a large expected effect size, Group 2 will need a larger sample size to achieve the same statistical power. Choice A reverses this relationship—it's the small effect size group that needs more participants, not the large effect size group. Choice B incorrectly suggests the comparison is impossible; while specific calculations depend on the test type, the directional relationship between effect size and required sample size is universal. Choice C ignores how effect size influences sample size requirements. Study tip: Remember the inverse relationship—big effects need small samples, small effects need big samples for equivalent power.

Question 14

A power analysis is conducted for a one-sample t-test with the goal of achieving 80% power to detect a mean difference of μμ0=10\mu - \mu_0 = 10 units. If the true mean difference is actually 15 units, how would the actual power of the test (using the sample size determined by the analysis) compare to the planned 80%?

  1. The actual power would be less than 80% because the assumed effect size was incorrect.
  2. The actual power would still be 80% because power is fixed by the sample size and significance level.
  3. The actual power would be greater than 80% because a larger effect is easier to detect than the one planned for. (correct answer)
  4. The power cannot be determined, as the p-value is unknown.
Explanation: Power is a function of effect size. The study was designed to have an 80% chance of detecting a 10-unit difference. A true difference of 15 units is a larger effect size. For a fixed sample size and significance level, the power of a test always increases as the true effect size increases. Since the actual effect size (15) is larger than the effect size the study was powered for (10), the actual power of the test will be greater than the planned 80%. Distractor A is incorrect. Distractor B is incorrect because power depends on the true effect size, not just n and α\alpha. Distractor D is incorrect; power is a property of the test design, not the resulting p-value.

Question 15

A researcher plans to decrease the significance level α\alpha of a hypothesis test from 0.05 to 0.01. To maintain the original level of statistical power, what corresponding change must be made to the study design?

  1. The sample size must be increased. (correct answer)
  2. The sample size must be decreased.
  3. The anticipated effect size must be decreased.
  4. No change is needed; changing α\alpha does not affect the factors that determine power.
Explanation: There is an inherent trade-off between the Type I error rate (α\alpha) and the Type II error rate (β\beta), and thus power (1β1-\beta). Decreasing α\alpha from 0.05 to 0.01 makes the criterion for rejection stricter, which reduces the power of the test, all else being equal. To counteract this loss of power and restore it to its original level, the researcher must make another change that increases power. The most direct way to do this is to increase the sample size. A larger sample size reduces the standard error and compensates for the stricter rejection criterion. Decreasing sample size (B) or anticipating a smaller effect size (C) would both further decrease power.

Question 16

An epidemiologist is studying the link between a rare environmental exposure and a certain disease. Due to the rarity of the exposure, the number of participants in the exposed group is unavoidably small. The researcher suspects the exposure has a moderate effect on the disease risk. Given these constraints, what is the most likely statistical challenge this study will face?

  1. A high probability of a Type I error, leading to a false claim of an association.
  2. A biased effect size estimate, because small sample sizes tend to overestimate true effects.
  3. Difficulty in calculating a p-value because of the non-normal distribution of the data.
  4. Low statistical power, leading to a high probability of failing to detect the true association. (correct answer)
Explanation: When you encounter a study design question involving small sample sizes and effect detection, you're dealing with statistical power—the probability of correctly identifying a true effect when it exists. Statistical power depends on four key factors: sample size, effect size, significance level, and variability. In this epidemiological study, the researcher faces a critical constraint: the exposed group is "unavoidably small" due to the rarity of the environmental exposure. Even though a "moderate effect" is suspected, small sample sizes dramatically reduce power, making it difficult to detect even real associations. This creates a high risk of Type II error—failing to reject a false null hypothesis when the exposure actually does increase disease risk. Choice A is incorrect because Type I error (false positive) isn't more likely with small samples; the significance level remains constant regardless of sample size. Choice B misunderstands how sample size affects estimates—while small samples increase variability, they don't systematically bias effect sizes upward. Choice C incorrectly assumes that p-value calculation requires normal distributions; many statistical tests work with non-normal data, and exact tests exist for small samples. Choice D correctly identifies that low statistical power is the primary concern. With insufficient participants in the exposed group, the study lacks the ability to detect the suspected moderate effect, leading to a high probability of missing a true association. Remember: small sample size = low power = high risk of missing real effects. When you see "small sample" paired with "detecting an effect," think power problems first.

Question 17

A team of psychologists conducted a pilot study with n=30n=30 participants to test a new intervention and found a non-significant result (p = 0.25). They suspect the intervention has a small but meaningful effect. They are now designing a larger, definitive study and want to maximize their chances of correctly identifying this effect if it truly exists. Which of the following strategies provides the most direct and substantial increase in statistical power, assuming all other factors are held constant?

  1. Increasing the sample size from n=30n=30 to n=300n=300. (correct answer)
  2. Changing the significance level from α=0.05\alpha = 0.05 to α=0.01\alpha = 0.01.
  3. Recruiting a more diverse and heterogeneous sample of participants.
  4. Using a two-tailed test instead of the one-tailed test used in the pilot study.
Explanation: Statistical power is heavily dependent on sample size. Increasing the sample size (n) reduces the standard error of the test statistic. This results in a narrower sampling distribution, which makes it easier to distinguish the alternative hypothesis distribution from the null hypothesis distribution, thus increasing the probability of correctly rejecting a false null hypothesis. Answer B is incorrect because decreasing α\alpha would decrease power. Answer C is incorrect because a more heterogeneous sample would likely increase population variance (σ2\sigma^2), which would decrease power. Answer D is incorrect because, for a fixed α\alpha and an effect in the expected direction, a one-tailed test is more powerful than a two-tailed test; switching to a two-tailed test would decrease power.

Question 18

A cognitive scientist has a strong theoretical reason to hypothesize that a specific memory-enhancing drug will improve recall scores, and has no reason to believe it could harm them. In designing an experiment to test this hypothesis against a placebo, which choice regarding the hypothesis test structure and its corresponding rationale is most appropriate for maximizing the ability to detect the anticipated effect?

  1. A two-tailed test should be used because it is more conservative and reduces the risk of a Type I error, which increases power.
  2. A one-tailed test should be used because it allocates the entire rejection region (α\alpha) to one side of the distribution, making the test more sensitive to an effect in the hypothesized direction. (correct answer)
  3. A two-tailed test should be used because it has greater power than a one-tailed test to detect any deviation from the null hypothesis, regardless of direction.
  4. A one-tailed test should be used, but the significance level α\alpha must be halved to maintain the same power as a two-tailed test.
Explanation: When there is a strong directional hypothesis (e.g., improvement only), a one-tailed test is more powerful than a two-tailed test for detecting an effect in that specific direction. This is because the entire significance level α\alpha is placed in one tail. For a given α\alpha, this results in a critical value that is less extreme than the critical values for a two-tailed test, making it easier to reject the null hypothesis if the effect is indeed in the predicted direction. Answer A is incorrect because being more conservative against Type I error (using a two-tailed test when a one-tailed is justified) decreases power. Answer C is incorrect; a one-tailed test is more powerful for the specified direction. Answer D is a misunderstanding of the relationship between test type and power.

Question 19

A medical research team is designing a trial for a new cancer therapy. They determine that a Type I error (concluding the therapy is effective when it is not) is dangerous, but a Type II error (failing to recognize an effective therapy) is also a major concern as it would shelve a potentially life-saving treatment. How does the desire to avoid a Type II error relate to the concept of statistical power?

  1. Avoiding a Type II error is unrelated to power, which only concerns the probability of a Type I error.
  2. To avoid a Type II error, the researchers should aim for a test with low power, as high power increases the chance of error.
  3. Minimizing the probability of a Type II error is statistically equivalent to maximizing the power of the test. (correct answer)
  4. The probability of a Type II error is equal to the power of the test, so they are the same concept.
Explanation: A Type II error is the failure to reject a null hypothesis when it is in fact false. The probability of a Type II error is denoted by β\beta. Statistical power is defined as the probability of correctly rejecting a null hypothesis when it is false. By definition, power is equal to 1β1 - \beta. Therefore, minimizing the probability of a Type II error (minimizing β\beta) is the same as maximizing the statistical power (maximizing 1β1 - \beta). These are two sides of the same coin. Distractor A is incorrect as power is directly related to Type II error. Distractor B is the opposite of the correct relationship. Distractor D is incorrect; power is 1β1 - \beta, not β\beta.

Question 20

Two research groups, Group 1 and Group 2, are planning studies to achieve a statistical power of 0.80. Group 1 is investigating a phenomenon with a large expected effect size. Group 2 is investigating a different phenomenon with a small expected effect size. Assuming both groups use the same significance level and have similar population variability, what can be concluded about the sample sizes required for their respective studies?

  1. Group 1 will need a larger sample size than Group 2.
  2. The required sample sizes cannot be compared without knowing the specific test being used.
  3. Both groups will require approximately the same sample size to achieve the same power.
  4. Group 2 will need a larger sample size than Group 1. (correct answer)
Explanation: When you encounter questions about statistical power and sample size, remember that these four factors are interconnected: sample size, effect size, significance level (α), and power. When three are held constant, the fourth must adjust accordingly. Statistical power represents your ability to detect a true effect when it exists. To achieve the same power level (0.80 in this case), you need enough data to reliably distinguish your effect from random noise. Here's the key relationship: larger effect sizes are easier to detect and require smaller sample sizes, while smaller effect sizes are harder to detect and require larger sample sizes. Think of it like trying to hear a whisper versus a shout in a noisy room. A large effect size is like a shout—you don't need many observations to confidently say it's real. A small effect size is like a whisper—you need many more observations to distinguish it from background noise with the same level of confidence. Since Group 2 is studying a phenomenon with a small expected effect size while Group 1 has a large expected effect size, Group 2 will need a larger sample size to achieve the same statistical power. Choice A reverses this relationship—it's the small effect size group that needs more participants, not the large effect size group. Choice B incorrectly suggests the comparison is impossible; while specific calculations depend on the test type, the directional relationship between effect size and required sample size is universal. Choice C ignores how effect size influences sample size requirements. Study tip: Remember the inverse relationship—big effects need small samples, small effects need big samples for equivalent power.