Biostatistics Quiz: Type I Ii Errors And Power
20 questions · exam conditions
0:00
Type I Ii Errors And PowerQuestion 1 of 20

In a one-tailed test with α = 0.05, a researcher finds that increasing the sample size from 30 to 120 changes the power from 0.65 to 0.90. If the researcher had instead used a two-tailed test with the same sample sizes and effect size, how would the power values compare?

Both power values would be higher because two-tailed tests are more sensitive
Both power values would be lower because the critical region is split between two tails
The power values would remain exactly the same since sample size determines power
Power would be higher for n=30 but lower for n=120 due to different scaling effects
The power comparison would depend on the direction of the true effect size
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: Type I Ii Errors And Power

Practice Type I Ii Errors And Power in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Type I Ii Errors And Power, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

In a one-tailed test with α = 0.05, a researcher finds that increasing the sample size from 30 to 120 changes the power from 0.65 to 0.90. If the researcher had instead used a two-tailed test with the same sample sizes and effect size, how would the power values compare?

  1. Both power values would be higher because two-tailed tests are more sensitive
  2. Both power values would be lower because the critical region is split between two tails (correct answer)
  3. The power values would remain exactly the same since sample size determines power
  4. Power would be higher for n=30 but lower for n=120 due to different scaling effects
  5. The power comparison would depend on the direction of the true effect size
Explanation: When you encounter questions about statistical power and test design, remember that power is the probability of correctly rejecting a false null hypothesis, and it's affected by sample size, effect size, significance level, and whether the test is one-tailed or two-tailed. The key insight here is how the critical region differs between one-tailed and two-tailed tests. In a one-tailed test with α = 0.05, the entire 5% rejection region is placed in one tail. In a two-tailed test with the same α, this 5% is split between both tails (2.5% each), making the critical values more extreme and harder to reach. Since power depends on how likely you are to fall into the rejection region when the alternative hypothesis is true, splitting the critical region between two tails reduces this probability. Both the n=30 and n=120 scenarios would have lower power with a two-tailed test compared to the one-tailed version, though the larger sample size would still provide greater power than the smaller one. Option A is incorrect because two-tailed tests are actually less sensitive for detecting effects in a specific direction. Option C misses the crucial point that test design (one-tailed vs. two-tailed) affects power independently of sample size. Option D incorrectly suggests different scaling effects between sample sizes, which isn't how statistical power works. Study tip: Remember that one-tailed tests are more powerful when you correctly predict the direction of an effect, but two-tailed tests are more conservative and commonly used when direction isn't predetermined. The trade-off between power and Type I error protection is fundamental to hypothesis testing.

Question 2

In a clinical trial with significance level α = 0.05, researchers want to detect a clinically meaningful difference with power = 0.80. If they decrease the significance level to α = 0.01 while keeping all other factors constant, what happens to the probability of Type II error?

  1. The probability of Type II error decreases because the test becomes more stringent
  2. The probability of Type II error increases because it becomes harder to reject the null hypothesis (correct answer)
  3. The probability of Type II error remains unchanged since it only depends on sample size
  4. The probability of Type II error becomes 0.01 since it equals the new significance level
  5. The probability of Type II error becomes 0.20 since power remains constant at 0.80
Explanation: When you encounter questions about statistical power and error types, focus on the fundamental relationship between Type I error (α), Type II error (β), and power (1-β). These elements are interconnected - changing one affects the others when sample size and effect size remain constant. Decreasing the significance level from α = 0.05 to α = 0.01 makes your rejection criteria much more stringent. You now need stronger evidence to reject the null hypothesis. This creates a higher bar for detecting the true effect, making it more likely you'll fail to reject a false null hypothesis - which is precisely the definition of Type II error. Since the original power was 0.80 (meaning β = 0.20), making the test more conservative will increase β above 0.20. Choice A incorrectly suggests that a more stringent test reduces Type II error. While the test becomes more rigorous against false positives, this actually makes it worse at detecting true effects. Choice C reflects a common misconception - while sample size does affect Type II error, it's not the only factor. The significance level also influences β when other parameters are held constant. Choice D confuses Type I and Type II errors entirely. The significance level α represents the probability of Type I error, not Type II error. Remember this inverse relationship: when you decrease α (making it harder to reject the null), you increase β (making it more likely to miss a true effect), assuming everything else stays the same. This trade-off is fundamental to hypothesis testing.

Question 3

A diagnostic test for a rare disease has a sensitivity of 0.90 and specificity of 0.85. In hypothesis testing terms, if the null hypothesis is 'patient does not have the disease' and the alternative is 'patient has the disease', what do sensitivity and specificity represent respectively?

  1. Sensitivity represents (1 - probability of Type I error), specificity represents power of the test
  2. Sensitivity represents power of the test, specificity represents (1 - probability of Type I error) (correct answer)
  3. Sensitivity represents probability of Type I error, specificity represents probability of Type II error
  4. Sensitivity represents (1 - probability of Type II error), specificity represents probability of Type I error
  5. Sensitivity represents probability of Type II error, specificity represents (1 - probability of Type I error)
Explanation: When you encounter questions linking diagnostic test characteristics to hypothesis testing, you need to map medical terminology onto statistical concepts by carefully defining what constitutes a "positive result" in the hypothesis testing framework. In this setup, the null hypothesis is "patient does not have the disease" and the alternative is "patient has the disease." A positive test result means rejecting the null hypothesis (concluding the patient has the disease). Sensitivity measures the probability of correctly identifying someone who actually has the disease - this means correctly rejecting a false null hypothesis. In hypothesis testing terms, this is the power of the test (the probability of correctly rejecting the null when it's false). Specificity measures the probability of correctly identifying someone who doesn't have the disease - this means correctly accepting a true null hypothesis, which equals (1 - probability of Type I error). Choice A reverses these concepts, incorrectly assigning sensitivity to (1 - Type I error) and specificity to power. Choice C confuses the probabilities themselves rather than their complements - sensitivity isn't the Type I error rate, and specificity isn't the Type II error rate. Choice D incorrectly states that specificity represents the Type I error probability itself, when it actually represents one minus that probability. The correct answer is B: sensitivity represents power, and specificity represents (1 - probability of Type I error). Study tip: Remember that sensitivity relates to detecting true positives (power), while specificity relates to avoiding false positives (avoiding Type I errors). Draw the 2×2 table linking disease status to test results to visualize these relationships.

Question 4

An environmental study tests whether pollution levels exceed safety standards. The consequences of incorrectly concluding pollution is safe when it's actually dangerous are severe, while incorrectly concluding pollution is dangerous when it's safe results in unnecessary but manageable costs. How should the researchers set their significance level?

  1. Use a large significance level (α = 0.10) to minimize Type II error since missing real danger is costly (correct answer)
  2. Use a small significance level (α = 0.01) to minimize Type I error since false alarms are costly
  3. Use the standard significance level (α = 0.05) since it balances both error types equally
  4. Use a very small significance level (α = 0.001) to minimize both error types simultaneously
  5. The significance level choice doesn't matter since power analysis determines error probabilities
Explanation: When you encounter questions about setting significance levels, think about the real-world consequences of each type of error. Type I error means rejecting a true null hypothesis (false positive), while Type II error means failing to reject a false null hypothesis (false negative). In this pollution study, the null hypothesis would be "pollution levels are safe." A Type I error means concluding pollution is dangerous when it's actually safe, leading to unnecessary cleanup costs. A Type II error means concluding pollution is safe when it's actually dangerous, potentially causing severe health consequences. Since missing real danger has severe consequences while false alarms only create manageable costs, you want to minimize Type II error. Using a larger significance level like α=0.10α = 0.10 makes it easier to detect actual pollution problems by increasing your chance of rejecting the null hypothesis when pollution truly is dangerous. Answer A correctly identifies this approach. Answer B is backwards—using α=0.01α = 0.01 minimizes Type I error (false alarms) but increases Type II error (missing real danger), which is exactly what you don't want here. Answer C incorrectly assumes α=0.05α = 0.05 provides optimal balance regardless of context—the consequences clearly aren't equal in this scenario. Answer D misunderstands the relationship between significance levels and error types; you cannot minimize both simultaneously by choosing a very small alpha, as this actually increases Type II error. Study tip: Always identify which error type has more serious real-world consequences. When missing a problem is worse than a false alarm, use a larger significance level to reduce Type II error.

Question 5

Two studies examine the same research question with identical effect sizes, but Study A uses n=50 participants while Study B uses n=200 participants. Both use α = 0.05. Compared to Study A, Study B has:

  1. Higher probability of Type I error and lower probability of Type II error
  2. Lower probability of Type I error and higher probability of Type II error
  3. Same probability of Type I error and lower probability of Type II error (correct answer)
  4. Same probability of Type I error and higher probability of Type II error
  5. Lower probability of both Type I and Type II errors due to increased precision
Explanation: When you encounter questions about sample size and error probabilities, focus on how Type I and Type II errors relate to study design parameters like sample size and significance level. Type I error probability (α) is the chance of rejecting a true null hypothesis - essentially finding a "false positive." This probability is set by the researcher when choosing the significance level. Since both studies use α = 0.05, they have identical Type I error probabilities of 5%, regardless of sample size. Type II error probability (β) is the chance of failing to reject a false null hypothesis - missing a real effect. Statistical power (1 - β) increases with larger sample sizes when the effect size and α remain constant. Study B's larger sample (n=200 vs n=50) provides greater power to detect the same effect size, meaning lower probability of Type II error. Answer C correctly identifies that Type I error probability stays the same (both studies use α = 0.05) while Type II error probability decreases with the larger sample size. Answer A incorrectly suggests Type I error probability changes with sample size - it doesn't, since α is fixed. Answer B wrongly claims Type I error decreases while Type II error increases, which contradicts how sample size affects power. Answer D correctly identifies unchanged Type I error probability but incorrectly states that Type II error increases - the opposite is true with larger samples. Remember: α (Type I error) is set by your significance level choice, while β (Type II error) decreases as sample size increases, assuming constant effect size and α.

Question 6

An educational intervention study uses a very large sample size (n = 5000) to test whether a new teaching method improves test scores. The researchers find a statistically significant result (p < 0.001) but the actual improvement is only 1 point on a 100-point scale. What concern should be raised about this finding?

  1. Type II error is likely because the effect size appears to be practically meaningless
  2. Type I error is likely because very large samples can detect trivial differences as significant (correct answer)
  3. No error concerns exist since the p-value provides strong evidence against the null hypothesis
  4. Both error types are problematic due to the mismatch between sample size and effect size
  5. Power was insufficient despite the large sample since the effect was barely detectable
Explanation: When you encounter a study with very large sample sizes and statistically significant but practically small effects, you're dealing with a classic issue in statistical interpretation: the relationship between sample size, statistical significance, and practical importance. With extremely large samples (n = 5000), statistical tests become incredibly sensitive and can detect even tiny differences between groups as "statistically significant." Here, a 1-point improvement on a 100-point scale represents just 1% improvement - practically negligible in educational terms. Yet the massive sample size gave enough statistical power to detect this trivial difference with high confidence (p < 0.001). This illustrates how statistical significance doesn't guarantee practical significance. Option B correctly identifies that large samples can make Type I errors more likely by detecting differences that, while real, are so small they're meaningless in practice. The statistical test is working "too well" - finding significance where none should matter. Option A confuses the error types. Type II error means failing to detect a real effect, but here we did detect an effect (though trivial). Option C misses the key issue entirely - a low p-value alone doesn't address whether the finding matters practically. The statistical evidence against the null hypothesis is strong, but the effect itself is worthless. Option D incorrectly suggests both error types are present, when only the concern about detecting trivial differences (Type I error territory) applies. Remember: with large samples, always evaluate effect size alongside statistical significance. A tiny effect that's statistically significant may still be practically meaningless.

Question 7

A pharmaceutical company tests 100 different drug compounds, each in a separate hypothesis test with α = 0.05. Assuming all null hypotheses are true (no compounds are effective), approximately how many Type I errors would be expected, and what does this illustrate about multiple testing?

  1. About 5 Type I errors would occur, illustrating why multiple testing requires power analysis
  2. About 5 Type I errors would occur, illustrating why multiple testing requires adjusted significance levels (correct answer)
  3. About 1 Type I error would occur, showing that multiple testing reduces overall error rates
  4. About 95 Type I errors would occur, demonstrating the multiplicative effect of repeated testing
  5. No Type I errors would occur since the null hypotheses are actually true in this scenario
Explanation: When you encounter multiple testing scenarios in biostatistics, think about the cumulative probability of making at least one error across all tests. This is a fundamental concept in controlling family-wise error rates. With 100 independent hypothesis tests at α = 0.05, and assuming all null hypotheses are true, each test has a 5% chance of producing a Type I error (false positive). The expected number of Type I errors is simply: 100×0.05=5100 \times 0.05 = 5 errors. This demonstrates why multiple testing is problematic—even when no real effects exist, you'll still get several "significant" results by chance alone, leading to false discoveries. Looking at the answer choices: Choice A correctly calculates 5 Type I errors but incorrectly suggests this illustrates the need for power analysis. Power analysis relates to Type II errors and detecting true effects, not controlling false positives. Choice C is mathematically wrong (1 error instead of 5) and incorrectly claims multiple testing reduces error rates—it actually inflates them. Choice D drastically overestimates the errors at 95, confusing the concept with the probability of avoiding any Type I errors. Choice B is correct because it properly calculates the expected 5 Type I errors and correctly identifies that this illustrates why multiple testing requires adjusted significance levels. Methods like Bonferroni correction, false discovery rate control, or other multiple comparison adjustments are essential to maintain the overall error rate at acceptable levels. Remember: multiple testing inflates your chance of false discoveries, so always consider whether significance level adjustments are needed when conducting numerous simultaneous tests.

Question 8

A nutrition study tests whether a dietary supplement affects cholesterol levels. The researchers plan for 80% power to detect a 15 mg/dL reduction in cholesterol. After the study, they find a non-significant result (p = 0.18) but the actual mean reduction was 12 mg/dL. What most likely explains this outcome?

  1. A Type I error occurred because the p-value exceeds the typical significance threshold
  2. A Type II error occurred because the actual effect was smaller than what the study was powered to detect (correct answer)
  3. No error occurred because the study was properly designed with adequate power for meaningful differences
  4. The power calculation was incorrect since any reduction in cholesterol should have been detected
  5. A Type I error was appropriately avoided because the effect size was clinically unimportant
Explanation: When you encounter questions about study outcomes and statistical errors, focus on the relationship between what the study was designed to detect versus what actually happened. This scenario involves understanding statistical power and Type II errors. The study was powered to detect a 15 mg/dL reduction with 80% power, meaning there was an 80% chance of finding a statistically significant result if the true effect was 15 mg/dL or larger. However, the actual effect was only 12 mg/dL—smaller than what the study was designed to detect. Since the effect was smaller than anticipated, the study lacked sufficient power to reliably detect it, making a non-significant result likely even when a true effect exists. This describes a Type II error: failing to detect a real effect because the study wasn't adequately powered for that effect size. Option A incorrectly identifies this as a Type I error, which would be finding a significant result when no true effect exists—the opposite of what happened here. Option C is wrong because while the study design was reasonable for detecting 15 mg/dL differences, it wasn't adequate for detecting the smaller 12 mg/dL effect that actually occurred. Option D misunderstands power calculations—studies are designed to detect clinically meaningful differences, not any detectable change, and smaller effects require larger sample sizes to detect reliably. Remember: Type II errors often occur when the actual effect size is smaller than what the study was powered to detect. Always compare the observed effect size to the target effect size used in power calculations.

Question 9

In a behavioral study, researchers test whether a mindfulness intervention reduces anxiety scores. They use α = 0.05 and achieve 90% power to detect a moderate effect. The intervention truly has no effect, but due to measurement error and random variation, they obtain p = 0.03. What error occurs, and what is notable about the power level?

  1. Type I error occurs; high power made this error more likely to happen
  2. Type I error occurs; high power is irrelevant since power only relates to Type II error (correct answer)
  3. Type II error occurs; high power should have prevented this error from happening
  4. No error occurs; high power ensures correct decisions regardless of the true effect
  5. Type I error occurs; high power indicates the study was over-designed for this research question
Explanation: When you encounter questions about statistical errors and power, focus on the fundamental definitions: Type I error occurs when you reject a true null hypothesis, while Type II error occurs when you fail to reject a false null hypothesis. Power is specifically the probability of correctly rejecting a false null hypothesis. In this scenario, the intervention truly has no effect (null hypothesis is true), but the researchers obtained p = 0.03 < 0.05, leading them to reject the null hypothesis. Since they rejected a true null hypothesis, this is a Type I error. The key insight is that power only relates to situations where the null hypothesis is false - it's the probability of avoiding Type II error. When the null hypothesis is actually true (as in this case), power becomes irrelevant to the outcome. Option A incorrectly suggests that high power increases the likelihood of Type I error. Power doesn't affect Type I error rates - that's controlled by your alpha level (0.05 here). Option C misidentifies this as Type II error, which would occur if they failed to detect a real effect. Option D is wrong because high power doesn't guarantee correct decisions when the null hypothesis is true - Type I errors can still occur at the rate specified by alpha. The correct answer is B because it properly identifies the Type I error and recognizes that power is irrelevant in this situation since power only applies when the alternative hypothesis is true. Study tip: Remember the power paradox - high power is great for detecting real effects, but it's completely irrelevant when there's no effect to detect. Type I error rates depend only on your alpha level, never on power.

Question 10

A researcher conducts a pilot study (n = 25) and finds a promising but non-significant trend (p = 0.08). She calculates that with the observed effect size, she would need n = 100 to achieve 80% power. If she conducts the larger study and the true effect size remains the same, what is the probability of obtaining a significant result?

  1. 0.05, since this equals the Type I error rate regardless of sample size
  2. 0.80, since this represents the power to detect the true effect with n = 100 (correct answer)
  3. 0.92, calculated as 0.80 + 0.05 + 0.07 accounting for all possible outcomes
  4. 0.95, representing one minus the Type I error rate for this study design
  5. 0.20, representing the Type II error rate for this specific scenario
Explanation: When you encounter power analysis questions, focus on understanding what each statistical concept represents in the context of detecting true effects. Statistical power is defined as the probability of correctly rejecting a false null hypothesis when there is a true effect present. In this scenario, the researcher has calculated that with n = 100, she would have 80% power to detect the observed effect size. This means that if the true effect size remains the same as observed in the pilot study, there's an 80% probability that the larger study will yield a statistically significant result (p < 0.05). The correct answer is B because power directly answers the question being asked - the probability of obtaining a significant result when a true effect exists is exactly what statistical power measures. Option A incorrectly confuses Type I error rate (α = 0.05) with power. Type I error is the probability of finding significance when there's no true effect, not when there is a true effect. Option C attempts to add probabilities inappropriately - you can't simply sum power, alpha, and an arbitrary value to get a meaningful probability. These aren't mutually exclusive outcomes that sum to 1. Option D represents 1 - α, which has no relevance to detecting true effects and would actually be the probability of not making a Type I error under the null hypothesis. Remember: Power analysis questions often test whether you can distinguish between Type I error (false positive rate) and power (true positive rate). Power tells you the probability of detecting real effects, while alpha tells you the probability of falsely detecting non-existent effects.

Question 11

An industrial engineer tests whether a process modification reduces defect rates from the current 8% level. She observes 12 defects in a sample of 200 items (6% defect rate) and obtains p = 0.15 using a one-tailed test with α = 0.05. Unknown to her, the modification actually reduces defects to 5%. What error occurred, and what was the approximate power for detecting this true 3 percentage point reduction?

  1. Type I error occurred; power was approximately 85% for detecting a 3 point reduction
  2. Type II error occurred; power was approximately 45% based on the non-significant result pattern (correct answer)
  3. Type I error occurred; power cannot be determined from the information provided in this scenario
  4. Type II error occurred; power was approximately 85% but sampling variation caused the miss
  5. No error occurred; the 6% observed rate correctly reflects the true 5% population rate
Explanation: When analyzing hypothesis testing outcomes, you need to distinguish between the statistical decision and the true state of nature. Here, the engineer failed to reject the null hypothesis (p = 0.15 > α = 0.05), concluding no significant reduction in defects. However, the modification actually does reduce defects from 8% to 5% - meaning the null hypothesis is false but wasn't rejected. This scenario represents a Type II error: failing to reject a false null hypothesis. The engineer missed detecting a real improvement. Since we know the true effect size (3 percentage point reduction), we can estimate the power. With a sample size of 200, baseline rate of 8%, and true rate of 5%, the power for this one-tailed test would be approximately 45% - meaning there was only about a 45% chance of detecting this difference. The non-significant result (p = 0.15) is consistent with this moderate power level. Answer choice A incorrectly identifies a Type I error (rejecting a true null hypothesis), which didn't occur since the null wasn't rejected. The power estimate of 85% is also too high for this scenario. Choice C correctly identifies the error type but wrongly claims power can't be determined - it can be estimated when you know the true effect size. Choice D correctly identifies the Type II error but overestimates the power at 85%, when the actual power for detecting a 3-point reduction with n=200 is much lower. Remember: Type I errors occur when you reject a true null (false positive), while Type II errors occur when you fail to reject a false null (false negative).

Question 12

A meta-analysis combines results from 15 studies testing the same intervention. Each individual study had 60% power to detect the target effect size. If the intervention truly has the target effect size, what is the approximate probability that the meta-analysis will fail to detect a significant overall effect?

  1. 0.40, since this represents the Type II error rate that applies to meta-analyses
  2. Much less than 0.40, since combining studies dramatically increases the effective power (correct answer)
  3. 0.60, since this represents the power level that carries over to meta-analyses
  4. Approximately 0.006, calculated as (0.40)^15 representing independent failure across all studies
  5. Cannot be determined without knowing how many individual studies achieved significance
Explanation: When you encounter meta-analysis questions, remember that combining multiple studies creates a much larger effective sample size, which dramatically increases statistical power beyond what any individual study could achieve. Each study has 60% power, meaning a 40% chance of failing to detect the true effect (Type II error). However, when you combine 15 studies in a meta-analysis, you're not just adding their individual power rates—you're pooling their data to create a much more sensitive test. The combined sample size becomes enormous compared to individual studies, and this substantially reduces the probability of missing a real effect. Think of it like multiple witnesses to an event: while each witness might miss certain details, combining all their observations gives you a much clearer picture than any single account. Option A incorrectly assumes the 40% Type II error rate transfers directly to meta-analyses, ignoring the power gained from combining studies. Option C makes the opposite error, suggesting the 60% power rate simply carries over unchanged. Option D attempts a mathematical approach by calculating (0.40)150.000001(0.40)^{15} ≈ 0.000001, treating each study as an independent trial that must fail for the meta-analysis to fail. While creative, this drastically underestimates the failure probability because it assumes perfect independence and doesn't account for how meta-analysis actually works. The correct answer is B—the probability of failing to detect the effect becomes much smaller than 0.40 due to the dramatically increased effective power from combining studies. Study tip: For meta-analysis questions, always remember that the key benefit is vastly increased power through larger effective sample sizes, not simple arithmetic combinations of individual study statistics.

Question 13

A screening program for a genetic disorder uses a test with 95% sensitivity and 90% specificity. Public health officials want to minimize the chance of missing affected individuals, even if this means more false positives. In hypothesis testing framework, what error type are they prioritizing, and how should they adjust their decision threshold?

  1. Minimizing Type I error by raising the threshold to require stronger evidence before declaring someone affected
  2. Minimizing Type II error by lowering the threshold to make it easier to declare someone affected (correct answer)
  3. Minimizing Type I error by lowering the threshold to reduce false negative diagnoses
  4. Minimizing Type II error by raising the threshold to improve the specificity of testing
  5. Balancing both error types equally since sensitivity and specificity are both already high
Explanation: When you encounter screening test questions, think about the relationship between hypothesis testing errors and diagnostic outcomes. Type I error means rejecting a true null hypothesis (false positive), while Type II error means failing to reject a false null hypothesis (false negative). In medical screening, the null hypothesis is typically "the person is not affected." A Type II error would mean failing to detect someone who actually has the condition—exactly what the officials want to minimize. When you want to catch more affected individuals, you need to lower the decision threshold, making it easier to declare someone positive. This increases sensitivity but decreases specificity, resulting in more false positives but fewer missed cases. Answer B correctly identifies that minimizing Type II error (missing affected individuals) requires lowering the threshold to make positive declarations more likely. This trade-off accepts more false positives to reduce false negatives. Answer A incorrectly suggests minimizing Type I error and raising the threshold, which would actually increase missed cases. Answer C correctly identifies minimizing Type I error as the wrong priority but incorrectly states that lowering the threshold reduces false negatives—while true, the goal isn't to minimize Type I error here. Answer D wrongly suggests minimizing Type II error requires raising the threshold, which would actually increase Type II errors by making it harder to detect affected individuals. Remember: In screening contexts, when officials prioritize "not missing cases," they're willing to accept false positives to minimize false negatives—that's minimizing Type II error through lowered thresholds.

Question 14

A quality control manager tests whether a manufacturing process produces defective items at a rate different from the historical 5% rate. She collects a sample and finds a test statistic that would occur in fewer than 3% of samples if the null hypothesis were true. However, the true defect rate is actually 5%. What error, if any, will the manager make?

  1. No error will be made since the test statistic correctly reflects the sampling distribution
  2. A Type I error will be made because a true null hypothesis will be rejected
  3. A Type II error will be made because the alternative hypothesis will not be supported
  4. Both Type I and Type II errors occur simultaneously in this scenario
  5. The error type depends on which significance level the manager chooses to use (correct answer)
Explanation: This question tests your understanding of statistical errors in hypothesis testing, specifically when you know the true population parameter. The key insight is recognizing what actually happens when the null hypothesis is true but gets rejected. Let's set up the scenario: The null hypothesis states the defect rate equals 5%, and we're told the true defect rate is actually 5%. This means the null hypothesis is true. However, the manager finds a test statistic so extreme it would occur in less than 3% of samples under the null hypothesis. If she's using a typical significance level (like α = 0.05), she would reject the null hypothesis because her p-value (< 0.03) falls below the threshold. When you reject a true null hypothesis, you commit a Type I error by definition. This is exactly what happens here - the manager will conclude the defect rate differs from 5% when it actually is 5%. The correct answer is B. Answer A is wrong because unusual test statistics can occur even when the null hypothesis is true - that's the nature of sampling variability. Answer C describes a Type II error, which occurs when you fail to reject a false null hypothesis, but that's not this scenario. Answer D is incorrect because you cannot make both types of errors simultaneously in a single test - they're mutually exclusive outcomes. Study tip: Remember that Type I errors occur when rejecting true null hypotheses, while Type II errors occur when failing to reject false null hypotheses. The type of error depends solely on the truth of the null hypothesis and your decision, not on how extreme your test statistic appears.

Question 15

A researcher plans a study with 80% power to detect a medium effect size at α = 0.05. After data collection, the p-value is 0.08. Given that the true effect size is indeed medium as hypothesized, what most likely occurred?

  1. A Type I error occurred because the p-value exceeds the significance level
  2. A Type II error occurred because a true effect was not detected as significant (correct answer)
  3. No error occurred because the study was properly powered for this effect size
  4. The power calculation was incorrect since the p-value should have been less than 0.05
  5. A Type I error was avoided because the null hypothesis was not rejected inappropriately
Explanation: When you encounter questions about statistical errors and study power, you need to distinguish between what the study was designed to detect versus what actually happened, and understand the relationship between power, effect size, and statistical significance. A Type II error occurs when there's truly an effect present, but your statistical test fails to detect it as significant. Here, you're told the true effect size is medium (as hypothesized), but the p-value of 0.08 exceeds the α = 0.05 threshold, so the result isn't statistically significant. This means a real effect went undetected – the definition of a Type II error. The 80% power tells us there was a 20% chance this would happen, and unfortunately, this study fell into that 20%. Looking at the wrong answers: (A) incorrectly identifies this as a Type I error, but Type I errors occur when you detect a significant effect that doesn't actually exist – the opposite situation. (C) misunderstands that proper powering doesn't guarantee significant results; even well-powered studies have a failure rate equal to (1 - power). (D) reflects a fundamental misconception that adequate power guarantees p < 0.05, but power only tells you the probability of detecting an effect if it exists. Remember that power calculations give you probabilities, not certainties. An 80% powered study still has a 20% chance of missing a true effect. When you see questions combining power, effect sizes, and p-values, always identify whether a true effect exists first, then determine if your test detected it.

Question 16

A researcher studying cognitive enhancement compares two groups and finds no significant difference (p = 0.12) using α = 0.05. Unknown to the researcher, the treatment actually produces a moderate effect size (Cohen's d = 0.5). If the study had 45% power to detect this effect size, what is the probability that this specific outcome occurred?

  1. 0.05, representing the probability of Type I error for this significance level
  2. 0.55, representing the probability of Type II error given the study's power (correct answer)
  3. 0.45, representing the power to detect the effect that was actually present
  4. 0.12, representing the exact p-value obtained in this particular study
  5. 0.95, representing the probability of correctly accepting the null hypothesis
Explanation: When you encounter a question about statistical errors and power, you need to identify what actually happened in the study and match it to the correct probability concept. The key facts are: the researcher found no significant difference (p = 0.12 > 0.05), but a true effect actually exists (Cohen's d = 0.5). This means the researcher failed to detect a real effect - the definition of a Type II error (β). Since the study had 45% power, this means β = 1 - 0.45 = 0.55, or a 55% chance of missing the true effect. Answer B is correct because 0.55 represents the probability of Type II error. When power = 0.45, the probability of failing to detect the true effect is 1 - 0.45 = 0.55. This matches exactly what happened in the study. Answer A incorrectly identifies this as Type I error probability. Type I error (α = 0.05) occurs when you find a significant result when there's no true effect - the opposite of this scenario. Answer C confuses the probability of the outcome with the power itself. Power (0.45) represents the probability of correctly detecting the effect, not the probability of the specific outcome that occurred. Answer D mistakes the p-value for the probability of this outcome. The p-value (0.12) tells you the probability of observing your data assuming no true effect, but it doesn't represent the probability of making a Type II error. Remember: when a study fails to detect a true effect, focus on Type II error probability, which equals 1 minus the study's power.

Question 17

A researcher tests a new medication against a placebo to determine if it reduces blood pressure. The null hypothesis states there is no difference between treatments. If the researcher concludes the medication is effective when it actually has no effect, and separately, if the medication is truly effective but the researcher fails to detect this effect, what types of errors occur respectively?

  1. Type I error occurs in the first case, Type II error occurs in the second case (correct answer)
  2. Type II error occurs in the first case, Type I error occurs in the second case
  3. Type I error occurs in both cases since the null hypothesis is incorrectly handled
  4. Type II error occurs in both cases since the alternative hypothesis is incorrectly assessed
  5. No error occurs in either case since these represent normal statistical variation
Explanation: When you encounter hypothesis testing questions, focus on the relationship between what's actually true versus what the researcher concludes. There are two fundamental errors that can occur when making statistical decisions. A Type I error happens when you reject a true null hypothesis - essentially finding an effect that doesn't actually exist (a "false positive"). A Type II error occurs when you fail to reject a false null hypothesis - missing a real effect that does exist (a "false negative"). In this blood pressure study, the first scenario describes concluding the medication works when it actually has no effect. Since the null hypothesis states "no difference between treatments," this null hypothesis is actually true, but the researcher incorrectly rejects it. This is a classic Type I error. The second scenario involves a truly effective medication that the researcher fails to detect. Here, the null hypothesis is false (there really is a difference), but the researcher fails to reject it, committing a Type II error. Answer A correctly identifies both error types. Answer B reverses them, which is a common mistake when students confuse "false positive" with "false negative." Answer C incorrectly suggests both scenarios involve Type I errors, missing that failing to detect a real effect is fundamentally different from falsely claiming an effect exists. Answer D makes the same mistake with Type II errors. Remember this simple framework: Type I = finding something that isn't there (false alarm), Type II = missing something that is there (missed detection). This distinction is crucial for interpreting research validity.

Question 18

A researcher studying memory enhancement conducts identical experiments in three different laboratories. Each lab uses α = 0.05 and has 70% power to detect the target effect size. The true effect size is exactly as hypothesized. If all three labs conduct their studies independently, what is the probability that at least one lab will fail to detect the significant effect?

  1. 0.30, representing the Type II error rate for each individual laboratory study
  2. 0.027, calculated as (0.30)³ representing failure in all three laboratories simultaneously
  3. 0.657, calculated as 1 - (0.70)³ representing at least one lab failing to detect the effect (correct answer)
  4. 0.90, calculated as 3 × 0.30 representing the cumulative Type II error across laboratories
  5. 0.343, calculated as (0.70)³ representing successful detection in all three laboratories
Explanation: When you encounter questions about multiple independent experiments, you're dealing with probability calculations involving complementary events. The key insight is recognizing whether the question asks for "at least one" outcome versus "all" or "none." Here, each lab has 70% power, meaning a 70% chance of detecting the true effect and a 30% chance of failing to detect it (Type II error). Since the labs work independently, you can multiply their individual probabilities. The correct approach is to use the complement rule. The probability that at least one lab fails equals 1 minus the probability that all labs succeed. Since each lab has a 0.70 probability of success, the probability all three succeed is (0.70)3=0.343(0.70)^3 = 0.343. Therefore, the probability that at least one fails is 10.343=0.6571 - 0.343 = 0.657, making answer C correct. Answer A (0.30) only gives the Type II error rate for a single lab, ignoring that we have three independent attempts. Answer B (0.027) calculates the probability that all three labs fail simultaneously—the opposite of what we want. This represents (0.30)3(0.30)^3, which is the most restrictive outcome. Answer D (0.90) incorrectly adds the individual failure probabilities (3 × 0.30), but probabilities of independent events multiply, not add. Remember this pattern: for "at least one" questions with independent events, use the complement rule (1 minus the probability of the opposite outcome). This approach is much simpler than calculating all the individual scenarios where exactly one, two, or three labs fail.

Question 19

A medical researcher conducts a clinical trial to test a new treatment. The study has 70% power to detect the minimum clinically important difference. After analyzing the data, she finds a statistically significant result (p = 0.03) favoring the new treatment. What can be concluded about the possibility of error?

  1. No error occurred since the result is statistically significant and the study was adequately powered
  2. A Type I error may have occurred with 5% probability based on the significance level used (correct answer)
  3. A Type II error occurred because the power was insufficient at only 70%
  4. The probability of error is 30% based on the complement of the study's power
  5. A Type I error may have occurred with 3% probability based on the obtained p-value
Explanation: When you encounter questions about statistical significance and study power, you need to distinguish between the two types of statistical errors and understand what each probability represents. Since the researcher found a statistically significant result (p = 0.03), she rejected the null hypothesis. However, there's always a chance this rejection was incorrect—this would be a Type I error (falsely concluding there's an effect when there isn't one). The probability of Type I error equals the significance level used in the study. With p = 0.03, the researcher likely used α = 0.05, meaning there's a 5% chance the significant result occurred by chance alone. Answer B correctly identifies this possibility. Let's examine why the other options miss the mark. Answer A incorrectly assumes that statistical significance and adequate power eliminate all possibility of error—this is never true in hypothesis testing. Answer C confuses the error types: Type II error occurs when you fail to detect a real effect, but here we have a significant result, so Type II error didn't occur. The 70% power is actually reasonable, not "insufficient." Answer D misinterprets what the 30% represents—this would be the Type II error probability if we had failed to find significance, but it doesn't apply to our significant result. Remember this key distinction: when you find statistical significance, you risk Type I error (false positive) with probability equal to your alpha level. When you fail to find significance, you risk Type II error (false negative) with probability equal to 1 minus the power.

Question 20

A pharmaceutical company conducts multiple hypothesis tests on the same dataset to evaluate drug safety. If they perform 20 independent tests, each at α = 0.05, and all null hypotheses are actually true, what is the approximate probability of making at least one Type I error?

  1. 0.05, since each individual test maintains the same significance level
  2. 0.64, calculated as 1 - (0.95)^20 accounting for multiple comparisons (correct answer)
  3. 0.36, calculated as (0.95)^20 representing no Type I errors
  4. 1.00, since performing multiple tests guarantees at least one significant result
  5. 0.25, representing the average probability across all tests performed simultaneously
Explanation: When you encounter multiple hypothesis testing problems, you're dealing with the family-wise error rate (FWER) - the probability of making at least one Type I error across all tests. This is different from the individual test error rate. To find the probability of at least one Type I error, it's easier to calculate the complement: the probability of making NO Type I errors across all tests. For each individual test with α = 0.05, the probability of correctly not rejecting a true null hypothesis is (1 - 0.05) = 0.95. Since the 20 tests are independent, the probability of making no Type I errors across all tests is (0.95)20=0.358(0.95)^{20} = 0.358. Therefore, the probability of making at least one Type I error is 1(0.95)20=10.358=0.6420.641 - (0.95)^{20} = 1 - 0.358 = 0.642 ≈ 0.64. Answer A incorrectly assumes the family-wise error rate equals the individual test significance level - a common misconception that ignores the cumulative effect of multiple testing. Answer C gives you the probability of making NO Type I errors, which is the complement of what the question asks. Answer D is wrong because even with multiple testing, there's still a chance (about 36%) of making no Type I errors. Remember this key pattern: as the number of tests increases, your chance of finding at least one "significant" result by chance alone increases dramatically, even when all null hypotheses are true. This is why multiple comparison corrections like Bonferroni are essential in practice.