All questions
Question 1
An initial experiment testing a new cognitive training program against a control group yields a statistically significant result, with a p-value of 0.03 for the improvement in test scores. Based on this p-value, what can be said about the probability that an exact replication of the experiment (with the same sample size and from the same population) will also produce a statistically significant result (p < 0.05)?
- The probability is 0.97.
- The probability is 0.95.
- The probability is exactly 1 minus the p-value from the replication study, which is unknown.
- The probability is high, but it cannot be determined from the p-value alone, as it depends on the study's statistical power. (correct answer)
Explanation: When you encounter questions about replicating statistically significant results, remember that a single p-value tells you about one specific experiment, not about future replications. The key concept here is statistical power - the probability of detecting a true effect when it actually exists.
The correct answer is D because the probability of replicating a significant result depends on the study's statistical power, which cannot be determined from the p-value alone. Statistical power depends on factors like effect size, sample size, and the significance level chosen. Even with a significant result (p = 0.03), you need to know the underlying effect size and study design to calculate the probability of replication success.
Option A incorrectly assumes the replication probability is 1 - 0.03 = 0.97. This reflects a common misconception that p-values directly translate to replication probabilities. Option B (0.95) might stem from confusing the significance level (α = 0.05) with replication probability, but these are entirely different concepts. Option C is problematic because it suggests the replication probability equals 1 minus the future study's p-value, which is circular reasoning - you can't predict a future p-value to calculate replication probability.
The fundamental issue is that p-values are calculated assuming the null hypothesis is true, while replication probability depends on the alternative hypothesis being true and the study's ability to detect that true effect.
Study tip: Remember that p-values describe the current data under the null hypothesis, while replication depends on statistical power under the alternative hypothesis. These are completely different statistical concepts that students often confuse.
Question 2
A 95% confidence interval for the difference in mean test scores between two teaching methods (Method A - Method B) is found to be [-2.4 points, 5.8 points]. Based solely on this interval, what can be concluded about the result of a two-sided hypothesis test of H₀: μ_A - μ_B = 0 versus Hₐ: μ_A - μ_B ≠ 0 at the α = 0.05 significance level?
- The null hypothesis would be rejected, and the p-value would be less than 0.05, because the interval is not symmetric around zero.
- The null hypothesis would not be rejected, and the p-value would be greater than 0.05, because the interval contains the value 0. (correct answer)
- The test is inconclusive because the interval contains both negative and positive values, so the p-value cannot be determined.
- The null hypothesis would be rejected, and the p-value would be less than 0.05, because the sample difference was likely positive.
Explanation: The correct answer accurately describes the duality between confidence intervals and two-sided hypothesis tests. If a (1-α) confidence interval for a parameter contains the value specified in the null hypothesis (in this case, 0), then the null hypothesis would not be rejected at the α significance level, and the corresponding p-value will be greater than α.
A is incorrect because the symmetry of the interval around zero is irrelevant to the conclusion. The key is whether the null value is contained within it.
C is incorrect because the test is not 'inconclusive'; the formal conclusion is 'fail to reject the null hypothesis.' The relationship between the CI and p-value is well-defined.
D is incorrect. While the sample difference was likely positive (since the interval's midpoint is 1.7), this does not lead to rejecting H₀. Rejection depends on whether the entire interval excludes the null value.
Question 3
A computer simulates drawing a sample and calculates a 95% confidence interval for the population mean, μ. The population mean is known to be μ = 100. The calculated interval for this one sample is (98.5, 101.2).
What is the probability that the true mean, μ = 100, is contained within this specific interval of (98.5, 101.2)?
- 0.95
- 1 (correct answer)
- 0.05
- It cannot be determined from the information given.
Explanation: This is a trick question that tests the fundamental definition of a confidence interval. Once a specific interval is calculated, the true parameter is either inside it or it is not. There is no longer any probability involved. Since the value 100 is clearly inside the numerical interval (98.5, 101.2), the probability that it is contained in this interval is 1 (i.e., it is a certainty).
A is the most common mistake, confusing the confidence level (a property of the procedure) with the probability for a specific outcome.
C is incorrect.
D is incorrect because the question can be answered by simple inspection.
Question 4
A study on workplace wellness programs yields a 95% confidence interval for the mean change in employee sick days per year of [-4.1 days, 0.5 days]. The company plans to replicate the study with a much larger sample size next year. Assume the new study finds the exact same sample mean change and sample standard deviation as the original study.
Compared to the original study, what is the most likely outcome for the new 95% confidence interval and the new p-value for testing H₀: μ_change = 0?
- The new confidence interval will be wider, and the p-value will be larger.
- The new confidence interval will be narrower, but the p-value will not change because the observed effect did not change.
- The new confidence interval will be narrower, and the p-value will be smaller. (correct answer)
- The new confidence interval will have a different center, and the p-value will be smaller.
Explanation: The correct answer correctly identifies the effect of increasing sample size (n). The width of a confidence interval is inversely proportional to the square root of n. Therefore, a larger n will result in a narrower, more precise interval. The test statistic (e.g., t or z) is also proportional to the square root of n. A larger n will lead to a larger test statistic for the same effect size, which in turn leads to a smaller p-value.
A is incorrect as it reverses the effect of sample size.
B incorrectly assumes the p-value is independent of sample size.
D is incorrect because the center of the interval is the sample mean, which the problem states is the same in the new study.
Question 5
A researcher constructs a 90% confidence interval for a population mean based on a sample of data. Unsatisfied with the level of confidence, the researcher decides to recalculate the interval at a 99% confidence level using the exact same sample data. Which of the following statements correctly describes the new interval and a related consequence?
- The 99% confidence interval will be narrower than the 90% interval, providing a more precise estimate of the mean.
- The 99% confidence interval will be wider than the 90% interval, and a corresponding hypothesis test would use a smaller significance level (α). (correct answer)
- The center of the 99% confidence interval will be a different value than the center of the 90% confidence interval.
- The 99% confidence interval will be wider than the 90% interval, increasing the chance of making a Type I error in a related hypothesis test.
Explanation: The correct answer identifies two key consequences of increasing the confidence level. First, to be more confident that the interval captures the true parameter, the interval must be made wider (less precise). Second, a (1-α) confidence level corresponds to a significance level of α. Thus, a 99% confidence level (α=0.01) corresponds to a smaller significance level than a 90% confidence level (α=0.10).
A is incorrect because it confuses higher confidence with greater precision; they are inversely related.
C is incorrect because the center of the interval is the sample mean, which does not change when the confidence level is changed.
D is incorrect because a higher confidence level (and thus a smaller α) decreases the chance of a Type I error, it does not increase it.
Question 6
Engineers test a new gasoline formula and find a 95% confidence interval for the improvement in fuel efficiency (New Formula - Old Formula) to be [0.8 mpg, 2.1 mpg]. Which of the following is an invalid interpretation of this result?
- The data provide statistically significant evidence at the 5% level that the new formula improves fuel efficiency.
- The procedure used has a 95% chance of generating an interval that contains the true mean improvement in fuel efficiency.
- A plausible range for the true mean improvement in fuel efficiency is from 0.8 mpg to 2.1 mpg.
- There is a 95% probability that the true mean improvement in fuel efficiency is between 0.8 mpg and 2.1 mpg. (correct answer)
Explanation: Confidence intervals are one of the most misunderstood concepts in statistics, particularly regarding what the "confidence level" actually means. When you encounter confidence interval interpretation questions, focus carefully on whether statements refer to the procedure versus this specific interval.
The correct answer is D because it commits the most common confidence interval fallacy. Once you've calculated a specific interval like [0.8, 2.1], the true parameter either is or isn't in that range—there's no probability involved anymore. The 95% refers to the long-run success rate of the procedure, not the probability for this particular interval.
Let's examine why the other options are valid interpretations:
A is correct because the entire interval lies above zero, indicating statistically significant evidence of improvement at the 5% level. When a confidence interval for a difference doesn't contain zero, it demonstrates statistical significance.
B correctly describes what "95% confidence" actually means—it's about the procedure's reliability. If you repeated this process many times, about 95% of the intervals would capture the true parameter.
C appropriately uses the word "plausible," acknowledging that this interval represents a reasonable range of values for the true mean improvement based on the sample data.
The key distinction is between the procedure (which has a 95% success rate) and any individual interval (which either contains the parameter or doesn't). Remember: confidence levels describe methods, not specific intervals. Watch for this subtle but crucial difference on statistics exams.
Question 7
A statistics professor asks 50 independent research groups to each collect a sample of data and construct a 90% confidence interval for a specific, known population proportion, π = 0.60.
After all 50 confidence intervals are constructed, which of the following outcomes is most plausible?
- Exactly 45 of the 50 intervals (90%) will contain the value 0.60.
- All 50 of the intervals will contain the value 0.60, as they are all estimating the same known parameter.
- Approximately 45 of the 50 intervals will contain the value 0.60, and the other intervals will not. (correct answer)
- Each of the 50 intervals will have a 90% probability of containing the value 0.60.
Explanation: The correct answer reflects the long-run frequency interpretation of the confidence level. The 90% confidence level implies that we expect 90% of the intervals to capture the true parameter. Due to random sampling variability, the actual number will be approximately 90%, not exactly 90%.
A is incorrect because it removes the crucial element of variability. The number of successful intervals is a binomial random variable with n=50 and p=0.90; the expected value is 45, but other outcomes are possible.
B is incorrect because it ignores sampling error. Different samples will produce different intervals, and some will, by chance, fail to capture the true parameter.
D is incorrect because it applies the 90% probability to individual intervals after they have been constructed. Once an interval is calculated, the true parameter is either in it or not (a probability of 1 or 0). The 90% refers to the procedure, not the outcome of a specific interval.
Question 8
A researcher investigates the effect of a new supplement on two separate outcomes in a group of athletes: average sprint time (Test A) and vertical jump height (Test B). The hypothesis test for the effect on sprint time yields a p-value of 0.01. The hypothesis test for the effect on vertical jump height yields a p-value of 0.04.
Based on a direct comparison of these two p-values, what is the most appropriate conclusion?
- The supplement has a stronger and more practically important effect on sprint time than it does on vertical jump height.
- There is stronger statistical evidence against the null hypothesis for sprint time than there is for vertical jump height. (correct answer)
- The probability that the supplement has no effect on sprint time is 1%, while the probability that it has no effect on jump height is 4%.
- Since both p-values are less than 0.05, the evidence for the effect on sprint time and vertical jump height is equally strong.
Explanation: The correct answer properly interprets the relative meaning of p-values. A smaller p-value indicates that the observed data are more surprising under the null hypothesis, which translates to stronger statistical evidence against the null hypothesis. However, this does not allow for conclusions about the relative sizes or practical importance of the effects.
A is incorrect because it confuses statistical evidence (p-value) with the magnitude or practical importance of the effect size.
C is incorrect as it misinterprets the p-value as the probability of the null hypothesis being true (P(H₀ | data)).
D is incorrect because it engages in 'dichotomous thinking' (significant/not significant) and fails to recognize that, within the significant range, a smaller p-value represents stronger evidence.
Question 9
Two independent studies investigate the same new antidepressant. Study A, with 100 participants, finds a p-value of 0.04 for the reduction in depression scores. Study B, with 1,000 participants, finds a p-value of 0.08. Both studies used a significance level of α = 0.05.
A science journalist reports that Study A 'proved the drug is effective' while Study B 'showed the drug is ineffective.' Which statement provides the most accurate critique of this reporting?
- The reporting is flawed; Study A found a statistically significant result, but this does not 'prove' effectiveness. Study B failed to reject the null hypothesis, which is not the same as proving the drug is ineffective. (correct answer)
- The reporting is accurate because Study A's p-value is below the 0.05 threshold, indicating a real effect, while Study B's is above it, indicating no effect.
- The reporting is flawed because Study A's result is less reliable due to the smaller sample size, and Study B's higher p-value suggests the drug is actually safe.
- The reporting is accurate because the effect of the drug was clearly stronger in Study A, as evidenced by its much lower p-value.
Explanation: The correct answer accurately points out two common pitfalls: equating statistical significance with 'proof' and interpreting 'failing to reject the null' as 'proving the null is true.' Statistical inference deals with evidence, not certainty.
B is incorrect because it endorses the reporter's flawed, black-and-white interpretation of the p < 0.05 threshold without acknowledging the nuances.
C is incorrect because a higher p-value has no bearing on the drug's safety, and while sample size affects power, it doesn't automatically make the result 'unreliable.'
D is incorrect because it confuses statistical significance (p-value) with effect size. A smaller p-value does not necessarily mean a larger or more important effect, especially when sample sizes differ.
Question 10
An environmental scientist conducts two studies to test for a certain pollutant in a river. Study X uses a sample of n=25 water specimens and finds a mean pollutant level higher than the safety standard, with p = 0.07. Study Y uses a sample of n=2500 specimens and finds a mean level just slightly higher than the standard, with p = 0.02. Which is the most valid conclusion?
- The pollutant level found in Study Y is of greater practical concern because its p-value is much smaller.
- The effect size (the actual amount by which the pollutant level exceeds the standard) is likely larger in Study X than in Study Y. (correct answer)
- The results from both studies are unreliable because they produced different p-values for the same river.
- The effect size is definitely larger in Study Y, as proven by the smaller p-value and larger sample size.
Explanation: The correct answer recognizes the relationship between sample size, effect size, and p-value. A very large sample size (as in Study Y) can make a very small, practically insignificant effect statistically significant (a small p-value). Conversely, achieving a p-value that is close to significant (p=0.07) with a very small sample (as in Study X) suggests that the observed effect size was likely substantial. Thus, it is probable that the effect size was larger in Study X.
A is incorrect because it confuses statistical significance with practical importance. A tiny, harmless effect can be statistically significant with a large enough sample.
C is incorrect because sampling variability means that different studies will almost always produce different p-values, even if they are studying the same phenomenon.
D is incorrect because it makes the classic error of equating a smaller p-value with a larger effect size.
Question 11
In a clinical trial for a new allergy medication, the null hypothesis is that the medication has no effect on allergy symptoms. The trial results in a p-value of 0.02. Which of the following statements provides the most precise interpretation of this p-value?
- There is a 2% chance that the medication is ineffective.
- There is a 98% chance that the medication is effective.
- If the medication were truly ineffective, there would be a 2% chance of observing a reduction in symptoms as large as, or larger than, what was found in this trial. (correct answer)
- The probability of concluding that the medication is effective when it is actually not (a Type I error) is exactly 2%.
Explanation: The correct answer provides the precise, formal definition of a p-value: the probability of observing the current data, or more extreme data, under the assumption that the null hypothesis is true.
A is incorrect because it misinterprets the p-value as P(H₀ | data), the probability that the null hypothesis is true given the data.
B is incorrect because it misinterprets 1-p as P(Hₐ | data), the probability that the alternative hypothesis is true given the data.
D is incorrect because it confuses the p-value with the significance level, α. The probability of a Type I error is α, which is a pre-set threshold. The p-value is a property of the data, not the decision rule.
Question 12
A well-designed randomized controlled trial compared a new reading curriculum to a standard curriculum. After a semester, the analysis found no statistically significant difference in reading comprehension scores between the two groups (p = 0.35). Which is the most appropriate conclusion to draw?
- The study provides conclusive evidence that the new curriculum and the standard curriculum are equally effective.
- There is a 35% probability that the two curricula are equally effective.
- The study did not find sufficient evidence to conclude that one curriculum is more effective than the other. (correct answer)
- The standard curriculum is likely superior to the new one, as the study failed to show any benefit from the new curriculum.
Explanation: The correct answer properly interprets a non-significant result. 'Failing to reject the null hypothesis' means the data did not provide enough evidence to support the alternative hypothesis. It does not prove the null hypothesis is true.
A is incorrect because it commits the error of accepting the null hypothesis. Absence of evidence is not evidence of absence.
B is incorrect because it misinterprets the p-value as the probability that the null hypothesis is true (P(H₀ | data)).
D is incorrect because a failure to find a significant difference does not imply the new curriculum is worse. It simply means the study couldn't detect a statistically significant difference, which could be because there is no difference, the difference is small, or the study lacked statistical power.
Question 13
A researcher is testing the null hypothesis H₀: μ = 30 based on sample data where the sample mean is x̄ = 34 and the standard error of the mean is 2. The resulting test statistic is z = (34 - 30) / 2 = 2.0.
Regarding the p-value for this test, which of the following statements is correct?
- The p-value will be smaller if the alternative hypothesis is Hₐ: μ > 30 than if it is Hₐ: μ ≠ 30. (correct answer)
- The p-value depends on the chosen significance level, α, not on the alternative hypothesis.
- The p-value represents the probability that the true population mean is actually 30.
- If the sample mean had been x̄ = 26, the p-value for the test against Hₐ: μ ≠ 30 would be different.
Explanation: The correct answer recognizes that the p-value calculation depends on the alternative hypothesis. For a positive test statistic (z=2.0), the p-value for a one-sided test (Hₐ: μ > 30) is P(Z ≥ 2.0). The p-value for a two-sided test (Hₐ: μ ≠ 30) is 2 * P(Z ≥ 2.0). Therefore, the one-sided p-value is half the size of the two-sided p-value.
B is incorrect; the p-value is calculated from the data and is then compared to α. It does not depend on α.
C is the common P(H₀ | data) fallacy.
D is incorrect because if x̄ = 26, the z-statistic would be (26-30)/2 = -2.0. For a two-sided test, the p-value is based on the magnitude of z, so P(|Z| ≥ 2.0) is the same as for z = +2.0.
Question 14
Two agencies, Agency A and Agency B, are evaluating the effectiveness of a new public safety initiative. They both analyze the exact same dataset, which yields a p-value of 0.03 for the reduction in crime rates. Agency A adheres to a strict α = 0.01 significance level due to the high cost of implementation. Agency B uses the more conventional α = 0.05 significance level.
Based on their respective standards, what will each agency conclude about the statistical significance of the initiative's effect?
- Agency A will conclude the effect is significant, while Agency B will conclude it is not significant.
- Both agencies will conclude the effect is significant because the p-value is small.
- Neither agency will conclude the effect is significant because the p-value is greater than 0.01.
- Agency A will conclude the effect is not significant, while Agency B will conclude it is significant. (correct answer)
Explanation: The correct answer applies the decision rule for hypothesis testing: reject the null hypothesis if p ≤ α. For Agency A, p = 0.03 is not less than or equal to α = 0.01, so they fail to reject H₀ and find the result not statistically significant. For Agency B, p = 0.03 is less than or equal to α = 0.05, so they reject H₀ and find the result statistically significant.
A reverses the conclusions.
B ignores Agency A's stated significance level.
C ignores Agency B's stated significance level.
Question 15
From a large survey, a 95% confidence interval for the mean number of books read annually by adults in a country is [10.4, 13.6]. A student reads this result and concludes, 'This means that 95% of the adults in this country read between 10.4 and 13.6 books per year.' What is the primary flaw in the student's interpretation?
- The student's interpretation is correct and represents the main purpose of constructing a confidence interval.
- The confidence level is too low; a 99% confidence interval would be required to make claims about the individuals in the population.
- The student is confusing a confidence interval for a population mean with a range that describes the variability of individuals in the population. (correct answer)
- The student should have used the term 'probability' instead of '95% of the adults' for the interpretation to be correct.
Explanation: The correct answer identifies the common confusion between a confidence interval and a prediction interval or a simple descriptive range. A confidence interval provides a range of plausible values for a population parameter (like the mean), not for individual data points. The variation among individuals is typically much larger than the uncertainty in the estimate of the mean.
A is incorrect because the student's interpretation is a classic mistake.
B is incorrect because changing the confidence level does not change what the interval estimates (it still estimates the mean, not individual values).
D is incorrect because swapping terms would create a different, but still incorrect, interpretation (the Bayesian fallacy).
Question 16
A city implemented a costly program to reduce emergency response times. The null hypothesis is that the program has no effect. An analysis after one year yields a p-value of 0.15 for the change in response times.
A city manager argues, 'The p-value of 0.15 means there is an 85% chance the program is working (1 - 0.15 = 0.85). We should continue to fund it.' Which statement best describes the flaw in the manager's reasoning?
- The manager's reasoning is correct; an 85% chance of effectiveness is a strong argument for continuing the program.
- The manager should have calculated a 95% confidence interval instead, which would directly show the probability of the program's success.
- The manager is incorrectly interpreting 1-p as the probability that the alternative hypothesis is true. The p-value is calculated assuming the program has no effect. (correct answer)
- The manager is correct that there is an 85% chance the program is working, but this is too low of a probability to justify continued funding.
Explanation: The correct answer pinpoints the exact logical error: interpreting 1-p as P(Hₐ | data). The p-value is P(data or more extreme | H₀ is true). These two quantities are not directly related in this way. The manager is committing a common p-value fallacy.
A incorrectly endorses the manager's flawed reasoning.
B suggests a confidence interval would provide a probability of success, which is also incorrect. A CI provides a range of plausible values for the effect, but the confidence level does not represent the probability of success for the program.
D accepts the manager's flawed probabilistic reasoning and only quarrels with the conclusion, thereby failing to identify the statistical mistake.
Question 17
Based on a very large national survey (n > 20,000), a 95% confidence interval for the proportion of the population that supports a particular piece of legislation is determined to be [0.512, 0.528]. Which of the following is the most astute interpretation of this finding?
- The legislation is extremely popular, as the lower bound of the confidence interval is well above 0.50.
- The result is not practically significant because even though the interval is above 0.50, the true level of support is very close to an even split.
- We can be 95% confident that the true proportion of support is exactly the midpoint of the interval, 0.520.
- The data provide strong statistical evidence that the true proportion of support is greater than 0.50, but the practical importance of this small majority may be debatable. (correct answer)
Explanation: When interpreting confidence intervals, you need to distinguish between statistical significance and practical significance, especially with large sample sizes like this one.
The confidence interval [0.512, 0.528] tells us several important things. Since the entire interval lies above 0.50, we have strong statistical evidence that true support exceeds 50% - this rules out the possibility of an even split or minority support. With n > 20,000, this finding is statistically robust and unlikely due to sampling error.
However, the interval only ranges from 51.2% to 52.8%, representing a fairly modest majority. While statistically significant, whether a 1-3 percentage point lead over 50% constitutes meaningful practical significance depends on context and perspective.
Answer D correctly captures both aspects: acknowledging the strong statistical evidence for majority support while recognizing that the practical importance of such a slim majority is debatable.
Answer A overstates the finding by calling legislation with 51-53% support "extremely popular" - this represents only a small majority. Answer B incorrectly dismisses the statistical significance; when an entire confidence interval excludes 0.50, this provides clear evidence against an even split. Answer C misunderstands confidence intervals fundamentally - we cannot be 95% confident about any single point value, including the midpoint. The interval represents a range of plausible values for the true proportion.
Remember: Large samples can detect small differences that are statistically significant but may not be practically meaningful. Always consider both the statistical evidence and the real-world importance of your findings.
Question 18
A 95% confidence interval for the difference in mean test scores between two teaching methods (Method A - Method B) is found to be [-2.4 points, 5.8 points]. Based solely on this interval, what can be concluded about the result of a two-sided hypothesis test of H₀: μ_A - μ_B = 0 versus Hₐ: μ_A - μ_B ≠ 0 at the α = 0.05 significance level?
- The null hypothesis would be rejected, and the p-value would be less than 0.05, because the interval is not symmetric around zero.
- The null hypothesis would not be rejected, and the p-value would be greater than 0.05, because the interval contains the value 0. (correct answer)
- The test is inconclusive because the interval contains both negative and positive values, so the p-value cannot be determined.
- The null hypothesis would be rejected, and the p-value would be less than 0.05, because the sample difference was likely positive.
Explanation: The correct answer accurately describes the duality between confidence intervals and two-sided hypothesis tests. If a (1-α) confidence interval for a parameter contains the value specified in the null hypothesis (in this case, 0), then the null hypothesis would not be rejected at the α significance level, and the corresponding p-value will be greater than α.
A is incorrect because the symmetry of the interval around zero is irrelevant to the conclusion. The key is whether the null value is contained within it.
C is incorrect because the test is not 'inconclusive'; the formal conclusion is 'fail to reject the null hypothesis.' The relationship between the CI and p-value is well-defined.
D is incorrect. While the sample difference was likely positive (since the interval's midpoint is 1.7), this does not lead to rejecting H₀. Rejection depends on whether the entire interval excludes the null value.
Question 19
In a clinical trial for a new allergy medication, the null hypothesis is that the medication has no effect on allergy symptoms. The trial results in a p-value of 0.02. Which of the following statements provides the most precise interpretation of this p-value?
- There is a 2% chance that the medication is ineffective.
- There is a 98% chance that the medication is effective.
- If the medication were truly ineffective, there would be a 2% chance of observing a reduction in symptoms as large as, or larger than, what was found in this trial. (correct answer)
- The probability of concluding that the medication is effective when it is actually not (a Type I error) is exactly 2%.
Explanation: The correct answer provides the precise, formal definition of a p-value: the probability of observing the current data, or more extreme data, under the assumption that the null hypothesis is true.
A is incorrect because it misinterprets the p-value as P(H₀ | data), the probability that the null hypothesis is true given the data.
B is incorrect because it misinterprets 1-p as P(Hₐ | data), the probability that the alternative hypothesis is true given the data.
D is incorrect because it confuses the p-value with the significance level, α. The probability of a Type I error is α, which is a pre-set threshold. The p-value is a property of the data, not the decision rule.
Question 20
A well-designed randomized controlled trial compared a new reading curriculum to a standard curriculum. After a semester, the analysis found no statistically significant difference in reading comprehension scores between the two groups (p = 0.35). Which is the most appropriate conclusion to draw?
- The study provides conclusive evidence that the new curriculum and the standard curriculum are equally effective.
- There is a 35% probability that the two curricula are equally effective.
- The study did not find sufficient evidence to conclude that one curriculum is more effective than the other. (correct answer)
- The standard curriculum is likely superior to the new one, as the study failed to show any benefit from the new curriculum.
Explanation: The correct answer properly interprets a non-significant result. 'Failing to reject the null hypothesis' means the data did not provide enough evidence to support the alternative hypothesis. It does not prove the null hypothesis is true.
A is incorrect because it commits the error of accepting the null hypothesis. Absence of evidence is not evidence of absence.
B is incorrect because it misinterprets the p-value as the probability that the null hypothesis is true (P(H₀ | data)).
D is incorrect because a failure to find a significant difference does not imply the new curriculum is worse. It simply means the study couldn't detect a statistically significant difference, which could be because there is no difference, the difference is small, or the study lacked statistical power.