All questions
Question 1
A bank is considering a new loan approval algorithm. In a test on 50,000 applications, the new algorithm increased the profitability per loan by an average of $2.50. The result is statistically significant (p = 0.04). Which piece of information is LEAST relevant when deciding whether this result is practically significant?
- The cost to develop and implement the new algorithm.
- The total number of loans the bank processes annually.
- The p-value from a replication study using a different sample of applications. (correct answer)
- The standard deviation of profitability per loan.
Explanation: When you encounter questions about practical significance versus statistical significance, remember that practical significance focuses on real-world impact and business value, while statistical significance only tells you whether a result is likely due to chance.
The correct answer is C because a p-value from a replication study doesn't help determine practical significance of the current result. While replication is valuable for confirming statistical validity, it doesn't change whether the $2.50 increase per loan is practically meaningful for this bank's business operations.
Let's examine why the other options are highly relevant to practical significance. Option A (implementation costs) is crucial because you need to weigh the 2.50perloanbenefitagainstdevelopmentexpensestodetermineifthealgorithmisworthpursuing.OptionB(annualloanvolume)isessentialforcalculatingtotalimpact—2.50 per loan means vastly different things if the bank processes 1,000 versus 100,000 loans annually. Option D (standard deviation of profitability) matters because it provides context for the $2.50 improvement; this amount might be practically negligible if typical loan profitability varies by hundreds of dollars, but highly significant if it typically varies by only a few dollars.
Remember this distinction: statistical significance (p-values, confidence intervals) tells you about the reliability of your findings, while practical significance requires business context like costs, scale, and variability. On business statistics exams, questions about practical significance will always point toward real-world business considerations rather than purely statistical measures. Question 2
A software company tested whether a new user interface design improved task completion rates. The study involved 400 users randomly assigned to either the current interface or the new design. Users with the new interface completed tasks successfully 78.5% of the time compared to 75.0% with the current interface (p = 0.08). Industry standards suggest that interface improvements should increase completion rates by at least 5 percentage points to justify the disruption of user workflow and retraining costs.
Based on these results, which statement about statistical and practical significance is most accurate?
- The results demonstrate practical significance meeting industry standards, but lack statistical significance, suggesting the need for a larger sample size to confirm the effect.
- The results show marginal statistical significance with strong practical significance, supporting implementation despite the borderline p-value given the meaningful improvement size.
- The results lack both statistical significance and practical significance, as the improvement falls short of industry standards and cannot be distinguished from random variation. (correct answer)
- The results demonstrate neither statistical nor practical significance definitively, requiring additional testing with modified designs to achieve both industry standards and statistical validation.
Explanation: When you encounter hypothesis testing results, you need to evaluate both statistical significance (can we trust the effect is real?) and practical significance (is the effect large enough to matter in practice). Statistical significance depends on the p-value relative to your significance level, while practical significance compares the observed effect to meaningful thresholds.
Let's analyze these results systematically. For statistical significance, the p-value is 0.08, which exceeds the conventional 0.05 threshold, meaning we cannot confidently conclude the difference is real rather than due to random sampling variation. For practical significance, the improvement is 78.5% - 75.0% = 3.5 percentage points, which falls short of the 5 percentage point industry standard needed to justify implementation costs.
Answer A incorrectly claims practical significance exists when the 3.5 point improvement doesn't meet the 5 point standard. Answer B makes two errors: calling p = 0.08 "marginal statistical significance" when it actually indicates lack of statistical significance, and wrongly asserting "strong practical significance" despite failing to meet industry standards. Answer D suggests the results are inconclusive, but we can definitively say both criteria are unmet based on the given standards.
Answer C correctly identifies that both significance types are lacking - the p-value of 0.08 means the effect cannot be distinguished from random variation, and the 3.5 percentage point improvement falls short of industry standards.
Study tip: Always check both significance types independently. A large sample can make tiny, meaningless differences statistically significant, while small samples can hide practically important effects.
Question 3
An environmental consultant evaluated two air filtration systems for a manufacturing facility. System X reduced airborne particulates by an average of 12% with p = 0.001 based on 6 months of monitoring. System Y reduced particulates by 8% with p = 0.15 over the same period. Environmental regulations require at least 10% reduction for compliance, and System X costs $200,000 while System Y costs $80,000. Both systems have similar maintenance requirements.
How should the facility manager interpret these findings when considering statistical and practical significance for regulatory compliance and cost management?
- Neither system demonstrates sufficient practical significance for implementation, as the cost per percentage point of reduction exceeds industry benchmarks regardless of statistical significance levels.
- Both systems show practical significance for cost-effectiveness, but only System X provides statistical significance, making the choice dependent on risk tolerance for regulatory compliance.
- System X shows statistical significance with questionable practical significance given costs, while System Y demonstrates practical significance without statistical validation, requiring further evaluation of both options.
- System X demonstrates both statistical and practical significance for regulatory compliance, while System Y offers cost advantages but lacks both statistical validation and regulatory compliance assurance. (correct answer)
Explanation: When evaluating business decisions involving statistical tests, you need to distinguish between statistical significance (reliability of results) and practical significance (real-world importance), then consider how both relate to your specific business constraints.
System X achieves 12% reduction with p = 0.001, meaning there's only a 0.1% chance these results occurred by random variation—this demonstrates strong statistical significance. Since 12% exceeds the 10% regulatory requirement, it also shows practical significance for compliance. System Y shows 8% reduction with p = 0.15, meaning a 15% chance the results are due to random variation—this fails to meet conventional statistical significance thresholds (typically p < 0.05) and falls short of the 10% regulatory requirement.
Answer D correctly identifies that System X demonstrates both statistical and practical significance for regulatory compliance, while System Y offers cost advantages but lacks statistical validation and regulatory compliance assurance.
Answer A incorrectly dismisses both systems based on undefined "industry benchmarks" not mentioned in the problem. Answer B wrongly claims both systems show practical significance—System Y doesn't meet regulatory requirements. Answer C incorrectly suggests System X has "questionable practical significance" when it clearly meets regulatory needs, and wrongly credits System Y with practical significance despite failing compliance requirements.
Study tip: In business statistics questions, always evaluate statistical significance (p-values), practical significance (meeting real-world thresholds), and business constraints separately before combining them into recommendations. Don't let cost considerations override fundamental statistical validity or regulatory compliance requirements.
Question 4
An educational researcher studying class sizes finds that reducing class size from 28 to 25 students results in a statistically significant improvement in test scores (p = 0.03), with mean scores increasing from 78.2 to 79.1 points. The cost to implement this reduction district-wide would be $2.3 million annually. Educational research suggests score improvements of at least 3 points are needed to impact long-term academic outcomes meaningfully. What challenge does this scenario best illustrate regarding significance interpretation?
- Large sample sizes can make trivially small differences appear statistically significant, requiring careful evaluation of whether effects justify implementation costs. (correct answer)
- Small sample sizes often fail to detect important practical differences, leading to Type II errors that prevent beneficial policy implementations.
- Statistical significance always aligns with practical significance when proper experimental controls are maintained and measurement instruments are validated.
- Practical significance can only be evaluated after statistical significance is established, making sequential testing the optimal approach for policy research.
Explanation: The correct answer is A. This scenario demonstrates how large samples can detect statistically significant but practically insignificant differences. The 0.9-point improvement is statistically significant but falls well below the 3-point threshold for meaningful academic impact, yet would cost $2.3 million annually. This illustrates the importance of considering effect size and practical importance, not just p-values. Choice B is incorrect because the issue here is detecting differences that are too small to matter, not failing to detect important ones. Choice C is wrong because statistical and practical significance often diverge. Choice D is incorrect because practical significance can and should be evaluated independently of statistical significance.
Question 5
A financial analyst discovers that a new investment strategy produces returns that are statistically significantly different from the market benchmark (p = 0.001) based on 5 years of data. However, after accounting for transaction costs and management fees, the strategy's net returns are only 0.2% higher annually than a low-cost index fund. The strategy requires active management costing 1.8% annually. Which statement best reflects the relationship between statistical and practical significance in this context?
- Strong statistical significance combined with negative practical significance demonstrates why p-values alone are insufficient for investment decision-making and cost-benefit analysis. (correct answer)
- Strong statistical significance with marginal practical significance suggests the strategy should be implemented with reduced management fees to improve net outcomes.
- The combination of statistical and practical significance provides compelling evidence for strategy adoption, with management costs justified by risk-adjusted performance improvements.
- Weak practical significance despite strong statistical significance indicates measurement error in the underlying data and suggests additional validation studies are needed.
Explanation: The correct answer is A. While the strategy shows strong statistical significance (p = 0.001), the practical significance is actually negative: the strategy yields 0.2% higher returns but costs 1.8% in fees, resulting in a net loss of 1.6% annually compared to the index fund. This perfectly illustrates why statistical significance doesn't guarantee practical value. Choice B incorrectly suggests positive practical significance exists. Choice C wrongly concludes both types of significance support adoption. Choice D incorrectly attributes the poor practical performance to measurement error rather than recognizing that statistically significant differences can still be practically worthless or harmful.
Question 6
A marketing analyst finds that a new advertising campaign increased website conversion rates from 2.1% to 2.3% with a p-value of 0.001 based on 50,000 website visitors. The company's revenue per conversion is $25, and the campaign costs $1,500 per month. Which conclusion about significance is most appropriate?
- Strong statistical significance with questionable practical significance requires cost-benefit analysis to determine campaign viability and business impact. (correct answer)
- Weak statistical significance combined with strong practical significance suggests the need for additional data collection before implementation decisions.
- Both statistical and practical significance are strong, providing clear justification for continuing the advertising campaign without further analysis.
- Neither statistical nor practical significance is demonstrated, indicating the campaign should be discontinued immediately to avoid further losses.
Explanation: The correct answer is A. The p-value of 0.001 shows strong statistical significance, but the 0.2 percentage point increase (from 2.1% to 2.3%) represents only about a 9.5% relative improvement. With 50,000 visitors, this means 100 additional conversions worth $2,500 monthly revenue against $1,500 costs - requiring careful cost-benefit analysis for practical significance. Choice B is wrong because statistical significance is strong (p = 0.001). Choice C is incorrect because practical significance isn't clearly established without considering all costs and long-term effects. Choice D is wrong because statistical significance is clearly demonstrated.
Question 7
A small firm pilots a new sales training program with 10 employees. After the program, the sales of these 10 employees increased by an average of 18%, a much larger increase than that of the 30 employees not in the program. However, a t-test comparing the groups yields a p-value of 0.15. What is the most likely interpretation of this outcome?
- The program may have a large, practically significant effect, but the study was likely underpowered to detect it statistically due to the small sample size. (correct answer)
- The program has no meaningful effect on sales, as demonstrated by the failure to achieve statistical significance at a conventional alpha level.
- The 18% increase is likely a result of random sampling error, and therefore the program should be considered practically insignificant.
- The lack of statistical significance implies that any observed sales increase was minor and does not have practical importance for the firm.
Explanation: The correct answer is A. This scenario highlights the classic tension between a large observed effect and a lack of statistical power. An 18% sales increase is almost certainly practically significant. However, with a very small sample size (n=10), the statistical test may not have enough power to rule out random chance as an explanation, leading to a p-value > 0.05. The most appropriate conclusion is that the finding is promising but requires more data.
- B is incorrect because failing to reject the null hypothesis does not prove the null hypothesis is true. It simply means there is insufficient evidence to reject it.
- C incorrectly equates statistical uncertainty (possibility of sampling error) with a lack of practical significance. The magnitude of the effect (18%) is what determines potential practical significance.
- D makes a common mistake by directly linking the lack of statistical significance to a lack of practical importance, which are separate concepts.
Question 8
An e-commerce company with 10 million monthly users tests a new checkout button color. After a one-month trial involving 2 million users (1 million per group), the new color yields a conversion rate of 2.15% compared to the old color's 2.11%. This difference is statistically significant (p = 0.008). Implementing the change requires a one-time IT cost of $50,000. Which conclusion is most appropriate for the business?
- The result is statistically significant, but its practical significance is uncertain until a cost-benefit analysis is performed on the conversion rate lift. (correct answer)
- Since the result is highly statistically significant (p < 0.01), the company should implement the change immediately to capitalize on the proven increase in conversions.
- The practical significance of the result cannot be determined without increasing the sample size to verify that the 0.04% lift is consistent across all users.
- The result lacks practical significance because a 0.04 percentage point increase is too small to represent a real effect on user behavior.
Explanation: The correct answer is A. Statistical significance (p = 0.008) indicates the observed difference is unlikely due to chance. However, practical significance depends on whether the magnitude of the effect is meaningful in a business context. A 0.04% increase in conversion rate might generate enough additional revenue to outweigh the $50,000 cost, or it might not. This requires a cost-benefit analysis, making practical significance uncertain without more information.
- B is incorrect because it mistakenly equates statistical significance with a mandate for business action, ignoring the practical assessment of costs vs. benefits.
- C is incorrect because the sample size is already exceptionally large; more data would only make the p-value smaller, not change the small effect size. The key issue is the business value of that small effect, not its statistical certainty.
- D is incorrect because statistical significance suggests the effect is real (not due to chance). It confuses the existence of an effect with its practical importance.
Question 9
A company is testing a new manufacturing process designed to reduce product defects. A large sample of products is tested, and the 95% confidence interval for the reduction in the defect rate (Old Rate - New Rate) is found to be [0.01%, 0.04%]. Retooling the factory for the new process will cost millions of dollars. Which statement represents the most sound business reasoning?
- The result is statistically significant, but the company must determine if a defect reduction between 0.01% and 0.04% provides a return that justifies the high cost. (correct answer)
- The confidence interval is very narrow, which proves the new process is superior and should be implemented regardless of the upfront cost.
- Since the interval contains values very close to zero, the effect is too small to be practically meaningful, and the process change should be rejected.
- The result is not practically significant because a much larger reduction is needed, so the company should re-run the test with an even larger sample size.
Explanation: The correct answer is A. The confidence interval does not contain zero, indicating the result is statistically significant at the α=0.05 level. The interval provides a range of plausible values for the true effect size: a reduction between 0.01% and 0.04%. The next step is to assess practical significance by comparing the financial benefits of such a reduction against the implementation costs.
- B incorrectly assumes that statistical evidence (a narrow CI not containing zero) automatically translates to a wise business decision, ignoring the crucial cost-benefit analysis.
- C makes a premature judgment on practical significance. For a high-volume product, even a tiny defect reduction could save millions, making it practically significant.
- D is incorrect because a larger sample size would likely only make the already narrow CI even narrower; it wouldn't change the small magnitude of the effect. The decision hinges on the value of that small effect, not on gathering more data.
Question 10
A consultant tells a CEO, 'Our employee satisfaction survey shows the new remote work policy led to a highly statistically significant increase in morale, with a p-value of 0.0001. This extremely low p-value proves that the policy has had a massive and profoundly positive impact.' Which statement best identifies the flaw in the consultant's reasoning?
- The consultant should have used a confidence interval instead of a p-value to prove the impact was massive and positive.
- The consultant cannot make this claim unless the sample included every employee in the company rather than a subset.
- The p-value only measures the probability that the null hypothesis is true, so it cannot be used to claim a positive impact.
- The consultant is confusing a low p-value (high statistical significance) with a large effect size (high practical significance). (correct answer)
Explanation: When evaluating statistical results, you need to distinguish between statistical significance (whether an effect exists) and practical significance (how large or meaningful that effect is). The p-value only tells you about the former.
The consultant's error lies in conflating a very low p-value with a "massive" impact. A p-value of 0.0001 simply means there's strong evidence that some change occurred - it could be a tiny improvement that's statistically detectable due to a large sample size, or it could indeed be a large improvement. The p-value alone cannot tell you which. To determine if the impact is truly "massive," you'd need to examine the effect size - the actual magnitude of change in employee satisfaction scores.
Answer D correctly identifies this fundamental confusion between statistical and practical significance.
Answer A is wrong because confidence intervals, while useful for showing effect size, aren't required to "prove" anything - and the consultant's main error isn't about which statistical tool to use. Answer B is incorrect because well-designed sampling can provide valid insights without surveying every employee; census data isn't required for statistical inference. Answer C misunderstands p-values entirely - they measure the probability of observing your data (or more extreme) assuming the null hypothesis is true, not the probability that the null hypothesis itself is true.
Remember: Low p-values indicate strong evidence of some effect, but tell you nothing about whether that effect is large enough to matter in practice. Always look beyond statistical significance to assess practical importance.
Question 11
A research study on a new productivity software reports that its use resulted in a statistically significant increase in tasks completed per day (p < .01), with an effect size of Cohen's d = 0.10. By convention, a Cohen's d of 0.2 is considered a 'small' effect, 0.5 'medium', and 0.8 'large'. How should a manager interpret these findings?
- The Cohen's d is too small to be meaningful, suggesting the statistically significant result is likely a Type I error.
- The high statistical significance (p < .01) indicates a strong effect, so the software is a worthwhile investment for the company.
- The results are contradictory; high statistical significance suggests a large effect while the Cohen's d suggests a small one.
- The software has a genuine, but very small, effect on productivity that may not be worth the cost of implementation and training. (correct answer)
Explanation: When interpreting research findings, you need to distinguish between statistical significance and practical significance. Statistical significance tells you whether an effect is likely real (not due to chance), while effect size tells you how meaningful that effect is in practice.
The study shows a statistically significant result (p < .01), meaning there's less than a 1% chance the observed increase happened by random chance alone. However, Cohen's d = 0.10 indicates the actual magnitude of improvement is very small—even smaller than the conventional "small" effect threshold of 0.2. This means the software does produce a genuine productivity increase, but it's quite modest.
Option A incorrectly suggests the small effect size indicates a Type I error. Type I errors occur when we falsely conclude there's an effect when none exists, but effect size doesn't determine this—the p-value does. Option B falls into the common trap of equating statistical significance with practical importance. A highly significant p-value doesn't automatically mean a strong or valuable effect. Option C misunderstands that there's no contradiction here—you can have high confidence in detecting a genuinely small effect, especially with large sample sizes.
Option D correctly interprets both pieces of information: the software works (statistical significance) but produces only tiny improvements (small effect size). For a manager, this means the minimal productivity gains might not justify costs like purchasing licenses, training employees, and disrupting workflows.
Remember: Always examine both statistical significance and effect size together. High significance with small effect sizes often signals results that are statistically real but practically questionable.
Question 12
A pilot study (n=40) of a new inventory management software showed it reduced stock-outs by an average of 10%, but the result was not statistically significant (p=0.20). The company then ran a larger study (n=1000) and found the software reduced stock-outs by an average of 2%, with a p-value of 0.01. What is the most accurate conclusion?
- The company should implement the software, as two separate studies both indicated a reduction in stock-outs, confirming its practical value.
- The second study's statistical significance proves the software is effective, and its result should be trusted over the flawed pilot study.
- The results are contradictory and no conclusion can be drawn, as one study shows a large effect and the other shows a small effect.
- The larger study provided a more precise estimate of the effect, which is statistically significant but much smaller than the pilot study suggested. (correct answer)
Explanation: When you encounter questions comparing studies with different sample sizes and statistical significance, focus on understanding what statistical power and precision tell you about effect sizes.
The correct answer is D because it accurately describes what happened across both studies. The pilot study (n=40) found a large effect (10% reduction) but lacked statistical power to detect it reliably (p=0.20). The larger study (n=1000) had much greater statistical power, providing a more precise estimate of the true effect size, which turned out to be smaller (2%) but statistically significant (p=0.01). This is a classic example of how larger samples give more precise estimates and can detect smaller effects.
Answer A is wrong because statistical significance, not just direction of effect, matters for implementation decisions. A 2% reduction may not justify the software's cost. Answer B incorrectly assumes the pilot study was "flawed" and misses that statistical significance doesn't automatically mean practical significance—a 2% improvement might be trivial. Answer C is wrong because the results aren't truly contradictory; they're consistent with the larger study providing a more accurate estimate of a smaller true effect than the pilot study suggested.
Remember this pattern: when comparing studies of different sizes, the larger study typically provides more precise estimates of effect size, while smaller studies often show more variable results. Always distinguish between statistical significance (reliable detection of an effect) and practical significance (whether the effect size matters in real-world terms).
Question 13
A hypothesis test for a new marketing strategy results in a p-value of 0.35. A business manager correctly concludes that there is no statistical significance. Which of the following is an incorrect conclusion to draw from this result?
- 'The observed effect could plausibly be due to random chance.'
- 'The study failed to provide sufficient evidence that the strategy has an effect.'
- 'The test proves that the new marketing strategy has no effect on sales.' (correct answer)
- 'A larger study might be needed to detect a small but potentially valuable effect.'
Explanation: When you encounter hypothesis testing questions with p-values, focus on what the results can and cannot definitively prove. A p-value tells you about the strength of evidence against the null hypothesis, but it has important limitations.
With a p-value of 0.35, you have weak evidence against the null hypothesis (typically, p-values above 0.05 indicate no statistical significance). This means you fail to reject the null hypothesis, but this is critically different from proving the null hypothesis is true.
Option C is incorrect because it commits the classic error of confusing "failure to reject" with "proof of no effect." A non-significant result doesn't prove the marketing strategy has zero effect—it simply means the study didn't detect a statistically significant effect. The true effect could be small, or the study might lack sufficient power to detect it.
Option A is correct because a high p-value (0.35) suggests the observed results could easily occur by random chance alone. Option B accurately describes what non-significant results mean—insufficient evidence for an effect. Option D correctly acknowledges that larger studies have more power to detect smaller effects that might still be practically important.
The key distinction is between statistical evidence and absolute proof. Hypothesis tests can provide evidence against a null hypothesis when p-values are low, but they cannot prove a null hypothesis is true when p-values are high.
Study tip: Remember that "failing to reject" ≠ "proving true." Non-significant results indicate insufficient evidence, not proof of no effect.
Question 14
A food company reformulates a snack to reduce its sodium content by 5mg per serving. A large consumer taste panel (n=2000) compares the new and old versions. The new version scores an average of 0.08 points lower on a 10-point taste scale. This difference, though small, is statistically significant (p=0.02). The company considers any drop greater than 0.5 points to be practically significant. What is the best course of action?
- Rerun the taste panel with a smaller sample, hoping to produce a result that is not statistically significant.
- Do not launch the new formula, because any statistically significant drop in taste scores is a major business risk.
- Launch the new formula, as the health benefit is achieved and the negative impact on taste is statistically significant but not practically significant. (correct answer)
- Launch the new formula, but only in a small test market, as the practical significance of the taste drop is still unknown.
Explanation: This question tests your understanding of the crucial distinction between statistical significance and practical significance—two concepts that often confuse students but are essential in business decision-making.
Statistical significance tells you whether an observed difference is likely real (not due to chance), while practical significance tells you whether that difference matters in the real world. Here, with a massive sample of 2000 people, even tiny differences become statistically significant. The 0.08-point taste drop is statistically significant (p=0.02), meaning it's probably real, but it's far below the company's 0.5-point threshold for practical concern.
The correct approach is C: launch the new formula. You achieve the health benefit of reduced sodium while the taste impact, though measurably real, is practically negligible—well within acceptable limits.
Option A is methodologically dishonest—you can't manipulate sample sizes to chase desired p-values. Option B represents a common business mistake: treating any statistically significant result as automatically important, regardless of magnitude. This misunderstanding costs companies valuable opportunities when they abandon beneficial changes due to trivial but "significant" effects. Option D is unnecessarily cautious since you already have clear evidence the taste drop is below your practical significance threshold.
Study tip: When you see large sample sizes producing small but statistically significant effects, immediately ask yourself about practical significance. Companies often waste resources worrying about statistically significant differences that customers would never notice. Always consider both the statistical evidence AND the business relevance of any finding.
Question 15
Two A/B tests are run to improve website conversion. Test 1 results in a conversion lift of 0.2% with p=0.03. Test 2 results in a conversion lift of 0.8% with p=0.12. Both tests had the same sample size. The manager decides to implement the change from Test 2, stating 'the potential payoff is much larger, even if the result isn't statistically certain yet.' This decision prioritizes:
- Demonstrated statistical significance over potential practical significance.
- Potential practical significance over demonstrated statistical significance. (correct answer)
- Type II error avoidance over Type I error avoidance.
- Short-term certainty over long-term potential gains.
Explanation: When analyzing A/B test results, you need to distinguish between statistical significance (whether results are likely real vs. due to chance) and practical significance (whether results matter in the real world). Statistical significance is measured by p-values, while practical significance relates to the actual size and business impact of the effect.
The manager chose Test 2 despite its higher p-value (0.12 > 0.05, not statistically significant) because the conversion lift of 0.8% is four times larger than Test 1's 0.2% lift. Even though Test 1 has demonstrated statistical significance (p=0.03 < 0.05), its tiny effect size may not justify implementation costs or generate meaningful business value. The manager is betting that Test 2's larger effect, while not yet proven statistically, has greater potential business impact.
This decision prioritizes potential practical significance over demonstrated statistical significance, making B correct. Choice A reverses the relationship—the manager is doing the opposite by choosing practical over statistical considerations. Choice C misapplies error types: Type I errors involve false positives (claiming an effect exists when it doesn't), while Type II errors involve false negatives (missing real effects). The manager isn't primarily concerned with error types here. Choice D incorrectly frames this as a time-based decision, when it's actually about weighing statistical certainty against effect size.
Remember: In business statistics, statistical significance tells you if something is real, but practical significance tells you if it matters. Sometimes a statistically uncertain but large effect can be more valuable than a statistically certain but trivial one.
Question 16
A utility company tests a new 'smart meter' in 100,000 households and finds it reduces average energy consumption by 0.8% compared to a control group. The result is statistically significant (p < 0.001). For an average customer, this 0.8% reduction amounts to a savings of $1.10 per month. The cost of installing a new meter is $150. Which factor is most critical in assessing the practical significance of this program for the company?
- The p-value, which indicates the high certainty of the energy savings.
- The payback period for the meter installation cost versus the accumulated customer savings. (correct answer)
- The Cohen's d effect size for the reduction in energy consumption.
- The total number of households in the company's service area.
Explanation: When you encounter questions about research results in business, you need to distinguish between statistical significance (whether an effect exists) and practical significance (whether it matters in the real world). Statistical significance tells you the result isn't due to chance, but practical significance determines if you should actually act on it.
The correct answer is B because practical significance for a business program centers on economic viability. With monthly savings of $1.10 per customer and installation costs of $150 per meter, the payback period is approximately 136 months (over 11 years). This extremely long payback period suggests the program may not be practically worthwhile, despite the statistically significant energy reduction.
Answer A is wrong because the p-value only confirms the result isn't due to random chance—it doesn't tell you if the effect is large enough to matter economically. Answer C is incorrect because Cohen's d measures effect size in standardized units, which doesn't directly translate to business value. While useful for research interpretation, it doesn't address the fundamental question of whether the program makes financial sense. Answer D is wrong because knowing the total service area helps with implementation planning but doesn't determine whether the program is worth implementing in the first place.
Study tip: In business statistics questions involving real-world implementation, always prioritize economic factors over statistical measures when assessing practical significance. A tiny but statistically significant effect may not justify the costs of implementation.
Question 17
A pharmaceutical company tested a new cholesterol medication on 5,000 patients over 6 months. The study found that patients taking the medication had a statistically significant reduction in LDL cholesterol compared to the placebo group (p = 0.02). The treatment group's mean LDL decreased by 8 mg/dL (from 180 to 172 mg/dL), while the control group's mean LDL increased by 2 mg/dL (from 178 to 180 mg/dL). The company's medical team noted that clinical guidelines suggest a reduction of at least 30 mg/dL is needed to meaningfully reduce cardiovascular risk.
Which statement best characterizes the study's findings regarding practical and statistical significance?
- The results are both statistically and practically significant, providing strong evidence for clinical implementation of the medication.
- The results are statistically significant but lack practical significance, as the observed reduction is below the clinically meaningful threshold. (correct answer)
- The results lack both statistical and practical significance due to the small effect size relative to baseline cholesterol levels.
- The results are practically significant but not statistically significant, suggesting the need for a larger sample size to detect the effect.
Explanation: The correct answer is B. The study shows statistical significance (p = 0.02 < 0.05), meaning the observed difference is unlikely due to chance. However, the 8 mg/dL reduction falls well short of the 30 mg/dL threshold considered clinically meaningful, indicating lack of practical significance. Choice A is incorrect because practical significance is lacking. Choice C is wrong because statistical significance was achieved (p = 0.02). Choice D is incorrect because the result was statistically significant, not the reverse.
Question 18
A pharmaceutical company conducts two studies on a new weight-loss supplement. A 4-pound weight loss is considered the minimum for practical significance.
- Study 1: A pilot study with n=30 participants finds an average weight loss of 8 pounds over 3 months, with p=0.09.
- Study 2: A large-scale study with n=5,000 participants finds an average weight loss of 0.5 pounds over 3 months, with p=0.01.
How should the company best interpret these combined results?
- Study 1 suggests a practically significant effect that may be real, warranting a larger follow-up study. Study 2 indicates a statistically significant but practically trivial effect. (correct answer)
- Study 2 is the more important result because of its low p-value and large sample size, so the supplement should be marketed as a statistically proven weight-loss aid.
- Neither study is useful; Study 1 failed to find a statistically significant result, and Study 2 found a result that is not practically significant.
- Study 1's high p-value indicates its result was an anomaly, and the true effect is the 0.5 pounds found in Study 2, which is both statistically and practically significant.
Explanation: The correct answer is A. This question requires comparing two scenarios. Study 1 shows a large effect size (8 pounds > 4 pounds), making it potentially practically significant. However, due to a small sample size, it is not statistically significant (p=0.09 > 0.05). This suggests the finding is promising but needs more powerful testing. Study 2, with its large sample, finds a statistically significant result (p=0.01 < 0.05) but the effect size (0.5 pounds) is well below the threshold for practical significance.
- B ignores the critical issue of practical significance in Study 2.
- C is too dismissive; Study 1's result, while not statistically proven, is promising and warrants further investigation.
- D incorrectly claims the effect in Study 2 is practically significant.
Question 19
A CEO is presented with a report stating that a new marketing campaign's results were 'statistically significant at the 99.9% confidence level (p < 0.001).' The CEO interprets this as proof of a highly successful and profitable campaign. What is the most important follow-up question a data-savvy board member should ask to validate this interpretation?
- What was the estimated return on investment or the effect size for the campaign? (correct answer)
- Was the sample size large enough to ensure the result wasn't a statistical fluke?
- Did you use a one-tailed or two-tailed test to arrive at that p-value?
- Could we replicate the study to confirm if the p-value remains this low?
Explanation: The correct answer is A. The CEO is making the common error of assuming high statistical significance equates to high business impact. The extremely low p-value only means there is very strong evidence that the campaign had some effect other than zero. It says nothing about whether that effect was large enough to be profitable or successful from a business perspective. Asking for the effect size or return on investment (ROI) directly addresses the question of practical significance.
- B is an unnecessary question; an extremely low p-value, especially from a business study, almost always implies a very large sample size was already used.
- C is a technical question about the test, but it doesn't address the core business issue of impact size.
- D is a good scientific practice, but before considering replication, the board needs to know if the initial result was even practically meaningful.
Question 20
A marketing manager's report on a new advertising campaign states, 'The campaign was a success, leading to a statistically significant increase in brand awareness (p=0.002).' As a business analyst reviewing this statement, what is the most critical piece of additional information needed to evaluate the business impact of the campaign?
- The effect size or the magnitude of the increase in brand awareness. (correct answer)
- The sample size used for the study that produced the p-value.
- The alpha level (significance level) that was pre-specified for the test.
- The statistical power of the test used to analyze the data.
Explanation: The correct answer is A. A p-value only indicates the strength of evidence against the null hypothesis (i.e., that there is no effect). It does not describe the size or importance of the effect. The campaign could have increased awareness by a trivial 0.1% or a substantial 10%; with a large enough sample, both could be statistically significant. To assess the business impact, you must know the magnitude of the change, which is measured by the effect size.
- B, sample size, is important for context but secondary to the effect size.
- C, the alpha level, is just the threshold for significance; it doesn't describe the result itself.
- D, statistical power, is a property of the study design, not a measure of the result's magnitude.