Business Analytics Quiz: Sample Size And Power
10 questions · exam conditions
0:00
Sample Size And PowerQuestion 1 of 10

A pilot experiment evaluating a new loyalty offer had only 3535 percent power to detect the smallest lift considered commercially worthwhile. The estimated lift was positive, but the test produced a p-value of 0.180.18. Its confidence interval ranged from a 11-percentage-point decline to a 55-percentage-point increase, while management considers a 33-point increase worthwhile.

Which recommendation is most defensible?

Abandon the offer because the nonsignificant result demonstrates that no positive lift exists.
Launch the offer because the interval's upper endpoint exceeds the worthwhile lift.
Run an adequately powered follow-up because the pilot did not rule out worthwhile effects.
Repeat identical small pilots until one produces a conventionally significant positive result.
← Back to quizzes

Business Analytics Quiz

Business Analytics Quiz: Sample Size And Power

Practice Sample Size And Power in Business Analytics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Sample Size And Power, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Analytics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A pilot experiment evaluating a new loyalty offer had only 3535 percent power to detect the smallest lift considered commercially worthwhile. The estimated lift was positive, but the test produced a p-value of 0.180.18. Its confidence interval ranged from a 11-percentage-point decline to a 55-percentage-point increase, while management considers a 33-point increase worthwhile.

Which recommendation is most defensible?

  1. Abandon the offer because the nonsignificant result demonstrates that no positive lift exists.
  2. Launch the offer because the interval's upper endpoint exceeds the worthwhile lift.
  3. Run an adequately powered follow-up because the pilot did not rule out worthwhile effects. (correct answer)
  4. Repeat identical small pilots until one produces a conventionally significant positive result.
Explanation: Whenever you see a question combining a nonsignificant p-value with a low-powered study, your first instinct should be to think carefully about what "nonsignificant" actually means — it does not mean "no effect exists." Statistical power of only 35%35\% means this pilot had a 65%65\% chance of missing a real, worthwhile effect even if one existed. With such a weak test, a p-value of 0.180.18 is essentially uninformative — you cannot distinguish "no effect" from "real effect that the study was too small to detect." Crucially, the confidence interval spans 1-1 to +5+5 percentage points, which includes the commercially meaningful +3+3-point threshold. This means a worthwhile lift remains entirely plausible. The only honest conclusion is that the pilot was inconclusive, and a properly powered study is needed — making C the most defensible recommendation. A commits the classic low-power fallacy: interpreting a nonsignificant result as proof of no effect. Absence of evidence is not evidence of absence, especially when power is this low. B is tempting but dangerously selective — the interval's upper bound exceeding the worthwhile threshold doesn't justify launch, because the interval equally includes zero and even a small decline. Cherry-picking the optimistic endpoint ignores real uncertainty. D describes p-hacking: repeatedly running small tests until chance produces a significant result inflates your false-positive rate and is both statistically invalid and ethically indefensible. As a study tip, remember this pattern: low power + nonsignificant result = inconclusive, not negative. On business analytics exams, questions pairing underpowered pilots with "no effect" conclusions are almost always traps.

Question 2

Before an experiment begins, a retailer specifies that a new recommendation engine will be adopted only if it increases revenue. Historical and business evidence make a decrease irrelevant to the adoption claim, though a decrease would still be monitored operationally. The team considers replacing its two-sided test with a one-sided test at the same significance level.

Which statement about the proposed change is correct?

  1. If specified in advance, it can increase power for an increase but will not support the same claim for a decrease. (correct answer)
  2. If selected after observing a positive estimate, it validly increases power without changing false-positive risk.
  3. It decreases power for an increase because one-sided tests divide the significance level across two tails.
  4. It guarantees that exactly half the original sample is sufficient for every possible positive effect.
Explanation: When a question asks about switching from a two-sided to a one-sided test, focus on two things: when the decision is made and what the test can and cannot conclude. A one-sided test concentrates all of your significance level (α) into one tail — the direction you care about. When you pre-specify a one-sided test before seeing any data, you gain statistical power in that direction compared to a two-sided test at the same α, because you're no longer "wasting" half the rejection region on the opposite tail. However, this comes with a strict trade-off: you formally surrender the ability to make a statistically supported claim in the other direction. In the retailer's scenario, a one-sided test increases power to detect a revenue increase, but if the engine actually decreased revenue, the test provides no statistical basis to claim a significant decrease. That's exactly what A describes — and it's correct precisely because the decision is made in advance. B is wrong because the timing matters enormously. Choosing a one-sided test after observing a positive estimate is p-hacking. It artificially inflates Type I error (false-positive risk) rather than controlling it, making the result invalid. C has the logic backwards. A one-sided test increases power for the specified direction by consolidating α into one tail, not decreasing it. D is a fabrication. No such guarantee exists — sample size requirements depend on effect size, variance, and desired power, not a fixed halving rule. Study tip: On hypothesis-testing questions, always ask: was the directionality decided before or after seeing the data? Pre-specification is what separates a legitimate one-sided test from a biased one.

Question 3

A marketing experiment is designed with 8080 percent power to detect a true sales lift of exactly 66 percent at the chosen significance level. A manager interprets this design specification as meaning there is an 8080 percent probability that any statistically significant result will be true.

Which interpretation of the power specification is correct?

  1. If a result is significant, there is an 8080 percent probability that the true lift is at least 66 percent.
  2. If the true lift is exactly 66 percent, repeated tests will reject the null about 8080 percent of the time. (correct answer)
  3. If a result is nonsignificant, there is a 2020 percent probability that the null hypothesis is actually true.
  4. If the campaign is deployed, approximately 8080 percent of customers will generate higher sales than before.
Explanation: Whenever you see a question about statistical power, anchor yourself to this definition: power is the probability of correctly rejecting a false null hypothesis, given a specific assumed effect size. It is a statement about what happens if the true effect exists — nothing more. With that foundation, B is the only valid interpretation. If the true sales lift is exactly 6%6\%, then across many repeated experiments of this design, the null hypothesis will be rejected approximately 80%80\% of the time. Power lives entirely in the world of "assuming the effect is real — how often do we catch it?" That is precisely what B states. A gets the conditional backwards. Power conditions on the true effect existing; A conditions on the result being significant and then makes a probability claim about the true lift. That's a positive predictive value calculation, which depends on the base rate of true effects and the false-positive rate — not on power alone. This is the exact confusion the manager in the passage makes. C confuses power with the false-negative rate (1power=20%1 - \text{power} = 20\%). A nonsignificant result means you failed to detect an effect — it says nothing about the probability that the null is actually true. Confusing "failing to reject" with "the null is true" is a classic logical error. D is simply a misreading of the word "power" in plain English, importing a consumer-behavior meaning that has nothing to do with hypothesis testing. Study tip: Always ask, "What is being conditioned on?" Power conditions on the true effect — not on the observed result. Questions that flip that conditional are testing whether you know the difference.

Question 4

A company is considering two balanced A/B tests with the same significance level and power. Test L is expected to increase conversion from 55 percent to 77 percent. Test H is expected to increase conversion from 5050 percent to 5252 percent. Assume the usual large-sample approximation for comparing two proportions.

Which comparison of required sample sizes is most accurate?

  1. Test L needs about one-quarter as many observations because its binomial outcome variance is lower. (correct answer)
  2. Test H needs about one-quarter as many observations because its relative lift is smaller.
  3. Both tests need approximately equal samples because both target a two-point absolute lift.
  4. Test H needs far fewer observations because a higher baseline rate produces more observed conversion events per unit of time.
Explanation: Whenever you see a sample-size comparison across A/B tests, your instinct should be to reach for the two-proportion z-test formula. Required sample size per group scales with p1(1p1)+p2(1p2)(p2p1)2\frac{p_1(1-p_1) + p_2(1-p_2)}{(p_2 - p_1)^2}. The denominator is identical for both tests — a two-point absolute lift means (0.070.05)2=(0.520.50)2=0.0004(0.07-0.05)^2 = (0.52-0.50)^2 = 0.0004 in both cases. So the difference comes entirely from the numerator, the sum of the two binomial variances p(1p)p(1-p). For Test L: 0.05(0.95)+0.07(0.93)0.0475+0.0651=0.11260.05(0.95) + 0.07(0.93) \approx 0.0475 + 0.0651 = 0.1126. For Test H: 0.50(0.50)+0.52(0.48)0.2500+0.2496=0.49960.50(0.50) + 0.52(0.48) \approx 0.2500 + 0.2496 = 0.4996. The ratio is roughly 0.4996/0.11264.40.4996 / 0.1126 \approx 4.4, meaning Test H requires about four times more observations — or equivalently, Test L needs roughly one-quarter the sample of Test H. That confirms answer A. Answer B gets the direction exactly backwards: Test H does not need fewer observations. Its higher baseline inflates variance, which increases required sample size. Answer C is the most seductive trap — students often think "same absolute lift = same sample size," but this ignores the variance term entirely. Answer D introduces an irrelevant consideration; events-per-unit-time affects how quickly you accumulate data in calendar time, but it has no bearing on the statistical sample size requirement. Your takeaway: never equate "same absolute lift" with "same required sample." Always check the variance p(1p)p(1-p), which peaks at p=0.50p = 0.50 — high-baseline tests are the most expensive to run.

Question 5

A training experiment needs 900900 analyzed employees in each arm. Based on prior programs, the treatment arm is expected to retain 9090 percent of randomized employees, while the control arm is expected to retain 7575 percent. The analyst wants to preserve the required analyzed sample in both arms.

How many employees should be randomized in total, ignoring any additional safety margin?

  1. Approximately 1,8001{,}800 employees
  2. Approximately 2,0002{,}000 employees
  3. Approximately 2,1152{,}115 employees
  4. Approximately 2,2002{,}200 employees (correct answer)
Explanation: When a study expects dropout, you must inflate the randomized sample to ensure the analyzed sample meets your target. The key insight is that each arm may have a different retention rate, so you must calculate the required randomized count separately for each arm, then sum them. Start with the treatment arm: you need 900900 analyzed employees, and you expect to retain 90%90\%. So you randomize 9000.90=1,000\frac{900}{0.90} = 1{,}000 employees. For the control arm: you need 900900 analyzed employees, but retention is only 75%75\%, so you randomize 9000.75=1,200\frac{900}{0.75} = 1{,}200 employees. The total randomized is 1,000+1,200=2,2001{,}000 + 1{,}200 = 2{,}200, confirming D is correct. Choice A (1,8001{,}800) is the trap of simply doubling the analyzed target without any dropout adjustment at all — it assumes perfect retention in both arms, which the passage explicitly contradicts. Choice B (2,0002{,}000) likely comes from applying a single average retention rate (roughly 82.5%82.5\%) to the total, but averaging retention rates across arms is mathematically incorrect when dropout rates differ. Choice C (2,1152{,}115) may result from applying the treatment arm's 90%90\% retention rate to both arms or from another partial adjustment, but it fails to correctly account for the control arm's lower 75%75\% retention. A reliable study tip: whenever arms have different dropout rates, always inflate each arm independently before summing. Never average retention rates across groups — the arm with the worst retention drives a disproportionately large inflation and must be handled separately.

Question 6

An online retailer originally planned a balanced A/B test requiring 1,6001{,}600 customers per group to detect a 44-percentage-point conversion lift with the desired significance level and power. Management now wants the test to detect a 33-percentage-point lift. Assume the baseline rate, variance, significance level, and target power remain unchanged.

Approximately how many customers per group should the revised test include?

  1. About 900900 customers per group
  2. About 2,1332{,}133 customers per group
  3. About 2,8442{,}844 customers per group (correct answer)
  4. About 3,2003{,}200 customers per group
Explanation: When a question asks how sample size changes after adjusting the minimum detectable effect, you need to remember the fundamental relationship: sample size scales inversely with the square of the effect size. Specifically, n1(Δ)2n \propto \frac{1}{(\Delta)^2}, where Δ\Delta is the lift you want to detect. Here's the logic. The original test was designed to detect a 44-percentage-point lift with 1,6001{,}600 customers per group. Now management wants to detect a 33-percentage-point lift — a smaller effect — which always requires a larger sample. The scaling factor is: nnew=nold×(ΔoldΔnew)2=1,600×(43)2=1,600×1692,844n_{\text{new}} = n_{\text{old}} \times \left(\frac{\Delta_{\text{old}}}{\Delta_{\text{new}}}\right)^2 = 1{,}600 \times \left(\frac{4}{3}\right)^2 = 1{,}600 \times \frac{16}{9} \approx 2{,}844 That confirms C is correct. Looking at why the other choices fail: A (900 customers) implies the sample decreases, which is backwards — detecting a smaller effect always demands more data, not less. B (2,133) corresponds to a linear scaling 1,600×431{,}600 \times \frac{4}{3}, forgetting that the relationship is squared, not linear. D (3,200) doubles the original sample, which would correspond to halving the detectable effect, not reducing it by one percentage point. A useful rule of thumb: whenever the detectable effect shrinks by a factor of kk, your required sample grows by k2k^2. Here, the effect shrank by 43\frac{4}{3}, so the sample grew by (43)21.78×\left(\frac{4}{3}\right)^2 \approx 1.78\times. Memorize this squared inverse relationship — it appears frequently in A/B testing and experimental design questions.

Question 7

A logistics company planned an experiment with a total sample of 2,0002{,}000 deliveries. After incorporating a strong pre-experiment predictor into the analysis, the residual standard deviation of delivery time is expected to fall from 2020 minutes to 1616 minutes. The target effect, significance level, power, and group allocation are unchanged.

Under the usual sample-size approximation, what total sample should now provide approximately the original power?

  1. About 800800 deliveries
  2. About 1,2801{,}280 deliveries (correct answer)
  3. About 1,6001{,}600 deliveries
  4. About 2,5002{,}500 deliveries
Explanation: When you reduce variability in an experiment by controlling for a strong covariate, you directly reduce the noise in your outcome measure. Sample size under the standard approximation scales with the square of the standard deviation — specifically, required nσ2n \propto \sigma^2. This means if you can shrink σ\sigma, you need proportionally fewer observations to achieve the same power. Here, the residual standard deviation drops from 2020 to 1616 minutes. The ratio of the new required sample to the original is: nnewnoriginal=σnew2σoriginal2=162202=256400=0.64\frac{n_{\text{new}}}{n_{\text{original}}} = \frac{\sigma_{\text{new}}^2}{\sigma_{\text{original}}^2} = \frac{16^2}{20^2} = \frac{256}{400} = 0.64 Applying this to the original sample of 2,0002{,}000: nnew=2,000×0.64=1,280n_{\text{new}} = 2{,}000 \times 0.64 = 1{,}280 That confirms B as the correct answer. Looking at the distractors: A (800) applies a ratio of 0.400.40, which might come from mistakenly using the ratio of standard deviations (16/20=0.816/20 = 0.8) and then squaring incorrectly or halving — a common algebraic slip. C (1,600) uses the linear ratio 16/20=0.816/20 = 0.8 directly without squaring, a classic error of forgetting that sample size scales with variance, not standard deviation. D (2,500) goes in the wrong direction entirely, implying you'd need more observations after reducing variability — the opposite of the truth. Study tip: Always remember that reducing σ\sigma by a factor of kk reduces required sample size by a factor of k2k^2. When you see covariate adjustment or blocking in a question, immediately think "variance reduction → squared savings in sample size."

Question 8

A subscription company can enroll 900900 customers in an experiment. Historical data suggest that the outcome's standard deviation will be 1818 units in the treatment group and 99 units in the control group. Assume equal enrollment costs and independent observations. The objective is to minimize the variance of the estimated difference in group means.

Which allocation is approximately optimal?

  1. 300300 treatment customers and 600600 control customers
  2. 450450 treatment customers and 450450 control customers
  3. 600600 treatment customers and 300300 control customers (correct answer)
  4. 720720 treatment customers and 180180 control customers
Explanation: When you see an experiment with unequal variances across groups, the key insight is that equal sample sizes are not automatically optimal. The goal here is to minimize the variance of the estimated difference in means, which equals σT2nT+σC2nC\frac{\sigma_T^2}{n_T} + \frac{\sigma_C^2}{n_C}, subject to the constraint nT+nC=900n_T + n_C = 900. The optimal allocation rule (Neyman allocation) says you should assign participants proportionally to each group's standard deviation: nTσTn_T \propto \sigma_T and nCσCn_C \propto \sigma_C. With σT=18\sigma_T = 18 and σC=9\sigma_C = 9, the ratio is 18:9=2:118:9 = 2:1. Splitting 900 in a 2:1 ratio gives nT=600n_T = 600 and nC=300n_C = 300, confirming that C is correct. You can verify: the resulting variance is 182600+92300=0.54+0.27=0.81\frac{18^2}{600} + \frac{9^2}{300} = 0.54 + 0.27 = 0.81, which is lower than any other split. A (300 treatment, 600 control) reverses the logic — it allocates more participants to the lower-variance group, which inflates the variance of the high-variance group unnecessarily, yielding 324300+81600=1.08+0.135=1.215\frac{324}{300} + \frac{81}{600} = 1.08 + 0.135 = 1.215, far worse. B (450/450) is the intuitive "fair split" trap — equal allocation ignores the variance difference entirely, producing 0.72+0.18=0.900.72 + 0.18 = 0.90, worse than C. D (720 treatment, 180 control) over-allocates to treatment beyond the optimal ratio, giving 0.45+0.45=0.900.45 + 0.45 = 0.90, also worse. Your study tip: whenever standard deviations differ across groups, remember Neyman allocation — assign proportionally to σ\sigma, not equally. Equal splits only minimize variance when σT=σC\sigma_T = \sigma_C.

Question 9

A grocery chain will randomize stores rather than individual shoppers to test a promotion. An individual-randomization calculation indicates that 1,2001{,}200 shoppers would be required. Each store contributes 5050 shoppers, and the anticipated intraclass correlation is 0.040.04. Assume equal cluster sizes and use the design effect 1+(m1)ρ1+(m-1)\rho.

What is the smallest whole number of stores that should be included using this approximation?

  1. 2424 stores, representing 1,2001{,}200 shoppers
  2. 6060 stores, representing 3,0003{,}000 shoppers
  3. 7171 stores, representing 3,5503{,}550 shoppers
  4. 7272 stores, representing 3,6003{,}600 shoppers (correct answer)
Explanation: When a study randomizes clusters (like stores) instead of individuals, the effective sample size must be inflated to account for the fact that people within the same cluster tend to behave similarly. The design effect (DEFF) captures this inflation: DEFF=1+(m1)ρ\text{DEFF} = 1 + (m-1)\rho, where mm is the cluster size and ρ\rho is the intraclass correlation. You multiply your original individual-level sample size by this factor to get the required number of individuals under cluster randomization, then divide by the cluster size to find the number of clusters needed. Here, m=50m = 50 shoppers per store and ρ=0.04\rho = 0.04, so: DEFF=1+(501)(0.04)=1+49×0.04=1+1.96=2.96\text{DEFF} = 1 + (50-1)(0.04) = 1 + 49 \times 0.04 = 1 + 1.96 = 2.96 The adjusted total number of shoppers needed is: 1,200×2.96=3,5521{,}200 \times 2.96 = 3{,}552 Dividing by 50 shoppers per store gives 3,552/50=71.043{,}552 / 50 = 71.04 stores. Since you must round up to the next whole number (you can't have a partial store), you need 72 stores, representing 3,600 shoppers — confirming answer D. Choice A ignores the design effect entirely, using the original individual-randomization figure as if clustering didn't matter. Choice B applies an incorrect design effect of 2.5, which would correspond to a different ρ\rho or mm. Choice C makes the critical rounding error — 71.04 rounds up to 72, not down to 71, because you must meet the minimum threshold. Remember: always round cluster counts up, never down. Rounding down would leave you underpowered, which defeats the purpose of the calculation.

Question 10

A bank plans a fixed-size experiment comparing two credit-card onboarding messages. Before collecting data, its compliance team changes the test from a significance level of 0.050.05 to 0.010.01. No other design feature changes.

Which statement best describes the consequences of this change?

  1. The false-positive risk decreases, power decreases, and the minimum detectable effect increases. (correct answer)
  2. The false-positive risk decreases, power increases, and the minimum detectable effect decreases.
  3. The false-positive risk increases, power decreases, and the minimum detectable effect remains unchanged.
  4. The false-positive risk remains unchanged, power increases, and the minimum detectable effect increases.
Explanation: Whenever you see a hypothesis testing question involving a change to the significance level, think about the three-way relationship between Type I error risk, statistical power, and minimum detectable effect (MDE) — they're all connected through the same underlying probability thresholds. Lowering α\alpha from 0.050.05 to 0.010.01 means you're demanding stronger evidence before rejecting the null hypothesis. This directly reduces your false-positive rate (Type I error), which is exactly what the compliance team wanted. But this strictness comes at a cost: with a fixed sample size, you're moving the critical value further into the tail of the distribution. That makes it harder for a true effect to clear the new, higher bar — so power (the probability of detecting a real effect) decreases. Because power falls, the smallest effect the test can reliably detect must be larger — meaning the MDE increases. Answer A captures all three of these consequences correctly. Answer B is wrong because it claims power increases, which is the opposite of what happens. Lower α\alpha reduces power when nothing else changes. Answer C incorrectly states that false-positive risk increases — lowering α\alpha does the opposite, it tightens the threshold. Answer D claims false-positive risk is unchanged, which contradicts the definition of α\alpha itself; changing α\alpha is literally changing the false-positive rate. A useful memory anchor: think of the significance level as a filter. A finer filter (smaller α\alpha) catches fewer false positives, but it also lets real signals slip through more often — so you need bigger effects to be seen.