Business Statistics Quiz: A B Testing Concepts
20 questions · exam conditions
0:00
A B Testing ConceptsQuestion 1 of 20

A streaming platform is testing a new content recommendation algorithm. They randomly assign 100,000 users to treatment and control groups. After two weeks, they find that treatment group users watch 8% more content (statistically significant), but they also notice that 12% of treatment users and 8% of control users have canceled their subscriptions during the test period. What does this pattern suggest about the algorithm's impact, and what analysis approach would be most informative?

The algorithm successfully increases engagement; the subscription cancellations are likely unrelated to the treatment and should be analyzed separately from the content consumption metric
The algorithm may be recommending more addictive but lower-quality content, leading to short-term engagement gains but long-term satisfaction decline requiring cohort retention analysis
The results demonstrate a successful treatment with acceptable churn levels; the 4% difference in cancellation rates is within normal variation for subscription services
The algorithm creates a bimodal user response where some users engage more while others disengage completely, requiring segmentation analysis to identify user characteristics predicting each response
← Back to quizzes

Business Statistics Quiz

Business Statistics Quiz: A B Testing Concepts

Practice A B Testing Concepts in Business Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on A B Testing Concepts, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A streaming platform is testing a new content recommendation algorithm. They randomly assign 100,000 users to treatment and control groups. After two weeks, they find that treatment group users watch 8% more content (statistically significant), but they also notice that 12% of treatment users and 8% of control users have canceled their subscriptions during the test period. What does this pattern suggest about the algorithm's impact, and what analysis approach would be most informative?

  1. The algorithm successfully increases engagement; the subscription cancellations are likely unrelated to the treatment and should be analyzed separately from the content consumption metric
  2. The algorithm may be recommending more addictive but lower-quality content, leading to short-term engagement gains but long-term satisfaction decline requiring cohort retention analysis (correct answer)
  3. The results demonstrate a successful treatment with acceptable churn levels; the 4% difference in cancellation rates is within normal variation for subscription services
  4. The algorithm creates a bimodal user response where some users engage more while others disengage completely, requiring segmentation analysis to identify user characteristics predicting each response
Explanation: The combination of increased short-term engagement but higher cancellation rates suggests the algorithm may be optimizing for immediate engagement at the expense of long-term satisfaction. Cohort retention analysis would reveal if this pattern continues over time. Option A incorrectly dismisses the cancellation pattern as unrelated when it's likely connected to the algorithm change. Option C minimizes a 50% relative increase in cancellation rates (from 8% to 12%). Option D suggests bimodal responses but doesn't address the concerning trend of higher cancellations despite increased engagement.

Question 2

An e-learning platform tests a new course recommendation system. The test runs for 8 weeks with 25,000 students per group. Results show that treatment group students complete 14% more courses (p = 0.02), but when analyzed by course difficulty level, easy courses show 28% higher completion while difficult courses show 12% lower completion. The platform's business model depends on students progressing through increasingly challenging content. How should this finding influence the implementation decision?

  1. Implement the new system because the overall increase in course completions will improve student engagement metrics and platform usage statistics
  2. Implement the system with monitoring for course difficulty progression patterns and student long-term outcomes to catch any negative trends early
  3. Modify the recommendation algorithm to weight difficult courses more heavily, then retest to see if the bias toward easy content can be corrected
  4. Reject the new system because it appears to steer students away from challenging content, potentially undermining educational outcomes and long-term business sustainability (correct answer)
Explanation: When analyzing A/B test results in business statistics, you must look beyond surface-level metrics to understand the underlying patterns and long-term implications. This question tests your ability to interpret segmented data and consider strategic business consequences. The correct answer is D because the segmented analysis reveals a critical flaw: while overall completions increased 14%, the system is actually steering students toward easier content (28% increase) and away from difficult courses (12% decrease). Since the business model explicitly depends on students progressing through increasingly challenging material, this recommendation system directly undermines the core value proposition and long-term sustainability. A is wrong because it focuses solely on the aggregate metric without considering the concerning breakdown by difficulty level. Raw completion numbers don't tell the whole story when the business strategy requires specific types of engagement. B is flawed because implementing a system you already know has fundamental problems, hoping to catch issues later, ignores clear warning signs in the data. The monitoring approach makes sense for unknown risks, not identified problems. C misses the point by assuming the algorithm can be easily fixed. The current results suggest the system may have inherent biases that could be difficult to correct, and retesting delays addressing a known strategic misalignment. Key strategy: In business statistics questions involving A/B tests, always examine segmented results and consider whether improvements in topline metrics align with underlying business objectives. Surface-level gains can mask deeper strategic problems.

Question 3

A food delivery platform implements an A/B test where treatment users receive personalized restaurant recommendations based on machine learning, while control users see the standard list sorted by delivery time. After 5 weeks with 80,000 users per group, treatment users place 11% more orders (p = 0.01) and have 7% higher average order values (p = 0.04). However, restaurant partners in the treatment group report 23% fewer orders on average compared to control group restaurants. What does this result pattern most likely indicate about the recommendation algorithm?

  1. The algorithm successfully optimizes user experience and business metrics; the restaurant partner concerns reflect a temporary adjustment period to the new system
  2. The algorithm demonstrates successful personalization by matching users with restaurants they prefer, and the restaurant partner feedback represents selection bias from less popular establishments
  3. The algorithm has a data quality issue where restaurant availability or delivery capacity isn't properly factored into recommendations, leading to suboptimal restaurant utilization
  4. The algorithm creates a concentration effect, heavily favoring popular restaurants and potentially creating long-term marketplace imbalances that could reduce restaurant diversity (correct answer)
Explanation: When analyzing A/B test results for recommendation systems, you need to examine outcomes across all stakeholders—users, the platform, and supply-side partners. Conflicting signals often reveal deeper algorithmic behaviors that aren't immediately apparent from user metrics alone. The data shows users are happier (11% more orders, 7% higher order values), but restaurants report 23% fewer orders on average. This pattern indicates the algorithm is concentrating demand among fewer restaurants rather than distributing it evenly. Popular restaurants likely receive disproportionate recommendations, while less popular ones see significant order drops—a classic concentration effect. Answer A incorrectly assumes this is temporary. A 23% decrease across restaurant partners after 5 weeks suggests a systematic issue, not an adjustment period. Answer B misinterprets the restaurant feedback as selection bias from unpopular establishments, but the magnitude (23% average decrease) indicates widespread impact beyond just struggling restaurants. Answer C focuses on data quality issues around availability or capacity, but the user satisfaction metrics contradict this—if restaurants couldn't fulfill orders due to capacity problems, users wouldn't be ordering more with higher values. Answer D correctly identifies the core issue: the algorithm optimizes for user preferences by heavily recommending popular restaurants, creating marketplace concentration that reduces opportunities for diverse restaurant partners. Study tip: In marketplace A/B tests, always analyze multi-sided effects. Strong user metrics paired with negative supplier metrics often signal concentration effects that can threaten long-term marketplace health, even when short-term user satisfaction improves.

Question 4

An e-commerce company tests a new checkout button design (Variant B) against the current design (Variant A). They monitor 10 key metrics. The primary metric, overall conversion rate, shows no statistically significant change. However, a secondary metric, 'average session duration,' shows a statistically significant increase with a p-value of 0.04. What is the most appropriate conclusion for the product manager to draw?

  1. The new button should be launched because it significantly improves user engagement, as measured by session duration, which is a positive business outcome.
  2. The experiment is inconclusive because the primary and secondary metrics provide conflicting signals, indicating a need to redesign the test.
  3. The significant result for session duration is likely a false positive due to the multiple testing problem, and it should be treated with skepticism unless confirmed by a follow-up experiment. (correct answer)
  4. The statistical significance level (α) for the experiment should have been retroactively lowered to 0.005 to properly account for the 10 metrics being tested.
Explanation: When multiple metrics are tested simultaneously, the probability of observing at least one significant result by random chance (a Type I error) increases. With 10 metrics and α = 0.05, the probability of at least one false positive is approximately 1 - (0.95)^10 ≈ 40%. Therefore, a single significant result on a secondary metric, especially when the primary metric shows no effect, is suspect. The most appropriate conclusion is that this might be a statistical anomaly due to multiple comparisons, not a true effect. Distractor A ignores the primary metric and the multiple testing problem. Distractor B is too general and doesn't identify the specific statistical issue. Distractor D suggests changing the significance level after seeing the results, which is improper statistical practice; adjustments for multiple testing like Bonferroni correction should be planned in advance.

Question 5

An e-commerce company tests a new feature on its product pages that recommends similar, but less expensive, alternative products. The primary metric, 'conversion rate on the main product,' decreases significantly. However, a key secondary metric, 'overall revenue per visitor,' shows a statistically significant increase. What is the best interpretation of this outcome?

  1. The test failed because the primary metric, conversion rate on the main product, showed a significant negative impact and should be the main focus.
  2. The results are contradictory and therefore inconclusive; the experiment should be rerun with a different design to obtain a clearer signal.
  3. The feature successfully increased overall business value by guiding users to alternatives they purchased, despite cannibalizing sales of the initial product. (correct answer)
  4. The feature should only be launched for products with very high prices, where the risk of revenue loss from cannibalization would be lowest.
Explanation: This scenario highlights the importance of choosing the right Overall Evaluation Criterion (OEC). While the feature 'failed' on the narrow primary metric (conversion of the main product), it 'succeeded' on the broader, more important business metric (overall revenue per visitor). The feature caused cannibalization—it moved sales from one product to another—but the net effect was positive. The correct interpretation is that this cannibalization was beneficial, leading to an overall lift. Distractor A focuses on the wrong metric. Distractor B fails to see that the results are not contradictory but rather tell a complete story about a business trade-off. Distractor D proposes a future action but is not the best interpretation of the current results, which show a net positive effect overall.

Question 6

A streaming service tests a new recommendation algorithm (Variant B). The primary metric, 'average daily watch time per user,' increases by 5% with statistical significance (p=0.01). However, a critical guardrail metric, 'rate of user-initiated content buffering events,' also increases significantly by 15% (p=0.005). What is the most prudent course of action for the company?

  1. Launch Variant B because the primary business objective of increasing watch time was successfully and significantly met.
  2. Do not launch Variant B, as the significant degradation in a key user experience metric likely outweighs the benefit and could lead to long-term churn. (correct answer)
  3. Rerun the experiment with a larger sample size to confirm if the 15% increase in buffering is a persistent and real effect.
  4. Launch Variant B to a small segment of users to gather qualitative feedback on whether the increased buffering is noticeable.
Explanation: Guardrail metrics are designed to catch unintended negative consequences. A significant negative change in a critical user experience metric like buffering can lead to user frustration and attrition over the long term, even if a primary engagement metric improves in the short term. The most prudent decision is to prioritize the long-term health of the user base and not launch a change that significantly worsens their experience. Distractor A exhibits 'metric fixation' by ignoring the negative guardrail. Distractor C is unnecessary as the negative effect is already statistically significant. Distractor D is risky; the quantitative data already provides a strong warning signal that should be heeded before exposing more users to a poor experience.

Question 7

A data analyst is running an A/B test on a new feature, with a pre-calculated required sample size of 20,000 users per variant. After only 5,000 users have been assigned, the analyst checks the results and finds a p-value of 0.04 for the primary metric. They decide to stop the test and declare a winner. This practice of 'peeking' is statistically problematic primarily because it:

  1. decreases the statistical power of the test, making it more difficult to detect a true effect if one exists.
  2. introduces selection bias, as the early adopters in the test may not be representative of the entire user population.
  3. violates the assumption of independence between data points that is fundamental to the statistical test being used.
  4. inflates the Type I error rate, increasing the probability of concluding there is an effect when one does not actually exist. (correct answer)
Explanation: Repeatedly checking results and stopping a test as soon as the p-value crosses the significance threshold (e.g., 0.05) dramatically increases the Type I error rate. Each check is an opportunity for random fluctuations to yield a 'significant' result by chance. By stopping early, the analyst capitalizes on this randomness. The correct procedure is to commit to a sample size in advance and only analyze the results once that sample size is reached. Distractor A is incorrect; peeking doesn't reduce power. Distractor B describes a potential separate issue (sample representativeness), but the core statistical sin of peeking is the inflation of the false positive rate. Distractor C is incorrect; the data points themselves are likely still independent, but the stopping rule is what is flawed.

Question 8

Before launching a series of A/B tests, a data science team runs an A/A test where both user groups see the exact same version of the website. After collecting data from 50,000 users, they find that the conversion rate for Group A is 5.2% and for Group B is 5.8%, with a p-value of 0.03. What is the most important implication of this A/A test result?

  1. The natural variability in user behavior is higher than anticipated, so future A/B tests will require a much larger sample size to achieve sufficient power.
  2. The A/A test was successful because it demonstrated that random chance can produce different outcomes even between identical groups.
  3. The experimentation platform has a systematic bias, possibly in the randomization algorithm or data collection, that must be investigated and fixed. (correct answer)
  4. The team should proceed with A/B tests but use a more conservative significance level (e.g., α = 0.01) to account for the platform's instability.
Explanation: The purpose of an A/A test is to validate the testing system. The null hypothesis is that there is no difference between the groups. In a well-calibrated system, we expect the p-value to be greater than 0.05 about 95% of the time. Finding a statistically significant difference (p=0.03) is a strong indication that something is wrong with the testing infrastructure, such as biased user assignment or faulty data logging. This is a critical failure that must be fixed before any real A/B tests can be trusted. Distractor A misattributes a systematic issue to random noise. Distractor B misunderstands the purpose of an A/A test; a significant result is a failure, not a success. Distractor D suggests a workaround that doesn't address the root cause of the problem.

Question 9

A company tests a new, simplified sign-up form. The primary metric, sign-up completion rate, shows a significant 10% improvement overall. However, a segment analysis reveals that for mobile users the completion rate improved by 15%, while for desktop users it decreased by 5% (a statistically significant drop). The company has a roughly 50/50 split between mobile and desktop users. What is the most critical conclusion from this result?

  1. The new form should be launched for all users, as the overall impact on the primary metric is positive and statistically significant.
  2. The test shows a positive result and the form should be launched for mobile users while the old form is retained for desktop users.
  3. The overall positive result is misleading because it masks a significant negative impact on a large user segment that must be investigated. (correct answer)
  4. The experiment is invalid due to Simpson's Paradox, and a new test must be designed that isolates mobile and desktop traffic from the start.
Explanation: A key principle of A/B testing is to 'do no harm.' While the overall result is positive, this is an aggregate figure that hides a significant problem. Launching the feature would actively harm the experience for half the user base (desktop users). The correct interpretation is that the overall 'win' is misleading. The critical next step is to understand why the experience degraded for desktop users before making any launch decision. Distractor A is a classic error of relying only on aggregate metrics. Distractor B is a potential long-term solution, but it creates technical debt and a fragmented user experience; the immediate step is investigation, not a partial launch. Distractor D correctly identifies the paradoxical nature of the result but incorrectly calls the experiment invalid; the data is valid and provides a crucial insight that requires further analysis.

Question 10

An A/B test on advertising copy results in a 95% confidence interval for the difference in conversion rates (Variant B - Variant A) of [+0.5%, +3.5%]. Which of the following statements is an incorrect interpretation of this result?

  1. We are 95% confident that the true difference in conversion rates between Variant B and Variant A is between 0.5% and 3.5%. (correct answer)
  2. Since the interval does not contain zero, the result is statistically significant at the α = 0.05 level.
  3. If we were to repeat this experiment many times, we would expect 95% of the calculated confidence intervals to contain the true difference.
  4. The result is consistent with Variant B having a true effect as small as 0.5% or as large as 3.5% over Variant A.
Explanation: This is a subtle but critical point in frequentist statistics. The confidence level (95%) refers to the reliability of the procedure used to create the interval, not the probability that a specific interval contains the true parameter. After an interval is calculated (e.g., [0.5%, 3.5%]), the true parameter is either in it or it is not. The probability is either 0 or 1. Therefore, stating there is a 95% probability the true value lies within this specific interval is incorrect. Choice C provides the correct frequentist interpretation. Choice B is correct; if a 95% CI for a difference does not include 0, the result is significant at the 5% level. Choice D is a correct and practical interpretation of the range of plausible values for the true effect.

Question 11

A B2B software company starts an A/B test on its pricing page on a Thursday morning. By Friday afternoon, the new page (Variant B) shows a 30% lift in demo requests with a p-value of 0.01. A product manager argues to stop the test and ship the change immediately. What is the most significant flaw in this reasoning?

  1. The sample size is likely too small to make a reliable decision, regardless of the p-value.
  2. The test has not run for a full weekly business cycle, and user behavior and intent may differ significantly between weekdays and weekends. (correct answer)
  3. The novelty effect is likely inflating the results, and the observed lift will almost certainly decrease if the test continues to run longer.
  4. The p-value is suspiciously low for such a short test, which may indicate a data collection error or a bug in the experiment setup.
Explanation: A fundamental rule for experiment duration is to run the test long enough to capture the natural cyclicality of the business. For a B2B company, user profiles and purchasing intent can vary dramatically between weekdays (when employees are working) and weekends. Making a decision based on data from only two weekdays is premature because it is not representative of the behavior of all users across a full week. While small sample size (A) and novelty effect (C) are also valid concerns, failing to account for the business cycle is the most critical and definite flaw in this scenario.

Question 12

A marketing team tests five new email subject lines (B, C, D, E, F) against the current one (A) in a single experiment. They want to maintain a family-wise error rate (FWER) of at most 5%. If they use the Bonferroni correction to adjust for multiple comparisons, what significance level (α) should they use for each of the five individual hypothesis tests (A vs. B, A vs. C, etc.)?

  1. 0.05
  2. 0.01 (correct answer)
  3. 0.0083
  4. 0.025
Explanation: The Bonferroni correction is a method to control the FWER when performing multiple hypothesis tests. It works by dividing the desired overall alpha level by the number of comparisons being made. In this scenario, there are five new subject lines being compared to the control, resulting in 5 comparisons. To maintain a FWER of 0.05, the significance level for each individual test should be α=αk=0.055=0.01\alpha' = \frac{\alpha}{k} = \frac{0.05}{5} = 0.01. Distractor A fails to make any correction. Distractor C incorrectly divides by the total number of groups (6) instead of the number of comparisons (5). Distractor D confuses the logic of two-tailed tests with the Bonferroni correction method.

Question 13

An A/B test is conducted, and the p-value for the difference in the primary metric is 0.25. The team concludes there is no difference between the variant and the control. However, a post-hoc power analysis reveals the test only had 30% power to detect the pre-specified minimum detectable effect (MDE). What is the most accurate interpretation of the test results?

  1. The conclusion is correct; a p-value of 0.25 provides strong evidence that the null hypothesis of no effect is true.
  2. The test was inconclusive; there was a 70% chance of failing to detect a real effect of the MDE size (a Type II error), so a true effect may well exist. (correct answer)
  3. The results are invalid, and the experiment must be re-run with a higher significance level (e.g., α = 0.10) to increase the power.
  4. The variant likely had a small negative effect on the metric, but it was not large enough to be statistically significant.
Explanation: Statistical power is the probability of detecting an effect of a certain size, if one truly exists. A power of 30% means that even if the variant had a real effect equal to the MDE, there was only a 30% chance of getting a statistically significant result. Conversely, there was a 70% chance of a Type II error (failing to reject the null when it is false). A high p-value combined with low power means the test was underpowered and therefore inconclusive. It does not prove the null hypothesis is true (A). While re-running with more power (likely from a larger sample) is a good idea, the interpretation of the current result is that it's inconclusive. Changing alpha (C) is not the standard way to fix power issues. There is no evidence to suggest a negative effect (D).

Question 14

During the planning of an A/B test, the Minimum Detectable Effect (MDE) was set to a 2% lift in conversion rate. The test concludes, and the results show a 3% lift with a p-value of 0.03. The 95% confidence interval for the lift is [0.5%, 5.5%]. Which statement represents the most accurate conclusion?

  1. The test is inconclusive because the lower bound of the confidence interval (0.5%) is below the pre-specified MDE of 2%.
  2. The test should be run longer to narrow the confidence interval and confirm that the true effect is definitively above the 2% MDE.
  3. The MDE was set incorrectly, as the results suggest the true effect is likely closer to 3% and the confidence interval is quite wide.
  4. The test was successful because it detected a statistically significant effect, and the observed lift of 3% exceeded the MDE. (correct answer)
Explanation: When evaluating A/B test results, you need to understand the relationship between statistical significance, practical significance, and the Minimum Detectable Effect (MDE). The MDE represents the smallest effect size you want to reliably detect, but it doesn't invalidate larger effects that meet statistical significance thresholds. The test successfully detected a statistically significant effect (p-value = 0.03 < 0.05) with an observed lift of 3% that exceeds the pre-specified MDE of 2%. This means you have sufficient evidence to conclude there's a real difference between the variants, making answer D correct. Answer A incorrectly suggests the test is inconclusive because the confidence interval's lower bound (0.5%) falls below the MDE. However, statistical significance is determined by whether the entire confidence interval excludes zero (no effect), not whether it excludes the MDE. Since the interval [0.5%, 5.5%] doesn't include zero, the result is statistically significant. Answer B misunderstands the purpose of the confidence interval. While a narrower interval would be nice, the current results already provide sufficient evidence for a conclusion based on the p-value and the fact that the observed effect exceeds the MDE. Answer C confuses the role of MDE in test design. The MDE isn't "incorrect" just because the actual effect was larger or the confidence interval is wide. It served its purpose in determining the appropriate sample size for the test. Remember: MDE is a planning tool for sample size calculation, not a threshold that invalidates significant results when confidence intervals extend below it.

Question 15

A website A/B tests a new homepage layout. Overall, the new layout (B) has a lower conversion rate than the old one (A). However, when the data is segmented by traffic source, the new layout (B) has a higher conversion rate for users from both Search and Social Media. This reversal is an example of Simpson's Paradox. What is the most likely underlying cause of this paradox?

  1. The randomization was flawed, sending more high-intent users to Layout A and more low-intent users to Layout B during the experiment.
  2. The sample size was too small, leading to random fluctuations that created the appearance of a paradox where none truly exists.
  3. The statistical test used for the overall comparison was inappropriate for data that has significant segment-level differences in performance.
  4. The proportion of traffic from different sources shifted during the test, with Layout B disproportionately receiving more traffic from the lower-converting source. (correct answer)
Explanation: Simpson's Paradox occurs when aggregated data shows a different trend than the data within individual subgroups. In A/B testing, this typically happens due to compositional differences between test groups rather than the treatment effect itself. Answer D correctly identifies the root cause: Layout B received a disproportionate amount of traffic from the lower-converting source. Here's how this creates the paradox: Imagine Search converts at 10% and Social Media at 2%. If Layout A gets 80% Search traffic and 20% Social traffic, while Layout B gets 30% Search and 70% Social, then even if Layout B performs better within each segment, its overall conversion rate will be dragged down by having more low-converting Social traffic. Answer A suggests flawed randomization based on user intent, but the paradox exists specifically because we can see Layout B performs better within defined segments (Search and Social) — this indicates the randomization worked correctly within those segments. Answer B blames small sample size, but Simpson's Paradox isn't about random fluctuations or statistical noise. It's a systematic effect that occurs even with large, statistically significant samples. Answer C incorrectly focuses on the statistical test choice. The paradox isn't about using the wrong test — it's about the underlying data composition that makes aggregate analysis misleading regardless of which test you use. Remember: When you see Simpson's Paradox in business contexts, look for compositional differences between groups. The "reversal" usually stems from one group having more exposure to a naturally lower-performing segment, not from the treatment itself being harmful.

Question 16

A mobile app team runs an A/B test on push notification frequency, measuring 7-day retention as the primary metric. After 3 weeks, they observe no significant difference in retention (p = 0.45) but notice that treatment group users have 18% higher in-app purchase revenue. The team wants to switch the primary metric to revenue and declare the test successful. What are the primary statistical and business risks of this approach?

  1. The main risk is Type I error inflation from post-hoc metric selection, potentially leading to false positive results and poor business decisions based on chance findings (correct answer)
  2. The main risk is confounding variables in the revenue metric, as purchase behavior may be influenced by factors outside the notification frequency treatment
  3. The main risk is insufficient statistical power for revenue metrics, which typically require larger sample sizes than retention metrics to detect meaningful differences
  4. The main risk is survivorship bias, as users who remain active enough to make purchases may not be representative of the broader user population
Explanation: Post-hoc metric switching ('HARKing' - Hypothesizing After Results are Known) inflates Type I error rates because it's essentially cherry-picking significant results from multiple comparisons. This can lead to implementing changes based on random variation rather than true effects. Option B mentions confounding but doesn't address the core statistical issue of metric switching. Option C discusses power but revenue differences are often more detectable than retention differences. Option D describes a real phenomenon but isn't the primary concern with switching metrics post-hoc.

Question 17

A company is conducting an A/B test on their checkout page to measure conversion rate improvements. The test shows a statistically significant 15% increase in conversion rate (p < 0.05) for the treatment group. However, when they segment the results by device type, they discover that mobile users show a 25% decrease in conversion rate while desktop users show a 45% increase. What statistical phenomenon does this illustrate, and what should be the primary concern for implementing the treatment?

  1. This demonstrates Simpson's paradox, and the primary concern should be ensuring mobile traffic represents a significant portion of overall traffic before rejecting the treatment
  2. This demonstrates interaction effects, and the primary concern should be developing separate treatments for mobile and desktop users rather than implementing a universal solution (correct answer)
  3. This demonstrates selection bias, and the primary concern should be re-randomizing users to ensure equal distribution across device types
  4. This demonstrates multiple testing problems, and the primary concern should be applying Bonferroni correction to maintain the overall significance level
Explanation: This scenario demonstrates interaction effects between the treatment and device type, where the treatment works differently for different subgroups. The primary concern should be developing device-specific solutions rather than implementing a one-size-fits-all treatment that harms mobile users. Option A incorrectly identifies this as Simpson's paradox and focuses on traffic volume rather than user experience. Option C misidentifies the issue as selection bias when proper randomization was likely conducted. Option D incorrectly suggests this is a multiple testing issue requiring statistical correction.

Question 18

A company runs two simultaneous but independent A/B tests. Test 1, a change to the homepage headline, shows a 10% lift in sign-ups. Test 2, a change to the 'Sign Up' button color, shows a 5% lift. Both are statistically significant. The company launches both changes. Three months later, the overall sign-up rate has increased by only 8%, less than the expected 15.5% (1.10 * 1.05 - 1). What is the most likely statistical explanation for this discrepancy?

  1. The novelty effect for both changes wore off after the launch, causing the real, long-term lift to be smaller than initially measured.
  2. A negative interaction effect occurred, where the combined impact of the two changes is less than the sum of their individual impacts. (correct answer)
  3. The initial results were likely Type I errors (false positives) that did not hold up over time due to random variation.
  4. Regression to the mean, as the initially high-performing test results naturally became more average after being launched to all users.
Explanation: When two changes are tested independently, their effects are measured in isolation. When deployed together, they can interact. A negative interaction effect means the two changes interfere with each other, or are redundant, making their combined effect smaller than what would be predicted by simply adding or multiplying their individual lifts. For example, a great headline might convince a user to sign up, making the button color less relevant. While novelty effect (A) or regression to the mean (D) could cause a general drop, the interaction effect (B) specifically explains why the combination underperforms relative to the sum of the parts. Type I error (C) is a possibility for any test, but interaction effects are a common reason for over-optimistic forecasts based on independent A/B tests.

Question 19

An online publisher wants to increase advertising revenue. They test a new page layout with more prominent ad slots (Variant B). The primary metric is 'ad revenue per user,' which increases by a statistically significant 8%. Which of the following is the most critical guardrail metric to monitor in this experiment?

  1. Click-through rate on the ads
  2. Number of unique visitors per group
  3. Average session duration or pages per session (correct answer)
  4. Total server costs associated with loading ads
Explanation: A guardrail metric is intended to monitor for unintended negative consequences, particularly to the user experience, that could harm the business in the long run. While increasing short-term ad revenue is the goal, bombarding users with ads could cause them to become frustrated, read fewer articles, and leave the site sooner (decreasing pages per session or session duration). This would harm long-term revenue potential. Therefore, engagement metrics are critical guardrails. Click-through rate (A) is a component of the primary metric, not a guardrail. Number of visitors (B) is a sanity check for randomization (a sample ratio mismatch check). Server costs (D) are a valid technical guardrail, but user flight is a more direct and critical business risk.

Question 20

An analyst identifies the 10% of customers with the lowest purchase frequency over the past year and targets this segment with a special promotional campaign. In the subsequent three months, the average purchase frequency of this group increases significantly. The analyst concludes the campaign was a success. Why is this conclusion potentially flawed, even if no other external factors changed?

  1. The promotional campaign may have caused the Hawthorne effect, where subjects changed their behavior simply because they knew they were being targeted.
  2. The observed increase is likely explained by regression to the mean, as a group selected for an extreme low value will tend to be less extreme on a subsequent measurement. (correct answer)
  3. A true A/B test was not conducted, as there was no control group of similarly low-frequency purchasers who did not receive the promotion.
  4. The three-month observation period was not long enough to establish a true change in purchasing behavior compared to the one-year baseline.
Explanation: Regression to the mean is a statistical phenomenon where an extreme result on one measurement is likely to be closer to the average on a second measurement. By selecting the lowest 10% of customers, the analyst selected a group whose performance was likely at a temporary, random low point. Some of this low performance was due to chance. In the next period, their performance would be expected to increase back toward their personal average and the overall population average, even without any intervention. While not using a control group (C) is the fundamental flaw in the experimental design, regression to the mean is the specific statistical phenomenon that explains why an increase would be observed even if the campaign had no effect.