Business Analytics Quiz: Randomization And Experiment Design
10 questions · exam conditions
0:00
Randomization And Experiment DesignQuestion 1 of 10

A software company recruits 8,000 customers who voluntarily opened an in-app invitation to test a new dashboard. Within this group, customers are randomly assigned with equal probability to the new dashboard or the existing dashboard. After four weeks, the new-dashboard group has higher retention.

Which conclusion is best supported by the experiment's design?

The dashboard caused higher retention for all customers because random assignment also makes the volunteer sample representative.
The dashboard caused higher retention among invitation responders, but extending the result to nonresponders requires additional assumptions.
The dashboard is associated with higher retention, but no causal conclusion is possible because participation was voluntary.
The dashboard caused higher retention among all invited customers because every invitee had an opportunity to participate.
← Back to quizzes

Business Analytics Quiz

Business Analytics Quiz: Randomization And Experiment Design

Practice Randomization And Experiment Design in Business Analytics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Randomization And Experiment Design, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Analytics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A software company recruits 8,000 customers who voluntarily opened an in-app invitation to test a new dashboard. Within this group, customers are randomly assigned with equal probability to the new dashboard or the existing dashboard. After four weeks, the new-dashboard group has higher retention.

Which conclusion is best supported by the experiment's design?

  1. The dashboard caused higher retention for all customers because random assignment also makes the volunteer sample representative.
  2. The dashboard caused higher retention among invitation responders, but extending the result to nonresponders requires additional assumptions. (correct answer)
  3. The dashboard is associated with higher retention, but no causal conclusion is possible because participation was voluntary.
  4. The dashboard caused higher retention among all invited customers because every invitee had an opportunity to participate.
Explanation: When you see a question mixing random assignment with voluntary participation, you need to track two separate issues: internal validity (can we claim causation?) and external validity (who can we generalize to?). Random assignment within the 8,000 volunteers is the critical feature here. Because customers were randomly split between dashboards, any systematic differences between the two groups are due to chance alone — this controls confounding and supports a causal conclusion. The dashboard caused the retention difference, not some pre-existing difference between the groups. That much is solid. However, the 8,000 participants all self-selected by opening an in-app invitation. They may be more engaged, tech-curious, or loyal than typical customers who ignored the invite. This self-selection limits generalizability — the causal finding applies confidently only to people like those who responded. B captures exactly this: causal inference is warranted within the volunteer group, but extending results to nonresponders requires additional assumptions about their similarity. A is wrong because random assignment within a sample does not make that sample representative of a broader population. These are independent concepts — representativeness comes from how you recruit, not how you assign. C is wrong because it overcorrects. Voluntary participation undermines external validity, not internal validity. Random assignment still permits causal conclusions; C incorrectly strips away that ability. D is wrong because being invited doesn't mean the results apply to all invitees — only those who opted in were actually studied. A useful rule of thumb: random assignment → causation; random sampling → generalization. A study can have one without the other. Keep these two concepts separate on exam day.

Question 2

A grocery-delivery platform wants to test a referral credit. A treated customer can send the credit to friends who may otherwise be assigned to the control group. Customers are naturally organized into mostly nonoverlapping neighborhood social networks.

Which design best addresses the principal threat created by the referral mechanism?

  1. Randomize individual customers and exclude control customers who receive referrals before estimating the effect.
  2. Randomize neighborhood networks and compare outcomes using an analysis that accounts for network-level assignment. (correct answer)
  3. Randomize individual customers and add the number of referrals received as a post-treatment control variable.
  4. Assign the largest neighborhood networks to treatment and the smallest networks to control to maximize exposure.
Explanation: When an experiment risks "contamination" between treatment and control groups — where treated units can influence control units — you're dealing with interference or spillover. The referral mechanism here is a classic spillover threat: a treated customer can pass the credit to someone assigned to control, muddying the comparison between groups. The standard fix is to randomize at a level where spillover between groups is unlikely to occur. That's exactly what B does. By randomizing entire neighborhood networks — which are described as mostly nonoverlapping — you ensure that treated and control units belong to different social clusters. Spillover stays within treatment networks rather than crossing into control networks, preserving a clean comparison. Analyzing at the network level (e.g., using cluster-robust methods) correctly accounts for the fact that individuals within a network share an assignment. A is tempting but flawed. Excluding control customers who received referrals introduces post-treatment selection bias — you're selectively dropping units based on something that happened after randomization, which distorts your control group and biases estimates. C commits the same post-treatment error by adding referrals received as a covariate. Referrals received are a downstream consequence of treatment, not a pre-treatment confounder. Controlling for them blocks part of the causal pathway and can introduce bias rather than remove it. D abandons random assignment entirely. Systematically assigning large networks to treatment creates imbalance — larger networks may differ from smaller ones in ways that confound the outcome. Study tip: Whenever you see interference or spillover in an experiment, ask "at what level should randomization occur?" Cluster-level randomization along natural social boundaries is the go-to solution.

Question 3

A bank randomly assigns customers to receive either a short or long loan-application form. Completion can be measured for all assigned customers, but a satisfaction survey appears only after an application is submitted. The long-form group submits fewer applications, and submitted applications in that group receive slightly higher satisfaction scores.

What is the most defensible interpretation of the satisfaction result?

  1. The long form increased satisfaction because random assignment makes the survey respondents comparable across conditions.
  2. The long form did not affect satisfaction because the difference is smaller than the difference in submission rates.
  3. The satisfaction comparison may be biased because treatment affected which customers became eligible to answer the survey. (correct answer)
  4. The satisfaction comparison is unbiased if the same survey questions and rating scale were used in both conditions.
Explanation: Whenever an experiment measures an outcome only for people who take a voluntary action — submitting a form, responding to a survey, showing up for a follow-up — you need to ask whether the treatment itself changed who ends up in that measured group. This is the core concept of post-treatment selection bias, and it's exactly what this question tests. Here, satisfaction scores are only observable for customers who actually submitted an application. The long form caused fewer people to submit, meaning the long-form group's survey respondents are a self-selected subset — likely more motivated, patient, or committed customers. The short-form group's respondents are a broader, less filtered pool. So even though customers were randomly assigned at the start, the two groups being compared on satisfaction are no longer comparable. That's why C is correct: the comparison is potentially biased because the treatment determined who became eligible to respond. A is the most tempting trap. Random assignment does make groups comparable at the point of assignment, but that protection evaporates once you condition on a post-treatment behavior (submitting an application). Randomization cannot rescue a comparison made on a non-random subset. B makes a methodological claim with no logical basis — the size of the satisfaction difference relative to the submission-rate difference is irrelevant to whether the comparison is valid or biased. D confuses measurement consistency with study-design validity. Using identical survey questions removes measurement bias, but it does nothing to fix the sample selection problem created by differential dropout. The strategy to remember: any time an outcome is only measured for people who self-select into a post-treatment step, flag it as a potential selection bias problem — randomization at intake doesn't protect you.

Question 4

A food-delivery marketplace can activate a new dispatch algorithm only for the entire city at one time. Demand and traffic vary sharply by hour. The company plans a four-week experiment using hourly treatment periods.

Which assignment plan best supports a causal comparison of delivery time?

  1. Randomize the algorithm within comparable time blocks and account for dependence between nearby hours. (correct answer)
  2. Alternate algorithms every hour in a fixed pattern, beginning each day with the new algorithm.
  3. Use the new algorithm during weekday peak hours and the old algorithm during off-peak hours.
  4. Use the old algorithm for two weeks and the new algorithm for the next two weeks.
Explanation: When designing experiments where treatment can only be applied at the city-wide level (no individual-level randomization), you're dealing with a time-series or switchover experiment. The core challenge is isolating the algorithm's effect from confounding factors like time of day, day of week, and traffic patterns — while also managing the statistical complication that consecutive hours are correlated with each other. Option A is correct because it randomizes which algorithm runs within comparable time blocks (e.g., randomly assigning treatment among similar peak-hour slots across days), which neutralizes systematic confounders. Crucially, it also acknowledges autocorrelation — the fact that conditions in hour 3 PM are related to conditions in hour 4 PM — and accounts for it in the analysis. This produces a valid causal comparison. Option B uses a fixed alternating pattern, which sounds balanced but isn't random. Starting every day with the new algorithm means the new algorithm always runs during morning conditions, systematically confounding algorithm type with time-of-day effects. Option C is a classic confounding trap: the new algorithm only runs during peak hours and the old one during off-peak hours. Any difference in delivery times could simply reflect the inherent difficulty of peak-hour delivery, not the algorithm's effectiveness. Option D — two weeks old, two weeks new — is a simple before/after design. It cannot separate the algorithm's effect from any external changes that happened between weeks (weather shifts, demand changes, seasonal trends), making causal attribution impossible. Study tip: When an experiment assigns treatment by time rather than by individual, ask yourself: "Are the comparison periods truly exchangeable?" If they differ systematically in any way, you have confounding, not causation.

Question 5

In a randomized email experiment, 10,000 customers are assigned equally to treatment and control. Before viewing outcomes, analysts discover that the treatment group contains 54 percent prior purchasers while the control group contains 50 percent. Prior purchase strongly predicts the conversion outcome. The assignment code is verified as random.

Which response is most appropriate?

  1. Discard the assignment and rerandomize until prior-purchaser percentages are exactly equal in both groups.
  2. Analyze only customers who were not prior purchasers because their counts are closer across groups.
  3. Move enough prior purchasers from treatment to control to balance the groups before sending the emails.
  4. Keep the assignment and use a prespecified analysis that adjusts for prior-purchase status to improve precision. (correct answer)
Explanation: Whenever you see a question about randomized experiments with pre-treatment imbalance, focus on one key principle: random assignment is valid even when it produces imperfect balance by chance. Imbalance between groups is expected to occur sometimes — that's the nature of probability, not evidence of a broken process. Because the assignment was verified as random, any observed difference in prior-purchaser rates (54% vs. 50%) is simply sampling variation, not a systematic error. The correct move is D: keep the assignment and use a prespecified covariate adjustment (such as regression controlling for prior-purchase status). This approach is statistically principled — it removes noise from the outcome estimate caused by the imbalance, improving precision without introducing bias. Crucially, this adjustment must be prespecified to avoid the appearance of p-hacking. A is wrong because rerandomizing until groups look "perfectly equal" on one covariate destroys the integrity of the randomization process. It also only balances what you can observe — you'd still have unknown imbalances elsewhere. Chasing exact equality is both impractical and scientifically unsound. B is wrong because subsetting to non-prior-purchasers discards half your data arbitrarily, reduces statistical power, and limits how broadly you can apply your findings. You'd no longer be studying the population you intended to study. C is wrong because manually moving participants between groups after randomization introduces selection bias. The treatment group would no longer be a random sample, invalidating the entire experiment. Your takeaway: in experimental design, covariate adjustment preserves randomization while improving precision — it's the go-to tool when chance imbalance appears on a known predictor.

Question 6

An insurer randomly offers an analytics-based coaching program to 2,000 policyholders and does not offer it to another 2,000. Only 1,200 of those offered enroll, while 100 control-group policyholders obtain similar coaching from another source. The primary outcome is claim cost.

For the primary analysis, which comparison best preserves the protection provided by randomization?

  1. Compare everyone assigned an offer with everyone assigned no offer, regardless of actual coaching received. (correct answer)
  2. Compare coached policyholders with uncoached policyholders after matching them on observed claim-risk variables.
  3. Exclude crossovers from both groups and compare only policyholders who followed their assigned condition.
  4. Compare enrollees offered the program with control policyholders who did not obtain outside coaching.
Explanation: When a question asks about "preserving randomization" in an experimental study, you should immediately think about Intent-to-Treat (ITT) analysis — the gold standard for protecting the integrity of randomized designs. Randomization works by balancing both observed and unobserved confounders across groups. This balance exists at the moment of assignment, not at the moment of participation. Once you start filtering participants based on what they actually did (enroll, seek outside coaching, comply), you reintroduce selection bias — because who chooses to act is rarely random. The ITT principle solves this by analyzing participants exactly as assigned, no matter what happened afterward. Answer A is correct because comparing all 2,000 offered the program against all 2,000 not offered preserves the original randomization. Any non-compliance (the 800 who declined, the 100 controls who found outside coaching) is absorbed symmetrically, and the groups remain unbiased. The estimated effect will be conservative — diluted by imperfect take-up — but it remains causally valid. Answer B is wrong because matching on observed risk variables only controls for known confounders. Randomization controls for everything, including variables you haven't measured. Matching discards that protection. Answer C is wrong because excluding non-compliers from both groups creates a self-selected sample. People who follow instructions may differ systematically from those who don't, reintroducing exactly the confounding randomization was designed to eliminate. Answer D is wrong because it compares a self-selected subgroup (enrollees) against compliant controls — mixing compliance behavior across groups in an asymmetric, biased way. Study tip: On any RCT analysis question, if an answer filters by what participants did rather than what they were assigned to, it almost certainly violates ITT and introduces selection bias.

Question 7

A retailer is testing a demand-forecasting tool across 120 stores. Historical forecast error differs substantially between urban and rural stores, and there are 60 stores of each type. The retailer expects the tool's effect may also differ by store type.

Which assignment procedure most directly improves treatment-control comparability while preserving the ability to estimate an overall causal effect?

  1. Rank all stores by historical error, assign the highest-error half to treatment, and assign the remaining half to control.
  2. Randomize 60 stores to treatment, then verify afterward that urban and rural stores are approximately balanced.
  3. Within urban and rural stores separately, randomly assign half to treatment and half to control. (correct answer)
  4. Assign urban stores to treatment and rural stores to control, then adjust statistically for historical forecast error.
Explanation: When an experiment involves groups known to differ on a key variable — and where that variable may interact with the treatment — stratified random assignment is the gold standard. Here, store type (urban vs. rural) affects both baseline forecast error and potentially the tool's effectiveness, so your assignment strategy needs to account for it upfront, not as an afterthought. Option C does exactly this. By randomizing separately within each store type — 30 urban to treatment, 30 urban to control, 30 rural to treatment, 30 rural to control — you guarantee balance on store type by design. This eliminates the risk that randomization happens to cluster one type into treatment, and it preserves your ability to estimate both the overall effect and a subgroup comparison between urban and rural stores. Option A is a deliberate mismatch: assigning the highest-error stores to treatment creates a systematic confound between historical error and treatment status. Any observed difference could reflect error levels, not the tool's effectiveness. Option B is a common trap. Post-hoc balance checks are reassuring but not guaranteed — especially with only 60 stores per type, a single randomization can produce meaningful imbalance by chance. You're relying on luck rather than structure. Option D completely sacrifices comparability by design. Assigning an entire group to treatment and another to control means store type and treatment are perfectly confounded — no statistical adjustment can fully recover a causal estimate from that structure. Study tip: When you see a question involving known subgroups that differ on background characteristics, your first instinct should be stratified (blocked) randomization — it solves the comparability problem before data collection, not after.

Question 8

A subscription company will compare a redesigned cancellation flow with its current flow. The outcome's standard deviation is expected to be 12 units under the redesign and 8 units under the current flow. Assignment costs are equal, observations are independent, and the total sample size is fixed at 1,200. The company wants to minimize the variance of the estimated difference in means.

Approximately how should the sample be allocated?

  1. Assign 600 customers to the redesign and 600 customers to the current flow.
  2. Assign 900 customers to the redesign and 300 customers to the current flow.
  3. Assign 480 customers to the redesign and 720 customers to the current flow.
  4. Assign 720 customers to the redesign and 480 customers to the current flow. (correct answer)
Explanation: When comparing two groups with unequal variances, equal splitting is rarely optimal. The principle tested here is Neyman allocation: to minimize the variance of an estimated difference in means under a fixed total sample size, you should allocate proportionally to each group's standard deviation. The formula is: niσini=Nσiσ1+σ2n_i \propto \sigma_i \quad \Rightarrow \quad n_i = N \cdot \frac{\sigma_i}{\sigma_1 + \sigma_2} With σredesign=12\sigma_{\text{redesign}} = 12 and σcurrent=8\sigma_{\text{current}} = 8, the total weight is 12+8=2012 + 8 = 20. So: nredesign=12001220=720,ncurrent=1200820=480n_{\text{redesign}} = 1200 \cdot \frac{12}{20} = 720, \quad n_{\text{current}} = 1200 \cdot \frac{8}{20} = 480 This confirms D is correct. You assign more observations to the noisier group because doing so reduces that group's contribution to overall variance most efficiently. A is wrong because a 50/50 split ignores the differing standard deviations — it only minimizes variance when σ1=σ2\sigma_1 = \sigma_2. B (900/300) dramatically over-allocates to the redesign beyond what Neyman allocation prescribes and would actually inflate total variance. C (480 redesign / 720 current) inverts the logic entirely — it assigns more subjects to the less variable group, which is the opposite of what minimizes variance. A useful memory device: more noise needs more observations. When you see a fixed-sample allocation problem with unequal variances, immediately reach for Neyman allocation rather than defaulting to equal splitting, which is a common but costly assumption.

Question 9

A call center wants each agent to use an AI script for one week and the standard script for another week. Agents are randomly assigned to AI-first or standard-first sequences. Managers expect that agents may continue using techniques learned from the AI script after they switch back to the standard script.

Which design issue is most important, and what is the best response?

  1. Selection bias is most important; assign the most experienced agents to use the AI script first.
  2. Carryover is most important; use parallel randomized groups or add a credible washout period. (correct answer)
  3. Sampling bias is most important; increase the number of calls handled during each agent-week.
  4. Measurement error is most important; conceal from agents which script is considered the treatment.
Explanation: When you see a question involving a within-subject (crossover) experiment — where the same participants receive both treatments in sequence — your first instinct should be to ask: Can experiencing one condition permanently alter how someone performs in the next? That potential contamination is called carryover, and it's the central threat here. The passage explicitly tells you that agents may retain habits learned from the AI script even after switching back to the standard script. This means the two measurement periods aren't truly independent — the "standard script week" is polluted by AI-script learning. The best fixes are either switching to parallel randomized groups (so each agent only sees one script) or inserting a washout period long enough for the carryover effect to fade. Answer B identifies this correctly and prescribes both remedies. Answer A is wrong on two levels: it misnames the problem as selection bias and proposes the wrong fix. Deliberately assigning experienced agents to AI-first would actually introduce selection bias rather than eliminate it — it creates a confound between agent skill and script type. Answer C conflates the issue with sampling bias, which concerns whether your sample represents a broader population. Increasing calls per week doesn't address the sequencing contamination described in the passage at all. Answer D raises measurement error and suggests blinding agents to which script is "treatment." While blinding is a valid general tool, agents clearly know which script they're reading — this isn't a perception or labeling problem, so it's irrelevant to the core issue. Study tip: In crossover designs, always ask whether the first treatment can "leak" into the second period. If yes, carryover is your primary concern, and washout or parallel groups are your go-to solutions.

Question 10

An online retailer simultaneously tests two binary changes: personalized recommendations versus standard recommendations, and free returns versus paid returns. Customers are randomly assigned to one of the four resulting combinations. The retailer suspects that free returns may make personalization more effective.

Which analysis directly tests that suspicion while retaining the factorial design's advantage?

  1. Compare all personalized customers with all standard-recommendation customers, ignoring return policy.
  2. Compare free-return customers with paid-return customers, ignoring the recommendation assignment.
  3. Estimate whether the personalization effect differs between free-return and paid-return conditions. (correct answer)
  4. Compare the personalization-with-free-returns group only with the standard-with-paid-returns group.
Explanation: When you see a factorial experiment — multiple factors tested simultaneously across all combinations — the key advantage is the ability to detect interaction effects: whether one factor's impact depends on the level of another. The retailer's suspicion ("free returns may make personalization more effective") is precisely an interaction hypothesis, so your analysis must directly test whether the effect of personalization changes depending on return policy. That's exactly what C does. Estimating whether the personalization effect differs between free-return and paid-return conditions is the definition of testing an interaction. You're comparing the personalization lift in one condition against the personalization lift in another — and you're using all four groups, which preserves the statistical power and balance that a factorial design is built to deliver. A collapses across return policy entirely, giving you only the main effect of personalization. You'd learn whether personalization works on average, but you'd completely miss whether free returns amplify it — the exact question being asked. B makes the same mistake in the other direction: it collapses across recommendation type to isolate the main effect of return policy. Again, no interaction is tested. D is a common trap. Comparing only personalization-with-free-returns against standard-with-paid-returns confounds both factors simultaneously. You can't tell whether any difference comes from personalization, return policy, or their combination — it's an apples-to-oranges comparison that discards the other two groups and surrenders the design's advantage. Study tip: In factorial design questions, watch for "suspects that X makes Y more effective" — that phrasing always signals an interaction, and the correct analysis must explicitly estimate that interaction using all experimental cells.