Business Analytics Quiz: Translating Results To Actions
10 questions · exam conditions
0:00
Translating Results To ActionsQuestion 1 of 10

A retailer estimates an observational regression using one year of store data. After controlling for store size, region, and local income, stores open one additional hour have sales that are 12%12\% higher on average. Management is considering extending hours chain-wide, but stores historically chose their own hours based partly on expected evening demand.

What is the most defensible recommendation based on this analysis?

Extend hours chain-wide because controlling for several store characteristics makes the estimated 12%12\% relationship sufficiently causal for implementation.
Run a randomized or staggered pilot and evaluate incremental profit because anticipated demand may influence both operating hours and observed sales.
Reject longer hours because observational models cannot provide any information useful for selecting or designing a business pilot.
Forecast a 12%12\% sales increase at every store and extend hours only where that increase exceeds the additional wage expense.
← Back to quizzes

Business Analytics Quiz

Business Analytics Quiz: Translating Results To Actions

Practice Translating Results To Actions in Business Analytics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Translating Results To Actions, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Analytics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A retailer estimates an observational regression using one year of store data. After controlling for store size, region, and local income, stores open one additional hour have sales that are 12%12\% higher on average. Management is considering extending hours chain-wide, but stores historically chose their own hours based partly on expected evening demand.

What is the most defensible recommendation based on this analysis?

  1. Extend hours chain-wide because controlling for several store characteristics makes the estimated 12%12\% relationship sufficiently causal for implementation.
  2. Run a randomized or staggered pilot and evaluate incremental profit because anticipated demand may influence both operating hours and observed sales. (correct answer)
  3. Reject longer hours because observational models cannot provide any information useful for selecting or designing a business pilot.
  4. Forecast a 12%12\% sales increase at every store and extend hours only where that increase exceeds the additional wage expense.
Explanation: Whenever you see a regression estimated from observational data where the treatment (here, store hours) was chosen by the units being studied, your first instinct should be to ask: what drove that choice? That question is the heart of this problem. The passage tells you stores set their own hours based partly on anticipated evening demand. This creates a classic endogeneity problem — stores expecting high demand chose longer hours and had higher sales. The estimated 12%12\% relationship therefore conflates the causal effect of extra hours with a selection effect: high-demand stores self-selected into longer hours. Controlling for size, region, and income doesn't fix this because anticipated demand is unobserved. The most defensible path forward is B: run a randomized or staggered pilot. Randomly assigning hour extensions breaks the link between anticipated demand and treatment status, letting you isolate the true incremental effect. Evaluating profit (not just sales) is also essential, since extra hours add wage costs. A is wrong because adding control variables doesn't cure omitted-variable bias from an unobserved confounder like anticipated demand. More controls help, but they can't substitute for random assignment when a key driver of treatment is unmeasured. C is wrong because it overcorrects. Observational data is valuable — it can reveal where a pilot is most promising, help design the experiment, and set realistic expectations. Dismissing it entirely misunderstands how analytics and experimentation complement each other. D is wrong because it takes the biased 12%12\% estimate at face value, applying it mechanically chain-wide without acknowledging selection bias or testing causality first. Study tip: On business analytics questions, whenever units self-select into a treatment, treat any observed effect as suspect until tested experimentally — that's when a pilot earns its value.

Question 2

A distribution optimization model maximizes weekly contribution profit subject to labor, vehicle, and inventory constraints. Warehouse labor is binding, and its shadow price is 1818 dollars per additional hour for increases of up to 600600 hours. Temporary warehouse labor is available for 1414 dollars per hour. Vehicle and inventory constraints are expected to remain unchanged.

What is the most appropriate operational recommendation?

  1. Purchase up to 600600 temporary labor hours, subject to validating model assumptions, because each hour has an estimated net gain of 44 dollars. (correct answer)
  2. Purchase unlimited temporary labor because a positive shadow price guarantees the same marginal benefit at every possible capacity level.
  3. Purchase no temporary labor because the 1818-dollar shadow price represents an added cost rather than the value of added capacity.
  4. Purchase exactly 1818 temporary labor hours because the numerical shadow price specifies the optimal quantity of added capacity.
Explanation: When you see a question involving shadow prices and external resource costs, you're being tested on marginal analysis: comparing the value of one more unit of a constrained resource against its cost to acquire it. A shadow price tells you how much the objective function (here, contribution profit) improves for each additional unit of a binding constraint — but only within a specific right-hand-side range. Here, the shadow price of $18\$18 per labor hour means each added hour generates $18\$18 in additional profit, and this holds for up to 600600 additional hours. Temporary labor costs $14\$14 per hour, so the net benefit is $18$14=$4\$18 - \$14 = \$4 per hour. That makes purchasing up to 600600 hours financially justified — which is exactly what A recommends, correctly adding the caveat of validating model assumptions before acting. B is wrong because shadow prices are only valid within a bounded range. Beyond 600600 hours, the shadow price may change as other constraints become binding — "unlimited" purchasing ignores this critical limitation. C confuses the shadow price's meaning. It isn't an added cost; it's the marginal value of relaxing the constraint. The real cost is the $14\$14 temporary labor rate, which is less than $18\$18, making the purchase beneficial. D misreads what the shadow price communicates. The $18\$18 figure is a value per unit, not a recommended quantity. Buying exactly 1818 hours has no analytical basis. Study tip: Always pair a shadow price with its valid range, then compare it against the acquisition cost. Net benefit = shadow price − cost. If positive and within range, act.

Question 3

A pre-registered randomized test evaluates a new onboarding flow. New customers and returning customers each represent half of test participants. Conversion increases by 66 percentage points for new customers but decreases by 11 percentage point for returning customers. Both subgroup effects are statistically reliable, the treatment has negligible variable cost, and the overall lift is 2.52.5 percentage points.

Which deployment recommendation best uses the experimental results?

  1. Deploy the new flow to every customer because the positive overall lift is the most stable estimate for all future users.
  2. Deploy the new flow only to returning customers because their smaller measured effect is less likely to be overstated.
  3. Retain the existing flow for everyone because opposite subgroup effects invalidate the positive overall experimental result.
  4. Deploy the new flow to new customers but retain the existing flow for returning customers, while monitoring segment outcomes. (correct answer)
Explanation: When an experiment reveals heterogeneous treatment effects — where the intervention helps one segment but hurts another — the right business decision is almost never a single blanket policy. Your job is to match the deployment strategy to what the data actually show, not to average away meaningful differences. Here, the new onboarding flow produces a +6+6 pp lift for new customers and a 1-1 pp drop for returning customers, and both effects are statistically reliable. That means the data give you strong, actionable signal at the segment level. The +2.5+2.5 pp overall lift is simply a weighted average (0.5×6)+(0.5×1)=2.5(0.5 \times 6) + (0.5 \times -1) = 2.5 that masks the opposing dynamics. D is correct because it acts on the disaggregated evidence: new customers get the superior flow, returning customers keep what works for them, and ongoing monitoring guards against surprises in production. A is wrong because deploying universally ignores the confirmed harm to returning customers — hiding bad news inside an average is not a sound business practice. B inverts the logic entirely; the flow hurts returning customers by 11 pp, so deploying it to them would deliberately damage that segment. The reasoning about "smaller effects being less overstated" doesn't justify deploying a treatment that reduces conversion. C is wrong because opposite subgroup effects don't invalidate the experiment — they enrich it. Discarding reliable, directional findings is a waste of valid experimental information. Study tip: When you see statistically significant subgroup effects pointing in opposite directions, always recommend a segmented deployment strategy rather than a one-size-fits-all decision based on the aggregate.

Question 4

An online retailer randomly assigns 50,00050{,}000 visitors to its current checkout and 50,00050{,}000 to a checkout offering a coupon. Conversion increases from 10.0%10.0\% to 10.6%10.6\%, with a two-sided significance test producing p=0.03p=0.03. Each order generates 3030 dollars of contribution before the coupon, and every treatment-group purchaser redeems a coupon costing the retailer 55 dollars.

Assuming no other material effects, which rollout decision is best supported?

  1. Roll out the coupon because the statistically significant conversion increase establishes that the treatment also increases contribution profit.
  2. Do not roll out because the incremental contribution from new orders is smaller than the coupon cost paid on treatment orders. (correct answer)
  3. Roll out only to half of visitors because the conversion lift is significant but the experiment used equal-sized treatment groups.
  4. Do not roll out because a statistically significant result requires a conversion increase of at least one full percentage point.
Explanation: Whenever a business experiment shows a statistically significant result, your instinct might be to roll out the change immediately — but statistical significance and economic significance are separate questions. You must always run the numbers on actual profit impact before recommending a decision. Here's the math. The treatment group had 50,000×10.6%=5,30050{,}000 \times 10.6\% = 5{,}300 purchasers, versus 50,000×10.0%=5,00050{,}000 \times 10.0\% = 5{,}000 in the control. The incremental orders from the lift are 5,3005,000=3005{,}300 - 5{,}000 = 300. Those 300 new orders generate 300×$30=$9,000300 \times \$30 = \$9{,}000 of additional contribution. However, the coupon is redeemed by every treatment purchaser — all 5,300 of them — costing 5,300×$5=$26,5005{,}300 \times \$5 = \$26{,}500. The net impact is $9,000$26,500=$17,500\$9{,}000 - \$26{,}500 = -\$17{,}500. Rolling out the coupon destroys profit, making B correct. A is wrong because statistical significance only tells you the conversion lift is real — it says nothing about whether contribution profit increases. Confusing these two is a classic trap. C is wrong because splitting rollout to "half of visitors" has no logical basis; treatment group size in an experiment doesn't dictate deployment strategy. D is wrong because there is no rule requiring a one-percentage-point minimum lift for significance — this threshold is invented and irrelevant. The key study tip: always separate the question "Is the effect real?" (statistics) from "Is the effect worth it?" (economics). On business-analytics exams, profitable rollout requires that incremental revenue exceeds all incremental costs, including costs paid on your existing customers.

Question 5

A staffing forecast has an overall mean absolute percentage error of 8%8\%. Validation by region shows errors of 5%5\% in Regions B and C, but 25%25\% in Region A. Region A's forecasts also systematically understate demand, while errors in Regions B and C are approximately unbiased. Management wants to automate next month's staffing decisions in all regions.

Which delivery plan best translates the validation results into action?

  1. Automate staffing in all regions because the overall 8%8\% error is low enough to represent each region's expected performance.
  2. Deploy only in Region A because systematic underforecasting can be corrected simply by adding the overall 8%8\% error.
  3. Reject the model for every region because one weak segment makes the aggregate performance estimate analytically invalid.
  4. Deploy in Regions B and C, use an override in Region A, and recalibrate Region A before broader automation. (correct answer)
Explanation: When validating a forecasting model, aggregate performance metrics can mask critical variation across subgroups — and deployment decisions should reflect segment-level performance, not just the overall average. Here, the overall 8%8\% MAPE looks acceptable, but it blends two very different realities: Regions B and C perform well at 5%5\% with unbiased errors, while Region A performs poorly at 25%25\% with systematic underestimation. These are distinct problems requiring distinct responses. The right move — confirmed by answer D — is to deploy automation where the model is trustworthy (Regions B and C), apply a human override in Region A to compensate for known underforecasting, and recalibrate Region A's model before extending automation there. This is responsible, evidence-based deployment. A is wrong because it treats the aggregate 8%8\% as representative of all regions. It isn't — Region A's 25%25\% error is obscured by averaging, and automating staffing there based on a misleading summary would propagate real operational harm. B is wrong on two counts: Region A is the weakest region, so deploying only there is backwards. Worse, correcting systematic bias by adding the overall error rate (8%8\%) rather than Region A's actual bias is statistically unsound — you'd be applying the wrong correction factor. C is wrong because it overcorrects. One underperforming segment doesn't invalidate the entire model. Regions B and C demonstrate strong, unbiased performance, and discarding them wastes a genuinely useful tool. Study tip: Whenever you see aggregate model metrics alongside segment-level breakdowns, always ask whether the aggregate is masking heterogeneity — exam questions frequently test exactly this trap.

Question 6

A fraud model can operate at either of two thresholds for each batch of 10,00010{,}000 transactions. The lower threshold catches 9090 fraudulent transactions and falsely flags 500500 legitimate transactions. The higher threshold catches 6565 fraudulent transactions and falsely flags 100100 legitimate transactions. Reviewing any flagged transaction costs 2020 dollars, catching a fraudulent transaction avoids an average loss of 2,0002{,}000 dollars, and the review team can process at most 400400 flags per batch.

Which action best reflects both the model economics and the operating constraint?

  1. Use the lower threshold because its larger number of caught frauds gives it the highest expected value regardless of review capacity.
  2. Use the higher threshold because it is feasible and yields positive expected value, while the lower threshold exceeds review capacity. (correct answer)
  3. Use neither threshold because the review cost must be compared with the total value of every transaction rather than avoided fraud loss.
  4. Use the lower threshold but review only 400400 randomly selected flags because random sampling preserves all of the model's expected benefit.
Explanation: When evaluating a classification model in a business context, you must always check two things: feasibility (can we actually execute this?) and economics (does the value outweigh the cost?). Both conditions must hold. Start with feasibility. The lower threshold flags 500500 transactions, but the review team can handle at most 400400. It's simply not executable as designed. The higher threshold flags only 100100 transactions — well within capacity. Now check the economics of the higher threshold. The cost of reviewing all flags is 100×$20=$2,000100 \times \$20 = \$2{,}000. The value of catching 6565 fraudulent transactions is 65×$2,000=$130,00065 \times \$2{,}000 = \$130{,}000. Net expected value: $130,000$2,000=$128,000\$130{,}000 - \$2{,}000 = \$128{,}000. That's strongly positive, confirming B is the right choice. A is tempting because catching 90 frauds sounds better than 65, but it ignores the hard capacity constraint of 400 reviews. A model that can't be operationalized has zero real-world value — more theoretical fraud-catching doesn't help if your team is overwhelmed. C introduces a nonsensical benchmark. You never compare review cost against the value of every transaction in a batch; you compare it against the avoided loss from the fraudulent ones you actually catch. This distractor misrepresents basic cost-benefit framing. D sounds clever but is fundamentally flawed. The model's value comes from its targeted identification of likely fraud. Randomly sampling from its flags throws away the discriminative power entirely — you'd be no better than random guessing among those 400 reviews. Study tip: On model-deployment questions, always resolve feasibility constraints before calculating expected value — a technically superior threshold that violates an operating constraint is never the right answer.

Question 7

After a price increase, a subscription company reports that quarterly revenue rose by 8%8\% and gross profit rose by 2%2\%. During the same period, units sold fell by 4%4\%, the six-month repeat-purchase rate declined by 66 percentage points, and billing-related complaints increased by 35%35\%. No control group or comparable untreated market is available.

What should the analyst recommend to senior management?

  1. Expand the price increase immediately because positive revenue and gross-profit changes are sufficient evidence of sustainable customer value.
  2. Reverse the increase immediately because declining units and repeat purchases prove that the new price caused long-term profit to fall.
  3. Hold broader expansion, analyze cohort lifetime value, and run a controlled pricing test with profit and retention guardrails. (correct answer)
  4. Ignore retention and complaint measures because financial metrics should always take priority over nonfinancial operational indicators.
Explanation: When you see mixed signals across financial and operational metrics — some positive, some negative — your job is to synthesize them carefully rather than cherry-pick. This question tests whether you can distinguish short-term revenue gains from sustainable business health, and whether you understand the limits of causal inference without a control group. The right move here is C. The revenue and gross-profit increases look encouraging, but the concurrent drop in units sold, the 6-6 percentage-point repeat-purchase rate, and the $$35%$ spike in billing complaints are serious warning signs about customer retention and long-term lifetime value (LTV). Crucially, without a control group or comparable untreated market, you cannot confidently attribute any of these changes to the price increase alone — confounders may be at play. The prudent recommendation is to pause broad rollout, model cohort-level LTV to see whether retained customers are worth more over time, and design a controlled pricing experiment with guardrails on both profit and retention. A is tempting but dangerous — it commits the trap of treating short-term revenue growth as proof of sustainable value while ignoring clear deterioration in customer behavior. Revenue can rise briefly even as you lose your customer base. B overcorrects in the opposite direction. Declining units and repeat-purchase rates are concerning signals, but they don't prove causation without a control group, and an immediate reversal foregoes the opportunity to learn through rigorous testing. D reflects a common but flawed heuristic. Nonfinancial metrics like retention and complaints are leading indicators of future financial performance — ignoring them is analytically unsound. On exam questions like this, watch for "some metrics up, some metrics down" setups — they almost always reward the answer that calls for further analysis over immediate action.

Question 8

A business-to-business supplier uses a model to rank overdue invoices by probability of default. Among the highest-ranked invoices, the estimated default probability is 36%36\%; among all remaining invoices, it is 8%8\%. A randomized pilot found that contacting an account reduces its default probability by 25%25\% on a relative basis. Each contact costs 77 dollars, and preventing a default avoids an average loss of 200200 dollars. The collections team can contact only 1,0001{,}000 accounts this month.

Which action best translates these results into a collections decision?

  1. Contact the 1,0001{,}000 highest-ranked accounts because their expected net benefit is positive, while contacting typical remaining accounts would destroy value. (correct answer)
  2. Contact a random 1,0001{,}000 accounts because the randomized pilot establishes effectiveness but makes the model ranking irrelevant to the decision.
  3. Contact the 1,0001{,}000 lowest-ranked accounts because preventing an unexpected default creates more value than preventing a likely default.
  4. Contact no accounts because a 25%25\% relative reduction is smaller than the 36%36\% predicted default probability for the highest-ranked group.
Explanation: When a model scores accounts by default probability, the key question is whether acting on that ranking creates positive expected value — and whether targeting the highest-risk group outperforms the alternative. Work through the math for each group. For the highest-ranked accounts: a 25% relative reduction on a 36% baseline means the contact reduces default probability by 0.25×0.36=0.090.25 \times 0.36 = 0.09, or 9 percentage points. The expected benefit per contact is 0.09×$200=$180.09 \times \$200 = \$18, against a cost of $7\$7. Net benefit: $18$7=+$11\$18 - \$7 = +\$11 per account. For typical remaining accounts: the same relative reduction on an 8% baseline gives 0.25×0.08=0.020.25 \times 0.08 = 0.02, or 2 percentage points saved. Expected benefit: 0.02×$200=$40.02 \times \$200 = \$4, which is less than the $7\$7 contact cost. Net benefit: $3-\$3 per account — value destruction. This confirms A is correct: contact the 1,000 highest-ranked accounts, where the net benefit is positive, and skip the rest, where it is negative. B is wrong because the pilot established that contacting works, not that who you contact doesn't matter. The model ranking is precisely what tells you where that effect generates positive returns. C is backwards — lower default probability means smaller expected savings per contact, making low-risk accounts worse targets, not better. D confuses a relative reduction (25%) with the baseline probability (36%); these aren't being compared to each other — the net-benefit calculation is what matters. Remember: always convert a relative treatment effect into an absolute probability change, then multiply by the avoided loss before comparing to cost.

Question 9

A service director reports that average case-resolution time declined from 1414 hours last quarter to 11.411.4 hours this quarter. Last quarter, standard cases averaged 88 hours and enterprise cases averaged 2020 hours, with each type representing half of all cases. This quarter, standard cases averaged 99 hours and enterprise cases averaged 2121 hours, while the enterprise share fell to 20%20\%.

Which interpretation should the analyst communicate to management?

  1. Operational performance improved because the overall average fell, so the current process should be standardized across both case types.
  2. Operational performance deteriorated within both case types; the lower overall average mainly reflects a shift toward faster standard cases. (correct answer)
  3. Enterprise performance improved because its lower case share reduced its contribution to the overall average resolution time.
  4. The evidence is inconclusive because an overall average cannot be reconciled with separate averages for standard and enterprise cases.
Explanation: Whenever you see overall averages changing alongside subgroup averages, you should immediately think about Simpson's Paradox — the phenomenon where a trend in aggregated data reverses or disappears when the data is disaggregated. Here, the math confirms the paradox precisely. Last quarter's overall average: 0.5(8)+0.5(20)=140.5(8) + 0.5(20) = 14 hours. This quarter: 0.8(9)+0.2(21)=7.2+4.2=11.40.8(9) + 0.2(21) = 7.2 + 4.2 = 11.4 hours. Notice that standard cases worsened from 898 \to 9 hours and enterprise cases worsened from 202120 \to 21 hours — yet the overall average fell. The drop is entirely explained by the mix shift: enterprise cases shrank from 50%50\% to 20%20\% of volume, pulling the overall average down despite both subgroups deteriorating. Answer B captures this exactly — operational performance declined within every case type, and the favorable headline number is a compositional illusion. Answer A is dangerously wrong because it treats a misleading aggregate improvement as a reason to standardize current processes, which would lock in worse per-type performance. Answer C misattributes causality — enterprise's smaller share did reduce its contribution to the average, but that is not the same as enterprise performance improving; its resolution time actually increased. Answer D is flatly incorrect: the separate averages are perfectly reconcilable with the overall average through weighted arithmetic, as shown above. Your strategy: whenever a passage gives you subgroup averages and an overall average, always verify the weighted calculation yourself. If the overall trend contradicts the subgroup trends, you're looking at a mix-shift effect — one of the most common analytical traps tested in business analytics.

Question 10

A seasonal product has forecast demand that is approximately normal with a mean of 1,0001{,}000 units and a standard deviation of 120120 units. An unsold unit creates a clearance loss of 1010 dollars, while a stockout loses 4040 dollars of contribution. The appropriate critical fractile is therefore 40/(40+10)=0.8040/(40+10)=0.80, and the standard-normal 8080th percentile is approximately 0.840.84.

Which inventory action best translates the forecast and cost asymmetry into a stocking decision?

  1. Stock approximately 1,0001{,}000 units because an unbiased point forecast minimizes every type of forecast-related business loss.
  2. Stock approximately 900900 units because the clearance loss requires subtracting one forecast standard deviation from expected demand.
  3. Stock approximately 1,1001{,}100 units because stockouts are more costly, making the 8080th demand percentile appropriate. (correct answer)
  4. Stock approximately 1,1201{,}120 units because one full standard deviation always provides the optimal service level under uncertainty.
Explanation: Whenever you see a seasonal or one-time stocking decision with asymmetric costs, you're in newsvendor territory. The core idea: because overstocking and understocking carry different penalties, the optimal stock level is not simply the mean demand — it's a specific percentile of the demand distribution determined by the cost ratio. The critical fractile formula tells you which percentile to target: CF=CuCu+CoCF = \frac{C_u}{C_u + C_o}, where CuC_u is the cost of understocking (stockout) and CoC_o is the cost of overstocking (clearance). Here, CF=4040+10=0.80CF = \frac{40}{40+10} = 0.80. This means you should stock enough to satisfy demand 80% of the time. The 80th percentile of a normal distribution with mean μ=1,000\mu = 1{,}000 and σ=120\sigma = 120 is 1,000+0.84×1201,1011{,}000 + 0.84 \times 120 \approx 1{,}101, which rounds to approximately 1,1001{,}100 units — making C correct. A is wrong because stocking at the mean (50th percentile) ignores the cost asymmetry entirely. An unbiased forecast minimizes squared error, not profit. B is wrong in both direction and logic. Subtracting a standard deviation targets the 16th percentile, which would be appropriate only if overstocking were far more costly than understocking — the opposite of this scenario. D is wrong because adding exactly one full standard deviation (1,1201{,}120) corresponds to the 84th percentile, not the 80th. The zz-score of 0.840.84 scales with the standard deviation; it doesn't equal it. Study tip: Memorize the newsvendor setup — draw the number line, place the mean, then shift up when Cu>CoC_u > C_o and down when Co>CuC_o > C_u. The direction of the shift always follows the more expensive mistake.