Business Analytics Quiz: Segment And Cohort Analysis
10 questions · exam conditions
0:00
Segment And Cohort AnalysisQuestion 1 of 10

On July 15, an analyst compares 60-day retention for customers acquired in April with 60-day retention for customers acquired in May. Acquisition occurred throughout each month, and 60-day retention means being active during days 3131 through 6060 after acquisition.

Which approach produces the most valid cohort comparison as of July 15?

Compare all April and May customers because both calendar acquisition months have already ended
Compare April's 60-day rate with May's current active rate and label both as retention
Include only customers in each cohort who have had the full 60-day observation opportunity
Exclude April customers so that only the more recent May acquisition cohort is evaluated
← Back to quizzes

Business Analytics Quiz

Business Analytics Quiz: Segment And Cohort Analysis

Practice Segment And Cohort Analysis in Business Analytics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Segment And Cohort Analysis, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Analytics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

On July 15, an analyst compares 60-day retention for customers acquired in April with 60-day retention for customers acquired in May. Acquisition occurred throughout each month, and 60-day retention means being active during days 3131 through 6060 after acquisition.

Which approach produces the most valid cohort comparison as of July 15?

  1. Compare all April and May customers because both calendar acquisition months have already ended
  2. Compare April's 60-day rate with May's current active rate and label both as retention
  3. Include only customers in each cohort who have had the full 60-day observation opportunity (correct answer)
  4. Exclude April customers so that only the more recent May acquisition cohort is evaluated
Explanation: Cohort analysis questions test whether you understand observation completeness — the idea that every customer in a comparison group must have had the same opportunity to exhibit the behavior being measured. When a metric like "60-day retention" requires a full 60-day window, any customer who hasn't yet reached day 60 simply cannot have a valid retention measurement yet. The right approach, C, ensures that only customers who have had the full 60-day observation window are included in each cohort. Customers acquired on April 1 reached day 60 by May 31, and even April 30 customers completed their window by June 29 — all well before July 15. For May, however, customers acquired after May 15 haven't yet reached day 60 as of July 15. Including them would mean measuring retention over fewer than 60 days for some customers, which artificially deflates May's rate and makes the comparison invalid. Answer A is wrong because a calendar month ending doesn't guarantee every customer in that month has completed a 60-day observation window — May 31 acquired customers don't finish until July 30. Answer B is doubly flawed: it mixes two different metrics (a completed 60-day rate vs. a real-time snapshot) and mislabels them identically, which would produce a comparison that's both methodologically invalid and misleading. Answer D throws away April data entirely, which wastes a perfectly complete cohort and eliminates the comparison altogether. The strategic takeaway: whenever you see a cohort comparison question, ask yourself "has every member of each group had equal exposure time to the measured event?" If the answer is no, the comparison is biased — and the fix is to restrict each cohort to fully-observed customers only.

Question 2

Two subscription plans are evaluated across customer segments. For Plan X, 9090 of 100100 small-business customers and 400400 of 900900 enterprise customers renew. For Plan Y, 720720 of 900900 small-business customers and 4040 of 100100 enterprise customers renew.

Which conclusion best explains the renewal results?

  1. Plan X has higher renewal within both segments, but Plan Y is higher overall because its customer mix favors small businesses (correct answer)
  2. Plan Y has higher renewal within both segments, so the overall comparison confirms a consistent plan advantage
  3. Plan X has higher renewal overall, but Plan Y appears stronger because enterprise customers receive excessive weight
  4. The plans have equivalent segment performance, and their overall difference is caused only by unequal total sample sizes
Explanation: When a dataset is divided into subgroups, the overall (aggregate) result can contradict the subgroup results — a classic statistical trap called Simpson's Paradox. Whenever you see renewal or success rates broken down by segment, always compute both the segment-level and overall rates before drawing conclusions. Here, Plan X renews 90100=90%\frac{90}{100} = 90\% of small-business and 40090044%\frac{400}{900} \approx 44\% of enterprise customers. Plan Y renews 720900=80%\frac{720}{900} = 80\% of small-business and 40100=40%\frac{40}{100} = 40\% of enterprise customers. Plan X wins within every segment. Yet overall, Plan X renews 4901000=49%\frac{490}{1000} = 49\%, while Plan Y renews 7601000=76%\frac{760}{1000} = 76\%. Plan Y looks stronger in aggregate because it serves mostly small-business customers — the high-renewal segment — making its customer mix, not its actual plan quality, drive the overall number. That's exactly what A describes. B is wrong because Plan Y does not outperform Plan X within both segments — Plan X holds higher renewal rates in each one. C has the direction backwards: Plan Y is higher overall, not Plan X, so this answer contradicts the actual numbers. D is wrong because the plans do not have equivalent segment performance — Plan X consistently outperforms Plan Y within each group, and the reversal is driven by mix, not merely sample size. Your study tip: whenever overall rates conflict with segment rates, ask who makes up the majority of each group? The segment with the larger share pulls the aggregate toward its rate — that's Simpson's Paradox in action.

Question 3

A cohort contains 100100 customers acquired in June. During the cohort's first full month after acquisition, 6060 customers remain active and collectively generate 9,0009{,}000 dollars in revenue.

If the dashboard defines cohort revenue per acquired customer using the original cohort size as the denominator, what value should it report for that month?

  1. 6060 dollars, calculated from the number of customers who remain active
  2. 150150 dollars, calculated as revenue per active customer in the month
  3. 9090 dollars, calculated as revenue divided by the original cohort size (correct answer)
  4. 9,0009{,}000 dollars, because cohort revenue should not be converted to an average
Explanation: When you see a cohort analysis question, the key is identifying which population serves as the denominator. Cohort metrics are specifically designed to track performance relative to the original group size, not the surviving active users in any given period. This design choice lets analysts make fair comparisons across cohorts over time, regardless of churn. Here, the dashboard explicitly uses the original cohort size as the denominator. With 9,0009{,}000 dollars in revenue and 100100 originally acquired customers, the calculation is straightforward: 9,000100=90 dollars per acquired customer\frac{9{,}000}{100} = 90 \text{ dollars per acquired customer} That confirms C is correct. A is wrong because it divides by the number of active customers (6060), not the original cohort. That gives 9,00060=150\frac{9{,}000}{60} = 150 dollars — which is actually what B reports. Both A and B describe revenue per active customer, just with slightly different framing. The flaw in both is that they ignore the 40 customers who churned, making the metric look artificially favorable and incomparable across cohorts with different retention rates. D is wrong because cohort revenue is routinely normalized into a per-customer average — that's a standard and useful practice. Reporting 9,0009{,}000 as a raw total would make it impossible to compare cohorts of different sizes. Study tip: On business analytics questions, watch for whether the denominator is the original cohort size versus current active users — this distinction is one of the most commonly tested traps in cohort metric questions.

Question 4

An A/B test is summarized for new and existing customers. In the treatment group, new customers convert at 12%12\% and existing customers at 24%24\%; 80%80\% of treatment observations are new customers. In the control group, new customers convert at 10%10\% and existing customers at 20%20\%; only 20%20\% of control observations are new customers.

The unadjusted overall conversion rate is lower for treatment than for control. What is the best analysis?

  1. Conclude that treatment is harmful because the overall conversion rate is the only metric that matters for the experiment
  2. Apply consistent segment weights to both groups because treatment improves conversion within each customer segment (correct answer)
  3. Analyze only existing customers because they have the highest conversion rates in both groups and drive the most revenue
  4. Average the four reported segment rates equally because each percentage carries the same statistical weight
Explanation: Whenever you see an A/B test where segment-level trends contradict the overall result, you're looking at Simpson's Paradox — a statistical phenomenon where aggregated data reverses the direction seen within every subgroup. Here, treatment improves conversion in both segments: new customers go from 10%10\% to 12%12\%, and existing customers go from 20%20\% to 24%24\%. Yet the treatment group's blended rate appears lower because it contains 80%80\% new customers (the low-converting group), while control contains only 20%20\% new customers. The mix difference, not the treatment itself, drives the overall gap. The correct fix is B: apply consistent segment weights to both groups — a technique called standardization or reweighting. For example, using a 50/50 split, both groups would reflect the true per-segment lift rather than the distorted aggregate. A is wrong because blindly trusting the unadjusted overall rate ignores a known confound (segment mix). Concluding harm when every subgroup shows improvement is precisely the trap Simpson's Paradox sets. C is wrong because cherry-picking only existing customers discards valid data and introduces selection bias — you'd misrepresent treatment effects for the full customer base. D is wrong because the four segment rates don't carry equal statistical weight; each is based on different sample sizes, so a simple average produces a meaningless number. Your study tip: whenever segment compositions differ between test groups, never trust the raw aggregate. Always check whether segment mix could explain a counterintuitive result — this is a classic business-analytics exam trap.

Question 5

To identify behaviors associated with long-term retention, an analyst selects customers who are active at the end of the year and compares their first-month product usage. Customers who canceled before year-end are absent from the dataset.

Which revision would most improve the validity of the resulting cohort summary?

  1. Begin with all customers acquired in defined periods and determine later retention for each eligible customer (correct answer)
  2. Retain only year-end active customers but separate them by their first-month level of product usage
  3. Restrict the analysis to the highest-usage active customers to make behavioral differences more visible
  4. Combine customers from every acquisition year so the active-customer sample becomes substantially larger
Explanation: Whenever you see a question about cohort analysis or behavioral research, your first instinct should be to ask: who is missing from this dataset, and why? This question tests your ability to recognize survivorship bias — a classic validity threat where your sample only includes subjects who "survived" some selection process, systematically excluding others. The analyst's fatal flaw is starting with year-end active customers. These are, by definition, people who already retained. Comparing their early behaviors to predict retention is circular reasoning — you've pre-selected the outcome you're trying to explain. Choice A corrects this by starting with all customers acquired during defined periods and then tracking which ones retained. This prospective, unbiased approach lets you legitimately compare early behaviors across customers with different outcomes, giving your findings real predictive validity. Choice B keeps the same biased sample — only active customers — and just segments them differently. Segmenting a flawed group more finely doesn't fix the fundamental selection problem. Choice C makes things worse by narrowing the sample even further to high-usage active customers, introducing an additional filter that compounds the survivorship bias and eliminates the very behavioral variation you need to study. Choice D pools customers across acquisition years, which might increase sample size but does nothing to address the core issue: canceled customers are still absent, and a larger biased sample is still biased. As a study tip, watch for research scenarios where the dataset is defined by an outcome (active, successful, surviving) rather than by acquisition or eligibility — that's your signal that survivorship bias may be invalidating the analysis.

Question 6

An online service reports 30-day retention for two acquisition cohorts. The January cohort acquired 800800 customers and retained 40%40\%. The February cohort acquired 200200 customers and retained 70%70\%.

What is the appropriate combined 30-day retention rate across the two cohorts?

  1. 55%55\%, because the two cohort retention percentages should receive equal weight
  2. 46%46\%, because retained customers should be divided by total acquired customers (correct answer)
  3. 52%52\%, because the difference in cohort sizes should be partially weighted
  4. 44%44\%, because the larger January cohort determines nearly the entire result
Explanation: Whenever you see a question asking you to combine rates across groups of different sizes, your instinct should be to work from raw counts — not percentages. Averaging percentages directly is one of the most common traps in business analytics, and this question is designed to test exactly that. To find the true combined retention rate, you need to calculate the total number of retained customers and divide by the total number of acquired customers. The January cohort retained 800×0.40=320800 \times 0.40 = 320 customers. The February cohort retained 200×0.70=140200 \times 0.70 = 140 customers. Combined, that's 320+140=460320 + 140 = 460 retained out of 800+200=1,000800 + 200 = 1{,}000 acquired, giving a retention rate of 4601,000=46%\frac{460}{1{,}000} = 46\%. That confirms B is correct. Choice A falls into the classic trap of simple-averaging the two percentages: 40%+70%2=55%\frac{40\% + 70\%}{2} = 55\%. This treats both cohorts as if they were the same size, which they are not — January is four times larger. Choice C suggests "partial weighting," which isn't a real statistical method and produces a fabricated result with no mathematical basis. Choice D arrives at 44%44\% by over-anchoring on the January cohort and essentially ignoring February's contribution almost entirely, which also misrepresents the data. The key study tip: whenever combining rates or percentages across groups, always reconstruct the numerator and denominator separately, then divide. Never average the percentages themselves unless the groups are identical in size.

Question 7

A company defines acquisition month as cohort month 00, the following month as cohort month 11, and the next month as cohort month 22. In March, revenue is generated by customers acquired in January, February, and March.

Which statement correctly distinguishes a cohort summary from a calendar-period summary?

  1. January cohort month 22 includes all revenue generated by every customer during March
  2. March calendar revenue includes only customers whose acquisition occurred during March
  3. January cohort month 22 and March calendar revenue are necessarily equal by definition
  4. January cohort month 22 includes March revenue from January-acquired customers only (correct answer)
Explanation: When analyzing customer data, businesses use two fundamentally different lenses: cohort analysis (tracking a specific group of customers over time) and calendar-period analysis (capturing all revenue within a given time window, regardless of when customers were acquired). Keeping these perspectives distinct is exactly what this question tests. In the passage's framework, January is cohort month 00 for January-acquired customers, making February their cohort month 11 and March their cohort month 22. So January cohort month 22 refers specifically to revenue generated during March but only by customers acquired in January. That is precisely what D states — it correctly isolates one customer segment (January-acquired) within a calendar month (March). D is the right answer. A is wrong because it conflates cohort revenue with total calendar revenue. January cohort month 22 does not include revenue from February- or March-acquired customers, even though all of them transact in March. B inverts the logic of calendar-period analysis. March calendar revenue actually includes customers acquired in January, February, and March — not just March-acquired customers. Restricting it to March acquisitions would itself be a cohort filter. C is wrong because the two figures are almost never equal. March calendar revenue aggregates all three cohorts, while January cohort month 22 captures only a subset of that total. Study tip: On cohort questions, always ask two things: which customers are included, and which time window applies. A cohort metric filters by acquisition date; a calendar metric filters by transaction date. Mixing these up is the classic trap.

Question 8

A company has two customer segments. The premium segment contains 200200 customers and produces average monthly contribution margin of 8080 dollars per customer. The standard segment contains 1,0001{,}000 customers and produces average monthly contribution margin of 2020 dollars per customer. Management can pursue an initiative expected to increase premium margin by 10%10\% or a different initiative expected to increase standard margin by 5%5\%, with equal implementation costs.

Based only on the expected aggregate monthly margin increase, which initiative should management select?

  1. The premium initiative, because its expected incremental gain is 1,6001{,}600 dollars versus 1,0001{,}000 dollars for the standard initiative (correct answer)
  2. The standard initiative, because its current total margin of 20,00020{,}000 dollars exceeds the premium total of 16,00016{,}000 dollars
  3. The premium initiative, because its per-customer margin of 8080 dollars is four times the standard per-customer margin of 2020 dollars
  4. The standard initiative, because its customer count of 1,0001{,}000 is five times the premium customer count of 200200
Explanation: When comparing initiatives in business analytics, your job is to quantify the total incremental impact — not just react to ratios, counts, or existing totals. Here, that means calculating the aggregate margin increase each initiative would generate. For the premium initiative: 200 customers×$80×10%=$1,600200 \text{ customers} \times \$80 \times 10\% = \$1{,}600 in additional monthly margin. For the standard initiative: 1,000 customers×$20×5%=$1,0001{,}000 \text{ customers} \times \$20 \times 5\% = \$1{,}000 in additional monthly margin. The premium initiative produces $600\$600 more in incremental gain, making A the correct answer — and exactly the right framework to apply. Choice B is tempting because the standard segment's current total margin ($20,000\$20{,}000) exceeds the premium total ($16,000\$16{,}000), but the question asks about incremental gain from the initiative, not which segment currently earns more. Existing totals are irrelevant here. Choice C makes the right decision but for the wrong reason — citing the $80\$80 per-customer margin ignores both the percentage increase applied and the number of customers, so it's an incomplete and unreliable basis for comparison. Choice D commits a similar error in the opposite direction: a larger customer count only matters when combined with the margin and percentage change. Raw headcount alone tells you nothing about incremental value. The key study tip: whenever a question asks you to compare initiatives, convert everything into a single comparable output — here, total dollar impact. Watch out for distractors that isolate one variable (size, rate, or base margin) and ignore how all the components multiply together.

Question 9

A marketing dashboard classifies customers by every campaign they interacted with before purchase. Because many customers interacted with multiple campaigns, the displayed segment counts sum to more than the number of acquired customers. Management wants to report each channel's share of acquired customers, with all shares summing to 100%100\%.

Which segmentation change best supports management's objective?

  1. Retain every campaign interaction and divide each campaign count by the total interaction count
  2. Assign each customer to one consistently defined acquisition channel, such as the first-touch channel (correct answer)
  3. Assign customers to all observed channels but cap each reported channel share at 100%100\%
  4. Remove customers with multiple campaign interactions before calculating acquisition-channel shares
Explanation: When a question asks how to make segment shares sum to 100%100\%, you should immediately recognize the core problem: if customers are counted in multiple segments simultaneously, the denominator and numerator become misaligned, and shares will always exceed 100%100\% in aggregate. The fix requires mutual exclusivity — each customer must belong to exactly one segment. That's precisely why B is correct. By assigning each customer to a single, consistently defined acquisition channel — such as first-touch (the first campaign they interacted with) — you create non-overlapping groups. Now each acquired customer is counted once, so dividing each channel's count by total acquired customers produces shares that naturally sum to 100%100\%. The rule is consistent and auditable, satisfying management's reporting goal cleanly. A is tempting but flawed. Dividing each campaign count by total interactions (not total customers) measures interaction volume, not customer acquisition share. A customer who clicked three campaigns contributes three interactions, so the denominator inflates and the resulting shares don't reflect how many customers each channel actually acquired. C doesn't solve anything. Capping individual shares at 100%100\% is an arbitrary constraint that doesn't make the shares collectively sum to 100%100\% — it just prevents any single channel from exceeding that cap while the total across channels remains meaningless. D discards real data. Removing multi-touch customers introduces selection bias, since these customers often represent your most engaged segment. You'd be reporting acquisition shares on an unrepresentative subset. Study tip: Whenever shares must sum to 100%100\%, check for mutual exclusivity first — overlapping segment definitions are almost always the culprit, and the solution is a single assignment rule, not a mathematical patch.

Question 10

A retailer tracks the first repeat purchase for a fully matured cohort of 500500 customers. Of these customers, 120120 first repeat within days 11 through 3030, another 9090 first repeat within days 3131 through 6060, and another 9090 first repeat within days 6161 through 9090.

What is the cumulative 60-day repeat-purchase rate for this cohort?

  1. 18%18\%, using only customers whose first repeat occurs during days 3131 through 6060
  2. 24%24\%, using only customers whose first repeat occurs during days 11 through 3030
  3. 42%42\%, using customers whose first repeat occurs by the end of day 6060 (correct answer)
  4. 60%60\%, using customers whose first repeat occurs by the end of day 9090
Explanation: When you see "cumulative rate" in a repeat-purchase question, that word cumulative is your signal: you must count everyone who performed the behavior up to and including that cutoff, not just those who acted within a single window. The 60-day cumulative repeat-purchase rate asks: of the original 500 customers, how many made their first repeat purchase by the end of day 60? That includes the 120 who repeated in days 1–30 and the 90 who repeated in days 31–60, giving you 120+90=210120 + 90 = 210 customers. Dividing by the cohort size: 210500=0.42=42%\frac{210}{500} = 0.42 = 42\%. Answer C is correct. Answer A makes the mistake of isolating only the days 31–60 window: 90500=18%\frac{90}{500} = 18\%. This is a period rate, not a cumulative one — it ignores customers who already repeated earlier. Answer B isolates only the days 1–30 window: 120500=24%\frac{120}{500} = 24\%. This would be the 30-day cumulative rate, not the 60-day rate. Answer D adds all three periods — including days 61–90 — giving 120+90+90500=60%\frac{120 + 90 + 90}{500} = 60\%. That's the 90-day cumulative rate, overshooting the 60-day cutoff the question specifies. A useful rule of thumb: whenever a question names a specific day cutoff and uses the word "cumulative," stack every interval from day 1 through that cutoff. If the question asked for a single period's rate, it would say "during days 31–60," not "by day 60."