The Overall Rate Improves While Every Customer Segment Gets Worse. What Should the Deck Show?
The Overall Rate Improves While Every Customer Segment Gets Worse. What Should the Deck Show?
Show both. The overall rate and the segment rates are not competing claims about the same thing, and a deck that picks one of them is not simplifying — it is deleting half of the sentence the room needs.
Here is the whole mechanism in one line: an overall rate is a weighted average of segment rates, and the weights are each segment's share of the observations. Move the shares and the average moves, even when every segment's own rate falls. The aggregate is not lying. It is answering "how did the program do, given who came through it," and that is a different question from "how did each kind of customer do."
The rest of this is about how to put both answers on a slide without letting either one impersonate the other.
The arithmetic, with numbers you can check by hand
The figures below are constructed for illustration, not an observed company result. They are stated as counts so every percentage can be verified.
| Period 1 | Eligible | Successes | Rate | Share of observations |
|---|---|---|---|---|
| Segment A | 100 | 90 | 90% | 10% |
| Segment B | 900 | 180 | 20% | 90% |
| Total | 1,000 | 270 | 27% | 100% |
| Period 2 | Eligible | Successes | Rate | Share of observations |
|---|---|---|---|---|
| Segment A | 900 | 720 | 80% | 90% |
| Segment B | 100 | 10 | 10% | 10% |
| Total | 1,000 | 730 | 73% | 100% |
Both segments fall by exactly 10 percentage points. The total rises by 46. Note also that the eligible count is 1,000 in both periods: nothing here reflects growth or contraction in volume. Only the composition moved, and it moved a long way — Segment A went from a tenth of the observations to nine tenths, and Segment B did the reverse.
First, check that the comparison is legal
Before any of this arithmetic is worth doing, four definitions have to hold across both periods: one eligible observation, one successful outcome, the segment boundaries, and the observation window. Let the counts carry this. If "completed" meant a full checkout in period 1 and a confirmed order in period 2, you do not have a mix problem, you have two different metrics wearing the same word, and no amount of weighting fixes it. That is a separate repair, and it comes first.
The same test applies to the segment boundaries. A shift this large — A going from 10% to 90% of observations — is worth a specific question before it goes on a slide. Did anything in the product, the funnel, or the segmentation rule change how observations get assigned to A and B? If segment membership is derived from a field that was redefined, then the segment rates are as incomparable as two different metric definitions, and the interesting finding is about the field, not the customers. Ask that question in the analysis, not in the meeting.
Only once those are settled does the reversal mean what it appears to mean: two comparable measurements of the same thing, moving in opposite directions at different levels of aggregation.
Rebuild both totals from the counts
Twenty-seven percent is not the average of 90% and 20%. That average is 55%, and it corresponds to no population.
What produces 27% is this:
0.10 × 90% + 0.90 × 20% = 9% + 18% = 27%
And 73% is this:
0.90 × 80% + 0.10 × 10% = 72% + 1% = 73%
Those are the same shares that appear in the tables above. Writing the total this way does two jobs at once: it shows that the aggregate is arithmetically correct, and it shows exactly where the movement came from. Segment A's rate fell, but A's weight in the calculation rose from 0.10 to 0.90 — nine times its former share. Segment B's rate fell and B's weight collapsed. The average followed the weights.
Here is the trap worth naming out loud, because someone in the room may reach for it. The unweighted mean of the two segment rates goes from 55% to 45% — it falls by 10 points, which happens to be exactly the segment story. So the wrong calculation produces the direction that feels right. That coincidence is precisely why it is dangerous: an indefensible number that agrees with your intuition is much harder to catch in review than one that disagrees. Never average subgroup percentages when the denominators differ. Here they differ by a factor of nine in each period.
This pattern has a standard name. Section 2 of the Stanford Encyclopedia of Philosophy's entry on Simpson's Paradox sets out the weighted-average characterization — how an aggregate association can diverge from the within-group associations when group weights differ. It is authored mathematical exposition, and it supplies the mechanism only; it offers no account of why a mix changes, and the entry's medical examples and causal discussion are not what is being borrowed here. The name is worth using once. The arithmetic is what does the work, because the name alone tells the audience nothing about which weight moved.
The slide that shows the reversal without arguing about it
Keep the total-only slide. Then put a counts-and-rates panel beside it. The panel is a table with a row per segment and a total row, and columns for eligible count, success count, rate, and share of observations, one block per period. That is the whole design. It is unglamorous and it ends the argument, because every number in it is measured.
Under the panel, one line of prose carries the change:
Overall rate 27% → 73%. Both segments fell 10 points. Segment A's share of observations rose from 10% to 90%.
That is the entire finding, stated so that both halves survive. Nothing in it says "despite." Nothing says "adjusted." The audience can hold two facts.
One display note, since this audience will read the numbers closely: 27% to 73% is +46 percentage points, which is a relative increase of about 170%. Say which one you mean, and put points in the axis label. The two readings differ by a factor of nearly four, and slide text is where that ambiguity usually starts.
If the room needs the change split rather than just stated, add a decomposition. Working from period-1 weights and period-1 rates as the base:
- within-segment change: −10 points
- mix change: +56 points
- interaction: 0 points
- total: +46 points
The mix term is the one people find striking, and it is the one that most needs a caution attached: it is named for where the arithmetic puts it, not for a discovered cause. Note also why the interaction is zero here. Both segments moved by exactly 10 points, which is a special case. With unequal segment changes the interaction is nonzero, and the split between "within" and "mix" then depends on whether you used period-1 weights or period-2 weights as the base. In this example either base gives the same answer — a consequence of the tidy construction, not a general property. Real data rarely cooperates. Declare the base, and expect a reviewer to redo it with the other one.
Add a fixed-mix row only when you can name its question
If the discussion turns to how the two periods compare as a matter of segment performance rather than composition, you can standardize: apply one declared set of weights to both periods' segment rates.
At period-1 shares: 0.10 × 80% + 0.90 × 10% = 17%. So period 2's rate would have been 17% had the mix stayed as it was, against period 1's actual 27%.
That comparison can also be run the other way. At period-2 shares, period 1's rate would have been 0.90 × 90% + 0.10 × 20% = 83%, against period 2's actual 73%.
Both directions give a 10-point decline — again because both segments moved by the same amount. When the within-segment changes differ, the two standardizations will disagree, and you will have to choose one and defend the choice, or show both. Standardizing to the earlier period's mix is the more common convention because it holds the starting composition fixed; that is a convention, not a truth.
Four rules make this row safe to display.
Name the weights in the row or its footnote. "Standardized to period-1 mix" plus the two shares. Not "adjusted," not "underlying." A reader who cannot see which weights were used cannot check the number.
Do not put it in the headline, and do not call it the real result. It answers: how would the two periods compare if the same composition of observations had come through both? It does not answer what the business earned, because it holds period 2's segment rates exactly as observed and assumes the shift in mix left them untouched. That assumption is untestable from these rows. On a slide, the tell is visible: every observed row has a numerator and a denominator; a standardized row has a rate and a weight set and no counts. If a number has no denominator, it was constructed.
Keep the observed total in the same view. The mix change may be the most operationally interesting fact in the deck. If volume moved toward a particular segment on purpose, or if it moved because something broke, the actual 73% is the number that reflects it. Replacing the observed figure with the standardized one would hide the outcome of whatever happened.
Say the question out loud before the number lands. The audience's default assumption about a large number on a slide is that it is what happened. A standardized figure violates that assumption unless you announce it. One sentence at the top of the slide, in the presenter's voice, is enough: "This next row answers a different question — what if the mix had not changed?"
What the arithmetic cannot say
The reversal is a fact about weights. It is not evidence about behavior, and the deck should not let the visual drama of two lines crossing imply more than the counts support.
Do not say the mix shift caused the segment declines, or that segment declines caused the mix shift. The decomposition splits a total change into pieces by construction; it does not order events or assign responsibility. A rising share for the higher-performing segment and falling rates inside both segments can coexist under any number of real stories, and these two rows cannot distinguish among them.
Do not attach significance language. Nothing here was designed as an experiment, the two periods may not be independent samples of anything, and the identical eligible counts are a convenience of the construction. "Significant" in a deck should mean a stated test with stated assumptions, not a large gap.
Do not describe the aggregate as misleading. It is an accurate summary of the program as it ran. The defect is not in the number; it is in presenting it as though it were also a summary of segment performance.
And do not let the standardized comparison drift into being the honest one. It is an honest computation built on a stated assumption about composition. That is a real thing to show, but it is a conditional statement, and conditionals belong on the slide with their condition attached.
What the finished pair looks like
Left slide, unchanged from the original draft: 27% → 73%, one number per period, large type. It is correct and it is the operating result.
Right slide: the counts-and-rates panel, with segment rows, total rows, and shares for both periods. Beneath it, the two labeled summary lines:
- Operating result (observed): 27% → 73%
- Standardized to period-1 mix (10% A / 90% B): 27% → 17%
Then the presenter's sentence, which is where the article's real answer lives: "Our overall rate went up because the mix of observations shifted heavily toward the segment with the higher rate. Inside both segments, performance fell by ten points. The first line is what happened. The second line is what the same mix of customers would have produced. Both are true, and they lead to different follow-up questions."
That last clause matters more than the numbers. A reversal does not tell you which figure to act on — that is a business judgment about whether the mix shift is a problem, a strategy, or an artifact. What the deck can do is make that judgment possible by putting both calculations in view, with counts underneath the observed ones and the weights written next to the constructed one. The conclusion worth reaching is not which number is better. It is that the room now knows exactly what each number answers and which of those questions they are deciding.
Frequently asked questions
How can the overall rate improve while every customer segment gets worse?
An overall rate is a weighted average of segment rates, and the weights are each segment's share of observations. If the mix shifts toward the higher-performing segment, the average can rise even while every segment's own rate falls. In the constructed example, Segment A falls from 90% to 80% and Segment B from 20% to 10%, yet the total rises from 27% to 73% because A's share goes from 10% to 90%.
What must be checked before treating the reversal as a legal comparison?
Four definitions must hold across both periods: one eligible observation, one successful outcome, the segment boundaries, and the observation window. If “completed” meant different things in the two periods, that is not a mix problem but two metrics wearing the same word. Also ask whether anything in the product, funnel, or segmentation rule changed how observations get assigned, because a redefined field can make segment rates incomparable.
Why is averaging the segment rates the wrong calculation?
Twenty-seven percent is not the average of 90% and 20%; that average is 55% and corresponds to no population. The correct total comes from weighting: 0.10 × 90% + 0.90 × 20% = 27%. The unweighted mean can fall in the direction that feels right, which makes it harder to catch in review. Never average subgroup percentages when the denominators differ.
How should the deck show both the overall rate and the segment rates?
Keep the total-only slide, then add a counts-and-rates panel with a row per segment and a total row, showing eligible count, success count, rate, and share of observations for each period. Under it, state the change in one line: overall rate, both segments' movement, and the share shift. Say percentage points rather than relative increase when that is what you mean. If adding a standardized row, name the weights, keep the observed total in the same view, avoid putting it in the headline, and announce that it answers a different question.
What can the arithmetic not say about the reversal?
It cannot say the mix shift caused the segment declines, or that the segment declines caused the mix shift. The decomposition splits a total change by construction; it does not order events or assign responsibility. It also does not support significance language without a stated test and assumptions. The aggregate is not misleading — it is an accurate summary of the program as it ran — and a standardized comparison is a conditional statement, not automatically the honest one.