Two Valid Metrics Tell Opposite Stories. What Should the Deck Lead With?
Two Valid Metrics Tell Opposite Stories. What Should the Deck Lead With?
The headline is accurate. That is what makes it expensive.
"Active users doubled" can be a true statement about a period in which the average active user did substantially less. Someone who reads only the headline leaves with a different picture of the business than someone who reads only the second figure, and both readers were handed correct numbers from the same table. They will walk into the next meeting prepared to make different decisions.
This is the situation worth writing about: not a bad number, not a broken definition, not a chart that needs better colors. Two measures, each doing its job, pointing in opposite directions, and a slide that has to lead with one of them.
The reflex is to lead with the flattering figure and footnote the other. A slightly more refined reflex is to lead with the figure you believe and treat the other as a caveat. Both are forms of editing, and both settle something for the room that the room has not settled itself.
The argument here is narrower than "always show everything." It is that a divergence between two valid measures is often the most decision-relevant content in the deck, and it is the content that a deck's structure is most likely to erase.
First, check that the disagreement is real
Everything below assumes the two measures mean what the presentation says they mean. If they don't, you have a consistency problem, not a tension. The neighbor article in this series covers that case: sales, product and finance using one metric name for three different quantities. That gets repaired before anything here applies.
So the first pass is mechanical. For each measure, establish the unit, the qualifying event, the population, the period and the denominator. Confirm the two periods are equal in length. Confirm both use the same eligibility rules. If one number counts distinct users completing at least one task, and the other averages over "engaged accounts" without saying which accounts count, you do not yet have two stories. You have one story and one rumor.
This check also decides who owns the answer. Definitions belong to whoever maintains the measurement, not to whoever is building the slide. If the two figures were assembled from different pipelines by different teams, ask each owner to confirm the definition in writing before you interpret a divergence. The cost is an email. The alternative is a strategy conversation built on a definition drift.
There is a public route for the disclosure side of this, if the deck also has to explain metrics outside the company: the SEC maintains a landing page for its guidance on management's discussion and analysis, which is where the question of how management explains a measure to outsiders lives. That page's title and description are the extent of what informed this piece. The linked guidance was not read here, nothing below depends on it, and it says nothing about how an internal deck should be built.
Both figures can be true at once, which is the whole problem
Take an invented case. Two equal-length periods. Active user means a distinct user who completed at least one task.
- Period A: 100 active users, 400 tasks.
- Period B: 200 active users, 500 tasks.
These numbers are arithmetic inputs, not company performance, and no cause is attached to them. Period B took in 25% more task volume than period A while its user count doubled. Average tasks per active user fell from 4.0 to 2.5, a decline of 37.5%.
Nothing here contradicts anything else, because the two measures are not independent witnesses. They are the same activity summarized two ways. The identity is simple:
tasks = active users × tasks per active user
A: 400 = 100 × 4.0
B: 500 = 200 × 2.5
The user count multiplied by 2.0. The rate multiplied by 0.625. Their product is 1.25, which is exactly the rise in total tasks. One factor grew faster than the other fell. That is the entire divergence.
This matters because a slide reading "growth up, engagement down" implies two signals arriving from two directions. In fact the second measure has the first measure inside its definition. The moment you change what counts as an active user, you change the denominator of tasks per active user. So the two numbers cannot be treated as separate evidence that happen to agree or disagree. They are one structure with a denominator.
A useful test when you're staring at a pair like this: does either figure sit inside the definition of the other? If yes, you are not reconciling two findings. You are reading one finding at two magnifications, and the only question is which magnification the decision needs.
What the aggregate cannot tell you
Here is the part that catches people.
Period A had 100 distinct active users. Period B had 200. Whatever else is true, at least 100 of period B's users were not active in period A, because period A only had 100 people in it. At minimum, half of the population behind that 2.5 average was absent from the comparison period. And "not active in period A" is not the same claim as "new to the product" — it means new to this window, which is a weaker and more honest statement.
That fact alone should stop the sentence "users became less engaged." It is not a claim the data supports, because the population changed.
Watch how much the aggregates leave open. Here are two different worlds that produce identical numbers:
- World one. The original 100 users held steady at 4.0 tasks each, contributing 400. Another 100 users arrived and completed one task each, contributing 100. Total 500, over 200 users, mean 2.5.
- World two. Two hundred users averaged 2.5 tasks, with no one holding steady at 4.0 and no group performing exactly once.
Same totals. Same mean. Same headline options. Completely different businesses. The presentation cannot distinguish them, and no amount of care in writing the headline will change that, because the distinction does not live in the aggregates.
Three families of explanation are consistent with the pattern, and they call for different responses:
Composition. The added users behave differently from the ones already there. Under this explanation, nothing about the original experience has changed.
Behavior. The people present in both periods are each doing less. Under this explanation, something about the product, the market or the task itself has changed for the core population.
Definition drift. What counts as a task, or as an active user, moved between periods. Under this explanation, the two figures may not be comparable at all, and you are back at the first section.
These are not mutually exclusive, and none of them is established by the two numbers. What separates them is user-level records: which of period B's users also appear in period A; the distribution of tasks per user rather than just its average; the share completing exactly one task; whether the task catalog changed; whether the eligible population — the people who could have been active — moved. Those records narrow the possibilities. They will not, on their own, establish a cause.
Choosing a headline is choosing what the room will argue about
Three candidate headlines for the invented case, all defensible:
Reach only. "Active users doubled."
Depth only. "The average active user completed 37.5% fewer tasks."
Both. "Active users doubled to 200, while tasks per active user fell from 4.0 to 2.5."
The first is true, and it steers the room toward acquisition, cost per user and capacity. The second is also true, and it steers the room toward retention, onboarding and the experience of the people already there. The third does not choose. It hands the room a fact with two handles and asks which one the decision is about.
The test worth applying is not "is this claim accurate." It is: if a reader saw only the other figure, would their sense of the business reverse? If yes, the second figure is not a caveat. It is part of the headline's job, because leaving it out is a selection, and selections change decisions.
That does not automatically mean a compound headline. Two-figure headlines get long, and long headlines get skimmed. The practical requirement is visibility: the counterpoint has to be on the same screen, in a size the reader cannot miss without trying. A footnote fails this test specifically because it requires the reader to already suspect there is a problem. Nobody reads footnotes on the way to a decision.
Two qualifications, both important.
Not every divergence earns the headline. If a measure drifts by an amount smaller than its ordinary period-to-period movement, promoting it to the headline is noise-taking. With 200 users and 500 tasks, small absolute changes move the mean a lot, and whether a 1.5-task swing exceeds normal variation is a question for whoever owns the measurement — ideally with a longer history than two periods. This article is not going to certify a threshold, because the threshold depends on the metric's behavior over time. The point is that "preserve the tension" is not "always lead with both." It is "lead with what changes the decision."
Prominence is the editorial variable, not inclusion. You almost never have to choose between showing a number and hiding it. You choose how loud it is. That choice is where the honest deck and the flattering deck separate, and it separates quietly, because both versions contain every figure.
Two routes, and the evidence that would separate them
Now the part the deck can actually help with.
Suppose the company faces a real choice. One route pushes reach: keep acquiring, accept that the marginal user may be a lighter user, and plan capacity around totals. The other protects depth: investigate why per-user intensity fell, hold acquisition steady, and invest in whatever the core population is doing.
State each as a route with a cost, not as a recommendation with a number attached.
The reach route assumes that the business constraint is adoption, and that per-user intensity is an interesting variable rather than a commitment. Its cost: if the decline reflects a genuine change for the people who were already there, more acquisition compounds a problem that the deck has just authorized spending against. Its supporting evidence would be cohort records showing the decline concentrated among users with no prior activity, a stable distribution among returning users, and a definition audit showing nothing changed in what counts as a task.
The depth route assumes that per-user intensity is where the product's value lives, and that a 37.5% decline is a warning. Its cost: if the decline is composition — a larger population of users who each did one thing — this route spends on a non-problem and slows a channel that is working. Its supporting evidence would be returning-user rates also falling, a distribution shifted left for users present in both periods, and the same cohort still low in a third period.
Notice that neither route is selected by the two numbers. The numbers constrain what you can claim. They do not rank the objectives. Whether reach or depth matters more is a statement about what the product is for, and that is a decision someone has to make out loud, not an inference available in a table.
Each route also carries a testable trade-off, and stating it as a hypothesis is the most useful thing the deck can do. If we push reach, we expect tasks per active user to fall further, because the additional users arrive at lower intensity. If we protect depth, we expect the user count to grow more slowly. The first half of that is an assumption about your product, not a law, and the deck should say so. It becomes testable the moment you track the next period's population separately and report the returning users' rate alongside the overall rate. Then the argument in three months is about evidence rather than about whose reading of one number was right.
What goes on the slide
The assembled answer for the invented case looks like this. A headline that names both movements. A line that states the decision the slide serves. And a closing line that names the evidence which would settle the question.
Something close to:
Active users doubled to 200. Tasks per active user fell from 4.0 to 2.5. We are deciding whether to fund the next acquisition push or investigate the decline in per-user intensity first. What we need: cohort overlap between the two periods, the distribution of tasks per user, and confirmation that the definition of a task did not change.
That is not a tidier slide than "Active users doubled." It is a harder one. It refuses to let two accurate numbers become one convenient story, and it puts the unresolved choice where the room can see it rather than where the presenter can manage it.
The temptation at this point is to end with a recommendation. Resist it. The deck can make a choice legible, bound it with evidence and name what would settle it. It cannot decide the business's priorities from a pair of figures, and pretending otherwise is the same error as burying the second number, just pointed in a more sophisticated direction.
The divergence is not in your way. It is the decision, showing up early.
Frequently asked questions
What is the first check when two metrics tell opposite stories?
Confirm the two measures mean what the presentation says they mean. For each measure, establish the unit, the qualifying event, the population, the period, and the denominator. Confirm the two periods are equal in length and that both use the same eligibility rules. If not, you have a consistency problem to repair before treating the divergence as a real tension.
Why are active users and tasks per active user not independent evidence?
Tasks per active user has active users inside its denominator. The identity is tasks = active users times tasks per active user. In the invented example, 400 = 100 times 4.0 and 500 = 200 times 2.5. One factor grew faster than the other fell; the two figures are one structure read at two magnifications.
What can two aggregate figures not distinguish?
They cannot distinguish composition, behavior, or definition drift. Two different worlds can produce the same 200 users, 500 tasks, and 2.5 mean. Separating those explanations requires user-level records: which period B users also appear in period A, the distribution of tasks per user, the share completing exactly one task, whether the task catalog changed, and whether the eligible population moved.
When should the counterpoint move into the headline?
If a reader seeing only the other figure would have their sense of the business reverse, the second figure is part of the headline's job. Still, not every divergence earns the headline. If a measure drifts by less than its ordinary period-to-period movement, promoting it is noise-taking. Prominence is the editorial variable, and a footnote usually fails because nobody reads it on the way to a decision.
Do the two numbers choose between a reach route and a depth route?
No. The numbers constrain what can be claimed, but they do not rank the objectives. State each route with its cost and supporting evidence. Track the next period's population separately and report the returning users' rate alongside the overall rate, so the argument becomes about evidence rather than whose reading of one number was right.