Skip to content

Is Retention Improving—or Are You Comparing Cohorts at Different Ages?

Business

Is Retention Improving—or Are You Comparing Cohorts at Different Ages?

A retention slide can be built entirely from accurate numbers and still be wrong. The usual version looks like this: last month's cohort on the left, a cohort from the spring on the right, arrows pointing up. Both figures came from a real dashboard. Both were computed correctly. The comparison between them is nevertheless a comparison of two different questions.

The repair fits in one sentence. Compare cohorts at the same elapsed age, under the same entry and return rules, and count only the accounts that have actually had the opportunity to do the thing your return metric asks about. A cohort that has not yet reached an interval has no completed observation for that cell. It does not have a zero.

What follows is the working-out: four decisions you make before any number reaches the slide, three arithmetic traps, and what a founder can honestly present while one cohort is still maturing.

Define entry, return, unit, and the observation clock

Four definitional choices determine everything downstream, and they need to be settled before you look at a chart, because a chart will happily impose its own.

Entry. The first qualifying event for the unit: account creation, first completed session, first invoice paid. Pick one, name it, and apply it to every cohort. If entry dates can be backfilled later—a signup imported from a CRM two weeks after the fact—the entry instant is unstable, and every elapsed age computed from it moves later when the record changes.

Return. The later action that counts. Opening the app? A session that includes a specific event? A payment? This is where metric definitions earn their keep, because "any session" and "session with a core action" produce different curves from the same traffic. Write down which.

Unit. Account or person. Ten accounts may be seven people, and if three of them belong to one power user, the percentage describes accounts, not adoption. Whichever you choose, the denominator is units, not events.

The clock. Say plainly what Day X means. A clean and common convention is the half-open window from X days to X+1 days after the entry instant, with at most one counted return per unit, so the denominator stays the entry count instead of drifting as heavy users log more sessions. Then name the observation cutoff—the as-of timestamp for the export—and treat a window as complete only when the cutoff falls at or past its end.

One distinction is easy to lose and expensive to lose: return within an interval and return on or after that age are different questions, not two settings for the same chart. Under the first rule, the value for a late interval can fall below an early one, and that is ordinary—fewer people come back on day 28 than on day 7. Under the second, the same cohort's value accumulates and cannot decline as the age grows. Amplitude's documentation for its Retention Analysis chart distinguishes these return-event rules and describes how a cohort that has not yet reached an interval affects its denominators (checked 18 September 2026). The definitions support a like-for-like comparison; they do not establish growth or its cause. If one chart in a deck uses a per-interval rule and another uses a cumulative one, a rising line between them is a rule change wearing a trend's clothing.

A smaller decision belongs here too, because it changes who can contribute later. When an account is deleted, merged, or otherwise unobservable, does it leave the denominator, or does it stay and count as a non-return? Both are defensible. They produce different late-age curves. State which one you used.

Align the rows by elapsed age

With the definitions fixed, the comparison rule is short: Day 7 against Day 7, Day 28 against Day 28, and nothing in the deck that silently spans both.

A worked fixture makes the stakes concrete. Three invented cohorts, ten accounts each, with every account in a cohort entering at a single shared instant within that cohort. Cohorts A and C have 35 completed days of observation. Cohort B has 14. Day X is the window from X to X+1 days after entry, with at most one counted return per account. These cohorts, accounts, and counts are fictional; they exist to make the arithmetic visible, and they establish nothing about any real product.

Cohort Day 7 returns Day 28 returns
A 6/10 = 60% 3/10 = 30%
B 6/10 = 60% Incomplete: window not reached
C 4/10 = 40% 0/10 = 0%

The tempting sentence is that B is outperforming A: 60% against 30%. That sentence changes two things at once. It changes which age is being measured, and it changes which population is being described. At the age both cohorts have actually reached, A and B are identical—60% and 60%. The Day 28 column measures a different interval, and B has no value there at all.

Note also that a per-interval rule makes the 60%-to-30% movement inside cohort A unremarkable. Under this definition, Day 28 is not "retention after four weeks"; it is the share of cohort A that returned specifically during the window from day 28 to day 29. That share can be lower than the Day 7 share without anything having deteriorated. Under a cumulative rule, cohort A's numbers would have to climb or hold. Same accounts, same behavior, different question.

The second alignment problem is quieter. Calendar months are not cohorts. A "January cohort" contains accounts that entered on 1 January and accounts that entered on 31 January; at a fixed as-of date, the first has had roughly thirty more days of opportunity than the last. Labeling a cell "Day 7" for that group averages over accounts whose seventh day fell on different dates, some of them weekends, some of them during a pricing change. Two options are available. Keep the absolute entry date in the record so maturity is checkable line by line, or truncate every member to a common elapsed age before computing the cell. What you cannot do is assume the monthly label gives everyone the same opportunity to return.

The fixture's shared entry instants are a simplifying fiction—an unusually clean one, and the reason A, B, and C can be compared at all.

Keep an immature cell separate from zero

Cohort B's Day 28 cell is not empty because B did badly. It is empty because the window from day 28 to day 29 has not closed. With 14 completed days of observation, B would need observation through its 29th completed day—fifteen more days—before that cell has a value of any kind.

Three states need three labels, and a dash in a spreadsheet is not enough to tell them apart.

  • Observed. The window closed and the defined return event was captured. This includes zero. Cohort C's 0/10 is a measured zero: the window closed, and among ten accounts, no qualifying return was recorded in it.
  • Incomplete. The window is still open. No value exists yet, and no value can be inferred. This is cohort B at Day 28.
  • Unknown. The window closed, but the measurement is unreliable—the return event was renamed mid-window, instrumentation lapsed, or the export has a gap. No value exists here either, and it is a different problem with a different remedy.

Filling an incomplete cell with zero is the single most common way a retention table starts lying. It converts "we don't know yet" into "nobody came back," and it does so in the exact direction that makes newer cohorts look weak. The reverse mistake is rarer but also real: treating a genuine zero as missing, and quietly dropping cohort C's mature 0% from the plot so the curve looks smoother than the data.

A measured zero carries its own caveats, which are worth one sentence on the slide rather than a footnote nobody reads. A 0/10 depends on the return event having actually fired during that window. If the event was renamed on day 26, the zero is an artifact of instrumentation, not of behavior. And three accounts separate 0% from 30% here. That difference is worth investigating. It is not a trend.

Compare aggregates only with their contributing populations visible

Aggregates are where age alignment quietly breaks. In the fixture, all three cohorts are mature at Day 7, so a combined figure has a clean denominator: 6 + 6 + 4 returns across 30 accounts, or 16/30, which is 53.3%.

At Day 28, only A and C are mature. The correct aggregate is 3/20—three returned accounts across the twenty accounts that had the opportunity—which is 15%. If B's ten accounts stay in the denominator while its cell is dropped as incomplete, the same returns are reported as 3/30, or 10%. Both figures describe the same three accounts. Five percentage points of the difference is nothing but the size of a population that should not have been there.

Then the more consequential point. The Day 7 aggregate covers 30 accounts; the Day 28 aggregate covers 20. Putting 53.3% and 15% on one line implies a population that shrank by a third while it was being measured. It did not. The population at the later age is a different and smaller set, and the line between those two points is partly a description of who was eligible, not of what anyone did.

The clean alternative is to show a curve whose endpoints come from the same accounts. Cohorts A and C together: 10/20 at Day 7, which is 50%, and 3/20 at Day 28, which is 15%. Every point on that line comes from the same twenty accounts. It excludes B entirely—which means B's 60% at Day 7, identical to A's, is not represented in it. That is not a flaw to hide. It is the definition of the population the curve describes, and it belongs in the label.

Two further habits follow.

First, know which cohorts dominate the long ages. At Day 28 in a three-cohort table, the aggregate is essentially A and C. In a real deck with twenty cohorts, the Day 180 figure may be dominated by two or three older groups acquired through a channel that no longer exists. A line that looks longitudinal is often cross-sectional: different people, different moment, presented as one history.

Second, present the older cohorts' longer histories separately, labeled as single cohorts, rather than blending them into the comparison. Cohort A's full 35 days are genuinely useful. They belong in their own chart, captioned as one cohort over time, not stacked into an all-users curve whose membership changes at every age.

And keep the honest limits attached to whatever survives. Cohort A's 30% and cohort C's 0% at Day 28 are a like-for-like comparison: same age, same rule, both mature. In the fixture, three accounts separate them. That is a difference to examine, not an established improvement, a channel effect, or a statistically reliable gap. A correct denominator establishes that a comparison is fair; it says nothing about why the numbers differ or whether the difference would recur.

One habit worth carrying over from the definitions: check the arithmetic you inherit. Vendor documentation that draws the interval distinction correctly can still print a worked division that does not check out—the Return On All Users paragraph records 48,219 divided by 125,665 as 72%, while the division gives approximately 38.37% (checked 18 September 2026). A wrong number inside a correct definition is easy to copy and hard to spot on a slide, because it sits where a real figure should be. Recompute anything you did not compute yourself.

What the slide can say now

The age-aligned comparison the fixture supports is narrow, and it is presentable.

At Day 7, all three cohorts are mature, and the honest figures are 60%, 60%, and 40% by cohort, or 53.3% for the thirty accounts together. At Day 28, only two cohorts can speak: A at 30%, C at 0%, and a combined 15% across the twenty accounts that had the opportunity. B contributes nothing to that age—not a zero, not a gap to be smoothed over, but an interval still awaiting observation, fifteen days short of the window closing.

Label the cells accordingly, in the cell itself rather than a legend: "incomplete—window open" is a sentence a reader can act on; a blank invites someone to type a zero into it. Show the denominator beside every percentage, because a changing denominator is the most persuasive false improvement a retention chart can offer. And keep B's one completed point where it belongs: alongside A and C at Day 7, where it is a real observation, rather than extrapolated forward into a Day 28 figure it has not earned.

The curve will eventually fill in. What it will not do on its own is tell you whether the people who arrived later are better retained than the people who arrived earlier, because that question requires the same age, the same rules, and the same eligible population on both sides. A smooth line is a property of the chart. It is not evidence that anything improved.

Frequently asked questions

Why can a retention comparison be wrong even when every number is accurate?

The comparison may measure different ages or populations. In the invented fixture, cohort B's Day 7 result is 60% and cohort A's Day 28 result is 30%; saying B outperforms A changes both the interval and who is being described. At the age both have reached, A and B are both 60% at Day 7.

What is an incomplete retention cell, and how should it be labeled?

It means the observation window has not closed, so no value exists yet and none can be inferred. It is not zero. Label it in the cell as incomplete or window open. A zero is an observed value where the window closed; unknown means the window closed but the measurement is unreliable.

How should aggregates treat cohorts that have not reached an interval?

Count only units that had the opportunity. In the fixture, Day 28 is mature only for A and C, so the correct aggregate is 3/20, or 15%. If B's ten accounts stay in the denominator while its incomplete cell is dropped, the same returns become 3/30, or 10%. Show the denominator beside every percentage.

What is the difference between return within an interval and return on or after that age?

Within an interval asks whether a unit returned during the window from X to X+1 days. On or after an age is cumulative and cannot decline as age grows. Mixing the two rules can make a rule change look like a trend. The definitions support like-for-like comparison; they do not establish growth or its cause.

What limits remain even after ages and denominators are aligned?

A fair comparison does not explain why numbers differ or whether a difference would recur. In the fixture, three accounts separate cohort A's 30% and cohort C's 0% at Day 28, so that is a difference to examine, not an established improvement or channel effect. Also recompute inherited arithmetic: the article notes a vendor example printing 48,219 divided by 125,665 as 72%, while the division gives about 38.37%.

More in Business Browse all articles