Retention: the definitive guide
How to measure retention honestly and actually move it — cohort curves and their artifacts, users vs. dollars, the instruments that fail replication, churn mechanics, billing architecture, and the AI churn wave. Every number sourced.
Retention is the share of a cohort still active — or still paying — a fixed interval after they started, measured separately in people and in dollars; the only trustworthy form of it is the cohorted curve and whether it flattens. There is no single good number: six-month user retention that counts as "good" runs about 25% for consumer social and about 70% for enterprise SaaS, and median B2B net revenue retention is 102% by survey or 82% by instrumented billing data depending on who measures. Before comparing any two figures, check the counting rule — N-day, unbounded and rolling retention are three different mathematical objects, so numbers from different tools are not comparable.
- Only cohorted curves tell the truth. Aggregate retention can drift from ~45% to ~70% with zero improvement in any cohort — Fader & Hardie's "ruse of heterogeneity."
- N-day, unbounded and rolling retention are three different mathematical objects. A benchmark quoted without its counting rule is not comparable to your number.
- The two biggest sources of median B2B NRR disagree by twenty points — 102% in SaaS Capital's survey, 82% in ChartMogul's billing data. Present both and name the gap; never average them.
- 22% of SaaS churn (≈35% across Recurly's network) is failed payments, not product — and ~70% of it is recoverable with retries, an account updater and dunning.
- NPS and CES both failed independent replication; in de Haan et al., CES correlated negatively with two-year retention (r = −.073) while plain satisfaction predicted it best.
Last reviewed 2026-08-22
Retention is where growth stops being a story you tell and becomes a verdict you receive. Andrew Chen, from staring at thousands of these curves: "a product stalls when its churn catches up with its customer acquisition" — and as the base grows, absolute churn grows with it, so a fixed acquisition rate mathematically stops producing net growth. The bucket metaphor everyone reaches for has real data behind its sharpest edge: across Amplitude's 2,600+ company dataset, "there's no meaningful correlation between how well you acquire users and how well you keep them." Acquisition skill does not transfer. Retention is its own discipline.
But before a single benchmark, the caveat that governs this entire guide: the same cohort produces materially different retention numbers depending on the counting rule. N-day retention (returned on day N), unbounded retention (returned on or after day N), and rolling retention are three different mathematical objects, and the vendors' own docs flag the artifacts — Amplitude's FAQ explains that a rising retention curve is usually a measurement artifact of incomplete cohort windows, not improving retention. N-day punishes weekly-frequency products; unbounded flatters everything because one later return counts forever. Cross-tool benchmarks are not comparable unless counting rules match, and one experienced practitioner's estimate is that "it takes up to six months to nail accurate retention reporting." Every number below is "retention as that vendor's customers happen to define it."
The measurement layer: curves, accounting, and two traps
The flattening test. The canonical instrument is the cohort retention curve, and the canonical question is whether it flattens. Tribe Capital's phrasing: "The focus on this graph is whether logo retention has reached an asymptote at long duration." Casey Winters' operational version: "A flattened retention curve of your key action at the designated frequency plus month over month growth in new customers is the best way I have found to measure true product/market fit." Chen adds the empirical rhythm of decay — "Whatever the D1, it drops by 50% on D7. Whatever the D7, it drops by another 50% at D30" — and the uncomfortable corollary: "You can't fix bad retention. No, adding more notifications will not fix your retention curve. You can't A/B test your way to good retention." The curve is largely a property of the market-product pair, set at launch.
Growth accounting. Jonathan Hsu's Social Capital framework decomposes any period's actives into additive identities — MAU(t) = new + retained + resurrected; last period's MAU = retained + churned — with the Quick Ratio, (new + resurrected) / churned, as the summary statistic. It was proposed explicitly as "GAAP for startups," because without standard definitions your retention numbers aren't comparable to anything, including your own past. Its limits are structural: the identities are definitional, so output depends entirely on the period length and the "active" definition, and the user-level version cannot represent expansion or contraction at all — which is why revenue got its own buckets.
Trap one: survivor bias. Aggregate (non-cohorted) retention drifts upward mechanically as the base becomes dominated by long-tenured survivors — one walkthrough shows aggregate retention "improving" from ~45% to ~70% with zero improvement in any cohort. This is the operational form of what Fader & Hardie call the "ruse of heterogeneity": "the phenomenon of retention rates increasing over time is simply due to heterogeneity (i.e., the high-churn customers drop out early)." Rising per-tenure retention rates are not customers becoming loyal; they are the disloyal ones already being gone. Only cohorted curves tell the truth.
Trap two: the blended churn number. Churn is strongly tenure-dependent — one operator reports ~25% churn in month one dropping to almost nothing after a few months, making any single quoted churn rate "fairly useless." A blended average mixes a fast-decaying new population with a slow-decaying survivor population and describes neither.
Sources: Amplitude retention docs ("Return On" vs "Return On or After") · Mixpanel retention docs (rolling intervals).
The base becomes dominated by long-tenured survivors, so the average rises mechanically — Fader & Hardie's "ruse of heterogeneity." Values stylized after the Retention-Led Growth walkthrough (~45% → ~70% with zero cohort improvement).
Users and dollars are different currencies — and they diverge
Retention must be measured twice: in people and in revenue. They routinely point in opposite directions, and Chen states the divergence as a rule: "Revenue retention expands, while usage retention shrinks." Expansion from surviving accounts papers over a decaying user base — the pattern known as negative churn, Skok's "ultimate solution to the churn problem", and simultaneously the mechanism of the NRR mirage.
The definitions: gross revenue retention (GRR) counts only churn and downgrades; net revenue retention (NRR) adds expansion back. Bessemer's benchmark framing — top cloud companies run annual logo churn below 7% and revenue churn below 5% — comes with their own counterexample: "we have seen successful cloud businesses that address SMB segments — such as HubSpot (in its 3 years pre-IPO) — with 70-80% gross retention." And NRR is partly a pricing-model property, not a product property: usage-based companies average roughly ten points higher NRR than traditional-pricing peers, so cross-model NRR comparison is invalid. SaaS Capital's advice on segmentation: "For retention, benchmarking by ACV is the best starting point. More than by company age, revenue level, or industry."
Now the finding your dashboard vendor won't volunteer: the two big sources of "median NRR" disagree by twenty points, and both are right. SaaS Capital's survey (self-reported, private B2B SaaS over $1M ARR) puts median NRR at 102%. ChartMogul's instrumented billing data (~2,700 B2B companies over $250K ARR) puts it at 82%. The gap is selection and self-report bias in the survey population plus different denominators — and the honest move is to present both and name the gap, never to average them. What NRR buys you is uncontested though: SaaS Capital's growth-by-NRR banding shows median growth of 15% below 90% NRR, 20% at 100–110%, 30% at 110–120%, and 44% above 130%. Retention compounds into growth rate — Bessemer's line is that "every percentage taken out of your retention is taken out of your growth rate."
Sources: SaaS Capital retention benchmarks (survey, 2025) · ChartMogul FY2025 (billing data, ~2,700 B2B cos) · SaaS Capital growth-rate brief (2025).
The numbers behind the figure
| Measure | Value | Method |
|---|---|---|
| Median B2B NRR | 102% (97–111 quartiles) | SaaS Capital survey |
| Median B2B NRR | 82% (97 upper quartile) | ChartMogul billing data |
| Growth at <90 / 100–110 / 110–120 / >130 NRR | 15% / 20% / 30% / 44% | SaaS Capital |
What "good" looks like — always relative to model and frequency
The only published benchmark set that respects business-model differences is Lenny Rachitsky and Casey Winters' expert-elicitation study — twenty experienced growth practitioners, six-month user retention by model. Carry its own caveats with it: the numbers are elicited opinion from operators of iconic businesses (explicit selection bias), undefined as to counting rule, and the authors say plainly that these bars describe venture-scale ambition — "this level of retention is not required for product-market fit or to build a sustainable business."
The spread between models is larger than good-to-great within one — a benchmark without a model label is noise. Source: Lenny's Newsletter × Casey Winters; elicited expert opinion, counting rule undefined, venture-scale bar.
The numbers behind the figure
| Model | Good | Great |
|---|---|---|
| Consumer social | ~25% | ~45% |
| Consumer transactional | ~30% | ~50% |
| Consumer SaaS | ~40% | ~70% |
| SMB / mid-market SaaS | ~60% | ~80% |
| Enterprise SaaS | ~70% | ~90% |
| 12-mo NRR: bottom-up / enterprise | ~100% / ~110% | ~120% / ~130% |
Two more calibration rules. First, natural frequency: Winters — "you may only look for a place to live once every few years… but you look for something to eat multiple times a day." Measure retention at the product's inherent cadence or you measure nothing. Elena Verna's habitual definitions put teeth on "active": WAU should mean "active 3 of the last 4 weeks", not "touched it once." For yearly-frequency products — Zillow, TurboTax, Thumbtack in Reforge's Use Case Frequency Spectrum — the retention-curve PMF test is simply the wrong instrument; these live in the "Forgettable Zone" and need a different strategy entirely. Judge an episodic product on DAU/MAU and you'll declare a working business dead.
Second, distributions beat averages. DAU/MAU proxies habit ("apps over 20% are said to be good, and 50%+ is world class") but Chen's own critique is canonical — "I've also not seen a 10% DAU/MAU product, through sheer effort, become 40% DAU/MAU," and pushing notifications to juice it "will actually lower your DAU/MAU." The Power User Curve — the full histogram of days-active-per-month — replaces the blended ratio with the shape: it "will 'smile' when things are good," with a16z's own caveat that not every company needs the smile and lower-frequency SaaS should use an L7 variant.
The instruments that don't survive scrutiny
A definitive guide owes you the demolition section, because several standard retention instruments have published refutations that practice ignores.
NPS. Reichheld's claim that likelihood-to-recommend is the "single most reliable indicator of a company's ability to grow" has failed replication in the peer-reviewed literature — Keiningham et al., Journal of Marketing 2007: "the research fails to replicate his assertions regarding the 'clear superiority' of Net Promoter… We find no support for the claim." A second refutation is blunter still: NPS "possesses few, if any, of the characteristics that might be regarded as highly desirable in a high-level market research metric; on the contrary, it has done considerable damage." The method problem: the original analysis correlated NPS with historical growth. At least one organization has publicly abandoned it on exactly these grounds, switching to the plain mean of the likelihood-to-recommend item because the research "failed to replicate, even when the replication was using the originally published data set (!)."
Customer Effort Score. The "Effortless Experience" claim that CES beats NPS and satisfaction reversed under independent replication: de Haan et al., 6,649 respondents across 93 firms, found satisfaction the best predictor of two-year retention (r = .184), NPS close behind (r = .170) — and CES negatively correlated (r = −.073). The original research also covered only contact-center interactions, which most customers never have.
Health scores. Composite scores with hand-set weights are lagging indicators wearing leading-indicator costumes. The critiques worth keeping come from the field's own architects: Skok — "Usage does not necessarily equate to deriving business benefits" — and Lincoln Murphy's sharper cut, that customers achieving the required outcome without an appropriate experience "are a churn threat, even though you may see them as 'healthy'." An operator corollary from the field: one founder killed a feature ~40% of accounts had touched — it drove half of support inquiries — and churn fell (self-reported, no controls; usage ≠ value, inverted into a decision rule).
The leaky-bucket doctrine itself. The "retention is cheaper than acquisition" family of claims deserves its asterisk: the famous "5% retention → 25–95% profit" range traces to a Bain brief that states 25%, for financial services only — no fetched source supports the 95% bound — and the "5× cheaper to retain" line is treated as a myth by Ipsos's own loyalty research. The full refutation is Byron Sharp's: "brand growth and profits depend largely on a company's ability to outperform similar-sized rivals in acquisition, while all brands suffer from close to expected rates of defection" — with Ehrenberg's 1974 observation that lapsed buyers "are merely relatively infrequent buyers." Sharp's context is repertoire consumer markets, not subscription software — but the discipline transfers: retention advantage must be demonstrated in your data, not assumed from folklore.
Churn mechanics: most of it isn't what your dashboard says
A quarter to a third of your churn is a payments problem, not a product problem. Across Stripe + Churnkey's 5.4M failed payments over 25M subscriptions, involuntary churn is 22% of all SaaS churn; Recurly's network split implies ~35%. It is also the most fixable churn there is: ~70% of detected involuntary churn is recoverable, dunning campaigns average 42% recovery, 90% of successful recoveries happen within 10 days — and the customers you save are real: 49.6% of recovered subscribers' total lifetime happens after the recovery event. Every framework that reads churn as a verdict on the product — curves, health scores, PMF surveys — is contaminated to the degree this bucket goes unmanaged.
The cheapest retention program in existence is retries + card updater + dunning — and unmanaged, this bucket falsifies every product-quality read of churn. Sources: Churnkey/Stripe · Recurly.
The numbers behind the figure
| Measure | Value | Source |
|---|---|---|
| Involuntary share of SaaS churn | 22% | Stripe + Churnkey |
| Involuntary share, all industries | ~35% | Recurly network |
| Recoverable share / dunning recovery | ~70% / 42% | Churnkey · Recurly |
| Lifetime after recovery event | 49.6% | Recurly |
The voluntary churn that is announced is not the churn that kills you. Stated cancellation reasons across ~3M cancellation sessions: budget 33%, infrequent usage 30.6% — and save offers work unevenly (discount accepted 53.9%, pause 19.2%, plan change 6.7%). But Skok's two top causes of churn are quieter: failure to onboard (the churn was created in week one and surfaces at renewal — a misdiagnosed activation failure) and champion churn — the value and political cover lived in one person, and the renewal died with their departure, at perfectly healthy usage. The operators asking "how do you catch churn before it happens" have noticed the same thing: the real early signals are relational — the champion goes quiet, response times stretch, an unknown stakeholder appears at the meeting — and no usage dashboard measures any of them. This is one of the largest unsolved instrumentation gaps in retention.
Save tactics have a compounding cost. The panic discount addresses price when the problem was value, and it teaches: one founder's 40%-off save bought two months, the customer cancelled anyway, and three other customers then asked for the same discount (self-reported; the pattern, not the number, is the lesson).
And cancel-flow friction is now a legal exposure, not a tactic. The FTC's amended complaint against Amazon describes the Prime "Iliad" flow as "designed to be labyrinthine" — a "four-page, six-click, fifteen-option cancellation process." Adobe drew an FTC action over hidden early-termination fees and a $150M settlement. The Click-to-Cancel Rule was vacated on procedural grounds in July 2025, but enforcement continued under §5/ROSCA, the FTC moved to revive the rule in March 2026, and ~30 states run their own automatic-renewal laws — some stricter than the vacated federal version.
Billing architecture is retention infrastructure
The least glamorous retention levers are contractual, and they are enormous. Annual plans run 10–20 percentage points higher NRR than monthly across 3,500 companies. In consumer subscriptions, RevenueCat's 115,000-app dataset puts year-1 retention at 1–2% for weekly plans, 6–14% for monthly, and 20–40% for yearly — the plan length is the retention curve — and every band is deteriorating cohort over cohort (yearly median 31%→28%, monthly 10%→8%). One operator ran the trade explicitly: annual-only pricing cut signups 60% and beat the old model's revenue by month six, having previously paid CAC on twenty monthly customers of whom eight were gone inside 90 days (self-reported, no controls — but directionally consistent with the ChartMogul figure).
Two cautions keep this honest. Price point sorts who you recruit: cheap tiers recruit customers who structurally cannot retain — visible at its most extreme in AI products under $50/month running 23% GRR against 70% above $250/month — so a blended retention number judges the product on a population it was never going to keep. And contract length can mask rather than fix: "Churn is a lagging indicator," the low-engagement death spiral that used to surface in year three now surfaces in year two, and ChartMogul documents AI-era contracts written with 3-month opt-outs where 70–80% of customers opt out. Gainsight's line for it: "Retention doesn't fail all at once. It erodes quietly until renewals expose it."
The AI churn wave
The starkest retention data of the era: ChartMogul's AI-native cohort (~200 companies over $250K ARR) shows median GRR of 40% and NRR of 48% — against 82% NRR for B2B SaaS overall. The mechanism in their words: "the curse of the AI wrapper… easy to buy is being easy to cancel." In consumer, AI apps churn 36% faster than non-AI apps (12-month annual-plan retention 21.1% vs 30.7%). And the cohorts behave like nothing in classic SaaS: a16z's "Cinderella glass-slipper effect" — a model briefly fits a use case, attracts a tourist cohort, and when a better fit ships, the cohort leaves wholesale; for commodity models, "no cohort ever formed a durable attachment."
Handle the era's coping frameworks with care. a16z proposes rebasing retention from month 0 to month 3 because "most often, the curve begins to flatten around M3" — analytically defensible for separating tourists from adopters, and also self-servingly flattering, since it excludes exactly the window where most AI churn occurs; an M12/M3 figure is not comparable to any classic M0-based benchmark. They also report a "smiling" retention curve for AI products ("ChatGPT's retention curve is the perfect example") — which directly contradicts Chen's long-standing rule ("What you never see is a curve that starts high, then goes low, then becomes high again. That's not possible"). Someone's rule is breaking; the honest read is that the data is too young to say whose. And the whole AI benchmark set is fragile: ~50 companies per price bucket, survivorship-filtered, self-described as "directional rather than statistically bulletproof," with AI GRR moving 13 points (27% → 40%) inside nine months. Cite AI retention numbers with their measurement month or not at all.
Price point sorts for commitment; the wave is real but stratified — and moving fast enough that every figure needs its measurement date. Source: ChartMogul, FY2025.
The numbers behind the figure
| Segment | GRR | NRR |
|---|---|---|
| B2B SaaS (median) | — | 82% |
| AI-native (median) | 40% | 48% |
| AI >$250/mo | 70% | 85% |
| AI $50–249/mo | 45% | 61% |
| AI <$50/mo | 23% | 32% |
Eight ways retention fails, mechanically
- Scaling acquisition while the curve never flattens. Churn volume grows with the base; acquisition substitutes for the harder work until it mathematically can't (Chen; a16z's restatement — "increasing GTM spend simply adds water to a leaky bucket").
- Reading survivor-biased aggregates as improvement. Aggregate retention rises while no cohort improves (the ruse of heterogeneity). Cohort everything.
- The NRR mirage. Expansion, seat growth, and price increases net against gross decay: "almost everyone saw NRR fall in 2023… Strip our price increases, it looks even worse." Report GRR and logo retention beside NRR, always.
- Usage read as value in B2B. The buyer's outcome is economic, not behavioral; health reads green while the account renews out (Skok; Murphy).
- Onboarding failure and champion loss misdiagnosed as "retention problems." The deficit was created in week one, or in one person's departure — and no amount of month-six lifecycle email touches either (Skok's top two).
- Involuntary churn left unmanaged and read as product churn. 22–35% of the total, ~70% recoverable, invisible in every product metric (Churnkey, Recurly).
- Retention assigned to people with no product authority. Reforge's documented lose-lose: "a company needs to improve retention, they assign a marketer or two to own the metric, the marketer can only change some emails" — while retention lives in the product, the pricing, and the payments stack.
- Over-notification burning the channel. LinkedIn's controlled experiment: dropping ~half of emails cost 2.6% of page views, while full-volume sending generated 45% more negative responses — and unsubscribes and spam flags are permanent channel loss. They shipped the cut: 50% fewer emails, complaints down 65%.
The counterexamples file
For every piece of standard retention advice, there's a documented case of the opposite working. The anecdote rows are marked because they're illustrations of a pattern, not effect sizes.
| Standard advice | What actually happened |
|---|---|
| Don't add friction to something users already have — mass churn. | Netflix rolled paid sharing to 100+ countries: "Revenue in each region is now higher than pre-launch, with sign-ups already exceeding cancellations" (SEC-filed). Q2 net adds: +5.9M vs −0.97M a year earlier. |
| Price increases raise churn; raise gently. | 37signals collapsed Basecamp to one ~$99 flat plan: "sign ups fell… only by a little bit. But average ticket price per sale went up significantly, and conversion rates went up as well." |
| Cancellation must always be one click. | Baremetrics, at ~13% churn, removed self-serve cancellation and required a conversation — "cut our churn in half… literally saved Baremetrics." Caveat: reinstated two years later, and the tactic predates today's FTC/ROSCA posture. The transferable lesson is talk to churning customers, not hide the button. |
| Track NPS as your retention leading indicator. | The Centre for Effective Altruism publicly dropped NPS because the research "failed to replicate, even when the replication was using the originally published data set (!)". |
| Gamification buys engagement now, costs retention later. | Duolingo's Energy system ("a carrot, not a stick," SEC-filed) "increased DAUs, median time spent learning well, and subscriber conversion… We have rarely seen a feature move more than one of these metrics, let alone all three." Hold beside it: the ACM study naming Duolingo's exit dark patterns — the variable is reward-vs-punishment framing, not gamification per se. |
| More lifecycle email, more retention. | LinkedIn: 50% fewer emails, complaints down 65%, at a cost of 2.6% of page views. |
| High-touch onboarding doesn't scale — automate early. | Superhuman ran human onboarding for thousands of paying customers per week for ~3 years before going self-serve; the manual phase taught self-serve what to do. Sequencing, not permanence. |
| A curve that doesn't flatten high means no PMF. | Reforge's ICED analysis places Zillow, TurboTax, Opendoor, Carvana in the yearly-or-less band — the retention-curve test is the wrong instrument for episodic products. (Primary financials unverified — flagged.) |
| Retention is a product-quality problem. | Balfour: "There are terrible products that have reached $1B+ and amazing products that never make it anywhere." |
| Maximize signups; offer monthly. | Annual-only at $468/yr: signups −60%, revenue ahead by month six (self-reported, no controls). |
| Run QBRs to protect renewals. | One CS team swapped QBRs for biweekly 15-minute async updates: renewal 87%→91% (single team, one quarter, no controls). |
The pattern: the winners treated retention as a system property — of pricing, contracts, payments, channel, and cohort composition — and were willing to trade top-line optics (signups, email volume, sharing accounts, QBR theater) for base quality.
Questions operators actually ask
What's a realistic churn rate?
There isn't one number — churn is tenure-dependent (~25% in month one falling to near zero for survivors is a normal shape), model-dependent, and ACV-dependent. Anchor on cohorted curves and the model-relative benchmarks above; any single blended rate you're quoted, including this question's answers, is close to meaningless.
How do I catch churn before it happens when the signals are relational?
Mostly unsolved — the honest state of the art. Usage dashboards miss champion decay by construction. The practical partial answers: track what accounts stopped doing (not what they never started), instrument champion behavior specifically (response latency, meeting attendance, login recency of the named champion), and build a second relationship into every account before you need it — champion loss is a top-two churn cause precisely because it has no backup.
Is NPS worth tracking?
As a growth predictor, the replications say no. If you keep the survey, use the plain mean of the likelihood-to-recommend item (the promoter/detractor bucketing throws away information), and treat it as one lagging sentiment signal — never as the retention system.
What's the minimum fix for failed payments?
Card retries with smart timing, an account updater, and a Day 1/7/14 dunning sequence — the ~70% recoverable share concentrates in the first ten days. If you do nothing else in retention this quarter, do this; it's the only churn that fixes itself with configuration.
How do I think about retention when my product's goal is for users to stop needing it?
Measure the episode, not the habit: completion of the outcome, satisfaction at exit, return on next need, and referral in between. ICED-style thinking applies — the asset you're retaining is memory and preference, not weekly usage. A "churned" user who succeeded is your acquisition channel.
How do I detect churn risk in AI/agentic products where logins are meaningless?
The instrumentation is unwritten — the operator threads propose tracking output quality signals like edit rate rather than sessions. The retention-relevant events are: outputs accepted vs discarded, workflows completed end-to-end without human rescue, and the product's outputs appearing in downstream systems. If the agent works while the user is absent, presence metrics are noise.
