IV. The Proof: Check Our Work

How to Misread a Benchmark

In Brief
Benchmark surveys are misread in recurring ways: treating the latest edition as the truth about every year it prints (each edition restates history, 79% of verified restatements downward), reading E-year estimates as results (optimistic in 10 of 12 checkable cases), and splicing trends across editions or providers, where different panels, statistics, definitions, bases, fielding windows, and renamed series manufacture movement that never happened: the same year’s NDR spans 102% to 108% across surveys. The defensible practice: pin each figure to one named edition, never splice, and show cross-source disagreement as a range.

Benchmark surveys are the most misquoted documents in SaaS. Not because operators are careless, but because the surveys look like something they are not: a measurement of the industry, stable across time, comparable across publishers. Seven editions of reconciliation work taught us how each of those assumptions fails, and this page collects the failures in one place. Every example is from this report’s own certified record. The goal is not to talk you out of using benchmarks. It is to keep the number you quote from being taken apart by the first person who checks it.

Misconception 1: The latest edition tells you the truth about every year it prints

Each edition restates the years it covers, because the respondent panel is redrawn annually and the new panel re-answers history. The 2024 edition put the median 2022 EBITDA margin at negative 26%; the 2025 edition put the same year at negative 47%. Median 2022 ARR was $22.5M in one and $17.0M in the next. Neither is a typo and neither is a correction; they are different panels remembering the same year. Across the 39 restatements we verified between those two editions, 79% moved the past in a worse direction. A single edition is one panel’s memory of history. Quote it as that, pinned to its edition, or a reader holding the other edition will end your argument with a page number.

Restatement register: Data Integrity & Corrections LogVerified

Misconception 2: A stable number means a stable business

Three different failures hide inside this one, and a single year’s survey cannot expose any of them.

A stable composite can mask opposite-moving parts: NDR holds near 101% while roughly 14 points of churn are offset by 15 points of expansion, a composition you cannot see in the headline (the reason this report reports the parts separately). A rising median can mask a thinning population: median EBITDA improved 48 points while the share of companies actually clearing Rule of 40 fell from 11% to 5%. And a metric that is stable within every edition can be falling across them: read alone, each edition reports retention as steady; the 109%-to-101% decline appears only when seven editions are laid side by side, each held to its own vintage. Stability inside one survey year is a fact about that snapshot. It is not evidence about the trajectory, and it cannot be.

Composition: What 101% Hides; population thinning: EBITDA and the Rule of 40Verified

Misconception 3: The two most recent years in the survey are data

They are the survey’s own forward estimates, marked E, printed because the survey closes mid-year before the current year’s financials exist. Estimates have a record, and the record is checkable: where the 2024 edition’s estimates could be measured against the actuals the 2025 edition reported, the actual came in worse in 10 of 12 cases, concentrated in profitability, ARR dollars, and growth. A reader who quotes an E-year as a result is quoting the panel’s optimism, and the panel’s optimism has a documented lean. Read forward columns as a ceiling.

Estimate accuracy: Why the ‘E’ Years Persist and the estimate table on the same pageVerified

Misconception 4: You can build a trend by hopping between editions, or between providers

This is the one that ruins otherwise careful arguments, so it gets the full treatment. The temptation is understandable: no single survey covers every year and every cut, so the analyst assembles a series from whatever is at hand: 2019 from one edition, 2021 from another, 2023 from a different publisher whose chart looked cleaner. The resulting line is not a trend. It is a splice of measurements that differ in ways the label does not show, and every difference is a place your conclusion can be dismantled.

The mechanisms, each one sufficient on its own to manufacture or erase movement:

Different panels answering. Respondent pools are redrawn annually within one provider and are entirely different populations across providers. In the one year this report could compare directly, 2022 NDR ranged from 102% to 108% depending on which survey measured it: a six-point spread, larger than the entire three-year decline this report documents, produced by respondent pools alone.

Different statistics under the same label. A median company and a dollar-weighted pool are different measurements that move independently. KBCM itself switched its expansion-share statistic mid-series, from a company median through 2021 to a pooled figure from 2022, and the adjacency reads as a fall from 46% to 42% that never happened. A cross-provider splice does this silently, because nothing on the chart tells you which statistic each publisher chose.

Different definitions under the same name. Metric names are not standards. Within a single publisher, this report documents “Fully-Loaded CAC” meaning an S&M-only efficiency ratio in the survey’s definitions and an all-in cost concept in common usage, and a printed CAC-payback formula that inverts its own definition. Across publishers, the same words (net retention, payback, expansion) sit on top of quietly different arithmetic, and the differences are rarely printed where a reader will find them.

Different bases and cuts. A median of companies above $5M ARR and an unqualified median are different populations wearing the same metric name. This report keeps its own 2021 retention anchor off the main trend line for exactly this reason: it was printed on a different basis than the years around it.

Different fielding windows. A survey that closes mid-year and one that closes at year-end can both print a figure for the same calendar year while describing different months of it, straddling different macro conditions.

Renamed series. Publishers rename metrics between editions while revising their history. The 2022 median contract value appears under one label in the 2024 edition and a slightly different label in the 2025 edition, $10K apart; matched on names, the two prints never register as the same series. A splice built by label-matching inherits every such break invisibly.

Survivorship, differently shaped. Every panel tilts toward survivors, but each provider’s panel tilts differently: by stage, by ownership, by who keeps answering in a downturn. Splicing providers splices their biases, and the biases do not cancel; they compound in unknown directions.

Self-reported, unaudited, differently asked. All of these figures are what a finance office said in response to a questionnaire. Question wording, ordering, and what the respondent’s systems can even produce differ across instruments, and none of it is audited.

The consequence for anyone trying to make a defensible point: a spliced series does not merely weaken your argument, it hands your counterparty the argument. If assembling editions freely is allowed, any story can be assembled, including the opposite of yours, and the debate collapses into dueling splices. The discipline this report follows exists because it is the only position that survives an audit: every figure pinned to one named edition, no values joined across editions or providers into one line, and where two sources genuinely measure the same year, the disagreement shown as a range rather than resolved by preference. Inference

Certified anchors: cross-survey spread (Cross-Survey Measurement Variation), statistic change (Two Regimes, Not One Trend), name collision and formula inversion (On “Fully-Loaded CAC”), renamed series (Reading KBCM Across Editions). Fielding-window and survivorship-shape mechanisms are stated generally, not as claims about any named outside publisher.Calculated

Misconception 5: More sources make the number more defensible

The instinct is to cite three surveys where one would do, on the theory that agreement is evidence. Agreement between instruments that measure different panels with different definitions is closer to coincidence than corroboration, and disagreement between them is not a scandal, it is expected. One source, named, pinned, with its basis and its limits stated, is a defensible citation. Five sources blended into a single confident number is a liability wearing the costume of rigor. Where this report cites outside the KBCM series, it cites for triangulation, states the source’s own methodology, and never averages across instruments.

Misconception 6: The benchmark measures the industry

It measures its panel, through its instrument, as answered by whoever fills it in. Sixty of 71 respondents to the 2025 edition sit in the CFO, finance, or accounting function, and the instrument shows it: dozens of sales and spend benchmarks, no Customer Success spend line in sixteen years, no owner for the renewal anywhere in the question set. What a survey does not ask about does not exist in its data, and analysts trained on the data inherit the omission. The fuller argument is in The Benchmark’s Blind Spots; the reading rule here is simpler: before quoting a benchmark on a topic, check whether the instrument actually asks about it, or whether the figure you are about to quote is the residue of questions pointed somewhere else.

Respondents by role: KBCM/Sapphire Survey 2025 p8; instrument coverage: The Benchmark’s Blind SpotsVerified

Misconception 7: If everyone quotes it, someone measured it

The most repeated numbers in SaaS are the least traceable. The claim that roughly 80% of future revenue comes from existing customers, and its companion that expansion reaches 60 to 70% of new ARR at maturity, trace to no primary source we could locate, and the certified record contradicts the picture they paint: the measured company median for expansion share never crossed 46%. Ubiquity is not provenance. A figure that arrives without an edition, a page, and a stated basis is folklore until proven otherwise, and folklore compounds: each citation of a citation adds confidence and subtracts nothing.

Folk-belief tracing and the measured median: The Expansion MythVerified

The Reader’s Rules

The misconceptions reduce to five working rules, which are the rules this report was built under. Pin every figure to one named edition and quote it with that pin. Never splice editions or providers into one series; show disagreement as a range. Treat E-years as the panel’s optimism, bounded above. Trust gradients measured inside a single edition and directions confirmed across several; hold levels loosely, and treat any single-year improvement as unproven until a second edition confirms it. And before quoting any benchmark, ask what the instrument can see, because you are not quoting the industry, you are quoting a questionnaire’s view of it.

Held to those rules, the benchmark record is the most valuable public evidence this industry has. Held loosely, it will prove anything, which is the same as proving nothing.

How these rules were arrived at: The Story Behind the Story; the full method: MethodologyCalculated

Frequently asked questions

Can you compare metrics across different SaaS benchmark surveys?

Not safely. Panels, statistics, definitions, bases, and fielding windows all differ across providers. In 2022, reported median NDR ranged from 102% to 108% depending on the survey, a spread larger than the three-year decline this report documents.

Why do SaaS surveys report different numbers for the same year?

Respondent panels are redrawn annually, so each edition re-answers history. The 2024 edition put median 2022 ARR at $22.5M; the 2025 edition put the same year at $17.0M. Both are correct for their own edition.

Are the current-year figures in the KBCM survey actuals or estimates?

The two most recent years in every edition are the survey’s own estimates, marked E, because fielding closes mid-year. Checked against later actuals, those estimates came in optimistic in 10 of 12 cases.

What is the safest way to quote a benchmark survey figure?

Pin it to one named edition and page, state the basis and cut, never splice editions or providers into one trend line, and treat estimate years as a ceiling. Show disagreement between sources as a range rather than averaging it away.

Last reviewed: July 2026

Previous
The Story Behind the Story