VIII. Your Move: The Role Assignments
The Benchmark's Blind Spots
What Sixteen Years of Excellent Data Still Cannot See
Start with what KeyBanc Capital Markets deserves: credit.
For sixteen years, KBCM has produced the only longitudinal, public-facing performance survey of private SaaS companies at this depth. Growth rates, retention, sales productivity, expense structure, pricing models: tracked year over year, across market cycles, through a pandemic, a correction, and an AI platform shift. Private-company operating data is the hardest data in software to get, and KBCM has given it away annually as a public good.
This report exists because that survey exists. Every retention series analyzed, every efficiency trend traced, came from KBCM’s instrument.
That is precisely why its gaps matter.
The KBCM survey is an incomplete picture of SaaS company performance. Not because it measures anything badly, but because its metric architecture was designed around the acquisition-led growth model and never expanded to measure the functions that now determine performance. The gaps are not randomly distributed. They cluster, without exception, on the retention side of the business. Which means the benchmark systematically under-informs exactly the decisions its own last four years of data show matter most.
Here is the arithmetic that makes the case. In the survey’s own series, expansion’s share of pooled new ARR rose from 42% (2022) to 45% (2023) to 52% (2024): a majority of new revenue, and above $25M ARR the majority is larger still (53 to 56%). Over the same window, NDR fell from 106% to 101%, the NDR–GDR spread compressed from 20 points toward a forecast 13, the net magic number sat frozen at 0.5 for four consecutive years, and new-logo-only CAC payback stretched from 31 to 35 to 37 months.
The survey documents, precisely and honestly, that the acquisition engine stalled and the retention engine became the growth engine. It just cannot tell you why, because it never built the instruments to measure the retention side’s inputs.
One Function Gets a Numerator. The Other Gets a Denominator.
The clearest evidence of the architecture’s bias is a comparison the survey itself makes possible.
For Sales, KBCM measures both capacity and output. AE Productivity, defined in the 2025 survey as (New Logo + Expansion ARR) divided by quota-carrying AEs, is tracked from 2022 data forward. AE Quota is tracked longitudinally ($795K in 2022, holding near $800K through the 2026 forecast). Quota attainment gets its own quartile distribution. Sales, in other words, has a numerator: you can see what the function produces per dollar and per head.
For Customer Success, KBCM measures capacity only: CSM headcount, book of business (roughly $2.1M overall in 2023 data), accounts covered per CSM (about 21 overall in 2024 data). There is no output metric. No revenue-retained-per-CSM. No expansion-sourced-per-CSM.
A benchmark reader can tell you how loaded a CSM is, but not what a CSM produces, which means CS investment cannot be evaluated in the same units as Sales investment. And in a budget meeting, the function with a productivity number beats the function with a headcount number every time.
The fix is one metric: Net Revenue Retained + Expansion ARR Sourced per FTE CSM. Same construction as AE Productivity, same units, same survey section. It would make the retention investment case comparable to the acquisition investment case for the first time in the survey’s history.
The asymmetry compounds from there.
Expansion ARR, now the majority of new ARR, appears in the survey only inside the AE Productivity formula. That placement embeds an assumption: that expansion is Sales-owned. In practice the industry runs at least two models, a split/handoff model where CS develops and Sales closes, and a CS-owned model where CS carries expansion end to end, and the survey cannot see the difference. There is no question asking who owns expansion quota and close.
There is no question about CSM compensation structure either: quota-carrying or not, at what rate, with what expansion-credit split. A single demographic item would resolve the ownership question empirically, the way the survey already resolves sales-motion questions.
And the org-structure curiosity exists; it was just never pointed at the retention motion. The survey asks how companies source leads (83% quality-based versus 17% volume-based in 2025) and how SDR teams are oriented (21% inbound, 56% outbound, 24% hybrid on average). It asks nothing about who executes renewals. The renewal, the single event where retained revenue is won or lost, has no owner in sixteen years of instrument design.
To defend the current design, you would have to believe that a benchmark should measure the majority revenue source through the lens of one possible owner, without ever asking which owner a respondent actually uses. That is not a benchmark position; it is an inherited default. Inference
The Cost Line That Cut Itself Invisibly
The survey’s expense architecture has the same shape.
Customer Success has never been its own OpEx line. It is allocated primarily inside S&M: 65% of CS cost on average in 2021 data (2022 Survey, p. 52), 61% in 2022 data (2023 Survey, p. 30). Now put that beside the survey’s most dramatic recent series: S&M spend falling from 54% of revenue in 2022 to a forecast 32% in 2026, a 22-point cut, the deepest of any expense line.
Breaking CS out as its own expense line would cost the survey one row. It would let readers see, for the first time, whether the industry funded or defunded its actual growth engine during the efficiency era.
Revenue Nobody Gets Credit For
Two more gaps share a root cause: the survey has no concept of defended revenue.
There is no churn-prevented or save-rate metric anywhere in sixteen years of instrument. A team that saves $2M of at-risk ARR has produced something economically identical to a team that sold $2M of new ARR at zero acquisition cost, and it appears in no one’s credit column, in the survey or in the compensation plans the survey implicitly models. Because the survey never counts the save, no compensation plan modeled on the survey rewards it.
Professional services get the same treatment from the other direction. The survey has repeatedly shown PS attach to be one of the strongest churn suppressors in its dataset: in the 2019 survey (2018 data), companies with no PS attach churned at 16.3% gross versus 5.8% for companies attaching PS to more than half of deals. The 2021 and 2022 surveys found the same gradient: 15% falling to 9%, then 13% falling to 5%.
Yet PS appears only as a churn-suppression input and a margin drag, never as a credited output. The survey’s own data says services-led delivery creates retention value; its metric architecture keeps that value in the cost column. Elevating services and CS-adjacent delivery into the value-creation narrative would simply bring the framing into line with the evidence the survey already publishes.
Aggregates Without Causes
GDR and NDR are the survey’s retention instruments, and they are lagging aggregates.
When GDR runs 86/85/86 across 2022 to 2024, the survey can say retention held; it cannot say why companies lost what they lost. There is no churn-reason decomposition: product gap, onboarding failure, ICP misfit, competitive loss, pricing. Without cause attribution, accountability cannot be assigned to the function that generated the loss, and every function can plausibly point at another.
The gap has a second edge, and it points backward into the sales cycle. Every churn lever the survey has ever documented is a variable fixed at or before the close: contract length, contract size, services attach. No post-close operational variable carries a measured churn gradient anywhere in seven editions, partly because, as this page documents, none was ever instrumented. A churn-reason decomposition would settle the question the pre-close gradients only imply: how much of each year’s loss was committed before kickoff, in deal fit, promise accuracy, and contract structure, and how much was genuinely lost afterward. Without it, the industry funds retention where churn surfaces rather than where it is caused, which is how six years of post-sale investment coexisted with a floor that never moved (see The Floor Is Set Upstream).
The Product side of this is its own gap. R&D spend is tracked carefully, 39% of revenue in 2022 falling to a forecast 26% in 2026, but no metric connects product investment to retention outcomes. No product-gap churn attribution, no outcome-delivery measure. The survey can show you that companies cut R&D by roughly a quarter through the actuals, and by a third at the forecast endpoint; it cannot show you what that cut did to the revenue base, because the causal wire between product investment and retention was never installed.
The stakes on this wire are rising, not falling. Forty percent of 2025 respondents still price primarily on seats (33% fixed, 7% variable). In an AI era where seat counts compress, seat-priced revenue defends itself only through demonstrated outcomes, which makes the absence of any outcome-delivery measure a forward-looking gap, not just a historical one.
The Functions That Aren’t There
And then there is Support: not under-measured but unmeasured.
In sixteen years, Support appears exactly twice: once as a cost-allocation split (87% of Support expense sits in COGS; 2023 Survey, p. 30), and once as an AI-opportunity ranking, where 55% of 2025 respondents named customer service and support as a top AI opportunity. No ticket volume, no escalation rate, no resolution time, no signal-quality metric, nothing that treats Support as an operating function with performance.
If the survey aims to represent the full SaaS operating model (Product, Marketing, Sales, CS, Support, PS), its GTM picture covers four of six, and its own best evidence points at one of the missing two. The respondents clearly believe Support matters; they ranked it the place AI will pay off first. The instrument has no way to measure whether they’re right. When AI-driven support transformation shows up in next year’s data, the survey will be able to see the cost effect and nothing else.
One more fact reframes all of the above. This is, in plain fact, a finance-office survey: 60 of 71 respondents in the 2025 edition sit in the CFO, finance, or accounting function (CFO 29, Finance 28, Accounting 3; 2025 Survey, p. 8). A survey answered from the finance seat reports what the finance systems of record can see, and those systems were built for the acquisition-led model. Inference
The Gap Inventory
Eleven gaps follow, and it is worth being precise about what they collectively are. They are not eleven separate oversights. They are one design decision, made repeatedly and consistently across sixteen years: when a question could be asked about either side of the revenue engine, it was asked about acquisition.
Sales got the output metric; CS got the capacity metric. Lead sourcing got the demographic question; renewal ownership didn’t. Acquisition cost got its own payback series; retention cost got folded into someone else’s line. That consistency is what makes the pattern structural rather than accidental, and what makes it fixable, because the templates for every missing instrument already exist elsewhere in the survey.
| # | Gap | Why it matters | Suggested addition | What it enables |
|---|---|---|---|---|
| 1 | No CS output/productivity metric (Sales gets AE Productivity + Quota; CS gets capacity only) | CS investment can’t be evaluated in the same units as Sales investment | Net Revenue Retained + Expansion ARR Sourced per FTE CSM | Retention vs. acquisition ROI comparison in one currency |
| 2 | No expansion-ownership question | Expansion (52% of new ARR) is measured through an implicit Sales-owned lens | Demographic item: who owns expansion quota/close | Benchmark expansion performance by ownership model |
| 3 | No CSM compensation/quota-structure question | Ownership model invisible; comp design unbenchmarkable | Quota-carrying Y/N, rate, expansion-credit split | Empirically resolves #2; comp benchmarking for CS |
| 4 | CS cost buried in S&M (65% allocated 2021; 61% 2022) | The 22-pt S&M cut (54%→32%) reduced retention capacity invisibly | CS as its own OpEx line | Shows whether retention was funded or defunded |
| 5 | No churn-prevented / save-rate metric | Defended ARR appears in no one’s credit column | Dollar-denominated saves / at-risk resolution rate | Values retention work at its economic worth |
| 6 | No churn-reason attribution | GDR/NDR are lagging aggregates without causes | Churn decomposition: product, onboarding, ICP, competitive, pricing | Assigns accountability to the causing function |
| 7 | No renewal-execution-ownership question | Survey asks SDR orientation and lead sourcing but not who runs renewals | Org-structure item for renewal ownership | Benchmark renewal performance by structure |
| 8 | PS attach tracked only as churn-suppression input (16.3% vs 5.8% churn gradient) | Value-creating delivery work stays in the cost column | Elevate services/CS-adjacent delivery as credited output | Retention value of services becomes visible |
| 9 | Support unmeasured (two signals in 16 years: expense split; 55% AI-opportunity rank) | Full-operating-model picture missing an entire function | Ticket/escalation/resolution + signal-quality metrics | Support benchmarkable; AI-impact claims testable |
| 10 | No product-investment-to-retention linkage (R&D 39%→26% with no outcome wire) | R&D cuts can’t be evaluated against revenue-base effects | Product-gap churn attribution; outcome-delivery measure | Makes the R&D line causally interpretable |
| 11 | PS measured as input, never as function (attach, revenue share, and the steepest churn gradient in the dataset, but no margin, headcount, ownership, or expansion linkage) | The survey’s strongest causal lever has no functional home, so it can’t be staffed, funded, or benchmarked as one | PS margin and headcount rows; org-ownership item; attach-to-expansion cut | The proven churn lever becomes an operable function |
The Strongest Objection, Taken Seriously
The best defense of the current design runs like this: survey length is the enemy of response rate; the respondents are largely finance leaders who know their expense lines and sales metrics cold but may not have CS operational data at hand; and a benchmark earns its longitudinal value precisely by not changing its questions.
Each point deserves an answer.
On length: the inventory above is three demographic questions, one expense row, and a handful of metrics built in the survey’s existing house style. The survey has absorbed larger additions before. AE Productivity itself first appears with 2022 data, and the AI section added an entire topic area in two cycles. The instrument demonstrably evolves; the question is only in which direction.
On respondent knowledge: the same CFOs who report accounts-per-CSM and CSM book of business, which the survey already collects, can report revenue retained per CSM from the same systems. The capacity data proves the operational visibility exists. Only the output question was never asked.
On longitudinal integrity: additions do not break series; they start new ones. Every series in the survey was new once. The alternative, never measuring the half of the business that now produces most new ARR in order to preserve comparability, keeps the instrument consistent and leaves its purpose unmet. Inference
What the Additions Would Actually Buy
None of this would make the survey a different product. Every suggested addition is a question or a row, in sections that already exist, built with the same instinct that produced AE Productivity, the SDR split, and the PS attach gradient, pointed, this time, at the other half of the revenue engine.
What the additions would buy is causal interpretability for the metrics KBCM already publishes. Today the survey documents that retention weakened: NDR from 106% to 101%, the expansion spread narrowing from 20 points toward 13, the net magic number flat at 0.5 for four years while new-logo payback stretched to 37 months. With the added inputs, it could show why, and by function: whether NDR fell because CS capacity was cut inside S&M, because expansion ownership sat with sellers compensated to hunt, because product gaps went unattributed, or because saves nobody measured stopped happening.
A benchmark that can assign cause enables better guidance. That is not a new mission; it is the survey’s stated one.
Who This Serves
The practical implications land differently by reader, and all of them favor the additions.
For operators, the current architecture forces a choice between benchmarking what the survey measures and managing what the business needs. A CS leader defending headcount has KBCM data on how loaded competitors’ CSMs are, and nothing on what those CSMs produce. With gap #1 closed, the same leader walks into planning with a productivity benchmark denominated in the same units as the AE number across the table.
For boards and investors, the additions convert a descriptive benchmark into a diagnostic one. An investor comparing two companies at 101% NDR today sees a tie. With ownership models, CS expense lines, and churn-reason decomposition visible, one of those companies may show a funded, CS-owned expansion motion losing revenue to product gaps, and the other an underfunded handoff model losing revenue to onboarding failure: identical aggregates, opposite interventions, different valuations.
For KBCM itself, the gaps mark data no benchmark yet collects. The survey’s franchise is being the instrument the private SaaS market trusts; the first benchmark to publish retention-side inputs by ownership model becomes the only place those comparisons exist. Sixteen years of longitudinal trust is exactly the asset required to collect data nobody else can.
For the industry’s data commons, one caution: every year the acquisition-side series extend and the retention-side series don’t, the interpretive default hardens. Analysts, LPs, and operators train on what the benchmark shows. An instrument this influential doesn’t just measure the industry’s model; over sixteen years, it quietly teaches it.
The asymmetry carries an interpretive cost for readers, too. When every output metric belongs to Sales, the eye assigns every trend to Sales; the survey’s architecture makes the Sales-framed reading of the expansion mix and the expense trajectories the default one, even for careful readers. That reading bias is the report’s argument for completing the instrument. Inference
Sixteen years of data earned KBCM the right to be the industry’s benchmark. The next sixteen depend on measuring the model the industry actually runs. The path forward is sequenced in The Playbook.
Frequently asked questions
What does the KBCM survey fail to measure?
The survey measures SaaS acquisition with sixteen years of precision but infers retention from residue. It contains 11 structural gaps, all clustered on the retention side, including no output metric for Customer Success and no churn-reason decomposition anywhere in the series.
Why does Customer Success get less benchmark data than Sales?
Sales gets both a capacity metric and an output metric, AE Productivity, tracked since 2022 data. Customer Success gets capacity only: headcount, book of business, accounts per CSM. There is no output metric, so CS investment cannot be evaluated in the same units as Sales investment.
What does KBCM's benchmark survey miss about Professional Services?
Professional Services is measured four separate ways, attach by GTM motion, share of first-year ARR, revenue composition, and churn by attach, yet recognized as a function zero times: no margin, headcount, ownership, or expansion-linkage question exists despite carrying the steepest churn gradient in the dataset.
Who actually answers the KBCM SaaS benchmark survey?
Sixty of 71 respondents in the 2025 edition sit in the CFO, finance, or accounting function, CFO 29, Finance 28, Accounting 3. A survey answered almost entirely from the finance seat reports what finance systems of record can see, and those systems were built for the acquisition-led model.
Last reviewed: July 2026
