BlogPlaybooks

PLAYBOOKS

Half the AI sales statistics you have seen come from the vendors selling it

Four published AI SDR reply rates disagree by more than a factor of two, and all four publishers sell into the category they are measuring.

SAGARISPlaybooks6 min
Half the AI sales statistics you have seen come from the vendors selling it

The AI SDR performance figures circulating in sales decks this year come from four pages: https://www.kwanzoo.com/blog/why-first-gen-ai-sdrs-failed, https://www.digitalapplied.com/blog/ai-sdr-statistics-2026-outbound-sales-data-points, https://www.devcommx.com/blogs/ai-sdr-reply-rates-roi and https://ziellab.com/post/ai-sdr-what-works-after-the-hype-2026-guide. Every one of those four is published by a company selling into the category the numbers are about.

Here is what that set of four pages contains. Outbound volume up 6.4x, with positive reply rates falling to 1.3 percent against a 2.1 percent human baseline. Per-rep monthly volume rising from 1,150 to 7,400, with reply rates falling from 4.7 percent to 2.9 percent. Roughly half of AI SDR pilots shut down within 90 days. Between 50 and 70 percent cancelled within a year. We cite the four URLs as a set rather than mapping each figure to one page, because the figures recirculate across all of them and we could not establish who published which first.

Now look at the two baselines in that set. One set of figures says a human SDR gets a 2.1 percent positive reply rate. The other says 4.7 percent. Those differ by more than a factor of two, and they are describing the same job in the same market in roughly the same period. The post-AI numbers disagree by a similar margin, 1.3 percent against 2.9 percent.

They cannot all be right, and none of the four publishers is disinterested.

Do not take any of those numbers out of this article. They are here as the specimen, not as evidence.

A note on the title

We have not counted. "Half" in the headline above is a figure of speech, not a measurement, and in a piece about unsourced numbers it would be dishonest to leave that implicit. What we can say precisely is narrower and still bad: for the AI SDR performance statistics we tried to trace, every figure we found was published by a company selling into the category, and they contradict each other by more than a factor of two on both the human baseline and the post-AI rate. A statement attributed to Bain Capital Ventures in April 2026 is quoted second-hand all over the category; we could not locate the primary post.

The best citation in this category is not a statistic

There is one piece of evidence about AI SDRs that needs no methodology at all, and it is stronger than any of the four numbers above.

Artisan ran the billboards that defined the category. "Stop Hiring Humans." On its own blog, at https://www.artisan.co/blog/stop-hiring-humans, originally dated 3 May 2026 and updated in August 2026, the company wrote: "In August 2026, we officially retired the slogan, and Artisan is hiring its first human BDR." And: "We told the world to stop hiring humans. Now we're hiring our first human BDR. Both were the right call."

That is a dated, first-party, self-implicating statement on the publisher's own domain. There is no panel, no sample frame, no definition to argue about. It is the vendor telling you what it did.

This is worth internalising as a sourcing habit. When you are choosing between a percentage from a party with an interest and a plain dated statement of what someone actually did, the plain statement is usually the better evidence, even though it feels less quantitative and therefore less serious.

Three more numbers, and where they come from

"88 percent of AI pilots fail." This is IDC research conducted in partnership with Lenovo, published in the Lenovo CIO Playbook 2025 and reported at https://www.cio.com/article/3850763/88-of-ai-pilots-fail-to-reach-production-but-thats-not-all-on-it.html in March 2025. Three things about it are routinely lost. It is more than a year old. It measured AI proofs of concept generally, not agentic pilots, which is the framing it now gets quoted under. And what it measured was conversion to production: for every 33 AI proofs of concept a company launched, four graduated. Separately, the claim that the figure was independently replicated by a16z and an MIT Sloan CIO panel appears only in SEO content, and our sourcing review located no primary source for either replication.

"95 percent of AI pilots fail." The MIT NANDA report is the origin, popularised by https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/. What the MIT NANDA report measured was that 95 percent of integrated generative-AI pilots showed no measurable profit-and-loss impact. That is a narrow financial claim, and "no measurable P&L impact" is not "failed". The finding rests on roughly 300 publicly disclosed initiatives whose sources were not cited, plus 52 structured interviews, which the report itself described as directionally accurate. Critiques at https://www.marketingaiinstitute.com/blog/mit-study-ai-pilots and https://www.futuriom.com/articles/news/why-we-dont-believe-mit-nandas-werid-ai-study/2025/08. It is a year old and is not currently a live topic; it stays in circulation because it is a good slide.

"76 percent of support requests closed out without a human." This one comes from an acquisition announcement: https://www.salesforce.com/news/press-releases/2026/06/15/salesforce-signs-definitive-agreement-to-acquire-fin/, 15 June 2026, announcing a signed agreement to acquire Fin. The release states the agent closes out roughly 76 percent of incoming support requests. There is no methodology and no definition of "closes out", and at the time of the release the deal was signed rather than closed. It is now being quoted as an industry benchmark, which is a considerable distance from where it started.

A survey we do cite, with its conflict stated

Not every funded number is useless, and it would be lazy to pretend otherwise.

TrustRadius published a 2026 B2B Buying Disconnect report on 15 July 2026 (https://www.prnewswire.com/news-releases/trustradius-2026-b2b-buying-disconnect-report-reveals-ai-has-changed-how-buyers-research-but-not-what-they-trust-302825792.html) based on 1,862 technology buyers and 444 vendors surveyed globally in January 2026. Among AI-using buyers, 94 percent fact-check its responses at least some of the time. Analyst reports are used by 13 percent. Transparent pricing has been the number one buyer wish-list item for four consecutive years.

That has a stated sample, a stated field window, and a stated population. It is a real instrument.

Now apply this article's own test to it. TrustRadius is a software review platform. Its report finds that 74 percent of buyers use reviews. That is a finding about the value of reviews, published by a company whose business is reviews. It does not make the number false, and we have no reason to think it is. It does mean that the review finding and the fact-checking finding are not equally independent, and anyone citing the first should say who published it in the same sentence.

The same discipline applies to a number we find genuinely useful elsewhere: Ahrefs reported on 4 February 2026 that the presence of an AI Overview correlates with a 58 percent lower average clickthrough rate for the top-ranking page across 300,000 keywords (https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/). It is correlational, and Ahrefs sells SEO tooling and has a commercial interest in exactly that narrative. We cite it as an observation with a declared interest attached, not as a fact about the web.

And the counter-example that proves funding is not destiny: Backlinko, an SEO publisher, retracted its own widely repeated finding that question headlines get more clicks. Its updated study at https://backlinko.com/google-ctr-stats, which the page states was last updated 16 April 2025 across 4 million results from 1,312,881 pages, reports that question-based title tags perform similarly to non-question ones, 15.5 percent against 16.3 percent, and calls the difference not significant. SEO blogs still recycle the old number. A vendor publishing against its own earlier claim is the strongest signal in this whole landscape, and it is rare enough to be worth naming when it happens.

Three questions, and one admission

Before a statistic goes into your deck:

Who paid for it, and does that party sell the thing the number recommends? Not disqualifying. Disclosing.

What would have happened if the result had gone the other way? If a study finding the opposite would never have been published, the publication tells you almost nothing.

Can the method be inspected? A sample size, a field window and a definition of the counted thing. If those three are missing, you are quoting a sentence, not a study.

The admission. We sell AI software for revenue teams. We are a vendor in exactly the category described above, and we have no performance numbers of our own to offer, because we have no results to report yet. Our own site says so in the heading over the only figures we do publish: "We have no customer results to show you yet. Here is what you can check instead." Those figures are things like the price of a seat and the number of tests in the repository, which are facts about us rather than claims about outcomes for anyone else.

That is not modesty. It is the only position consistent with the argument above. If we told you that AI SDRs fail 70 percent of the time and that ours are different, we would be the fifth URL in the opening paragraph.

SAGARIS

Written by the SAGARIS team.

See the engine run on your pipeline.

Thirty minutes, your own data, no setup.

Book a demo

Get the next one in your inbox.

SAGARIS opens fully in October 2026. Join the waitlist and we will be in touch before launch.

We use these details to contact you about SAGARIS. See our privacy policy.

Book a demo