AI Is Talking Buyers Out of Purchases

The GEO Show
September 8, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Semrush and Exploding Topics report that 57.5% of AI users said chatbot information caused them not to buy something, while 65% said AI had at least partially replaced Google for product research. On episode 23 of The GEO Show, recorded September 8, 2026, Paris Childress covers purchase exclusion as a GEO risk, Latent Space's mentioned-versus-chosen gap, SurfacedBy single-check volatility and Reddit citation splits, Intender's 8% cross-engine source overlap, Nicholas Sitter's Google AI Mode citation architecture by question type, querying.ai's production collection API, a directional Claude retrieval-versus-citation study, Similarweb's historical AI visibility moat, and Mastercard's agentic shopping forecast.

Key takeaways

  • Semrush and Exploding Topics: among AI users, 57.5% said chatbot information caused them not to buy; 65% said AI partially replaced Google for product research; weekly users bought on organic AI recommendations 73.6% of the time.
  • Latent Space Frontier tracker: Kysely appeared in 42 of 42 answers but was first choice once; ChatGPT 5.6 Sol and ChatGPT 6 Astra differed on leaders in 33 of 121 categories.
  • SurfacedBy: 80% of appearing brand combinations were inconsistent across checks; a single check matched the majority outcome only 72.2% of the time on inconsistent combinations.
  • SurfacedBy Reddit window: ChatGPT Reddit citations fell from 37.9% to 3.4%, while Perplexity and Gemini did not collapse the same way.
  • Intender's 27,924-citation study found only 8% source overlap across five engines, with 71% of domains appearing on one engine only.
  • Nicholas Sitter: entity-wrapped Google AI Mode citations hit 87.6% on recommendation prompts and 0% on generic informational queries; Mastercard forecasts 300M+ shoppers could use AI agents to shop and pay by 2030.

Semrush and Exploding Topics report that 57.5% of AI users said chatbot information caused them not to buy something. Among the same AI users, 65% said AI had at least partially replaced Google for product research, and weekly AI users bought on organic recommendations 73.6% of the time.

Paris Childress recorded episode 23 of The GEO Show solo on September 8, 2026. He treats purchase exclusion and recommendation role as first-order GEO risks, then walks mentioned-versus-chosen metrics, single-check volatility, engine-specific Reddit citations, 8% cross-engine source overlap, Google AI Mode citation architecture by question type, production collection APIs, Claude retrieval gaps, historical AI visibility archives, and Mastercard's agentic shopping forecast.

Why can AI visibility include talking buyers out of purchases?

Because generative engines do not only discover products. They characterize them. Semrush and Exploding Topics surveyed 2,338 US adults (vendor-sponsored, ±2 percentage points). Among AI users, more than half said chatbot information caused them not to buy. Paris argues GEO is moving from pure visibility into reputation and buyer risk, and that negative characterization can be commercially worse than not appearing at all.

Being negatively characterized can be commercially worse than not appearing at all.

Why does mention share of voice overstate commercial visibility?

Latent Space's Frontier tracker holds 6,762 saved answers across 161 categories. Kysely appeared in every answer (42/42) and was first choice once. Category leaders also differed between ChatGPT 5.6 Sol and ChatGPT 6 Astra in 33 of 121 comparable categories. Paris separates first choice, recommended alternative, neutral mention, conditional mention, and objection as roles that SEO never forced brands to measure.

  • SurfacedBy found 80% inconsistency across checks for brands that appeared at least once, with a single check matching majority outcome only 72.2% of the time on inconsistent combinations.
  • Paris estimates roughly 30 to 40 runs per prompt before volatility curves flatten below a two-point margin of error.
  • ChatGPT Reddit citations in SurfacedBy's window fell from 37.9% to 3.4%, while Perplexity and Gemini did not mirror that collapse.

Is there one AI citation ecosystem across engines?

No. Intender's Australian study produced 27,924 citations from 3,600 searches and found only 8% source overlap across five engines. Of 4,625 unique domains, 71% appeared on only one engine and just 65 appeared in all five. Cross-engine source persistence metrics matter more than a blended AI visibility score.

How does question type change Google AI Mode citations?

Nicholas Sitter's expanded Google AI Mode study (1,868 captures plus a 1,496-capture replicate) found entity-wrapped citations in 87.6% of "best cafes in Berlin"-style recommendation prompts, 19.8% of explanatory "why so many cafes" prompts, and 0% of generic informational queries. Prompt intent can change citation architecture itself. Paris points to SignalForge tagging inside GEOforge by buyer stage and intent so bottom-funnel and top-funnel prompts are reported separately.

What comes after recommendation in the AI shopping funnel?

querying.ai sells production-surface collection across seven AI interfaces from $20 per month for 10,000 credits, arguing live-surface answers differ from model API pulls. Is My Brand in AI's Claude API web search test (directional: API tool, three runs) found named-brand sites retrieved in only 33% of run-brand pairs, with citation following retrieval only part of the time. Similarweb argues historical AI answer archives are a structural data moat because re-asking today cannot reconstruct yesterday's answer. Mastercard forecasts that more than 300 million online shoppers could routinely use AI agents to shop and pay by 2030, with teens more willing than parents to trust fully AI-run shopping assistants.

Notable moments

00:35. Semrush/Exploding Topics: 57.5% of AI users said a chatbot caused them not to buy.

02:38. Latent Space: Kysely mentioned 42/42, first choice once.

05:02. SurfacedBy: 80% inconsistent combinations; single check 72.2% majority agreement.

08:14. Intender: only 8% source overlap across five engines.

16:00. Mastercard: 300M+ shoppers could use AI agents to shop and pay by 2030.

The throughline is commercial role, not mere presence: AI can veto a purchase, rank mentions without choosing a winner, and diverge sharply by engine and question type on the way to agentic transactions.

Full transcript

Hi, everybody, and welcome to The GEO Show, episode 23. Today, I will attempt to screen share my top stories here on the screen so that you can follow along more closely. The GEO Show is brought to you by GEOforge, full self-driving for AI visibility. Let's get into today's top stories.

First up, AI is becoming a purchase gatekeeper, including talking buyers out of purchases. According to Semrush, a vendor-sponsored consumer research report from Semrush and Exploding Topics published on September 7th surveyed two thousand three hundred and thirty-eight US adults. Among the respondents who use AI, fifty-seven and a half percent said chatbot information had caused them not to buy something.

Sixty-five percent said AI had at least partially replaced Google for product research. Among weekly AI users, seventy-three point six percent reported buying something based on an organic AI recommendation. The margin of error here was plus or minus two percentage points.

So what are we seeing here? GEO is increasingly about purchase exclusion as much as it is about discovery, and being negatively characterized can be commercially worse than not appearing at all. So we're moving from purely just visibility into reputation and even buyer risk, and I think this is going to become a more and more important aspect of the AI visibility funnel: what happens after you're visible. Are you getting actually recommended positively, or is AI actually talking people out of buying your product? This is something that never happened in SEO.

Second story. Latent Space demonstrates that mentioned and chosen are completely different metrics. So Latent Space's Frontier AI visibility tracker, which was updated yesterday on September seventh, contains six thousand seven hundred and sixty-two saved answers across one hundred and sixty-one product categories, seven model configurations, and six question frames.

Several products appeared in virtually every answer but were almost never the first recommendation. Kysely, spelled K-Y-S-E-L-Y, for example, appeared in forty-two out of forty-two answers, but it was the first choice only once. The tracker also found different category leaders between ChatGPT 5.6 Sol and ChatGPT 6 Astra in thirty-three out of a hundred and twenty-one comparable categories.

So this is the first study where we have now seen some results from GPT 6 Astra, and it differs quite a lot from GPT 5.6 Sol. That is interesting. But the headline here is that there really are no rankings in GEO. The share of voice based on mentions can substantially overstate commercial visibility, and the recommendation role matters.

So I think we have to be clear here to distinguish being the first choice from being a recommended alternative or a neutral mention, sometimes even a conditional mention or even an objection. So these are all part of an emerging new trend within GEO, which is called sentiment analysis, and it's related to the first story as well. This is new for people who are transitioning from classic SEO to GEO because SEO just presented you the links in an objective way, and it was up to you to dig into those links and, let's say, read the reviews or read the articles and pages behind those links.

But now AI is proactively presenting brands favorably or negatively or sometimes in a neutral way. So this is important to track and to measure now.

Next story. Fresh repeat run study shows one AI check can misclassify visibility. This is something I have been preaching for months here. Surfaced by analyzed five hundred and thirty-six platform brand question combinations over sixty days, requiring at least five checks per combination.

Of the hundred and fifty-three combinations where a brand appeared at least once, eighty percent were inconsistent across checks. On those inconsistent combinations, a single check agreed with the majority outcome only seventy-two point two percent of the time. The authors explicitly described the sample as directional, not a universal estimate.

All right. Again, single response GEO monitoring is weak. It is basically a coin toss, and most of the tracking tools out there do it that way. And here is another study that did five checks per combination or per prompt, and even then they saw a massive amount of inconsistency.

In our experience, you need to get up to around thirty to forty runs or thirty to forty checks per prompt before you can see that the volatility curves start to flatten and where you can get a margin of error below two percentage points. But running once or even five times is simply not enough to have any meaningful, trustworthy AI visibility data.

On to the next story. ChatGPT's Reddit citation collapse in one controlled benchmark, while Gemini and Perplexity did not. This is by Surfaced by again, and they compared fixed brand and question cohorts across two 14-day windows. In its sample, ChatGPT answers citing Reddit fell from 37.9% to 3.4%. Perplexity moved from 50% to 44.7%, while Gemini moved from 30.8% up to 33.7%.

So that's interesting because the highly reported Reddit citation collapse is quite nuanced. I think it is most pronounced within ChatGPT, but you can see it's not that pronounced within Perplexity and within Gemini; actually, its share has increased according to this study. So off-site GEO strategies can lose effectiveness very abruptly when you're measuring only one engine without making changes on any others. So another reminder here that off-site recommendations should be engine-specific and evidence-based.

Next story. A twenty-seven thousand nine hundred and twenty-four citation study finds only eight percent source overlap across five AI engines. An Australian agency called Intender ran thirty-six hundred searches covering eighty buying topics, three phrasings, five engines, and multiple measurement dates, producing just shy of twenty-eight thousand citations.

Only eight percent of the cited sources overlapped across all five platforms. The platforms are not named here, but I'm assuming that's going to be ChatGPT, Google's AI Overviews and AI Mode, Gemini, Perplexity, and perhaps Claude. So of four thousand six hundred and twenty-five unique domains, seventy-one percent appeared on only one engine, and just sixty-five domains out of the four thousand six hundred total domains appeared in all five.

It's a tiny fraction. So really, there is no AI citation ecosystem. All the engines are very, very different. And again, this theme just keeps recurring throughout this podcast. Now you really need to check cross-engine source persistent metrics and not assume broad AI visibility across all models because the models are behaving very, very differently.

Next story. Google AI Mode citation behavior may depend heavily on the type of question. This is from Nicholas Sitter. Nicholas Sitter expanded an earlier Google AI Mode study into one thousand eight hundred and sixty-eight captures plus a one thousand four hundred and ninety-six capture replicate spanning forty-four Google Places categories, ten question forms, and four cities.

Entity wrapped citations appeared in eighty-seven point six percent of citations for best cafes in Berlin style recommendation prompts versus just nineteen point eight percent for why are there so many cafes in Berlin, and zero in generic informational queries. So the prompt intent can change not merely the answer, but Google's citation architecture itself.

So one piece of advice and one thing that we have done within SignalForge, inside of GEOforge, is that we have a tagging system that allows a user to tag each prompt by buyer stage and by intent, because it's very different when you have a prompt that is bottom of the funnel with high purchase intent or conversion intent versus a prompt at the top of the funnel with simply educational or informational intent. And it's very important, I think, to segment and report on these buyer stage prompts separately because they differ quite a bit.

Next up, querying.ai exposes production AI search collection as a cheap API. So querying.ai markets itself explicitly as the API for GEO. It retrieves real customer surface responses rather than standard model API output across seven surfaces. That includes ChatGPT, Gemini, Perplexity, Google AI Mode and AI Overviews, and Naver. I've never heard of that one. The responses can include fan-out search queries, citations, shopping cards, model identification, and source influence endpoint mapping answer claims to support passages. The pricing begins at twenty dollars a month for ten thousand credits.

This is something that I want to dig deeper into because the big difference here is that this tool is not using the APIs, but it's rather surfacing the real user's queries that the users would see on their actual screens. So it's raw collection from production AI interfaces rather than pulling from the model's APIs.

And what we've seen in our testing is that you get very different results when you simulate an actual user's prompt on a screen versus pulling that same prompt response from an API. And most of the tracking tools that are out there now are pulling from the model's APIs, and I think that presents a relatively misleading AI visibility picture.

Next story. A Claude study exposes retrieval as a separate failure point from citations. So a company called Is My Brand in AI ran 50 B2B SaaS buyer questions three times each, not enough, against Claude Sonnet 5's API web search tool on September 7th. So they're using the API tool here. When a query explicitly named a brand, that brand's own site appeared in the retrieved set in only thirty-six out of one hundred and eight run brand pairs, which is thirty-three percent.

Once domains were retrieved, sixty-two point seven percent were cited on average. Repeated runs cited domain sets had mean Jaccard overlap of zero point five five five. I don't know what Jaccard overlap is, probably a statistical term. But the author clearly cautions that the test used Claude's API web search tool, not the consumer Claude.ai interface, and only three runs per prompt.

So two weak aspects of this report: using Claude's API instead of the Claude.ai interface, and only doing three runs per prompt is not going to reduce the volatility enough to get you a reasonable level of confidence or a reasonably low margin of error. So I think it's a shaky study, but what it is finding is that the mentions and citations are not highly correlated here and those need to be tracked separately.

Okay, our next story is about Similarweb. Similarweb argues that historical AI visibility is becoming a structural data moat, and I agree very much with this. Similarweb published a September sixth comparison of historical AI search datasets. It says that its own sampled brand market AI data extends to January twenty-twenty-four, while Comscore's broader generative search collection dates to May of twenty-twenty-three.

The core methodological point is sound, independent of vendor rankings. If nobody sampled an AI answer when it happened, asking the same prompt today does not reconstruct that historical answer. So historical depth will increasingly separate the data companies from the prompt trackers.

And historical data is becoming a competitive advantage. It's actually becoming a structural data moat. So it's important to always preserve the historical ranking so you get that long trend line for your AI visibility. And that data itself becomes something that you can't replicate easily, so that becomes a data moat.

The last story comes to us from Mastercard. Mastercard today published a report forecasting that more than three hundred million online shoppers could routinely use AI agents to shop and pay by twenty thirty. This is a futurist forecast, not current adoption. The supporting customer research surveyed twenty-six thousand parents and teenagers across thirteen European markets in June through July of twenty twenty-six.

Twenty-seven percent of teenagers said they were likely to use a fully AI-run shopping assistant versus just sixteen percent of parents. And thirty-one percent said they would trust an AI product recommendation over a friend's recommendation. Very interesting. AI agentic shopping is definitely coming. It's just a question of time, and I think now we need to look more holistically at this entire funnel that right now is focused a lot on visibility, but as we've talked about in this episode, it's moving towards recommendation, and then it's going to further go on towards agent selection and transaction.

So we need to start to think about websites that are agent-ready. That means that they are machine-readable, their claims are machine-readable with verifiable evidence, clear and transparent pricing, product and service facts, and structured information that an agent needs to qualify a vendor. This is the continuum, the full spectrum from getting discovered to being actually purchased by an AI agent and everything in between.

All right. That wraps it up for episode twenty-three. Thank you for listening, and we'll see you in the next one.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.