Google AIO & AI Mode only overlap 10%

The GEO Show
September 22, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Google AI Overviews and AI Mode remain separate citation ecosystems, with SE Ranking measuring 10.7% URL and 16% domain overlap and Ahrefs finding 13.7% citation overlap. Paris Childress covers agentic search that varies by model (~6X Claude vs ChatGPT), Server.ai's referral and crawler-log expansion, first-party logs that undercut raw crawler counts, Share of Model, open-source citation tooling, and a Catalyst recommendation benchmark.

Key takeaways

  • Claude Sonnet 4.6 invoked search 825 times per 1,000 prompts versus 140 for ChatGPT 5.3 (roughly 6X); used evidence and citations are not the same set.
  • Server.ai now documents AI referral tracking across 13 platforms plus crawler log import and a REST API for visibility, citations, mentions, and competitive data.
  • SE Ranking measured 10.7% URL and 16% domain overlap between Google AI Mode and AI Overviews; Ahrefs found 13.7% citation overlap on a larger paired set.
  • In one publisher's first-party logs, ChatGPT referred 613 human sessions while crawling the site only 55 times, so raw AI crawler counts are a weak proxy for discovery.
  • Share of Model (Monroyia) reframes blended Share of Voice as per-model brand naming rates across buyer questions and stages.
  • Catalyst's 618-answer board showed Peak named in 58.1% of answers versus Profound 15.9% and AirOps 9%, with ~15 percentage points of one-day noise.

Paris Childress hosts episode 40 of The GEO Show on why Google AI Overviews and AI Mode remain separate citation ecosystems, with SE Ranking reporting 10.7% URL and 16% domain overlap and Ahrefs measuring 13.7% citation overlap. The through-line: AI visibility must be measured surface by surface, even inside a single company.

The episode also covers agentic web search that varies by model, Server.ai's move into referrals and crawler logs, first-party logs that undercut raw crawler counts, Share of Model as a KPI, an open-source Python package for citation measurement on PyPI, and a Catalyst recommendation benchmark where Peak outpaces Profound and AirOps.

Next story, which I think is probably our top story of the day.

Are Google AI Overviews and AI Mode the same citation surface?

No. A September 20 analysis highlighted persistent low overlap. SE Ranking found 10.7% URL overlap and 16% domain overlap between AI Mode and AI Overviews. Ahrefs independently measured 13.7% citation overlap across a larger paired response set. Paris notes that SignalForge inside GEOforge already separates the two, because blended Google AI numbers hide surface-level differences.

How much does agentic search behavior vary by model?

In a controlled 1,000-prompt experiment summarized with academic research on ChatGPT, Claude, Grok, and DeepSeek, Claude Sonnet 4.6 invoked search 825 times versus 140 for ChatGPT 5.3, roughly a 6X gap. More frequent search did not reliably mean better answers, and retrieved URLs used as evidence were not always the same as cited URLs.

Why are raw AI crawler counts a poor proxy?

E-commerce Fast Lane analyzed 14 days of server logs totaling 49,559 crawler visits, with roughly 100 from OpenAI, Anthropic, and Perplexity bots. ChatGPT referred 613 human sessions while crawling the site only 55 times. Assistants can retrieve through conventional search indices, so crawler frequency alone should not equal visibility opportunity.

Do commercial leaders always win AI recommendations?

Not in Catalyst's September 20 leaderboard of 618 answers across 12 buyer questions and five platforms. Seven-day averages named Peak in 58.1% of answers, Profound in 15.9%, and AirOps in 9%. Catalyst flagged roughly 15 percentage points of one-day noise. Category revenue leadership and recommendation share can diverge.

Notable moments

  • Agentic search: Claude 825 vs ChatGPT 140 searches per 1,000 prompts.
  • Server.ai: referrals across 13 platforms plus crawler logs and API.
  • Google split: 10.7% URL / 16% domain / 13.7% citation overlap.
  • Fast Lane logs: 55 ChatGPT crawls vs 613 referrals.
  • Share of Model: model-level KPI over blended Share of Voice.
  • Catalyst: Peak 58.1% vs Profound 15.9% vs AirOps 9%.

Watch the full episode on YouTube: https://www.youtube.com/watch?v=pp6AzglZgFs. Audio: https://share.transistor.fm/s/61cadbff.

Full transcript

[00:00:00] **Paris Childress-1:** Hi, everybody. Welcome back to another episode of "The GEO Show," brought to you by GEO Forge: Full Self-Driving for AI Visibility. I'm your host, Parrish Childress, recording episode 40 on September 21st, 2026, and we have some interesting stories to run through today. Let's get right to the very first one, which is about agentic web search behavior.

[00:00:31] Academic research shows agentic search behavior varies radically by model. Researchers studying ChatGPT, Claude, Grok, and DeepSeek found substantial differences in when agents search, how they formulate queries, which domains their platform search engines return, and how retrieved evidence becomes final answers.

[00:00:57] More frequent search did not reliably mean better answers, and some claims relied on retrieved URLs that were never cited. In a controlled 1,000 prompt experiment summarized alongside the paper, Claude Sonnet 4.6 invoked search 825 times versus just 140 times for ChatGPT 5.3, a roughly 6X difference. The wider study included real-world interaction traces from more than 600 users.

[00:01:32] So let's unpack this. There's a lot going on here. The overall takeaway really is that There, the models, there is no reliable data here on how models are retrieving their answers, and it varies radically by engine, which we've talked about m-many times on this show. So in this controlled experiment with a thousand prompts, we saw that Claude Sonnet is evoking search, meaning it's going out and doing a real-time search 82 and a half percent of the time.

[00:02:10] ChatGPT is only doing it 14% of the time. That's a massive difference, and that to me signals that perhaps Claude has a shallower knowledge base of pre-training data, meaning that it's, it's forced to go out and do real-time web searches more often because it does not have a ready answer to, to dish out in real time, so it, it evokes search.

[00:02:40] Whereas ChatGPT has a deeper set of training data to pull from, and it relies less on real-time search. So we should think about this long-term measurement object in terms of the retrieval life cycle, not just the final answer. You have search invocation with queries, you have return domains, there's used evidence, there's cited evidence, which here is not necessarily the same thing, and then there's an answer.

[00:03:14] Another interesting nugget from this study is that sometimes the used evidence-- S-sometimes the citations were not in the used evidence and vice versa. So the, it doesn't necessarily mean that the citations are a strict set of sources used in the answer Next story. Server.ai expands from monitoring into referrals, crawler logs, and API access. This is updated on September twentieth. Server.ai now documents AI referral tracking across 13 platforms, AI crawler logs, citation tracking, competitor monitoring, and REST API exposing visibility, citations, mentions, and competitive data.

[00:04:05] Its crawler log import, its crawler log import identifies AI requests from server logs using crawler pattern user agents. So this stack with this company is getting more and more technical. It's moving from just prompts and citations and traffic into... Well, it's moving from prompt tracking and citation tracking into referral traffic tracking, crawler logs, and tracking how the crawlers are digesting the site.

[00:04:44] And with this API, it's ra- the API is rapidly becoming a commodity functionality even amongst smaller vendors. So I think what that means for this space is that we need to start really paying attention more to verified crawler identity and repeated sampling, confidence intervals in our reporting source influence, proprietary knowledge, causal diagnosis, and execution All right.

[00:05:16] Next story, which I think is probably our top story of the day. Google AI Overviews and AI Mode remain separate citation ecosystems, and this is debunking a very popular and common sense belief that AI Overviews and AI Mode are essentially the same thing. But in fact, a September 20th brand-cited analysis highlighted the persistent low overlap between Google's two AI surfaces.

[00:05:46] SE Ranking's underlying study found just a 10.7% URL overlap and a 16% domain overlap between AI Mode and AI Overviews. Ahrefs independently measured a 13.7% citation overlap across a much larger paired response data set. So even within Google, AI visibility is not one surface, and we should keep in mind that we really need to measure platform by platform, and that means even separating Google AI Overviews and Google AI Mode, which is what we do in SignalForge within GEOForge.

[00:06:28] And it does make me wonder if AI Overviews is essentially a bridge to Google AI Mode, which is gonna become the default Google experience at some point, probably in the near future, then why aren't they returning essentially the same answers with the same citation sets? I think right now that Gemini is what is p-powering Google AI Mode, and there is a different system and more of a, I guess, the ba-- the traditional search infrastructure is powering Google AI Mode.

[00:07:04] So I'm not really sure how to explain that other than these are the facts. These are still very different surfaces. All right, onto the next story. First-party logs challenge the value of raw AI crawler counts. E-commerce Fast Lane analyzed 14 days of its own server logs, 49,559 crawler visits. Roughly 100 of them were from OpenAI, Anthropic, and Perplexity bots.

[00:07:39] Yet ChatGPT referred 613 human sessions. ChatGPT itself crawled the site only 55 times. The publisher explicit-explicitly cautions that its site is not a Shopify store, and the ratios should be treated as directional So what is this saying? The, the direct AI bot traffic really could be a very poor proxy for actual AI discovery because the AI assistants can retrieve answers through conventional search in- indices.

[00:08:19] So here, this, this debunks what we've been looking at previously, I, I think. This is almost 50,000 crawler visits

[00:08:28] that were reported in server logs, a much smaller number coming from the AI crawler accounts, and ChatGPT, which referred six hundred and thirteen human sessions itself only crawled the site fifty-five times. So I don't think we should by default equate crawler frequency, AI crawler frequency with visibility opportunity.

[00:08:56] Um, I think a lot of people are making that conclusion now, but we shouldn't rush to that conclusion. It's more important really to m-mathematically correlate the bot activity with index eligibility and indexation, citation, appearance, answer appearance, and referrals before assigning any meaning. All right.

[00:09:23] Next story. Share of Model emerges as a competing measurement vocabulary. Monroyia's current dataset covers twenty thousand nine hundred and ninety-six buyer questions, one hundred and forty-three thousand two hundred and ninety-eight cited sources, forty-one thousand three hundred and ninety-seven distinct domains, and four models through September twenty-first.

[00:09:47] It defines share of model as the percentage of sampled answers naming a brand and recommends breaking it down by model and by buyer stage. Monroyia also claims vendor-owned pages account for well under one percent of sources in its observed dataset. That is vendor-owned research and that is-- This is vendor-owned research and requires independent replication.

[00:10:14] Okay. So a little bit more shifting of KPI terminology. The most popular KPI term right now to describe measuring AI visibility is this share of voice. It's also the one that we're using right now. But as I'm seeing more and more evidence that each model is very, very different in its, its brand mention activity and its citation rates, I do think share of model is actually better because it implies that you do need to measure each model separately.

[00:10:49] So share of voice implies a blended aggregate share of voice across all of the AI engines like ChatGPT and Google AI Overviews and Gemini, Claude, Perplexity, et cetera. Whereas share of model implies what is the specific share of voice, let's say, within ChatGPT or within Gemini. And I think share of model is actually, I hope, a term that is eventually gonna overtake share of voice.

[00:11:19] All right, on to the next story AEO is becoming an open source Python primitive. What does this mean? This is reported by PyPI, and it says that version zero point two nine point zero of the open source AEO Python package was released yesterday on September 20th. It describes itself as software for measuring and improving citations across ChatGPT, Google AI Overviews/AI Mode, Perplexity, and Claude.

[00:11:58] The project is explicitly labeled an early release with an API still evolving. So now AEO is, is getting built into open source Python. Basic AE- basic AI search measurement is moving from a SaaS feature to a developer building block. So just becoming, uh, more and more commoditized, I would say. All right, onto our last story of the day Peak materially outperforms Profound and AirOps in one live AI recommendation benchmark Catalyst's AI search visibility leaderboard measured six hundred and eighteen answers across twelve buyer questions and five platforms through September twentieth.

[00:12:49] Its seven-day averages show Peak named in fifty-eight point one percent of answers, Profound in just fi- uh, fifteen point nine percent, and AirOps in only nine percent. Catalyst explicitly notes that sample is small and that one-day results carry a ru- roughly fifteen percentage points of noise, or what we refer to as the margin of error.

[00:13:18] And fifteen percentage points margin of error is quite large, which tells me that they did not really do enough runs for it to be statistically significant. But anyway, it is still interesting that the market leader by revenue and, of course, by valuation is Profound, and they're way out ahead in terms of market share.

[00:13:42] And that was reported recently by Ramp. But in this study, when you look at the percentage of answers from an AI platform when people are asking about buying in this category, uh, Peak is recommended much more often. So Peak seems to have the recommend-- the same recommendation rate or, I mean, relative to Profound Peak has a much larger share of recommendation rate, of course, but not the same market share.

[00:14:12] So it is interesting here that these two things are not always aligned. Commercial traction and AI recommendation or share of voice are not the same thing. A category leader can still be weakly recommended by the systems that are influencing future buyers. All right. That wraps up episode forty.

[00:14:33] Thank you all for listening, and thank you to those listeners who have hung with me so far. I hope this has been valuable. If you do like this podcast, please go onto Apple or Spotify and give it a five-star rating. That will help me very much. And until next time, uh, we'll see you all on the next episode.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.