Four AI engines agree on the number one recommendation only 4% of the time, and a blended GEO score can rise while ChatGPT, Google AI Mode, and Google AI Overviews contradict each other. On episode 25 of The GEO Show, Paris Childress covers a sixteen-week GEOforge graded-answer case study, Yext engine consensus, Semrush mention-citation-traffic layers, fashion treating AI visibility as PR, Claude's B2B Similarweb audience, Similarweb inside Perplexity Computer, AirOps GTM repositioning, Cited Index's 114 tools with 97 monitoring-only, Surfer MCP, and Surva.ai's $39 full-funnel bundle.
Four AI engines agree on the number one recommendation only 4% of the time. On episode 25 of The GEO Show, Paris Childress argues that blended GEO scores and single-engine dashboards conceal that fragmentation, then walks ten stories spanning graded answers, manufacturing traffic, fashion PR, proprietary datasets, and a tool market that mostly monitors without executing.
The episode is the show's first with slides. The throughline is engine divergence: ChatGPT, Google AI Mode, Google AI Overviews, Claude, Gemini, and Perplexity do not share one ranking reality, and brands that optimize to a single blend will misread risk.
GEOforge graded AI answers over sixteen weeks with thirty buyer prompts across thirteen categories, each run forty times against ChatGPT, Google AI Mode, and Google AI Overviews from July 5 through September 7. Share of voice moved differently on each surface while the blended score rose. Engine, topic, and recommendation content change the verdict, so topic-level and engine-level diagnosis belong in product and category positioning.
We see all the time with clients that the blended score may say up or down, but you need to dive deeper into each individual AI engine.
Yext's research brief finds agreement on the number one recommendation only 4% of the time across four engines. OpenAI leaned more on directories; Anthropic used review platforms more frequently. Those patterns are observational associations, not causal ranking proof. Explicit engine consensus and divergence scores matter because blended share of voice conceals where platforms disagree.
No. Semrush manufacturing data shows Google AI Overviews rising from 38% to 57% of tracked search volume from January to July while AI Mode and AI assistance still generated only a small percentage of sessions. More people read the answer and do not click through. Measure mention and citation, then recommendation, visit drop-off, and revenue impact as separate stages.
Fashion brands via Launchmetrics (reported in Vogue) track ChatGPT, Claude, Gemini, and Perplexity across ten markets and event windows. PR, communications, and brand insights budgets are entering the addressable market because what AI says about a brand often comes from earned media rather than owned pages alone. AirOps' September 9 webinar, titled "AI Search is Bigger than SEO," pushes the same upward GTM move: AI search as a buying channel and decision-influence problem, not only content optimization.
Cited Index catalogs 114 AI visibility and GEO products; 97 are monitoring-only. Surfer MCP makes an SEO workspace callable from ChatGPT and Claude. Surva.ai ships monitoring, content, publishing, and citation analysis from $39/month across five platforms, with WordPress and Webflow publish paths on higher tiers. Similarweb's proprietary data live inside Perplexity Computer shows answers drawing from connected datasets, not only public web retrieval. Claude's Similarweb audience also skews toward B2B ICPs, with mobile growth far faster than ChatGPT's, so absolute audience size understates Claude for SaaS and developer tools.
00:25. GEOforge case study: blended score rose while three engines diverged.
03:22. Yext: four engines agree on #1 only 4% of the time.
05:28. Semrush: AI Overviews 38% to 57% of manufacturing search volume; thin session traffic.
07:42. Launchmetrics / Vogue: AI visibility as a fashion PR metric.
15:24. Cited Index: 114 tools, 97 monitoring-only.
19:48. Surva.ai at $39/month with monitoring, content, and publishing.
The throughline is engine fragmentation: treat consensus and divergence as first-class metrics, and stop trusting a single blended GEO score as strategy.
Paris: Hi, everybody, and welcome back to another exciting episode of The GEO Show, brought to you by GEOforge, full self-driving for AI visibility. I'm your host, Paris Childress. Today, we are going to experiment with a slightly different format. I have prepared some visual slides, which I'll run through as we go through the stories to hopefully liven things up a bit.
Paris: And thank you for that YouTube comment on making these podcasts a little bit more visual. So we've got 10 great stories. Let's get right into it, starting with a case study of our own from GEOforge. One blended GEO score hid three contradictory engine outcomes. GEOforge's new case study covers graded AI answers over a sixteen-week period.
Paris: It uses thirty buyer prompts across thirteen categories. Each one ran forty times against ChatGPT, Google AI Mode, and Google AI Overviews from July fifth through September seventh. During that period, share of voice changed on Google AI Overviews, ChatGPT, and Google AI Mode. So the blended score actually rose. The study also found a spread between the best and worst-performing topic clusters, a median time from publication to first AI citation, and cited pieces receiving their first citation within fourteen days.
Paris: Engine topic and recommendation content materially change the verdict, so it is very important to turn engine divergence plus topic-level diagnosis into a prominent product and category positioning advantage. We see all the time with clients that the blended score may say up or down, but you need to dive deeper into each individual AI engine.
Paris: So one popular clustering or tagging method would be funnel stage. You can have prompts tagged as top of funnel, middle of funnel, and bottom of funnel, and look to see how prompts are performing at each stage and within different answer engines. There's a whole lot of nuance there. All right, moving on to the next story from Yext. Yext independently finds four AI engines agree on the number one recommendation only four percent of the time. This is coming from a Yext research brief.
Paris: OpenAI relied more heavily on directories, while Anthropic used review platforms much more frequently. These are observational associations, not proof that any source type causes rankings. There is no single AI search ranking reality, and different engines pull different types of citations depending on the situation.
Paris: We really need an explicit engine consensus and engine divergence score. That shows where AI platforms agree and where the blended share of voice conceals risk and nuance. The next story comes to us from Semrush. Semrush data separates AI discovery into mention, citation, and traffic layers.
Paris: Semrush analyzed manufacturing industrial keywords plus clickstream and AI visibility datasets. Google AI Overviews increased from thirty-eight percent to fifty-seven percent of tracked search volume from January to July. Yet AI assistance and AI mode generated only a small percentage of sessions to manufacturing sites.
Paris: More people are reading the answers, getting their answers, and not clicking through. AI visibility is not a one-funnel stage, but a multi-stage funnel. It starts with a mention and a citation, then moves to recommendation, and then to visit, where there is a drop-off, and finally to revenue. It is important to measure that entire journey through all stages, concluding with revenue impact. All right. Let's go to the next story from the fashion industry. Fashion is turning AI visibility into a PR and brand metric.
Paris: Launchmetrics tracks ChatGPT, Claude, Gemini, and Perplexity across ten markets, benchmarks weekly, monitors narratives, and measures how specific events affect AI visibility before and after. This was reported in Vogue magazine itself. GEO is escaping the SEO department. PR, communications, and brand insights budgets are becoming part of the addressable market. GEO is not simply competing for SEO budgets because it is the natural next version of SEO. We can pull in GEO brand budgets and PR communications because it impacts brand positioning.
Paris: What AI says about a brand comes from earned media, and that is what other people are saying about you, not what you are publishing on your own owned properties. Let's move on to story number five, reported by Similarweb. Claude's audience is disproportionately relevant to B2B ICPs.
Paris: It is a huge difference here. Claude's visits grew much faster. Its worldwide mobile audience increased roughly fourteen times versus thirty-one percent for ChatGPT. Similarweb reports that Claude users are more likely than average searchers to visit business service sites and programming sites. It probably started from around December last year with Claude Code suddenly hitting the scene. The absolute audience size of ChatGPT versus Claude underestimates Claude's importance for B2B services like SaaS, cybersecurity, and developer tools.
Paris: Similarweb's proprietary market data is now live inside Perplexity, enabling users to benchmark traffic, audience, and market share through natural language prompts. This is an integration with Perplexity Computer. It is not evidence that Similarweb data now influences every Perplexity answer. The AI answers are increasingly able to draw from connected proprietary datasets, not merely public web retrieval. You are going to have more partnerships where large data providers make partnerships with LLM and AI engines, allowing them to draw their data from datasets instead of doing public web retrieval.
Paris: That is going to make things even more fragmented and it is important that your data is represented accurately and fully in non-web, non-public web data sources. Next story is about AirOps and their latest webinar, which is happening today on September ninth, and its title is AI Search is Bigger than SEO.
Paris: AirOps is explicitly repositioning AI search as something bigger than SEO, not just the next version of SEO, but bigger. Its messaging says AI search is becoming a core buying channel, and content visibility alone no longer explains discovery, citations, or buying decisions. This is fundamentally a go-to-market positioning signal. AirOps is expanding the problem definition upward from content optimization toward the broader buyer journey. We do not want to compete as AI SEO software; we want to position as AI decision influence.
Paris: Cited Index says the AI visibility tool market has already reached one hundred and fourteen indexed products. Cited Index now catalogs AI visibility and GEO tools, including classified products for AI visibility monitoring and citation services. The categories overlap. Ninety-seven out of a hundred and fourteen are classified as AI visibility monitoring. That means they are read-only and reporting on what they see. They are probably not measuring with enough statistical rigor, and it is not helping you actually take action. There is no execution layer.
Paris: Third-party citations influence responses more than a site's own content. A small minority of the category is trying to tackle the most important factor of success, which is citation building. Surfer is making its SEO workspace callable from ChatGPT and Claude through MCP. Surfer is introducing Surfer MCP, connecting its workspace to Claude, ChatGPT, and other MCP-capable assistants. Users should be able to access and act on Surfer data through assistants they already use instead of navigating to the product manually.
Paris: MCP is the real deal. It is going to be how most of us actually use our SaaS stack. The large leaders of SaaS will survive because of MCP, while the mid-tail and long tail will likely be replaced by features and capabilities within the LLMs themselves.
Paris: And now on to our last study, our last story of the day, which is about another new entrant, another low-priced new entrant into this GEO platform category called Surva, S-U-R-V-A.ai. They are bundling monitoring, content, and publishing from just thirty-nine dollars a month.
Paris: It includes thirty prompts, which is the same number that we track across five platforms. We track only three. Unlimited competitors with five deeply tracked. Fifteen AI-generated articles per month. Citation source analysis. They have a citation module such as Site Forge and an AI SEO audit. Surva.ai higher tiers will add crawler logs, CMS auto-publishing, APIs, and agency capabilities. Its AI visibility documentation was updated yesterday on September eighth.
Paris: They are also advertising a WordPress and a Webflow integration, which can automatically publish to a CMS. They are tracking across ChatGPT, Perplexity, Claude, Gemini, and Google AI Overviews. We do not track Perplexity or Claude or Gemini, but we track Google AI Overviews and Google AI Mode in addition to ChatGPT.
Paris: Another new player is always fun to compete with and to see what other players are doing in the space. This baseline bundle of new entrants is no longer just visibility. People are getting into this game with larger feature sets like monitoring and gap analysis, content generation, publishing, and citation analysis. So be interested to see the path that Surva.ai takes. All right, everybody. That'll do it for episode twenty-five. Thanks for tuning in, and we'll see you all in the next one.