Low Confidence AI Visibility Scores are Plaguing GEO

The GEO Show
September 12, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Low-confidence AI visibility scores from tiny prompt sets and one-run sampling are plaguing GEO. On episode 29 of The GEO Show, Paris Childress covers Otterly's evidence that 10→100 prompts cut coverage variation ~73%, Peak full-funnel attribution, Ahrefs Topics, Rank Masters citation concentration, Brand Ghost engine disagreement, Pallas AI, M11 Labs proof ranking, Instacart Clementine, Google AI Max, and Elmo open-source tracking.

Key takeaways

  • Otterly: moving from 10 to 100 prompts cut brand-coverage variation about 73% across 252,407 answers / 520 prompts / 7 engines; one-run sampling still leaves day-to-day volatility.
  • GEOforge bar for statistical significance is about 30-40 runs per prompt for 95% confidence / +/-2% margin of error; cheap tools often price on one-run depth.
  • Rank Masters B2B case: 5 URLs produced 52% of 6,831 owned AI citations; 20 URLs produced 90% (Gemini 64% / Perplexity 20% / ChatGPT 15%).
  • Brand Ghost: recommendation overlap averaged 11.2% across four engines; 50.7% of comparable sets shared no brands; blended SOV can hide engine risk.
  • Peak wires GA4 into AI referrals and page-level crawl to cite to visit to conversion; Ahrefs Topics groups prompts by parent topic for owns/misses.
  • Prompt tracking is becoming commodity (Elmo open-source); moat sits in proprietary knowledge, statistical certainty, decision intelligence, execution, and off-site citation building.

Low-confidence AI visibility scores are plaguing Generative Engine Optimization. On episode 29 of The GEO Show, Paris Childress walks through independent Otterly evidence that tiny prompt sets and one-run sampling create false confidence: moving from 10 to 100 prompts cut brand-coverage variation about 73% across seven engines, while GEOforge still targets about 30-40 runs for statistical significance.

Around that throughline he covers Peak full-funnel attribution, Ahrefs Brand Radar Topics, Rank Masters citation concentration, Pallas AI optimization loop, Brand Ghost engine disagreement, M11 Labs on agentic commerce ranking proof, Instacart Clementine recommendation-to-cart, Google AI Max versus dynamic search ads, and Elmo open-source GEO tracking.

Why do small prompt sets fake AI visibility?

Because coverage estimates swing wildly when the sample is tiny. Otterly analyzed 252,407 answers across 520 prompts and seven engines. Ten to one hundred prompts reduced variation in reported brand coverage by about 73%. Forty-eight percent of distinct cited URLs appeared only once, while the top one percent generated forty-four percent of citations.

Shallow sampling produces false confidence.
  • Otterly still used one run per prompt, per engine, per day.
  • That measures day-to-day volatility, not within-session variance.
  • GEOforge experience: about 30-40 runs for 95% confidence / +/-2% margin of error.

Where is citation value concentrating?

A Rank Masters B2B study of 6,831 owned AI citations across 64 commercial prompts over 90 days found five URLs produced 52% of citations and twenty URLs produced 90%. Gemini supplied 64%, Perplexity 20%, and ChatGPT 15%. Paris notes the single-site caveat and still reads a power-law pattern: a small set of canonical pages can drive disproportionate GEO value.

Why is blended share of voice dangerous now?

Brand Ghost September 10 run across 7,473 citations found about 12.9% cross-engine source overlap, 11.2% average recommendation overlap, and 50.7% of comparable sets sharing no brands at all. A GEOforge client case showed AI Overviews up about 50% while ChatGPT and Google AI Mode were slightly negative. Unlike SEO Google monopoly, ChatGPT versus Google AI Mode and AI Overviews must be split.

What sits above commodity prompt tracking?

Peak is wiring GA4 into AI referrals and page-level crawl-to-conversion timelines. Ahrefs is adding topic-aware Brand Radar reporting. Pallas AI launched with fetchable/chosen/extractable gates and a watch-decide-act-review loop. Elmo MIT self-hostable tracker shows basic monitoring is becoming open infrastructure. Paris argues the moat is proprietary knowledge, statistical certainty, decision intelligence, execution, and off-site citation building.

Where is agentic commerce heading?

M11 Labs forthcoming skincare evidence study claims 47% of brands lacked independently verifiable evidence and only one in twenty evidence links confirmed the judged claim. Instacart Clementine turns conversation, recipe, or list into a ready-to-buy cart, with possible no referral click between recommendation and transaction. Google is auto-migrating parts of DSA into AI Max with about +7% conversions in internal non-retail data.

Notable moments

00:15. Peak AI Referrals + My Website: full-funnel attribution beyond visibility dashboards.

03:33. Throughline: Otterly 10 to 100 prompts cut coverage variation ~73%; ~30-40 runs for significance.

05:36. Rank Masters: 5 URLs = 52% of owned citations.

08:30. Brand Ghost: 11.2% recommendation overlap; 50.7% share no brands.

12:07. Instacart Clementine: recommendation becomes cart; agentic commerce prep.

The throughline is measurement quality itself: until AI visibility sampling is deep enough for statistical confidence, GEO teams risk optimizing noise dressed up as a score.

Full transcript

Hi, everybody, and welcome back to another episode of The GEO Show, brought to you by GEOforge, full self-driving for AI visibility. I'm your host, Paris Childress, and it's the weekend, and we've got some great stories on September 12th. So why don't we get right into it?

First up, Peak closes the gap from AI visibility to revenue attribution. Peak launched AI Referrals on September 11th, connecting Google Analytics 4, GA4, to show sessions, engagement, conversions, revenue, landing pages, and assistance. It simultaneously launched My Website, combining server logs, prompt tracking, and GA4 at the page level from crawl to source usage to visit and conversion.

Peak is on the move. Peak started as a visibility platform like most, and now they are layering on much more of the full funnel here. The GA4 integration is pretty big. By the way, we also have a GA4 integration with GEOforge because this is actually what allows us to have a closer look at AI referral traffic and the behavior of that traffic. So visibility-only dashboards are becoming insufficient, really. Full funnel attribution is where the game is moving towards. That means looking at how AI bots are crawling, retrieving, citing, and then recommending, referring, and finally the conversion in one full timeline.

All right, let's move to the next story. Ahrefs is turning Brand Radar into a programmable topic-aware intelligence layer. Ahrefs shipped a Brand Radar AI response API on September 8th, entities and upgraded exports on September 10th, Report Builder integration on September 11th, and today added a Topics report grouping queries by parent topic to expose angles a brand owns or misses.

So this is similar to the way they've been grouping keywords into parent topics. They're now doing this with queries and prompts. So Ahrefs, I would consider them an enterprise SEO incumbent. They are rapidly shifting over to GEO. I think it's existential for them to do so, and they're absorbing more standalone GEO monitoring functionality and borrowing some of their methodology from their legacy SEO tool, such as grouping queries into parent topics. So that's a nice, interesting feature layer that they've added.

Moving on. Otterly provides strong independent evidence that small prompt sets are noisy. Otterly analyzed 252,407 answers covering 520 prompts across seven engines. Moving from ten to one hundred prompts reduced variation in reported brand coverage by about seventy-three percent across engines. Forty-eight percent of distinct cited URLs appeared only once, while the top one percent generated forty-four percent of all citations. So its methodology still used only one run per prompt, per engine, per day, so it measures day-to-day volatility rather than within-session probabilistic variance.

And I think this really strengthens our own argument at GEOforge that shallow sampling produces false confidence. And one of the reasons why Otterly and so many other tools can maintain such low price points is because they just do one run per prompt, and sometimes they do that every day. But in order to reach statistical significance, and really statistical significance within one session, one run is definitely not enough. You have massive variance, and our experience shows that you need to do upwards of thirty or even forty runs before you get to ninety-five percent confidence, or stated another way, a plus or minus two percent margin of error. So they went from ten to a hundred prompts, although that was done over a long period of time, but that reduced variation by seventy-three percent.

All right, let's move on to the next story. A fresh B2B case study shows citation performance is highly concentrated. Rank Masters published a September 11 study covering 6,831 owned AI citations across 64 commercial prompts over 90 days. Five URLs produced 52% of citations, and 20 URLs produced 90%. Gemini supplied 64% of the measured citations, Perplexity 20%, and ChatGPT just 15%. That's quite interesting. Very concentrated here.

The authors explicitly caution that this is only one site in one B2B services category, so it's not a universal benchmark, but it is one of the starkest examples we've seen yet of concentration of citations among a small number of URLs. So GEO content may exhibit a power law effect, a small number of canonical pages producing disproportionate value. And we don't know more on this other than, for example, I don't know if this was impacted by maybe the domain authority or what types of pages. But it is interesting behavior of serious concentration of citations in a small number of pages.

Moving to the next story, Pallas AI launches with an autonomous GEO architecture very close to other tools like GEOforge. This is a new entrant, Pallas AI. They've launched September 11th around nine-engine monitoring, Shopify integration, a unified marketing context OS, three visibility gates which are fetchable, chosen, and extractable. That's interesting. And an autonomous watch, decide, act, and review optimization loop.

I do like the notion of an optimization loop. Theirs is of course watching, which probably means measurement. Deciding, which is obvious. Acting probably means either generating content for publishing on an owned property or acting outside offsite with some sort of offsite citation action, and then reviewing, meaning just remeasuring and observing what happened as a result of that intervention. It's a very smart loop, I think.

All right. Let's move on to the next story, which is about Brand Ghost. Brand Ghost's latest run shows recommendation agreement remains extremely low. Brand Ghost's September 10th observatory run analyzed 7,473 citations across four engines. Cross-engine source overlap was about 12.9%. Average recommendation overlap was just 11.2%. And 50.7% of comparable recommendation sets shared no brands at all.

All right. We have seen this so many times now. Different LLM engines, different AI engines have very, very different results. So any sort of blended share of voice reporting can really hide a major engine-specific risk. We've just published a case study on our own site recently showing that one of our clients had a big jump in visibility and share of voice for Google AI Overviews and something like a fifty percent increase. And on ChatGPT, and interestingly on AI Mode, Google AI Mode, they were actually slightly negative.

And the overall blended share of voice, of course, was still looking very positive, but it was all entirely driven by Google AI Overviews. And it's important that we all look deeper so that we can tell those stories about what AI engines are driving either growth or which ones are disproportionately weighing down on the blended picture. And it's very different than SEO because SEO, we only really cared about Google because they were so dominant, and we only really needed to report on Google rankings, Google organic traffic. That's really all that mattered.

But this is very, very different now. If nothing else, we need to look at ChatGPT versus Google's properties of AI Mode and AI Overviews. All right, moving on to the next story. M11 Labs says that agentic commerce will rank proof, not just brand presence. M11 Labs emerged from stealth on September 9th with an agentic trust platform and marketplace. Its forthcoming skincare study analyzed 600-plus brands and 15,000-plus evidence captures. M11 says that 47% lacked independently verifiable evidence. 58% of claims referenced research that was not publicly accessible, and only one in 20 associated evidence links directly confirmed the judged claim.

All right. The full report is not due until September 16th, so we'll hopefully be able to dive a little bit deeper. What's happening here? AI visibility could really evolve more towards machine-verifiable brand claims, and that's, I believe actually that AI visibility is just the very beginning or the early innings of GEO, and reporting on AI visibility is just the starting point. Things are gonna get much more complex, and we're gonna look back on AI visibility ultimately as a vanity metric, the same way that we used to look back on rank tracking as a vanity metric for SEO.

On to the next story. Instacart turns an AI recommendation directly into a transaction. Instacart launched Clementine on September 9th across the US and Canada. A conversation, recipe, or grocery list can become a personalized, ready-to-buy cart. Instacart is also deploying the underlying cart assistant technology on retailers' own properties.

All right. We all saw this coming. It's starting to happen already with Instacart. Simon and I talked about this on the last episode of what happens when, after the recommendation, your AI agent allows you to then make the purchase, and that could happen, of course, with Instacart. That Instacart is getting ready for that future. I think it's actually gonna be happening within the LLMs themselves because they are turning into AI assistants, personalized assistants that you never really should need to go outside of to do anything.

And specifically for agentic commerce there, there may be no referral click anymore between the recommendation and the transaction. And what happens then? I still think it could be an affiliate, revenue-share type of arrangement. But it's very important, especially for e-commerce sites, to really start preparing for this agentic commerce future, where agentic recommendation visibility and particularly product data availability, product attributes, and other types of evidence becomes served in the proper formats to agents.

All right, next story. Google is moving more search execution into AI Max this month. Google postponed the dynamic search ads sunset to February 2027. But automatically created assets and campaign-level broad match still begin automatic migration to AI Max in September 2026. Google reports the full AI Max suite produced an average 7% more conversions or conversion value at similar cost per acquisition / return on ad spend in its internal non-retail advertiser data.

All right. Really interesting. So this is internal non-retailer advertiser data. Google is showing that migrating to AI Max from dynamic search ads drove conversion value up by 7%. So for those of you who know dynamic search ads, basically for very large sites, it allows Google to crawl the website and then serve the ad to the appropriate page and even create the ad itself. So that is getting sunsetted. It's being replaced by AI Max, where I think that's gonna be covering all of Google's properties, whereas dynamic search ads are covering just the search. So Google is here increasingly letting AI choose the queries, creative, and the landing page destinations.

All right. Last story of the day. Open source GEO tracking is becoming credible enough to commoditize basic monitoring. This is reported by Elmo on GitHub. Elmo describes itself as an MIT-licensed self-hostable GEO platform, tracking mentions and citations across ChatGPT, Claude, Perplexity, Gemini, Copilot, Grok, and Google AI Overviews. Its public repository was updated September 11th and showed roughly 319 stars and 75 forks when indexed.

So another new entrant into the game of tracking, GEO tracking, and this is open-sourced on GitHub, so anybody can get it. So prompt tracking is the entry point here. It is definitely becoming commodity infrastructure. And again, I believe that the moat needs to sit above just the collection and observation of data, but it needs to be associated with proprietary knowledge, statistical certainty, decision intelligence, execution, and off-site citation building. Prompt tracking alone is really table stakes at this point.

All right. That's all we've got for today. Thank you for tuning in, and we'll see you all in the next one.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.