Stop Blocking Bots. Start Getting Paid.

The GEO Show
September 19, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Cloudflare now lets sites reject AI training crawlers without sacrificing search indexing. Paris Childress argues most brands maximizing AI visibility should let training bots learn proprietary knowledge, alongside Azoma shopping-citation data, BrandGhost's 11.8% cross-engine overlap, and Machine Relations' long-tail map where half of citations need 1,357 domains.

Key takeaways

  • Profound flagged Google AI mode citation data as still variable after model-layer changes.
  • Azoma: 86.5% of Alexa shopping citations and 76% of Walmart Sparky citations came from earned or social media.
  • BrandGhost measured about 11.8% cross-engine source overlap across ChatGPT, Claude, Gemini, and Perplexity.
  • Machine Relations: half of AI citations require 1,357 domains; Reddit leads at 1.81%.
  • AppScribed saw live ChatGPT cite a site in 5/10 runs while an API workflow returned 0/30.
  • Cloudflare Disallow AI Training separates training preference from search indexing; Paris says most brands should let training bots in.

Paris Childress hosts episode 38 of The GEO Show on Cloudflare Disallow AI Training, Azoma shopping citations, BrandGhost 11.8% cross-engine overlap, and Machine Relations long-tail map. Watch: https://youtu.be/pidvOA4pl_c Audio: https://share.transistor.fm/s/e550f347

Full transcript

[00:00:00] Paris Childress: Hey, everybody. Welcome back to another episode of The GEO Show, brought to you by GEO Forge: full self-driving for AI visibility. I'm your host, Paris Childress. We're on episode 38. It's a lovely weekend, and I am excited as always to bring you the most important news and updates in the world of GEO. So let's get right to the first story, which is coming from Profound.

[00:00:30] Profound publicly flags Google AI mode citation instability. Profound's status page says its Google AI mode citation data remains affected by variability. On September 18th, Profound said it had made progress understanding the issue and was validating an updated approach to improve stability. All right.

[00:00:50] Lots of news about Profound in the last few days. Here they are admitting that they can't really figure out what's going on with Google AI mode citations. One of the category's largest vendors the largest vendor really, is publicly acknowledging that AI search measurement itself can become unstable when the underlying platform changes.

[00:01:11] So again, this is reinforcing the theme that you always really need to keep an eye on what's happening at the model layer here, because changes to the model, updates to the models can really, really impact your citation rate, your mention rate, your visibility, and your share of voice. And I think it's good b-by Profound to get out ahead of this and to publicly state here, uh, in their status page of their website that the data from Google AI mode now remains affected by variability, and therefore, you probably should treat it with a grain of salt Second story of the day, Dunnhumby invests in Amazon's...

[00:01:54] Excuse me, Dunnhumby invests in Azoma's agentic commerce optimization platform. Dunnhubby-- Dunnhumby Ventures, that's a tongue twister, invested in Azoma, which measures and improves how AI shopping agents discover and recommend products. Azoma says its Q2 analysis found eighty-six point five percent of citations behind Alexa for shopping recommendations came from earned or social media, and seventy-six percent for Walmart Sparky.

[00:02:28] Okay, GEO verticalizing into commerce-specific optimization here, and eighty-six percent. We are seeing a common theme here, which is that the large majority of citations on Alexa for shopping and also Walmart Sparky, eighty-six percent on Alexa and seventy-six percent on Walmart are coming from earned or social media sources, not from the brand's own websites.

[00:02:57] And this is commerce, so it's very interesting. So even in e-commerce, it's very important to do the off-site work to build the off-site citations so that you can make it into these shopping recommendations for Alexa and Walmart Sparky. Next story. Nexi embeds agentic commerce optimization into its merchant network European payments company Nexi partnered with Refi Buy to help merchants structure p- product catalogs for AI agents.

[00:03:33] Nexi says for-- A forthcoming survey of roughly twenty-eight thousand consumers in eleven countries found that more than half had used AI in the prior month to search, find, or buy products and services online. All right. Payments and merchant platforms are starting to treat AI discoverbi- discoverability as infrastructure, not just a marketing experiment, and this is coming over to Europe.

[00:04:00] So what they're trying, what Nexi partnering with Refi Buy, helping merchants structure product catalogs for AI agents. This is be-becoming a major theme in e-commerce AI visibility. It's how do you structure your product catalogs in the optimal way to ensure that AI agents can read them, they can understand about your products, they know who's recommending them, they know when to recommend your products.

[00:04:26] And as that next step after recommendation moves into AI assistants being able to help you complete the purchase, the discoverability is key. And these product catalogs being machine-readable is going to be an absolute must for any e-commerce company or brand that wants to sell through AI agents. All right, next story.

[00:04:50] Mastercard gives AI agents one-time payment credentials. So more on AI agents for shopping. Mastercard and Alchemy are integrating Mastercard Agent Pay into Agent Card, enabling approved AI agents to use one-time tokenized payment credentials within user-set constraints at ordinary online merchants.

[00:05:14] Agentic commerce now is moving from recommendation into transaction execution, and this is gonna happen faster than we all realize. This is all about trust, and I am old enough to remember when the majority of people did not trust putting their credit card online and completing transactions and, and basically buying things online with their credit card.

[00:05:37] They did not trust the merchants. They did not trust the internet in general. But after a few years, all that went away, and now you can see the level of trust is there. And I think that agentic commerce is gonna go through a similar cycle where right now, most people are skeptical. Should I give my AI agent my credit card?

[00:05:57] What if it goes rogue? What if it spends a lot of money that it's not authorized to do? What if somebody, uh, what if it interacts with another agent who might steal my card information? All sorts of fears like that are very reminiscent of, of the internet maybe around twenty, twenty-odd years ago, maybe twenty-five or so.

[00:06:16] All that got resolved and it's gonna be resolved here. And I think most of the buying people are gonna do in the next few years will be done through agents via agent to agent, meaning that my personal agent will go and interact with, maybe even negotiate with the seller's agent, agree on a price, and complete the transaction, all without me really having to leave my chat or my AI environment, which is probably gonna be mostly voice in the future.

[00:06:45] All right, back to the s- stories. Cross-engine source overlap is only about 11.8%. BrandGhost's September 18th AI Discovery Observatory analyzed seven thousand nine hundred and two citations across ChatGPT, Claude, Gemini, and Perplexity. It reports roughly 11.8% cross-engine source overlap. That's not much, and it says that 41.8% of comparable recommendation sets had no brand overlap at all.

[00:07:20] No meaningful universal AI ranking service sur- surface, excuse me. So getting citations on one platform really does not mean that you're getting citations on another platform. This is particularly true now as we analyze ChatGPT versus Google's Gemini AI mode and AI Overviews. So it's entirely different.

[00:07:46] So unlike the world of SEO, where you really just had to optimize for Google, which had, ninety plus percent market share, here you do have to optimize at least for two different platforms, ChatGPT, and I would, I would say right now, Google AI Overviews. But I do believe that Google AI Mo- Overviews is just a bridge to Google AI mode, which will become the permanent Google experience maybe even within the next six to twelve months.

[00:08:13] And those are gonna be the two dominant places, but GEO is gonna be at least a two, two-model optimization game instead of a one-model optimization game, or a one channel or one platform. All right, next story Half of AI citations are spread across one thousand three hundred and fifty-seven domains.

[00:08:37] Machine Relations' September 18th index covers twenty-two thousand hundred and seventy-nine cited domains and fifteen thousand seven hundred and eighty-two answer runs. The top domains account for only seven point four percent of citations. The top ten domains only make up seven percent of citations.

[00:08:59] Reaching fifty percent of the citations requires one thousand three hundred and fifty-seven domains. Wow. And eighty percent, getting to eighty percent of the citation share requires six thousand six hundred and forty-three domains. Reddit is the most cited domain, and that accounts by itself for one point eight one percent.

[00:09:25] So what are we seeing here? If we map this out as a graph, it's basically all long tail from what I can see. If you just decide I'm gonna go all in on Reddit and I'm gonna put all my efforts into my Reddit presence because I wanna win citations, yeah, maybe you can, maybe you can break into one and a half, two percent of the citations of your prompt set.

[00:09:49] That's just not enough. So the strategy really here is to try to figure out how to exploit the long tail in a scalable way. And I think what that really means here with citations is creating a lot of content, publishing a lot of content on your own properties, but also doing a whole lot of outreach and getting mentioned and getting published on as many different unique domains as possible.

[00:10:15] I think this is a volume game now more than it is about authority, 'cause it used to be in the AI world that if, let's say if, if I got one link from The New York Times or, you know, something like a domain authority, domain rating of ninety plus, that would be worth more than a thousand links which had low domain rating.

[00:10:37] But that's not true here, it seems. It seems that it is really about volume. Getting as-- Getting your brand in as many domains as possible is the name of the game for GEO, regardless of domain rating or anything else. All right, next up One study finds live ChatGPT and API-based tracking can disagree completely.

[00:11:04] Now, this one is very interesting because most of the tools that are out there, including our own, are using API-based tracking for tracking prompts. That's really the best way to do it, and it's very, very difficult to try to simulate real users, and I don't think anybody is really doing this at any degree of real scale yet.

[00:11:28] So a company called AppScribed reports that one buyer query cited its site in five out of ten live logged out ChatGPT runs. These are live runs, logged out users, live runs. While an automated API workflow returned zero citations in thirty runs. Its broader dataset logged one hundred and ninety-six thousand five hundred and fifty-four citations over ninety-one days, and it also observed pages ranking number eighty or outside Google's top one hundred being cited All right.

[00:12:13] Lots of interesting data to take in here. Let's see here. I-- The first thing that jumps out at me here is that there is a major difference in tracking your prompts. If you're tracking AI visibility from a set of prompts, there's a major difference from using the API, which is what most tools do. So most likely, I mean, Profound, Peak, Air Offs Scrunch, Athena to my knowledge, these are all using the API.

[00:12:44] So Chat- uh, OpenAI's API to, to basically scrape those answer pages and to assess, uh, brand mentions and citations and overall visibility. That's very, very different when you have a live logged out ChatGPT run. And we're not even talking about actual users who are logged in with accounts, which is the majority.

[00:13:08] So with logged in users, you have another layer of personalization and memory. And for those of you who are heavy ChatGPT users, you'll know just how personalized the answers are getting. A lot of your answers sometimes refer back to things that you've mentioned to ChatGPT and facts that you've given it about yourself and your life, and it could go way back.

[00:13:31] So i-it-- That-- This is really just another proof that prompt tracking is inherently very different, uh, very, very difficult, much different than rank tracking ever was in the SEO days. Rank tracking was more or less stable and deter-deterministic. Prompt tracking, on the other hand, is extremely probabilistic.

[00:13:54] It is much more personalized. And then, of course, you just get very different responses from the API versus simulating live users And there's also another little interesting tidbit here, which is that when the big dataset with 196,000 citations over a 91-day period, that pages, uh, that it res- it observed pages ranking number 80 or outside Google's top 100 being cited.

[00:14:26] So this means that also underscoring the fact that you really don't need to be ranked, not even on the top 10 or top 20. Here, there were citations that were not even ranking in the top 80 or the top 100, top 10 pages of Google. Still, the that query fan-out process and that real-time retrieval, that RAG process can go very, very deep.

[00:14:50] It can go very far into the search results. So the key for now in my understanding is you do have to get indexed because that real-time search is gonna go into Google, into Bing, and you do need to be ranked somewhere, but you definitely don't need to be ranked in the top 10 to have a good shot at winning a citation Next story.

[00:15:13] Cloudflare now lets sites reject AI training without sacrificing search. Here we're going to remind everyone on September 15th, there was a major change in Cloudflare's policy. The new Disallow AI Training setting publishes a no training pref- preference while allowing qualified mixed-use crawlers to continue traditional search indexing.

[00:15:37] Apple, Google, and Microsoft are classified as accountable mixed-use operators that honor or have committed to honor the separation. This is getting very refined. I will remind everybody who's listening again, if you haven't-- if you have a Cloudflare account, either free or paid, you really need to go into your settings and look at your AI bot crawl activity, look at those settings, and really make a decision.

[00:16:02] Do you wanna allow the search bots and the training bots? Uh, because by default, most likely, Cloudflare has reset the crawl permissions to disallow the training bots and to allow only the search bots. So what does, what does that mean? Disallowing the training bots means that the training bots that come to update the training datasets of the LLMs.

[00:16:26] So, uh, I think it's OpenAI's training data bot. I can't remember the exact name. It's not going there to, to to conduct inference as part of a real-time search to fulfill a user's prompt in the moment. That's what the search bots are doing. The training bots are going there really to just enhance and enrich their own memory, their training datasets.

[00:16:48] And I think it's still very, very important to, to have your brand's proprietary knowledge made available for these training bots so that AI can learn more about your brand, not only fetch it in real-time through search, but it can actually learn and commit the accurate facts about your brand into its training datasets.

[00:17:08] This is really important. I think that in most cases, there are few good reasons, unless you're a large publisher that doesn't want your content scraped and res- and, uh, resynthesized in LLM responses. If that's not you, and I think you're a brand that wants to maximize your surface area and your overall visibility in AI search, let those training bots in.

[00:17:31] Let them-- Tell them all about your brand. The more they know, the better. All right. Next story. Unsealed documents expose an AI content supply chain problem. Newly unsealed filings in The New York Times case against OpenAI and Microsoft include internal discussions warning of a doom loop in which AI products weaken the economics of publishers whose content models depend on, whose content models depend on.

[00:18:04] Microsoft says some quoted comments represented individual employee views rather than company legal positions. All right, this builds on the last story that we mentioned here, and of course, yeah, this doom loop thing is, is a little bit disturbing here because if basically what it, this, what it's saying here the the discussions of this doom loop is that effectively if no one reads The New York Times anymore, because you're able to basically to just get that content.

[00:18:38] If, let's just say if The New York Times allowed all of these training bots to come in and train on all of its data, including the search bots, then you would really be, be able to access all of The New York Times content through ChatGPT, and you could bypass The New York Times entirely. You wouldn't need to go to their website.

[00:18:58] You wouldn't need to subscribe as long as you were subscribed to ChatGPT. This is clearly problematic. But then the doom loop means that if New York Times then went out of business as a result of that, then that rich, that rich reporting, that, that first-hand reporting, that rich content would also cease to exist, and then that would degrade the quality of what ChatGPT could provide.

[00:19:21] And that, I think that's an interesting concept here. For most of us in, uh, B2B marketing, this is not a concern really. But I think what we should be focused on is, um, really giving as much proprietary brand knowledge and data as possible without crossing the line into the secret sauce of the alpha of the business All right, last story of the day comes-- is about a new brand called Pyra.

[00:19:54] Pyra launches a low-cost AI visibility race engineer. Pyra launched Primey, an AI search race engineer layered on top of its Prime AI visibility product. It uses stored answers and citation evidence to explain competitor wins and recommended next actions. Their played-- paid plans start at just sixty-nine dollars a month, with enterprise, enterprise plans starting at fourteen ninety-nine, uh, one thousand four hundred and ninety-nine per month.

[00:20:28] Uh, another example of the diagnosis and recommendation layer moving rapidly down market. So what they're doing, it appears here, is digging into the competitor wins. So it's looking at the citation evidence, and it's seeing which competitors are gaining ground. So maybe you're in the lead, but which competitors may be closing in, or, or if you're not in the lead, which competitors may be pulling further away, and then what you should do about it.

[00:20:57] This is, I have to say, one area where GEO is similar to SEO in that it is very much about watching the competition, building your strategy directly based on what competitors are doing and basically making sure that whatev-- what you're doing has to be better than your nearest competitors, whether that be you have to outproduce better content, you have to, to out citation build them, whatever it is.

[00:21:31] But it is very much a competitive game, just like with SEO. All right. That'll do it for today. Thank you all for listening, and see you all on the next one.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.