Cloudflare's September 15 AI crawler defaults make infrastructure configuration a real GEO variable: search crawlers allowed by default, training and mixed agents blocked. On episode 31 of The GEO Show, Paris Childress covers AirOps revenue attribution, Perplexity frontier-model retrieval shifts, Lantad robots.txt training-versus-search splits, a 47% cross-engine invisibility finding, revised Wikipedia AI Overview referral estimates, Serva's data-provider stack, tourism and automotive GEO verticals, and marketing-tool consolidation.
Cloudflare's September 15 AI crawler defaults make infrastructure configuration a real Generative Engine Optimization variable. On episode 31 of The GEO Show, Paris Childress explains that new domains and free customers without a preference will allow AI search crawlers by default while blocking training and mixed search-training agents on ad pages, so brands can gain or lose retrievability because a setting changed rather than because content quality did.
Around that throughline he covers AirOps Page 360 revenue attribution, Perplexity on frontier models changing retrieval, Lantad robots.txt training-versus-search splits, a Start Solutions finding that 47% of businesses were invisible across five engines, a revised University of Washington Wikipedia AI Overview referral estimate, Serva's SERP API and DataForSEO stack, Japan Inbound AI Visibility for tourism, Horizon Creator embedding GEO in automotive content systems, and Workaholic consolidating legacy marketing tools.
Because crawl policy is now part of retrievability. Cloudflare will allow search crawlers by default, block training by default, and apply the stricter rule to mixed-purpose agents. Free accounts that never set a preference inherit those defaults. Paris argues that monitoring AI bot crawling is becoming a critical GEO step and that Cloudflare is among the best places to do it. Brands that want training inclusion must open training bots intentionally.
Infrastructure configuration is becoming a real GEO variable here. And brands may gain or lose retrievability because a Cloudflare setting has changed rather than because content quality or authority has changed.
AirOps' Page 360 framework joins AI visibility, GA4, and Google Search Console, and claims that 70% of AI-referred sessions can appear as direct in GA4. Paris walks the failure modes: click-through LLM referrals, unlinked recommendations that become branded organic search, and app sessions mislabeled as direct. Software attribution alone is incomplete. Self-reported "How did you hear about us?" still matters through the pipeline.
Perplexity co-founder Johnny Ho says frontier coding gains help the product write better search programs, summarize, and operate workflows more autonomously. Paris ties that to ChatGPT 6-class capability: retrieval, source selection, query fan-outs, and synthesis all shift when the model changes. Annotating model-update dates is required when diagnosing visibility swings.
Yes. Lantad's robots.txt pass across 1,027 hostnames (718 parsed) found dozens of root blocks for GPTBot, CCBot, Bytespider, and ClaudeBot, and 59 hosts that blocked GPTBot while still allowing OpenAI's search bot. Paris compares blocking search bots to de-indexing from live RAG and says most non-publisher brands should allow both training and search crawlers if AI visibility is the goal.
Start Solutions tested 36 businesses across five engines and found 47% named by none of them, with only two brands appearing in all five. Aggregate inclusion rates looked relatively close while brand overlap was poor. Paris calls aggregate-only share of voice pointless and insists on model-by-model reporting. The same episode also shows GEO verticalizing into tourism and automotive content systems, and platforms consolidating around SERP API / DataForSEO data providers as legacy marketing tools bundle under pressure from AirOps- and Profound-class systems of record.
00:28. Cloudflare: search allowed, training and mixed agents blocked by default starting September 15.
03:18. AirOps: 70% of AI-referred sessions can appear as direct in GA4.
07:15. Lantad: 59/718 hosts blocked GPTBot while allowing OpenAI search.
08:46. Start Solutions: 47% of businesses invisible across five engines.
13:07. Japan Inbound AI Visibility: tourism vertical product for regional GEO.
The throughline is access and measurement: crawler defaults, attribution gaps, and engine divergence now decide whether brands appear in AI answers as much as content quality does.
Hi, everybody, and welcome back to another episode of The GEO Show, brought to you by GEOforge: Full Self-Driving for AI Visibility. I'm your host, Paris Childress, and we are on episode 31 on September 14th. Let's get right into it. We have an impending deadline coming up from Cloudflare tomorrow on September 15th.
Cloudflare is going to change their AI crawler defaults starting tomorrow. Cloudflare says that new domains will default to allowing AI search crawlers while blocking training and agent activity on pages that are displaying ads. Mixed-purpose crawlers combining search and training inherit the more restrictive policy. Cloudflare also says existing free customers who have not selected a preference will receive the new defaults. Cloudflare is explicitly positioning attribution and business insights around AI bot consumption, referral traffic, and AI answer visibility.
This is really, really important. I think that this is becoming a critical monitoring step of GEO, which is monitoring your AI bot crawling activity. Cloudflare is probably the best tool to do so. And tomorrow, for those who are on free accounts, the default is going to change, and you may have some AI agents that are blocked by default starting from tomorrow, and these could include some training bots. You have to be very careful to configure this.
So infrastructure configuration is becoming a real GEO variable here. And brands may gain or lose retrievability because a Cloudflare setting has changed rather than because content quality or authority has changed. Search will be allowed by default. Training will be blocked by default. But any agent that mixes search and training will also be blocked by default.
If you do want your data to be used to train AI agents and to update LLM training datasets, which in most cases you will want, because if you want to appear in LLM responses you have to open up your content to training bots, that is going to be blocked by default. So pay close attention to that.
Next up, AirOps moves explicitly into AI search revenue attribution. As of September 13th, AirOps has published a framework for connecting AI visibility to revenue, and it says its Page 360 framework brings AI visibility, Google Analytics 4, and Google Search Console into one single view. AirOps claims that 70% of AI-referred sessions can appear as direct in GA4 and argues that teams need software attribution plus self-reported attribution.
AirOps is really doing the right thing here. They are bringing together GA4 data, and some of the most important data there for GEO is going to be your AI referral traffic. And also, we have Google Search Console data. I think what is important inside of there is your generative AI impressions, which are basically your citations in Google's properties, as well as your brand search trend.
Because many people will see a recommendation in an LLM response, but there is no link there, so they will copy that name, paste it into Google, and land on your site via brand search. And that needs to be attributed the right way. And this is why the theme of self-reported attribution now comes back into focus. It is very important regardless of how people arrive at your website. If they are clicking through an LLM response, great, then that is going to show up as AI referral traffic in GA4.
If they are copy-pasting your brand name into Google, that is going to show up as brand search, as an organic branded search. And also, if they are using mobile apps or even desktop apps, then a lot of the AI-referred sessions will be incorrectly labeled as direct. So ultimately, the way that you fix all of this is you have to ask the lead in the first form that they fill out, "How did you hear about us?" And hope that they can answer accurately, and then you need to attribute this all the way down through the pipeline.
Let's move on to the next story. Perplexity says better frontier models directly improve its search engine. Perplexity co-founder Johnny Ho says improvements in GPT-6 Astra's coding capabilities improve Perplexity's search engine because the model can write better programs to search the web and internal sources. It can better summarize information, test systems, and operate production workflows more autonomously.
We are hearing it yet again, and we are hearing it specifically for ChatGPT 6. Because of its improved capabilities, it is changing retrieval behavior, source selection, query fan-outs, and synthesis. It is not just affecting the quality of the final answer, but it is affecting the fundamental infrastructure of how ChatGPT is synthesizing that answer, how it is doing its retrieval, its source selection, fan-outs, and so on. What is really important here is to know and to annotate when a model changes, because changes in your AI visibility can be as a result simply of model updates like GPT-6.
Next story. Real-world robots.txt data shows publishers separating training from search. Lantad requested robots.txt from 1,027 hostnames on September 12th. Seven hundred eighteen parsed successfully. GPTBot was blocked at the root by 82. CCBot was blocked by 80. Bytespider was blocked by 76, and ClaudeBot was blocked by 75. More strategically, 59 of the 718 blocked GPTBot while still allowing OpenAI's search bot, indicating deliberate training opt-out without search citation opt-out.
So here again, we see this theme of managing bots. Do you want search bots? I think if you block search bots, then you are effectively almost de-indexing your pages from the RAG process, the real-time search and retrieval process, and you will almost definitely become invisible in LLM responses. On top of that, if you opt out of training bots, then you are saying that you do not want the LLMs to use your information to update their training data. So whenever they are supplying a direct answer without a real-time search, you do not want to be part of that training dataset that allows that to happen.
I think for most people here, unless you are a publisher that is monetizing editorial content, like a large news media company, it makes sense to let all of these bots crawl to give yourself the best chance for AI visibility, whether it be through real-time search or through the instant answers.
Next story. A small cross-engine study finds 47% of businesses are invisible everywhere. Start Solutions AI tested 36 businesses across ChatGPT, Gemini, Copilot, Google AI Mode, and Perplexity. Seventeen, 47%, were named by none of the five. Only two appeared across all five. Aggregate inclusion rates were relatively close: 33% for ChatGPT, 28% for Gemini, 28% for Copilot, 22% for Google AI Mode, 17% for Perplexity, while the overlap was poor.
This is a relatively small dataset, but what we are seeing here is that similar aggregate engine visibility rates can hide radically different brand sets underneath. Engine divergence is a real thing, and reporting on share of voice at the aggregate level is pointless in my opinion. You really have to do it model by model.
Next story. A major AI traffic estimate fell from roughly fifteen percent to five percent after revision. University of Washington researchers now estimate default Google AI Overview availability reduced external search referrals to English Wikipedia by 5.45% versus German Wikipedia and 4.82% versus French Wikipedia. The updated estimate replaced a substantially larger estimate from the earlier version. So Wikipedia's own AI traffic is now starting to diverge among its own language versions for AI Overview availability. Even serious AI search impact studies can change dramatically when specifications and controls improve.
Next story is coming from Serva. Serva.ai updated its AI visibility documentation on September 12th. Its brand mentions view combines SERP API tracked prompt responses with DataForSEO-discovered LLM mentions, covering ChatGPT, Perplexity, Claude, Gemini, and Google AI Overviews. It also exposes estimated AI search volume, citations, competitors, and discovered prompts.
I took a look at this in more detail, and what I am taking away here is that there is a new emerging layer supporting this tools category, the GEO platform category, and these are data providers such as SERP API and DataForSEO. We, for example, use DataForSEO to get our visibility metrics from Google AI Overviews and Google AI Mode. It is basically scraping those results for us, and we are partnering with DataForSEO to get those results.
Here, Serva.ai is using both DataForSEO for LLM mentions and SERP API for tracked prompt responses. So I took a closer look and they are pulling from SERP API the prompt responses and the visibility, and from DataForSEO they are trying to estimate the prompt volume. And I think that this is fundamentally flawed by trying to estimate prompt volume based on keyword research, as I mentioned in a previous episode. But regardless, DataForSEO is still attempting to provide some sort of estimate for prompt volume.
But the more important story here is that there is a new set of data providers that are supporting this category, and more and more share of voice reports are going to be derived not from the LLM engine's APIs directly, but from these data providers like SERP API and DataForSEO.
Next story. AI visibility gets a vertical tourism product. Japanese firm 56 Product launched Japan Inbound AI Visibility on September 13th. It runs eight fixed English travel questions daily against 14 tourism regions using the web-enabled OpenAI Responses API. In the initial run, Hokuriku appeared in seven of eight queries. Twelve of fourteen regions appeared at least once. The company plans 30, 60, and 90-day tracking, more models, source comparison, and custom dashboards for municipalities and tourism organizations.
That is very interesting to see now that GEO is starting to verticalize into different industries. And here, this is tourism-specific. And customers that are buying industry-specific outcomes rather than generic AI search dashboards is going to become more of a thing. And I think this category is going to start to splinter now into industry verticals, such as tourism or maybe cybersecurity or SaaS. I think it is inevitable given the sheer number of players that are already in this category.
Next story. Horizon, H-R-I-Z-N, embeds GEO inside an automotive content operating system. Horizon released Horizon Creator on September 13th for automotive dealerships. It uses live vehicle identification numbers, VINs, and dealership data to generate scripts, run compliance review before recording, and connects frontline-created material into a larger content operating system. Horizon explicitly frames the output as supporting SEO and GEO, and it describes its broader objective as durable visibility across search, social, and AI-mediated discovery.
Very interesting. Here we have another vertical play for automotive dealerships from this company, Horizon. And we can see that vertical software companies now are increasingly treating GEO as one output of a workflow-specific content system rather than as a standalone discipline. So GEO is now starting to become a major consideration for other software providers that handle content because that content needs to be GEO-friendly, and we are seeing a direct example of this here in the automotive space with Horizon Creator. Interesting stuff.
Onto our last story. AI marketing platforms continue consolidating previously separate tools. Workaholic announced on September 13th that its AI content, SEO, email marketing, pay-per-click, and analytics tools are now unified inside one operating system sharing data between functions. This is not a GEO-specific launch, so its relevance is structural rather than directly competitive.
Point solutions here like Workaholic are facing growing pressure to bundle as broader marketing systems are absorbing previously separate capabilities. What we are seeing in the leaders of the GEO category space like AirOps and Profound is we are seeing absorption of these capabilities, and these systems want to become the new marketing systems of record. AirOps and Profound want to replace all these point solutions like Workaholic that has email marketing and pay-per-click management and analytics built in. They want to unify all that.
So what that is forcing some of these legacy point solutions to do, these pre-GEO era legacy marketing solutions, is that they now have to bundle, and they have to start unifying their features, because there is pressure from this new breed of AI-native marketing platforms.
All right. That will do it for today. Thank you all for listening, and we will see you on the next one.