ChatGPT Is Building Its Own Search Index

The GEO Show
September 7, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Peak AI reports OpenAI is developing internal ChatGPT search infrastructure, including a system referred to as Labrador and specialized indexes for news, shopping, PDFs, YouTube, and academic content, so ChatGPT can retrieve from indexed and cached web pages rather than relying only on Bing-style third-party results. On episode 22 of The GEO Show, recorded September 7, 2026, Paris Childress covers why crawl and index behavior become first-order GEO variables, then Profound Context Manager, Semrush Spotlight after Adobe, Cloudflare's September 15 crawler defaults, Google AI Mode at 1 billion monthly users, AirOps budget lag, Metrisque and Rankcaster prediction layers, GEO Defender, Ahrefs engine-specific sources, Perplexity retrieval, ChatGPT SaaS plugins, Rebid ChatGPT Ads, Microsoft Advertising MCP agents, and First Page Agency's off-site citation case.

Key takeaways

  • Peak AI documents OpenAI building ChatGPT search infrastructure (Labrador) plus vertical indexes for news, shopping, PDFs, YouTube, and academic content, enabling retrieval from indexed/cached web content rather than Bing-only third-party results.
  • ChatGPT crawl and index behavior becomes a first-order GEO variable, reducing dependence on Microsoft/Bing and expanding OpenAI control over which sources get surfaced and cited.
  • Profound Context Manager structures brand memory across six records for agentic recommendations; Semrush Spotlight (Oct 13, London) doubles down on GEO after Adobe's April 28, 2026 Semrush close.
  • Google AI Mode surpassed 1 billion monthly users with query volume more than doubling each quarter; Cloudflare's Sept 15 defaults will block mixed-use AI crawlers on new sites just as OpenAI indexes the web itself.
  • Metrisque reported 92% recommendation-territory prediction accuracy; Rankcaster shipped P-citation prediction; GEO Defender cut attack success from 50% to 6%.
  • Ahrefs shows Reddit/YouTube citation share differs by engine; First Page Agency reported 800-plus citations in under four months via off-site contextual coverage rather than backlinks alone.

Peak AI reports that OpenAI is developing internal ChatGPT search infrastructure, including a system referred to as Labrador and specialized indexes for news, shopping, PDFs, YouTube, and academic content. ChatGPT can retrieve from indexed and cached web content rather than relying solely on Bing or Google-style third-party results.

Paris Childress recorded episode 22 of The GEO Show solo on September 7, 2026. He treats proprietary crawl and index behavior as a first-order GEO variable, then walks brand-memory tools, conference-scale GEO positioning, crawler defaults, Google AI Mode scale, budget lag, prediction layers, anti-manipulation defenses, engine-specific sources, SaaS plugins, paid ChatGPT visibility, ad agents, and off-site citation acquisition.

Why does ChatGPT building its own search index matter for GEO?

Because crawl and index behavior, not only answer-time ranking, becomes a controllable surface for AI visibility. If OpenAI builds proprietary crawling rather than leaning only on Bing, brands must treat ChatGPT bots like first-class indexers. Paris notes this also reduces ChatGPT dependence on Microsoft and Bing and expands OpenAI control over which sources get surfaced and cited.

If OpenAI is building proprietary crawling and indexing infrastructure rather than leaning on Bing, ChatGPT's own crawl and index behavior, not just its retrieval-time ranking, becomes a first-order variable in AI visibility, similar to how Googlebot behavior has driven SEO for two decades.

How are GEO vendors moving beyond share-of-voice dashboards?

Profound's Context Manager (September 1) ingests documents, transcripts, and conversations into six brand records for more consistent AI Marketer recommendations, live for enterprise customers. Semrush, now under Adobe after the April 28, 2026 close, is hosting Spotlight in London on October 13 for more than a thousand CMOs, following Adobe Brand Visibility's combined AI visibility datasets. Paris is speaking at BrightonSEO on October 8 and may attend Spotlight.

  • Context layers raise switching costs and push vendors toward systems of record.
  • Cloudflare's September 15 defaults will block mixed-use AI crawlers on new sites; audit allowlists as OpenAI indexes for ChatGPT itself.
  • Google AI Mode passed 1 billion monthly users; AI Overviews keep absorbing AI Mode capability on a Gemini 3.7 Flash-class model.

Are AI recommendations becoming predictable enough to score before they appear?

Metrisque's pre-registered study found predictions before roughly 1,100 AI recommendations correctly identified territory 92% of the time across 1,916 products, while conventional rankings and mentions explained little. Rankcaster added a September 6 machine learning layer that predicts citeable URLs using embeddings plus crawl freshness, schema coverage, domain age, and link graph features into a P-citation score. Paris sees the stack moving toward opportunity identification, expected impact, recommended action, and validation.

What happens when black-hat GEO meets engine-specific sources?

An arXiv GEO Defender paper (September 2) cut average success of seven citation-preference attacks from 50% to 6% while retaining 94% benign evidence use across five models. Ahrefs' September data shows Reddit at 16.8% of ChatGPT top source share and 28.5% in Gemini, while Google AI Mode splits Reddit 17.9% and YouTube 17.8%. Perplexity discloses proprietary embedding, retrieval, and re-ranking on an exabyte-scale index. There is no universal GEO publication list.

Where do paid ads, agents, and off-site corroboration fit?

OpenAI added Zendesk and OneNote plugins so ChatGPT can use SaaS knowledge, not only recommend software. Rebid put ChatGPT Ads beside Google, Meta, LinkedIn, X, and DB360, so earned GEO and paid visibility need separate attribution. Microsoft Advertising's MCP server lets Copilot, Claude, ChatGPT, and peers query live campaign data, with Stagwell and Conversios already auditing accounts. First Page Agency reported taking a brand from effectively zero AI visibility to 800-plus citations and mentions in under four months via LinkedIn, Medium, Reddit, and community coverage rather than backlinks alone.

Notable moments

00:21. Peak AI: ChatGPT Labrador index and vertical indexes.

02:15. Profound Context Manager as brand memory system of record.

06:31. Google AI Mode hits 1 billion monthly users.

10:30. Metrisque 92% recommendation predictability; Rankcaster P-citation scores.

19:31. First Page Agency: 800-plus citations via off-site context.

The throughline is proprietary retrieval: ChatGPT is indexing the web itself while GEO tools, defenses, paid layers, and off-site corroboration race to catch up.

Full transcript

Hey, everybody. Welcome back to another exciting episode of The GEO Show, brought to you by GEOforge: full self-driving for AI visibility. I'm your host, Paris Childress. We're recording episode 22 and got a bunch of hot stories to get into, so let's get right to it.

First up, ChatGPT is quietly building its own search index according to new research from Peak AI. This is from September 4th, and Peak AI has documented that OpenAI is developing internal search infrastructure for ChatGPT, including an internal system referred to as Labrador and specialized indexes reportedly in development for news, shopping, PDFs, YouTube, and academic content.

Peak's confirmed observation is that ChatGPT can retrieve information from indexed and cached web content rather than relying solely on Bing/Google-style third-party results. Really interesting news here, and not a big surprise, because I think the large majority of ChatGPT answers are produced via real-time search retrieval, and it's natural that they don't want to rely exclusively on other third-party search engines and that they're going to start to build their own index.

And it's interesting that they're doing even vertical indices in news, shopping, YouTube, and academic content. So if OpenAI is building proprietary crawling and indexing infrastructure rather than leaning on Bing, ChatGPT's own crawl and index behavior, not just its retrieval-time ranking, becomes a first-order variable in AI visibility, similar to how Googlebot behavior has driven SEO for two decades.

And of course, this reduces ChatGPT's dependence on Microsoft and Bing and gives OpenAI much more control over which sources get surfaced and cited. So very important development.

Next story. Profound ships Context Manager, turning AI Marketer into a persistent brand memory system. Profound announced on September 1st the Context Manager feature, which ingests documents, transcripts, and conversations into six structured brand records: positioning, market dynamics, product, audience, strategic priorities, and execution methods, which its AI Marketer agent then uses to generate more specific brand-consistent recommendations instead of just generic ones.

Profound confirms that it's live today for enterprise customers, with self-serve customers directed to request a demo. All right, this is a direct move up the value chain for Profound, moving from visibility monitoring towards agentic marketing execution. Profound here is trying to become the system of record for brands' other AI tools to plug into, not just merely a dashboard that reports share of voice.

And of course, these types of moves raise the switching costs for customers and they make their product much stickier. So very interesting move, not surprising by Profound, to try to become your AI marketing system of record.

And I think a lot of players are trying to do that, including the next one up here, Semrush, which is now an Adobe company. They have announced an event called Spotlight, a flagship AI visibility conference, doubling down on GEO after their Adobe acquisition. So the acquisition was closed officially on April 28th of this year, Adobe's acquisition of Semrush, and now Semrush is hosting a conference called Spotlight in London on October 13th, 2026, focused on AI visibility.

They're expecting more than a thousand CMOs and marketing directors with a confirmed speaker lineup that includes executives from OpenAI, LinkedIn, Wix, WPP, and Adobe. And this follows Adobe's earlier launch of the Adobe Brand Visibility product, which combines Adobe's own LLM visibility tooling with Semrush's AI visibility tool and its claimed database of roughly 300 million AI search prompts, 28 billion keywords, and 43 trillion backlinks.

So this is a marriage of two massive datasets across Adobe and Semrush, but the headline here is a new conference coming up in October, just next month. I'm going to be, by the way, speaking at BrightonSEO in Brighton, UK on October 8th, and I may just hang around a little bit and swing by this London conference on October 13th. It looks great. So if any of you are going to be at either of these conferences, at BrightonSEO in Brighton on October 8th and 9th or at Spotlight in London on October 13th, please shoot me a quick message. My email is [email protected].

All right, let's get on to our next story. It's a reminder that Cloudflare's September 15th deadline to default-block mixed-use AI crawlers is now one week away. I'll just mention this quickly because we have covered it before. But on September 15th, all new websites that sign up for Cloudflare will, by default, be blocking certain crawlers. They will be blocking crawlers that mix training and model-building activity.

And it's very important to monitor this and to make sure that you're not accidentally blocking one of the LLM bots that will get you more visible. And in the case of ChatGPT and OpenAI, those bots are now crawling to build their own index.

Next story. Google confirms that AI Mode has passed one billion monthly users, and GEO vendors are already repositioning around it for the holiday season. Wow, we're already talking about the holidays. Google stated at Google I/O 2026, and it was reported by TechCrunch on August 11th, that AI Mode has surpassed one billion monthly users, with AI Mode query volume more than doubling every quarter since launch.

And Search Engine Roundtable's September 2026 Webmaster report confirms that Google has continued pushing AI Mode capability into AI Overviews, which is now on a Gemini 3.7 Flash-class model, with new features like travel booking, price tracking, and stock displays. Separately, on September 1st, AI shopping vendor Azoma published commentary citing this billion-user milestone and recommending brands audit citation source mix.

The main takeaway here is that AI Mode is the real deal. AI Mode is the future of Google search. It is merging with AI Overviews. I've always thought that Google AI Overviews is just a temporary bridge to take users on a more smooth path from traditional Google to Google AI Mode. And I would not be surprised if the full transition to Google AI Mode as the default Google experience happens this year.

And of course, the billion-user count, compared to how many people are also using Google AI Overviews, because I don't believe that there are a billion users who are going right to Google AI Mode and bypassing Google.

Let's move on to our next story. This is from AirOps. AirOps' survey of 300-plus marketing leaders finds 86.6% are prioritizing AI search channels, but budget and maturity are lagging. AirOps has published its 2026 Marketing Leaders Reality Index, a self-commissioned survey of over 300 marketing leaders.

Reported findings include 86.6% of respondents prioritizing AI search channels, 88.4% describing their organization's AI maturity as emerging or operational rather than advanced, 75% facing higher pipeline and revenue targets, and only 43% receiving budget increases, and fewer than 25% receiving full budget approval for their plans. Teams with full approval were reportedly 2.3 times more likely to expand headcount.

This is vendor-commissioned research from AirOps; we always have to take that into account. What's interesting here is that, as we've seen in many aspects of marketing, there is a lot of interest and a high priority for this new channel. Almost 87% are now prioritizing AI search, and 88% describe their maturity as emerging or operational, but only 43% have received budget increases.

So I think that the budgets are lagging behind the interest and the enthusiasm, but I think the budgets will close the gap once we start to see more case studies and proof of return on investment in this emerging channel start to surface.

All right. New research says that AI recommendation may be more predictable than prompt tracking suggests. So this is cutting against some of the common logic so far. A company called Metrisque, spelled M-E-T-R-I-S-Q-U-E, launched an AI visibility measurement platform alongside a pre-registered study. It says predictions made before roughly 1,100 AI recommendations correctly identified the eventual recommendation territory 92% of the time.

The company's broader dataset covers 1,916 products and says conventional rankings and mentions explained relatively little of recommendation behavior. So the interesting question may be shifting from how often did we appear to why does the model classify us as relevant to this buyer need? This is a very interesting finding, and their predictions were pretty accurate. 92% of the time, they identified the eventual recommendation. So it's going to be interesting to see how this develops because companies are now starting to build models that will predict what comes out of an LLM response.

GEO software is moving from citation tracking to citation prediction. This is directly building on our last story. On September 6th, Rankcaster added a machine learning system that predicts which URLs are likely to be cited for specific prompts. Its model combines embeddings with signals such as crawl freshness, schema coverage, domain age, and link graph features to generate a P-citation score.

Again, it's now becoming prediction territory here. The next product layer is going to be opportunity identification, moving to expected impact, recommended action, and finally validation. Next story. Researchers are already building their defenses against GEO manipulation.

This is reported by arXiv on September 2nd. A new paper tests defenses against content deliberately rewritten to exploit generative engine citation preferences. Its GEO Defender system reduces the average success rate of seven tested attacks from 50% all the way down to 6% while retaining 94% of benign evidence use across five models.

Just like we saw the emergence of manipulative SEO, probably 15 to 20 years ago, and black-hat SEO, I think we are now starting to see the early signs of black-hat GEO and efforts to manipulate GEO by forcing certain types of content, these LLM-friendly rewriting patterns, onto the web. Let's see how GEO Defender is going to handle this, and it's going to be very interesting how the AI engines themselves are going to guard against this manipulation the same way that Google did back in the day when they launched their anti-spam efforts behind Matt Cutts and team.

September data reinforces that every AI engine has a different source ecosystem. This is from Ahrefs. Ahrefs' new monthly dataset puts Reddit at 16.8% of top source citation share in ChatGPT, 28.5% in Gemini, while Google AI Mode splits heavily between Reddit at 17.9% and YouTube at 17.8%.

So again, there is no universal best GEO publication list. We see that Reddit is still hanging strong across ChatGPT, Gemini, and Google AI Mode. And of course, YouTube is very strong in Google AI Mode. So we always need to build engine-specific source influence maps by prompt cluster and not just look at universal citations and share of voice across all models.

Next story. Perplexity just gave us a useful look inside its source selection machinery. Perplexity says its search stack uses proprietary embedding and ranking models to identify relevant results from an exabyte-scale search index. It has built its own serving infrastructure around those retrieval and re-ranking models. Traditional Google ranking is only one upstream signal. The AI engines can independently embed, retrieve, and re-rank documents.

We're seeing it directly here with Perplexity, and we now have the early signs of this happening with ChatGPT as well. So just to remind everyone, we're not optimizing for Google organic SEO here in order to be retrieved by GEO engines. We're optimizing more broadly for semantic retrievability and evidence relevance, and we should continue to measure conventional SEO separately.

All right, next story. OpenAI keeps turning SaaS products into native ChatGPT capabilities. OpenAI added Zendesk and OneNote plugins to ChatGPT and Codex. Zendesk users can now retrieve customer history, support knowledge, and tickets. OneNote users can retrieve and update organizational knowledge.

So AI visibility is expanding from can the model recommend my software to can the model actually use my software? And I think this trend has continued. We saw a major announcement a few weeks back by Salesforce that through MCP connectors, effectively AI is going to become the new user interface that sits between human users and their software stack. We've seen it at CRM level, and now we're seeing it across lots of other major SaaS tools like Zendesk and OneNote.

All right, next story. ChatGPT Ads is already becoming part of the mainstream media stack. A company called Rebid added ChatGPT Ads activation, analytics, and optimization alongside Google, Meta, LinkedIn, X, and DB360 inside its agentic marketing platform. So AI discovery is now developing the same two-layer structure as search, where we have earned visibility, which is organic, and we have paid visibility. So we need to make sure to manage GEO and ChatGPT Ads both together, but also strategically keeping their attribution separate.

Okay, last story of the day. Microsoft is letting agencies operate paid search through AI agents. This is from Microsoft Advertising's blog. Microsoft Advertising's MCP server lets Copilot, Claude, ChatGPT, and other compatible agents query live campaign data. Stagwell is using it for campaign audits while Conversios built an audit dashboard covering 120-plus accounts.

All right. I saw this coming, and it is definitely happening. AI agents will start to actively manage ad accounts inside of Google Ads and Microsoft Ads, and probably also LinkedIn ads and Meta ads as well. And now Microsoft is the first one who is expressly facilitating this move. So that means that PPC agencies, digital ad management agencies such as ours, we need to really start to see that the value is moving away from daily execution and optimization within the platform and more towards broader judgment, experimentation, strategy, and proprietary workflows beyond just pulling data and writing weekly summaries.

Okay. We do have one final story that I don't want to leave out. A fresh field case strengthens the argument for off-site context over backlinks alone. First Page Agency reports taking a brand from effectively zero AI visibility to 800-plus citations and mentions in under four months across ChatGPT, Gemini, Perplexity, Google AI experiences, and Grok.

Its program emphasized third-party contextual coverage across LinkedIn, Medium, Reddit, and communities rather than traditional backlink acquisition alone. All right. So the backlink game is definitely being replaced by the citation acquisition game as reported by this agency. GEO authority is much bigger than link building.

It's increasingly looking like distributed corroboration, not merely PageRank-style link authority. So all these link-building programs and people that are paying for links according to domain authority and domain rating, I think that whole game is going away, and it's going to be replaced by more of this distributed corroboration of getting brand mentions more broadly around the web.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.