Comparison Pages Win B2B AI Citations

The GEO Show
October 1, 2026
Watch this episode on YouTubeListen on Apple Podcasts, Spotify and more

Comparison pages and listicles dominate B2B AI citations: in a Foundation and AirOps study of B2B software buyer prompts, they made up 84 of the 100 most cited URLs. Paris Childress also covers Peec AI's finding that tracking more prompts beats more repeat runs, Emplifi and Semrush data on AI product research, BrightEdge on how the cited source changes with the question, Graphite's look at ChatGPT ads, the AI visibility strategy gap behind Yext's Expanded Scout, Google's AI Max migration, and G2's crowded category of AI visibility software.

Key takeaways

  • In Foundation and AirOps' study of about 3 million citation events across 81 B2B software categories, comparison pages and listicles made up 84 of the 100 most cited URLs, and vendor domains were the largest source class.
  • Peec AI found that with the same 1,500 monthly runs, 50 prompts run once daily gave about plus or minus 6.1 points of monthly uncertainty, versus 13.2 points for 10 prompts run five times daily. Add prompts before adding runs.
  • Group branded and non-branded prompts, freeze your prompt set, and judge AI visibility by the trend over time, not a single snapshot.
  • BrightEdge found an insurer's own domain won the citation only for branded queries; government, clinical, and HR sources owned the other question types.
  • Graphite's analysis of Similarweb panel data found business software was almost 20% of ChatGPT ad impressions from March to August.
  • A Corporate Ink survey cited by Yext found 88% of CMOs and VP-level marketers are being asked about AI visibility, while only 34% of marketers surveyed have a defined strategy.

Comparison pages and listicles dominate B2B AI citations. In a new Foundation and AirOps study, they made up 84 of the 100 most cited URLs across B2B software buyer prompts. The bigger theme this week: the category is moving from "can you track AI" to "can you measure the right market and influence the right sources."

In this solo roundup, Paris Childress, founder of Hop AI and co-founder of GEOforge, covers eight stories.

Do more repeat runs make AI visibility data more accurate?

Not as much as more prompts do, according to Peec AI. It analyzed about 120 million chats collected from February to August 2026 across major AI assistants. In a controlled comparison where both scenarios used 1,500 monthly runs, 50 prompts run once daily produced about plus or minus 6.1 points of monthly portfolio uncertainty, versus 13.2 points for 10 prompts run five times daily. Broader prompt coverage improves portfolio-level estimates more than repeatedly sampling a narrow set, while repeated observations over time still matter for detecting real change.

Peec's five recommendations: group prompts into branded and non-branded, don't judge visibility on single prompts, judge movement on the portfolio across time, add prompts before adding runs, and freeze the prompt set so the trend line compares like with like.

"This is not deterministic in kind of the way that SEO rank tracking used to be." (Paris Childress)

Which pages get cited in B2B software answers?

Foundation and AirOps tracked about 380,000 AI answers and 3 million citation events across 81 B2B software categories and six AI surfaces. Product and vendor domains were the largest source class on all six, and comparison pages or listicles made up 84 of the 100 most cited URLs. The source mix is both engine-specific and category-specific, so a universal prescription such as "do Reddit" or "get on G2" is structurally weak. For B2B SaaS, comparison pages and listicles still work, even though many listicles are self-serving.

How many people use AI to research products?

In an Emplifi survey of 1,638 US and UK consumers, 66% said they had used AI tools to research products or discover gift ideas, and 11% now start product research directly with AI. Semrush estimates AI referrals to a basket of 30 major US retailers grew 190 times from June 2023 to June 2026 and now equal roughly 8% of the basket's social referral traffic. Referral traffic is still small, but it is growing at an enormous clip.

Does better content always win the citation?

No. BrightEdge grouped a US insurance prompt panel into four question types. The insurer's own domain was the cited source only for branded queries. Government sources dominated definitional and compliance questions, official and clinical sources dominated drug and procedure coverage, and job and HR platforms dominated workplace benefits. Some query classes have structural authority owners that are extremely difficult to displace. BrightEdge does not disclose the panel size, and the findings are specific to this vertical.

Who is advertising on ChatGPT?

Mostly B2B so far. Business Insider reports that Graphite analyzed Similarweb panel data on ChatGPT ads from March through August. Business software was almost 20% of observed ad impressions, professional and business services 8.5%, and developer tools 4.6%. Click-through rate more than doubled from late March to early September, and fewer than half of ad-eligible users had been shown an ad at all. These are panel-derived observations, not OpenAI ad server data.

Do companies have an AI visibility strategy yet?

Most don't. Yext's Expanded Scout, which adds brand-level AI visibility optimization across more cited sources, was due for broader availability on September 30. Yext cited a Corporate Ink survey in which 88% of CMOs and VP-level marketers said leadership or the board was asking about AI visibility, while only 34% of marketers surveyed reported a defined strategy. That gap between awareness and strategy is a big opening for consultants and platforms with a clear workflow and measurement in place.

What else moved this week?

  • Google closes the AI Max migration window. Eligible campaigns moved to AI Max during September. Google says its full AI Max suite delivered an average of 7% more conversions or conversion value at a similar CPA or ROAS versus search term matching alone, in its internal non-retail advertiser data. Even paid search is moving toward machines deciding which queries, text, and landing pages match intent.
  • A crowded category on G2. G2's September snapshot lists 642 products in its category for optimizing brand visibility in AI answers, from SEO incumbents to reputation platforms and dedicated AI search vendors. G2 reviews are some of the most important citations a B2B SaaS company can build.

Notable moments

  • [02:15] Prompt breadth vs repeat runs: two different statistical questions
  • [04:04] Judge the trend, not the snapshot
  • [07:39] Why comparison pages and listicles still work
  • [10:34] Query classes with structural authority owners
  • [13:33] Boards are asking about AI visibility
  • [16:44] How crowded the category has become

Subscribe to The GEO Show wherever you get your podcasts, and watch the full episode on YouTube: https://www.youtube.com/watch?v=IgvqLJV5lEY

Audio: https://share.transistor.fm/s/70d0b237

Full transcript

Paris Childress: Hi, everybody. Welcome back to another episode of The GEO Show brought to you by GEOforge, full self-driving for AI visibility. We've got some great stories today on September thirtieth, episode fifty-two, and the exclusive theme here for today is gonna be that the category is moving from can you track AI to can you measure the right market and influence the right sources?

Paris Childress: So let's get into the top story, a very important analysis here coming from Peec AI. Peec AI challenges the more repeat runs equals better measurement assumption. And by the way, that is definitely our assumption at GEOforge, which is that more runs leads to higher confidence and a smaller margin of error in AI visibility reporting. So let's see what this is all about. Peec AI analyzed roughly one hundred and twenty million chats collected between February and August twenty twenty-six across major AI assistants.

Paris Childress: In its controlled comparison, both scenarios used fifteen hundred monthly runs, fifty prompts run once daily, produced a plus or minus six point one percentage points of monthly portfolio uncertainty versus thirteen point two points for ten prompt runs five times daily. So let's, let's take another look at that. That's a big difference, more than two times difference in confidence here.

Paris Childress: Peec AI concludes that broader prompt coverage improves estimates of portfolio-level visibility more than repeatedly sampling a narrow prompt set, and separately argues that repeated observations over time remain important for detecting actual change. This doesn't necessarily invalidate the repeated run testing at the individual prompt level argument, but it separates two different statistical questions, which is, one, estimating the behavior of one stochastic prompt versus estimating visibility across the broader market of possible buyer questions. And these are two different things.

Paris Childress: And what I'd like to do now is pull up the recommendations from Peec AI when they announced this. There are five steps for putting this into practice. Number one is grouping the prompts, branded versus non-branded, and we also have a tagging system in GEOforge within SignalForge that allows you to tag in any way that you want. And sometimes it's also helpful to tag your prompts by funnel stage: top of funnel, middle of funnel, bottom of funnel.

Paris Childress: But definitely you wanna separate branded versus non-branded because your visibility for branded prompts will almost always be one hundred percent naturally. Second recommendation is don't look at single prompts for visibility, and that is the real crux of this argument. If you're only looking at one prompt or a small set of prompts, it's just a weaker sample of the whole universe of potential prompts. It's just a smaller sample, so you have lower surface area.

Paris Childress: And Peec AI is recommending that if you want to get more accurate visibility across your portfolio, then you should increase the size of the portfolio, increase the number of prompts that you're tracking, rather than doing more repeated runs on a relatively smaller set of prompts. I generally do agree with that. Number three is to judge movement on the portfolio across time.

Paris Childress: Because what we're doing here essentially is probabilistic sampling, it is important not to look at any snapshot of one particular day or one particular week of visibility, but keep your measurement keep your measurement approach consistent over time and then look at the trend over time. If you're trending up or trending down then directionally you can see where your AI visibility is trending.

Paris Childress: And you don't need to focus as much, or you shouldn't focus as much on the actual percentages that are reported because they always will imply some degree of error, some margin of error, and this is not deterministic in kind of the way that SEO rank tracking used to be. The fourth recommendation is to add prompts before adding runs, and that is the point that this article is making. If you want more accuracy, then increase the size of the prompts.

Paris Childress: Increase the number of prompts that you're tracking, not the runs per prompt. And recommendation number five is to freeze the prompt set. That makes a lot of sense. You don't wanna be constantly changing the prompts that you're measuring because then you're just measuring apples and oranges, and that will really screw up your trend line. So very interesting analysis. In some ways, I agree with this but in other ways I don't.

Paris Childress: I mean, we, we have a policy of doing 40 runs per prompt every week across a portfolio set of 30 prompts, and that's what we've locked in on for GEOforge, and I think we have more statistical rigor in taking that approach than virtually any other competitor that we're going up against.

Paris Childress: Next up, AirOps maps three million B2B citation events and comparison content dominates. Foundation and AirOps tracked roughly three hundred and eighty thousand AI answers and three million citation events across eighty-one B2B software categories and six AI surfaces. Product vendor domains were the largest source class across all six. AirOps webinar says that eighty-four out of the one hundred most cited URLs were comparison pages or listicles. Eighty-four out of a hundred. Now, this is for B2B software, and that's important.

Paris Childress: This webinar, by the way, was from yesterday, September twenty-ninth, so this is very fresh research. And of course, we know that most of these listicles are self-serving. Social citations were only seven percent on Google's AI surfaces, but zero point three percent or lower on ChatGPT, Gemini, and Claude. ChatGPT's community share was nine point two six percent versus Claude's at two point six seven percent.

Paris Childress: So the source strategy is both engine-specific and category-specific, and a universal prescription, such as do Reddit or get on G2 or publish more owned content, is structurally weak here. What this is showing for B2B SaaS is that a very important strategy is to be generating comparison pages, me versus competitor X and competitor Y, et cetera, and as well as listicles.

Paris Childress: I think that there's really only space for one listicle per domain, and these are generally self-serving, but unfortunately, they do still work, and this research supports that.

Paris Childress: Next story. 66% of surveyed consumers have used AI for product research. Emplifi survey plus Semrush's modeled traffic data, a September 29th joint report surveyed sixteen hundred and thirty-eight US and UK consumers, and 66% said they had used AI tools to research products or discover gift ideas. And now 11% start product research directly with AI. I would have guessed higher than 11%.

Paris Childress: Semrush estimates AI referrals to a basket of 30 major US retailers increased one hundred and ninety times between the months of June 2023 and June 2026. Okay, that's a three-year period, so I guess that is plausible. And it now equals roughly 8% of the basket's social referral traffic. So very interesting. Two out of three people are now using AI tools to research products and discover gift ideas and referral traffic, while still low, is growing at an enormous clip.

Paris Childress: Next story. BrightEdge shows that the right source changes with the question. BrightEdge grouped a US insurance prompt panel into four question types. The insurer's own domain was the cited source in exactly one of four, which was the branded queries. No surprise there. Government sources dominated definitional and compliance questions. Official and clinical sources dominated drug and procedure coverage. Job and HR platforms dominated workplace benefits questions, and insurers could still rank among leading mentions while losing the citation to government domains.

Paris Childress: BrightEdge does not disclose the panel size, and the findings are specific to this vertical. So citation acquisition here is not merely about producing better content. Some query classes have They have structural authority owners that are extremely difficult to displace. And if you think about it, that does make sense. I mean, government sources dominating in compliance questions clinical sources dominating in drug procedure coverage, job and HR platforms dominating workplace benefits.

Paris Childress: That all makes sense, but it does give us a new lens in which to think about query classifications and the types of sites that would naturally dominate them.

Paris Childress: Next up, ChatGPT ads are already heavily B2B. Business Insider reports that Graphite analyzed Similarweb panel data covering ChatGPT ads from March through August. Business software represented almost twenty percent of observed ad impressions. Professional and business services, just eight point five percent, and developer tools, four point six percent. Those are all B2B. Graphite also found click-through rate more than doubled from late March to early September, while fewer than half of ad-eligible users had not been shown an ad. These are panel-derived observations, not OpenAI ad server data.

Paris Childress: All right. It's very interesting here that B2B is dominating ChatGPT Ads because I would guess that the ChatGPT Ads audience are not as much business users because these are people that are not paying for ChatGPT subscriptions. These are the free ChatGPT users. But despite that B2B is still very strong. In fact, B2B SaaS business software alone is almost twenty percent. Very interesting to see that.

Paris Childress: Next up, Yext, Yext's Expanded Scout reaches its announced September thirtieth availability milestone. Yext announced on September first that Expanded Scout capabilities would add brand-level AI visibility and answer engine optimization across more cited sources, initially piloting with enterprise customers ahead of a broader availability on September thirtieth, today. Yext cited a Corporate Ink survey in which eighty-eight percent of CMO/VP-level marketers said leadership or boards were asking about AI visibility versus only thirty-four percent of marketers reporting a defined strategy.

Paris Childress: As of this briefing there's no surface yet of a separate same-day Yext announcement confirming that the broader rollout is fully live. What jumps out here mostly is this statistic: eighty-eight percent of CMOs and marketing VPs said that their leadership and/or their boards were asking about AI visibility, so there's a high degree now of awareness and curiosity at the most senior levels, the most senior decision-making levels in organizations, corporate companies, versus just thirty-four percent of marketers reporting that they have a defined strategy.

Paris Childress: So we're in an early stage here of this channel where there is now a lot of awareness, there's a lot of buzz, but about two-thirds of people don't have a defined strategy. And I believe that the lack of a strategy really creates a massive opportunity for consultants and for platforms like GEOforge to move in and have a very clear strategy, a very clear workflow, and measurement infrastructure in place so that collectively, we can all close this gap between awareness and demand being generated, but a lack, a clear lack of defined strategies.

Paris Childress: Next story. Google closes September's AI Max migration window. Google says eligible campaigns using automatically created assets and campaign-level broad match would automatically migrate to AI Max during September, with eligible upgrades expected to conclude by the end of September. Google reports that its complete AI Max feature suite, search term matching, text customization, and final URL expansion delivered an average of seven percent more conversions or conversion value at a similar CPA/return on ad spend versus search term matching alone in its internal non-retail advertiser data. All right.

Paris Childress: So we can see now that even traditional paid search is moving towards machines deciding which queries, text, and landing pages best match intent.

Paris Childress: And on to our last story of the day, G2 now tracks six hundred and forty-two plus answer engine optimization products. Wow. G2's September category snapshot lists six hundred and forty-two answer engine optimization products with twenty-seven thousand three hundred plus reviews. And that includes Semrush, Profound, Similarweb, Peec AI, Brandlight, Birdeye, SE Ranking, and others represented here. Separately, Brandi AI announced on September twenty-ninth that it was named a high performer in G2's Fall twenty twenty-six answer engine optimization grid.

Paris Childress: The six hundred and forty-two product figure is G2's September first category snapshot. All right, so this is very sobering news for me as a competitor in this category. It's already gotta be much, much bigger than GEO. And I'm coming upon new names practically every day. Six hundred and forty-two answer engine optimization products in this category. And we've been trying to get G2 reviews. This has been a major focus for us of trying to get to the top of this list.

Paris Childress: I think the G2 reviews are one of the most important types of citations that B2B SaaS should be building. And of course, where else would it be more competitive for answer engine optimization than within its own category? Answer engine optimization, or we call it GEO. But answer engine optimization has clearly become a procurement category. It's already extremely crowded. It includes SEO incumbents also reputation platforms, and dedicated AI search vendors. All right. That's it for today's episode 52 of The GEO Show.

Paris Childress: Thank you all for tuning in, and see you on the next one.

See how AI answers describe your brand

GEOforge measures how often ChatGPT, Google AI Overviews and AI Mode cite you, then builds the content and citations that change it.