Three AI Engines, Three Different Verdicts

A longevity diagnostics brand ran 46,672 measured AI answers over sixteen weeks. Share of Voice rose 16% on Google's AI surfaces and fell 11% on ChatGPT — and the category-level spread was 54 points. What a single number would have hidden.

Client
Longevity diagnostics brand (anonymised)
Industry
Consumer & clinical biomarker testing
Period
19 May – 7 Sep 2026
Content Published
52 pieces
46,672
AI answers graded
19 cycles · 3 engines
+16.1%
Share of Voice, Google AI surfaces
27.8% → 32.3%
2.0×
Lead over tracked competitor set
was 0.9× in July
5.9
Median days, publish → first AI citation
82% cited inside 14 days

Over sixteen weeks this brand's AI visibility went up 51% on one engine and down 11% on another. Both numbers are true. A platform that sampled one engine, once, would have sold the client either a victory lap or a crisis — and had the data to back whichever it picked.

Context & Starting Point

The client sells an at-home biomarker test, direct to consumers and through a network of functional-medicine, hormone and longevity clinics. Scientific credibility is not their problem: the science behind the test is peer-reviewed and the company has university research partners. Their problem is that the buying journey now begins inside an AI answer, and that answer is assembled from other people's pages.

The Opportunity

Their category — biological age, chronic inflammation, hormone and menopause testing — is exactly the kind of high-consideration, high-anxiety question buyers now take to an AI assistant long before they visit a website. Whoever the model names, wins the shortlist.

Starting Conditions

A large, mature content library and real scientific authority — but a Share of Voice on ChatGPT of 22.4% at first measurement in May 2026, drifting down week over week. No structured view of which topics were losing ground, or to whom.

THE CORE CHALLENGE

Nobody could say whether AI visibility was improving — because it was improving and deteriorating at the same time, on different engines, in different topic clusters. Any single headline number would have been a coin flip wearing a metric's clothes.

What We Actually Ran
01
Knowledge ingestion

141 source documents — research summaries, clinical protocols, practitioner material, product documentation — chunked into 2,798 embedded passages, roughly 1.05 million tokens. Everything written afterwards was written from this, not from the open web.

02
Measurement design

30 buyer prompts across 13 topic categories, each run 40 times against ChatGPT, Google AI Mode and Google AI Overviews on a weekly cycle. 3,600 graded answers per cycle. No single-shot spot checks.

03
Content production

52 pieces published in 16 weeks — 47 FAQ-format answer pages and 5 long-form articles. Average information gain 7.9/10, average factual accuracy 9.95/10, every claim traceable to an ingested source.

04
Citation programme

38,736 cited URLs harvested from the measured answers and scored for reachability and influence. 20,180 flagged as citation opportunities. 69 outreach emails sent so far, 9 replies.

Why 40 Runs and Not One

Ask an AI engine the same question twice and you can get two different shortlists. At one sample per prompt, the margin of error on a Share of Voice reading is wider than any weekly movement you would ever want to report. At 40 runs per prompt across 30 prompts, a single cycle rests on 1,200 graded answers per engine — enough that a 2-point move means something.

Matched sampling. Cycles are only compared with cycles run at the same depth. This programme ran at 50 runs per prompt from May to June, then settled at 40 from 5 July onward. Every before-and-after figure on this page is drawn from the 40-run series, so nothing here is an artefact of changing how hard we looked. Blended figures are row-weighted across all graded answers in the cycle, not an average of engine averages. Tracked engines are ChatGPT, Google AI Mode and Google AI Overviews.
Three Engines, Three Verdicts
0% 10% 20% 30% 40% 5–6 Jul 20–24 Jul 3 Aug 17 Aug 31 Aug 7 Sep
Google AI Overviews
Google AI Mode
ChatGPT

Share of Voice by engine, weekly cycles at 40 runs per prompt. The AI Overviews line carries one gap: there was no completed AI Overviews cycle in the week of 27 July. The 20 July batch finished at partial depth and was re-run on 24 July; the re-run is what is plotted.

Engine5–6 Jul7 SepChange
Google AI Overviews26.2%39.5%+50.6%
ChatGPT19.1%17.0%−11.3%
Google AI Mode29.4%25.2%−14.5%
Blended, all three24.9%27.2%+9.0%
What this means

Google AI Mode did not simply decline. It climbed to 40.0% on 10 August and then gave most of it back. A spot check that morning would have shown a 36% gain. The same check on 7 September shows a 15% loss. Neither is wrong; both are useless on their own.

The blended figure — up 9.0% — is the only number here that survives contact with the next cycle, and it is the least dramatic of the four.

Competitive Position

On Google's two AI surfaces, where the tracked competitor set appears most often, the client's share rose while the field's fell away. The interesting part is that the client did not appear in more answers — presence stayed flat at about 35.5%. What changed is how much of each answer they held: where the client is named at all, their share of the answer went from 77.7% to 90.9%.

The mechanism matters, and it is not that rivals lost ground inside the answers they were in. Where a competitor appears, the average number of competitors named is 1.25 — identical in July and September. What halved is the number of answers that name a rival at all: 1,011 of 2,400, down to 503 of 2,388.

Client vs. tracked set — Google AI surfaces
Client Share of Voice27.8% → 32.3% +16.1%
Tracked competitors, combined32.2% → 16.5% −48.9%
Largest single competitor14.1% → 7.6% −46.0%
Client lead over the field0.9× → 2.0× +127%
Answer composition — Google AI surfaces
Answers graded2,400 → 2,388 matched
Answers naming the client35.8% → 35.5% flat
Answers naming any rival1,011 → 503 −50.2%
Rivals named, where any appear1.25 → 1.25 unchanged
Client share of answers it appears in77.7% → 90.9% +17.0%
Where the Movement Actually Happened

Blended across all three engines, the 30 tracked prompts break into 13 topic categories. Eight gained, five lost. The spread between the best and worst category is 54 points — which is why a brand-level average, however carefully sampled, is a planning instrument and not a diagnosis.

Topic category
Change in Share of Voice
Delta
Menopause & hormones10.2% → 46.5%
+36.3 pts
Longevity trends33.5% → 56.1%
+22.6 pts
Clinical integration32.4% → 54.0%
+21.6 pts
Corporate wellness18.7% → 23.9%
+5.2 pts
Biomarker comparison55.4% → 59.4%
+4.0 pts
Disease risk & prevention35.1% → 38.5%
+3.4 pts
Inflammation & ageing7.3% → 8.1%
+0.8 pts
How-to queries0.0% → 0.4%
+0.4 pts
Intervention tracking8.8% → 3.9%
−4.9 pts
Biological age education10.1% → 1.5%
−8.6 pts
Immune ageing36.7% → 26.3%
−10.4 pts
At-home testing40.0% → 23.6%
−16.4 pts
Product comparison45.3% → 27.7%
−17.6 pts

Share of Voice by topic category, blended across the three tracked engines. First full 40-run cycle (5–6 Jul) versus latest (7 Sep).

Why Those Categories, and Not the Others

The obvious explanation — we published, so we won — does not survive the data. Eighteen of the 52 pieces were inflammation-related, and that category moved 0.8 points. Twelve of the client's published pages were cited under product-comparison prompts, and that category lost 17.6 points. Volume was not the variable.

What separated the winners from the losers was who already owned the citation surface. In menopause and hormones, only 13% of the pages AI engines cited belonged to a competitor — the surface was open, and owned content walked into it. In inflammation and ageing, 45% of the cited pages were competitor-owned. Publishing into a surface a rival already occupies moves almost nothing.

Open surface — content wins
Menopause & hormones+36.3 pts
Cited pages owned by a rival13%
Third-party roundups in the surface1%
Client pages cited here12
Contested surface — content stalls
Product comparison−17.6 pts
Cited pages owned by a rival16%
Third-party roundups in the surface20%
Client pages cited here12
The uncomfortable finding

In the comparison categories the client's own pages were being cited — 4,775 citation events across 12 pages — and the client still lost share. Being cited is not the same as being recommended.

When a buyer asks which test is best, the engine builds its answer out of third-party roundups: 664 editorial listicles sit in that citation surface. You cannot publish your way onto someone else's “best of” list. You have to get onto it.

From Published to Cited

The loop that did work, worked fast. Of the 52 pieces published, 35 have been cited in a measured AI answer — and the median piece took under six days to get there.

Content → citation
Pieces published52
Cited in a measured AI answer35 67%
Citation events traced to them7,052
Share of all client-domain citations16.5%
Time to first citation
Median5.9 days
Fastest2.8 days
Cited within 14 days28 of 34 82%
Slowest62.7 days

Across the whole programme the engines cited 38,736 distinct URLs from 5,105 domains. 759 of those domains mention the client somewhere in the answer they support. That gap — 5,105 domains shaping the category's answers, 759 of them aware the client exists — is the size of the remaining opportunity, and it is what the outreach programme exists to close.

What Didn't Work
The comparison categories

Product comparison and at-home testing both fell double digits. These are the highest-intent prompts in the set and the hardest to win with owned content alone. They need third-party placement, and that work has barely started.

Outreach under-run

20,180 scored citation opportunities, 69 emails sent. Nine replies came back — a 13% reply rate that suggests the targeting is sound and the volume is the constraint. This is the single largest unexploited lever in the account.

A stale knowledge base

The last document upload was 39 days before this report. The corpus that made the content credible is ageing, and the categories built on the oldest material are the ones drifting.

ChatGPT is still flat

Sixteen weeks of publishing moved ChatGPT Share of Voice from 19.1% to 17.0%. Whatever is working on Google's surfaces is not transferring, and the reason is not yet established.

What This Programme Proves
One number cannot hold three engines

A 51% gain and an 11% loss ran side by side for four months. Reporting either alone would have been accurate and misleading at once.

Category is the unit of work

The gap between the best and worst topic cluster was 54 points. Brand-level Share of Voice tells you how you are doing; category-level tells you what to do on Monday.

Content wins open surfaces, not contested ones

Where rivals owned 13% of the cited pages, owned content added 36 points. Where they owned 45%, it added 0.8. Check who holds the surface before you commission the brief.

Citation velocity is measurable in days

Median 5.9 days from publish to first AI citation, 82% inside two weeks. The feedback loop in AI search is short enough to steer by.

See what your category's answers are made of
GEOforge measures Share of Voice across ChatGPT, Google AI Mode and Google AI Overviews — then builds the content and the citations that change it.
Start your free trial