A longevity diagnostics brand ran 46,672 measured AI answers over sixteen weeks. Share of Voice rose 16% on Google's AI surfaces and fell 11% on ChatGPT — and the category-level spread was 54 points. What a single number would have hidden.
Over sixteen weeks this brand's AI visibility went up 51% on one engine and down 11% on another. Both numbers are true. A platform that sampled one engine, once, would have sold the client either a victory lap or a crisis — and had the data to back whichever it picked.
The client sells an at-home biomarker test, direct to consumers and through a network of functional-medicine, hormone and longevity clinics. Scientific credibility is not their problem: the science behind the test is peer-reviewed and the company has university research partners. Their problem is that the buying journey now begins inside an AI answer, and that answer is assembled from other people's pages.
Their category — biological age, chronic inflammation, hormone and menopause testing — is exactly the kind of high-consideration, high-anxiety question buyers now take to an AI assistant long before they visit a website. Whoever the model names, wins the shortlist.
A large, mature content library and real scientific authority — but a Share of Voice on ChatGPT of 22.4% at first measurement in May 2026, drifting down week over week. No structured view of which topics were losing ground, or to whom.
Nobody could say whether AI visibility was improving — because it was improving and deteriorating at the same time, on different engines, in different topic clusters. Any single headline number would have been a coin flip wearing a metric's clothes.
141 source documents — research summaries, clinical protocols, practitioner material, product documentation — chunked into 2,798 embedded passages, roughly 1.05 million tokens. Everything written afterwards was written from this, not from the open web.
30 buyer prompts across 13 topic categories, each run 40 times against ChatGPT, Google AI Mode and Google AI Overviews on a weekly cycle. 3,600 graded answers per cycle. No single-shot spot checks.
52 pieces published in 16 weeks — 47 FAQ-format answer pages and 5 long-form articles. Average information gain 7.9/10, average factual accuracy 9.95/10, every claim traceable to an ingested source.
38,736 cited URLs harvested from the measured answers and scored for reachability and influence. 20,180 flagged as citation opportunities. 69 outreach emails sent so far, 9 replies.
Ask an AI engine the same question twice and you can get two different shortlists. At one sample per prompt, the margin of error on a Share of Voice reading is wider than any weekly movement you would ever want to report. At 40 runs per prompt across 30 prompts, a single cycle rests on 1,200 graded answers per engine — enough that a 2-point move means something.
Share of Voice by engine, weekly cycles at 40 runs per prompt. The AI Overviews line carries one gap: there was no completed AI Overviews cycle in the week of 27 July. The 20 July batch finished at partial depth and was re-run on 24 July; the re-run is what is plotted.
| Engine | 5–6 Jul | 7 Sep | Change |
|---|---|---|---|
| Google AI Overviews | 26.2% | 39.5% | +50.6% |
| ChatGPT | 19.1% | 17.0% | −11.3% |
| Google AI Mode | 29.4% | 25.2% | −14.5% |
| Blended, all three | 24.9% | 27.2% | +9.0% |
Google AI Mode did not simply decline. It climbed to 40.0% on 10 August and then gave most of it back. A spot check that morning would have shown a 36% gain. The same check on 7 September shows a 15% loss. Neither is wrong; both are useless on their own.
The blended figure — up 9.0% — is the only number here that survives contact with the next cycle, and it is the least dramatic of the four.
On Google's two AI surfaces, where the tracked competitor set appears most often, the client's share rose while the field's fell away. The interesting part is that the client did not appear in more answers — presence stayed flat at about 35.5%. What changed is how much of each answer they held: where the client is named at all, their share of the answer went from 77.7% to 90.9%.
The mechanism matters, and it is not that rivals lost ground inside the answers they were in. Where a competitor appears, the average number of competitors named is 1.25 — identical in July and September. What halved is the number of answers that name a rival at all: 1,011 of 2,400, down to 503 of 2,388.
Blended across all three engines, the 30 tracked prompts break into 13 topic categories. Eight gained, five lost. The spread between the best and worst category is 54 points — which is why a brand-level average, however carefully sampled, is a planning instrument and not a diagnosis.
Share of Voice by topic category, blended across the three tracked engines. First full 40-run cycle (5–6 Jul) versus latest (7 Sep).
The obvious explanation — we published, so we won — does not survive the data. Eighteen of the 52 pieces were inflammation-related, and that category moved 0.8 points. Twelve of the client's published pages were cited under product-comparison prompts, and that category lost 17.6 points. Volume was not the variable.
What separated the winners from the losers was who already owned the citation surface. In menopause and hormones, only 13% of the pages AI engines cited belonged to a competitor — the surface was open, and owned content walked into it. In inflammation and ageing, 45% of the cited pages were competitor-owned. Publishing into a surface a rival already occupies moves almost nothing.
In the comparison categories the client's own pages were being cited — 4,775 citation events across 12 pages — and the client still lost share. Being cited is not the same as being recommended.
When a buyer asks which test is best, the engine builds its answer out of third-party roundups: 664 editorial listicles sit in that citation surface. You cannot publish your way onto someone else's “best of” list. You have to get onto it.
The loop that did work, worked fast. Of the 52 pieces published, 35 have been cited in a measured AI answer — and the median piece took under six days to get there.
Across the whole programme the engines cited 38,736 distinct URLs from 5,105 domains. 759 of those domains mention the client somewhere in the answer they support. That gap — 5,105 domains shaping the category's answers, 759 of them aware the client exists — is the size of the remaining opportunity, and it is what the outreach programme exists to close.
Product comparison and at-home testing both fell double digits. These are the highest-intent prompts in the set and the hardest to win with owned content alone. They need third-party placement, and that work has barely started.
20,180 scored citation opportunities, 69 emails sent. Nine replies came back — a 13% reply rate that suggests the targeting is sound and the volume is the constraint. This is the single largest unexploited lever in the account.
The last document upload was 39 days before this report. The corpus that made the content credible is ageing, and the categories built on the oldest material are the ones drifting.
Sixteen weeks of publishing moved ChatGPT Share of Voice from 19.1% to 17.0%. Whatever is working on Google's surfaces is not transferring, and the reason is not yet established.
A 51% gain and an 11% loss ran side by side for four months. Reporting either alone would have been accurate and misleading at once.
The gap between the best and worst topic cluster was 54 points. Brand-level Share of Voice tells you how you are doing; category-level tells you what to do on Monday.
Where rivals owned 13% of the cited pages, owned content added 36 points. Where they owned 45%, it added 0.8. Check who holds the surface before you commission the brief.
Median 5.9 days from publish to first AI citation, 82% inside two weeks. The feedback loop in AI search is short enough to steer by.