GEOforge Product · Version 2.0
We audited our own Share of Voice metric across 111 production measurement cycles. Then we shipped the fixes that made our own numbers look worse.
GEOforge 2.0 is live, and it changes two things at once. Anyone can now sign up and start measuring their brand’s visibility inside AI assistants without waiting for an invitation. And the number they’ll see when they get there has been rebuilt to be honest — including in the places where honesty made it look worse.
That second part is the story we actually want to tell, because it is the part almost nobody in this category will tell you.
Every tool in this market will hand you a percentage. “You have 14% share of voice in ChatGPT.” It looks precise. It goes in the board deck.
Here is what that number usually hides.
AI answers are not stable. Ask an assistant the same buyer question twice and you can get two different shortlists. In an internal variance study across 15,580 AI responses, 45 questions and 5 different models, we found the answers never fully settle — and that a given brand appeared in as few as 0.65% of responses. When the thing you’re counting is that rare, a handful of samples tells you essentially nothing.
Most tools ask the AI to grade itself. The common shortcut is to send a question to the model and, in the same breath, ask it to score how well it recommended your brand. A grader marking its own homework is not a measurement.
And the error bar is usually decoration. If a dashboard shows a fixed “±2%” that never changes regardless of how much data sits behind it, that band is a design element, not a statistic. We know, because until last week ours was one too. We removed it.
None of this is malice. It’s what happens when you build a measurement product quickly. The difference is what you do when you find it.
At the end of July we ran a full audit of the GEOforge Share of Voice pipeline — every line of measurement code, and every claim we made about precision — against 111 completed measurement cycles spanning 18 real brands in production. We wrote the findings down, including the findings that were embarrassing, and shipped the fixes in 2.0.
Four things came out of it.
GEOforge splits measurement into three separate jobs that don’t get to influence each other.
| Step | What happens | Who does it |
|---|---|---|
| Ask | Each buyer question goes to the AI assistant as a normal question, with live web access, and gets a natural answer. | The answering model (e.g. ChatGPT) |
| Review | A different AI, from a different company, reads that answer and records the facts: was your brand recommended by name? Was it only cited as a source? Which competitors were named? Was the tone positive or negative? | An independent reviewer model |
| Score | Plain arithmetic turns those facts into one number. No AI involved, so the same answers always produce the same score. | Code |
Two consequences matter commercially. First, the reviewer shares no incentive and no training lineage with the answerer, so it has no reason to flatter it. Second, because scoring is pure arithmetic over recorded facts, we can re-score your entire measurement history when we improve the formula — and show you exactly why a number is what it is, question by question.
The competitor set is yours, and it is closed. The competitors you configure are the entire denominator your Share of Voice is measured against — no hidden exclusions, nothing quietly dropped. That sounds obvious. It wasn’t: fixing it is one of the corrections below.
“Best CRM for startups? I’d suggest Acme” and “…according to acme.com” are both technically a mention. They are worth wildly different amounts to your pipeline.
GEOforge weights them differently: a direct recommendation counts a full point, a citation-only appearance counts 0.3, and appearing both ways counts most of all. One headline number, on a weighting that is fixed, documented, and applied identically to every brand on the chart — so the score means the same thing this week as last, and the same thing for you as for your rivals.
And you can see inside it. Mention share and citation share are reported separately, to every account — on the score card, as their own three-line view on the trend chart alongside the combined score, and as a balance meter on your dashboard. That matters because the two move for different reasons and are fixed by different work. Being cited but rarely recommended means the AI trusts your pages as a source but doesn’t put you on the shortlist; being recommended but rarely cited means the opposite. One number tells you where you stand. The split tells you which problem you have.
The important 2.0 change is that your competitors are now scored on exactly that same formula. Previously the competitor lines on your trend chart showed something simpler — how often a rival was mentioned at all — while your own line showed the weighted score. Two different measures, same axis, and every comparison you drew from that chart was subtly wrong. Now it’s like-for-like.
This is the finding we least enjoyed. Our old margin of error pooled every answer we had collected and divided by the square root of the total. That treats 40 answers to the same question as 40 independent readings of your brand’s visibility. They aren’t independent — answers to one question cluster tightly together — so the arithmetic was wrong for either question a customer might reasonably ask of an error bar.
We replaced it with a band that answers one specific, useful question: if we re-ran this week’s measurement on the same set of questions, how close would we land? It’s built only from how much your answers to each individual question actually varied, and it moves as your own data moves.
“Re-run this week’s measurement on the same questions and you’d land within ±X.” That’s what the number next to your score means now — and where we can’t compute it honestly, we show nothing at all.
Where there genuinely isn’t enough data to compute it — fewer than two questions with enough answers to estimate variation at all — we show no band, rather than a reassuring ±0. We also stopped hiding it: the uncertainty display used to be gated to certain accounts, and it’s now on for everyone, because a bare number with no error bound at all is worse than an honest one.
And one thing that band deliberately does not claim. It isn’t a statement about whether a different set of questions would have produced a different score. Your tracked questions genuinely differ from one another — you might dominate “best tool for enterprise teams” and be invisible in “cheapest option for a startup” — but that spread is a property of what you chose to track, not error in the measurement. Blending the two into a single number would make the ± look authoritative while answering neither question. We keep them separate.
“How precise is this number?” and “did this number move?” are two different questions, and they need two different statistics. Reporting one band and letting it stand in for both is how a measurement product ends up either crying wolf or missing real movement.
So the comparison gets its own statistic, computed on the paired per-question differences between two cycles. Because the question set is held fixed from one week to the next, everything that is a permanent property of that set — the fact that some questions are simply better for you than others — is common to both cycles and cancels out. What’s left is the part that can actually change.
We measured how tight that comparison is, on live production data:
In other words: the measurement itself is stable to about a point. On ChatGPT specifically, our week-over-week change detection has a standard error of roughly 1 percentage point at full measurement depth. That’s the number that licenses a claim like “our visibility went up 4 points after that content push.”
There’s a third statistic for a third question. For a brand that appears in well under 1% of answers, the most defensible thing to report isn’t a weighted score at all — it’s the raw share of answers you appear in, with a confidence interval built for rare events. GEOforge reports that alongside the rest, because for a brand starting from near-invisibility it’s the number that will move first.
This is what we mean by Visibility Delta: not “here’s your score,” but “here is the defensible, measured impact of the work you did.”
Every full-depth GEOforge brand gets, per AI surface, per cycle:
For comparison, a tool sampling 5 or 10 times per question is reading a slot machine and reporting the first pull.
Depth is also what makes the week-over-week comparison usable. Cut measurement depth to a quarter and the band on “did it move?” widens by roughly half again on ChatGPT — the source where change detection is most feasible in the first place. And for a brand appearing in under 1% of answers, estimating that rate at all takes hundreds of samples, however you arrange them. That’s why the depth stays where it is.
An incomplete week now looks incomplete. If a measurement cycle only managed to collect part of its answers, it used to appear on your chart identical to a full week — same dot, same confident comparison arrow. Now it’s flagged, and its comparison arrows are withheld. You will occasionally see a gap where you used to see a number. The number was the problem.
Your measured history can no longer be rewritten. Editing your tracked questions or clearing out an old measurement run used to be able to reach backwards and alter charts you’d already reported to your board. Measured history is now immutable by default. What was measured on a date stays measured on that date, permanently.
Google AI Overviews are graded correctly when there’s nothing there. Google doesn’t always serve an AI Overview. An empty response is not the same as “nobody was visible” — one is a failed observation, the other is real data. Treating the first as a zero drags your score down for no reason. GEOforge now tells them apart.
And a correction we made downward. We found a case where a competitor could silently drop out of the comparison, which made a small number of brands’ Share of Voice look better than it was. We fixed it and re-scored the history, which moved those numbers down. We’d rather hand a client a smaller true number than defend a flattering false one.
The measurement work is why the version number changed to 2.0 in spirit. Self-serve is why it changed on paper.
GEOforge is now self-serve. Sign up, start a 14-day free trial with access to the whole platform, and pick a plan when you’re ready — no sales call required, no invitation to wait for. Sign in with Google if you’d rather not manage another password.
If you’re already a GEOforge customer, nothing about your account changes. No quotas, no paywall, no new charge. This is new-customer news, not a pricing change.
Elsewhere in the release: branded email signatures that attach to every outreach pitch and reply automatically; bounce tracking, so you know when an outreach email couldn’t be delivered and can find another contact; AI crawler tracking that works on any website, behind any CDN, proxy or application, closing the loop from published, to crawled by an AI, to cited in an answer; light and dark mode; and review feedback that lands in the editor’s Comments panel with a link straight to the passage a reviewer meant.
Because “AI visibility” is about to be full of confident percentages with nothing behind them, and the only durable way to be distinguishable is to show your work.
If you’re evaluating any tool in this category — ours included — these are the five questions we think you should insist on an answer to.
We had the wrong answer to some of those a month ago. We wrote down what was wrong, measured the fix against real production data, and shipped it — including the parts that made our own numbers look less impressive. That’s the product. The percentage is just the part you can see.
GEOforge measures your brand across ChatGPT, Google AI Overviews and AI Mode at roughly 1,200 graded answers per surface per cycle — then builds the content and citations that move it. 14-day free trial, no sales call.
Start a free trial →Sources & method. Figures verified against the GEOforge measurement database on 31 July 2026. Corpus: 111 completed measurement cycles across 18 tracked brands, measured June–July 2026, restricted to cycles with more than 200 graded answers and more than 4 questions. Split-half and quarter-split agreement computed on cycles of 40 or more answers per question, comparing brand-level Share of Voice from disjoint subsets of runs. The change-detection standard error is measured on consecutive production cycle pairs, recomputing each delta from a reduced subsample against all runs. The 0.65% brand-appearance rate and the 15,580-response variance study are from earlier internal analysis across 45 questions and 5 models. Figures are specific to the sources named; results differ by AI surface.