See What AI Says
METHOD v1.0 · frozen 2026-08-27

How we measure

Published in full, on purpose. If you can't check how a number was produced, it isn't a measurement — it's a claim. Every report prints the version of this document it was run under.

The engines

3 surfaces, chosen because they are what your customers actually use. 3 are assistants we ask directly; 0 are read off live Google results.

SystemModeWhy it's in
Gemini with Google Search The default assistant on Android and across Google's apps
Claude from memory, no web access Heavily used by business buyers
Gemini from memory, no web access The control arm — the gap against grounded Gemini is the finding

What we don't query yet

We would rather list this than let you assume it. These are built and disclosed but not running, so they are counted in none of the totals on this page and no report claims them:

SystemStatus
ChatGPT (with web search)Not available on Vertex, so it needs a second vendor. Reached through a scraper of the real interface rather than the OpenAI API, so every answer comes with a link you can open yourself.
Google AI Overviews (live search results)
Google AI Mode (live search results)

Until it runs, treat every conclusion in a report as a statement about the surfaces in the first table. They lean towards Google's models, and answers there do not automatically transfer to an assistant we have not measured.

We also run every question a second time without web access. The gap between the two is the useful part: if a model names you only when it can search, the internet knows you but the model doesn't. If it names you both ways, you're in the model's baseline knowledge — a much stronger position, and a much slower one to earn.

The 23 questions

Generated from your site, then shown to you for correction before we run anything.

Google has described how its AI features issue multiple related sub-queries behind a single search — what it calls query fan-out. One question is therefore not one test. Asking 23 across the full intent ladder is how we cover the spread rather than sampling one point of it.

Do you query the API or the real chat interface?

Today, the API. All 3 of our live surfaces are queried through official model APIs. Reading the Google results page as a customer sees it is built and disclosed — it is in the table above — but it needs a paid search-data vendor we have not switched on, so no report currently contains it. That makes the argument below apply to everything we publish, which is why we would rather have it out here than in a footnote.

Some vendors market interface scraping as strictly better, and on one point they're right: a chat interface carries personalisation, session history and A/B tests that an API call doesn't see. If what you want to know is "what did one specific logged-in person see on Tuesday", an API cannot tell you and we won't pretend otherwise.

There is a number attached to that, published by the vendor with the strongest commercial reason to publish it. Surfer ran 1,000 prompts through ChatGPT's real interface and through the OpenAI API side by side and found that only 24% of the brands named overlapped between the two, and only 4% of the cited sources. On Perplexity they put source overlap at 8%. Taken at face value that says an API answer and a screen answer are close to different populations, and we are not going to bury a figure because it is aimed at us. (Surfer, 3 Feb 2026.)

Two things we would want you to know before you weight it. It is a study of ChatGPT against the OpenAI API; our model surfaces are Gemini and Claude, and nobody has published the equivalent for any of them, so the 24% is suggestive for us rather than measured on us. And Surfer did not publish how they decided two brand names or two URLs were the same thing, which is most of the work in any overlap figure. We have not reproduced it.

We query the API deliberately, for the thing it is better at: reproducibility. Same prompt, same parameters, no retrieval noise from someone else's session — which is the only way a second run three months later means anything, and re-measurement is what we're actually selling. Scraped interface results cannot be reproduced even in principle, which is also why no vendor doing it publishes a confidence interval.

So read our model-surface numbers as the model's tendency under controlled conditions, not as a transcript of any real user's screen. We have never claimed the second thing and this page is where we say so. There is no section of the report that gives you the "what a user sees" reading, because the surfaces that would carry it are not running. When they are, this paragraph will point you at them.

What the argument does not touch: the blocking findings. Whether your pages are marked noindex or nosnippet, whether an AI crawler is disallowed in your robots.txt, whether your address and phone number agree across the web — none of that changes depending on how a model was queried, and those are the findings that most often turn out to be the reason a business is invisible. When your report opens with blocked, no scraping-versus-API dispute affects it. When we add ChatGPT it will be through a scraper of the real interface rather than the OpenAI API, for the reason on this page.

The verdict, and the rule that picks it

Your report opens with one of exactly four answers to the question you bought it to settle: is this worth your money right now? The answer is chosen mechanically from the numbers, not written by hand, so the same run always produces the same verdict and you can check the working.

VerdictWhen we say it
Leave this alone You were named in 33.0% or more of the answers to questions where a customer is choosing a supplier (tiers L3 and L4). This is checked first, because the most valuable thing this report can tell you is to keep your money.
Something is blocking you We found one of 6 specific settings that make a page ineligible for AI answers outright — a noindex or nosnippet directive, an AI crawler blocked in robots.txt, or a page that only exists after JavaScript runs. A known mechanical cause outranks any statistical reading of the same run, because you can verify it yourself and fix it once.
Invisible, but the causes look cheap You are not being named, and there is evidence the engines can already reach your material — they read your site and recommended someone else, or you surface occasionally. Better position than it feels like.
Invisible, and we don't think spending fixes it No mechanical cause, no sign the engines are reaching you, competitors named throughout. We will say so. It is the least comfortable of the four and the one most likely to save you a year of retainer fees.

Each verdict carries a confidence level with the reason attached. We borrow the floor from Trakkr, who publish the only public statistical rule we could find in this field: under 20 mentions, a rate is a direction and not a measurement. Below that line your report says so rather than rounding up to confidence it hasn't earned. If any part of the method could not run for your site, the confidence drops a step automatically, whatever the numbers say.

How a mention is counted

String matching is not enough, and this is where most free tools go wrong. Ask “tell me about Acme Dental” and the answer will contain “Acme Dental” whether or not the model knows anything. That's an echo, not recognition.

Every response is classified by a separate model pass into one of:

ResultMeaning
RecommendedNamed as an option someone should consider
Mentioned onlyNamed, but not as a recommendation
Cited, not recommendedYour site was used as a source while a competitor got recommended. This one hurts, and it's common.
EchoedYour name appears only because it was in the question
AbsentNot there

Alongside each: your position in the list when named, and your share of voice against every other business that came up.

Repeat sampling

AI answers drift. Independent tracking puts week-to-week change in Google AI Overview content near 70%, and citation sets shift with every model update. You pick the 3 questions that matter most to you; we ask each of those 3 times and report how consistent the answer was. A finding that held all 3 times and a finding that appeared once are labelled differently in your report.

We do not repeat all 23 questions — that would roughly triple the cost of the report without changing any decision you'd make. We also don't sample across multiple days, because you'd be waiting a week. Both limits are stated in your report.

If you serve a local area

Google describes local ranking as relevance, distance and prominence. We measure the first and third directly. For distance, we run your key queries from several coordinates around your address and record where in the map pack you appear from each — which shows you the practical edge of your reach rather than a theoretical radius.

We also check your Google Business Profile primary category (the single highest-leverage field there is), your review volume, rating, recency and response rate, and whether your name, address and phone number agree across your own site and the directories AI reads.

What Google itself says

There are “no additional requirements … nor other special optimizations necessary” for a page to appear in Google's AI features.
— Google Search Central, AI features documentation

We quote this on purpose, even though it undercuts the premise of the entire GEO industry. Google's position is that a page needs to be indexed and eligible to show a snippet, and that you don't need to create new machine-readable files or special markup. We think that's broadly right.

Which is exactly why the valuable thing is not optimization advice — it's finding out what AI currently says about you. Nobody can tell you that without running the queries, and running them properly costs money and takes 150 API calls. That's the product.

What we cannot measure

Stated here rather than buried, and repeated in an appendix in every report:

Version policy

This method is frozen. Question counts, the classification rules, and every metric definition are fixed at v1.0. If we change any of them, the version number changes, and reports run under the old version keep saying so — so two reports carrying the same version stamp are always comparable.

The best way to judge the method is to read a report produced by it.

Read a complete sample report

We'd like to use Google Analytics to count how many people read this page and how many go on to buy a report — nothing beyond that. It writes a cookie to your browser and sends your IP address to Google, so we have to ask first. Saying no changes nothing about how this site works for you. What it collects.