CrawlCheck

Findings · 2026-08-20 · By · 0 views

AI visibility scores measure an API, not what anyone sees in ChatGPT

The answers these tools record are real. The label on top of them — “your visibility in AI search” — describes something that was never measured.

They measure something real: what a model API returned for a chosen prompt list on a given day. What they usually claim is different — your visibility in ChatGPT, or in AI Overviews. Three substitutions sit between the two. An API is not the consumer app, which adds retrieval, tools, personalisation and memory. A hand-picked prompt list is not a property of your brand. And one answer is one draw from a distribution that changes between runs. The number is not fake; the sentence wrapped around it is.

An AI-visibility tool asks the Gemini API who is the best tree service in Denver, records what comes back, counts how often your brand appears, and reports a percentage. Every step of that is real. The model genuinely produced those words. Nothing is fabricated.

What is false is the sentence wrapped around the number: here is your visibility in AI search. The measurement is real; the claim about what it represents is not. Three substitutions sit between the two.

1. The surface: an API is not the app #

The OpenAI API is not ChatGPT. ChatGPT performs retrieval, calls tools, and carries personalisation and memory that the raw API does not reproduce. Google's AI Overviews is a different system again from the Gemini API — different retrieval, different ranking, different presentation, sitting inside a search results page rather than a chat window.

Measuring one and reporting the other is the substitution nobody names out loud. It is not fraud; there is no public API for the consumer surfaces, so every vendor in this category measures the thing they can reach and labels it as the thing you asked about.

2. The sample: the prompt list is somebody's guess #

Even a perfect measurement only covers the prompts somebody chose to type. Nobody has the distribution of what buyers actually ask an assistant — not the vendors, not the model providers, not us. A share-of-voice percentage is a statistic over a hand-picked prompt list, presented as a property of your brand.

Change the list and the number changes. That is not a flaw in the arithmetic. It is what the arithmetic was always describing.

3. The moment: one answer is one draw #

Ask the same model the same question twice and you can get different answers. Ask it next week, after a model update, and the difference is larger. A single run tells you almost nothing on its own. The churn between runs is the signal — how stable your presence is across dated repetitions — and it is the part most tools do not report, because it makes the headline number look unsteady.

The test that separates a measurement from a label #

Ask what would have to be true for the number to be wrong, and how you would find out. If there is no answer, it is not a measurement. A percentage you cannot check is a claim about a thing nobody measured.

What survives that test is narrow and unglamorous: name the model, name the mode (grounded or from weights), name the date, report presence or absence rather than a rank, and publish the churn. Then a reader can disagree with you on evidence.

What a defensible substitute looks like #

If the prompt-sampling number cannot carry the claim, something has to. The half of the question that is directly observable is the retrieval side: not whether an assistant named you, but whether it could have. That is a fetch, and a fetch can be repeated.

The measurement is three questions asked of your own site, each with a yes-or-no answer and a number behind it. Reach: does a response arrive at all for a non-browser client. Read: is there readable text in the bytes, before anything executes. Quote: is there a self-contained span in that text worth lifting into an answer. CrawlCheck fetches your page as each named AI crawler and compares the replies.

What that produces is a status code, a ratio and a diff — three things a reader can re-run against their own site and disagree with. What it does not produce, and must not be sold as, is evidence of citation. It shows the input was available. It says nothing about the output.

We took a Review node off our own directory #

This site publishes a measured score for a list of vendor websites. For a while, the structured data under that list emitted each score as a schema.org Review carrying a reviewRating.

That markup was the exact substitution this post is about. A review is a claim that a person formed an opinion and attached a rating to it. Nobody did. A scanner fetched a URL and computed a number, and the markup dressed it as a verdict — the machine-readable version of writing “your visibility in AI search” over an API sample.

It now emits an Observation instead, with the score and its components as PropertyValue measures and an isBasedOn link to the evidence report the number came from. Same number, honest type, and a reader who follows the link lands on the measurement rather than on a rating. The rule that came out of it: if the markup would let a reader believe a person formed an opinion, and no person did, the markup is wrong — however accurate the number inside it.

A number that moves needs a date and a version #

Two further rules follow from the same principle, and both cost features.

The comparison view refuses to rank sites unless the readings are actually comparable. When they are not, it prints “Limited comparison” and names the condition that failed, rather than rendering a table that looks authoritative and is not. A ranking assembled from readings taken weeks apart under different instruments is a chart of when each site happened to be scanned.

Published snapshots of that dataset are never rewritten. If a published number later looks wrong, the record of what was published stands and a correction is logged beside it. A figure you cannot re-fetch exactly as it stood is not a measurement; it is a memory of one.

Five questions, and four of them have factual answers #

For any AI-visibility number, from anyone, including this site:

QuestionWhat a straight answer sounds like
Which surface?“The Gemini API” — not “AI search”
Which model, and which mode?A version string, and grounded or from weights
Which date?A date, not “recently”
What was the sample?The prompt list, or the URL and the client identity
Can I re-run it?The method, in enough detail to disagree with

The first four are facts a vendor either has or does not. The fifth decides whether you were sold a measurement or a label.

Why the API number still has a use #

None of this makes prompt sampling worthless. Run against a fixed list, with the model and mode pinned, on a schedule, it is a reasonable change detector: the absolute percentage is close to meaningless, but a series of them can tell you that something moved. That is a genuinely useful signal, and it is a smaller claim than the one usually printed above it.

The failure is never the measurement. It is the rename between the measurement and the headline — and the rename is easy to catch, because it always involves a word the method never used.

What a citation claim would actually require #

It is worth spelling out what would have to be observable for “you appear in AI answers” to be a measurement rather than an inference, because the list is short and none of it is available from outside.

You would need the surface itself, not an API standing in for it. You would need the retrieval corpus that surface drew from at that moment, including whatever personalisation and memory applied to that user. And you would need attribution for the composed answer — which source contributed which claim — which the systems themselves do not expose and in most architectures cannot.

Nobody outside the providers holds any of the three. That is not a criticism of the tools; it is the reason the honest version of the product is narrower than the marketed one. A vendor who says “we sample a model and report what came back” is describing something true. A vendor who says “we measure your share of AI answers” is describing something that would require access nobody has.

The distinction is not academic. It decides what you can do when the number moves: with a sample you can ask whether the prompt list, the model version or the mode changed, and re-run it. With a label you can only ask the vendor.

Three substitutions between the number and the claim #

What was measuredWhat the label saysWhy they are not the same
The Gemini or OpenAI APIyour visibility in ChatGPT or AI Overviewsthe consumer surfaces add retrieval, tools, personalisation and memory that the raw API does not reproduce
A hand-picked prompt lista property of your brandnobody holds the distribution of what buyers actually ask; change the list and the number changes
One answer, on one daya stable share of voicethe same prompt gives different answers between runs, and larger differences after a model update

Ask what would have to be true for the number to be wrong, and how you would find out. If there is no answer, it is not a measurement.

Where this leaves us, including the part that is ours #

Our own answer ledger is not exempt. It queries an API too. When it produces figures, they will carry the model name, the mode and the date on every row, and they will describe what that model said on that day — not what a person saw. Today it produces nothing at all, because the pipeline is not running, and we would rather say that than publish a number we cannot defend.

The half of this problem that is directly observable is the audit. What your site serves to a crawler is not a proxy or a sample: it is the bytes, fetched as each named crawler. The causes are measurable. Only the effects are estimated — and a tool that tells you which is which is worth more than one that reports a confident percentage for both.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Do AI visibility tools measure what people see in ChatGPT?

No. They query provider APIs, and an API is not the consumer app. ChatGPT adds retrieval, tools and personalisation the raw API does not reproduce, and AI Overviews is a different system again.

Are AI visibility scores fake?

The recorded answers are real model output. What is unsupported is the label placed on them, because the surface, the prompt sample and the moment all differ from what the number claims to describe.

What makes an AI visibility number checkable?

Naming the model, the mode, and the date, reporting presence or absence rather than a rank, and publishing the churn between dated runs so a reader can see how stable the result is.

Can any tool measure my visibility in ChatGPT?

Not from an API. The consumer app adds retrieval, tools, personalisation and memory that the raw model API does not reproduce, so an API sample describes the API. A tool can honestly report what a named model returned, in a named mode, for a named prompt list, on a date, and that is a useful change detector. It is not the same statement as visibility in the product.

What should an AI visibility report contain?

The surface queried, the model and the mode, the date, the sample it was computed over, and enough method to re-run it. Four of those five are facts a vendor either has or does not. The fifth decides whether the number can be disagreed with, which is the working definition of a measurement.

Is share of voice a real metric for AI search?

It is real about the prompt list it was computed over, on the day it was computed. It is not a property of a brand, because nobody holds the distribution of what buyers actually ask and the same prompt returns different answers between runs. Treated as a series against a fixed list it can show movement. Treated as a percentage it invites a conclusion the method does not support.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works