Findings · 2026-08-19 · By VSNARY | Emmanuel Orta · 0 views
Firewall or website? When an AI visibility scanner grades your WAF
We published a failing grade for a competitor. The grade was wrong, and the way it was wrong is the most useful thing this scanner has taught us.
For several hours, this site published a report grading a competitor 15 of 100 — an F, with findings like ‘no sitemap’ and ‘robots.txt not served’. All of it was false. From an ordinary residential connection, their robots.txt is a normal 251-byte file that declares a sitemap. What our scanner had measured was their hosting firewall challenging our datacentre address on every path, at HTTP 200, in identical ~200-byte responses.
Running it twice proved nothing #
We re-ran the scan and got the same 15. That felt like confirmation. It was not: a deterministic block reproduces perfectly, so two identical readings from one vantage point are one reading. Nothing about running it twice tested the thing that was actually wrong.
The tell #
It was hiding in plain sight: the site answered a URL that cannot exist exactly the way it answered its homepage. No real website does that. A site that returns the same bytes for /pricing and /this-path-was-generated-at-random is not describing itself — something in front of it is speaking on its behalf.
The scanner now checks for that on every scan. When it finds it, the scan stops and the report says the site refused measurement — level refused, no grade, no score — because a refusal is not a bad site. It is a fact about our vantage point, and publishing it as their defect would be exactly the error we made.
What we see, what it means, what we say #
The same bytes can mean three different things depending on where they turn up, and the report has to say a different thing in each case. Collapsing them is how the 15 happened.
| What the scanner received | What it means | What the report says |
|---|---|---|
| Homepage and a random path answer with the same challenge body | The edge is refusing this address, not describing the site | Refused. No grade, no score, no findings. Named as our limit. |
Homepage is real; robots.txt is a challenge page at 200 | Machine files are behind a wall the page is not. Real crawlers get this too. | A finding on the site, because it is what any crawler receives. |
That challenge page carries cf-cache-status: HIT | A transient challenge was cached and is now served to everyone | A more serious finding: the challenge is pinned at the edge. |
| One named crawler identity is refused, the browser identity is not | Policy, not breakage | Reported per identity, not counted against the site; a deliberate block is a choice. |
The second row is the one people expect to be unfair and is not. A robots.txt that answers 200 with a verification interstitial is served that way to Googlebot as well; a validator reads it as an empty policy, which is to say everything allowed. That is a defect on the site, and the scanner says so. The first row is the one that was unfair, and it is the one that now produces no number at all.
Three outcomes, never two #
Every check in this scanner can end one of three ways: true, false, or unverifiable. The 15 was what happens when unverifiable is silently filed as false. Since then, an unverifiable result never counts against a site, and the report states which checks could not be verified and why. What each crawler receives is measured the same way: a 403 to one identity from our address is reported as a refusal of us, never as proof of what the real crawler gets.
One reading is an anecdote #
The other half of the fix is time. A single refusal can be a transient fault at the edge, so a refusal seen once is reported — the measurement happened and the reader is entitled to it — but labelled seen once, not yet reproduced. Every scan records the day it looked, so the report can say refused on 3 of the last 4 days we looked with a denominator rather than a rate over an unstated sample.
The exception is deliberate: a refusal caused by identity verification — a site that checks whether a request claiming to be Googlebot came from Google’s published ranges — reproduces every day, because that is the point of it. Recurring does not turn that into a block; it stays reported as a correct refusal of an address that is not Google’s.
Where it stands today #
We deleted the false report and the corpus record behind it. The corrected behaviour is public and it has a page: the vendor directory lists every SEO and AI-visibility tool on its own measurement, and as of this writing two of the twenty-five sit at refused this scanner — not graded, with no number next to them. They are not marked down for it. A vendor whose edge challenges datacentre traffic has made a defensible choice, and a scanner that punished it would be measuring its own address.
If your security layer challenges datacentre traffic, our scanner will tell you that it was refused — and it will not pretend the refusal was your website.
The lesson generalises past this scanner #
Any tool that fetches your site from cloud infrastructure can be fed a challenge page and grade it as content — most will. The audit that told you your sitemap was missing may have been reading a captcha. Ask your auditor one question: how do you know what you measured was my website?
Live, as you read this: the corpus now holds 189,169 domains across 3,137 scans. The figures in this piece were measured on the date above; this line is not.
How the refusal is detected, and what it is not #
The control is a path generated at random for that scan. If it answers with the same status, the same byte count within a small tolerance, and the same challenge language as the homepage, nothing we fetched was the site. The record is stored with refused: true, no grade, no section scores, and a finding that names our limit. Two kinds of wall are deliberately not called refusals any more: a site that challenges only identities it can verify by address, and one that also challenges the mobile browser identity, both read as severity-2 unconfirmed rather than as a block, because in both cases the site is applying a policy to everyone rather than refusing a scanner.
How to test any auditor #
Ask for the raw bytes the tool received for your homepage and for a path you know does not exist. If they match, the tool graded your firewall. Most will not be able to answer, which is its own answer. Is Cloudflare blocking AI crawlers on your site covers the most common source of the wall; the glossary defines nonexistent-path control and uniform refusal.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
Can a scanner grade a firewall instead of a website?
Yes. If the origin refuses the scanner's address, every check fails for one reason and the resulting grade describes the refusal, not the site.
How should a scanner handle being refused?
By reporting the refusal as a refusal. A grade computed from blocked requests is a measurement of access, and publishing it as a quality score is an error.
What did this change here?
Refused scans are now identified and are not published as failing grades.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.