CrawlCheck

Library · Research

Research

In one sentence

Findings are short case studies written from real measurements — a robots.txt that answered 200 with a challenge page, a site whose GPTBot traffic came from one address wearing seven crawler names, a homepage that was two percent readable text, a contact email that bounced for one published character — each with the bytes that showed it, and none naming a scanned third-party domain.

Measured across the corpus or the telemetry, not one site.

71findings
34guides
47finding codes explained
2,474scans behind the live figures

2026-09-25 · Research9 min

Google's Maps stack names two US business-data suppliers. You can check one of them in a minute.

The recovered 793-name provider list has no Yelp. It has Data Axle, under two old names, and Acxiom. What that list is, why a server cannot read Data Axle's search, how to check and fix your own record, and what nine businesses showed.

NAP_UNDECLARED_LISTING5.4%

Read the measurements →

2026-09-25 · Research4 min

Google Places API rankings are not what Google Maps shows

Same query, four Denver-area cities, one day. The API and Maps agreed on who competes, disagreed on the order, and the API could not find two businesses Maps ranks.

Read the measurements →

2026-09-23 · Research4 min

Our homepage says 6,557 sites measured. Six people ran a scan this week.

Both numbers are true. One is the corpus the scanner built to have something to benchmark against; the other is human demand. A week of edge logs shows who actually visits a site about crawlers: mostly crawlers.

Read the measurements →

2026-09-20 · Research5 min

Our llms.txt is linked from every page. Crawlers fetch it about as often as a blog post

Sixteen AI and search crawlers made 12,831 requests to crawlcheck.io in two weeks. Eight were for llms.txt, linked from every page. Per URL, that is an ordinary page's rate, not a front door's.

Read the measurements →

2026-09-17 · Research4 min

The web server that gained and lost 43 million sites

Netcraft counts 66 million sites running OpenResty. W3Techs does not list it at all. Both are measuring honestly, and the gap between them is the lesson.

Read the measurements →

2026-09-17 · Research8 min

AI visibility scores measure an API, not what anyone sees in ChatGPT

The answers these tools record are real. The label on top of them — “your visibility in AI search” — describes something that was never measured.

Read the measurements →

2026-09-17 · Research9 min

Which AI crawlers publish IP ranges, and which cannot be verified

251 of 950 requests claiming a named crawler came from operators who publish no address range at all. Those requests are not forged — they are unfalsifiable, which is a different and more permanent problem.

Read the measurements →

2026-09-17 · Research6 min

GPTBot, ClaudeBot, PerplexityBot: what each AI crawler is actually for

The eight most-blocked AI crawlers on the web sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent choices.

Read the measurements →

2026-09-17 · Research5 min

llms.txt adoption rate: how many sites actually publish one

Declaring an entity is common. Declaring the specific facts an assistant is asked for is an order of magnitude rarer.

Read the measurements →

2026-09-17 · Research10 min

What answer engines receive from local business sites: mostly not text

Three findings from a corpus of sites measured byte by byte, and why none of them appear in a validator.

PAGE_IS_MOSTLY_CODE22.6%

Read the measurements →

2026-09-16 · Research5 min

Google's Maps provider list has 793 names. Yelp is not one of them.

A recovered Geostore enumeration names every source a fact about a place can come from. The consumer directories that citation services sell by the dozen are absent. The business's own website is present, with the same standing as a purchased feed.

Read the measurements →

2026-09-12 · Research5 min

1,582 requests arrived under a name Google says it never sends

Google-Extended is a robots.txt control token. Google's own documentation says it has no HTTP user-agent string. Over seven days, 1,582 requests reached our zones carrying it anyway.

Read the measurements →

2026-09-02 · Research6 min

The state of AI visibility, September 2026: what 1,364 scans say a crawler actually receives

Seven findings from the open dataset: the most common defect is a stale cache, a quarter of sites carry a grade-capping finding, opt-outs travel by copy-paste, and a third of crawler traffic is forged.

Read the measurements →

2026-09-01 · Research5 min

33 AI visibility tools scanned: a third publish llms.txt, none an entity map

Thirty-three SEO and AI-visibility vendors, run through the same scanner they would point at you.

Read the measurements →

2026-09-01 · Research6 min

How much AI crawler traffic is forged? Checked against IP ranges

A user-agent string is a claim, not an identity. Every request here that claims a named crawler is checked against the IP ranges that operator publishes — and a measurable share of claims fail the check.

Read the measurements →

2026-09-01 · Research5 min

Site audits go stale: the day you audited it is not the finding

Measuring the same sites every day produces a number a one-off audit structurally cannot: how many were broken at least once. The gap between “broken today” and “ever broken” is the case for a series over a snapshot.

Read the measurements →

2026-09-01 · Research5 min

Which AI crawlers actually visit a small business site

Not a probe: real, unasked-for visits to sites we operate, each checked against the operator's published IP ranges. Which AI crawlers turned up, how often, and which were forged.

Read the measurements →

2026-08-29 · Research3 min

93% of our arrivals sent no referrer

Across five sites and 1,104 recorded arrivals, 1,026 carried no referrer at all. Every claim about where AI traffic comes from is being made on the remaining seven per cent.

Read the measurements →

2026-08-26 · Research4 min

We measured a 225% spread between two Lighthouse runs of the same page

Nothing on the page changed. Blocking time came back 2,643 ms, then 812 ms. First paint moved 69%. If your report shows one lab number as a fact, it is overclaiming.

Read the measurements →

2026-08-26 · Research5 min

AI text watermarks only work if the model's provider opted in

Watermarking is real and the mechanism is sound. But it covers one vendor's models, you cannot run the detector, and paraphrasing removes most of the signal.

Read the measurements →

2026-08-21 · Research8 min

LLMs ignored our schema and still identified us: what actually got read

We publish a four-node JSON-LD graph. Two unrelated tools that turn pages into text for language models both dropped the whole of <head>, and with it every node. Both still named the company correctly. Then we measured our own ten sites: 87% of our structured data sits where neither of them looks.

Read the measurements →

2026-08-15 · Research5 min

Generic AEO scores penalise sites for checks that do not apply to them

A public checker deducted seven points from a tree service for having no OAuth server, no payment endpoint and no MCP card — after its own output had already said the site was not a commerce site.

Read the measurements →

2026-08-14 · Research3 min

The meta tag 1.65 million sites carry and the top 10,000 refuse

MobileOptimized is a Microsoft tag for Internet Explorer Mobile. Adoption across the whole web is climbing. In the top million it is falling, and in the top ten thousand it has never appeared at all.

Read the measurements →

Questions about this page

QAre these real sites?
Yes, every finding is a measurement taken on a live site. Domains are not named unless the author owns them.
QHow is a finding different from a guide?
A finding reports one measured case and what it showed. A guide explains how to run and read the check yourself.

Keep reading