CrawlCheck

Findings · 2026-08-19 · By · 0 views

Is your host blocking ClaudeBot? Group the 403s by IP first

A 57% block rate in our own access logs looked like a hosting provider shutting out AI crawlers. Grouped by source address, almost all of it was one machine wearing seven different crawler names.

We were one edit from publishing that a managed host was blocking Claude across five applications. Grouped by user-agent, the refusals looked like seven crawler operators being turned away. Grouped by source address, they were one machine sending seven different crawler names — which means at least six were forged, and the refusals were a defence working correctly. A smaller set was real, concentrated on the machine files, and was fixed by putting those endpoints behind an edge cache. Measured across nine domains on 19 August, eighteen of eighteen requests returned 200.

In July, AI crawlers were being refused across five WordPress applications we run on one managed host. This is not a guess from a sample: it is in the access logs. On one site ClaudeBot received 553 refusals out of 642 requests that month. On another, 136 refusals against 29 successes — 82% of everything Anthropic's crawler asked for.

On 30 July that stopped. The August logs for the same two sites read 23 refusals out of 171, and 117 successes against 15 refusals. The change lands on one date across applications that share nothing but a host, which is the signature of a platform default rather than anything in our configuration. The host is Cloudways. We never saw the rule itself and cannot show you its text — what we can show is five applications changing behaviour on the same day, in logs we did not write.

The part where we nearly got it wrong #

The refusals did not go to zero. On one site a single day in August showed 303 crawler requests and 172 refusals — 57%. Written down that way it reads like the block never really lifted, and that sentence was close to being published.

Then we grouped the refusals by source address. All 172 came from one IP. That address presented seven different crawler identities in a single day: Amazonbot, ChatGPT-User, ClaudeBot, OAI-SearchBot, GPTBot, Google-Extended and PerplexityBot. It sits in no range any of those operators publishes. On a second site the same shape appeared — 151 of roughly 229 refusals from one Google Cloud address, again wearing seven names.

One host cannot be seven companies. What the 57% actually measured was a scraper being correctly refused, and counting it as blocked crawler traffic did not just inflate the number — it inverted the conclusion. The genuine crawlers were fine in the same logs: Googlebot refused zero times, Bingbot zero.

The refusals that did matter #

A smaller set was real, and it was concentrated in the worst possible place. Of just over a thousand AI-crawler refusals on one site, 316 were the sitemap and 24 were /robots.txt. A crawler that cannot read your sitemap cannot enumerate your site and falls back to following links. Worse, most compliant crawlers treat a refusal on robots.txt as disallow everything — not as no rules found. A few dozen refusals on one small file explains more lost coverage than any amount of theorising about crawl demand.

Those were fixed by putting the machine files behind an edge cache, so the high-frequency endpoints stop touching the origin at all. The remaining refusals are vulnerability probes — requests for /.ssh/, /.claude.json, /.env — where refusing is the correct behaviour and unblocking would be a mistake.

Where it stands today #

Measured on 19 August 2026, across nine domains, requesting the homepage and /robots.txt as ClaudeBot and as a browser: eighteen of eighteen returned HTTP 200, with no challenge page in any response body. Nothing on this network refuses Claude today.

We had a headline available in July and it would have been true then. Published this week it would have been false, about a named company, on stale data. The check that makes it a story rather than an accusation is the boring one: re-run the measurement before you publish it.

Why one machine looked like seven companies #

The chart that nearly became a headline was refusals grouped by user-agent. That grouping is the error, and it is an easy one to make because every log tool offers it and the resulting bar chart looks like a finding.

A user-agent is a field in the request. Grouping by it counts names sent, not senders, and those are the same number only if every sender uses one name. Regroup the identical rows by source address and the seven bars collapse into one, because they were one client cycling through identities.

The general form is worth keeping: a rate is only a rate if its denominator counts independent events. Seven refusals to one address inside a minute is one event observed seven times. Presented as seven, it supports a sentence about seven operators that never happened.

This is the same arithmetic that makes small-sample forgery rates useless: without a minimum sample, the loudest single misbehaving client sets the headline.

The layer underneath: those names were claims, not identities #

Once the rows collapse to one address, a second fact becomes unavoidable. A single machine cannot be seven crawler operators. At most one of those names was true, so at least six were forged — and both large operators involved publish address ranges, which means the claims were checkable and failed.

That inverts the story completely. The chart was not evidence of a host blocking AI crawlers. It was evidence of something impersonating AI crawlers and being correctly refused, which is the defence doing exactly what it exists to do. Had we published, we would have reported a working security control as an outage, on a host we were about to name.

It also explains why the refusal count did not fall to zero when the real problem was fixed on 30 July, and why we should have treated that residue as a question rather than as a remainder. A number that stops falling when the cause is removed is telling you there was more than one cause.

What the genuine refusals had in common #

The smaller, real set was not distributed across the sites. It was concentrated almost entirely on the machine files — robots.txt, llms.txt, the sitemap — and the reason is mechanical rather than sinister.

Those endpoints are the highest-frequency and lowest-cost requests any crawler makes. A well-behaved operator re-reads robots.txt often, by design. So on a per-address rate limiter they are the first requests to cross a threshold, and they are the requests whose refusal does the most damage, because a crawler that cannot read your policy file either stops or proceeds without it.

The fix was to put the machine files behind an edge cache so those requests stop touching the origin at all. Worth saying why that is the right fix rather than allowlisting the crawlers: an allowlist keyed on a user-agent is unenforceable, for the reason the section above demonstrates in our own data. Caching the file removes the origin from the path regardless of who is asking, and it is enforceable because it does not depend on the request being honest about itself.

The publishing rule this produced #

We had a headline available in July and it would have been true in July. It stopped being true on 30 July, and the version we nearly published in August was a story about a host that had already fixed the problem, inflated by a forgery artefact we had not checked for.

The rule now: any finding that accuses a named third party gets re-verified by hand, on the host the scanner actually read, immediately before publication. Not on the domain you typed — on the host the request resolved to, which is a different thing whenever an apex redirects to www, and has caught a near-miss here for exactly that reason.

The cost is that some true stories go out late, and one or two die because they stopped being true while we were checking. That is the correct trade for a site whose proposition is that claims should be checkable. A measurement business that publishes a wrong accusation has not made a small error; it has disproved its own product.

What to do with your own logs #

If you are about to conclude that something is blocking AI crawlers on your site, do these three things first, in this order:

Group refusals by source IP. One address responsible for most of them means a scraper, not a policy. Count identities per address. Any single IP claiming more than one named crawler is forging at least one of them. Check the paths. Refusals on robots.txt and your sitemap cost you real coverage; refusals on credential paths are your server doing its job.

A blocked-rate with none of that behind it is not a measurement. It is a number that happens to be large, and large numbers travel further than correct ones.

Live, as you read this: the corpus now holds 189,169 domains across 3,137 scans. The figures in this piece were measured on the date above; this line is not.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Is my host blocking ClaudeBot?

Group your 403s by source IP before answering. A high block rate can be one machine wearing several crawler names rather than a hosting policy.

What did the grouping reveal here?

Almost all of the apparent block came from a single address presenting as seven different crawlers. The host was not the cause.

Which refusals did matter?

The ones on robots.txt and the sitemap. A 403 on robots.txt reads to a parser as disallow everything.

How do I tell if my host is blocking AI crawlers?

Request your homepage and your robots.txt from outside your network, once as the crawler identity and once as an ordinary browser, and compare the status and the body of the final response. A refusal that appears for the crawler and not the browser is a block. Anything you conclude from a log chart alone, without that paired request, is an inference about a field the sender controls.

Why did grouping refusals by user-agent give the wrong answer?

Because a user-agent counts names sent rather than senders. One client cycling through seven crawler names produces seven bars and one actual source. Regrouping the same rows by source address collapsed them into a single machine, which also meant at least six of the seven names were forged.

What should I check before concluding a crawler is being blocked?

Three things, in order. Group the refusals by source address rather than by user-agent. Check whether those addresses fall inside the ranges the named operator publishes, where the operator publishes any. Then reproduce the refusal yourself with a direct request from outside your network. A blocked rate with none of that behind it is not a measurement.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works