CrawlCheck

Findings · 2026-08-18 · By · 0 views

Fake GPTBot and ClaudeBot requests are scanning for .env and .ssh files

Of 874 recorded requests from clients claiming to be a named AI crawler, 69 asked for a file that holds secrets — cloud credentials, Terraform state, .env files. Across 66 distinct paths, none of which any real crawler requests.

A forged user-agent is usually discussed as a statistics problem: someone inflating a crawl budget, or a scraper hiding behind a familiar name. On this site it is also a reconnaissance tool.

What arrived #

Measured 18 August 2026. Of 874 recorded path requests from clients claiming a named AI crawler, 69 were requests for files that hold secrets, spread across at least 66 distinct paths. That is a floor rather than a count: our record keeps a limited number of paths per identity, and several identities had already filled it, so later paths were never recorded. A sample, reproduced exactly as received:

The /@fs/ prefix is the giveaway. It is a path served by a popular JavaScript development server, and a misconfigured deployment will happily read files outside the project root through it. The doubled URL-encoding in the last example is a traversal attempt dressed to survive one round of decoding. No search or AI crawler has ever requested a Terraform state file.

How a claim is adjudicated, so you can run the same check #

Every request that carries a named crawler’s user-agent is one of three things and the record keeps them apart. Verified: the connecting address is inside the ranges the operator publishes for that name. Forged: the operator publishes ranges and the address is not in them. Unverifiable: the operator publishes nothing to check against, so the claim is neither confirmed nor accused. The third class is excluded from every forgery rate on this site rather than folded into either side, because a rate that counts the unknowable as fraud is the same defect as a scanner that scores a check that does not apply.

Name claimedChecked againstOutcome when the address is elsewhere
GPTBot, OAI-SearchBot, ChatGPT-UserOpenAI’s published JSON range files, one per agentforged
ClaudeBot, Claude-User, Claude-SearchBotAnthropic’s published rangesforged
Googlebot, Google-ExtendedGoogle’s published crawler rangesforged
An agent whose operator publishes no rangesnothingunverifiable — never counted as forged

The same adjudication runs on your own log lines through the verify endpoint, and the first-party tally of what reached this site is public, split into verified, forged and unverifiable.

Two rules that keep the rate honest #

First, a per-crawler forgery rate is not named or charted until there are enough adjudicable claims to mean something, because one spoofed request out of one is a 100% rate pretending to be a pattern. Second, a failed verification on a report URL, an API path or a tool page is counted separately: that is somebody checking a report this site produced, not a crawler impersonating an operator on a page, and quoting it in a forgery rate would inflate the figure with our own readers.

The one identity where attribution is unambiguous #

Where an identity has both verified and failed traffic, we cannot say which specific request came from whom — and we will not pretend otherwise.

One case has no such ambiguity. Every request claiming Google-Extended failed verification: 10 claims, none from Google's published ranges. Nine of the ten paths those requests asked for were credential files — every one except /: /.env, /production/.env, /@fs/.env.development, /@fs/var/www/.env, /@fs/root/.config/gcloud/application_default_credentials.json, /.aws/config, /.codex/config.toml, /openai.json and /private-key. Google-Extended is the identity a site owner uses to control whether their content trains Google's models — a name chosen, presumably, because it is one a cautious operator is least likely to block.

What 874 crawler-claiming requests actually asked for #

Measured 18 August 2026Count
Requests claiming a named AI crawler874
Requests for files that hold secrets69
Distinct secret-file pathsat least 66
Google-Extended claims that verified against Google’s published ranges0 of 10

The path count is a floor, not a total: paths beyond what the record keeps per identity were never recorded.

What this changes about blocking #

The usual advice is to decide which AI crawlers you allow and write that into robots.txt. That advice is sound and it is also inert against this: robots.txt is a request, honoured by the crawlers that were never your problem. The client asking for your .env is not reading it.

What does work is the same check that produces every figure on this site: adjudicate the name against the operator’s published address ranges, and treat a failure as untrusted traffic rather than as a crawler. That is a rule you can enforce at the edge in one place, and unlike a robots directive, it does not depend on the other party’s goodwill. How to block by range rather than by name is its own guide; the short version is that a name is a string anyone can type and an address range is not.

Why a 403 here is correct, and a robots line is not #

A 403 to a forged claim costs a real crawler nothing, because a real crawler never fails the check. A robots.txt disallow costs the real crawler its access and costs the forger nothing, because the forger does not read the file. So the two instruments point in opposite directions: the directive governs the well-behaved and the range check governs everyone else. The live dataset carries the running count of both kinds of traffic against this domain, so the sample above can be checked against what has arrived since.

Traffic to this site only, read from the live record on 18 August 2026. Paths are reproduced as received. A request for a credential file is not evidence that any file was exposed — none of the paths above returned one here.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Do fake AI crawlers scan for credential files?

Yes. Of 874 recorded requests claiming a named AI crawler, 69 asked for files that hold secrets such as .env, .ssh material or Terraform state.

How do you know those requests were not the real crawler?

The source addresses fell outside the ranges the named operators publish, so the identity claim fails against the operator's own feed.

Does this change how I should block crawlers?

It argues for verifying by address rather than name. A robots.txt rule is honoured by the real operator and ignored entirely by whoever is wearing its name.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works