CrawlCheck

Policy

What CrawlCheck collects, and how to opt out

In one sentence

CrawlCheck reads public URLs only (robots.txt, sitemaps, llms.txt, entitymap.json, agents.md and the requested page) and never signs in or submits a form. Machine-file requests identify themselves as CrawlCheck/1.0 and are signed; the homepage is also fetched once under each named crawler user-agent from CrawlCheck’s own address, which is the measurement. It keeps a shareable report for 90 days and one measurement row per domain, publishes aggregate statistics without naming domains, names domains only in the directory and the registry, and honours a CrawlCheck disallow or the line CrawlCheck: none in robots.txt on the next scan: the domain leaves the dataset, the registry and the crawl queue, and a directory vendor stays listed as opted out with no score.

If our crawler reached your server and you want to know what it is, this is the page.

Everything one scan sends to your server
Named crawler user-agentshomepage only, once each12Browser and client controlshomepage, to compare against3As CrawlCheck, signedrobots.txt, llms.txt, two sitemaps, entity map × apex and www10Made-up control pathstwo per host, to see if the server answers everything429 GET requests, read-only, from one address, in about a second — then nothing until someone asks again.

Who is fetching

Requests carry the user-agent Mozilla/5.0 (compatible; CrawlCheck/1.0; +https://crawlcheck.io/bot). Every request that fetches your machine files and control paths identifies itself this way. The measurement itself is the one exception: to see what each crawler is served, a scan also asks for your homepage once under each of 12 named crawler user-agents (GPTBot, ClaudeBot, Googlebot and the rest), from our own address. That address is not in any operator’s published range, so a site that verifies crawlers properly will refuse those requests, and the report says so. We never use those strings to get past a block, and we never send them anywhere but the homepage.

Signed, so the name can be checked. Every request we send under that user-agent carries an Ed25519 signature (HTTP Message Signatures, the Web Bot Auth profile) over your hostname, and a Signature-Agent header naming this site. The public keys are at /.well-known/http-message-signatures-directory, and that response is itself signed by every key it lists. What the bot is and how it behaves is at /.well-known/signature-agent-card; key changes are recorded at /.well-known/security-key-rotation.json. We own no IP range — requests leave through Cloudflare’s shared network — so /.well-known/ips.json lists none: an address is never proof a request was ours, the signature is. Check any signed request, ours or anyone’s, with the signature checker.

One exception, stated plainly because it matters: when a report measures what each agent receives, we deliberately send those user-agent strings to your public URL and record what comes back. Those requests come from our own address, which is not in the range those operators publish, so a site that verifies crawlers properly should refuse them. That is a correct result and the report says so. Those probes are sent unsigned: a signature would let a site that verifies Web Bot Auth recognise the probe as us and serve it something the real crawler never receives, and the measurement would stop measuring.

What we read

What we keep

The dataset page publishes aggregates only. It has never named a scanned domain and is not going to.

Google accounts you connect

Two features read data from Google on your behalf, and only after you sign in and grant them. Neither runs unless you do.

What we keep: a refresh token stored against your licence, the profile fields above, and the daily and monthly series. The token is never shown on any page or API. Disconnect on the licence page deletes the token and every series read with it. Where Google reports a keyword count only as “below a threshold”, we store the threshold and say so; we do not turn it into a number.

If page text is ever sent to a third party for analysis — the content-entity check is designed to use Google’s Natural Language API — that check will be named here before it is switched on, with what is sent and what comes back. As of this page’s date it is not on, and nothing a scan reads leaves this service.

Requests to this site

Separate from scanning, and stated because a page listing what we keep should list all of it: this site keeps a short operational record of requests made to it, the way any web server keeps an access log. It holds the last 200 requests and a per-address count that expires after eight days, and it is readable only by us.

Performance timings from this site’s own pages

These pages load a small script that measures how fast this site was for you — how long the main content took to appear, how quickly it answered your first interaction, and how much the layout moved — and sends one message when you leave the page. We run it on ourselves because we ask other people to run it, and it would be poor form to publish numbers we had not collected the same way.

How to opt out

Either of these stops future measurement, deletes the domain’s stored measurements, and removes it from the crawl queue, the registry and every public count from then on. In the directory the vendor’s name stays, marked opted out with the effective date and no score, so an absence never looks like an editorial exclusion. No email, no form, no account — you declare it in a file you already control and we honour it on the next scan: that scan deletes the stored record, its score history and the corpus records built from them. A total published before then is not recomputed — it is a number, not a list, and it names no one.

User-agent: CrawlCheck
Disallow: /

or a single line anywhere in robots.txt:

CrawlCheck: none

A page you request yourself can still be scanned — you may be the one asking — but nothing about it enters the dataset or any count. The report you asked for is kept for its link, for 90 days, like every report.

The directory

The directory publishes a score for named companies, which is different from every other measurement here, and it runs under stated rules.

What we do not do

Questions about this page

QHow do I opt my domain out of CrawlCheck?
Add User-agent: CrawlCheck with Disallow: / to robots.txt, or the single line CrawlCheck: none anywhere in it. No email or form. The exclusion is honoured on the next scan: that scan deletes the stored record and its history, and the domain leaves the dataset, the registry and the crawl queue, and adds nothing to any public count after that. A vendor in the directory stays listed by name, marked opted out with no score, so an absence never looks like an editorial exclusion.
QDoes CrawlCheck pretend to be Googlebot or GPTBot?
Never to gain access. When a report measures what each agent receives, it deliberately sends those user-agent strings from its own address and records what comes back; a site that verifies crawler IPs should refuse them, and the report says that is a correct result.
QWhat is stored about my site?
A report at /r/<id> for 90 days, one row per domain holding the measurements and scores so a re-scan can show what changed, and aggregate counters behind the dataset page. Nothing behind authentication is ever fetched.
QDoes the directory accept payment for placement or removal?
No. No vendor is listed for payment, no listing is removed for payment, and no score is edited on request. A dispute is a re-scan published as measured.

Keep reading