CrawlCheck

Policy

What CrawlCheck collects, and how to opt out

In one sentence

CrawlCheck identifies itself as CrawlCheck/1.0 on every request, reads public URLs only (robots.txt, sitemaps, llms.txt, entitymap.json, agents.md and the requested page), never signs in or submits a form, keeps a shareable report for 90 days and one measurement row per domain, publishes aggregate counts without ever naming a scanned domain, and excludes any domain from the dataset, the crawl queue and the directory on the next scan when its robots.txt carries a CrawlCheck disallow or the line CrawlCheck: none.

If our crawler reached your server and you want to know what it is, this is the page.

Who is fetching

Requests carry the user-agent Mozilla/5.0 (compatible; CrawlCheck/1.0; +https://crawlcheck.io/bot). We identify ourselves on every request and never send a user-agent belonging to Googlebot, GPTBot, ClaudeBot or any other operator in order to gain access.

One exception, stated plainly because it matters: when a report measures what each agent receives, we deliberately send those user-agent strings to your public URL and record what comes back. Those requests come from our own address, which is not in the range those operators publish, so a site that verifies crawlers properly should refuse them. That is a correct result and the report says so.

What we read

What we keep

The dataset page publishes aggregates only. It has never named a scanned domain and is not going to.

Google accounts you connect

Two features read data from Google on your behalf, and only after you sign in and grant them. Neither runs unless you do.

What we keep: a refresh token stored against your licence, the profile fields above, and the daily and monthly series. The token is never shown on any page or API. Disconnect on the licence page deletes the token and every series read with it. Where Google reports a keyword count only as “below a threshold”, we store the threshold and say so; we do not turn it into a number.

If page text is ever sent to a third party for analysis — the content-entity check is designed to use Google’s Natural Language API — that check will be named here before it is switched on, with what is sent and what comes back. As of this page’s date it is not on, and nothing a scan reads leaves this service.

Requests to this site

Separate from scanning, and stated because a page listing what we keep should list all of it: this site keeps a short operational record of requests made to it, the way any web server keeps an access log. It holds the last 200 requests and a per-address count that expires after eight days, and it is readable only by us.

Performance timings from this site’s own pages

These pages load a small script that measures how fast this site was for you — how long the main content took to appear, how quickly it answered your first interaction, and how much the layout moved — and sends one message when you leave the page. We run it on ourselves because we ask other people to run it, and it would be poor form to publish numbers we had not collected the same way.

How to opt out

Either of these excludes your domain from the dataset, the crawl queue, the registry and the directory, and from every public count from then on. No email, no form, no account — you declare it in a file you already control and we honour it on the next scan: that scan deletes the stored record, its score history and the corpus records built from them. A total published before then is not recomputed — it is a number, not a list, and it names no one.

User-agent: CrawlCheck
Disallow: /

or a single line anywhere in robots.txt:

CrawlCheck: none

A page you request yourself can still be scanned — you may be the one asking — but nothing about it enters the dataset or any count. The report you asked for is kept for its link, for 90 days, like every report.

The directory

The directory publishes a score for named companies, which is different from every other measurement here, and it runs under stated rules.

What we do not do

Questions about this page

QHow do I opt my domain out of CrawlCheck?
Add User-agent: CrawlCheck with Disallow: / to robots.txt, or the single line CrawlCheck: none anywhere in it. No email or form. The exclusion is honoured on the next scan: that scan deletes the stored record and its history, and the domain leaves the dataset, the registry, the directory and the crawl queue, and adds nothing to any public count after that.
QDoes CrawlCheck pretend to be Googlebot or GPTBot?
Never to gain access. When a report measures what each agent receives, it deliberately sends those user-agent strings from its own address and records what comes back; a site that verifies crawler IPs should refuse them, and the report says that is a correct result.
QWhat is stored about my site?
A report at /r/<id> for 90 days, one row per domain holding the measurements and scores so a re-scan can show what changed, and aggregate counters behind the dataset page. Nothing behind authentication is ever fetched.
QDoes the directory accept payment for placement or removal?
No. No vendor is listed for payment, no listing is removed for payment, and no score is edited on request. A dispute is a re-scan published as measured.

Keep reading