CrawlCheck

Glossary · area 2 of 19

Crawlers and access

Who is fetching, whether they are who they claim, and what your rules actually permit.

38 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

38terms in this area
10with a live finding rate

retrieval crawler

A crawler that fetches pages to answer a question being asked right now. Distinct from a training crawler by purpose, not by capability, and the two are often operated by the same company under different user-agents.

user-triggered fetch

A request made because a person asked the assistant about a specific URL. Volume reflects user curiosity rather than crawl policy, so it is a poor proxy for how well a site is indexed.

verified crawler

A request whose claimed crawler identity is confirmed by checking the source address against the operator's own published IP ranges. Reverse DNS and user-agent matching are weaker: only the operator's feed is authoritative.

forged crawler identity

A request whose user-agent names a known crawler while its source address falls outside the ranges that operator publishes. Verified against a published feed it is a fact, not an accusation; without a feed it is unverifiable.

unverifiable

The third outcome, and the one most tools omit. A profile behind a login wall, a 429, or a shell page under 512 bytes can be neither confirmed nor accused, and folding it into either bucket produces a number that overstates or understates by construction.

robots.txt

A file declaring which paths each user-agent may fetch. It controls access, not indexing, and a 403 on robots.txt itself is read by most compliant crawlers as disallow-everything.

Measured: ROBOTS_NOT_200 4.1% · ROBOTS_BLOCKED 0.4% · ROBOTS_IS_HTML 0.2%

user-agent group

A block of robots.txt rules addressed to named agents. Most parsers apply only the most specific matching group, so naming an agent to allow it can silently remove every rule it would otherwise have followed.

AI opt-out

Any published instruction asking that a site's content not be used for model training. It is a request rather than a control, and whether it is honoured varies by operator.

Measured: AI_OPTOUT_SET 6.3%

Content-Signal

A robots.txt directive expressing separate permissions for search indexing, AI input and AI training. It communicates intent to operators who choose to honour it and carries no enforcement of its own.

crawl budget

The effort a crawler will spend on a site before moving on. Wasting it on redirect chains and unreachable URLs is measurable; the budget itself is not published by any operator.

cloaking

Serving materially different content to a crawler than to a browser. A server refusing a datacentre address while serving a residential one is not cloaking - it is IP-based access control, and calling it cloaking is a false accusation.

challenge page at 200

A bot-protection interstitial returned with a success status. Any validator that trusts the status code records the file as present and readable when nothing readable was served.

Measured: CHALLENGE_SERVED_200 0.8% · CHALLENGE_PINNED_AT_EDGE 0.6%

uniform refusal

An origin answering the homepage and a path that cannot exist the same way: same status, same size, or the same challenge or block page. A site does not serve a missing page identically to its home page, so the answers describe a wall rather than the site. CrawlCheck records exactly one finding for such a scan and grades nothing; a score produced through a wall would be a number about someone else's product.

Measured: UNIFORM_REFUSAL 3.3%

nonexistent-path control

A request for a randomised URL that no site would ever publish, made alongside the real requests. It is the control that tells a scanner whether it reached the site at all: a real site answers it differently from its homepage. Tightened on 1 September 2026 to catch walls that answer with non-HTML, a challenge page of any size, or a block page byte-identical across paths.

blocked render resource

A stylesheet or script the homepage references that the site's own robots.txt disallows to every crawler. A rendering crawler obeys the rule and draws the page without it, so layout and any script-inserted content are judged from a page the site never intended to show. Reported as RENDER_RESOURCE_BLOCKED from the * group only.

Measured: RENDER_RESOURCE_BLOCKED 1.6%

off-host redirect

A homepage that answers with a redirect to a hostname outside its own www/apex pair, such as a regional storefront or a checkout host. The scanner does not follow it: whatever a crawler reads on the other host is attributed to that host, and the requested domain reads as a site with no homepage. Reported as HOMEPAGE_REDIRECTS_OFF_HOST.

Measured: HOMEPAGE_REDIRECTS_OFF_HOST 1.8%

homepage refused

A homepage answering 401, 403, 429 or similar to this client while robots.txt is served normally. The machine layer stays scored; every section that needs the homepage is unmeasured rather than scored, and the report says why. It is marked unconfirmed until checked from a second address, because a datacentre refusal may be about the scanner rather than about crawlers. Reported as HOMEPAGE_REFUSED.

Measured: HOMEPAGE_REFUSED 0.8%

FCrDNS

Forward-confirmed reverse DNS: resolve the requesting IP to a hostname, then resolve that hostname back and check it returns the same IP. It is the verification method for crawlers that publish no IP list, and it is the only honest way to test a Googlebot claim.

Crawler trap

An unbounded set of generated URLs — faceted filters, infinite calendars, recursive relative paths — that a crawler can follow forever. It consumes crawl allocation without exposing new content.

Conditional request

A fetch carrying If-None-Match or If-Modified-Since, which lets the server answer 304 Not Modified instead of resending the body. A site that never answers 304 pays full bandwidth for every re-crawl.

ASN blocking

Refusing traffic by hosting network rather than by identity. It blocks retrieval agents running on cloud infrastructure — which is most of them — while leaving residential proxies untouched, so it usually costs more visibility than it prevents scraping.

operator feed

The publisher-side list of source addresses a crawler operator maintains, usually an IP-range JSON, a reverse-DNS suffix, or both. It is the only authoritative basis for calling a hit verified or forged: a request from an address in the feed is the operator's, a request outside it is not, whatever the user-agent string says. Meta, Apple and Amazon publish only a documentation page, so their hits can be checked by reverse DNS but not by range. A stale feed indicts the operator's publishing, not the site being measured.

residential proxy

A request relayed through a consumer ISP address so that it is indistinguishable, at the network layer, from a person at home. It defeats ASN blocking and it defeats feed verification, because the address is in nobody's operator feed and belongs to nobody's data centre. Any forged-crawler rate computed against operator feeds is therefore a floor: the forgeries that used residential addresses were never counted, and no server-side method can count them.

Google-Extended

The robots.txt token that controls whether Google may use a site's content to train and ground its Gemini models. It is not a crawler; Googlebot still fetches the page. Blocking it does not remove a site from Google Search, from Discover, or from the retrieval path behind AI Overviews and AI Mode, which run on the Search index. It is the most commonly misread token on the web: sites block it expecting to leave AI answers and instead leave only the training set.

Google-Agent

A user-triggered fetcher Google added to its documentation on 20 March 2026 for agents that run on Google infrastructure and browse on a person's behalf, Project Mariner among them. It carries a Chrome-like user-agent string with the token inside it and publishes its own range file, user-triggered-agents.json, separate from the other fetcher feeds. Google classes it as a user proxy, so it generally ignores robots.txt. It is the concrete case where the caveat on user-triggered fetch stops being theoretical: a rule addressed to it is a request the fetcher is documented not to read.

GPTBot vs OAI-SearchBot

Two OpenAI crawlers with one operator and different purposes. GPTBot collects for model training; OAI-SearchBot fetches for live search retrieval and is what a ChatGPT search citation depends on; a third identity, ChatGPT-User, fetches when a person asks about a URL. Each has its own range in OpenAI's published feed and its own robots.txt token. Blocking one says nothing about the others, which is why allow and block decisions are per agent, not per company, and why a rule addressed to GPTBot alone leaves search retrieval untouched.

ClaudeBot vs Claude-User

The same split for Anthropic. ClaudeBot crawls autonomously to improve models; Claude-User fetches a page because a person asked Claude about it, and Claude-SearchBot fetches to build and refresh search results. All three verify against Anthropic's published IP ranges. Blocking ClaudeBot does not stop a user-triggered fetch, and counting a Claude-User hit as a crawl inflates a crawl figure with visits that were, functionally, referrals.

PerplexityBot vs Perplexity-User

The same split for Perplexity. PerplexityBot indexes for the answer engine and honours robots.txt; Perplexity-User fetches a page when a person's query requires it and, by Perplexity's own documentation, generally does not honour robots.txt because the request is treated as the user's. Both publish IP ranges. A site that sees Perplexity fetches after blocking PerplexityBot is usually seeing the second identity, not a violation by the first.

CCBot

The Common Crawl crawler. Its open corpus is downloaded and used by many model builders, so a page can enter training sets without ever being fetched by any model operator's own crawler. That is the case that breaks the sentence we blocked GPTBot, so we are out of the training data: the copy that trained the model may have been taken by a nonprofit archive years earlier. CCBot honours robots.txt and publishes its identity, so exclusion is possible; it is only retroactive removal that is not.

Bytespider

ByteDance's training crawler, reported over several years by publishers and by edge providers to fetch without regard to robots.txt exclusions. It is the standing example that a robots.txt rule is a request, honoured by operators who choose to honour it and enforced by nothing. The practical consequence for measurement is ordering: verify who actually fetched before attributing volume to a name, because a crawler that ignores rules is also the easiest name to forge.

agentic traffic

Requests from agentic browsers, Comet, Atlas and their successors, in which a model drives a real browser session on a person's behalf. They arrive with an ordinary browser user-agent, cookies, and human-looking session patterns, from consumer addresses. No user-agent match, operator feed or reverse-DNS check can classify them, so they are invisible to every verification method in this glossary. This is a definable limit rather than a solved problem: a verified count is a count of what identified itself, and this traffic does not.

crawler classification

The model edge providers adopted in 2026, Cloudflare from July with new-domain defaults changing on 15 September: verification confirms only who a crawler is, and access is then decided per behaviour class, Search, Agent, or Training, rather than per operator. A multi-purpose crawler is judged under every class it belongs to, which is how a rule that blocks Training can block a search crawler that also trains. It is the industry name for the identity-versus-permission distinction that verified crawler already implies. The classes are the provider's judgment of a crawler's purpose, and that judgment is not published in a form a site can audit.

user agent

The string a client sends naming itself, in the User-Agent request header. It is a claim, not an identity: anyone can send any string, so treating a crawler's name as proof of the crawler is the mistake behind most forged-crawler traffic. The name is what a site filters on; the source address is what verifies it.

datacenter IP

A source address belonging to a hosting provider rather than a consumer network. Many protection layers refuse or challenge it outright, which is why a scanner, an API client and an answer engine's crawler are often treated identically at the door: by where they come from, not by what they are.

Measured: ORIGIN_REFUSED_SCANNER 0.2% · UNIFORM_REFUSAL 3.3%

vantage point

The network position a measurement is taken from. A datacenter vantage sees what a datacenter client is served, which can differ from what a residential browser or an operator's own crawler receives. A finding that depends on the vantage is stated with it; a second vantage, such as a browser extension, is how the dependence is measured rather than assumed.

Measured: ORIGIN_REFUSED_SCANNER 0.2%

rate limit

A cap on how many requests a client may make in a window, answered with HTTP 429 when exceeded. For robots.txt, engines treat 429 as a server error rather than a missing file. A limit keyed on address rather than identity throttles a verified crawler and a forger alike.

delivery comparison

Fetching the same URL as each named crawler identity and as an unnamed browser control within the same second, then comparing status, redirect target and word count. It shows whether the site treats a crawler's name differently from a person; it cannot show whether the real crawler, arriving from its own address range, is treated the same way.

Measured: CRAWLER_SERVED_LESS 0.5% · CRAWLER_REDIRECTED_AWAY 0.1% · ANSWER_ENGINE_REFUSED 3.5%

honeypot

A field, link or path that no human visitor would use, placed so that automated clients reveal themselves by touching it. On a form it catches bots without a challenge the visitor sees; on a site it can catch crawlers that ignore robots.txt. It identifies behaviour, not intent, and a real crawler that follows every link will trip a link honeypot too.

← Answer engines and retrieval  ·  Machine files →

All 668 terms across 19 areas.