Guides · 2026-09-30 · By VSNARY | Emmanuel Orta · 0 views
Crawler delivery: refused, redirected away, or served less
How to tell whether a search or answer crawler receives the page a person receives: fetch as each identity and as a control, compare, and report only what reproduces.
CrawlCheck fetches the homepage as each named search and answer crawler and as an unnamed browser control, then compares status, redirect target and word count. A crawler refused with a non-200 is ANSWER_ENGINE_REFUSED; a crawler redirected to a target the control did not get, reproduced, is CRAWLER_REDIRECTED_AWAY; a crawler served under 70 percent of the control's words, reproduced, is CRAWLER_SERVED_LESS. A 202 homepage is HOMEPAGE_IS_INTERSTITIAL and the scan is refused; an identical wall on every path is UNIFORM_REFUSAL and the record carries only that.
The question a scanner should answer is not whether a page loads. It is whether the page a search or answer crawler receives is the page a browser receives. Sites split those two constantly and almost never on purpose: a bot-management rule, a caching layer keyed on user-agent, a redirect written for one client, a template that serves a lighter page to anything without JavaScript. The method for finding a split is simple to state and awkward to run by hand: fetch the same URL as each named crawler and as an unnamed control, in the same second, and compare. This guide is that comparison, the four outcomes it can produce, and the rule that keeps it from accusing a site of a split that was really a site-wide decision.
The identities #
CrawlCheck fetches the homepage as 15 identities: the named user-agents of the crawlers that feed search and answer engines, and a small set of unnamed browser controls. Each fetch records status, bytes, word count, redirect target and whether a known blocker signature appeared in the body. The comparison is always against a control, never against another crawler, because the control is what a person sees. Where two identities differ, the scanner re-runs the pair before reporting, so that a transient answer is not published as a policy. The identity list itself is on the policy page. Two details of the fetch matter for reading the results. Every identity asks for the same URL with the same Accept header, so the only variable is the name and the address is held constant. And the fetches are made within seconds of each other, so a page that changes by the hour cannot produce a false split; a difference between identities is a difference in treatment, not in time.
Outcome one: refused #
A named identity gets a non-200 that the control did not get. This is the loud case and the easy one: the site is refusing the crawler by name, or by the address range the crawler comes from. Reported as ANSWER_ENGINE_REFUSED (3.5% of scans) with the identity and the status. The Cloudflare guide covers the most common source of it.
Outcome two: sent somewhere else #
A named identity is redirected to a target the control was not redirected to, and the redirect reproduces on a second run. Reported as CRAWLER_REDIRECTED_AWAY (0.1% of scans), with each identity and where it went. The qualifier matters: a redirect that everyone follows, control included, is a site decision and is never reported here. Only a redirect that splits crawler from person counts.
Outcome three: served less #
A named identity gets 200 and a page, and the page is thinner than the control's. This is the quiet case, and it is the copy an engine actually quotes from. The rule has thresholds so that it does not fire on noise: the control must have returned more than 120 words, the crawler's copy must be under 70 percent of the control's word count, no blocker signature may be present, and the thinness must reproduce on a second fetch. Non-200 answers are excluded because they are already the refused case and must not double-count. Reported as CRAWLER_SERVED_LESS (0.5% of scans), naming each identity's word count against the control's.
Outcome four: nobody got a page #
Two cases stop the comparison entirely. If the homepage answers HTTP 202, the scanner refuses to grade: 202 is an acknowledgement, not a page, and a verification or queue step is answering in place of the site. Reported as HOMEPAGE_IS_INTERSTITIAL (0% of scans), and the report says which named identities, if any, received a real page, because that is the useful fact. If the homepage and a path that cannot exist answer with the same status, near-identical sizes and either a challenge signature or byte-identical bodies, every path is a wall and nothing measured through it describes the site. Reported as UNIFORM_REFUSAL (3.3% of scans), and a refused record carries exactly that one finding; anything raised before the wall was detected is discarded, because it was measured against the wall.
| Named identity | Control | Reproduced | Report |
|---|---|---|---|
| 403 | 200 | — | ANSWER_ENGINE_REFUSED |
| 301 to /bot-check | 200 on / | yes | CRAWLER_REDIRECTED_AWAY |
| 301 to /home | 301 to /home | — | nothing: a site decision |
| 200, 210 words | 200, 1,400 words | yes | CRAWLER_SERVED_LESS |
| 200, 1,100 words | 200, 1,400 words | — | nothing: over 70% |
| 202, 3,962 bytes | 202 | — | HOMEPAGE_IS_INTERSTITIAL, refused to grade |
| same tiny answer on / and /impossible | same | — | UNIFORM_REFUSAL, one finding only |
Why the control is a browser and not another crawler #
Comparing GPTBot to Googlebot tells you the two crawlers are treated differently and nothing about which one is right. Comparing each to an unnamed browser control tells you which one is getting something other than the page a person gets, and that is the only comparison with a consequence: an engine quoting from the thinner copy is quoting from a page no visitor has seen. The control is also why the method can be run from one vantage point; the thing being tested is the site's treatment of the identity, and the scanner's own address is held constant across the pair.
What it cannot see #
The comparison reads what the server sends. It does not execute JavaScript, so a page that renders its content client-side looks thin to every identity equally and is reported under a different rule, CONTENT_NEEDS_JAVASCRIPT, covered in the rendering guide. It also cannot verify that the site's treatment of a named user-agent is the treatment the real crawler receives from its own address range; a site that filters by IP rather than by name will pass this test and still refuse the real crawler. The post on what a crawler actually receives covers the reverse problem: forged user-agents in your own logs. The share of scans carrying each finding above is on the dataset page.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What does CrawlCheck compare when it checks crawler delivery?
The homepage fetched as each named search and answer crawler user-agent against the same page fetched as an unnamed browser control, in the same run: status, redirect target, word count and blocker signatures. Differences are re-run before they are reported.
Why is a redirect only a finding when the control was not redirected?
A redirect that every client follows is a site decision, such as a move to a new homepage. Only a redirect that sends a named crawler somewhere the control did not go is a split between crawler and person.
What thresholds does CRAWLER_SERVED_LESS use?
The control must return over 120 words, the crawler's copy must be under 70 percent of the control's word count with no blocker signature, and the thinness must reproduce on a second fetch. Non-200 answers are excluded because they are the refused case.
Why does a 202 homepage stop the scan?
202 Accepted is an acknowledgement, not a page; a verification or queue step is answering in place of the site. Grading it would grade the interstitial, so the scanner refuses and reports which identities received a real page instead.
What is a uniform refusal?
The homepage and a path that cannot exist answer with the same status, near-identical size and either a challenge signature or byte-identical bodies. Every path is a wall, so the record carries that one finding and nothing measured against the wall is kept.
Does this test prove the real crawler gets the page?
No. It tests how the site treats the crawler's name from one address. A site that filters by address range can pass this and still refuse the real crawler; log verification covers that side.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.