CrawlCheck

Second vantage

The half a server cannot measure

In one sentence

The CrawlCheck browser extension measures the half a server cannot: what each named AI crawler is served from your own residential address in a real browser, what exists only after your JavaScript runs, field vitals (LCP, CLS, TTFB) on your connection, and the consent or session wall an anonymous client is shown. Combined with a server scan it settles whether a refusal is a name block or correct IP verification. It sends nothing back.

A scan from here tells you what a crawler is served from a data centre. It cannot tell you what one is served from your address, in a real browser, after your JavaScript has run. The extension measures that half.

What it measures that a scan cannot

Identity from a residential address

The same page requested as each named AI crawler and as your own browser, from your network rather than a cloud range. When an edge treats the two differently, the difference is the finding.

Render delta

What exists only after JavaScript runs — headings, prices, schema and links that are on your screen and absent from the delivered HTML.

Field vitals

LCP, CLS, TTFB and long tasks from this page load on this connection. No API key, and it works on sites too small to have public field data.

Consent and session delta

The wall an anonymous client is served and you never see, because your session is already past it.

The join is the point

Neither vantage can settle a refusal on its own. Server refused and browser refused means the site is blocking that crawler by name. Server refused and browser served means the site is verifying crawler IP addresses — which is correct behaviour, not cloaking, and calling it a defect would be wrong.

Browser served the pageBrowser refusedServer served
Reachable everywhereThe crawler gets the same page from a datacentre and from your address. Nothing to fix at this stage.
Your network is the wallRare: the datacentre got in and you did not. A local filter, VPN or corporate proxy is in the way, not the site.
Server refused
IP verification, working as intendedRefused from a cloud range, served from a residential one: the site checks crawler IP addresses. Correct behaviour, not cloaking, and the report will not call it a defect.
Blocked by nameRefused from both addresses: the site is turning that crawler away by its user-agent. Whether that is policy or accident is yours to decide; the report names it.

That distinction needs two vantage points at once. It is the reason this is one product and not two.

What it sends back

Nothing. There is no POST, no sendBeacon and no XMLHttpRequest anywhere in the extension; every network call it makes is a GET. Measurements stay in the browser's own storage on your device. Every probe is sent with credentials: "omit", so your cookies and sessions never reach a site being measured.

One request reaches us and only if you hold a licence: a check of the key itself, at most once a day. It carries the key and nothing else.

How a measurement session runs

Open the site you want to measure, open the extension, pick the crawler identities to test, and run. Each probe is an ordinary fetch from your browser wearing that crawler’s user-agent string — sent with credentials: "omit" so your session never leaks into the measurement. The result is a table: identity, status, bytes, and whether the body differs from what your own browser was served. A difference is not automatically a defect; the next section is the part that matters.

Reading the result honestly

One probe from one connection is one reading. It settles what your network is served; it does not establish what every visitor everywhere sees. The server scan and the extension disagree sometimes — when they do, the disagreement is the finding, not an error in either tool.

Getting it

The Chrome Web Store listing is not up yet. Until it is, ask us for the build at hello@crawlcheck.io and we will send it with loading instructions.

Requires Chrome 121 or later. Free, and it works without a licence — a licence adds the scored report and twice-daily monitoring here. What the extension does with your data.

Questions about this page

QWhat does the extension measure that the server scan cannot?
Identity from a residential address (the same page as each AI crawler and as your browser, from your network), the render delta after JavaScript, field vitals without an API key, and the consent or session wall an anonymous visitor sees that your logged-in session hides.
QDoes the extension send my browsing data anywhere?
No. Every probe uses your own browser with credentials omitted; the only call home is a licence check at most once a day. Nothing it measures is transmitted or stored. The privacy policy states this in full.
QHow do the two vantage points settle a crawler refusal?
Server refused and browser refused means the site blocks that crawler by name. Server refused and browser served means the site verifies crawler IP addresses, which is correct behaviour and not cloaking. Only both readings together can tell those apart.
QWhich browsers are supported?
Chrome and Chromium-based browsers. The extension is version 0.8.0; the documentation on this page describes what a measurement session runs and how to read the result.

Keep reading