CrawlCheck

Glossary · area 6 of 19

Delivery and rendering

What arrives on the wire, and what only exists after something else runs.

16 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

16terms in this area
4with a live finding rate

payload ratio

The share of a delivered page that is text a machine can quote, against markup, scripts and styling that it cannot. A page can be large and still carry almost nothing an answer engine can use.

render dependence

How much of a page's meaning exists only after JavaScript runs. Crawlers that do not execute scripts see whatever the server sent, which on some sites is a shell.

Measured: CONTENT_NEEDS_JAVASCRIPT 0.7%

render dependence gap

The difference between what is in the HTML and what appears after scripts run. A non-rendering crawler receives the first, which is why a service list that only exists after JavaScript can be invisible to a fetcher that reached the page successfully.

Measured: CONTENT_NEEDS_JAVASCRIPT 0.7%

edge cache pinning

A cached response continuing to be served after the origin has changed. A purge API returning 200 is a statement about the API call, not evidence that anything was evicted.

Measured: CHALLENGE_PINNED_AT_EDGE 0.6% · MACHINE_FILE_CACHE_SPLIT 0.1%

TTFB

Time to first byte: how long a server takes to begin answering. It is the speed measurement that matters most to a crawler, because a crawler that times out records nothing at all.

Core Web Vitals

Google's field measurements of loading, interactivity and layout stability, drawn from real Chrome users. Two of the three describe rendering and interaction, which no AI crawler performs.

CrUX

The Chrome User Experience Report, which publishes real-user performance for origins with enough Chrome traffic. Absence of data is a statement about traffic volume, never about speed.

field data

Performance measured from real visits, as opposed to a laboratory run. It is the more honest number and it does not exist for most small sites.

prompt injection

Text on a page written to instruct a model reading it rather than the person. Relevant to publishers because instructions can arrive on a domain the owner did not write, through user content or a compromised dependency.

Hydration mismatch

A difference between the server-rendered HTML and the DOM the client script produces. A fetcher reading only the first response records the server version, which may not be what any person saw.

Dynamic rendering

Serving pre-rendered HTML to crawlers and the client-rendered app to browsers. Two code paths to keep in sync, and the crawler path is the one nobody looks at, so it goes stale quietly.

DOM depth

How deeply nested the document is. Deep trees slow headless rendering and can hit parse or time limits before the content near the bottom is reached.

Edge transformation

Rewriting HTML, headers or status codes at the CDN before the response reaches the client. It can mask an origin misconfiguration from anyone auditing from inside, and hide an edge fault from anyone auditing the origin.

202 interstitial

A homepage answering HTTP 202 Accepted. 202 is an acknowledgement, not a page: a queue or verification step is answering in place of the site, so nothing measured through it describes the site. The scan is refused rather than graded, and the per-identity table shows which named crawler, if any, was handed the real page. Reported as HOMEPAGE_IS_INTERSTITIAL.

Measured: HOMEPAGE_IS_INTERSTITIAL 0%

rate limiting

A server's answer to a client asking too fast, usually HTTP 429, sometimes a silent slowdown. For measurement the honest consequence is unverifiable: a 429 says the client was throttled and nothing about what it would have received. A scanner that records a throttled fetch as a refusal has invented a finding, and one that retries until it gets through has measured a different moment than the one it reports. The verdict for a throttled path is that it was not measured.

bot management platform

The product category, sometimes called a bot shield, that stands between the origin and every automated client and decides what each one gets. It is the source of challenge page at 200, uniform refusal, and graded through a wall: the shield answers a scanner and most AI crawlers identically, and the site owner may not know it is answering at all, because the owner's own browser passes. What it decides is set by configuration and by the provider's classification, so a scan reports the wall as the wall and does not attribute the response to the site's content.

← Indexing and discovery  ·  Measurement and evidence →

All 668 terms across 19 areas.