CrawlCheck

Glossary · area 12 of 19

HTTP, edge and rendering

The delivery layer between an origin and a client: proxies, caches, redirects, status codes and the headers that decide what a crawler actually received.

60 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

60terms in this area
10with a live finding rate

Origin server

The authoritative application or host responsible for a resource before intermediary handling. External clients may never reach it directly when a CDN, proxy, or firewall answers first.

Reverse proxy

A server that receives requests on behalf of upstream applications. It can cache, route, transform, authenticate, or block responses, making observed behavior different from origin behavior.

CDN

A distributed network that serves content near clients and absorbs traffic from the origin. It improves delivery but can preserve stale, transformed, or client-specific variants.

Edge node

A CDN or platform location that processes a request near the requester. Two nodes can hold different cache states or policies at the same moment.

WAF

A web application firewall that evaluates HTTP requests against security rules. It can stop attacks and legitimate crawlers, and a block page may still carry HTTP 200.

Bot score

A provider-specific estimate that traffic is automated. It is a classification signal, not proof of identity, intent, safety, or robots compliance.

TLS handshake

The negotiation establishing an encrypted HTTPS connection and server authentication. Success proves transport setup, not that the application will serve the requested resource.

DNS resolution

Converting a hostname into network addresses. Successful resolution identifies where to connect; it does not prove the resulting server is healthy or authoritative for content.

DNS timeout

Failure to obtain a DNS answer within the client’s limit. It leaves the resource unmeasured rather than proving the host does not exist.

Connection timeout

Failure to establish a transport connection within a configured period. It can reflect routing, firewall, congestion, or server conditions, not page quality.

Read timeout

A client ending a request because response data did not arrive in time. The result is incomplete measurement, not necessarily an origin error.

Redirect chain

Two or more redirects before a final response. Each hop adds failure and latency opportunities, while clients impose different maximums.

Redirect loop

Redirects that eventually revisit a previous URL or repeat indefinitely. A crawler stops after its own threshold, so no final page is retrieved.

301 redirect

A permanent redirect whose target is intended to replace the requested resource. Clients and indexes may still take time to consolidate signals.

302 redirect

A temporary redirect that preserves the possibility that the original URL remains authoritative. Long-term use can still be interpreted differently by search systems.

307 redirect

A temporary redirect that requires preserving the request method and content. It controls HTTP replay semantics, not indexing treatment by every engine.

308 redirect

A permanent redirect that preserves the request method and content. It does not guarantee immediate cache, canonical, or index consolidation.

304 response

“Not Modified,” telling a conditional requester to reuse its stored representation. It carries no new body and is meaningful only with the client’s cached copy.

404 response

A response indicating no current representation was found or disclosed for the target. It does not state whether the absence is permanent.

410 response

“Gone,” indicating the resource is intentionally unavailable and likely permanently removed. Crawlers may process it faster than a 404, but timing is not guaranteed.

429 response

“Too Many Requests,” indicating rate limiting. It says the client was throttled, not whether the resource would otherwise be allowed or readable.

500 response

A generic unexpected server failure. It identifies server-side failure but not the responsible component or whether retrying will succeed.

502 response

“Bad Gateway,” returned when an intermediary receives an invalid response from an upstream server. It can originate at a CDN or proxy rather than the website application.

503 response

“Service Unavailable,” indicating temporary inability to handle the request. A `Retry-After` value can guide clients but cannot force them to return.

504 response

“Gateway Timeout,” indicating an intermediary waited too long for upstream. It measures that request path, not universal origin availability.

HEAD request

A request for the headers that a corresponding GET would return, without transferring response content. It cannot verify body text, markup, challenge pages, or payload parity.

GET request

A request to retrieve a current representation of a target resource. HTTP success does not guarantee that the representation is the intended human page.

Content type

The declared media type describing how a response body should be interpreted. A correct header can accompany invalid content, and an incorrect header can hide a valid payload from parsers.

MIME sniffing

Inferring content format from bytes when metadata is absent or distrusted. Different clients can infer differently, producing parser and security inconsistencies.

Content encoding

A representation transformation such as gzip or Brotli declared for transfer. A client must decode it before assessing the underlying body.

Compression

Reducing transferred bytes through an encoding negotiated between client and server. It changes bandwidth and latency, not the semantic amount of content.

Brotli

A compression format commonly used for web responses. Support and compression level affect transfer size and processing cost, not crawlability by themselves.

ETag

An opaque validator identifying a selected representation version. Matching tags support cache validation; they do not prove semantic equivalence across URLs or encodings.

Last-Modified

A server-declared timestamp for when a representation changed. Its granularity and accuracy depend on the application generating it.

If-None-Match

A conditional request header asking whether the current ETag differs from one already stored. A matching validator permits a 304 instead of a new body.

Cache hit

A request served from stored intermediary or client state. It improves speed but may bypass the origin and expose stale or incorrectly keyed content.

Cache miss

A request that cannot use an existing cache entry and must obtain or generate a representation. It does not necessarily mean the resource was never cached before.

Cache key

The values a cache uses to decide which stored response matches a request. Missing a relevant dimension can serve one client’s variant to another.

Vary header

Response metadata naming request headers that affected representation selection. It helps caches separate variants but does not account for unstated edge rules.

Stale response

A cached response whose age exceeds its freshness lifetime. It may still be served under defined directives or failure policies.

Measured: STALE_CACHE_SERVED 16.7%

Stale-while-revalidate

A cache directive permitting stale reuse while an asynchronous refresh occurs. It reduces latency while creating a window where clients receive old content.

Edge redirect

A redirect generated by CDN or proxy configuration before the origin handles the request. Origin code inspection will not reveal it.

Origin bypass

Reaching an origin through a route that avoids its normal edge layer. It isolates origin behavior but does not reproduce what public crawlers receive.

Response parity

Material agreement between responses served to different identities, locations, or clients. Equal status codes alone are insufficient; bodies, headers, redirects, and readable text must also be compared.

Header parity

Agreement in response metadata across compared clients. It can reveal policy differences while missing body-level cloaking, challenges, or render dependence.

hreflang

A link annotation that tells a search engine which language or regional version of a page a URL is, and where the other versions live. It groups alternates; it does not redirect anyone, does not translate anything, and a set that is not reciprocal across every version is ignored by the engine that reads it.

canonical URL

The one address a site declares as the authoritative copy of a page, in a link rel=canonical tag or an HTTP header. It is a hint that engines usually follow, not a redirect: the duplicate stays reachable, and a canonical that points at a page returning 404 or a different host is discarded rather than obeyed.

Measured: HOST_DUPLICATE_200 3.3% · DOMAIN_IS_ALIAS 0.8%

HSTS

Strict-Transport-Security, a response header that tells a browser which has already reached a site over HTTPS to refuse plain HTTP for max-age seconds. It protects every visit after the first; the first request is only covered by the browser preload list, and the header is ignored if it arrives over HTTP.

Measured: HSTS_MISSING 7.6% · HSTS_PRELOAD_INELIGIBLE 1.9% · HSTS_PRELOAD_REMOVED 0.3%

HSTS preload list

A list compiled into browsers of hosts that must only be reached over HTTPS, so that even a first request is upgraded. Entry requires a one-year max-age, includeSubDomains and the preload token, and removal takes months to propagate through shipped browsers, which is why sending the token without qualifying is worse than not sending it.

Measured: HSTS_PRELOAD_INELIGIBLE 1.9% · HSTS_PRELOAD_REMOVED 0.3%

Content-Security-Policy

A response header that tells a browser which sources a page may load scripts, styles, images, frames and connections from. It is the one security header that can break a page at HTTP 200, which is why it has a report-only mode; a policy that permits 'unsafe-inline' for scripts removes most of the protection it was written for.

Measured: CSP_REPORTING_BROKEN 0.3%

CSP nonce

A random value generated per response, placed on the Content-Security-Policy header and on every inline script the server writes, so that only scripts carrying that value run. It defeats injected scripts without allowlisting hosts; it fails silently when any response path skips the swap, because the header then names a nonce the page does not carry.

edge cache

A copy of a response stored at a CDN's point of presence and served to later requests for the same key without consulting the origin. It is keyed on the URL, so a request with an unfamiliar query string usually misses it; the copy can outlive a deleted path, a changed file or an origin that has started refusing, and nothing about the bare URL reveals which.

Measured: STALE_CACHE_SERVED 16.7% · ORIGIN_BLOCKED_BEHIND_CACHE 0.3% · MACHINE_FILE_CACHE_STALE 0.7%

cache split

The state in which a bare request for a URL and a cache-busted request for the same URL return different answers: a different status, a different content type, or different bytes. A single fetch cannot see it; two can. A split on a machine file means what a crawler receives depends on cache state rather than on the site.

Measured: MACHINE_FILE_CACHE_STALE 0.7% · MACHINE_FILE_CACHE_SPLIT 0.1%

cache-buster

A random query parameter appended to a URL so that the request misses every cache keyed on the URL and reaches the origin. It is a measurement tool, not a fix: it shows what the origin serves now, and it says nothing about what the edge is still serving to everyone who asks for the bare URL.

Measured: MACHINE_FILE_CACHE_STALE 0.7% · MACHINE_FILE_CACHE_SPLIT 0.1% · ORIGIN_BLOCKED_BEHIND_CACHE 0.3%

Cache-Control

The response header that states how long, and by whom, a response may be stored: max-age for the browser, s-maxage for shared caches, no-store to forbid storing, private to keep it out of shared caches. It governs future stores only; a copy already sitting at an edge is unaffected by a header the origin starts sending later.

Age header

A response header, in seconds, saying how long ago the served copy was built. It is the one direct signal that a response came from a cache rather than the origin, and its value is what a scanner reads before deciding whether an old copy is stale: an old copy that matches a fresh one is a note, an old copy that differs is an edit stuck behind the cache.

Measured: STALE_CACHE_SERVED 16.7%

render-blocking resource

A stylesheet or synchronous script the browser must fetch and process before it can paint anything below it. Each one adds a network round trip to the first paint; deferring scripts and inlining or preloading critical CSS are the fixes. The count is measured from the delivered HTML, not from a rendered timeline.

payload text share

The share of a delivered page's bytes that are visible text once markup, scripts, styles and inline data are removed. A page under about ten percent text is mostly code: a reader of the bytes, which is what a crawler is, has to work through nine parts machinery for one part content. It is a ratio, so a long page with heavy tooling can still read low.

Measured: PAGE_IS_MOSTLY_CODE 22.6%

lazy loading

Deferring an image or frame until it is about to enter the viewport, by the loading=lazy attribute or by script. It saves bandwidth below the fold and costs paint time above it: an image the page opens with that is marked lazy arrives later than it would have unmarked, and a script-only lazy loader shows a crawler nothing at all.

404 and 410 status codes

404 says the resource was not found; 410 says it is gone and will not return. Engines drop a 410 faster than a 404, which they keep retrying for a while in case it was an error. A page that is missing but returns 200 with an error message, a soft 404, is worse than either, because nothing tells the engine to stop.

← Crawling and indexing  ·  Linked data and entities →

All 668 terms across 19 areas.