Origin server
The authoritative application or host responsible for a resource before intermediary handling. External clients may never reach it directly when a CDN, proxy, or firewall answers first.
Glossary · area 12 of 19
The delivery layer between an origin and a client: proxies, caches, redirects, status codes and the headers that decide what a crawler actually received.
60 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
The authoritative application or host responsible for a resource before intermediary handling. External clients may never reach it directly when a CDN, proxy, or firewall answers first.
A server that receives requests on behalf of upstream applications. It can cache, route, transform, authenticate, or block responses, making observed behavior different from origin behavior.
A distributed network that serves content near clients and absorbs traffic from the origin. It improves delivery but can preserve stale, transformed, or client-specific variants.
A CDN or platform location that processes a request near the requester. Two nodes can hold different cache states or policies at the same moment.
A web application firewall that evaluates HTTP requests against security rules. It can stop attacks and legitimate crawlers, and a block page may still carry HTTP 200.
A provider-specific estimate that traffic is automated. It is a classification signal, not proof of identity, intent, safety, or robots compliance.
The negotiation establishing an encrypted HTTPS connection and server authentication. Success proves transport setup, not that the application will serve the requested resource.
Converting a hostname into network addresses. Successful resolution identifies where to connect; it does not prove the resulting server is healthy or authoritative for content.
Failure to obtain a DNS answer within the client’s limit. It leaves the resource unmeasured rather than proving the host does not exist.
Failure to establish a transport connection within a configured period. It can reflect routing, firewall, congestion, or server conditions, not page quality.
A client ending a request because response data did not arrive in time. The result is incomplete measurement, not necessarily an origin error.
Two or more redirects before a final response. Each hop adds failure and latency opportunities, while clients impose different maximums.
Redirects that eventually revisit a previous URL or repeat indefinitely. A crawler stops after its own threshold, so no final page is retrieved.
A permanent redirect whose target is intended to replace the requested resource. Clients and indexes may still take time to consolidate signals.
A temporary redirect that preserves the possibility that the original URL remains authoritative. Long-term use can still be interpreted differently by search systems.
A temporary redirect that requires preserving the request method and content. It controls HTTP replay semantics, not indexing treatment by every engine.
A permanent redirect that preserves the request method and content. It does not guarantee immediate cache, canonical, or index consolidation.
“Not Modified,” telling a conditional requester to reuse its stored representation. It carries no new body and is meaningful only with the client’s cached copy.
A response indicating no current representation was found or disclosed for the target. It does not state whether the absence is permanent.
“Gone,” indicating the resource is intentionally unavailable and likely permanently removed. Crawlers may process it faster than a 404, but timing is not guaranteed.
“Too Many Requests,” indicating rate limiting. It says the client was throttled, not whether the resource would otherwise be allowed or readable.
A generic unexpected server failure. It identifies server-side failure but not the responsible component or whether retrying will succeed.
“Bad Gateway,” returned when an intermediary receives an invalid response from an upstream server. It can originate at a CDN or proxy rather than the website application.
“Service Unavailable,” indicating temporary inability to handle the request. A `Retry-After` value can guide clients but cannot force them to return.
“Gateway Timeout,” indicating an intermediary waited too long for upstream. It measures that request path, not universal origin availability.
A request for the headers that a corresponding GET would return, without transferring response content. It cannot verify body text, markup, challenge pages, or payload parity.
A request to retrieve a current representation of a target resource. HTTP success does not guarantee that the representation is the intended human page.
The declared media type describing how a response body should be interpreted. A correct header can accompany invalid content, and an incorrect header can hide a valid payload from parsers.
Inferring content format from bytes when metadata is absent or distrusted. Different clients can infer differently, producing parser and security inconsistencies.
A representation transformation such as gzip or Brotli declared for transfer. A client must decode it before assessing the underlying body.
Reducing transferred bytes through an encoding negotiated between client and server. It changes bandwidth and latency, not the semantic amount of content.
A compression format commonly used for web responses. Support and compression level affect transfer size and processing cost, not crawlability by themselves.
An opaque validator identifying a selected representation version. Matching tags support cache validation; they do not prove semantic equivalence across URLs or encodings.
A server-declared timestamp for when a representation changed. Its granularity and accuracy depend on the application generating it.
A conditional request header asking whether the current ETag differs from one already stored. A matching validator permits a 304 instead of a new body.
A request served from stored intermediary or client state. It improves speed but may bypass the origin and expose stale or incorrectly keyed content.
A request that cannot use an existing cache entry and must obtain or generate a representation. It does not necessarily mean the resource was never cached before.
The values a cache uses to decide which stored response matches a request. Missing a relevant dimension can serve one client’s variant to another.
Response metadata naming request headers that affected representation selection. It helps caches separate variants but does not account for unstated edge rules.
A cached response whose age exceeds its freshness lifetime. It may still be served under defined directives or failure policies.
Measured: STALE_CACHE_SERVED 16.7%
A cache directive permitting stale reuse while an asynchronous refresh occurs. It reduces latency while creating a window where clients receive old content.
A redirect generated by CDN or proxy configuration before the origin handles the request. Origin code inspection will not reveal it.
Reaching an origin through a route that avoids its normal edge layer. It isolates origin behavior but does not reproduce what public crawlers receive.
Material agreement between responses served to different identities, locations, or clients. Equal status codes alone are insufficient; bodies, headers, redirects, and readable text must also be compared.
Agreement in response metadata across compared clients. It can reveal policy differences while missing body-level cloaking, challenges, or render dependence.
A link annotation that tells a search engine which language or regional version of a page a URL is, and where the other versions live. It groups alternates; it does not redirect anyone, does not translate anything, and a set that is not reciprocal across every version is ignored by the engine that reads it.
The one address a site declares as the authoritative copy of a page, in a link rel=canonical tag or an HTTP header. It is a hint that engines usually follow, not a redirect: the duplicate stays reachable, and a canonical that points at a page returning 404 or a different host is discarded rather than obeyed.
Measured: HOST_DUPLICATE_200 3.3% · DOMAIN_IS_ALIAS 0.8%
Strict-Transport-Security, a response header that tells a browser which has already reached a site over HTTPS to refuse plain HTTP for max-age seconds. It protects every visit after the first; the first request is only covered by the browser preload list, and the header is ignored if it arrives over HTTP.
Measured: HSTS_MISSING 7.6% · HSTS_PRELOAD_INELIGIBLE 1.9% · HSTS_PRELOAD_REMOVED 0.3%
A list compiled into browsers of hosts that must only be reached over HTTPS, so that even a first request is upgraded. Entry requires a one-year max-age, includeSubDomains and the preload token, and removal takes months to propagate through shipped browsers, which is why sending the token without qualifying is worse than not sending it.
Measured: HSTS_PRELOAD_INELIGIBLE 1.9% · HSTS_PRELOAD_REMOVED 0.3%
A response header that tells a browser which sources a page may load scripts, styles, images, frames and connections from. It is the one security header that can break a page at HTTP 200, which is why it has a report-only mode; a policy that permits 'unsafe-inline' for scripts removes most of the protection it was written for.
Measured: CSP_REPORTING_BROKEN 0.3%
A random value generated per response, placed on the Content-Security-Policy header and on every inline script the server writes, so that only scripts carrying that value run. It defeats injected scripts without allowlisting hosts; it fails silently when any response path skips the swap, because the header then names a nonce the page does not carry.
A copy of a response stored at a CDN's point of presence and served to later requests for the same key without consulting the origin. It is keyed on the URL, so a request with an unfamiliar query string usually misses it; the copy can outlive a deleted path, a changed file or an origin that has started refusing, and nothing about the bare URL reveals which.
Measured: STALE_CACHE_SERVED 16.7% · ORIGIN_BLOCKED_BEHIND_CACHE 0.3% · MACHINE_FILE_CACHE_STALE 0.7%
The state in which a bare request for a URL and a cache-busted request for the same URL return different answers: a different status, a different content type, or different bytes. A single fetch cannot see it; two can. A split on a machine file means what a crawler receives depends on cache state rather than on the site.
Measured: MACHINE_FILE_CACHE_STALE 0.7% · MACHINE_FILE_CACHE_SPLIT 0.1%
A random query parameter appended to a URL so that the request misses every cache keyed on the URL and reaches the origin. It is a measurement tool, not a fix: it shows what the origin serves now, and it says nothing about what the edge is still serving to everyone who asks for the bare URL.
Measured: MACHINE_FILE_CACHE_STALE 0.7% · MACHINE_FILE_CACHE_SPLIT 0.1% · ORIGIN_BLOCKED_BEHIND_CACHE 0.3%
The response header that states how long, and by whom, a response may be stored: max-age for the browser, s-maxage for shared caches, no-store to forbid storing, private to keep it out of shared caches. It governs future stores only; a copy already sitting at an edge is unaffected by a header the origin starts sending later.
A response header, in seconds, saying how long ago the served copy was built. It is the one direct signal that a response came from a cache rather than the origin, and its value is what a scanner reads before deciding whether an old copy is stale: an old copy that matches a fresh one is a note, an old copy that differs is an edit stuck behind the cache.
Measured: STALE_CACHE_SERVED 16.7%
A stylesheet or synchronous script the browser must fetch and process before it can paint anything below it. Each one adds a network round trip to the first paint; deferring scripts and inlining or preloading critical CSS are the fixes. The count is measured from the delivered HTML, not from a rendered timeline.
The share of a delivered page's bytes that are visible text once markup, scripts, styles and inline data are removed. A page under about ten percent text is mostly code: a reader of the bytes, which is what a crawler is, has to work through nine parts machinery for one part content. It is a ratio, so a long page with heavy tooling can still read low.
Measured: PAGE_IS_MOSTLY_CODE 22.6%
Deferring an image or frame until it is about to enter the viewport, by the loading=lazy attribute or by script. It saves bandwidth below the fold and costs paint time above it: an image the page opens with that is marked lazy arrives later than it would have unmarked, and a script-only lazy loader shows a crawler nothing at all.
404 says the resource was not found; 410 says it is gone and will not return. Engines drop a 410 faster than a 404, which they keep retrying for a while in case it was an error. A page that is missing but returns 200 with an error message, a soft 404, is worse than either, because nothing tells the engine to stop.