Crawler identity
The operator and purpose attributed to an automated requester. A user-agent string asserts identity; verification requires network, DNS, or cryptographic evidence tied to the operator.
Glossary · area 11 of 19
How a crawler discovers, schedules, fetches and selects pages, and why a fetched page is not an indexed one.
45 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
The operator and purpose attributed to an automated requester. A user-agent string asserts identity; verification requires network, DNS, or cryptographic evidence tied to the operator.
The crawler name used to select a robots.txt group. RFC 9309 limits its characters and expects it to appear within the crawler’s HTTP identification string.
The complete User-Agent header sent with a crawler request. It can include version and documentation information, but any client can copy it.
The working collection of discovered URLs waiting for crawl decisions. Presence on the frontier proves discovery, not that the URL will be fetched.
The ordered subset of URLs scheduled for fetching. Priority, host limits, failures, and budgets can indefinitely delay an item in the queue.
Logic deciding when and in what order URLs are requested. Its decisions reflect crawler objectives and resource limits unavailable to the site owner.
A starting address from which a crawler discovers additional resources. A seed helps initiate traversal but does not guarantee coverage beyond reachable links.
Learning that an address exists through links, sitemaps, submissions, feeds, redirects, or prior records. Discovery is earlier than crawling and indexing.
Parsing documents to collect destination URLs. A string that resembles a link may be ignored if it is not exposed in a crawlable element or rendered state.
Time between requests for the same resource. It is chosen by the crawler and can differ by URL, change rate, importance, and observed server behavior.
How often a crawler requests a site or path during a stated window. It measures activity, not index coverage, ranking, or business value.
Delaying or limiting requests so one origin is not overloaded. It is an operator behavior rather than a guarantee in robots.txt.
Request volume per time unit for a crawler, host, or path. A rate without window, identity verification, and cache context is not comparable.
A non-standard robots.txt directive requesting time between crawler requests. Support varies, and RFC 9309 does not define it.
Combining rules from every group that matches the same product token. Under RFC 9309, matching groups are combined rather than only the first being used.
Measured: ROBOTS_RULES_SHADOWED 7.2%
Choosing the Allow or Disallow path with the most matching octets. File order is not the deciding factor when path specificity differs.
A robots path pattern using `*` to match a sequence of characters. Implementations should be tested because extensions and encoding details can change apparent matches.
A robots path ending in `$` to require the match at the end of the URI path. It narrows a rule but remains advisory access policy.
A crawler’s stored copy of robots.txt. A corrected file may not affect requests until that copy expires or is refreshed.
Syntax or encoding that prevents some directives from being interpreted as intended. A successful HTTP response does not prove that any crawler applied the file.
Measured: ROBOTS_IS_HTML 0.2% · ROBOTS_CTYPE 0.1%
An HTTP redirect encountered while requesting `/robots.txt`. Limited redirect following may be supported, but chains and off-host targets make behavior operator-specific.
A robots.txt request that cannot be completed because of server or network failure. RFC behavior depends on failure class and duration; it is not equivalent to an empty file.
Measured: ROBOTS_NOT_200 4% · ROBOTS_BLOCKED 0.4%
A machine-readable list of canonical URLs and optional metadata offered for discovery. Submission helps crawlers learn URLs but does not guarantee crawling, indexing, or ranking.
Measured: NO_SITEMAP_FOUND 6.8% · SITEMAP_BLOCKED 1.8% · SITEMAP_IS_HTML 0.9%
An XML sitemap document whose `urlset` contains page records. Its URL count describes declarations, not unique indexed pages.
A declared resource modification time. It is useful only when it reflects meaningful page changes rather than every build or request.
Measured: SITEMAP_NO_LASTMOD 4.6%
Language and regional alternates declared within sitemap entries. Valid syntax does not prove reciprocal, canonical, or content-level consistency.
Sitemap extensions describing images associated with pages. They support discovery, not ownership, licensing, image indexing, or appearance in results.
Sitemap extensions providing video metadata and locations. The crawler may ignore unsupported, inaccessible, contradictory, or low-quality declarations.
A sitemap limited to recent news publication URLs and metadata. It supports discovery for eligible publishers, not acceptance into a news surface.
A group of duplicate or near-duplicate URLs an engine treats as versions of one document. The engine, not the site’s canonical tag alone, selects the representative.
URLs grouped because their content is judged substantially equivalent. Incorrect clustering can suppress a distinct page; the site cannot directly inspect every engine’s cluster.
The decision to store or retain a crawled resource in an index. Fetch success is necessary for many pages but never sufficient for selection.
An engine’s estimated value in revisiting a URL based on importance, expected change, and other signals. It is inferred externally, not published as a per-page score.
Crawled pages waiting for a rendering service to execute supported scripts. Entry does not guarantee timely completion or equivalence with a user’s browser.
The DOM state after a renderer executes supported page code. It can differ by timing, viewport, cookies, geography, blocked resources, and runtime limits.
The stages that parse, render, canonicalize, classify, and store crawled resources. Observing one stage does not reveal whether later stages accepted the page.
How recently an indexed representation reflects the live resource. A recent crawl timestamp does not guarantee that visible results use the new version.
Delay between a page change and its appearance in an engine’s stored or served representation. It combines recrawl, processing, and serving delays that external tools cannot fully separate.
Removing a URL or representation from searchable storage or results. It can result from directives, deletion, quality decisions, legal action, canonicalization, or temporary system state.
A submitted request to hide or delete a URL from a search product. Temporary concealment is not the same as deletion from every underlying index or cache.
A diagnostic view of an engine’s known state for one URL. It is a sampled product report, not a complete history of all crawler and index decisions.
Request behavior that departs from a site or crawler baseline, such as a spike, new path pattern, or status shift. Anomaly does not identify cause without logs and controls.
Sending a request with another crawler’s claimed identity. User-agent matching detects the claim; operator-linked verification is needed to establish the spoof.
Network space used by several products or request purposes from one operator. Range membership can verify operator infrastructure without proving which internal product initiated a request.
A method for bots to cryptographically sign HTTP requests so recipients can authenticate identity without relying only on network location. A valid signature proves key control, not benign purpose.