CrawlCheck

Glossary · area 11 of 19

Crawling and indexing

How a crawler discovers, schedules, fetches and selects pages, and why a fetched page is not an indexed one.

45 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

45terms in this area
5with a live finding rate

Crawler identity

The operator and purpose attributed to an automated requester. A user-agent string asserts identity; verification requires network, DNS, or cryptographic evidence tied to the operator.

Product token

The crawler name used to select a robots.txt group. RFC 9309 limits its characters and expects it to appear within the crawler’s HTTP identification string.

Identification string

The complete User-Agent header sent with a crawler request. It can include version and documentation information, but any client can copy it.

Crawl frontier

The working collection of discovered URLs waiting for crawl decisions. Presence on the frontier proves discovery, not that the URL will be fetched.

Crawl queue

The ordered subset of URLs scheduled for fetching. Priority, host limits, failures, and budgets can indefinitely delay an item in the queue.

Crawl scheduler

Logic deciding when and in what order URLs are requested. Its decisions reflect crawler objectives and resource limits unavailable to the site owner.

Seed URL

A starting address from which a crawler discovers additional resources. A seed helps initiate traversal but does not guarantee coverage beyond reachable links.

URL discovery

Learning that an address exists through links, sitemaps, submissions, feeds, redirects, or prior records. Discovery is earlier than crawling and indexing.

Recrawl interval

Time between requests for the same resource. It is chosen by the crawler and can differ by URL, change rate, importance, and observed server behavior.

Crawl frequency

How often a crawler requests a site or path during a stated window. It measures activity, not index coverage, ranking, or business value.

Host politeness

Delaying or limiting requests so one origin is not overloaded. It is an operator behavior rather than a guarantee in robots.txt.

Crawl rate

Request volume per time unit for a crawler, host, or path. A rate without window, identity verification, and cache context is not comparable.

Crawl-delay

A non-standard robots.txt directive requesting time between crawler requests. Support varies, and RFC 9309 does not define it.

Robots group merging

Combining rules from every group that matches the same product token. Under RFC 9309, matching groups are combined rather than only the first being used.

Measured: ROBOTS_RULES_SHADOWED 7.2%

Longest-match rule

Choosing the Allow or Disallow path with the most matching octets. File order is not the deciding factor when path specificity differs.

Wildcard rule

A robots path pattern using `*` to match a sequence of characters. Implementations should be tested because extensions and encoding details can change apparent matches.

End-anchor rule

A robots path ending in `$` to require the match at the end of the URI path. It narrows a rule but remains advisory access policy.

Robots cache

A crawler’s stored copy of robots.txt. A corrected file may not affect requests until that copy expires or is refreshed.

Robots parse error

Syntax or encoding that prevents some directives from being interpreted as intended. A successful HTTP response does not prove that any crawler applied the file.

Measured: ROBOTS_IS_HTML 0.2% · ROBOTS_CTYPE 0.1%

Robots redirect

An HTTP redirect encountered while requesting `/robots.txt`. Limited redirect following may be supported, but chains and off-host targets make behavior operator-specific.

Robots unavailable

A robots.txt request that cannot be completed because of server or network failure. RFC behavior depends on failure class and duration; it is not equivalent to an empty file.

Measured: ROBOTS_NOT_200 4% · ROBOTS_BLOCKED 0.4%

Sitemap

A machine-readable list of canonical URLs and optional metadata offered for discovery. Submission helps crawlers learn URLs but does not guarantee crawling, indexing, or ranking.

Measured: NO_SITEMAP_FOUND 6.8% · SITEMAP_BLOCKED 1.8% · SITEMAP_IS_HTML 0.9%

Sitemap URL set

An XML sitemap document whose `urlset` contains page records. Its URL count describes declarations, not unique indexed pages.

Sitemap lastmod

A declared resource modification time. It is useful only when it reflects meaningful page changes rather than every build or request.

Measured: SITEMAP_NO_LASTMOD 4.6%

Sitemap hreflang

Language and regional alternates declared within sitemap entries. Valid syntax does not prove reciprocal, canonical, or content-level consistency.

Image sitemap

Sitemap extensions describing images associated with pages. They support discovery, not ownership, licensing, image indexing, or appearance in results.

Video sitemap

Sitemap extensions providing video metadata and locations. The crawler may ignore unsupported, inaccessible, contradictory, or low-quality declarations.

News sitemap

A sitemap limited to recent news publication URLs and metadata. It supports discovery for eligible publishers, not acceptance into a news surface.

Canonical cluster

A group of duplicate or near-duplicate URLs an engine treats as versions of one document. The engine, not the site’s canonical tag alone, selects the representative.

Duplicate cluster

URLs grouped because their content is judged substantially equivalent. Incorrect clustering can suppress a distinct page; the site cannot directly inspect every engine’s cluster.

Index selection

The decision to store or retain a crawled resource in an index. Fetch success is necessary for many pages but never sufficient for selection.

Crawl demand

An engine’s estimated value in revisiting a URL based on importance, expected change, and other signals. It is inferred externally, not published as a per-page score.

Rendering queue

Crawled pages waiting for a rendering service to execute supported scripts. Entry does not guarantee timely completion or equivalence with a user’s browser.

Rendered HTML

The DOM state after a renderer executes supported page code. It can differ by timing, viewport, cookies, geography, blocked resources, and runtime limits.

Indexing pipeline

The stages that parse, render, canonicalize, classify, and store crawled resources. Observing one stage does not reveal whether later stages accepted the page.

Index freshness

How recently an indexed representation reflects the live resource. A recent crawl timestamp does not guarantee that visible results use the new version.

Index lag

Delay between a page change and its appearance in an engine’s stored or served representation. It combines recrawl, processing, and serving delays that external tools cannot fully separate.

Deindexing

Removing a URL or representation from searchable storage or results. It can result from directives, deletion, quality decisions, legal action, canonicalization, or temporary system state.

Removal request

A submitted request to hide or delete a URL from a search product. Temporary concealment is not the same as deletion from every underlying index or cache.

URL inspection

A diagnostic view of an engine’s known state for one URL. It is a sampled product report, not a complete history of all crawler and index decisions.

Crawl anomaly

Request behavior that departs from a site or crawler baseline, such as a spike, new path pattern, or status shift. Anomaly does not identify cause without logs and controls.

Bot spoofing

Sending a request with another crawler’s claimed identity. User-agent matching detects the claim; operator-linked verification is needed to establish the spoof.

Shared IP range

Network space used by several products or request purposes from one operator. Range membership can verify operator infrastructure without proving which internal product initiated a request.

Web Bot Auth

A method for bots to cryptographically sign HTTP requests so recipients can authenticate identity without relying only on network location. A valid signature proves key control, not benign purpose.

← Retrieval engineering  ·  HTTP, edge and rendering →

All 668 terms across 19 areas.