Glossary · Crawling and indexing
Indexing pipeline
The stages that parse, render, canonicalize, classify, and store crawled resources. Observing one stage does not reveal whether later stages accepted the page.
Terms this definition uses
Crawling and indexing
How a crawler discovers, schedules, fetches and selects pages, and why a fetched page is not an indexed one.
Crawler identity · Product token · Identification string · Crawl frontier · Crawl queue · Crawl scheduler · Seed URL · URL discovery · Link extraction · Recrawl interval · Crawl frequency · Host politeness · Crawl rate · Crawl-delay · Robots group merging · Longest-match rule · Wildcard rule · End-anchor rule · Robots cache · Robots parse error · Robots redirect · Robots unavailable · Sitemap · Sitemap URL set · Sitemap lastmod · Sitemap hreflang · Image sitemap · Video sitemap · News sitemap · Canonical cluster · Duplicate cluster · Index selection · Crawl demand · Rendering queue · Rendered HTML · Index freshness · Index lag · Deindexing · Removal request · URL inspection · Crawl anomaly · Bot spoofing · Shared IP range · Web Bot Auth
← Rendered HTML · Index freshness →
See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.