CrawlCheck

Glossary · Crawling and indexing

Indexing pipeline

The stages that parse, render, canonicalize, classify, and store crawled resources. Observing one stage does not reveal whether later stages accepted the page.

Terms this definition uses

canonical

Crawling and indexing

How a crawler discovers, schedules, fetches and selects pages, and why a fetched page is not an indexed one.

Crawler identity · Product token · Identification string · Crawl frontier · Crawl queue · Crawl scheduler · Seed URL · URL discovery · Link extraction · Recrawl interval · Crawl frequency · Host politeness · Crawl rate · Crawl-delay · Robots group merging · Longest-match rule · Wildcard rule · End-anchor rule · Robots cache · Robots parse error · Robots redirect · Robots unavailable · Sitemap · Sitemap URL set · Sitemap lastmod · Sitemap hreflang · Image sitemap · Video sitemap · News sitemap · Canonical cluster · Duplicate cluster · Index selection · Crawl demand · Rendering queue · Rendered HTML · Index freshness · Index lag · Deindexing · Removal request · URL inspection · Crawl anomaly · Bot spoofing · Shared IP range · Web Bot Auth

Rendered HTML  ·  Index freshness

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.