CrawlCheck

Glossary · Crawling and indexing

Crawl scheduler

Logic deciding when and in what order URLs are requested. Its decisions reflect crawler objectives and resource limits unavailable to the site owner.

Terms this definition uses

Object

Crawling and indexing

How a crawler discovers, schedules, fetches and selects pages, and why a fetched page is not an indexed one.

Crawler identity · Product token · Identification string · Crawl frontier · Crawl queue · Seed URL · URL discovery · Link extraction · Recrawl interval · Crawl frequency · Host politeness · Crawl rate · Crawl-delay · Robots group merging · Longest-match rule · Wildcard rule · End-anchor rule · Robots cache · Robots parse error · Robots redirect · Robots unavailable · Sitemap · Sitemap URL set · Sitemap lastmod · Sitemap hreflang · Image sitemap · Video sitemap · News sitemap · Canonical cluster · Duplicate cluster · Index selection · Crawl demand · Rendering queue · Rendered HTML · Indexing pipeline · Index freshness · Index lag · Deindexing · Removal request · URL inspection · Crawl anomaly · Bot spoofing · Shared IP range · Web Bot Auth

Crawl queue  ·  Seed URL

See it in the full glossary · 579 terms across 19 areas. Scan a site to see which of these apply to it.