CrawlCheck

Glossary · Indexing and discovery

duplicate content

The same or substantially the same text reachable at more than one URL, on one site or across sites. Engines pick one copy to index and the choice is theirs unless a canonical, a redirect or a noindex makes it.

What duplicate content means

The same or substantially the same text reachable at more than one URL, on one site or across sites. Engines pick one copy to index and the choice is theirs unless a canonical, a redirect or a noindex makes it. Most audit tools accuse sites of it by heuristics; measuring it means comparing text, which almost none of them do.

Terms this definition uses

canonical · noindex

Related terms in Indexing and discovery

Whether a page can be found and kept, separately from whether it can be fetched.

index coverage · canonical · noindex · X-Robots-Tag · orphan page · E-E-A-T · E-E-A-T proxies · topical authority · Soft 404 · case consistency · accessibility structure · landmark region · accessible name · skip link · main landmark · alt attribute · form control name · duplicate id · IndexNow · faceted navigation · thin content · near-duplicate · boilerplate · MinHash · shingle · dead anchor · placeholder content · image placeholder · alt text · EXIF

Questions about duplicate content

What is duplicate content?

The same or substantially the same text reachable at more than one URL, on one site or across sites. Engines pick one copy to index and the choice is theirs unless a canonical, a redirect or a noindex makes it.

Which area of the glossary does duplicate content belong to?

Indexing and discovery: Whether a page can be found and kept, separately from whether it can be fetched.

← thin content  ·  near-duplicate →

See it in the full glossary · 668 terms across 19 areas.