CrawlCheck

Glossary · Indexing and discovery

near-duplicate

Two pages whose core text overlaps substantially once shared navigation, footer and boilerplate are removed.

What near-duplicate means

Two pages whose core text overlaps substantially once shared navigation, footer and boilerplate are removed. Measured by comparing sets of overlapping word sequences, and reported in bands rather than percentages, because the estimate quantises and a precise-looking number would overstate what the method knows.

What near-duplicate can and cannot support

It can supportIt cannot support
Two pages whose core text overlaps substantially once shared navigation, footer and boilerplate are removed.Measured by comparing sets of overlapping word sequences, and reported in bands rather than percentages, because the estimate quantises and a precise-looking number would overstate what the method knows.

Terms this definition uses

boilerplate

Definitions that name near-duplicate

Canonical cluster

Related terms in Indexing and discovery

Whether a page can be found and kept, separately from whether it can be fetched.

index coverage · canonical · noindex · X-Robots-Tag · orphan page · E-E-A-T · E-E-A-T proxies · topical authority · Soft 404 · case consistency · accessibility structure · landmark region · accessible name · skip link · main landmark · alt attribute · form control name · duplicate id · IndexNow · faceted navigation · thin content · duplicate content · boilerplate · MinHash · shingle · dead anchor · placeholder content · image placeholder · alt text · EXIF

Questions about near-duplicate

What is near-duplicate?

Two pages whose core text overlaps substantially once shared navigation, footer and boilerplate are removed.

What does near-duplicate not show or guarantee?

Measured by comparing sets of overlapping word sequences, and reported in bands rather than percentages, because the estimate quantises and a precise-looking number would overstate what the method knows.

Which area of the glossary does near-duplicate belong to?

Indexing and discovery: Whether a page can be found and kept, separately from whether it can be fetched.

← duplicate content  ·  boilerplate →

See it in the full glossary · 668 terms across 19 areas.