CrawlCheck

Glossary · Indexing and discovery

MinHash

A technique for estimating how similar two sets are from small fixed-size sketches of each, so that page overlap can be compared without storing the pages.

What MinHash means

A technique for estimating how similar two sets are from small fixed-size sketches of each, so that page overlap can be compared without storing the pages. With 64 hashes the estimate is within a few percent of the exact value, which is enough to place a pair in a band and not enough to publish as a number.

Terms this definition uses

Place

Related terms in Indexing and discovery

Whether a page can be found and kept, separately from whether it can be fetched.

index coverage · canonical · noindex · X-Robots-Tag · orphan page · E-E-A-T · E-E-A-T proxies · topical authority · Soft 404 · case consistency · accessibility structure · landmark region · accessible name · skip link · main landmark · alt attribute · form control name · duplicate id · IndexNow · faceted navigation · thin content · duplicate content · near-duplicate · boilerplate · shingle · dead anchor · placeholder content · image placeholder · alt text · EXIF

Questions about MinHash

What is MinHash?

A technique for estimating how similar two sets are from small fixed-size sketches of each, so that page overlap can be compared without storing the pages.

What does MinHash not show or guarantee?

A technique for estimating how similar two sets are from small fixed-size sketches of each, so that page overlap can be compared without storing the pages. With 64 hashes the estimate is within a few percent of the exact value, which is enough to place a pair in a band and not enough to publish as a number.

Which area of the glossary does MinHash belong to?

Indexing and discovery: Whether a page can be found and kept, separately from whether it can be fetched.

← boilerplate  ·  shingle →

See it in the full glossary · 668 terms across 19 areas.