CrawlCheck

Findings · 2026-08-21 · By · 0 views

Thin-content checks miss parking pages: 783 words, still parked

Scanners grade parked domains as if they were sites. The obvious fix — flag pages with almost no text — fails, because commercial landers are not thin. What works is identifying the operator serving the page.

A word count cannot tell a real page from a parked one. The parking page we measured carried 783 words of generated filler and passed every thin-content threshold in common use, because those checks count words rather than asking whether the words are about the business. What separates the two is specificity — named services, a real address, facts only this business could state — and none of that is measurable by length.

A scanner pointed at a domain that is parked will produce a grade. Every check will run, every section will score, and the report will describe a page that nobody built — confidently, and in detail.

We had this gap. Our scanner already refuses to grade a site that answers a challenge page, on the principle that a refusal is not a measurement. But it would happily grade a domain-sale lander, because a lander is a real page that returns a real 200 with real content.

This is not hypothetical for us. One of our own domains lapsed once and served a parking page for a period. A scan during that window would have handed us a grade for a page belonging to the registrar.

The heuristic that does not work #

The obvious detector is thin content: a parking page is mostly nothing, so count the words and flag anything nearly empty.

We tested that against a live parking page — a HugeDomains sale lander on a domain in our own product niche:

RequestBytesVisible words
Parking lander, browser identity43,891783
Parking lander, GPTBot identity43,891783
Real site on a rebuilt expired domain (control)68,5181,341

783 words. Not an empty page — a full commercial page with navigation, an FAQ, a price, a payment-plan offer, contact details and a search box. Any threshold low enough to catch it would flag a large number of legitimate small business sites, which routinely sit in the same range.

Note also that the lander returns identical bytes to a browser and to a named AI crawler. There is no cloaking to detect, no difference to exploit. It is the same page for everyone.

What does work: the operator, not the page #

A parking page is not written by the domain owner. It is served by an operator, and the operator leaves its name in the markup — in a script host, an asset path, a form action, a link. That marker is present because the operator put it there for its own purposes, and it is absent from every page that is not one of theirs.

Detection therefore starts from the operator, not the content:

The false positive this must never produce #

Buying an expired domain and building a real site on it is a common and entirely legitimate practice. Such a domain is often still listed for sale somewhere, because the owner left the listing up or the marketplace kept the record.

Our control is exactly that case: a 58-engine reference site, 1,341 words of original content, running on a domain that a technology profiler still flags as listed with a domain marketplace. It carries zero operator markers, because no operator is serving it — the owner is.

That is why the rule is written against what the origin actually serves, never against what a third party says about the domain. A listing is a fact about a marketplace. A lander is a fact about the response. Only the second one is ours to measure.

Refuse, do not score zero #

When a domain is parked, the scan stops and returns no grade at all.

A zero would be a claim about a site. There is no site. The report says what the domain served, names the operator, and stops — the same discipline applied to challenge pages and to components that could not be measured.

What transfers #

Method #

Both identities were sent with redirects followed, from outside the networks serving the sites.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Why does word count fail as a thin-content test?

A parking lander can carry hundreds of words of generic filler. What separates content from filler is provenance: who is serving the page and whether the words are about anything.

What does the scanner do with a parked domain?

It refuses to grade it. A score would describe the lander, not a site anyone built, so the record says parked, names the operator where it can, and asks for a rescan once the site is live.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works