CrawlCheck

Findings · 2026-09-30 · By · 0 views

Eight contradictions in our own copy, rechecked today

An outside audit listed eight self-contradictions on crawlcheck.io. Each one rechecked against the live site: what the copy says now, which were real, and the one still wrong this morning.

On 25 September 2026 an outside audit listed eight self-contradictions in crawlcheck.io's copy. Rechecked against the live Worker code and pages: identity count and section count now derive from one source; 'no daily cap' (site) and the monthly quota (API) are two products; 'regenerates' and 'kept 90 days' are one sentence; 'identifies on every request' was rewritten on 23 September; Agency 40 and Network 150+ are consistent; the app is in build. One was still false this morning, a post claiming no scanned domain is ever named in anything published, and it is corrected in this build.

On 25 September an outside audit of crawlcheck.io listed eight places where the site's own copy contradicted itself: one page said one number and another page said a different one, or a claim on the homepage was disproved by a feature two clicks away. Most were fixed the next day. This post rechecks all eight against the live site today, quotes what the copy now says for each, and reports the one that was still wrong this morning and is corrected in the build that publishes this post. The point of doing it in public is the same as the point of the scanner: a claim that cannot be checked is not worth much, and that applies to claims about ourselves.

How each one was checked #

Not from memory. The deployed Worker bundle was searched for each phrase the audit quoted and for its rivals, and the public counts endpoint and the pages themselves were read. Where two numbers had to agree, both were read from the live source: the identity list in the code, the section weights table, the plan limits object. The table is the result; the sections after it are the detail for the ones that needed judgement.

#The audit's claimWhat the copy says todayState
1identities: 14 on one page, 15 on another15 everywhere: the identity list has 15 entries, the counts endpoint reports 15, the homepage prints 15consistent
2sections: '24 scored' vs '35''35 measured sections with 24 of them scored'; the counts endpoint reports 35 and 24consistent — both true, now said together
3Agency 'no daily cap' vs 5,000 API scans a month'no daily cap' describes scanning on the site; the monthly figure is the API quota of a different productconsistent — two products
4'no scanned domain is ever named publicly' vs a directory that names domainsone post still said 'in anything we publish — including this article'was false; fixed in this build
5report 'regenerates, never stale' vs 90-day retention'a URL, kept 90 days, that regenerates so it stays true'consistent — both facts in one sentence
6'every request identifies as CrawlCheck' vs named-UA probesthe phrase no longer exists; the policy page says the scanner fetches as 15 named identitiesfixed 23 September
7'>40 sites: Agency, quoted' vs a plan called NetworkAgency is 40 domains; Network is 150+ domains, quotedconsistent
8app 'in build' vs an installable web app'in build for iPhone and Android'; no web app manifest is servedconsistent as written

The one that was still wrong #

Item four. The findings post on what a crawler actually receives ended a section with the sentence no scanned domain is ever named in anything we publish — including this article. The second half was true; the first half stopped being true on 10 September, when the directory and the registry began listing vendors and verified sites by name with their grades, and it was contradicted again by a post that names a newspaper. The dataset page's own FAQ had the scoped version right: no scanned domain is ever named there, meaning on the dataset page. The post's sentence now says that the article names no site and the dataset's aggregate figures name none, and that the directory and registry do name sites, by a stated method with a dispute route. Same facts, no overreach.

The three that were not contradictions #

Items two, three and five are cases where two true statements were read as one false one. A site has 35 measured sections and 24 scored ones; the copy that said 24 was not wrong, it was incomplete, and it now says both. The licensed lane on the site has no daily cap, and the API has a monthly quota; those are two products with two limits, and the pricing copy now names which one each sentence is about. A report URL is kept for 90 days and regenerates when opened during those 90 days; the sentence that says so says both halves. The audit's instinct was right in each case, because a reader who meets the two statements on different pages cannot tell they are compatible. The fix is to put them in one place.

The two that were real and were fixed first #

Item six was the policy page, which said the scanner identifies itself as CrawlCheck on every request. It does not: the delivery comparison fetches the homepage as fifteen identities, twelve of them the user-agent strings of named crawlers, because that is the measurement. The policy page was rewritten on 23 September to say exactly what is sent, and the policy page now lists the identities. Item one was a count of those identities that had drifted between pages as the list grew; it is now read from the list itself wherever it is printed, which is the only fix for a number that changes.

The two that stand as written #

Item seven and item eight. The pricing page says a customer with more than forty sites is quoted, and the plan that covers it is called Network, from 150 domains; the audit read the absence of the word Network beside the forty as a gap, and the pricing page now says it in one line: More than 40 domains: Network, quoted. Item eight is the app: the copy says it is in build for iPhone and Android, and the audit believed an installable web app existed that the copy should mention. No web app manifest is served from crawlcheck.io, so a browser cannot install anything, and the copy is accurate as written. It will change when there is something to install, and the check that decides it is the manifest, not the wording.

What was not rechecked #

Whether the six fixes made on 26 September were the right fixes at the time. That is history, and the deployed code from that day is in the repository for anyone who wants it. What this post can say is what the copy says now, read from the running code this morning, and that after this build all eight items read consistently. A ninth item found on the way, the identity count being printed from a hand-typed string in one place, was changed to read the list length in the same pass; it is not in the table because the audit did not list it.

The rule this produces #

A number that appears on more than one page is derived from one source or it will drift. The counts endpoint at /api/public/counts is that source for identities, sections and scored sections, and the pages read it. A claim that begins never or every is checked against every feature, not the feature it was written beside. And an outside audit that lists eight items gets a public answer to all eight, including the ones it was wrong about, because that is the standard the scanner holds other sites to. The rest of what changed in that week is in the release notes.

Live, as you read this: the corpus now holds 7,808 domains across 2,468 scans. The figures in this piece were measured on the date above; this line is not.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What were the eight contradictions?

Identity count, scored-section count, Agency daily cap versus API monthly quota, 'no domain is ever named' versus the directory, 'regenerates' versus 90-day retention, 'identifies on every request' versus named-UA probes, Agency versus Network above 40 sites, and the app described as in build.

How many were real?

Rechecked today: three were two true statements read as one false one and are now stated together; two were real and fixed on 23 and 26 September; one was still wrong this morning and is corrected in this build; two were consistent as written.

Which one was still wrong?

A post said no scanned domain is ever named in anything CrawlCheck publishes. The directory, the registry and a post about a newspaper all name sites. The sentence now scopes the claim to the article and the dataset's aggregate figures.

Does CrawlCheck identify itself on every request?

No. The delivery comparison fetches the homepage as 15 identities, 12 of them the user-agent strings of named crawlers, because the measurement is whether the site treats them differently. The policy page lists them.

Is 'no daily cap' compatible with a monthly API quota?

Yes. The cap-free scanning is the licensed lane on the site; the monthly quota is the API product. The pricing copy names which product each limit belongs to.

How were the claims checked?

By searching the deployed Worker code for each quoted phrase and its rivals, and reading the public counts endpoint and the pages, rather than from memory of what was fixed.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works