CrawlCheck

Compare

Two sites, one scan, side by side

The same measurement a customer gets, run on both homepages: reach, read, quote. Where they differ is where the evidence differs.

Two scans, ~20 seconds each if either site is new to us. Nothing is stored beyond the reports themselves. Need the whole field? One site against up to five competitors.

Example, from stored records denverpost.com against speakeasy.com, the top-scoring site in the directory. Enter any two domains above to run your own.

denverpost.com

82
Cgrade
81Reach
79Read
85Quote

8 findings · 2026-09-24 · full report

vs

speakeasy.com

90
Cgrade
96Reach
85Read
86Quote

4 findings · 2026-09-26 · full report

In plain English

  • speakeasy.com shows answer engines more usable evidence, by 8 points.
  • Crawler permissions: speakeasy.com is 67 ahead.
  • Page weight & text: speakeasy.com is 60 ahead.
  • Structured data: speakeasy.com is 40 ahead.

Every section, side by side

Sectiondenverpost.comspeakeasy.comGap
Machine files
robots.txt, llms.txt, sitemap and the entity graph: present, readable and served as the right type
6067-7
Agent surfaces
Emerging standards (agents.md, API catalog, MCP card). Measured for information only
n/an/a—
Server & hosting
What the edge and origin told us: cache state, compression, the software behind the page
80100-20
Structured data
The JSON-LD entity graph: stable @ids, connected nodes, credentials and locations declared
60100-40
Linked data
Whether the structured data parses as proper RDF and nothing dangles
1001000
Authors & trust signals
Named people, dated articles, an organisation with contact details: the proxies an engine uses for trust
57570
Accessibility structure
Language, alt text, labelled controls and one main landmark, read from the delivered HTML
8886+2
Entity corroboration
Whether the identity links the site declares point back at it, and whether a Wikidata record names this domain
1001000
Content entities
Named things in the page text that a knowledge graph recognises. Measured, not scored
n/an/a—
Quotable content
Whether the sentences an answer engine would lift are on the page in a form it can take: a defining lead, short factual paragraphs, answered questions, few first-person openers
67670
Service area map
Every place the schema declares, placed on one map around the declared centre and radius: which have a page, which fall outside the circle. Measured, not scored
n/an/a—
Whole site
Up to 60 declared pages, read once a day after a scan: not-200s, thin and mostly-code pages, missing JSON-LD, canonical elsewhere, noindex, missing or duplicate h1, orphans, undeclared pages, duplicate titles, and the median words, text share and response time across them. Every count is free; Watch names which pages. Measured, not scored
n/an/a—
Freshness
Every date the site publishes about itself: sitemap lastmod, Last-Modified, dateModified, dated elements, and where they disagree. Measured, not scored
n/an/a—
Crawl waste
Fetches that produce no new page: parameter and case variants, http links, redirecting or gone declared URLs, canonical-elsewhere and noindex pages. Scored from version 23
10067+33
Page subject
Which business entity the page is about and how every other business entity on it is related: chain, subsidiary, department, or undeclared. Scored from version 23
1001000
Local intents
Which of the 446 recovered Maps intent types the site names about itself, and which of those have a page the homepage links to. Measured, not scored
n/an/a—
Conversion path
Phone link, form and call-to-action on the page, and whether they sit before the fold. Measured, not scored
n/an/a—
Crawler permissions
What robots.txt actually tells each named AI and search crawler it may do
33100-67
Speed
How fast the page answered our fetch, plus real-visitor field data where Google has it
6789-22
Page weight & text
How many of the delivered bytes are readable words, and how much is markup, scripts and styling
2080-60
Machine-file chain
robots.txt names the sitemap, llms.txt points onward, the entity graph points home: the chain a crawler follows
10086+14
Claim consistency
Name, phone, city and years-in-business match between the structured data and the visible page
10075+25
Instructions for agents
Text aimed at AI agents rather than people: present, visible, and not hidden
1001000
What each crawler gets
Fifteen client identities fetched the page; this is whether they all got the same thing
88100-12
Mobile
Viewport, zoom, icons and manifest: whether a phone gets a page it can use
1001000
Local presence & NAP
Street address, coordinates, hours, phone, and the directory listings the site itself declares
92n/a—
Sitemap
A sitemap exists, resolves, and declares the pages the homepage links to
1001000
Internal links
Anchors carry words, targets resolve, and the link graph is not mostly repeats
1001000
URL naming
Lowercase hosts and paths, no case-only duplicates, machine files spelt correctly
80100-20
Image signals
Alt text, declared dimensions, lazy-loading, modern formats, hero priority and the entity image in JSON-LD.
80800

Gap is denverpost.com minus speakeasy.com. n/a means that section could not be measured on that site; it is not scored as zero.

Findings

Only on denverpost.com

  • ROBOTS_RULES_SHADOWED robots.txt rules do not apply to named agents
  • ENTITY_COLLISION Two business records on this page claim the same identity
  • RENDER_RESOURCE_BLOCKED The homepage depends on files this site's own robots.txt disallows
  • PAGE_IS_MOSTLY_CODE Almost nothing a machine receives from this page is readable text
  • GEO_COUNTRY_MISMATCH The published coordinates are not in the country the address names
  • NO_LLMS_TXT No llms.txt
  • NAP_UNDECLARED_LISTING 4 off-site listings carry this phone number and the site never links them

Only on speakeasy.com

  • HOST_DUPLICATE_200 www and the apex both serve the site instead of one redirecting
  • ENTITY_NO_COORDINATES A business record gives an address but no coordinates
  • STALE_CACHE_SERVED Visitors and crawlers are being served an old copy of this page

On both

  • MACHINE_CHAIN_HEAVY The machine files cost more than the page

Both sites measured by the same scan from the same vantage. A higher number means more of the evidence an answer engine needs is present, not that one site ranks above the other. JSON