This sample is shown unlocked, so you can see what a licence includes. On your own domain the free scan gives you the grade, the score, which stage fails and every finding by name, with the top one in full; the detail, the fix and the bytes behind each one are what a licence buys.
Answer engines can reach and read this site. It needs more it can quote.
AI visibility 82/100. An engine can reach and read this site. What is left is giving it something specific to say.
Held at C The score alone would grade higher; a medium-severity finding sets the letter. Clear it and the grade lifts on its own.
What this does not prove
This is what crawlers receive from the site today — reach, read, quote — measured from CrawlCheck’s own network. It does not show whether any engine cites the site, where it ranks, or what traffic or leads follow, and a crawler-identity probe is not a fetch from that operator’s verified addresses. What a number can and cannot say
8
findings
13/23
sections passing
+20
points available
8/60
publisher reference sites measured so far - the comparison appears at 30
11
scans on record
The short version
The website's robots.txt rules do not apply to all named agents, which can cause issues with crawlers. The site also has an identity collision in its structured data, where two nodes share the same URL. Additionally, the homepage depends on files that are disallowed by the site's own robots.txt, which can prevent proper rendering.
In plain English
- 13 of 23 scored sections pass. 7 need work, 3 are failing. The other 7 of the 30 sections are measured for information and never scored.
- Fix first: robots.txt rules do not apply to named agents 40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignores `*` entirely, so those rules do not… This one is holding the grade at C. See the fix list.
- Fix the 24 failing measures in the fix list and the score reads 100. Now 80, a gain of 20 points. That is arithmetic on measurements already taken, not a forecast. Each measure is a named row with its point value.
- Every named answer engine was served the page. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot. Disallowed by policy in robots.txt: Yahoo Slurp, Yandex (umbrella token), ChatGPT search index, ChatGPT live fetch, Claude search index, Claude live fetch, Perplexity index, Perplexity live fetch (a choice, not a defect).
Start here
The three findings holding the grade down most. Open one for what was seen and how to fix it.
serious
robots.txt rules do not apply to named agents
/robots.txt · Open the fix →
serious
Two business records on this page claim the same identity
Open the fix →
serious
The homepage depends on files this site's own robots.txt disallows
/robots.txt · Open the fix →
The three stages, in the order an engine hits a site
Each stage gates the next. An engine that cannot reach a page never reads it; one that cannot read it never quotes it. That is why reach is a ceiling and not just a weight.
Stage 1 · 40% of the score
Reach
Can a named answer-engine crawler get your pages at all?
At least one answer engine was refused, challenged or served less than a browser.
Stage 2 · 30% of the score
Read
Once it has the bytes, can it find the words?
The content is buried in markup, or the files that guide a crawler are missing.
Stage 3 · 30% of the score
Quote
Is there a specific fact it can state and attribute?
The facts an assistant is asked for are declared and consistent.
The stage to fix first is read, because the three run in order.
How this compares
Ranked against 60 sites in the same cohort (wordpress/publisher). This corpus is sites people chose to scan, not the web, so a rank here is a comparison, not a verdict.
Percentile: the share of comparable sites this one scores better than. Lower-is-better metrics (page size, render-blocking) are already inverted.
15
client identities sent, replies compared
5
machine files fetched on both hosts
114
named agents resolved from robots.txt
23
of 30 sections scored
42
invariants checked on this report itself
8
findings
A randomised path that cannot exist was also requested — the control that distinguishes your website from a wall in front of it.
What to fix, in order
Findings first: each one holds the grade down until it is gone (the top one sets the letter). Then headroom: each one lifts the score by the points shown. Open a card for what we saw, why it matters and how to fix it.
Measured 13 times, the oldest finding here first seen 2026-08-20. 8 of them have been here on an earlier reading.
Of 5 tagged findings, 2 are setting changes, 1 is a content edit and 2 need a developer. 3 findings are not tagged. Setting changes are usually the same afternoon; nothing here is ranked by effort, only labelled.
Files, generated from this scan
Each file says at its top exactly what was changed or filled and what was left as a TODO. Publish, then re-scan: done is measured.
1robots.txt rules do not apply to named agents40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignoresMediumSetting
first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignores `*` entirely, so those rules do not apply to any of the named agents (some declare their own, different rules). Worst group: `ai2bot` bypasses /wp-admin/, /cgi-bin/, /wp-includes/, /xmlrpc.php, /wp-content/plugins/, /wp-content/cache/, /trackback/, /comments/. This includes a SEARCH crawler (applebot) — the bypassed paths are being crawled and can enter the index. /robots.txt
Why it matters
A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.
How to fix it
Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.
Setting: a setting on the server, CDN, DNS or a plugin — no code and no writing
Why this was decided
| Rule | ROBOTS_RULES_SHADOWED · rule pack sv26 · scope core |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:75c45dfa8412ab147a92 · finding f1:cbfcabeaf2c8fae39c15 resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 4,425 | c18df974780a | already_held |
21 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
ROBOTS_RULES_SHADOWED · full transcript
2Two business records on this page claim the same identity1 identity collision in this page's structured data: two nodes share the same url.MediumDeveloper
first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
1 identity collision in this page's structured data: two nodes share the same url.
Why it matters
A system resolving which business this page is about has more than one candidate for one identity. Two LocalBusiness or Organization nodes sharing a phone number, an @id or a coordinate pair give a retrieval system no way to decide which record is the entity, so it may pick either or neither. This reads only what the page declares - it does not assert the business is duplicated or wrongly located.
How to fix it
Merge the duplicate business nodes into one, keep a single @id, and remove any plugin that emits a second empty LocalBusiness record.
Developer: a template, theme or application change — a developer touches it
Why this was decided
| Rule | ENTITY_COLLISION · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:f3f323b98a898391d2f1 · finding f1:94cc84cf5ff4efe5d75b resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 298,064 | 9a8581d03483 | kept |
| anon | 200 | 298,064 | 672b7a289653 | kept |
| gptbot | 200 | 298,064 | 672b7a289653 | kept |
| claudebot | 200 | 298,064 | 672b7a289653 | kept |
| oaisearch | 200 | 298,064 | 672b7a289653 | kept |
| claudesearch | 200 | 298,064 | 672b7a289653 | kept |
| perplexity | 200 | 298,064 | 9a8581d03483 | kept |
| googlebot | 200 | 298,064 | 9a8581d03483 | kept |
14 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
ENTITY_COLLISION · full transcript
3The homepage depends on files this site's own robots.txt disallows8 render resources referenced by the homepage are disallowed to every crawler by this site's own robots.txt: /wp-content/plugins/dfm-trust-indicators/dist/css/sMediumSetting
first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
8 render resources referenced by the homepage are disallowed to every crawler by this site's own robots.txt: /wp-content/plugins/dfm-trust-indicators/dist/css/style.min.css (/wp-content/plugins/), /wp-content/plugins/mng-digisubs/static/mng-digisubs.styles.css (/wp-content/plugins/), /wp-content/plugins/loader-wp/static/loader.min.js (/wp-content/plugins/), /wp-content/plugins/mng-digisubs/static/mng-digisubs.sophi.bundle.js (/wp-content/plugins/). A crawler that obeys robots.txt will not fetch these, so it renders the page without them. /robots.txt
Why it matters
Googlebot renders pages, and a stylesheet or script it is forbidden to fetch is rendered without. Layout, visibility and any content a script inserts are then judged from a page the site never intended to show. Only rules in the * group count here - a rule aimed at one named crawler is a narrower question.
How to fix it
Allow crawlers to fetch the CSS and JavaScript the page depends on. Remove the Disallow rules covering those paths, or serve the content without them.
Setting: a setting on the server, CDN, DNS or a plugin — no code and no writing
Why this was decided
| Rule | RENDER_RESOURCE_BLOCKED · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:b51d1319719ba988cbd3 · finding f1:2262efb943db36f97ff3 resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 4,425 | c18df974780a | already_held |
21 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
RENDER_RESOURCE_BLOCKED · full transcript
4Almost nothing a machine receives from this page is readable text4.4% of 297,605 decompressed bytes is visible text, and the largest single block is inline css at 44,011 bytesMediumDeveloper
first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
4.4% of 297,605 decompressed bytes is visible text, and the largest single block is inline css at 44,011 bytes
Why it matters
A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.
How to fix it
Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.
Developer: a template, theme or application change — a developer touches it
Why this was decided
| Rule | PAGE_IS_MOSTLY_CODE · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:304cde5e6ab450ac41f5 · finding f1:f4babb4a5f2416a1db5f resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 298,064 | 9a8581d03483 | kept |
| anon | 200 | 298,064 | 672b7a289653 | kept |
| gptbot | 200 | 298,064 | 672b7a289653 | kept |
| claudebot | 200 | 298,064 | 672b7a289653 | kept |
| oaisearch | 200 | 298,064 | 672b7a289653 | kept |
| claudesearch | 200 | 298,064 | 672b7a289653 | kept |
| perplexity | 200 | 298,064 | 9a8581d03483 | kept |
| googlebot | 200 | 298,064 | 9a8581d03483 | kept |
14 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
PAGE_IS_MOSTLY_CODE · full transcript
5The published coordinates are not in the country the address namesThis page declares latitude 39.7392, longitude 104.9903 beside an address in US (read from addressRegion). negating the longitude puts it inside US - a missing Low
first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
This page declares latitude 39.7392, longitude 104.9903 beside an address in US (read from addressRegion). negating the longitude puts it inside US - a missing minus sign.
Why it matters
Two facts in the same record disagree about where the business is. Anything that trusts the coordinate over the address is sent to the wrong place, and a range check cannot catch it because the value is a legal longitude. Nobody proofreads a number that never renders on the page.
How to fix it
Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.
Why this was decided
| Rule | GEO_COUNTRY_MISMATCH · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:2d023163a3d6e24168d5 · finding f1:b0dfa22e680744be8515 resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 298,064 | 9a8581d03483 | kept |
| anon | 200 | 298,064 | 672b7a289653 | kept |
| gptbot | 200 | 298,064 | 672b7a289653 | kept |
| claudebot | 200 | 298,064 | 672b7a289653 | kept |
| oaisearch | 200 | 298,064 | 672b7a289653 | kept |
| claudesearch | 200 | 298,064 | 672b7a289653 | kept |
| perplexity | 200 | 298,064 | 9a8581d03483 | kept |
| googlebot | 200 | 298,064 | 9a8581d03483 | kept |
14 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
GEO_COUNTRY_MISMATCH · full transcript
6No llms.txtNot present.NoteContent
first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
Not present. /llms.txt
Why it matters
Not a defect. llms.txt is an emerging convention for telling AI systems what a site is and which pages matter.
How to fix it
Publish /llms.txt: a short Markdown file naming the business, what it does, and the pages worth reading, served as text/plain or text/markdown.
Content: words, markup or images on a page — no deploy
Why this was decided
| Rule | NO_LLMS_TXT · rule pack sv26 · scope core |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:7c3893305bf52fff4f3f · finding f1:02b589098746b639fa40 resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 404 | 88,151 | ea92252b1c26 | kept |
21 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
NO_LLMS_TXT · full transcript
7The machine files cost more than the pageThe machine files a crawler reads before the page total 888,560 bytes, 2.98x the homepage HTML (298,064 B); the heaviest are /sitemap.xml 884,135 B, /robots.txtNote
first measured 2026-09-16 · 8 days ago · seen across 5 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
The machine files a crawler reads before the page total 888,560 bytes, 2.98x the homepage HTML (298,064 B); the heaviest are /sitemap.xml 884,135 B, /robots.txt 4,425 B. An absent agent file answers 404 with 88,202 bytes of HTML on average (/agents.md, /.well-known/agent-skills/index.json, /.well-known/mcp/server-card.json, /.well-known/api-catalog, /.well-known/media-kit.json) - a crawler probing for a file that is not there pays for a page each time; a short 404 body costs nothing to serve.
Why it matters
A crawler reads robots.txt, the sitemap, llms.txt and the agent files before the first page, and pays for every 404 that answers with a full page. Keep machine files as small as their job allows and serve absent paths with a short 404 body.
How to fix it
Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.
Why this was decided
| Rule | MACHINE_CHAIN_HEAVY · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:de8bab803db4c75a3403 · finding f1:5f60febd53b8621e833e resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 298,064 | 9a8581d03483 | kept |
| anon | 200 | 298,064 | 672b7a289653 | kept |
| gptbot | 200 | 298,064 | 672b7a289653 | kept |
| claudebot | 200 | 298,064 | 672b7a289653 | kept |
| oaisearch | 200 | 298,064 | 672b7a289653 | kept |
| claudesearch | 200 | 298,064 | 672b7a289653 | kept |
| perplexity | 200 | 298,064 | 9a8581d03483 | kept |
| googlebot | 200 | 298,064 | 9a8581d03483 | kept |
14 more in the manifest.
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
MACHINE_CHAIN_HEAVY · full transcript
84 off-site listings carry this phone number and the site never links themSearched YellowPages, DexKnows, BBB by (303) 954-1010. A listing the site does not declare is one an engine corroborates without you: whatever it prints becomesNote
first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.
What we saw
Searched YellowPages, DexKnows, BBB by (303) 954-1010. A listing the site does not declare is one an engine corroborates without you: whatever it prints becomes your name, hours and category in the answer.
Why it matters
Declare the ones you own in sameAs so the entity graph and the directory agree, and correct or claim the ones you do not.
How to fix it
Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.
Why this was decided
| Rule | NAP_UNDECLARED_LISTING · rule pack sv26 · scope route |
|---|---|
| Evaluated | 2026-09-24 19:13:40 UTC |
| Decision | d1:31cf56f1c3ab24725d3f · finding f1:a3b71814478a0c1117bf resolve either at /api/graph/<id> with a key |
| Observer | crawlcheck-worker build fb4f0b7e from ATL · audit v4 |
What it rests on
| Identity | Status | Bytes | sha256 | Kept |
|---|---|---|---|---|
| crawlcheck | 200 | 298,064 | 9a8581d03483 | kept |
| anon | 200 | 298,064 | 672b7a289653 | kept |
| gptbot | 200 | 298,064 | 672b7a289653 | kept |
| claudebot | 200 | 298,064 | 672b7a289653 | kept |
| oaisearch | 200 | 298,064 | 672b7a289653 | kept |
| claudesearch | 200 | 298,064 | 672b7a289653 | kept |
| perplexity | 200 | 298,064 | 9a8581d03483 | kept |
| googlebot | 200 | 298,064 | 9a8581d03483 | kept |
14 more in the manifest.
Where each value was read
| Field | Value | Read from |
|---|---|---|
| name | The Denver Post | JSON-LD LocalBusiness |
| phone | +13039541010 | JSON-LD LocalBusiness |
| address | 5990 Washington St., Denver, CO, 80216 | JSON-LD LocalBusiness |
| hours | Mo 8AM-4PM; Tu 8AM-4PM; We 8AM-4PM; Th 8AM-4PM; Fr 8AM-4PM | JSON-LD LocalBusiness |
Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89
NAP_UNDECLARED_LISTING · full transcript
Headroom +20 points if every row passes
Nothing here is broken. Each row is a measure that currently fails its optimal range, and what fixing it is worth to the score.
Watch this domain
Twice-daily scans, the date each finding first appeared, and a change receipt when something moves. Free for 30 days, no card.
Start watching denverpost.comGet this fixed for you
Every finding above fixed on your site, then re-measured — done is measured, not asserted. One-time, $749 for one local-business site.
Get this fixed — $749More detail — headroom arithmetic, priorities and clusters
Optimisation headroom
80 now → 100 with the 24 failing measure(s) fixed — a gain of 20 points. Each number below is the weighted contribution of that one row to the overall score; it is arithmetic on measurements already taken, not a forecast. The LETTER is separately capped at C by finding severity, so clearing that finding lifts the grade independently of the score. 128 measure(s) could not be measured; they are excluded, not counted against the site, and are not headroom.
| Fix this | Section | Currently | Points |
|---|---|---|---|
| Answer engines allowed | AI and search agents | 8 of 17 | +1.8 |
| Search indexes allowed | AI and search agents | 16 of 24 | +1.8 |
| Cost of a miss | Machine layer | 3 missing file(s) return >20KB of HTML | +1.3 |
| entitymap.json | Machine layer | absent (404) | +1.3 |
| Lead paragraph defines the subject | Quotable content | does not name it, no defining verb | +1.2 |
| Content ratio | Payload and render path | 4.4% | +0.9 |
| Deferred vs blocking scripts | Payload and render path | 2 deferred / 14 blocking | +0.9 |
| Inline JavaScript | Payload and render path | 38,121B | +0.9 |
| Render-blocking resources | Payload and render path | 24 | +0.9 |
| Stated robots policy matches actual behaviour | What each agent receives | OAI-SearchBot: disallowed but served, Claude-SearchBot: disallowed but served, PerplexityBot: disallowed but served, Applebot: disallowed but served | +0.8 |
| Cumulative Layout Shift (p75) | Speed a crawler sees, and field data where it exists | 0.13 | +0.7 |
| Largest Contentful Paint (p75) | Speed a crawler sees, and field data where it exists | 3.45s | +0.7 |
| A named person is declared | Authorship and trust signals (E-E-A-T proxies) | none | +0.6 |
| About page linked from the homepage | Authorship and trust signals (E-E-A-T proxies) | no | +0.6 |
| Organization declares who runs it | Authorship and trust signals (E-E-A-T proxies) | 0 of 2 | +0.6 |
| Credential edges | Schema and entity graph | 0 | +0.5 |
| Edge / CDN | Server and hosting | none detected | +0.5 |
| H1 count | Schema and entity graph | 2 | +0.5 |
| Stable @id coverage | Schema and entity graph | 33.3% | +0.5 |
| Width and height declared | Image signals | 0 of 41 | +0.5 |
| sameAs pointing at a place | Schema and entity graph | 1 | +0.5 |
| Form controls have a name | Accessibility structure (from the delivered HTML) | 1 of 2 (missing: checkbox) | +0.4 |
| Found off-site, not declared in sameAs | Local presence, citations and NAP | 4 — which ones, with a licence | +0.4 |
| Trailing slash | Naming and case consistency | 1 target linked both with and without a slash (/contact-us) | +0.4 |
What to fix first
Ordered by contextual priority, not by severity alone. Two findings of the same severity are not equally urgent — one that breaks a file every agent reads before anything else outranks one on a single page. Every multiplier below is derived from this scan and shown with the reason it applied.
5.8
Two business records on this page claim the same identity
/ · severity 3
| Prerequisite impact | ×1.1 | affects the homepage |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.75 | open for 39 days, first seen 2026-08-20 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
5.8
Almost nothing a machine receives from this page is readable text
/ · severity 3
| Prerequisite impact | ×1.1 | affects the homepage |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.75 | open for 39 days, first seen 2026-08-20 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
5.6
robots.txt rules do not apply to named agents
/robots.txt · severity 3
| Prerequisite impact | ×1.5 | /robots.txt is a discovery prerequisite - agents read it before anything else |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.25 | confirmed on 12 scans of this domain since 2026-09-10 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
5.6
The homepage depends on files this site's own robots.txt disallows
/robots.txt · severity 3
| Prerequisite impact | ×1.5 | /robots.txt is a discovery prerequisite - agents read it before anything else |
| Scope | ×1 | one crawler identity affected |
| Persistence | ×1.25 | confirmed on 12 scans of this domain since 2026-09-10 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
2.8
The published coordinates are not in the country the address names
/ · severity 2
| Prerequisite impact | ×1.1 | affects the homepage |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.25 | confirmed on 10 scans of this domain since 2026-09-10 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
2.6
No llms.txt
/llms.txt · severity 1
| Prerequisite impact | ×1.5 | /llms.txt is a discovery prerequisite - agents read it before anything else |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.75 | open for 39 days, first seen 2026-08-20 |
| Confidence | ×1 | derived from the response itself |
1.4
The machine files cost more than the page
/ · severity 1
| Prerequisite impact | ×1.1 | affects the homepage |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.25 | confirmed on 8 scans of this domain since 2026-09-16 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
1.4
4 off-site listings carry this phone number and the site never links them
/ · severity 1
| Prerequisite impact | ×1.1 | affects the homepage |
| Scope | ×1 | not an access-class finding - scope not multiplied |
| Persistence | ×1.25 | confirmed on 11 scans of this domain since 2026-09-10 |
| Confidence | ×1 | the bytes behind this finding are attached to it |
Priority = severity × prerequisite impact × scope × persistence × confidence. A factor the record cannot support stays at ×1 rather than being guessed upward.
These findings share a cause
Findings are symptoms. Where two or more point at one underlying cause, they are grouped here — fixing the cause closes all of them. A single finding is never presented as a cluster, because one symptom does not establish a diagnosis.
The site's own robots.txt withholds what its pages need to render
ROBOTS_RULES_SHADOWED RENDER_RESOURCE_BLOCKED · highest priority in group 5.6
A stylesheet or script the homepage depends on sits under a path the * group disallows. A rendering crawler obeys the rule and draws the page without it.
What to check: Allow the asset paths the page references, or move the assets out of the disallowed directory. Only the * group is in scope; a rule aimed at one named crawler is a separate question.
Every finding also appears on its own below, with its full evidence. Grouping changes the order of work, not the record.
What each crawler was served
14 requests for the same URL, from one CrawlCheck address, at 2026-09-24 19:13 UTC: each named crawler, a mobile browser and an unnamed client. Robots.txt is what the site asks for; the status and word count are what its server actually sent.
| Requested as | Robots.txt | Server sent | Words | Same as browser | |
|---|---|---|---|---|---|
| Unnamed client (control) | — | 200 | 2,036 | 100% | baseline |
| GPTBot | disallowed | 200 | 2,036 | 100% | disallowed but served |
| ClaudeBot | disallowed | 200 | 2,036 | 100% | disallowed but served |
| OAI-SearchBot | disallowed | 200 | 2,036 | 100% | disallowed but served |
| Claude-SearchBot | disallowed | 200 | 2,036 | 100% | disallowed but served |
| PerplexityBot | disallowed | 200 | 2,036 | 91% | disallowed but served |
| Googlebot | allowed | 200 | 2,036 | 91% | same page |
| Mobile browser (control) | — | 200 | 2,036 | 100% | baseline |
| Bingbot | allowed | 200 | 2,036 | 91% | same page |
| Applebot | disallowed | 200 | 2,036 | 91% | disallowed but served |
| Amazonbot | disallowed | 200 | 2,036 | 100% | disallowed but served |
| Bytespider | disallowed | 403 | 5 | — | blocked by the server |
| Meta-ExternalAgent | disallowed | 200 | 2,036 | 91% | disallowed but served |
| CCBot | disallowed | 200 | 2,036 | 91% | disallowed but served |
13
passing
7
need work
3
failing
7
information only
Sections, not individual checks: 30 in all, 23 scored. A scored section that could not be read is counted separately and never as a pass; information-only sections are measured and never scored.
Which crawlers are let in
73 of 114 named crawlers are allowed by robots.txt at /; 41 are disallowed. Robots.txt is a request; the edge decides. Where this scan also fetched the page as that crawler, the chip says what the edge did.
Answer engines 8/17 allowed
The crawlers behind ChatGPT, Claude, Perplexity and Google’s AI answers
Search engines 16/24 allowed
Google, Bing, Apple and the rest of classic search
Training crawlers 22/46 allowed
Collect pages to train models. Blocking these while allowing answer engines is coherent policy, not a defect
Show all 46
Some named groups declare no rules of their own, which means “allow everything” for that crawler regardless of the general rules. Those chips are marked.
How crawlers move through this site
4 of 14 identities reach the page. 9 are turned away by the site’s own robots rules — a policy, not a fault. 1 is refused at the edge before any rule applies, for a request carrying the name from our address — an edge that admits crawlers by their published IP ranges may still let the real one through. Missing: entitymap.json, llms.txt. Of the 12 others that got through, 1 received exactly the browser’s bytes, 5 the same text in different bytes, 6 different text.
Who asked
- Unnamed clientreaches the page · same text, different bytes
- GPTBotturned away by the site's robots rules · same text, different bytes
- ClaudeBotturned away by the site's robots rules · same text, different bytes
- OAI-SearchBotturned away by the site's robots rules · same text, different bytes
- Claude-SearchBotturned away by the site's robots rules · same text, different bytes
- PerplexityBotturned away by the site's robots rules · different text (91% of sentences shared)
- Googlebotreaches the page · different text (91% of sentences shared)
- Mobile browserreaches the page
- Bingbotreaches the page · different text (91% of sentences shared)
- Applebotturned away by the site's robots rules · different text (91% of sentences shared)
- Amazonbotturned away by the site's robots rules · identical bytes
- Bytespiderrefused at the edge (HTTP 403) to our request
- Meta-ExternalAgentturned away by the site's robots rules · different text (91% of sentences shared)
- CCBotturned away by the site's robots rules · different text (91% of sentences shared)
Files fetched
- robots.txtpresent · 4.3 KB
- sitemappresent · 863.4 KB
- llms.txtmissing
- entitymap.jsonmissing
- Page HTMLpresent · 291.1 KB
Derived
- JSON-LDpresent
- Quotable textpresent
- Enquiry pathpresent
- Declared identity (NAP)present
Answer engine output — not measured. No server can observe what an engine says.
Evidence
- Bytes kept22 of 22 responses
- Manifesteb69e87d5daf…
- Record digestfixed at the seal
- Day rootseals after 09-24
- Bitcoinafter the seal
✓ reached · ⊘ turned away by the site’s own rules · ✗ technical break · dashed box = missing or not measured. Drawn from this scan’s own fetches; nothing here is estimated. Fetched 2026-09-24 19:13 UTC from Cloudflare ATL by worker version fb4f0b7e under scoring rules v26. Each identity was sent by CrawlCheck from its own network: ✓ means our request under that name got through, not that the company’s own crawler did. The flow as JSON.
What each identity received — 22 responses captured, 22 kept as evidence
| Who asked | HTTP | Size | Compared with a browser | Kept as evidence |
|---|---|---|---|---|
| Unnamed client | 200 | 291.1 KB | same text, different bytes | kept · 672b7a28 |
| GPTBot | 200 | 291.1 KB | same text, different bytes | kept · 672b7a28 |
| ClaudeBot | 200 | 291.1 KB | same text, different bytes | kept · 672b7a28 |
| OAI-SearchBot | 200 | 291.1 KB | same text, different bytes | kept · 672b7a28 |
| Claude-SearchBot | 200 | 291.1 KB | same text, different bytes | kept · 672b7a28 |
| PerplexityBot | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| Googlebot | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| Mobile browser | 200 | 291.1 KB | the baseline | kept · e1063e72 |
| Mobile browser (control) | 200 | 291.1 KB | identical bytes | kept · e1063e72 |
| Bingbot | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| Applebot | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| Amazonbot | 200 | 291.1 KB | identical bytes | kept · e1063e72 |
| Bytespider | 403 | 146 B | refused (HTTP 403) | kept · 32f2fa94 |
| Meta-ExternalAgent | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| CCBot | 200 | 291.1 KB | different text (91% of sentences shared) | kept · 9a8581d0 |
| CrawlCheck scanner (home page) | 200 | 291.1 KB | — | kept · 9a8581d0 |
| Files the scanner fetched | ||||
| /robots.txt | 200 | 4.3 KB | — | kept · c18df974 |
| /sitemap.xml | 200 | 863.4 KB | — | kept · e1a6b687 |
| /sitemap_index.xml | 404 | 86.1 KB | — | kept · bd47eaac |
| /wp-sitemap.xml | 200 | 138.8 KB | — | kept · f2ff03d4 |
| /llms.txt | 404 | 86.1 KB | — | kept · ea92252b |
| /entitymap.json | 404 | 86.1 KB | — | kept · 44962592 |
Sizes are each response body as the fetch returned it: transfer compression removed, before any text decoding. A “kept” link re-hashes the stored copy inside our worker and says whether it still matches its digest; the bytes themselves are not republished.
Discovery chain breaks at llms.txt: everything after it is reachable only by guessing the conventional path. Rows behind this.
Where the bytes go
Of the 297,605 decompressed bytes the homepage delivers, 4.4% is text a reader or a model can actually use. Most of what a crawler downloads here is not words.
- Readable text 4.4% · 13,081 B
- Markup & attributes 65.7% · 195,658 B
- Structured data (JSON-LD) 0.8% · 2,238 B
- Inline CSS 15.3% · 45,571 B
- Inline JavaScript 12.8% · 38,121 B
- Inline SVG 0.1% · 372 B
- HTML comments 0.9% · 2,564 B
The individual blocks that weigh the most, so the fix is a search rather than a hunt:
- Inline CSS · 44,011 B (14.8%) — unnamedopens
@import url(https://fonts.googleapis.com/css2?family=Inter:ital,opsz,w - Inline JavaScript · 14,462 B (4.9%) —
#digisubs-social-login-js-afteropens(function () { var logPrefix = '[digisubs-social-login]'; var readyTim - Inline JavaScript · 3,089 B (1%) — WordPress emoji scriptopens
/*! This file is auto-generated */ var e="script#wp-emoji-settings",t=
Every section
23 of 30 carry weight in the grade. Grey means measured and deliberately unweighted - it counts for nothing rather than zero, and each grey card says why. Open any section for the rows behind it: what was measured, what it is now, and the range that passes.
What the site states about itself
The facts an answer engine could repeat about this business, taken only from what this scan fetched, with where each was found. Anything absent is [not stated]: an engine asked for it would have to guess.
| Fact | Stated as | Found in | Agreement | First seen |
|---|---|---|---|---|
| Name | The Denver Post | 2026-09-19 | ||
| Phone | (303) 954-1010 | structured data only | not shown to readers | 2026-09-19 |
| Address | 5990 Washington St., Denver, CO, 80216 | 2026-09-19 | ||
| Map point | 39.7392, 104.9903 | structured data | negating the longitude puts it inside US - a missing minus sign | 2026-09-19 |
| Opening hours | stated | structured data | 2026-09-19 | |
| Business type | LocalBusiness | structured data | 2026-09-19 | |
| Profiles it links | 3 (Facebook, X, LinkedIn) | links on the site | 3 of 3 point back |
Everything else that was measured
Measured for information, not scored
Section detail
Each row: the measure, what it is now, the range that passes, and whether this site is inside it. n/a means the input could not be measured and the row is excluded, not failed.
Machine layer — 60/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| robots.txt a crawler that cannot read robots.txt often treats that as disallow-everything | 200 · 4425B | 200, text/plain, 100–4,000B | ✓ |
| XML sitemap without it a crawler enumerates the site by following links only | live | 200, valid XML, at a path robots.txt names | ✓ |
| llms.txt reported, not scored. Some agents fetch it and it costs nothing to publish, so the row stays. It is not graded because nothing we have measured, and nothing we have read, links its presence to a citation or a ranking — a study across roughly 300,000 domains found no correlation, and dropping the variable made the model more accurate, not less. Grading a file we cannot show works would price it for our customers on our say-so | absent (404) | 200, text/plain, 2,000–50,000B | n/a |
| entitymap.json machine-readable entity layer; optional, but it is the strongest identity signal available | absent (404) | 200, application/json | ✗ |
| How rare is what this site publishes adoption counts are third-party aggregates read in August 2026 and count sites in that index rather than the whole web; they are here to size the opportunity, not to score the site | llms.txt — roughly 70,200 sites publish one; this is not among them. entitymap.json — absent, and no third-party adoption figure exists for it. for scale, roughly 677,748 sites publish something built for agents at all, and about 107,200 now disallow GPTBot outright, so an AI policy is mainstream while a machine-readable map of the site is not. | publish what almost nobody publishes, and keep it correct | n/a |
| Host agreement a crawler resolving www must not get a different machine layer | hosts agree | 0 divergent files | ✓ |
| Cost of a miss every absent file bills the crawler for a full HTML 404 | 3 missing file(s) return >20KB of HTML | a missing machine file should 404 small, not ship a full page | ✗ |
Hosts and URLs
/robots.txt | www.denverpost.com: 200 · 4,425B live denverpost.com: 301 → https://www.denverpost.com/robots.txt |
/llms.txt | www.denverpost.com: 404 · 88,151B of HTML (not the file) denverpost.com: 301 → https://www.denverpost.com/llms.txt |
/entitymap.json | www.denverpost.com: 404 · 88,169B of HTML (not the file) denverpost.com: 301 → https://www.denverpost.com/entitymap.json |
/sitemap.xml | www.denverpost.com: 200 · 884,135B live denverpost.com: 200 · 884,135B live |
/sitemap_index.xml | www.denverpost.com: 404 · 88,178B of HTML (not the file) denverpost.com: 301 → https://www.denverpost.com/sitemap_index.xml |
Scanned on www.denverpost.com. denverpost.com 301s to www.denverpost.com, so that is the host a crawler ends up reading. Probed with redirects not followed, so a 200 here means the file is served at that exact URL — a followed redirect would report a file that lives somewhere else. The denverpost.com host serves its own machine layer. A crawler that resolves that hostname reads those files, not the ones on www.denverpost.com.
Server and hosting — 80/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Compression our fetch decompresses the body and drops the content-encoding header before we can read it — this is our limitation, not a finding about the site. Check it in a browser’s network panel, or with curl --compressed -I | not visible from here | br preferred, gzip acceptable | n/a |
| Edge / CDN absorbs crawler load and keeps machine files fast | none detected | a CDN in front of the origin | ✗ |
| Cache state DYNAMIC means every crawler hit reaches the origin. Shown, not scored: the speed rows below score what that costs, and scoring the mechanism too would count it twice | 4 of 5 machine files DYNAMIC | HIT or MISS on machine files; DYNAMIC is Cloudflare's default for .txt and .xml | n/a |
| Version disclosure a public CMS version tells an attacker exactly which exploits to try | not disclosed | no generator version in the HTML | ✓ |
| Runtime disclosure same reason as above | not disclosed | no x-powered-by header | ✓ |
| Vary governs whether a cached copy is reused correctly | accept-encoding | Accept-Encoding at minimum | ✓ |
| Strict-Transport-Security max-age 365 days, subdomains included · hstspreload.org: unknown. Shown, not scored: HSTS changes nothing about what crawlers receive | max-age=31536000;includeSubdomains | max-age with includeSubDomains; preload only if you mean it | n/a |
| Webmaster verification tags meta tags only — DNS and file verification leave nothing in the HTML, so a missing tag is not evidence of no registration. Shown, not scored | Google no tag visible · Bing tag present · Yandex no tag visible | Bing at least: its index feeds ChatGPT search and Copilot | n/a |
| Client-side rendering framework present, HTML delivered with text. Measured on the bytes the server sent, not on a rendered view | script bundle, no known framework · 2036 words delivered | text present in the delivered HTML | ✓ |
| Application | WordPress detected from the delivered HTML, not from a header that can be removed |
| Varies on | accept-encoding a response that varies can be cached differently per client |
Schema and entity graph — 60/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Entity nodes with no schema the site has no declared identity at all | 3 | at least 3 typed nodes | ✓ |
| Stable @id coverage without an @id every page declares a brand-new unrelated entity | 33.3% | 80–100% | ✗ |
| Credential edges the machine-readable form of why this business is qualified; a publisher, a shop or a SaaS is not held to it | 0 | at least 1 (licence, membership or award) for a local or professional service | ✗ |
| Locations missing a street a service-area business legitimately omits this — not scored either way | 0 | 0 for a premises business; expected for service-area | ✓ |
| sameAs claims sameAs asserts identity, so a WRONG target claims the business is something else. A large number of correct profiles is not a defect; a single target that does not lead back here is. | 3 (3 corroborated, 0 not linking back, 0 dead) | at least 2, none dead, at least one linking back | ✓ |
| sameAs pointing at a place a map link in sameAs claims the business IS that location | 1 | 0 | ✗ |
| H1 count more than one h1 leaves no single subject for the page | 2 | exactly 1 | ✗ |
| Heading skips a skipped level breaks the document outline a parser builds | 0 | 0 | ✓ |
| Canonical without it duplicates compete with each other | present | present on every page | ✓ |
| Meta description over 160 it is truncated; under 50 it is a label, not a description. 95 chars is fine | 157 chars | 50–160 chars | ✓ |
| Entities declared | 3 nodes, 2 distinct types |
| Stable identity | 1 of 3 carry an @id (33.3%) Without an @id a node cannot be referenced or merged — every page re-declares a new, unrelated entity. |
| Credential edges | 0 Licences, memberships and awards expressed as graph edges rather than prose. |
| Topic edges | 0 A high count with generic values usually means the graph is being used to store keywords. |
| Locations | 1 declared, 0 without a street address |
| sameAs claims | 3 total · 1 map links · 0 resolve to this entity · 1 resolve to a place sameAs asserts identity. A map link pointing at a city claims the business is that city. No validator reports this — it only appears if the targets are resolved. |
| Document structure | 2 h1 · 120 headings · 0 level skips · canonical present · meta description 157 chars |
RDF and linked-data surface — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| JSON-LD blocks that parse a block that does not parse contributes no triples and looks, to a validator that counts tags, like structured data that is there | 2 of 2 | every block | ✓ |
| @context is schema.org a graph in another vocabulary is valid RDF that no answer-engine resolver is built to read | 2 of 2 | every block | ✓ |
| Triples declared (approximate) the size of what a linked-data reader actually receives; a single Organization with a name and a URL is three statements | 68 | at least 20 on a homepage | ✓ |
| Dangling @id references a reference to an @id no node declares is a statement about nothing; the resolver drops the edge and the entity loses that relation | 0 | 0 | ✓ |
| Top-level nodes carrying a subject IRI a root or @graph node without an @id is a blank subject that no other page or site can reference; nested values (an address, an offer, a list item) are not held to this | 0 of 2 (12 nodes in all) | every entity the page declares — shown; the schema section scores @id coverage | n/a |
| RDFa and Microdata Secondary serialisations are read by some parsers and are a second place for the facts to disagree. Not scored while JSON-LD is present; when a page declares no JSON-LD node, its schema.org Microdata and RDFa are read as its graph and every entity check runs on them | no RDFa, no Microdata, 23 Open Graph meta | optional; one serialisation kept accurate beats three that drift | n/a |
| RDF content negotiation An RDF-aware client sent Accept: text/turtle. Almost no commercial site answers with a graph and absence is not a defect; a site that does is handing the machine the same facts in a tenth of the bytes. Not scored | text/html served (HTTP 200), Vary: Accept | optional: text/turtle or application/ld+json to a client that asks for it | n/a |
| Alternate RDF representation linked Tells a linked-data client where the graph lives without negotiating for it. Not scored | none | optional: <link rel=alternate type=text/turtle> or a Link header | n/a |
JSON-LD is an RDF serialisation, so every schema block above is already a graph of triples. This section reads it as one: 68 statements about 0 identified subjects, 1 reference between them, none dangling. The triple count is approximate: each property value is one statement and each @type one more; nested objects without an @id are counted as values, not expanded. Weight 0.8 in the overall, below the schema section it overlaps.
Authorship and trust signals (E-E-A-T proxies) — 57/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| A named person is declared an evaluator asking who is behind this finds a machine-readable answer or nothing; a site with no Person node is anonymous to a resolver | none | at least 1 Person node with a name | ✗ |
| Person carries expertise properties jobTitle, hasCredential, alumniOf, memberOf, knowsAbout or award: the statements that say WHY this person is qualified, in the form a reader can check | — | every named person | n/a |
| Person is corroborated elsewhere (sameAs) a person who exists only on this site is a claim; a person with profiles that link back is a fact someone else asserts too | — | every named person | n/a |
| Person has a page of their own the page an evaluator lands on to judge experience; without it the person is a name in a script tag | — | a url or mainEntityOfPage on each | n/a |
| Articles name a typed author author as a string is a label; author as a Person or Organization object is an entity that can carry the expertise above | no Article nodes on this page | every article | n/a |
| Articles carry datePublished and dateModified freshness is read from dateModified; an article with only a publish date reads as never maintained | — | both on every article | n/a |
| Organization declares who runs it connects the business entity to the people entities; without the edge the two graphs never meet | 0 of 2 | founder, employee or member on the organisation | ✗ |
| Organization declares foundingDate the one claim behind every 'years of experience' counter that a reader can check against a registry | 1 of 2 | on the primary organisation | ✓ |
| Contact route declared in schema a business that cannot be reached by a stated channel is unverifiable by definition | 1 of 2 | telephone, email or contactPoint | ✓ |
| About page linked from the homepage the page an evaluator opens first; if the homepage does not link it, most readers never find it | no | linked | ✗ |
| Contact page or direct contact linked a contact path in the delivered HTML, not only in a footer image or a widget that renders later | yes | linked | ✓ |
| Privacy policy or terms linked the cheapest trust signal on the list and the one small sites most often omit | yes | linked | ✓ |
| Ratings carry a value and a count a rating with no count is a number nobody can weigh; a count with no reviews behind it is the shape of a fabricated one | no aggregateRating | ratingValue with reviewCount or ratingCount | n/a |
| Outbound references Citations outward are part of how authority is read, and a count cannot say whether they are good ones. Not scored | 18 external hosts linked (excluding social platforms): checkout.denverpost.com, accuweather.com, enewspaper.denverpost.com | cite what the content relies on; there is no right number | n/a |
What this can and cannot say. E-E-A-T is a judgement an evaluator makes about experience, expertise, authoritativeness and trust; it is not on the wire and this scanner does not pretend to score it. Each row above is one machine-readable statement a site makes that a reader could use as evidence toward that judgement, measured from the delivered homepage and its JSON-LD. A full row of passes does not create expertise, and a row of fails does not disprove it; what they decide is whether a machine reading this site can find the evidence at all. Weight 1.0 in the overall, the same as payload: what these rows measure is whether the evidence is findable by a machine, and that is the part this scanner can stand behind.
Accessibility structure (from the delivered HTML) — 88/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Document language declared a screen reader picks its voice and pronunciation from this; without it every word is read in the user's default language | lang=en-US | a lang attribute on <html> | ✓ |
| Page has a title the first thing announced when a page opens, and the label on every tab and bookmark | 89 chars | a non-empty <title> | ✓ |
| Images carry an alt attribute a missing alt attribute makes the reader announce the file name; alt="" marks an image as decorative and is correct | 41 of 41 | every image (alt="" for decoration) | ✓ |
| Form controls have a name a field with no name is announced as "edit text"; a placeholder is not a label, it disappears on the first keystroke | 1 of 2 (missing: checkbox) | a <label for>, wrapping label, aria-label or title on each | ✗ |
| Buttons have an accessible name an icon-only button with no name is announced as "button" and nothing else | 3 of 3 | text, aria-label, or an image with alt inside each | ✓ |
| Frames are titled a frame with no title is a hole in the reading order the user cannot identify or skip on purpose | no frames | a title on every iframe | n/a |
| One main landmark the region a reader jumps to for the content; two of them is as unusable as none | one | exactly one <main> or role=main | ✓ |
| A way past the navigation without one, every page starts by tabbing through the whole menu | skip link | a skip link, or nav and main landmarks | ✓ |
| No duplicate ids labels, aria-labelledby and skip links all resolve by id; a duplicate sends them to the first match, which is usually the wrong element | 0 | every id unique | ✓ |
What this is not. Not a WCAG audit: contrast, focus order, keyboard traps, motion and the experience of using the page with assistive technology need a rendered page and a person, and none of that is measured here. These rows read the document the server returned and ask whether the structure a screen reader depends on is present at all. Heading order, zoom lock and link text are scored in their own sections and are not counted again here. Weight 0.8 in the overall. Counts in the Read stage from scoring version 26, because parsers find content through the same structure a screen reader uses; reports issued earlier kept it outside the stages.
Entity corroboration — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Identity claim — facebook.com the target page references denverpost.com back | corroborated | corroborated | ✓ |
| Identity claim — x.com the target page references denverpost.com back | corroborated | corroborated | ✓ |
| Identity claim — linkedin.com the target page references denverpost.com back | corroborated | corroborated | ✓ |
| Wikidata Q2668654 the strongest third-party identity signal a machine can read | official website matches this domain | P856 names this domain | n/a |
3 corroborated · 0 unreciprocated · 0 unverifiable · 0 dead. A sameAs link is an identity CLAIM, and an identity claim can be reciprocated — the target either points back at this domain or it does not. Platforms that refuse datacentre readers are counted as unverifiable, never as failures: the claim can be neither confirmed nor accused from here, and pretending otherwise would manufacture findings. Corroborated claims pass, unreciprocated and dead claims fail, and unverifiable or declared-only claims are left out of the score — so a profile page that blocks readers can never cost a point, and a link that is dead or never mentions this site does.
Quotable content — 67/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Lead paragraph defines the subject the first paragraph is the one most often lifted whole; if it is a slogan, the engine has to assemble the answer from fragments. Naming the subject is shown, not scored | does not name it, no defining verb | a first sentence that says what this is (corpus: 30% of homepages do) | ✗ |
| Paragraphs an engine can quote whole a self-contained paragraph of 15-70 words with a fact in it is the unit answer engines extract; a 200-word paragraph gets summarised instead, and the summary is theirs | 9 of 12 (75%) | at least 25% (corpus median 29%) | ✓ |
| Question headings with an answer under them the shape a query has when it arrives; a page that already carries the question and a short answer is quoted in that order. Having none is not scored | no question headings | at least half, where question headings exist | n/a |
| FAQ markup matches the visible text FAQ markup for text a visitor cannot see is a claim with nothing behind it; engines that compare the two drop the markup | no FAQPage | every declared question readable on the page | n/a |
| Headings carry an id a heading with an id is a stable address for one passage; without it a citation can only point at the whole page | 0 of 27 | an id on each, so a passage can be cited by fragment (not scored: corpus median 0%) | n/a |
| Sentence length quotes are short; a 40-word sentence is paraphrased, and the paraphrase carries the engine's wording, not yours | median 13 words, 63% under 26 | a median of 22 words or fewer (corpus median 15) | n/a |
| Sentences with a checkable fact 8-30 words with a number, date or name, in the third person: the sentence an engine can attribute to you without editing it | 5 of 8 (63%) | at least 15% (corpus median 20%) | n/a |
| Paragraphs that open in the first person "We offer" needs rewriting before it can be quoted about you; "Acme offers" does not | 0 of 12 (0%) | a third or fewer (corpus p90 24%) | ✓ |
| Tables with a header row and at least two data rows a table is the one structure every extractor reads the same way; the same facts in prose are re-assembled differently by each engine | 0 | one per set of comparable facts (not scored: rare on homepages) | n/a |
| A visible updated or published date engines weigh freshness from the text as well as from schema; a date the reader can see is the one they trust | none found | a dated line in the copy (not scored: 4% of homepages carry one) | n/a |
Sentences an engine could lift as they stand:
- “Here’s what to expect in the final 3 years.”
- “Much of the I-70 Floyd Hill Project is taking place above drivers’ heads.”
- “Here’s what to expect in the final 3 years.”
The lead, as delivered: “Sign up for Newsletters and Alerts Sign Up”
What this scores, and what it is not. These rows read the delivered HTML for the shape of quotable text: nothing here judges whether the copy is good, true or wanted. Counts come from body paragraphs and list items with the navigation, header, footer and forms removed; a page with fewer than five is not scored here. Thresholds were set from a 176-homepage corpus pass and each row states the corpus figure it was set against. Weight 0.8 in the overall and part of the Quote stage, from score version 15.
Crawl waste — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Internal links carrying a query string each parameter form is a separate URL to a crawler; tracking and sort parameters on internal links multiply the pages it thinks you have | 1 of 297 - ntv_adpz | 0, unless the parameter changes the page | ✓ |
| Internal links to http:// every one is a redirect before the page, and a redirect is a fetch that returned nothing | 4 of 297 | 0 | n/a |
| Declared URLs that redirect a sitemap entry that redirects sends the crawler somewhere the sitemap should have named in the first place | 0 of 6 sampled | 0 | ✓ |
| Declared URLs that do not answer 200 a declared page that is gone is a fetch spent on a page that will not be indexed | 0 of 6 sampled | 0 | n/a |
| Declared pages whose canonical points elsewhere declaring a page and then telling the engine it is really another page is two fetches for one result | 11 of 60 | 0 | n/a |
| Declared pages marked noindex a page in the sitemap asking not to be indexed is a contradiction the crawler resolves by fetching it anyway | 0 of 60 | 0 | n/a |
| Duplicate titles across the read pages pages that look the same from the outside get fetched, compared and mostly discarded | 0 of 60 | 0 | n/a |
| Apex and www both answer 200 two hosts serving the same site is every page twice, and the engine has to guess which copy is the real one | no, one redirects | one host, the other redirects | ✓ |
A crawler arrives with a budget. Every row here is a fetch that produced no new page: a redirect, a parameter variant, a case variant, a declared page that points elsewhere. Counted from the delivered homepage, the sitemap, the sampled declared URLs and the daily whole-site read where one exists. None of this moves the grade.
Page subject, the Webref question — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Business entities declared on the page Webref decides which entity a page is about and how much; a page that names several unrelated businesses splits its topicality between them | 1 | one subject, others only by declared relation | n/a |
| The page's subject mainEntity is the explicit declaration; a url on this host is the next best evidence; guessing from the first node is what this check refuses to do | The Denver Post - found by url on this host | declared as WebPage.mainEntity | n/a |
| Title and h1 name the subject the two strongest on-page statements of what the document is about | title yes, h1 yes | both | n/a |
| Places named in the graph areaServed is read as places the business serves; a locality the business is IN is a different relation, and both belong on the record | 2 (Denver, Denver Metropolitan Area and the state of Colorado) | the branch's own locality plus a declared service area | n/a |
| Verdict single: one entity. related: several, all connected. diluted: unrelated businesses share the page. ambiguous: no subject can be resolved | single - one business entity, and it is the subject | single or related | ✓ |
A document does not compete for a keyword alone; the index links it to entities by their Knowledge Graph identity and ranks documents against each other for that entity (the Webref layer in the recovered Geostore schema). This section reads the structured data on the delivered homepage and asks the question that layer asks: which business is this page about, and how is every other business on it related? Chain membership, subsidiary, department and containment are declared relations; a second business with no relation is a claim the engine must arbitrate. Measured, not scored, until the corpus shows how often each verdict occurs.
AI and search agents — 33/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Answer engines allowed these fetch a page while composing a live reply | 8 of 17 | all of them | ✗ |
| Search indexes allowed classic index coverage still drives most discovery | 16 of 24 | all of them | ✗ |
| Training crawlers allowed blocking trainers while allowing answer engines is coherent, not a defect | 22 of 46 | your policy choice — not scored | n/a |
| Regional search engines allowed Naver, Sogou, 360 and Yisou matter if you sell into those markets and are irrelevant if you do not; that is a business fact we cannot read off the page | 7 of 7 | your policy choice — not scored | n/a |
| Social link previews allowed blocking these does not touch answer engines, it just makes your links render as bare URLs when anyone shares them | 6 of 6 | your policy choice — not scored | n/a |
| SEO and research crawlers allowed blocking a link-index crawler costs you competitor visibility, not answer-engine visibility — a different trade from the one above | 14 of 14 | your policy choice — not scored | n/a |
| robots.txt groups a second * group is invisible to parsers that stop at the first match | 68 | 1 group per user-agent, no duplicate * group | ✓ |
| Answer engines — fetch at question time — 8 of 17 allowed | |
| OAI-SearchBot | BLOCKED at / · ChatGPT search index · named rule |
| ChatGPT-User | BLOCKED at / · ChatGPT live fetch · named rule |
| Claude-SearchBot | BLOCKED at / · Claude search index · named rule |
| Claude-User | BLOCKED at / · Claude live fetch · named rule |
| PerplexityBot | BLOCKED at / · Perplexity index · named rule |
| Perplexity-User | BLOCKED at / · Perplexity live fetch · named rule |
| Gemini-Deep-Research | allowed · Gemini research agent · wildcard rule |
| MistralAI-User | BLOCKED at / · Le Chat live fetch · named rule |
| DuckAssistBot | BLOCKED at / · DuckDuckGo AI assist · named rule |
| YouBot | BLOCKED at / · You.com · named rule |
| PhindBot | allowed · Phind · wildcard rule |
| Kagibot | allowed · Kagi · wildcard rule |
| Copilot-User | allowed · Microsoft Copilot fetch · wildcard rule |
| Meta-ExternalFetcher | allowed · Meta AI live fetch · wildcard rule |
| Amzn-User | allowed · Amazon live fetch for Alexa questions · wildcard rule |
| Google-Agent | allowed · Google user-triggered agent, acts on the web · wildcard rule |
| Google-GeminiNotebook | allowed · Gemini Notebook user-supplied URL fetch · wildcard rule |
| Search indexes — 16 of 24 allowed | |
| Googlebot-Image | allowed · Google image crawler · wildcard rule |
| Googlebot-News | allowed · Google News · wildcard rule |
| Googlebot-Video | allowed · Google video crawler · wildcard rule |
| Storebot-Google | allowed · Google Shopping · wildcard rule |
| Mediapartners-Google | allowed · Google AdSense · wildcard rule |
| APIs-Google | allowed · Google push delivery · wildcard rule |
| BingPreview | allowed · Bing page preview · wildcard rule |
| msnbot | allowed · Microsoft legacy crawler · wildcard rule |
| Slurp | BLOCKED at / · Yahoo Slurp · named rule |
| Yandex | BLOCKED at / · Yandex (umbrella token) · named rule |
| coccocbot-web | allowed · Coc Coc · wildcard rule |
| Googlebot | allowed · Google Search and AI Overviews · wildcard rule |
| bingbot | allowed · Bing and Copilot index · wildcard rule |
| Applebot | BLOCKED at / · Apple and Siri · named rule |
| Amazonbot | BLOCKED at / · Amazon · named rule |
| DuckDuckBot | allowed · DuckDuckGo · wildcard rule |
| YandexBot | BLOCKED at / · Yandex · named rule |
| Baiduspider | BLOCKED at / · Baidu · named rule |
| Seznambot | allowed · Seznam · wildcard rule |
| Neevabot | allowed · Neeva · wildcard rule |
| PetalBot | BLOCKED at / · Huawei Petal · named rule |
| Amzn-SearchBot | BLOCKED at / · Amazon search eligibility, separate from Amazonbot · named rule |
| MistralAI-Index | allowed · Mistral search index for Vibe · wildcard rule |
| ExaSearchBot | allowed · Exa AI search index, Web Bot Auth signed · wildcard rule |
| Training / corpus crawlers — 22 of 46 allowed | |
| Magpie-crawler | allowed · Magpie AI · wildcard rule |
| img2dataset | allowed · img2dataset image corpus · wildcard rule |
| AwarioRssBot | allowed · Awario RSS · wildcard rule |
| AwarioSmartBot | allowed · Awario smart · wildcard rule |
| TurnitinBot | allowed · Turnitin · wildcard rule |
| archive.org_bot | BLOCKED at / · Internet Archive · named rule |
| ia_archiver | BLOCKED at / · Internet Archive (legacy) · named rule |
| meta-webindexer | allowed · Meta web index · wildcard rule |
| omgili | BLOCKED at / · Webz.io omgili · named rule |
| cohere-training-data-crawler | BLOCKED at / · Cohere training corpus · named rule |
| PanguBot | BLOCKED at / · Huawei PanGu · named rule |
| Ai2Bot-Dolma | BLOCKED at / · Allen Institute Dolma · named rule |
| FriendlyCrawler | allowed · FriendlyCrawler ML · wildcard rule |
| VelenPublicWebCrawler | BLOCKED at / · Velen · named rule |
| MyCentralAIScraperBot | allowed · MyCentral AI · wildcard rule |
| DeepSeekBot | allowed · DeepSeek · wildcard rule |
| ICC-Crawler | BLOCKED at / · NICT ICC · named rule |
| GoogleOther | allowed · Google non-search fetch · wildcard rule |
| Google-CloudVertexBot | allowed · Vertex AI agent build · wildcard rule |
| GPTBot | BLOCKED at / · OpenAI training and index · named rule |
| ClaudeBot | BLOCKED at / · Anthropic training · named rule |
| anthropic-ai | BLOCKED at / · Anthropic legacy agent · named rule |
| Claude-Web | BLOCKED at / · Anthropic legacy agent · named rule |
| Google-Extended | BLOCKED at / · Gemini training · named rule |
| Applebot-Extended | BLOCKED at / · Apple training · named rule |
| CCBot | BLOCKED at / · Common Crawl · named rule |
| Bytespider | BLOCKED at / · ByteDance · named rule |
| meta-externalagent | BLOCKED at / · Meta AI training · named rule |
| FacebookBot | BLOCKED at / · Meta legacy · named rule |
| cohere-ai | allowed · Cohere · wildcard rule |
| Diffbot | BLOCKED at / · Diffbot knowledge graph · named rule |
| Omgilibot | BLOCKED at / · Webz.io · named rule |
| ImagesiftBot | allowed · Imagesift · wildcard rule |
| Timpibot | BLOCKED at / · Timpi · named rule |
| AI2Bot | BLOCKED at / · Allen Institute · named rule |
| Scrapy | allowed · Generic scraper framework · wildcard rule |
| SemrushBot-OCOB | allowed · Semrush AI corpus · wildcard rule |
| Applebot-Extended-Ads | BLOCKED at / · Apple ads corpus · named rule |
| TikTokSpider | allowed · TikTok · wildcard rule |
| QuillBot | allowed · QuillBot · wildcard rule |
| Webzio-Extended | BLOCKED at / · Webz.io extended · named rule |
| ProRataInc | allowed · ProRata · wildcard rule |
| AwarioBot | allowed · Awario · wildcard rule |
| MistralAI-Training | allowed · Mistral training crawl · wildcard rule |
| GoogleOther-Image | allowed · Google common crawler, public image URLs · wildcard rule |
| GoogleOther-Video | allowed · Google common crawler, public video URLs · wildcard rule |
| SEO and market-research crawlers — 14 of 14 allowed | |
| AhrefsBot | allowed · Ahrefs link index · wildcard rule |
| SemrushBot | allowed · Semrush crawler · wildcard rule |
| MJ12bot | allowed · Majestic link index · wildcard rule |
| DotBot | allowed · Moz link index · wildcard rule |
| rogerbot | allowed · Moz site crawler · wildcard rule |
| DataForSeoBot | allowed · DataForSEO · wildcard rule |
| BLEXBot | allowed · WebMeUp link index · wildcard rule |
| CloudflareBrowserRenderingCrawler | allowed · Cloudflare Browser Run /crawl · wildcard rule |
| Cloudflare-AutoRAG | allowed · Cloudflare AutoRAG · wildcard rule |
| Peer39_Crawler | allowed · Peer39 ad context · wildcard rule |
| AdsBot-Google | allowed · Google Ads quality · wildcard rule |
| AmazonAdBot | allowed · Amazon Ads · named rule |
| AdIdxBot | allowed · Microsoft Ads · wildcard rule |
| OAI-AdsBot | allowed · OpenAI ChatGPT ads page validation · wildcard rule |
| Regional search engines — 7 of 7 allowed | |
| Yeti | allowed · Naver · wildcard rule |
| YoudaoBot | allowed · Youdao · wildcard rule |
| Exabot | allowed · Exalead · wildcard rule |
| Sogou web spider | allowed · Sogou · wildcard rule |
| YisouSpider | allowed · Yisou · wildcard rule |
| 360Spider | allowed · 360 Search · wildcard rule |
| Sosospider | allowed · Soso · wildcard rule |
| Social link previews — 6 of 6 allowed | |
| Facebot | allowed · Meta link crawler · wildcard rule |
| Twitterbot | allowed · X link preview · named rule |
| LinkedInBot | allowed · LinkedIn link preview · wildcard rule |
| facebookexternalhit | allowed · Facebook link preview · named rule |
| Pinterestbot | allowed · Pinterest · wildcard rule |
| Slackbot-LinkExpanding | allowed · Slack unfurl · wildcard rule |
9 answer engine(s) are blocked. These are the agents that fetch a page while composing a reply, so this directly removes the site from live answers.
Speed a crawler sees, and field data where it exists — 67/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Largest Contentful Paint (p75) how long a real visitor waits before the main thing on the page appears | 3.45s | 2.5s or less | ✗ |
| Interaction to Next Paint (p75) how long the page takes to respond after a real visitor taps something | 154ms | 200ms or less | ✓ |
| Cumulative Layout Shift (p75) how much the page moves under a reader mid-read | 0.13 | 0.1 or less | ✗ |
| Time to First Byte (p75) the server half of every other number on this list | 522ms | 800ms or less | ✓ |
| Machine files answer quickly (measured here) the median time this scanner waited for your own robots, sitemap and entity files. A crawler that times out records nothing at all, so this is the speed number that decides whether you are read | 56 ms median of 5 | under 500 ms | ✓ |
| No slow outlier among them one slow file is enough to lose a crawl. This is the worst single response of the same set, not an average that hides it | 449 ms slowest | under 1500 ms | ✓ |
| Lighthouse lab run a lab run is a synthetic test from Google’s servers, not a reading of what your visitors experienced | the Lighthouse run is queued; its numbers appear on your next scan of this URL | a completed lab run where field data is absent | n/a |
Measured on this exact URL. These are the 75th percentile of what real Chrome users experienced over the last 28 days, not a test run from here — the thresholds are Google’s published ones, the only ones that are.
Payload and render path — 20/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Content ratio the rest is markup a model must read and discard | 4.4% | 10–100% of decompressed bytes are visible text | ✗ |
| Decompressed bytes large pages are fetched less often and truncated more | 297,605B | under 500,000B | ✓ |
| Inline JavaScript inline JS is pure overhead to a text-extracting crawler | 38,121B | under 20,000B | ✗ |
| Render-blocking resources each one delays first paint and the crawler's render budget | 24 | 0–5 | ✗ |
| Deferred vs blocking scripts a blocking script stops HTML parsing dead | 2 deferred / 14 blocking | every script deferred or async | ✗ |
Where the bytes go
| Visible text | 13,081 B | 4.4% |
|---|---|---|
| Inline CSS | 45,571 B | 15.3% |
| Inline JavaScript | 38,121 B | 12.8% |
| Structured data | 2,238 B | 0.8% |
| Markup and attributes | 195,658 B | 65.7% |
| Render-blocking resources | 24 | 10 stylesheets, 14 scripts |
| Entities with a stable @id | 1 / 3 | 33.3% |
| Credential edges | 0 | none declared |
1 of 1 map identity links point at places, not this business — 5990 Washington St, Denver, CO 80216. sameAs asserts that two URLs describe the same entity.
Machine-file trust chain — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| robots.txt was readable every other file in the chain is normally discovered through robots.txt | yes | 200, text/plain | ✓ |
| robots.txt names a sitemap a crawler that has to guess the sitemap path often does not find it | 1 Sitemap: line(s) | at least one | ✓ |
| robots.txt points at llms.txt an llms.txt nothing links to is only reachable by guessing the conventional path | — | named when the file exists | n/a |
| llms.txt links stay on this host an off-host URL in your llms.txt sends the model to someone else's page as if it were yours | — | most links on this host | n/a |
| llms.txt points onward the chain should keep going: llms.txt is a map, not a terminus | — | names the sitemap or the entity graph | n/a |
| entitymap.json parses as JSON a machine file that does not parse is worth less than one that is absent, because it looks present | — | valid JSON | n/a |
| entity graph references this host a graph that never names this site is describing something else | — | yes | n/a |
| Canonical points at this host a canonical on another host hands the page's standing to that host | same host | same host as the one serving the page | ✓ |
| Canonical uses https an http canonical invites a redirect chain on every crawl | https | https | ✓ |
robots.txt ✓ → sitemap ✓ → llms.txt ✗ → entitymap.json ✗
A ✗ breaks the chain at that point: everything downstream is only reachable by a crawler guessing the conventional path.
Claim consistency — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| The schema declares an entity for this site an entity graph that names only other companies gives an answer engine nothing to attach this site to | yes, an organisation | one Organization, LocalBusiness or Person whose url is this host | ✓ |
| Schema business name appears on the page an answer engine ingests the assertion and never compares it to the page, so a stale one is repeated for months | — | the name the schema asserts is the name a reader sees | n/a |
| Schema strings are not HTML-escaped a JSON string is not an HTML context: the entity is read literally, so the business name contains the characters a-m-p | — | no & or ' inside a JSON-LD value | n/a |
| Schema phone appears on the page a phone number that exists only in the markup is the one an assistant will read out | — | the same 10 digits | n/a |
| Schema locality appears on the page a locality nobody states on the page is a claim with no support behind it | — | the declared town or city is named in the copy | n/a |
| Experience claim is backed by the schema a datable claim in the copy that the structured data contradicts is the cheapest thing in an audit to disprove | — | the copy and foundingDate agree within a year | n/a |
Checked against the page’s own visible text, with no extra request. Node type: NewsMediaOrganization.
Agent instruction surface — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| No agent-directed instructions in machine-only surfaces hidden text, comments, alt attributes and llms.txt are read by a machine and proofread by nobody | 38 surface(s) read, nothing matched | 0 matches | ✓ |
| No hidden block over 50 words an agent ingests hidden copy at full weight while a reader never sees it | none | 0 blocks | ✓ |
| Agent-instruction file (agents.md) the file agents are told to obey, as distinct from llms.txt which is the content map. Every Shopify store now ships one; adoption elsewhere is early, so its absence is not a defect — but if you publish one, everything in it is read as instruction | 404 | your choice — not scored | n/a |
This reports EXPOSURE, not intent. Most hidden text is an old SEO habit or a collapsed menu, and a match here is a prompt to go and read it — not a finding that someone attacked the site.
What each agent receives — 88/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Every identity gets a response a request that dies is indistinguishable from a site that is down, to the agent making it | all 14 answered | 0 silent | ✓ |
| No answer engine is refused a 403 to GPTBot is the whole answer to why a site is never cited | none refused | 0 refusals | ✓ |
| No answer engine is handed an interstitial a challenge at HTTP 200 looks fine to a status check and contains no content at all | none | 0 challenge pages | ✓ |
| Crawlers get what an unnamed client gets these fetches send a crawler's user-agent from OUR address, which is not in the range that operator publishes — so a split has two readings, cloaking or correct spoof-rejection, and only the operator can settle which | none | no identity split | ✓ |
| No identity is sent to a different URL a bot-only redirect quietly removes the page an answer engine was asked to read | none | 0 redirected | ✓ |
| Same x-robots-tag for every identity a bot-only noindex removes the page from search while the site looks perfectly fine in a browser, and nothing else checks it | consistent | no identity-specific header | ✓ |
| Answer engines get the same text as a browser the page a browser renders is not evidence about the page an answer engine was given | 100% of the browser's words | 90-100% | ✓ |
| Training crawlers reaching the site these collect pages for model training rather than answering questions, so refusing them is a policy decision and not a defect — it is reported because the edge may be doing it without anyone having decided, and a 403 here is not the same fact as a 429 | Bytespider 403 of 6 turned away | your choice — not scored | n/a |
| Stated robots policy matches actual behaviour robots.txt is a promise and the edge is the behaviour; neither one on its own can tell you they disagree | OAI-SearchBot: disallowed but served, Claude-SearchBot: disallowed but served, PerplexityBot: disallowed but served, Applebot: disallowed but served | 0 conflicts | ✗ |
| Served the same content from every region tested a site that serves different content by country is telling different engines different things about itself | — | identical text worldwide | n/a |
One URL, fifteen fetches in the same second — an unnamed client, twelve named crawlers, a mobile browser, and the browser again as a control. The control fetches agreed (self-similarity 1), so text differences below are differences, not noise. A user-agent is a claim, including when we are the one making it.
| Identity | Status | Words | Same text | What happened |
|---|---|---|---|---|
| Unnamed client | 200 | 2036 | 1 | |
| GPTBot | 200 | 2036 | 1 | |
| ClaudeBot | 200 | 2036 | 1 | |
| OAI-SearchBot | 200 | 2036 | 1 | |
| Claude-SearchBot | 200 | 2036 | 1 | |
| PerplexityBot | 200 | 2036 | 0.91 | different body text |
| Googlebot | 200 | 2036 | 0.91 | different body text |
| Mobile browser | 200 | 2036 | 1 | |
| Bingbot | 200 | 2036 | 0.91 | different body text |
| Applebot | 200 | 2036 | 0.91 | different body text |
| Amazonbot | 200 | 2036 | 1 | |
| Bytespider | 403 | 5 | — | refused HTTP 403 — vendor unidentified; reproduced on a second request |
| Meta-ExternalAgent | 200 | 2036 | 0.91 | different body text |
| CCBot | 200 | 2036 | 0.91 | different body text |
| Agent | Conflict | Detail |
|---|---|---|
| OAI-SearchBot | disallowed but served | HTTP 200 with 2036 words |
| Claude-SearchBot | disallowed but served | HTTP 200 with 2036 words |
| PerplexityBot | disallowed but served | HTTP 200 with 2036 words |
| Applebot | disallowed but served | HTTP 200 with 2036 words |
Per-engine retrieval. Each answer engine reads through named crawlers with different jobs — one builds the index, one fetches live when a user asks, one collects training data. They are separate permissions and a site commonly grants one and refuses another. Nothing here is scored: these same facts are already scored once above, and this measures what an engine is permitted and given — never what a model has retained or would cite, which no scanner can see from outside.
| Engine | Index | Live fetch | Training | What the edge actually did | Addressed by name |
|---|---|---|---|---|---|
| ChatGPT | ✗ blocked | ✗ blocked | ✗ blocked | ✓ served 2036 words | — neither file |
| Claude | ✗ blocked | ✗ blocked | ✗ blocked | ✓ served 2036 words | — neither file |
| Perplexity | ✗ blocked | ✗ blocked | — — | ✓ served 2036 words | — neither file |
| Google AI Overviews / Gemini | ✓ allowed | ✓ allowed | ✗ blocked | ✓ served 2036 words | — neither file |
| Microsoft Copilot | ✓ allowed | ✓ allowed | — — | ✓ served 2036 words | — neither file |
| Apple Intelligence | ✗ blocked | — — | ✗ blocked | ✓ served 2036 words | — neither file |
| Meta AI | ✗ blocked | ✓ allowed | ✗ blocked | ✓ served 2036 words | — neither file |
| Amazon | ✗ blocked | — — | — — | ✓ served 2036 words | — neither file |
| DuckAssist | ✗ blocked | — — | — — | — not probed | — neither file |
| Mistral | — — | ✗ blocked | — — | — not probed | — neither file |
| Common Crawl (feeds many models) | — — | — — | ✗ blocked | ✓ served 2036 words | — neither file |
A blocked training column beside an allowed index column is coherent policy, not a defect: it says "answer with me, do not train on me." The last column is whether your llms.txt or agents.md addresses that engine by name — robots.txt is a permission, those two files are where an operator actually talks to an agent. Naming nobody is the norm and is not scored.
no vantage provider is configured and no extension reading exists for this site in the last 14 days, so this was not measured
Mobile surface — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Images and embeds with dimensions declared Media without width and height reserves no space, so everything below it moves when the image lands. This is read from the markup and is a risk indicator, not a measured CLS value — the real number needs a real page load. | 1 of 42 | 97.6% unsized — not scored | n/a |
| Viewport meta without it a phone renders the desktop layout scaled down, and that is what a mobile crawler records. Roughly 218 million live sites declare one (BuiltWith, August 2026), so its absence is the exception, not the norm. | width=device-width | width=device-width, initial-scale=1 | ✓ |
| Zoom is not locked locking zoom is an accessibility failure and a one-line fix | readers can zoom | no user-scalable=no, no maximum-scale under 1.5 | ✓ |
| Apple touch icon what a saved-to-homescreen shortcut and several share surfaces use | declared | one apple-touch-icon link — not scored | n/a |
| Theme colour sets the browser chrome on mobile; its absence is the cheapest visible gap on this list | absent | a theme-color meta — not scored | n/a |
| Beyond the basics almost every site declares a viewport and almost none declare the rest, so this row is where a site separates itself rather than a place it loses points | none beyond viewport | not scored — these are differentiators, not defects | n/a |
Local presence, citations and NAP — 92/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Links to its Google Business Profile the profile and the site are one entity to an answer engine, and this link is the only bridge between them it can see | yes | a g.page, maps.app.goo.gl or maps place URL | ✓ |
| That profile link resolves a dead profile link is worse than none: it asserts an identity that cannot be checked. A 429 or 403 here is Google throttling our check, not a fault on your site, and is left unscored | 200 → www.google.com | 200 on a Google host | ✓ |
| Street address in the structured data a service-area business legitimately omits this, so an absence is reported and not scored against you | declared | declared for a premises business | ✓ |
| Geo coordinates coordinates are how a machine ties the entity to a place without parsing an address string | declared | latitude and longitude | ✓ |
| Opening hours “are they open now” is one of the most common questions an assistant is asked about a local business | declared | openingHours or openingHoursSpecification | ✓ |
| Telephone in the structured data the number an assistant reads out comes from here, not from the page | declared | declared | ✓ |
| Coordinates are usable two decimal places is about a kilometre — fine for a city, useless for a storefront someone is being driven to | 1 pair(s), full precision | in range, not 0,0, at least 3 decimal places | ✓ |
| Each location has its own coordinates one centroid copied onto every branch tells an assistant they are all the same place | — | no two locations share a point | n/a |
| Coordinates agree with the declared country negating the longitude puts it inside US - a missing minus sign | outside US — longitude sign | the point falls inside the country the address names — reported, not scored | n/a |
| Coordinates match the declared address the address and the point are two independent claims about one place, and nothing else on the web compares them | — | within 500m | n/a |
| Location entities declared more location entities than real listings is the single most common way a local entity graph goes wrong | 1 | 1 per real premises — not scored | n/a |
| City on the page vs declared address the place a page markets and the place it declares are the same, so an engine has nothing to reconcile | both say Denver | reported, not scored | n/a |
| Business name in structured data two spellings of a legal name are two entities to a retrieval system, and it cannot tell which one you are | The Denver Post | exactly one spelling | ✓ |
| Schema and page agree on the phone number with only one source there is nothing to compare, which is not the same as agreement | only structured data declares one | both places state the same number | n/a |
| Phone numbers a machine can read an assistant dictating a number has to pick one; a second number is usually an old one that still rings somewhere | +13039541010 | exactly one number | ✓ |
| Postal address in structured data two addresses split the entity across two places | 5990 Washington St., Denver, CO, 80216 | one address, or none for a service-area business | ✓ |
| Profiles the site claims sameAs is an identity claim: every link says this business IS the thing at that URL | 3 | the profiles you actually own | ✓ |
| Off-site listings found by your phone number searched by the one key a service-area business cannot hide; Yelp, Angi, Thumbtack, Houzz, Facebook refuse an outside client and are recorded as our limit, not as an absence | 4 on YellowPages, DexKnows, BBB | every one you own, declared | n/a |
| Found off-site, not declared in sameAs a listing you never linked is corroborated without you: whatever it prints becomes your name and category in the answer | 4 — which ones, with a licence | 0 | ✗ |
| Listings printing a different business name two names on one phone number are two entities to a retrieval system | 0 | 0 | ✓ |
| Data Axle carries this phone number Data Axle (formerly Infogroup / infoUSA) is one of two US business-data suppliers Google’s own provider list names (Acxiom is the other); a business it does not carry is missing from one named feed. Reported, not scored. | refused (403) — our limit, not an absence | listed | n/a |
| Directory listings declared reported, not scored — how many citations a business needs is a marketing judgement, not a measurement | 0 | the ones that matter for your trade | n/a |
Checked because this site declares a local business entity (LocalBusiness). This measures the site side of Google Business Profile alignment — whether the business points at its own profile and carries the fields a profile is matched on. It does not read the profile itself: that needs an API key, and inventing facts about a listing we cannot see would be worse than reporting nothing. Profile link found: https://maps.app.goo.gl/aK261HiEe923wgBV7.
Every profile below is one the site itself declares in sameAs. Nothing here was discovered by guessing at directories — undeclared listings need an index this scan does not have.
| Profile | Answered | Your phone | Your name |
|---|---|---|---|
| Facebook https://www.facebook.com/denverpost | not checked | — | — |
| X https://x.com/denverpost | not checked | — | — |
| LinkedIn https://www.linkedin.com/company/denver-post-media/ | not checked | — | — |
Directory coverage
Ten sources an answer engine is likely to reach for when asked about a local business, checked against what this site declares. Declared means the site names the profile in its own structured data — the only thing this scan can verify. A listing that exists but is not declared will read as missing here, and that is itself worth fixing: an engine reading your site has no way to find it either.
| Yelp the single most-cited local source in AI answers after Google itself | not declared |
| BBB trust signal, and one of the few directories with a verification process an engine can lean on | not declared |
| Facebook carries hours and phone, and is read by several engines as a primary source | declared |
| Nextdoor hyperlocal, and disproportionately cited for home services | not declared |
| Thumbtack category-specific lead surface for trades | not declared |
| Angi category-specific, still heavily indexed | not declared |
| Bing Places feeds Copilot and, historically, several ChatGPT retrievals | not declared |
| Apple Maps the default map on every iPhone, and invisible to most SEO tooling | not declared |
| Trustpilot review corpus that answer engines quote directly | not declared |
| Yellow Pages low value alone, but a cheap consistency anchor | not declared |
9 of 10 are not declared. Each one is a place a retrieval system could have found a second, independent statement of your name, address and phone — and the agreement between those statements is what makes any of them trustworthy.
Machine layer over time. Our own history of a site starts the first time we scanned it; the Internet Archive holds what came before. Nothing here is scored — a site’s past is not a defect, and the Archive’s coverage is uneven, so a missing snapshot says nothing about the site.
Archived robots.txt versions | 40 distinct, 2002-09-23 → 2004-04-12 |
| Earliest copy named an AI crawler | no — the site names 33 today, so that policy was written after 2002-09-23 |
| Earliest copy, first line | <!-- libOpenCDA reports: OID = 36~11~ --> <html> <head> <title>The Denver Post Online - De |
Archived llms.txt | none archived |
Sitemap coverage — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| A sitemap was readable a sitemap is the only place a crawler learns about pages nothing links to | 790 URLs declared | at least one sitemap resolves and parses | ✓ |
| Sitemaps named in robots.txt resolve a robots.txt pointing at a dead sitemap sends every crawler to a 404 | 1 of 1 | every named sitemap returns 200 | ✓ |
| Homepage links are declared in the sitemap a page missing from the sitemap is still findable by following links; a page missing from both is findable by nothing | not measured — only 20 of 3000 child sitemaps were read | 0 undeclared | n/a |
| Sampled declared URLs resolve a sitemap that lists dead URLs spends a crawler's budget on nothing | 0 of 6 broken | 0 broken | ✓ |
| Sampled declared URLs are final declaring the pre-redirect URL makes every crawl pay an extra hop | 0 of 6 redirect | 0 redirects | ✓ |
| URLs declared in the sitemap | 790 20 of 3000 child sitemaps read — the count is a floor, not a total |
| Internal links on the homepage | 193 |
| Linked but not declared | not measured only 20 of 3000 child sitemaps were read, so a page missing from what we read may be declared in a part we did not read |
| Declared URLs sampled | 6 checked · 0 did not resolve · 0 redirected |
This compares the sitemap against the links on the homepage only, so it finds pages the sitemap omits — it cannot prove a declared page is unreachable, because that needs a full crawl. The reverse gap (declared but linked from nowhere) is real and is not measured here. A page missing from the sitemap is still findable by following links; a page missing from both is findable by nothing.
Internal linking — 100/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Internal links on this page a page that links nowhere is a dead end for a crawler following your own structure | 297 | enough to reach the pages that matter | ✓ |
| Distinct destinations a navigation repeated in a header and a footer counts twice and reaches the same places | 194 of 297 links | most links going somewhere different | ✓ |
| Anchor text that says nothing anchor text is the one description of a destination that a machine reads before deciding to follow it | 0 | zero | ✓ |
| Links with no words at all an icon or bare image link is silent — a screen reader and a crawler both get nothing | 0 | zero | ✓ |
| Internal links marked nofollow telling a crawler not to follow your own pages is almost always a plugin default nobody chose | 0 | zero | ✓ |
| External links outbound links are an editorial choice, not a defect | 44 | reported, not scored | n/a |
Naming and case consistency — 80/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Hostname case hostnames are case-insensitive but a mixed-case host splits logs and analytics | lowercase | always lowercase | ✓ |
| Internal link paths paths ARE case-sensitive on most origins: /About and /about are two URLs, two cache keys and two index entries | all 297 lowercase | 98–100% lowercase | ✓ |
| Trailing slash a page linked as /x and /x/ is two URLs to a crawler; both get fetched and both compete for the same canonical | 1 target linked both with and without a slash (/contact-us) | one form per URL, and the other form redirects to it | ✗ |
| Host form in absolute links links to both www and the bare domain send a crawler to two copies of the site whatever the canonical says | consistent | one hostname in every self-link | ✓ |
| Query strings in internal links a query string makes a distinct URL; navigation that carries one produces near-duplicate pages | none | under 2% of internal links (and fewer than 3) | ✓ |
Image signals — 80/100
| Measure | Currently | Optimal | |
|---|---|---|---|
| Images on the page counted from <img> tags in the delivered HTML | 41 (3 SVG) | n/a | |
| Missing alt text alt is the only text an engine gets from a picture; empty alt is correct for decoration, missing alt is a hole | 0 missing, 1 empty (0% missing) | 0 missing; empty only on decorative images | ✓ |
| Width and height declared without both, the browser cannot reserve space and the page shifts as pictures load (the CLS score) | 0 of 41 | all | ✗ |
| Lazy-loaded lazy-loading the hero image delays the largest paint; lazy-loading the rest speeds everything else | 0 of 41 | below the fold only | ✓ |
| Modern formats (WebP/AVIF) smaller files, same picture | 0 of 41 | most | n/a |
| Hero image prioritised the one image the browser should fetch first | 4 fetchpriority=high, 4 preloaded | one, when the largest paint is an image | n/a |
| Entity image in structured data what an engine shows next to the name; a URL in image or logo, ideally an ImageObject with dimensions | yes | an image or logo on the organisation node | ✓ |
| og:image the picture a link preview uses when the page is shared | yes | present | ✓ |
What an engine can learn from the pictures. Five rows score: missing alt, declared dimensions, nothing lazy-loaded above the fold, an entity image in JSON-LD and an og:image. Empty alt, file format and hero priority are reported and never scored — empty alt is correct on a decorative image, and format is a preference, not something an engine fails to read.
What this scan could not see 23 of 24 scored sections measured · 13 of 14 identities answered · 0 stage failures · 1 vantage · 4 not observed
| Subject | State | Attempted | Reason |
|---|---|---|---|
nl_entities | skipped | no | the deep entity read runs for domains under continuous record; this scan used the keyless resolver |
vantage | unmeasured | no | no vantage provider is configured and no extension reading exists for this site in the last 14 days, so this was not measured |
identity:Bytespider | refused | yes | HTTP 403 |
section:commerce | unmeasured | unknown | no row of this section was produced for this page; it did not enter the grade |
A non-observation is recorded, never scored: an unmeasured row is excluded from its section, not counted as a fail. Counts only; there is no composite confidence number.
Where each fact on this page comes from 16 facts · 21 statements · no disagreements
| Fact | Value stated | Stated in | |
|---|---|---|---|
name | The Denver Post | JSON-LD LocalBusiness meta og:site_name | agrees |
telephone | +13039541010 | JSON-LD LocalBusiness | one source |
email | newsroom@denverpost.com | JSON-LD LocalBusiness | one source |
address | 5990 Washington St., Denver, CO, 80216 | JSON-LD LocalBusiness | one source |
hours | Mo 8AM-4PM; Tu 8AM-4PM; We 8AM-4PM; Th 8AM-4PM; Fr 8AM-4PM | JSON-LD LocalBusiness | one source |
geo | 39.7392,104.9903 | JSON-LD LocalBusiness | one source |
url | https://www.denverpost.com/ | JSON-LD NewsMediaOrganization JSON-LD LocalBusiness | agrees |
logo | https://i0.wp.com/www.denverpost.com/wp-content/uploads/2020/11/denverpost.jpg?w=1200&crop=00px100630px&ssl=1 | JSON-LD LocalBusiness | one source |
foundingDate | 1892 | JSON-LD LocalBusiness | one source |
areaServed | 1 area: Denver Metropolitan Area and the state of Colorado | JSON-LD LocalBusiness | one source |
sameAs | 3 profiles | JSON-LD LocalBusiness | one source |
title | The Denver Post – Colorado breaking news, sports, business, weather, entertainment. | <title> | agrees |
| The Denver Post | meta og:title | ||
description | Colorado breaking news, sports, business, weather, entertainment. | meta og:description | one source |
canonical | https://www.denverpost.com | meta og:url | one source |
language | en-US | <html lang> | agrees |
| en_US | meta og:locale | ||
image | https://www.denverpost.com/wp-content/uploads/2020/11/denverpost.jpg | meta og:imagemeta twitter:image | agrees |
Read from the page this scan received: its schema nodes, meta tags, link elements and page links, each value kept with the element that stated it. A fact stated in several places should say the same thing in each. Recorded, not scored.
If you would rather not do this yourself
| Fix everything in this report The 8 findings above, including 5 scored serious, corrected on your site and re-measured afterwards. You get a change receipt: the before and after, each sealed and independently verifiable on your own machine, so the fix is evidenced rather than asserted.
| Get this fixed — $749 one local-business site · what that covers |
The fix includes the machine layer: an entity map, an agents.md and an llms.txt authored for this business rather than templated — 6 of the 7 agent surfaces checked above are absent here. A bigger site, or several domains at once, is quoted against the report first: email hello@crawlcheck.io with this address.
Measured, not in the grade (7)
These are measured on every scan and shown in full. Each says why it carries no weight yet: most are waiting on enough sites to calibrate against, and some depend on an integration a site may not have, which must never decide a grade.
Agent surfaces (emerging, not scored) — 10 of 10 rows read
Not in the grade: these are published standards at very different stages of adoption, and several do not apply to every business.
| Measure | Currently | Optimal | |
|---|---|---|---|
| agents.md The one surface here that pays off for an ordinary business today. Almost every agents.md on the web is a platform default its owner never wrote | absent | An authored file at /agents.md: what the business does, the area it serves, what an agent must collect before acting, and what it must not promise | n/a |
| SKILL.md The agent-skills convention. A 200 that carries HTML is a soft 404 and counts as absent - a naive check would call it present | absent (404) | A markdown file at /SKILL.md with YAML frontmatter naming the skill, then what an agent can do here and how | n/a |
| Markdown for agents Cuts what an agent must parse. A toggle on some edges rather than a build | absent (text/html; charset=utf-8) | Accept: text/markdown returns a markdown body; HTML stays the default for browsers | n/a |
| Link response headers Useful where a relation is real. Inventing one to score a point is the cargo-culting these lists encourage | present: <https://www.denverpost.com/wp-json/>; rel="https://api.w.org/", <https://wp.me/7yQli>; rel=shortlink | Registered relations only - canonical, alternate, describedby. Never an invented rel to satisfy a checker | n/a |
| Agent skills index For sites that expose actions an agent can perform. A brochure site has no skills to declare | absent | /.well-known/agent-skills/index.json listing each skill with a name, description, url and sha256 | n/a |
| API catalog NOT APPLICABLE: no public API found. A catalog of nothing is not an improvement | absent | Nothing to publish - no API was discoverable on this site | n/a |
| MCP server card Applies once you run an MCP server. Publishing a card without one advertises an endpoint that does not answer | absent | /.well-known/mcp/server-card.json with serverInfo, transport endpoint and capabilities | n/a |
| Agent payment protocols NOT APPLICABLE: no Product, Offer or checkout found. A scanner that docks a service business for this is measuring the wrong site | not applicable | Nothing to publish - no commerce signals on this site | n/a |
| OAuth / protected-resource metadata NOT APPLICABLE: nothing to authenticate against | not applicable | Nothing to publish - no protected API on this site | n/a |
| Web Bot Auth Identifies you as a caller, not as a destination. Irrelevant to a site that only receives traffic | informational | A JWKS at /.well-known/http-message-signatures-directory - only if THIS site sends signed agent requests to others | n/a |
None of these affect the grade. They are published standards at very different stages of adoption, and several do not apply to every business — where that is true this table says so and why. A checklist that counts them all is measuring the wrong site.
Content entities (measured, not yet scored) — 4 of 4 rows read
Not in the grade: the reading is sound but the scoring rule is not settled, and a rule that may move should not move a grade.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Entities read from the content declared entities come from the page's own JSON-LD; the rest are a HEURISTIC read of the prose and are labelled as such | 24 | the things this page is about | n/a |
| Resolve to a public entity record resolution means Wikidata holds an item for this name. It is not a ranking signal and does not mean the page ranks for it | 19 of 24 | a named thing the public record knows | n/a |
| Sense is the primary one, not the page's a bare name resolves to whatever the public record treats as its main sense — "Aurora" reads as the light display, not the Colorado city. The description is printed beside every resolution so you can see which sense you got. Resolving in CONTEXT needs a contextual model and is not measurable from a keyless lookup | not disambiguated | resolved in context | n/a |
| Resolved and absent from the entity graph a subject the prose is about, that the public record recognises, and that this site's own structured data never names — the writer's gap, not the machine's | 18 | 0 | n/a |
Entities are resolved against Wikidata — free, keyless, and it answers the question a writer actually has: does this thing exist as a public entity record, or is it just a phrase. Two evidence classes, kept apart: declared entities are read from the page's own JSON-LD; detected entities are a heuristic read of the prose and will contain some noise. Resolution is not a ranking signal.
Written about, recognised by the public record, absent from this site's structured data:
Colorado— Q1261 · state of the United States of America · 17 mentionsOpinion— Q3962655 · judgement, viewpoint, or statement that is not conclusive · 7 mentionsDeion Sanders— Q954184 · American football and baseball player and football coach · 4 mentionsAurora— Q40609 · natural light display that occurs in the sky, primarily at high latitudes (near the Arctic and Antarctic on Earth) or even on other planets · 3 mentionsEastern Plains— Q5148801 · region of the U.S. state of Colorado east of the Rocky Mountains · 3 mentionsHill Project— Q140360689 · video game · 3 mentionsRed Rocks— Q2182648 · concert venue near Morrison, Colorado, United States of America · 3 mentionsAmendment— Q1269627 · legal act proposed in a bill or motion, that adds, changes, substitutes, omits or abolishes clauses to one or several other acts which previously adopted or projected · 2 mentionsCalifornia— Q99 · state of the United States of America · 2 mentionsCoors Field— Q1129916 · baseball stadium in Denver, Colorado, USA; home venue of the Colorado Rockies · 2 mentionsDenver Broncos— Q223507 · National Football League franchise in Denver, Colorado · 2 mentionsDenver Pride— Q7242731 · an annual Gay pride event held each June in Denver · 2 mentions
6 more not listed.
Nothing in this section affects the grade — the corpus measures the distribution first.
Service area, as declared (measured, not scored) — 4 of 6 rows read
Not in the grade: a declared service area can be right or wrong only against facts this scan cannot see.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Service areas declared in schema the list an answer engine reads when asked "do they cover X"; a business that serves it but never declares it is invisible for that question | 1 | every place the business actually serves, as City or Place nodes | n/a |
| Declared radius the radius is the one claim that bounds all the others | no GeoCircle | a GeoCircle around the real centre | n/a |
| Areas that could be placed on a map a name that no gazetteer resolves is a name an engine cannot place either | 0 of 1 | every declared name resolves to a place | n/a |
| Areas with a page of their own declaring an area and having a page for it are different claims; the page is what gets cited | 0 of 1 | a location page per area that matters | n/a |
| Declared areas outside the declared radius the graph contradicts itself: the circle says no, the list says yes, and a reader has to guess which one is true | none | zero | n/a |
| Farthest declared area the distance the business is committing to on the record | — | inside the radius | n/a |
declared, has its own page · declared in schema only · outside the declared radius · ring = GeoCircle
Not placed (no gazetteer match near the centre): Denver Metropolitan Area and the state of Colorado
Not scored. Positions are OpenStreetMap place centroids resolved once per name; a page counts when a sitemap URL contains the place name. What this shows is what the site declares, laid out so the contradictions are visible; whether the business really works there is not on the wire.
The whole site, not just the front door (measured, not scored) — 21 of 21 rows read
Not in the grade: the grade stays anchored on the homepage so it is comparable across sites and days.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Pages declared, pages read what the sitemap promises versus what this pass could fetch; the cap is stated, never hidden | 774 declared, 60 read (first 60) | every declared page read | n/a |
| Pages not answering 200 a declared page that redirects, errors or is gone is a promise the sitemap breaks on every crawl | 0 of 60 | 0 | n/a |
| Thin pages (under 150 words) a page with a title and no body is indexed as nothing; location pages are the usual offenders | 0 of 60 | 0 | n/a |
| Near-duplicate pages Every audit tool accuses sites of duplicate content; almost none measure it. This compares the pages against each other after removing the text they all share, and reports a band rather than a percentage because a percentage would imply a precision the method does not have. | 0 near-identical, 65 substantially overlapping, of 1770 pairs | 0 near-identical pairs | n/a |
| Template text removed before comparing Shared navigation and footers appear on every page and would put a floor under every comparison. Passages carried by more than 90% of the crawled pages are dropped first, so what is left is each page’s own copy. | 303 of 38719 distinct passages | — | n/a |
| Pages with no JSON-LD the homepage graph does not travel; a service page without its own node is anonymous to a resolver | 0 of 60 | 0 on pages that describe a service or place | n/a |
| Canonical points elsewhere a page telling crawlers to index a different page is either a deliberate merge or a copied template | 11 of 60 | 0 unintended | n/a |
| Orphans (declared, linked from nowhere read) a page the sitemap declares but no page links to is reachable only by the sitemap, and weighted accordingly | 0 of 60 | 0 | n/a |
| Undeclared pages (linked, not in the sitemap) linked pages the sitemap forgot; pagination is normal, a service page is not | 100 | 0 that matter | n/a |
| noindex on declared pages declaring a page and then telling crawlers not to index it is two files disagreeing | 0 of 60 | 0 | n/a |
| Duplicate titles two pages with one title compete with each other for the same query | 0 | 0 | n/a |
| Lane: article (59 pages) measured against the rows that apply to article pages only: answers 200; canonical points at itself; not noindex; exactly one h1; 150+ words of its own; an Article node; a publication date. A lane is read from what each page declares (its JSON-LD types, its path, a tappable phone number), so a city page and a docs page are no longer failed with the same ruler. | 11 of 59 missing canonical points at itself | 0 missing | n/a |
| Lane: home (1 page) measured against the rows that apply to home pages only: answers 200; canonical points at itself; not noindex; exactly one h1. A lane is read from what each page declares (its JSON-LD types, its path, a tappable phone number), so a city page and a docs page are no longer failed with the same ruler. | 1 of 1 missing exactly one h1 | 0 missing | n/a |
| Facts that disagree: phone number a resolver that reads two values for the same fact on the same site has to choose or hedge; the pages carrying each value are on the record. Example: https://www.denverpost.com/ versus https://www.denverpost.com/ | “+13039541010” on 1 page · “+13039541133” on 1 page | one value | n/a |
| Facts that disagree: email address a resolver that reads two values for the same fact on the same site has to choose or hedge; the pages carrying each value are on the record. Example: https://www.denverpost.com/ versus https://www.denverpost.com/ | “newsroom@denverpost.com” on 1 page · “cmoser@denverpostmedia.com” on 1 page | one value | n/a |
| Template families pages grouped by the assets they ship (stylesheets, scripts, inline blocks of 2 KB or more), not by their words. A defect in a family is one fix, however many pages carry it. | 6 families across the pages read, 2 one-offs | informational | n/a |
| Family 1 (27 pages: 27 article) pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/2026/09/24/daily-horoscope-for-sept-24-2026/ | 48 KB inline CSS and 35 KB inline script on a typical page, 5.5% text — 3 inline blocks byte-identical on every page (63 KB each, 1709 KB across the family) | a block every page carries belongs in one cached file | n/a |
| Family 2 (15 pages: 15 article) pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/obituaries/phillip-thomas-longo-littleton-co/ | 1.4 KB inline CSS and 33 KB inline script on a typical page, 3.3% text — 2 inline blocks byte-identical on every page (17 KB each, 256 KB across the family) | a block every page carries belongs in one cached file | n/a |
| Family 3 (14 pages: 14 article) pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/obituaries/jeanette-nail-denver-co/ | 1.4 KB inline CSS and 32 KB inline script on a typical page, 3.2% text — 2 inline blocks byte-identical on every page (17 KB each, 239 KB across the family) | a block every page carries belongs in one cached file | n/a |
| Family 4 (2 pages: 2 article) pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/2026/09/23/best-fall-clothes-for-children/ | 48 KB inline CSS and 34 KB inline script on a typical page, 6.8% text — 3 inline blocks byte-identical on every page (63 KB each, 127 KB across the family) | a block every page carries belongs in one cached file | n/a |
| Median words per page the shape of the site as delivered, not just its front door | 1242 · 4.7% text | informational | n/a |
60 pages read in 11 s, 2026-09-24 14:10 UTC. Read once a day after a scan somebody asked for; the cron sweeps never pay for it.
Which pages: the lists open with Watch — free for 30 days from /pricing. The counts above are the full measurement; the list is the part that saves the walk through the sitemap.
Not scored: the grade stays anchored on the homepage so it is comparable across sites and days. These counts are the same walk the depth tier runs at full length (300 pages, with the link graph) via /api/depth.
Freshness, as the site declares it (measured, not scored) — 6 of 7 rows read
Not in the grade: every date here is one the site published, so this reports what it claims rather than what is true. Calibrated 2026-09-16 on 114 records: the only row most sites carry (a sitemap that declares lastmod, 68%) is not a defect when absent, and the rows that would be defects (every lastmod stamped on one day, 8%; the sitemap and dateModified disagreeing by over a month, 2%) can be measured on too few sites to carry weight.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Sitemap URLs with a change date lastmod is the only change signal an engine can read without an account. A sitemap without dates tells a crawler nothing about what to re-fetch first | 790 of 790 | every URL, with its real date | n/a |
| Newest sitemap date what the site says was touched most recently; the gap to today is a claim about activity, not a defect on its own | 2026-09-24 (today) | within the last month for a site that publishes | n/a |
| Spread of sitemap dates a spread of dates is what a real edit history looks like | 26 distinct days across 790 dated URLs (2026-08-26 to 2026-09-24) | many days - real edit history | n/a |
| Homepage Last-Modified header the HTTP-level change date; a cache in front usually sets it, so it describes the copy served, not the edit | 2026-09-24 (today) | sent, and true | n/a |
| Homepage dateModified in JSON-LD the date the structured data asserts; an engine can quote it beside a fact, which is why a stale one is worse than none | none | present, and moved when the page changed | n/a |
| Dated elements on the page a <time datetime> element is a date a machine can read without guessing at the prose | 7 (newest 2026-09-24, today) | every date a person can see also marked up | n/a |
| robots.txt history in the public archive how long the site has been keeping a machine-readable front door; time depth is the one claim a new site cannot manufacture | first seen 2002-09-23, 40 versions | a long record | n/a |
Every date in this section is one the site published about itself. The section reports the claims and where they disagree; it does not decide whether a page is out of date, because only the business knows that. None of this moves the grade.
Local intents, the Maps vocabulary (measured, not scored) — 2 of 3 rows read
Not in the grade: the intent vocabulary is recovered names with no weights, and a claimed intent without a page is a content decision, not a defect; reported until the corpus says how it distributes.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Local intents the site's own declarations name a local query is resolved into one of 446 intent types before any place is retrieved; the site's business types, Service nodes, title and headings are what name them here | 1 (post offices) | every intent you serve, and none you do not | n/a |
| Claimed intents with a page linked from the homepage an intent named only in a paragraph competes with every page on the web that names it; a page whose address carries the intent is what a query lands on | 0 of 1 - no page for: post offices | all of them | n/a |
| Intents searchers used to find the Business Profile no Business Profile search terms on record for this domain | not measured | read from the profile's search terms | n/a |
Conversion path on the page (measured, not scored) — 5 of 6 rows read
Not in the grade: whether a visitor can act is a business judgement, and the grade is for what a machine receives.
| Measure | Currently | Optimal | |
|---|---|---|---|
| Phone link on the page on a phone, a tel: link is the shortest path from reading to calling; a printed number is not tappable | none | at least one tap-to-call link | n/a |
| Phone link before the fold the visitor who arrived from an answer engine has already decided; the number should be where they land | no | in the first quarter of the page | n/a |
| Quote or contact form a form is the only path that works when the office is closed | 1 found, one before the fold | one form, reachable without scrolling | n/a |
| Calls to action counted from link and button text: quote, estimate, call, book, schedule, contact | 2 (first one before the fold) | one clear ask, early | n/a |
| Contact or quote page linked the page a visitor goes to when the hero did not convert | yes | linked from every page | n/a |
| Sticky call button (mobile) a fixed call bar keeps the number reachable at every scroll position on a phone | not detected | present on service sites | n/a |
The SXO layer: whether a visitor who arrived ready to act can act. Counted from the delivered HTML, so a form inside a third-party iframe shows as a form only if the iframe names one. None of this moves the grade.
Common Crawl visibility — measured, not scored
CCBot is blocked by robots.txt, so future crawls will skip this site. Checked against September 2026 Index, from a stored result.
0 of 6 pages checked were captured.
| Page | In the dataset | Captured | Status |
|---|---|---|---|
| / | not found | — | — |
| /sitemap.xml?yyyy=2026&mm=09&dd=24 | not found | — | — |
| /sitemap.xml?yyyy=2026&mm=09&dd=23 | not found | — | — |
| /sitemap.xml?yyyy=2026&mm=09&dd=22 | not found | — | — |
| /sitemap.xml?yyyy=2026&mm=09&dd=21 | not found | — | — |
| /sitemap.xml?yyyy=2026&mm=09&dd=20 | not found | — | — |
A capture means the page is in an open dataset. It does not establish that any model trained on it, and nothing here claims otherwise.
Deep entity read — measured, not scored
the deep entity read runs for domains under continuous record; this scan used the keyless resolver
Identity fingerprint — measured, not scored
The identity this site publishes is unchanged since 2026-09-24. Same name, same phone, same address, same declared profiles.
Schema type newsmediaorganization. We hold no population figure for this type, so no rarity is claimed.
A fingerprint of the identity this site publishes, recomputed on every scan. A change here is a change the site made, not a judgement about it. Anyone can recompute this from the published page: the canonical form and the hash are both public at /api/entity-drift/latest.
How this compares — 60 wordpress/publisher sites, of 866 measured
| Measure | This site | Corpus median | Rank |
|---|---|---|---|
| AI visibility score (the grade shown) | 82 | 78 | 62nd percentile |
| Section average (weighted) | 80 | 82 | 30th percentile |
| Content ratio (% visible text) | 4.4 | 4 | 52nd percentile |
| Stable @id coverage (%) | 33.3 | 81.8 | 10th percentile |
| Render-blocking resources | 24 | 6.5 | 12th percentile |
| Delivered page size (KB) | 291 | 176.5 | 8th percentile |
Percentiles come from sites this scanner has measured itself, not from a published study — so they describe this corpus, not the web. The corpus is not a random sample and is weighted toward sites that were submitted or seeded, which is why the rank is shown next to the raw number rather than instead of it. A metric is left blank rather than ranked when fewer than 8 peers carry it. This rank is against the wordpress/publisher cohort (60 sites), not the whole corpus of 866, because a WordPress local-service site and a static personal site are not the same measurement problem.
Reach, read, quote — the three stages
An engine can reach and read this site. What is left is giving it something specific to say.
| Stage | Question | Score | |
|---|---|---|---|
| Reach 40% of the score | Can a named answer-engine crawler get your pages at all? At least one answer engine was refused, challenged or served less than a browser. | 81 | |
| Read 30% of the score | Once it has the bytes, can it find the words? The content is buried in markup, or the files that guide a crawler are missing. | 79 | |
| Quote 30% of the score | Is there a specific fact it can state and attribute? The facts an assistant is asked for are declared and consistent. | 85 |
Which engines, by name
From the probes in this scan — the same fetch each engine’s crawler makes, from our address.
| Served the page | OAI-SearchBot · Claude-SearchBot · PerplexityBot · Googlebot · Bingbot · Applebot |
| Disallowed in robots.txt | Yahoo Slurp · Yandex (umbrella token) · ChatGPT search index · ChatGPT live fetch · Claude search index · Claude live fetch · Perplexity index · Perplexity live fetch a policy refusal, which is a decision rather than a defect — it belongs here so it is visible, not because it is wrong |
Fix the 24 items this report already lists and the same arithmetic reads 100. That is addition on the table below, not a forecast — every point has a named component behind it:
- Answer engines allowed — now
8 of 17, worth 1.8 points - Search indexes allowed — now
16 of 24, worth 1.8 points - Cost of a miss — now
3 missing file(s) return >20KB of HTML, worth 1.3 points - entitymap.json — now
absent (404), worth 1.3 points
The stage to fix first is read, because the three run in order: an engine cannot read what it was refused, and it cannot quote what it could not read.
What this number covers. Everything above was measured on pages this scan fetched from your own domain. An answer engine also reads sources you do not control — directory listings, review corpora, and your Google Business Profile — and none of those are in this score. The profiles this site declares are listed below but were not fetched on this scan, so nothing here reflects what those listings actually say. A site can score well here and still be described wrongly by an engine reading a listing that disagrees with it.
What this number is not. It does not measure whether ChatGPT, Perplexity or Google’s AI answers actually cite you. No server can measure that, and a score that implied otherwise would be invented. This measures whether an engine can — reach, read, quote — which is the part you control and the part that has to be true first.
Since the last scan
Nothing. Every finding on this scan was present on the previous one at the same severity, and nothing that was open has cleared. Compared against 13 recorded scans of this domain.
Change since last scan — 11 distinct states across 13 scans
Comparing against the last reading that differed, taken 2026-09-24T14:10:19.087Z.
| Measure | Then | Now | Change |
|---|---|---|---|
| Overall score | — | 80/100 | not comparable |
| Content ratio | 4.4% | 4.4% | no change |
| Stable @id coverage | 33.3% | 33.3% | no change |
| Render-blocking resources | 24 | 24 | no change |
| Page size | 292 KB | 291 KB | -1 KB ✓ |
| Findings | 8 | 8 | no change |
History keeps the last 24 changes on this domain — repeated identical readings collapse into one entry that records when the state began and how many scans confirmed it, so re-running a scan can never push real history out of the window. A finding that comes and goes is the pattern worth acting on: a permanent failure gets noticed, an intermittent one only shows up if something is looking at the right moment.
Since the first scan
Measured on 7 separate days between 2026-08-20 and 2026-09-24. 1 defect has stopped being measurable and 8 are still open.
What this does and does not say. Each row below is two dates: the day a defect was first measured here, and the first day it could no longer be measured. That is evidence the defect is gone. It is not evidence that we caused it to go, and nothing on this page claims otherwise — a site can be fixed for reasons that have nothing to do with a report. What is verifiable is the pair of measurements, and either date can be checked against the sealed record without asking us.
Headline measurements, first reading against latest
| Measure | At first scan | Now | |
|---|---|---|---|
| Grade | F | C | worse |
| Overall score | 73 | 82 | better |
| Section average (weighted) | 77 | 80 | better |
| Visible text | 4.6% | 4.4% | worse |
| Render-blocking resources | 24 | 24 | no change |
| Page weight | 279 KB | 291 KB | worse |
| Open findings | 4 | 8 | worse |
Defects that stopped being measurable
| Finding | First measured | Last measured | Days |
|---|---|---|---|
| The sitemap returns HTML, not XML SITEMAP_IS_HTML | 2026-08-20 | 2026-09-10 | 21 |
Median time from first measurement to no longer measurable: 21 days.
Still open
| Finding | First measured | Days open |
|---|---|---|
| Two different entities share the same identifying data ENTITY_COLLISION | 2026-08-20 | 35 |
| Almost nothing a machine receives from this page is readable text PAGE_IS_MOSTLY_CODE | 2026-08-20 | 35 |
| robots.txt rules do not apply to named agents ROBOTS_RULES_SHADOWED | 2026-09-10 | 14 |
| RENDER_RESOURCE_BLOCKED RENDER_RESOURCE_BLOCKED | 2026-09-10 | 14 |
| Published coordinates are not in the country the address names GEO_COUNTRY_MISMATCH | 2026-09-10 | 14 |
| No llms.txt NO_LLMS_TXT | 2026-08-20 | 35 |
| NAP_UNDECLARED_LISTING NAP_UNDECLARED_LISTING | 2026-09-10 | 14 |
| The machine files cost more than the page MACHINE_CHAIN_HEAVY | 2026-09-16 | 8 |
This section is reported and never scored: a site that has improved a lot and a site that was always fine should get the same grade for the same measurements today.
Traffic, enquiries and sales
Not measured here, and not measurable here. Everything else in this report was produced by fetching your pages from outside. Clicks, enquiries and orders happen where no scanner can see them — in your Search Console, your inbox and your checkout. Connect one of those and this section fills in against the dates above. Until then it stays empty rather than showing an estimate, because an invented traffic figure on a report that measures honesty would be the whole product undone.
Cite this report
This report regenerates when it is opened, so both records carry the scan timestamp as the version and the access date as the retrieval. A citation that implied fixed content would be wrong.
BibTeX download
@techreport{crawlcheck_denverpost_com_iahhc6uf5p,
author = {{CrawlCheck}},
title = {Machine-layer measurement of {denverpost.com}},
institution = {CrawlCheck},
type = {Scan report},
number = {iahhc6uf5p},
year = {2026},
url = {https://crawlcheck.io/r/iahhc6uf5p},
urldate = {2026-09-28},
note = {Scanned 2026-09-24; report regenerates on access, figures are as of the access date}
}
RIS download
TY - RPRT TI - Machine-layer measurement of denverpost.com AU - CrawlCheck PB - CrawlCheck PY - 2026 DA - 2026/09/24 M1 - iahhc6uf5p UR - https://crawlcheck.io/r/iahhc6uf5p Y2 - 2026/09/28 N1 - Report regenerates on access; figures are as of the access date ER -
Type is RPRT — a dated report about one named site, not the aggregate dataset at /data. Note that /r/ is disallowed in our robots.txt, so this markup serves readers and reference managers, not search engines.
13
passing
7
needs work
3
failing
7
not measurable
Sections, not individual checks. Not measurable is counted separately and never as a pass — a section we could not read is not a section that passed.
Show this on your site
The badge renders the number as measured, whatever it is — it updates when this report is re-run. It is keyed on this report’s id, so only someone holding this link can produce it.
Paste this where you want it:
<a href="https://crawlcheck.io/r/iahhc6uf5p?via=badge"><img src="https://crawlcheck.io/badge/iahhc6uf5p.svg" alt="AI visibility 82/100, measured by CrawlCheck" width="228" height="64"></a>No landing has arrived from this badge yet. The link carries via=badge, so any that do are counted here.
Reading the transcript
CrawlCheck does not claim to predict rankings. It measures the inputs a crawler and a retrieval system must successfully receive before ranking, indexing or quoting is possible at all.
1
Observed
This scanner received this exact status, header, body and rendered-text result. Everything in the sections below is at this level unless it says otherwise — it is a fact about a real response, with the bytes attached.
2
Protocol consequence
What that response necessarily means for a compliant crawler. A disallow rule, a non-text robots.txt or a JavaScript dependency can prevent or impair access by construction — not because we ranked anything, but because the protocol says so.
3
Search implication
Possible or unknown, and labelled as such. No finding here asserts a ranking effect, a citation formula, or a weighting. If a claim cannot be reduced to level 1 or 2, it is not made.
An important limit, stated plainly: the 15 identities above are sent from this scanner’s own address wearing each crawler’s user-agent string. That is an edge-policy simulation — it shows what your edge does with a request that declares itself as that crawler. It is not a request from the real crawler, and it cannot be. A WAF may treat a verified Googlebot differently from anything merely claiming the name, which is exactly what Google recommends it do. Where a result says a crawler was refused, read it as a client declaring that name was refused — to prove what the real crawler received, check your own server logs against the operator’s published address ranges, which the free log verifier does line by line.
Transcript of the machine-file fetches
| Path | Status | Content type | Bytes | Edge | |
|---|---|---|---|---|---|
| /robots.txt | 200 | text/plain | 4,425 | BYPASS | ok |
| /sitemap.xml | 200 | application/xml | 884,135 | DYNAMIC | ok |
| /sitemap_index.xml | 404 | text/html | 88,178 | DYNAMIC | not present |
| /wp-sitemap.xml | 200 | text/html | 142,160 | DYNAMIC | not present |
| /llms.txt | 404 | text/html | 88,151 | DYNAMIC | not present |
Reading the scores
- The two scores are not a grade
- Crawl clarity is whether a machine can fetch, read and tell your pages apart. User clarity is whether a reader can tell what this is and who it belongs to. They fail independently, so read them separately rather than averaging them.
- A dash means not measured, not zero
- Components shown as
—and excluded from the score. If a scan was rate-limited or the pages render with JavaScript, that measurement is withheld. A false clean bill of health is more dangerous than a false defect. - Unconfirmed findings are separated on purpose
- Anything that could describe our access rather than your site — a refusal, a timeout — is listed apart and never counted toward the grade.
- Content ratio is bytes, not tokens
- The share of decompressed bytes that are visible text - the body as a parser reads it, not the smaller gzipped transfer. The deep audit measures tokens per KB with a real tokenizer; the two are named differently because they are different measurements.
- What to do with it
- Prefer fixes that are one setting over ones that need a sprint. Inline CSS is usually the largest cheap win, because it cannot be cached between pages — a crawler pays for it again on every URL.
What this scan does NOT measure
This is a machine-layer probe of one URL and the files that describe the whole site. Everything below is outside that scope by design, and no score on this page is adjusted for it.
- Not a site-wide crawler
- There is no full URL graph here, no broken-link inventory across every page, and no duplicate-content comparison. Sitemap coverage is checked against the homepage’s links only. For an exhaustive crawl of thousands of URLs, use a dedicated crawler — Screaming Frog, Sitebulb, or the site audit in Ahrefs or Semrush — and use this to check the machine layer they do not read.
- Not a log analysis platform
- Crawler behaviour is sampled in one tightly controlled interval, not tracked over weeks. Crawl budget, bot paths, and which pages a crawler actually spends its time on can only come from your own server logs. The free log verifier on this site reads those line by line against the operators’ published address ranges.
- Not a replacement for Search Console
- Index coverage, canonical selection, manual actions and Google-specific diagnostics come from Search Console and nowhere else. A page can be perfectly readable to every crawler measured here and still not be indexed, for reasons only Google can report.
- Extractability, and arrivals - but not citation
- This measures whether the bytes an engine receives can be fetched, parsed and attributed to you. Where the first-party beacon is installed it also counts arrivals: humans who clicked through from an AI assistant, recorded from your own site rather than guessed. What it still does not do is ask ChatGPT, Claude, Perplexity or Google whether they mention you when nobody clicks. A citation that produces no click is invisible here, and measuring that needs prompt monitoring, which is a different instrument.
- Our vantage point is not the crawler’s
- Every request here is honest and unauthenticated: it is sent from a datacentre address under our own name and never impersonates a verified crawler’s IP. A refusal therefore describes what a client declaring that name receives — which a WAF may treat differently from the real crawler, exactly as Google recommends it should.
- Not a WCAG audit
- The accessibility section reads whether the structure a screen reader depends on exists in the delivered document: a language, a title, names on images, inputs, buttons and frames, a main landmark, a way past the navigation, unique ids. Contrast, focus order, keyboard traps and motion need a rendered page and a person, and are not measured.
- One rendering, no JavaScript execution
- What is measured is the document the server returns. Content injected by client-side JavaScript is not executed here — which is deliberate, because most crawlers measured on this page do not execute it either, but it means a JavaScript-rendered site will read thinner here than it looks in a browser.
What we could not measure — 2
Every row in this report is either a measurement of your site or nothing at all. These lookups did not return a measurement, so they were left unscored rather than counted against you. They are listed because a silent absence is indistinguishable from a clean result, and that is the difference this tool exists to keep.
This covers the third-party and enrichment lookups only — CrUX, the Google profile link, the Internet Archive, the geocoder, and the two well-known files. Your own server’s refusals are never listed here: a 403 to GPTBot is the finding, not a gap in our instrument, and it is reported as a finding above.
| Lookup | Outcome | Whose limitation | Why |
|---|---|---|---|
| Off-site presence (directories searched by phone number) | reached | yours | Yelp, Angi, Thumbtack, Houzz, Facebook refuse datacentre clients and are our limit, never an absence. Google Business Profile and review counts are still not read from here |
| Response compression | unreached | ours | our fetch decompresses the body and drops the content-encoding header before we can read it — check it in a browser network panel or with curl --compressed -I |
Every figure here was measured with no cookies and no session. That is deliberate: it is what an answer engine gets. If you open your own site while logged into its admin, you are looking at a different document — caching plugins, page builders and membership tools routinely skip their output filters for signed-in users, and admin bars and edit links are added on top.
To see what this report describes, open your site in a private window. If the two disagree, the private window is the one a crawler sees.
| component | status | score | measured |
|---|---|---|---|
| Machine files | scored | 60 | robots.txt, llms.txt, sitemap and entitymap as scored above |
| Payload — content ratio | scored | 26 | 4.4% of decompressed bytes are visible text (byte-level; the deep audit scores tokens per KB separately) |
| Render path | scored | 0 | 24 render-blocking resources |
| Document structure | scored | 80 | 0 heading skips, 2 h1 |
| component | status | score | measured |
|---|---|---|---|
| Entity identity | scored | 33 | 33.3% of 3 entities carry a stable @id |
| Declared standing | scored | 40 | no credential edges declared |
| Location coherence | scored | 100 | 1/1 locations have a street address |
| Identity claims | scored | 0 | 1 of 1 map links point at places, not the business |
| Page labelling | scored | 100 | 157-char description, canonical present |
0 components not measured and excluded from the score — an unmeasured component is never counted as zero.
Score composition — The letter was set by the worst finding, not by the average — that is why a high score can carry a low grade.
| Section | Score | Weight |
|---|---|---|
| Machine layer | 60/100 | 1.5 |
| Server and hosting | 80/100 | 0.6 |
| Schema and entity graph | 60/100 | 1.2 |
| RDF and linked-data surface | 100/100 | 0.8 |
| Authorship and trust signals (E-E-A-T proxies) | 57/100 | 1 |
| Accessibility structure (from the delivered HTML) | 88/100 | 0.8 |
| Entity corroboration | 100/100 | 0.6 |
| Quotable content | 67/100 | 0.8 |
| Crawl waste | 100/100 | 0.6 |
| Page subject, the Webref question | 100/100 | 0.6 |
| AI and search agents | 33/100 | 1.2 |
| Speed a crawler sees, and field data where it exists | 67/100 | 1 |
| Payload and render path | 20/100 | 1 |
| Machine-file trust chain | 100/100 | 1.5 |
| Claim consistency | 100/100 | 1.3 |
| Agent instruction surface | 100/100 | 1 |
| What each agent receives | 88/100 | 1.5 |
| Mobile surface | 100/100 | 0.8 |
| Local presence, citations and NAP | 92/100 | 1.3 |
| Sitemap coverage | 100/100 | 1.2 |
| Internal linking | 100/100 | 1 |
| Naming and case consistency | 80/100 | 0.4 |
| Image signals | 80/100 | 0.6 |
Not scored, and therefore excluded rather than counted as zero: Agent surfaces (emerging, not scored), Content entities (measured, not yet scored), Service area, as declared (measured, not scored), The whole site, not just the front door (measured, not scored), Freshness, as the site declares it (measured, not scored), Local intents, the Maps vocabulary (measured, not scored), Conversion path on the page (measured, not scored).
Reach, then read, then quote
Reach, then read, then quote — each stage gates the next
Watch this site
These break silently and come back on their own. We re-check this site twice a day and record every change against a fingerprint of the last result, so nothing is missed between visits.
Changes are recorded from the moment you sign up, and you get one email when a result moves — nothing otherwise. Every alert carries a one-click stop link. A trial or Watch licence lets you read the record in full.
Export this result: JSON · CSV Same stored record as this page, so an export can never disagree with the report.