CrawlCheck

This sample is shown unlocked, so you can see what a licence includes. On your own domain the free scan gives you the grade, the score, which stage fails and every finding by name, with the top one in full; the detail, the fix and the bytes behind each one are what a licence buys.

Scan your own domain What a licence costs

Scan result

denverpost.com

2026-09-24 · 3 days agowordpressid iahhc6uf5p

82
Cgrade

Answer engines can reach and read this site. It needs more it can quote.

AI visibility 82/100. An engine can reach and read this site. What is left is giving it something specific to say.

Held at C The score alone would grade higher; a medium-severity finding sets the letter. Clear it and the grade lifts on its own.

What this does not prove

This is what crawlers receive from the site today — reach, read, quote — measured from CrawlCheck’s own network. It does not show whether any engine cites the site, where it ranks, or what traffic or leads follow, and a crawler-identity probe is not a fetch from that operator’s verified addresses. What a number can and cannot say

8

findings

13/23

sections passing

+20

points available

8/60

publisher reference sites measured so far - the comparison appears at 30

11

scans on record

The short version

The website's robots.txt rules do not apply to all named agents, which can cause issues with crawlers. The site also has an identity collision in its structured data, where two nodes share the same URL. Additionally, the homepage depends on files that are disallowed by the site's own robots.txt, which can prevent proper rendering.

In plain English

  • 13 of 23 scored sections pass. 7 need work, 3 are failing. The other 7 of the 30 sections are measured for information and never scored.
  • Fix first: robots.txt rules do not apply to named agents 40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignores `*` entirely, so those rules do not… This one is holding the grade at C. See the fix list.
  • Fix the 24 failing measures in the fix list and the score reads 100. Now 80, a gain of 20 points. That is arithmetic on measurements already taken, not a forecast. Each measure is a named row with its point value.
  • Every named answer engine was served the page. OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot. Disallowed by policy in robots.txt: Yahoo Slurp, Yandex (umbrella token), ChatGPT search index, ChatGPT live fetch, Claude search index, Claude live fetch, Perplexity index, Perplexity live fetch (a choice, not a defect).

Start here

The three findings holding the grade down most. Open one for what was seen and how to fix it.

The three stages, in the order an engine hits a site

Each stage gates the next. An engine that cannot reach a page never reads it; one that cannot read it never quotes it. That is why reach is a ceiling and not just a weight.

81

Stage 1 · 40% of the score

Reach

Can a named answer-engine crawler get your pages at all?

At least one answer engine was refused, challenged or served less than a browser.

79

Stage 2 · 30% of the score

Read

Once it has the bytes, can it find the words?

The content is buried in markup, or the files that guide a crawler are missing.

85

Stage 3 · 30% of the score

Quote

Is there a specific fact it can state and attribute?

The facts an assistant is asked for are declared and consistent.

The stage to fix first is read, because the three run in order.

How this compares

Ranked against 60 sites in the same cohort (wordpress/publisher). This corpus is sites people chose to scan, not the web, so a rank here is a comparison, not a verdict.

AI visibility score (the grade shown)you 82 · median 78
62th
Section average (weighted)you 80 · median 82
30th
Content ratio (% visible text)you 4.4 · median 4
52th
Stable @id coverage (%)you 33.3 · median 81.8
10th
Render-blocking resourcesyou 24 · median 6.5
12th
Delivered page size (KB)you 291 · median 176.5
8th

Percentile: the share of comparable sites this one scores better than. Lower-is-better metrics (page size, render-blocking) are already inverted.

What this scan attempted — read from its own record, not asserted

15

client identities sent, replies compared

5

machine files fetched on both hosts

114

named agents resolved from robots.txt

23

of 30 sections scored

42

invariants checked on this report itself

8

findings

A randomised path that cannot exist was also requested — the control that distinguishes your website from a wall in front of it.

What to fix, in order

Findings first: each one holds the grade down until it is gone (the top one sets the letter). Then headroom: each one lifts the score by the points shown. Open a card for what we saw, why it matters and how to fix it.

Measured 13 times, the oldest finding here first seen 2026-08-20. 8 of them have been here on an earlier reading.

Of 5 tagged findings, 2 are setting changes, 1 is a content edit and 2 need a developer. 3 findings are not tagged. Setting changes are usually the same afternoon; nothing here is ranked by effort, only labelled.

1robots.txt rules do not apply to named agents40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignoresMediumSetting

first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

40 named User-agent groups (40 agents) do not repeat 8 of the Disallow rules the `*` group carries. A crawler obeys only its own most-specific group and ignores `*` entirely, so those rules do not apply to any of the named agents (some declare their own, different rules). Worst group: `ai2bot` bypasses /wp-admin/, /cgi-bin/, /wp-includes/, /xmlrpc.php, /wp-content/plugins/, /wp-content/cache/, /trackback/, /comments/. This includes a SEARCH crawler (applebot) — the bypassed paths are being crawled and can enter the index. /robots.txt

Why it matters

A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.

How to fix it

Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.

Setting: a setting on the server, CDN, DNS or a plugin — no code and no writing

Why this was decided
RuleROBOTS_RULES_SHADOWED · rule pack sv26 · scope core
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:75c45dfa8412ab147a92 · finding f1:cbfcabeaf2c8fae39c15 resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck2004,425c18df974780aalready_held

21 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

ROBOTS_RULES_SHADOWED · full transcript

2Two business records on this page claim the same identity1 identity collision in this page's structured data: two nodes share the same url.MediumDeveloper

first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

1 identity collision in this page's structured data: two nodes share the same url.

Why it matters

A system resolving which business this page is about has more than one candidate for one identity. Two LocalBusiness or Organization nodes sharing a phone number, an @id or a coordinate pair give a retrieval system no way to decide which record is the entity, so it may pick either or neither. This reads only what the page declares - it does not assert the business is duplicated or wrongly located.

How to fix it

Merge the duplicate business nodes into one, keep a single @id, and remove any plugin that emits a second empty LocalBusiness record.

Developer: a template, theme or application change — a developer touches it

Why this was decided
RuleENTITY_COLLISION · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:f3f323b98a898391d2f1 · finding f1:94cc84cf5ff4efe5d75b resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck200298,0649a8581d03483kept
anon200298,064672b7a289653kept
gptbot200298,064672b7a289653kept
claudebot200298,064672b7a289653kept
oaisearch200298,064672b7a289653kept
claudesearch200298,064672b7a289653kept
perplexity200298,0649a8581d03483kept
googlebot200298,0649a8581d03483kept

14 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

ENTITY_COLLISION · full transcript

3The homepage depends on files this site's own robots.txt disallows8 render resources referenced by the homepage are disallowed to every crawler by this site's own robots.txt: /wp-content/plugins/dfm-trust-indicators/dist/css/sMediumSetting

first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

8 render resources referenced by the homepage are disallowed to every crawler by this site's own robots.txt: /wp-content/plugins/dfm-trust-indicators/dist/css/style.min.css (/wp-content/plugins/), /wp-content/plugins/mng-digisubs/static/mng-digisubs.styles.css (/wp-content/plugins/), /wp-content/plugins/loader-wp/static/loader.min.js (/wp-content/plugins/), /wp-content/plugins/mng-digisubs/static/mng-digisubs.sophi.bundle.js (/wp-content/plugins/). A crawler that obeys robots.txt will not fetch these, so it renders the page without them. /robots.txt

Why it matters

Googlebot renders pages, and a stylesheet or script it is forbidden to fetch is rendered without. Layout, visibility and any content a script inserts are then judged from a page the site never intended to show. Only rules in the * group count here - a rule aimed at one named crawler is a narrower question.

How to fix it

Allow crawlers to fetch the CSS and JavaScript the page depends on. Remove the Disallow rules covering those paths, or serve the content without them.

Setting: a setting on the server, CDN, DNS or a plugin — no code and no writing

Why this was decided
RuleRENDER_RESOURCE_BLOCKED · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:b51d1319719ba988cbd3 · finding f1:2262efb943db36f97ff3 resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck2004,425c18df974780aalready_held

21 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

RENDER_RESOURCE_BLOCKED · full transcript

4Almost nothing a machine receives from this page is readable text4.4% of 297,605 decompressed bytes is visible text, and the largest single block is inline css at 44,011 bytesMediumDeveloper

first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

4.4% of 297,605 decompressed bytes is visible text, and the largest single block is inline css at 44,011 bytes

Why it matters

A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.

How to fix it

Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.

Developer: a template, theme or application change — a developer touches it

Why this was decided
RulePAGE_IS_MOSTLY_CODE · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:304cde5e6ab450ac41f5 · finding f1:f4babb4a5f2416a1db5f resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck200298,0649a8581d03483kept
anon200298,064672b7a289653kept
gptbot200298,064672b7a289653kept
claudebot200298,064672b7a289653kept
oaisearch200298,064672b7a289653kept
claudesearch200298,064672b7a289653kept
perplexity200298,0649a8581d03483kept
googlebot200298,0649a8581d03483kept

14 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

PAGE_IS_MOSTLY_CODE · full transcript

5The published coordinates are not in the country the address namesThis page declares latitude 39.7392, longitude 104.9903 beside an address in US (read from addressRegion). negating the longitude puts it inside US - a missing Low

first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

This page declares latitude 39.7392, longitude 104.9903 beside an address in US (read from addressRegion). negating the longitude puts it inside US - a missing minus sign.

Why it matters

Two facts in the same record disagree about where the business is. Anything that trusts the coordinate over the address is sent to the wrong place, and a range check cannot catch it because the value is a legal longitude. Nobody proofreads a number that never renders on the page.

How to fix it

Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.

Why this was decided
RuleGEO_COUNTRY_MISMATCH · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:2d023163a3d6e24168d5 · finding f1:b0dfa22e680744be8515 resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck200298,0649a8581d03483kept
anon200298,064672b7a289653kept
gptbot200298,064672b7a289653kept
claudebot200298,064672b7a289653kept
oaisearch200298,064672b7a289653kept
claudesearch200298,064672b7a289653kept
perplexity200298,0649a8581d03483kept
googlebot200298,0649a8581d03483kept

14 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

GEO_COUNTRY_MISMATCH · full transcript

6No llms.txtNot present.NoteContent

first measured 2026-08-20 · 35 days ago · seen across 7 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

Not present. /llms.txt

Why it matters

Not a defect. llms.txt is an emerging convention for telling AI systems what a site is and which pages matter.

How to fix it

Publish /llms.txt: a short Markdown file naming the business, what it does, and the pages worth reading, served as text/plain or text/markdown.

Content: words, markup or images on a page — no deploy

Why this was decided
RuleNO_LLMS_TXT · rule pack sv26 · scope core
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:7c3893305bf52fff4f3f · finding f1:02b589098746b639fa40 resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck40488,151ea92252b1c26kept

21 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

NO_LLMS_TXT · full transcript

7The machine files cost more than the pageThe machine files a crawler reads before the page total 888,560 bytes, 2.98x the homepage HTML (298,064 B); the heaviest are /sitemap.xml 884,135 B, /robots.txtNote

first measured 2026-09-16 · 8 days ago · seen across 5 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

The machine files a crawler reads before the page total 888,560 bytes, 2.98x the homepage HTML (298,064 B); the heaviest are /sitemap.xml 884,135 B, /robots.txt 4,425 B. An absent agent file answers 404 with 88,202 bytes of HTML on average (/agents.md, /.well-known/agent-skills/index.json, /.well-known/mcp/server-card.json, /.well-known/api-catalog, /.well-known/media-kit.json) - a crawler probing for a file that is not there pays for a page each time; a short 404 body costs nothing to serve.

Why it matters

A crawler reads robots.txt, the sitemap, llms.txt and the agent files before the first page, and pays for every 404 that answers with a full page. Keep machine files as small as their job allows and serve absent paths with a short 404 body.

How to fix it

Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.

Why this was decided
RuleMACHINE_CHAIN_HEAVY · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:de8bab803db4c75a3403 · finding f1:5f60febd53b8621e833e resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck200298,0649a8581d03483kept
anon200298,064672b7a289653kept
gptbot200298,064672b7a289653kept
claudebot200298,064672b7a289653kept
oaisearch200298,064672b7a289653kept
claudesearch200298,064672b7a289653kept
perplexity200298,0649a8581d03483kept
googlebot200298,0649a8581d03483kept

14 more in the manifest.

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

MACHINE_CHAIN_HEAVY · full transcript

84 off-site listings carry this phone number and the site never links themSearched YellowPages, DexKnows, BBB by (303) 954-1010. A listing the site does not declare is one an engine corroborates without you: whatever it prints becomesNote

first measured 2026-09-10 · 14 days ago · seen across 6 measured days. A one-off scan cannot produce this line; it comes from the record.

What we saw

Searched YellowPages, DexKnows, BBB by (303) 954-1010. A listing the site does not declare is one an engine corroborates without you: whatever it prints becomes your name, hours and category in the answer.

Why it matters

Declare the ones you own in sameAs so the entity graph and the directory agree, and correct or claim the ones you do not.

How to fix it

Read the evidence below, then change what the server sends at that path. Re-scan to confirm: done is measured, not asserted.

Why this was decided
RuleNAP_UNDECLARED_LISTING · rule pack sv26 · scope route
Evaluated2026-09-24 19:13:40 UTC
Decisiond1:31cf56f1c3ab24725d3f · finding f1:a3b71814478a0c1117bf resolve either at /api/graph/<id> with a key
Observercrawlcheck-worker build fb4f0b7e from ATL · audit v4

What it rests on

IdentityStatusBytessha256Kept
crawlcheck200298,0649a8581d03483kept
anon200298,064672b7a289653kept
gptbot200298,064672b7a289653kept
claudebot200298,064672b7a289653kept
oaisearch200298,064672b7a289653kept
claudesearch200298,064672b7a289653kept
perplexity200298,0649a8581d03483kept
googlebot200298,0649a8581d03483kept

14 more in the manifest.

Where each value was read

FieldValueRead from
nameThe Denver PostJSON-LD LocalBusiness
phone+13039541010JSON-LD LocalBusiness
address5990 Washington St., Denver, CO, 80216JSON-LD LocalBusiness
hoursMo 8AM-4PM; Tu 8AM-4PM; We 8AM-4PM; Th 8AM-4PM; Fr 8AM-4PMJSON-LD LocalBusiness

Artifacts: manifest:eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89

NAP_UNDECLARED_LISTING · full transcript

Headroom +20 points if every row passes

Nothing here is broken. Each row is a measure that currently fails its optimal range, and what fixing it is worth to the score.

+1.8Answer engines allowednow 8 of 17 · Crawler permissions+1.8Search indexes allowednow 16 of 24 · Crawler permissions+1.3Cost of a missnow 3 missing file(s) return >20KB of HTML · Machine files+1.3entitymap.jsonnow absent (404) · Machine files+1.2Lead paragraph defines the subjectnow does not name it, no defining verb · Quotable content+0.9Content rationow 4.4% · Page weight & text+0.9Deferred vs blocking scriptsnow 2 deferred / 14 blocking · Page weight & text+0.9Inline JavaScriptnow 38,121B · Page weight & text+0.9Render-blocking resourcesnow 24 · Page weight & text+0.8Stated robots policy matches actual behaviournow OAI-SearchBot: disallowed but served, Claude-SearchBot: disallowed but served, PerplexityBot: disallowed but served, Applebot: disallowed but served · What each crawler gets+0.7Cumulative Layout Shift (p75)now 0.13 · Speed+0.7Largest Contentful Paint (p75)now 3.45s · Speed+0.6A named person is declarednow none · Authors & trust signals+0.6About page linked from the homepagenow no · Authors & trust signals+0.6Organization declares who runs itnow 0 of 2 · Authors & trust signals+0.5Credential edgesnow 0 · Structured data+0.5Edge / CDNnow none detected · Server & hosting+0.5H1 countnow 2 · Structured data+0.5Stable @id coveragenow 33.3% · Structured data+0.5Width and height declarednow 0 of 41 · Image signals+0.5sameAs pointing at a placenow 1 · Structured data+0.4Form controls have a namenow 1 of 2 (missing: checkbox) · Accessibility structure+0.4Found off-site, not declared in sameAsnow 4 — which ones, with a licence · Local presence & NAP+0.4Trailing slashnow 1 target linked both with and without a slash (/contact-us) · URL naming

Watch this domain

Twice-daily scans, the date each finding first appeared, and a change receipt when something moves. Free for 30 days, no card.

Start watching denverpost.com

Get this fixed for you

Every finding above fixed on your site, then re-measured — done is measured, not asserted. One-time, $749 for one local-business site.

Get this fixed — $749

What that covers

More detail — headroom arithmetic, priorities and clusters

Optimisation headroom

80 now → 100 with the 24 failing measure(s) fixed — a gain of 20 points. Each number below is the weighted contribution of that one row to the overall score; it is arithmetic on measurements already taken, not a forecast. The LETTER is separately capped at C by finding severity, so clearing that finding lifts the grade independently of the score. 128 measure(s) could not be measured; they are excluded, not counted against the site, and are not headroom.

Fix thisSectionCurrentlyPoints
Answer engines allowedAI and search agents8 of 17+1.8
Search indexes allowedAI and search agents16 of 24+1.8
Cost of a missMachine layer3 missing file(s) return >20KB of HTML+1.3
entitymap.jsonMachine layerabsent (404)+1.3
Lead paragraph defines the subjectQuotable contentdoes not name it, no defining verb+1.2
Content ratioPayload and render path4.4%+0.9
Deferred vs blocking scriptsPayload and render path2 deferred / 14 blocking+0.9
Inline JavaScriptPayload and render path38,121B+0.9
Render-blocking resourcesPayload and render path24+0.9
Stated robots policy matches actual behaviourWhat each agent receivesOAI-SearchBot: disallowed but served, Claude-SearchBot: disallowed but served, PerplexityBot: disallowed but served, Applebot: disallowed but served+0.8
Cumulative Layout Shift (p75)Speed a crawler sees, and field data where it exists0.13+0.7
Largest Contentful Paint (p75)Speed a crawler sees, and field data where it exists3.45s+0.7
A named person is declaredAuthorship and trust signals (E-E-A-T proxies)none+0.6
About page linked from the homepageAuthorship and trust signals (E-E-A-T proxies)no+0.6
Organization declares who runs itAuthorship and trust signals (E-E-A-T proxies)0 of 2+0.6
Credential edgesSchema and entity graph0+0.5
Edge / CDNServer and hostingnone detected+0.5
H1 countSchema and entity graph2+0.5
Stable @id coverageSchema and entity graph33.3%+0.5
Width and height declaredImage signals0 of 41+0.5
sameAs pointing at a placeSchema and entity graph1+0.5
Form controls have a nameAccessibility structure (from the delivered HTML)1 of 2 (missing: checkbox)+0.4
Found off-site, not declared in sameAsLocal presence, citations and NAP4 — which ones, with a licence+0.4
Trailing slashNaming and case consistency1 target linked both with and without a slash (/contact-us)+0.4

What to fix first

Ordered by contextual priority, not by severity alone. Two findings of the same severity are not equally urgent — one that breaks a file every agent reads before anything else outranks one on a single page. Every multiplier below is derived from this scan and shown with the reason it applied.

5.8

Two business records on this page claim the same identity

/ · severity 3

Prerequisite impact×1.1affects the homepage
Scope×1not an access-class finding - scope not multiplied
Persistence×1.75open for 39 days, first seen 2026-08-20
Confidence×1the bytes behind this finding are attached to it

5.8

Almost nothing a machine receives from this page is readable text

/ · severity 3

Prerequisite impact×1.1affects the homepage
Scope×1not an access-class finding - scope not multiplied
Persistence×1.75open for 39 days, first seen 2026-08-20
Confidence×1the bytes behind this finding are attached to it

5.6

robots.txt rules do not apply to named agents

/robots.txt · severity 3

Prerequisite impact×1.5/robots.txt is a discovery prerequisite - agents read it before anything else
Scope×1not an access-class finding - scope not multiplied
Persistence×1.25confirmed on 12 scans of this domain since 2026-09-10
Confidence×1the bytes behind this finding are attached to it

5.6

The homepage depends on files this site's own robots.txt disallows

/robots.txt · severity 3

Prerequisite impact×1.5/robots.txt is a discovery prerequisite - agents read it before anything else
Scope×1one crawler identity affected
Persistence×1.25confirmed on 12 scans of this domain since 2026-09-10
Confidence×1the bytes behind this finding are attached to it

2.8

The published coordinates are not in the country the address names

/ · severity 2

Prerequisite impact×1.1affects the homepage
Scope×1not an access-class finding - scope not multiplied
Persistence×1.25confirmed on 10 scans of this domain since 2026-09-10
Confidence×1the bytes behind this finding are attached to it

2.6

No llms.txt

/llms.txt · severity 1

Prerequisite impact×1.5/llms.txt is a discovery prerequisite - agents read it before anything else
Scope×1not an access-class finding - scope not multiplied
Persistence×1.75open for 39 days, first seen 2026-08-20
Confidence×1derived from the response itself

1.4

The machine files cost more than the page

/ · severity 1

Prerequisite impact×1.1affects the homepage
Scope×1not an access-class finding - scope not multiplied
Persistence×1.25confirmed on 8 scans of this domain since 2026-09-16
Confidence×1the bytes behind this finding are attached to it

1.4

4 off-site listings carry this phone number and the site never links them

/ · severity 1

Prerequisite impact×1.1affects the homepage
Scope×1not an access-class finding - scope not multiplied
Persistence×1.25confirmed on 11 scans of this domain since 2026-09-10
Confidence×1the bytes behind this finding are attached to it

Priority = severity × prerequisite impact × scope × persistence × confidence. A factor the record cannot support stays at ×1 rather than being guessed upward.

These findings share a cause

Findings are symptoms. Where two or more point at one underlying cause, they are grouped here — fixing the cause closes all of them. A single finding is never presented as a cluster, because one symptom does not establish a diagnosis.

The site's own robots.txt withholds what its pages need to render

ROBOTS_RULES_SHADOWED RENDER_RESOURCE_BLOCKED · highest priority in group 5.6

A stylesheet or script the homepage depends on sits under a path the * group disallows. A rendering crawler obeys the rule and draws the page without it.

What to check: Allow the asset paths the page references, or move the assets out of the disallowed directory. Only the * group is in scope; a rule aimed at one named crawler is a separate question.

Every finding also appears on its own below, with its full evidence. Grouping changes the order of work, not the record.

What each crawler was served

14 requests for the same URL, from one CrawlCheck address, at 2026-09-24 19:13 UTC: each named crawler, a mobile browser and an unnamed client. Robots.txt is what the site asks for; the status and word count are what its server actually sent.

Requested asRobots.txtServer sentWordsSame as browser
Unnamed client (control)—2002,036100%baseline
GPTBotdisallowed2002,036100%disallowed but served
ClaudeBotdisallowed2002,036100%disallowed but served
OAI-SearchBotdisallowed2002,036100%disallowed but served
Claude-SearchBotdisallowed2002,036100%disallowed but served
PerplexityBotdisallowed2002,03691%disallowed but served
Googlebotallowed2002,03691%same page
Mobile browser (control)—2002,036100%baseline
Bingbotallowed2002,03691%same page
Applebotdisallowed2002,03691%disallowed but served
Amazonbotdisallowed2002,036100%disallowed but served
Bytespiderdisallowed4035—blocked by the server
Meta-ExternalAgentdisallowed2002,03691%disallowed but served
CCBotdisallowed2002,03691%disallowed but served

13

passing

7

need work

3

failing

7

information only

Sections, not individual checks: 30 in all, 23 scored. A scored section that could not be read is counted separately and never as a pass; information-only sections are measured and never scored.

Which crawlers are let in

73 of 114 named crawlers are allowed by robots.txt at /; 41 are disallowed. Robots.txt is a request; the edge decides. Where this scan also fetched the page as that crawler, the chip says what the edge did.

Answer engines 8/17 allowed

The crawlers behind ChatGPT, Claude, Perplexity and Google’s AI answers

✗ChatGPT search indexrules shadowed✗ChatGPT live fetchrules shadowed✗Claude search indexrules shadowed✗Claude live fetchrules shadowed✗Perplexity indexrules shadowed✗Perplexity live fetchrules shadowed✓Gemini research agent✗Le Chat live fetchrules shadowed✗DuckDuckGo AI assistrules shadowed✗You.comrules shadowed✓Phind✓Kagi✓Microsoft Copilot fetch✓Meta AI live fetch✓Amazon live fetch for Alexa questions✓Google user-triggered agent, acts on the web✓Gemini Notebook user-supplied URL fetch

Search engines 16/24 allowed

Google, Bing, Apple and the rest of classic search

✓Google image crawler✓Google News✓Google video crawler✓Google Shopping✓Google AdSense✓Google push delivery✓Bing page preview✓Microsoft legacy crawler✗Yahoo Slurprules shadowed✗Yandex (umbrella token)✓Coc Coc✓Google Search and AI Overviews✓Bing and Copilot index✗Apple and Sirirules shadowed✗Amazonrules shadowed✓DuckDuckGo✗Yandexrules shadowed✗Baidurules shadowed✓Seznam✓Neeva✗Huawei Petalrules shadowed✗Amazon search eligibility, separate from Amazonbotrules shadowed✓Mistral search index for Vibe✓Exa AI search index, Web Bot Auth signed

Training crawlers 22/46 allowed

Collect pages to train models. Blocking these while allowing answer engines is coherent policy, not a defect

Show all 46
✓Magpie AI✓img2dataset image corpus✓Awario RSS✓Awario smart✓Turnitin✗Internet Archiverules shadowed✗Internet Archive (legacy)✓Meta web index✗Webz.io omgilirules shadowed✗Cohere training corpusrules shadowed✗Huawei PanGurules shadowed✗Allen Institute Dolmarules shadowed✓FriendlyCrawler ML✗Velenrules shadowed✓MyCentral AI✓DeepSeek✗NICT ICCrules shadowed✓Google non-search fetch✓Vertex AI agent build✗OpenAI training and indexrules shadowed✗Anthropic trainingrules shadowed✗Anthropic legacy agentrules shadowed✗Anthropic legacy agentrules shadowed✗Gemini trainingrules shadowed✗Apple trainingrules shadowed✗Common Crawlrules shadowed✗ByteDancerules shadowed✗Meta AI trainingrules shadowed✗Meta legacyrules shadowed✓Cohere✗Diffbot knowledge graphrules shadowed✗Webz.io✓Imagesift✗Timpirules shadowed✗Allen Instituterules shadowed✓Generic scraper framework✓Semrush AI corpus✗Apple ads corpus✓TikTok✓QuillBot✗Webz.io extendedrules shadowed✓ProRata✓Awario✓Mistral training crawl✓Google common crawler, public image URLs✓Google common crawler, public video URLs

Some named groups declare no rules of their own, which means “allow everything” for that crawler regardless of the general rules. Those chips are marked.

How crawlers move through this site

4 of 14 identities reach the page. 9 are turned away by the site’s own robots rules — a policy, not a fault. 1 is refused at the edge before any rule applies, for a request carrying the name from our address — an edge that admits crawlers by their published IP ranges may still let the real one through. Missing: entitymap.json, llms.txt. Of the 12 others that got through, 1 received exactly the browser’s bytes, 5 the same text in different bytes, 6 different text.

WHO ASKEDEDGEFILES FETCHEDDERIVEDPROCESSEDENGINEUnnamed client✓GPTBot⊘ClaudeBot⊘OAI-SearchBot⊘Claude-SearchBot⊘PerplexityBot⊘Googlebot✓Mobile browser✓Bingbot✓Applebot⊘Amazonbot⊘Bytespider✗Meta-ExternalAgent⊘CCBot⊘entitymap.json (JSON) — absentllms.txt (text) — absentpage HTML — openrobots.txt (text/plain) — opensitemap (XML) — openname, phone, address — openanswer text — unmeasureddeclared entities — absenttyped nodes — openenquiry — openidentity to corroborate — openembedded JSON-LD — openforms and call links — openvisible text — openquotable units — openrules per user-agent — openEdgedecision3 distinct response bodies served to the identities that were answered, by SHA-2563 versionsservedrobots.txt — present · 4.3 KB · sha256 c18df974780af520cf061caf636c9602ae711d8ab61a5c6a50fc40fb3a1e6a1d · keptrobots.txtpresent · 4.3 KBsitemap — present · 863.4 KB · sha256 e1a6b687dda2b279862b271def73006b6244a1699bf9cd61bbadfee6b17a7783 · keptsitemappresent · 863.4 KBllms.txt — missing · 404 · sha256 ea92252b1c26011e275db7d2bddcdfce4e341e83b82e5d8640f2c367971a02f3 · keptllms.txtmissing · 404entitymap.json — missing · 404 · sha256 44962592a38391cf86c8c78a305e1a4ac10f2132cdfb38290e922881db38bd6a · keptentitymap.jsonmissing · 404Page HTML — present · 291.1 KB · sha256 e1063e72c11937d7a0b253cfb47a053547b27b272adf666108415a9b2f1892bd · keptPage HTMLpresent · 291.1 KBJSON-LD — presentJSON-LDpresentDeclared identity (NAP) — presentIdentity (NAP)presentQuotable text — presentQuotable textpresentEnquiry path — presentEnquiry pathpresentrobots.txt resolutionrobots rulesParse and entity graphEntity graphName, address, phone corroborationNAP checkQuote extractionQuote extractionEnquiry (form or call)EnquiryAnswer engine output — not measuredEngineoutput notmeasuredEVIDENCEevery response above, kept by digest and sealedBytes kept — 10 distinct objects, 2.1 MB, each stored under its own SHA-256Bytes kept22 of 22 responsesManifest — sha256 eb69e87d5daf467ae8c9018f0bd59734b51eec9371d7ef25fc6f8dcb68430b89: every response above, listed by digestManifesteb69e87d5daf…Record digest — the manifest’s hash is part of this record’s digestRecord digestfixed at the sealDay root — the day is sealed after it closes (UTC); this digest is already recordedDay rootseals after 09-24Bitcoin — the root goes to the OpenTimestamps calendars at the seal; a Bitcoin block follows, usually within hoursBitcoinafter the seal

Who asked

  • Unnamed clientreaches the page · same text, different bytes
  • GPTBotturned away by the site's robots rules · same text, different bytes
  • ClaudeBotturned away by the site's robots rules · same text, different bytes
  • OAI-SearchBotturned away by the site's robots rules · same text, different bytes
  • Claude-SearchBotturned away by the site's robots rules · same text, different bytes
  • PerplexityBotturned away by the site's robots rules · different text (91% of sentences shared)
  • Googlebotreaches the page · different text (91% of sentences shared)
  • Mobile browserreaches the page
  • Bingbotreaches the page · different text (91% of sentences shared)
  • Applebotturned away by the site's robots rules · different text (91% of sentences shared)
  • Amazonbotturned away by the site's robots rules · identical bytes
  • Bytespiderrefused at the edge (HTTP 403) to our request
  • Meta-ExternalAgentturned away by the site's robots rules · different text (91% of sentences shared)
  • CCBotturned away by the site's robots rules · different text (91% of sentences shared)

Files fetched

  • robots.txtpresent · 4.3 KB
  • sitemappresent · 863.4 KB
  • llms.txtmissing
  • entitymap.jsonmissing
  • Page HTMLpresent · 291.1 KB

Derived

  • JSON-LDpresent
  • Quotable textpresent
  • Enquiry pathpresent
  • Declared identity (NAP)present

Answer engine output — not measured. No server can observe what an engine says.

Evidence

  • Bytes kept22 of 22 responses
  • Manifesteb69e87d5daf…
  • Record digestfixed at the seal
  • Day rootseals after 09-24
  • Bitcoinafter the seal

✓ reached · ⊘ turned away by the site’s own rules · ✗ technical break · dashed box = missing or not measured. Drawn from this scan’s own fetches; nothing here is estimated. Fetched 2026-09-24 19:13 UTC from Cloudflare ATL by worker version fb4f0b7e under scoring rules v26. Each identity was sent by CrawlCheck from its own network: ✓ means our request under that name got through, not that the company’s own crawler did. The flow as JSON.

What each identity received — 22 responses captured, 22 kept as evidence
Who askedHTTPSizeCompared with a browserKept as evidence
Unnamed client200291.1 KBsame text, different byteskept · 672b7a28
GPTBot200291.1 KBsame text, different byteskept · 672b7a28
ClaudeBot200291.1 KBsame text, different byteskept · 672b7a28
OAI-SearchBot200291.1 KBsame text, different byteskept · 672b7a28
Claude-SearchBot200291.1 KBsame text, different byteskept · 672b7a28
PerplexityBot200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
Googlebot200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
Mobile browser200291.1 KBthe baselinekept · e1063e72
Mobile browser (control)200291.1 KBidentical byteskept · e1063e72
Bingbot200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
Applebot200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
Amazonbot200291.1 KBidentical byteskept · e1063e72
Bytespider403146 Brefused (HTTP 403)kept · 32f2fa94
Meta-ExternalAgent200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
CCBot200291.1 KBdifferent text (91% of sentences shared)kept · 9a8581d0
CrawlCheck scanner (home page)200291.1 KB—kept · 9a8581d0
Files the scanner fetched
/robots.txt2004.3 KB—kept · c18df974
/sitemap.xml200863.4 KB—kept · e1a6b687
/sitemap_index.xml40486.1 KB—kept · bd47eaac
/wp-sitemap.xml200138.8 KB—kept · f2ff03d4
/llms.txt40486.1 KB—kept · ea92252b
/entitymap.json40486.1 KB—kept · 44962592

Sizes are each response body as the fetch returned it: transfer compression removed, before any text decoding. A “kept” link re-hashes the stored copy inside our worker and says whether it still matches its digest; the bytes themselves are not republished.

Discovery chain breaks at llms.txt: everything after it is reachable only by guessing the conventional path. Rows behind this.

Where the bytes go

Of the 297,605 decompressed bytes the homepage delivers, 4.4% is text a reader or a model can actually use. Most of what a crawler downloads here is not words.

  • Readable text 4.4% · 13,081 B
  • Markup & attributes 65.7% · 195,658 B
  • Structured data (JSON-LD) 0.8% · 2,238 B
  • Inline CSS 15.3% · 45,571 B
  • Inline JavaScript 12.8% · 38,121 B
  • Inline SVG 0.1% · 372 B
  • HTML comments 0.9% · 2,564 B

The individual blocks that weigh the most, so the fix is a search rather than a hunt:

  • Inline CSS · 44,011 B (14.8%) — unnamedopens @import url(https://fonts.googleapis.com/css2?family=Inter:ital,opsz,w
  • Inline JavaScript · 14,462 B (4.9%) — #digisubs-social-login-js-afteropens (function () { var logPrefix = '[digisubs-social-login]'; var readyTim
  • Inline JavaScript · 3,089 B (1%) — WordPress emoji scriptopens /*! This file is auto-generated */ var e="script#wp-emoji-settings",t=

Every section

23 of 30 carry weight in the grade. Grey means measured and deliberately unweighted - it counts for nothing rather than zero, and each grey card says why. Open any section for the rows behind it: what was measured, what it is now, and the range that passes.

Machine files60robots.txt, llms.txt, sitemap and the entity graph: present, readable and served as the right type scoredAgent surfaces—Emerging standards (agents.md, API catalog, MCP card). Measured for information only not scored: these are published standards at very different stages of adoption, and several do not apply to every businessServer & hosting80What the edge and origin told us: cache state, compression, the software behind the page scoredStructured data60The JSON-LD entity graph: stable @ids, connected nodes, credentials and locations declared scoredLinked data100Whether the structured data parses as proper RDF and nothing dangles scoredAuthors & trust signals57Named people, dated articles, an organisation with contact details: the proxies an engine uses for trust scoredAccessibility structure88Language, alt text, labelled controls and one main landmark, read from the delivered HTML scoredEntity corroboration100Whether the identity links the site declares point back at it, and whether a Wikidata record names this domain scoredContent entities—Named things in the page text that a knowledge graph recognises. Measured, not scored not scored: the reading is sound but the scoring rule is not settled, and a rule that may move should not move a gradeQuotable content67Whether the sentences an answer engine would lift are on the page in a form it can take: a defining lead, short factual paragraphs, answered questions, few first-person openers scoredService area map—Every place the schema declares, placed on one map around the declared centre and radius: which have a page, which fall outside the circle. Measured, not scored not scored: a declared service area can be right or wrong only against facts this scan cannot seeWhole site—Up to 60 declared pages, read once a day after a scan: not-200s, thin and mostly-code pages, missing JSON-LD, canonical elsewhere, noindex, missing or duplicate h1, orphans, undeclared pages, duplicate titles, and the median words, text share and response time across them. Every count is free; Watch names which pages. Measured, not scored not scored: the grade stays anchored on the homepage so it is comparable across sites and daysFreshness—Every date the site publishes about itself: sitemap lastmod, Last-Modified, dateModified, dated elements, and where they disagree. Measured, not scored not scored: every date here is one the site published, so this reports what it claims rather than what is true. Calibrated 2026-09-16 on 114 records: the only row most sites carry (a sitemap that declares lastmod, 68%) is not a defect when absent, and the rows that would be defects (every lastmod stamped on one day, 8%; the sitemap and dateModified disagreeing by over a month, 2%) can be measured on too few sites to carry weightCrawl waste100Fetches that produce no new page: parameter and case variants, http links, redirecting or gone declared URLs, canonical-elsewhere and noindex pages. Scored from version 23 scoredPage subject100Which business entity the page is about and how every other business entity on it is related: chain, subsidiary, department, or undeclared. Scored from version 23 scoredLocal intents—Which of the 446 recovered Maps intent types the site names about itself, and which of those have a page the homepage links to. Measured, not scored not scored: the intent vocabulary is recovered names with no weights, and a claimed intent without a page is a content decision, not a defect; reported until the corpus says how it distributesConversion path—Phone link, form and call-to-action on the page, and whether they sit before the fold. Measured, not scored not scored: whether a visitor can act is a business judgement, and the grade is for what a machine receivesCrawler permissions33What robots.txt actually tells each named AI and search crawler it may do scoredSpeed67How fast the page answered our fetch, plus real-visitor field data where Google has it scoredPage weight & text20How many of the delivered bytes are readable words, and how much is markup, scripts and styling scoredMachine-file chain100robots.txt names the sitemap, llms.txt points onward, the entity graph points home: the chain a crawler follows scoredClaim consistency100Name, phone, city and years-in-business match between the structured data and the visible page scoredInstructions for agents100Text aimed at AI agents rather than people: present, visible, and not hidden scoredWhat each crawler gets88Fifteen client identities fetched the page; this is whether they all got the same thing scoredMobile100Viewport, zoom, icons and manifest: whether a phone gets a page it can use scoredLocal presence & NAP92Street address, coordinates, hours, phone, and the directory listings the site itself declares scoredSitemap100A sitemap exists, resolves, and declares the pages the homepage links to scoredInternal links100Anchors carry words, targets resolve, and the link graph is not mostly repeats scoredURL naming80Lowercase hosts and paths, no case-only duplicates, machine files spelt correctly scoredImage signals80Alt text, declared dimensions, lazy-loading, modern formats, hero priority and the entity image in JSON-LD. scored

What the site states about itself

The facts an answer engine could repeat about this business, taken only from what this scan fetched, with where each was found. Anything absent is [not stated]: an engine asked for it would have to guess.

FactStated asFound inAgreementFirst seen
NameThe Denver Post2026-09-19
Phone(303) 954-1010structured data onlynot shown to readers2026-09-19
Address5990 Washington St., Denver, CO, 802162026-09-19
Map point39.7392, 104.9903structured datanegating the longitude puts it inside US - a missing minus sign2026-09-19
Opening hoursstatedstructured data2026-09-19
Business typeLocalBusinessstructured data2026-09-19
Profiles it links3 (Facebook, X, LinkedIn)links on the site3 of 3 point back

Everything else that was measured

Measured for information, not scored

Section detail

Each row: the measure, what it is now, the range that passes, and whether this site is inside it. n/a means the input could not be measured and the row is excluded, not failed.

Machine layer — 60/100
Section score 60 / 100
MeasureCurrentlyOptimal
robots.txt
a crawler that cannot read robots.txt often treats that as disallow-everything
200 · 4425B200, text/plain, 100–4,000B✓
XML sitemap
without it a crawler enumerates the site by following links only
live200, valid XML, at a path robots.txt names✓
llms.txt
reported, not scored. Some agents fetch it and it costs nothing to publish, so the row stays. It is not graded because nothing we have measured, and nothing we have read, links its presence to a citation or a ranking — a study across roughly 300,000 domains found no correlation, and dropping the variable made the model more accurate, not less. Grading a file we cannot show works would price it for our customers on our say-so
absent (404)200, text/plain, 2,000–50,000Bn/a
entitymap.json
machine-readable entity layer; optional, but it is the strongest identity signal available
absent (404)200, application/json✗
How rare is what this site publishes
adoption counts are third-party aggregates read in August 2026 and count sites in that index rather than the whole web; they are here to size the opportunity, not to score the site
llms.txt — roughly 70,200 sites publish one; this is not among them. entitymap.json — absent, and no third-party adoption figure exists for it. for scale, roughly 677,748 sites publish something built for agents at all, and about 107,200 now disallow GPTBot outright, so an AI policy is mainstream while a machine-readable map of the site is not.publish what almost nobody publishes, and keep it correctn/a
Host agreement
a crawler resolving www must not get a different machine layer
hosts agree0 divergent files✓
Cost of a miss
every absent file bills the crawler for a full HTML 404
3 missing file(s) return >20KB of HTMLa missing machine file should 404 small, not ship a full page✗

Hosts and URLs

/robots.txtwww.denverpost.com: 200 · 4,425B live
denverpost.com: 301 → https://www.denverpost.com/robots.txt
/llms.txtwww.denverpost.com: 404 · 88,151B of HTML (not the file)
denverpost.com: 301 → https://www.denverpost.com/llms.txt
/entitymap.jsonwww.denverpost.com: 404 · 88,169B of HTML (not the file)
denverpost.com: 301 → https://www.denverpost.com/entitymap.json
/sitemap.xmlwww.denverpost.com: 200 · 884,135B live
denverpost.com: 200 · 884,135B live
/sitemap_index.xmlwww.denverpost.com: 404 · 88,178B of HTML (not the file)
denverpost.com: 301 → https://www.denverpost.com/sitemap_index.xml

Scanned on www.denverpost.com. denverpost.com 301s to www.denverpost.com, so that is the host a crawler ends up reading. Probed with redirects not followed, so a 200 here means the file is served at that exact URL — a followed redirect would report a file that lives somewhere else. The denverpost.com host serves its own machine layer. A crawler that resolves that hostname reads those files, not the ones on www.denverpost.com.

Server and hosting — 80/100
Section score 80 / 100
MeasureCurrentlyOptimal
Compression
our fetch decompresses the body and drops the content-encoding header before we can read it — this is our limitation, not a finding about the site. Check it in a browser’s network panel, or with curl --compressed -I
not visible from herebr preferred, gzip acceptablen/a
Edge / CDN
absorbs crawler load and keeps machine files fast
none detecteda CDN in front of the origin✗
Cache state
DYNAMIC means every crawler hit reaches the origin. Shown, not scored: the speed rows below score what that costs, and scoring the mechanism too would count it twice
4 of 5 machine files DYNAMICHIT or MISS on machine files; DYNAMIC is Cloudflare's default for .txt and .xmln/a
Version disclosure
a public CMS version tells an attacker exactly which exploits to try
not disclosedno generator version in the HTML✓
Runtime disclosure
same reason as above
not disclosedno x-powered-by header✓
Vary
governs whether a cached copy is reused correctly
accept-encodingAccept-Encoding at minimum✓
Strict-Transport-Security
max-age 365 days, subdomains included · hstspreload.org: unknown. Shown, not scored: HSTS changes nothing about what crawlers receive
max-age=31536000;includeSubdomainsmax-age with includeSubDomains; preload only if you mean itn/a
Webmaster verification tags
meta tags only — DNS and file verification leave nothing in the HTML, so a missing tag is not evidence of no registration. Shown, not scored
Google no tag visible · Bing tag present · Yandex no tag visibleBing at least: its index feeds ChatGPT search and Copilotn/a
Client-side rendering
framework present, HTML delivered with text. Measured on the bytes the server sent, not on a rendered view
script bundle, no known framework · 2036 words deliveredtext present in the delivered HTML✓
ApplicationWordPress
detected from the delivered HTML, not from a header that can be removed
Varies onaccept-encoding
a response that varies can be cached differently per client
Schema and entity graph — 60/100
Section score 60 / 100
MeasureCurrentlyOptimal
Entity nodes
with no schema the site has no declared identity at all
3at least 3 typed nodes✓
Stable @id coverage
without an @id every page declares a brand-new unrelated entity
33.3%80–100%✗
Credential edges
the machine-readable form of why this business is qualified; a publisher, a shop or a SaaS is not held to it
0at least 1 (licence, membership or award) for a local or professional service✗
Locations missing a street
a service-area business legitimately omits this — not scored either way
00 for a premises business; expected for service-area✓
sameAs claims
sameAs asserts identity, so a WRONG target claims the business is something else. A large number of correct profiles is not a defect; a single target that does not lead back here is.
3 (3 corroborated, 0 not linking back, 0 dead)at least 2, none dead, at least one linking back✓
sameAs pointing at a place
a map link in sameAs claims the business IS that location
10✗
H1 count
more than one h1 leaves no single subject for the page
2exactly 1✗
Heading skips
a skipped level breaks the document outline a parser builds
00✓
Canonical
without it duplicates compete with each other
presentpresent on every page✓
Meta description
over 160 it is truncated; under 50 it is a label, not a description. 95 chars is fine
157 chars50–160 chars✓
Entities declared3 nodes, 2 distinct types
Stable identity1 of 3 carry an @id (33.3%)
Without an @id a node cannot be referenced or merged — every page re-declares a new, unrelated entity.
Credential edges0
Licences, memberships and awards expressed as graph edges rather than prose.
Topic edges0
A high count with generic values usually means the graph is being used to store keywords.
Locations1 declared, 0 without a street address
sameAs claims3 total · 1 map links · 0 resolve to this entity · 1 resolve to a place
sameAs asserts identity. A map link pointing at a city claims the business is that city. No validator reports this — it only appears if the targets are resolved.
Document structure2 h1 · 120 headings · 0 level skips · canonical present · meta description 157 chars
RDF and linked-data surface — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
JSON-LD blocks that parse
a block that does not parse contributes no triples and looks, to a validator that counts tags, like structured data that is there
2 of 2every block✓
@context is schema.org
a graph in another vocabulary is valid RDF that no answer-engine resolver is built to read
2 of 2every block✓
Triples declared (approximate)
the size of what a linked-data reader actually receives; a single Organization with a name and a URL is three statements
68at least 20 on a homepage✓
Dangling @id references
a reference to an @id no node declares is a statement about nothing; the resolver drops the edge and the entity loses that relation
00✓
Top-level nodes carrying a subject IRI
a root or @graph node without an @id is a blank subject that no other page or site can reference; nested values (an address, an offer, a list item) are not held to this
0 of 2 (12 nodes in all)every entity the page declares — shown; the schema section scores @id coveragen/a
RDFa and Microdata
Secondary serialisations are read by some parsers and are a second place for the facts to disagree. Not scored while JSON-LD is present; when a page declares no JSON-LD node, its schema.org Microdata and RDFa are read as its graph and every entity check runs on them
no RDFa, no Microdata, 23 Open Graph metaoptional; one serialisation kept accurate beats three that driftn/a
RDF content negotiation
An RDF-aware client sent Accept: text/turtle. Almost no commercial site answers with a graph and absence is not a defect; a site that does is handing the machine the same facts in a tenth of the bytes. Not scored
text/html served (HTTP 200), Vary: Acceptoptional: text/turtle or application/ld+json to a client that asks for itn/a
Alternate RDF representation linked
Tells a linked-data client where the graph lives without negotiating for it. Not scored
noneoptional: <link rel=alternate type=text/turtle> or a Link headern/a

JSON-LD is an RDF serialisation, so every schema block above is already a graph of triples. This section reads it as one: 68 statements about 0 identified subjects, 1 reference between them, none dangling. The triple count is approximate: each property value is one statement and each @type one more; nested objects without an @id are counted as values, not expanded. Weight 0.8 in the overall, below the schema section it overlaps.

Authorship and trust signals (E-E-A-T proxies) — 57/100
Section score 57 / 100
MeasureCurrentlyOptimal
A named person is declared
an evaluator asking who is behind this finds a machine-readable answer or nothing; a site with no Person node is anonymous to a resolver
noneat least 1 Person node with a name✗
Person carries expertise properties
jobTitle, hasCredential, alumniOf, memberOf, knowsAbout or award: the statements that say WHY this person is qualified, in the form a reader can check
—every named personn/a
Person is corroborated elsewhere (sameAs)
a person who exists only on this site is a claim; a person with profiles that link back is a fact someone else asserts too
—every named personn/a
Person has a page of their own
the page an evaluator lands on to judge experience; without it the person is a name in a script tag
—a url or mainEntityOfPage on eachn/a
Articles name a typed author
author as a string is a label; author as a Person or Organization object is an entity that can carry the expertise above
no Article nodes on this pageevery articlen/a
Articles carry datePublished and dateModified
freshness is read from dateModified; an article with only a publish date reads as never maintained
—both on every articlen/a
Organization declares who runs it
connects the business entity to the people entities; without the edge the two graphs never meet
0 of 2founder, employee or member on the organisation✗
Organization declares foundingDate
the one claim behind every 'years of experience' counter that a reader can check against a registry
1 of 2on the primary organisation✓
Contact route declared in schema
a business that cannot be reached by a stated channel is unverifiable by definition
1 of 2telephone, email or contactPoint✓
About page linked from the homepage
the page an evaluator opens first; if the homepage does not link it, most readers never find it
nolinked✗
Contact page or direct contact linked
a contact path in the delivered HTML, not only in a footer image or a widget that renders later
yeslinked✓
Privacy policy or terms linked
the cheapest trust signal on the list and the one small sites most often omit
yeslinked✓
Ratings carry a value and a count
a rating with no count is a number nobody can weigh; a count with no reviews behind it is the shape of a fabricated one
no aggregateRatingratingValue with reviewCount or ratingCountn/a
Outbound references
Citations outward are part of how authority is read, and a count cannot say whether they are good ones. Not scored
18 external hosts linked (excluding social platforms): checkout.denverpost.com, accuweather.com, enewspaper.denverpost.comcite what the content relies on; there is no right numbern/a

What this can and cannot say. E-E-A-T is a judgement an evaluator makes about experience, expertise, authoritativeness and trust; it is not on the wire and this scanner does not pretend to score it. Each row above is one machine-readable statement a site makes that a reader could use as evidence toward that judgement, measured from the delivered homepage and its JSON-LD. A full row of passes does not create expertise, and a row of fails does not disprove it; what they decide is whether a machine reading this site can find the evidence at all. Weight 1.0 in the overall, the same as payload: what these rows measure is whether the evidence is findable by a machine, and that is the part this scanner can stand behind.

Accessibility structure (from the delivered HTML) — 88/100
Section score 88 / 100
MeasureCurrentlyOptimal
Document language declared
a screen reader picks its voice and pronunciation from this; without it every word is read in the user's default language
lang=en-USa lang attribute on <html>✓
Page has a title
the first thing announced when a page opens, and the label on every tab and bookmark
89 charsa non-empty <title>✓
Images carry an alt attribute
a missing alt attribute makes the reader announce the file name; alt="" marks an image as decorative and is correct
41 of 41every image (alt="" for decoration)✓
Form controls have a name
a field with no name is announced as "edit text"; a placeholder is not a label, it disappears on the first keystroke
1 of 2 (missing: checkbox)a <label for>, wrapping label, aria-label or title on each✗
Buttons have an accessible name
an icon-only button with no name is announced as "button" and nothing else
3 of 3text, aria-label, or an image with alt inside each✓
Frames are titled
a frame with no title is a hole in the reading order the user cannot identify or skip on purpose
no framesa title on every iframen/a
One main landmark
the region a reader jumps to for the content; two of them is as unusable as none
oneexactly one <main> or role=main✓
A way past the navigation
without one, every page starts by tabbing through the whole menu
skip linka skip link, or nav and main landmarks✓
No duplicate ids
labels, aria-labelledby and skip links all resolve by id; a duplicate sends them to the first match, which is usually the wrong element
0every id unique✓

What this is not. Not a WCAG audit: contrast, focus order, keyboard traps, motion and the experience of using the page with assistive technology need a rendered page and a person, and none of that is measured here. These rows read the document the server returned and ask whether the structure a screen reader depends on is present at all. Heading order, zoom lock and link text are scored in their own sections and are not counted again here. Weight 0.8 in the overall. Counts in the Read stage from scoring version 26, because parsers find content through the same structure a screen reader uses; reports issued earlier kept it outside the stages.

Entity corroboration — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
Identity claim — facebook.com
the target page references denverpost.com back
corroboratedcorroborated✓
Identity claim — x.com
the target page references denverpost.com back
corroboratedcorroborated✓
Identity claim — linkedin.com
the target page references denverpost.com back
corroboratedcorroborated✓
Wikidata Q2668654
the strongest third-party identity signal a machine can read
official website matches this domainP856 names this domainn/a

3 corroborated · 0 unreciprocated · 0 unverifiable · 0 dead. A sameAs link is an identity CLAIM, and an identity claim can be reciprocated — the target either points back at this domain or it does not. Platforms that refuse datacentre readers are counted as unverifiable, never as failures: the claim can be neither confirmed nor accused from here, and pretending otherwise would manufacture findings. Corroborated claims pass, unreciprocated and dead claims fail, and unverifiable or declared-only claims are left out of the score — so a profile page that blocks readers can never cost a point, and a link that is dead or never mentions this site does.

Quotable content — 67/100
Section score 67 / 100
MeasureCurrentlyOptimal
Lead paragraph defines the subject
the first paragraph is the one most often lifted whole; if it is a slogan, the engine has to assemble the answer from fragments. Naming the subject is shown, not scored
does not name it, no defining verba first sentence that says what this is (corpus: 30% of homepages do)✗
Paragraphs an engine can quote whole
a self-contained paragraph of 15-70 words with a fact in it is the unit answer engines extract; a 200-word paragraph gets summarised instead, and the summary is theirs
9 of 12 (75%)at least 25% (corpus median 29%)✓
Question headings with an answer under them
the shape a query has when it arrives; a page that already carries the question and a short answer is quoted in that order. Having none is not scored
no question headingsat least half, where question headings existn/a
FAQ markup matches the visible text
FAQ markup for text a visitor cannot see is a claim with nothing behind it; engines that compare the two drop the markup
no FAQPageevery declared question readable on the pagen/a
Headings carry an id
a heading with an id is a stable address for one passage; without it a citation can only point at the whole page
0 of 27an id on each, so a passage can be cited by fragment (not scored: corpus median 0%)n/a
Sentence length
quotes are short; a 40-word sentence is paraphrased, and the paraphrase carries the engine's wording, not yours
median 13 words, 63% under 26a median of 22 words or fewer (corpus median 15)n/a
Sentences with a checkable fact
8-30 words with a number, date or name, in the third person: the sentence an engine can attribute to you without editing it
5 of 8 (63%)at least 15% (corpus median 20%)n/a
Paragraphs that open in the first person
"We offer" needs rewriting before it can be quoted about you; "Acme offers" does not
0 of 12 (0%)a third or fewer (corpus p90 24%)✓
Tables with a header row and at least two data rows
a table is the one structure every extractor reads the same way; the same facts in prose are re-assembled differently by each engine
0one per set of comparable facts (not scored: rare on homepages)n/a
A visible updated or published date
engines weigh freshness from the text as well as from schema; a date the reader can see is the one they trust
none founda dated line in the copy (not scored: 4% of homepages carry one)n/a

Sentences an engine could lift as they stand:

  • “Here’s what to expect in the final 3 years.”
  • “Much of the I-70 Floyd Hill Project is taking place above drivers’ heads.”
  • “Here’s what to expect in the final 3 years.”

The lead, as delivered: “Sign up for Newsletters and Alerts Sign Up”

What this scores, and what it is not. These rows read the delivered HTML for the shape of quotable text: nothing here judges whether the copy is good, true or wanted. Counts come from body paragraphs and list items with the navigation, header, footer and forms removed; a page with fewer than five is not scored here. Thresholds were set from a 176-homepage corpus pass and each row states the corpus figure it was set against. Weight 0.8 in the overall and part of the Quote stage, from score version 15.

Crawl waste — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
Internal links carrying a query string
each parameter form is a separate URL to a crawler; tracking and sort parameters on internal links multiply the pages it thinks you have
1 of 297 - ntv_adpz0, unless the parameter changes the page✓
Internal links to http://
every one is a redirect before the page, and a redirect is a fetch that returned nothing
4 of 2970n/a
Declared URLs that redirect
a sitemap entry that redirects sends the crawler somewhere the sitemap should have named in the first place
0 of 6 sampled0✓
Declared URLs that do not answer 200
a declared page that is gone is a fetch spent on a page that will not be indexed
0 of 6 sampled0n/a
Declared pages whose canonical points elsewhere
declaring a page and then telling the engine it is really another page is two fetches for one result
11 of 600n/a
Declared pages marked noindex
a page in the sitemap asking not to be indexed is a contradiction the crawler resolves by fetching it anyway
0 of 600n/a
Duplicate titles across the read pages
pages that look the same from the outside get fetched, compared and mostly discarded
0 of 600n/a
Apex and www both answer 200
two hosts serving the same site is every page twice, and the engine has to guess which copy is the real one
no, one redirectsone host, the other redirects✓

A crawler arrives with a budget. Every row here is a fetch that produced no new page: a redirect, a parameter variant, a case variant, a declared page that points elsewhere. Counted from the delivered homepage, the sitemap, the sampled declared URLs and the daily whole-site read where one exists. None of this moves the grade.

Page subject, the Webref question — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
Business entities declared on the page
Webref decides which entity a page is about and how much; a page that names several unrelated businesses splits its topicality between them
1one subject, others only by declared relationn/a
The page's subject
mainEntity is the explicit declaration; a url on this host is the next best evidence; guessing from the first node is what this check refuses to do
The Denver Post - found by url on this hostdeclared as WebPage.mainEntityn/a
Title and h1 name the subject
the two strongest on-page statements of what the document is about
title yes, h1 yesbothn/a
Places named in the graph
areaServed is read as places the business serves; a locality the business is IN is a different relation, and both belong on the record
2 (Denver, Denver Metropolitan Area and the state of Colorado)the branch's own locality plus a declared service arean/a
Verdict
single: one entity. related: several, all connected. diluted: unrelated businesses share the page. ambiguous: no subject can be resolved
single - one business entity, and it is the subjectsingle or related✓

A document does not compete for a keyword alone; the index links it to entities by their Knowledge Graph identity and ranks documents against each other for that entity (the Webref layer in the recovered Geostore schema). This section reads the structured data on the delivered homepage and asks the question that layer asks: which business is this page about, and how is every other business on it related? Chain membership, subsidiary, department and containment are declared relations; a second business with no relation is a claim the engine must arbitrate. Measured, not scored, until the corpus shows how often each verdict occurs.

AI and search agents — 33/100
Section score 33 / 100
MeasureCurrentlyOptimal
Answer engines allowed
these fetch a page while composing a live reply
8 of 17all of them✗
Search indexes allowed
classic index coverage still drives most discovery
16 of 24all of them✗
Training crawlers allowed
blocking trainers while allowing answer engines is coherent, not a defect
22 of 46your policy choice — not scoredn/a
Regional search engines allowed
Naver, Sogou, 360 and Yisou matter if you sell into those markets and are irrelevant if you do not; that is a business fact we cannot read off the page
7 of 7your policy choice — not scoredn/a
Social link previews allowed
blocking these does not touch answer engines, it just makes your links render as bare URLs when anyone shares them
6 of 6your policy choice — not scoredn/a
SEO and research crawlers allowed
blocking a link-index crawler costs you competitor visibility, not answer-engine visibility — a different trade from the one above
14 of 14your policy choice — not scoredn/a
robots.txt groups
a second * group is invisible to parsers that stop at the first match
681 group per user-agent, no duplicate * group✓
Answer engines — fetch at question time — 8 of 17 allowed
OAI-SearchBotBLOCKED at / · ChatGPT search index · named rule
ChatGPT-UserBLOCKED at / · ChatGPT live fetch · named rule
Claude-SearchBotBLOCKED at / · Claude search index · named rule
Claude-UserBLOCKED at / · Claude live fetch · named rule
PerplexityBotBLOCKED at / · Perplexity index · named rule
Perplexity-UserBLOCKED at / · Perplexity live fetch · named rule
Gemini-Deep-Researchallowed · Gemini research agent · wildcard rule
MistralAI-UserBLOCKED at / · Le Chat live fetch · named rule
DuckAssistBotBLOCKED at / · DuckDuckGo AI assist · named rule
YouBotBLOCKED at / · You.com · named rule
PhindBotallowed · Phind · wildcard rule
Kagibotallowed · Kagi · wildcard rule
Copilot-Userallowed · Microsoft Copilot fetch · wildcard rule
Meta-ExternalFetcherallowed · Meta AI live fetch · wildcard rule
Amzn-Userallowed · Amazon live fetch for Alexa questions · wildcard rule
Google-Agentallowed · Google user-triggered agent, acts on the web · wildcard rule
Google-GeminiNotebookallowed · Gemini Notebook user-supplied URL fetch · wildcard rule
Search indexes — 16 of 24 allowed
Googlebot-Imageallowed · Google image crawler · wildcard rule
Googlebot-Newsallowed · Google News · wildcard rule
Googlebot-Videoallowed · Google video crawler · wildcard rule
Storebot-Googleallowed · Google Shopping · wildcard rule
Mediapartners-Googleallowed · Google AdSense · wildcard rule
APIs-Googleallowed · Google push delivery · wildcard rule
BingPreviewallowed · Bing page preview · wildcard rule
msnbotallowed · Microsoft legacy crawler · wildcard rule
SlurpBLOCKED at / · Yahoo Slurp · named rule
YandexBLOCKED at / · Yandex (umbrella token) · named rule
coccocbot-weballowed · Coc Coc · wildcard rule
Googlebotallowed · Google Search and AI Overviews · wildcard rule
bingbotallowed · Bing and Copilot index · wildcard rule
ApplebotBLOCKED at / · Apple and Siri · named rule
AmazonbotBLOCKED at / · Amazon · named rule
DuckDuckBotallowed · DuckDuckGo · wildcard rule
YandexBotBLOCKED at / · Yandex · named rule
BaiduspiderBLOCKED at / · Baidu · named rule
Seznambotallowed · Seznam · wildcard rule
Neevabotallowed · Neeva · wildcard rule
PetalBotBLOCKED at / · Huawei Petal · named rule
Amzn-SearchBotBLOCKED at / · Amazon search eligibility, separate from Amazonbot · named rule
MistralAI-Indexallowed · Mistral search index for Vibe · wildcard rule
ExaSearchBotallowed · Exa AI search index, Web Bot Auth signed · wildcard rule
Training / corpus crawlers — 22 of 46 allowed
Magpie-crawlerallowed · Magpie AI · wildcard rule
img2datasetallowed · img2dataset image corpus · wildcard rule
AwarioRssBotallowed · Awario RSS · wildcard rule
AwarioSmartBotallowed · Awario smart · wildcard rule
TurnitinBotallowed · Turnitin · wildcard rule
archive.org_botBLOCKED at / · Internet Archive · named rule
ia_archiverBLOCKED at / · Internet Archive (legacy) · named rule
meta-webindexerallowed · Meta web index · wildcard rule
omgiliBLOCKED at / · Webz.io omgili · named rule
cohere-training-data-crawlerBLOCKED at / · Cohere training corpus · named rule
PanguBotBLOCKED at / · Huawei PanGu · named rule
Ai2Bot-DolmaBLOCKED at / · Allen Institute Dolma · named rule
FriendlyCrawlerallowed · FriendlyCrawler ML · wildcard rule
VelenPublicWebCrawlerBLOCKED at / · Velen · named rule
MyCentralAIScraperBotallowed · MyCentral AI · wildcard rule
DeepSeekBotallowed · DeepSeek · wildcard rule
ICC-CrawlerBLOCKED at / · NICT ICC · named rule
GoogleOtherallowed · Google non-search fetch · wildcard rule
Google-CloudVertexBotallowed · Vertex AI agent build · wildcard rule
GPTBotBLOCKED at / · OpenAI training and index · named rule
ClaudeBotBLOCKED at / · Anthropic training · named rule
anthropic-aiBLOCKED at / · Anthropic legacy agent · named rule
Claude-WebBLOCKED at / · Anthropic legacy agent · named rule
Google-ExtendedBLOCKED at / · Gemini training · named rule
Applebot-ExtendedBLOCKED at / · Apple training · named rule
CCBotBLOCKED at / · Common Crawl · named rule
BytespiderBLOCKED at / · ByteDance · named rule
meta-externalagentBLOCKED at / · Meta AI training · named rule
FacebookBotBLOCKED at / · Meta legacy · named rule
cohere-aiallowed · Cohere · wildcard rule
DiffbotBLOCKED at / · Diffbot knowledge graph · named rule
OmgilibotBLOCKED at / · Webz.io · named rule
ImagesiftBotallowed · Imagesift · wildcard rule
TimpibotBLOCKED at / · Timpi · named rule
AI2BotBLOCKED at / · Allen Institute · named rule
Scrapyallowed · Generic scraper framework · wildcard rule
SemrushBot-OCOBallowed · Semrush AI corpus · wildcard rule
Applebot-Extended-AdsBLOCKED at / · Apple ads corpus · named rule
TikTokSpiderallowed · TikTok · wildcard rule
QuillBotallowed · QuillBot · wildcard rule
Webzio-ExtendedBLOCKED at / · Webz.io extended · named rule
ProRataIncallowed · ProRata · wildcard rule
AwarioBotallowed · Awario · wildcard rule
MistralAI-Trainingallowed · Mistral training crawl · wildcard rule
GoogleOther-Imageallowed · Google common crawler, public image URLs · wildcard rule
GoogleOther-Videoallowed · Google common crawler, public video URLs · wildcard rule
SEO and market-research crawlers — 14 of 14 allowed
AhrefsBotallowed · Ahrefs link index · wildcard rule
SemrushBotallowed · Semrush crawler · wildcard rule
MJ12botallowed · Majestic link index · wildcard rule
DotBotallowed · Moz link index · wildcard rule
rogerbotallowed · Moz site crawler · wildcard rule
DataForSeoBotallowed · DataForSEO · wildcard rule
BLEXBotallowed · WebMeUp link index · wildcard rule
CloudflareBrowserRenderingCrawlerallowed · Cloudflare Browser Run /crawl · wildcard rule
Cloudflare-AutoRAGallowed · Cloudflare AutoRAG · wildcard rule
Peer39_Crawlerallowed · Peer39 ad context · wildcard rule
AdsBot-Googleallowed · Google Ads quality · wildcard rule
AmazonAdBotallowed · Amazon Ads · named rule
AdIdxBotallowed · Microsoft Ads · wildcard rule
OAI-AdsBotallowed · OpenAI ChatGPT ads page validation · wildcard rule
Regional search engines — 7 of 7 allowed
Yetiallowed · Naver · wildcard rule
YoudaoBotallowed · Youdao · wildcard rule
Exabotallowed · Exalead · wildcard rule
Sogou web spiderallowed · Sogou · wildcard rule
YisouSpiderallowed · Yisou · wildcard rule
360Spiderallowed · 360 Search · wildcard rule
Sosospiderallowed · Soso · wildcard rule
Social link previews — 6 of 6 allowed
Facebotallowed · Meta link crawler · wildcard rule
Twitterbotallowed · X link preview · named rule
LinkedInBotallowed · LinkedIn link preview · wildcard rule
facebookexternalhitallowed · Facebook link preview · named rule
Pinterestbotallowed · Pinterest · wildcard rule
Slackbot-LinkExpandingallowed · Slack unfurl · wildcard rule

9 answer engine(s) are blocked. These are the agents that fetch a page while composing a reply, so this directly removes the site from live answers.

Speed a crawler sees, and field data where it exists — 67/100
Section score 67 / 100
MeasureCurrentlyOptimal
Largest Contentful Paint (p75)
how long a real visitor waits before the main thing on the page appears
3.45s2.5s or less✗
Interaction to Next Paint (p75)
how long the page takes to respond after a real visitor taps something
154ms200ms or less✓
Cumulative Layout Shift (p75)
how much the page moves under a reader mid-read
0.130.1 or less✗
Time to First Byte (p75)
the server half of every other number on this list
522ms800ms or less✓
Machine files answer quickly (measured here)
the median time this scanner waited for your own robots, sitemap and entity files. A crawler that times out records nothing at all, so this is the speed number that decides whether you are read
56 ms median of 5under 500 ms✓
No slow outlier among them
one slow file is enough to lose a crawl. This is the worst single response of the same set, not an average that hides it
449 ms slowestunder 1500 ms✓
Lighthouse lab run
a lab run is a synthetic test from Google’s servers, not a reading of what your visitors experienced
the Lighthouse run is queued; its numbers appear on your next scan of this URLa completed lab run where field data is absentn/a

Measured on this exact URL. These are the 75th percentile of what real Chrome users experienced over the last 28 days, not a test run from here — the thresholds are Google’s published ones, the only ones that are.

Payload and render path — 20/100
Section score 20 / 100
MeasureCurrentlyOptimal
Content ratio
the rest is markup a model must read and discard
4.4%10–100% of decompressed bytes are visible text✗
Decompressed bytes
large pages are fetched less often and truncated more
297,605Bunder 500,000B✓
Inline JavaScript
inline JS is pure overhead to a text-extracting crawler
38,121Bunder 20,000B✗
Render-blocking resources
each one delays first paint and the crawler's render budget
240–5✗
Deferred vs blocking scripts
a blocking script stops HTML parsing dead
2 deferred / 14 blockingevery script deferred or async✗

Where the bytes go

Visible text13,081 B 4.4%
Inline CSS45,571 B 15.3%
Inline JavaScript38,121 B 12.8%
Structured data2,238 B 0.8%
Markup and attributes195,658 B 65.7%
Render-blocking resources24 10 stylesheets, 14 scripts
Entities with a stable @id1 / 3 33.3%
Credential edges0 none declared

1 of 1 map identity links point at places, not this business — 5990 Washington St, Denver, CO 80216. sameAs asserts that two URLs describe the same entity.

Machine-file trust chain — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
robots.txt was readable
every other file in the chain is normally discovered through robots.txt
yes200, text/plain✓
robots.txt names a sitemap
a crawler that has to guess the sitemap path often does not find it
1 Sitemap: line(s)at least one✓
robots.txt points at llms.txt
an llms.txt nothing links to is only reachable by guessing the conventional path
—named when the file existsn/a
llms.txt links stay on this host
an off-host URL in your llms.txt sends the model to someone else's page as if it were yours
—most links on this hostn/a
llms.txt points onward
the chain should keep going: llms.txt is a map, not a terminus
—names the sitemap or the entity graphn/a
entitymap.json parses as JSON
a machine file that does not parse is worth less than one that is absent, because it looks present
—valid JSONn/a
entity graph references this host
a graph that never names this site is describing something else
—yesn/a
Canonical points at this host
a canonical on another host hands the page's standing to that host
same hostsame host as the one serving the page✓
Canonical uses https
an http canonical invites a redirect chain on every crawl
httpshttps✓

robots.txt ✓ → sitemap ✓ → llms.txt ✗ → entitymap.json ✗

A ✗ breaks the chain at that point: everything downstream is only reachable by a crawler guessing the conventional path.

Claim consistency — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
The schema declares an entity for this site
an entity graph that names only other companies gives an answer engine nothing to attach this site to
yes, an organisationone Organization, LocalBusiness or Person whose url is this host✓
Schema business name appears on the page
an answer engine ingests the assertion and never compares it to the page, so a stale one is repeated for months
—the name the schema asserts is the name a reader seesn/a
Schema strings are not HTML-escaped
a JSON string is not an HTML context: the entity is read literally, so the business name contains the characters a-m-p
—no &amp; or &#39; inside a JSON-LD valuen/a
Schema phone appears on the page
a phone number that exists only in the markup is the one an assistant will read out
—the same 10 digitsn/a
Schema locality appears on the page
a locality nobody states on the page is a claim with no support behind it
—the declared town or city is named in the copyn/a
Experience claim is backed by the schema
a datable claim in the copy that the structured data contradicts is the cheapest thing in an audit to disprove
—the copy and foundingDate agree within a yearn/a

Checked against the page’s own visible text, with no extra request. Node type: NewsMediaOrganization.

Agent instruction surface — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
No agent-directed instructions in machine-only surfaces
hidden text, comments, alt attributes and llms.txt are read by a machine and proofread by nobody
38 surface(s) read, nothing matched0 matches✓
No hidden block over 50 words
an agent ingests hidden copy at full weight while a reader never sees it
none0 blocks✓
Agent-instruction file (agents.md)
the file agents are told to obey, as distinct from llms.txt which is the content map. Every Shopify store now ships one; adoption elsewhere is early, so its absence is not a defect — but if you publish one, everything in it is read as instruction
404your choice — not scoredn/a

This reports EXPOSURE, not intent. Most hidden text is an old SEO habit or a collapsed menu, and a match here is a prompt to go and read it — not a finding that someone attacked the site.

What each agent receives — 88/100
Section score 88 / 100
MeasureCurrentlyOptimal
Every identity gets a response
a request that dies is indistinguishable from a site that is down, to the agent making it
all 14 answered0 silent✓
No answer engine is refused
a 403 to GPTBot is the whole answer to why a site is never cited
none refused0 refusals✓
No answer engine is handed an interstitial
a challenge at HTTP 200 looks fine to a status check and contains no content at all
none0 challenge pages✓
Crawlers get what an unnamed client gets
these fetches send a crawler's user-agent from OUR address, which is not in the range that operator publishes — so a split has two readings, cloaking or correct spoof-rejection, and only the operator can settle which
noneno identity split✓
No identity is sent to a different URL
a bot-only redirect quietly removes the page an answer engine was asked to read
none0 redirected✓
Same x-robots-tag for every identity
a bot-only noindex removes the page from search while the site looks perfectly fine in a browser, and nothing else checks it
consistentno identity-specific header✓
Answer engines get the same text as a browser
the page a browser renders is not evidence about the page an answer engine was given
100% of the browser's words90-100%✓
Training crawlers reaching the site
these collect pages for model training rather than answering questions, so refusing them is a policy decision and not a defect — it is reported because the edge may be doing it without anyone having decided, and a 403 here is not the same fact as a 429
Bytespider 403 of 6 turned awayyour choice — not scoredn/a
Stated robots policy matches actual behaviour
robots.txt is a promise and the edge is the behaviour; neither one on its own can tell you they disagree
OAI-SearchBot: disallowed but served, Claude-SearchBot: disallowed but served, PerplexityBot: disallowed but served, Applebot: disallowed but served0 conflicts✗
Served the same content from every region tested
a site that serves different content by country is telling different engines different things about itself
—identical text worldwiden/a

One URL, fifteen fetches in the same second — an unnamed client, twelve named crawlers, a mobile browser, and the browser again as a control. The control fetches agreed (self-similarity 1), so text differences below are differences, not noise. A user-agent is a claim, including when we are the one making it.

IdentityStatusWordsSame textWhat happened
Unnamed client20020361
GPTBot20020361
ClaudeBot20020361
OAI-SearchBot20020361
Claude-SearchBot20020361
PerplexityBot20020360.91different body text
Googlebot20020360.91different body text
Mobile browser20020361
Bingbot20020360.91different body text
Applebot20020360.91different body text
Amazonbot20020361
Bytespider4035—refused HTTP 403 — vendor unidentified; reproduced on a second request
Meta-ExternalAgent20020360.91different body text
CCBot20020360.91different body text
AgentConflictDetail
OAI-SearchBotdisallowed but servedHTTP 200 with 2036 words
Claude-SearchBotdisallowed but servedHTTP 200 with 2036 words
PerplexityBotdisallowed but servedHTTP 200 with 2036 words
Applebotdisallowed but servedHTTP 200 with 2036 words

Per-engine retrieval. Each answer engine reads through named crawlers with different jobs — one builds the index, one fetches live when a user asks, one collects training data. They are separate permissions and a site commonly grants one and refuses another. Nothing here is scored: these same facts are already scored once above, and this measures what an engine is permitted and given — never what a model has retained or would cite, which no scanner can see from outside.

EngineIndexLive fetchTrainingWhat the edge actually didAddressed by name
ChatGPT✗ blocked✗ blocked✗ blocked✓ served 2036 words— neither file
Claude✗ blocked✗ blocked✗ blocked✓ served 2036 words— neither file
Perplexity✗ blocked✗ blocked— —✓ served 2036 words— neither file
Google AI Overviews / Gemini✓ allowed✓ allowed✗ blocked✓ served 2036 words— neither file
Microsoft Copilot✓ allowed✓ allowed— —✓ served 2036 words— neither file
Apple Intelligence✗ blocked— —✗ blocked✓ served 2036 words— neither file
Meta AI✗ blocked✓ allowed✗ blocked✓ served 2036 words— neither file
Amazon✗ blocked— —— —✓ served 2036 words— neither file
DuckAssist✗ blocked— —— —— not probed— neither file
Mistral— —✗ blocked— —— not probed— neither file
Common Crawl (feeds many models)— —— —✗ blocked✓ served 2036 words— neither file

A blocked training column beside an allowed index column is coherent policy, not a defect: it says "answer with me, do not train on me." The last column is whether your llms.txt or agents.md addresses that engine by name — robots.txt is a permission, those two files are where an operator actually talks to an agent. Naming nobody is the norm and is not scored.

no vantage provider is configured and no extension reading exists for this site in the last 14 days, so this was not measured

Mobile surface — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
Images and embeds with dimensions declared
Media without width and height reserves no space, so everything below it moves when the image lands. This is read from the markup and is a risk indicator, not a measured CLS value — the real number needs a real page load.
1 of 4297.6% unsized — not scoredn/a
Viewport meta
without it a phone renders the desktop layout scaled down, and that is what a mobile crawler records. Roughly 218 million live sites declare one (BuiltWith, August 2026), so its absence is the exception, not the norm.
width=device-widthwidth=device-width, initial-scale=1✓
Zoom is not locked
locking zoom is an accessibility failure and a one-line fix
readers can zoomno user-scalable=no, no maximum-scale under 1.5✓
Apple touch icon
what a saved-to-homescreen shortcut and several share surfaces use
declaredone apple-touch-icon link — not scoredn/a
Theme colour
sets the browser chrome on mobile; its absence is the cheapest visible gap on this list
absenta theme-color meta — not scoredn/a
Beyond the basics
almost every site declares a viewport and almost none declare the rest, so this row is where a site separates itself rather than a place it loses points
none beyond viewportnot scored — these are differentiators, not defectsn/a
Local presence, citations and NAP — 92/100
Section score 92 / 100
MeasureCurrentlyOptimal
Links to its Google Business Profile
the profile and the site are one entity to an answer engine, and this link is the only bridge between them it can see
yesa g.page, maps.app.goo.gl or maps place URL✓
That profile link resolves
a dead profile link is worse than none: it asserts an identity that cannot be checked. A 429 or 403 here is Google throttling our check, not a fault on your site, and is left unscored
200 → www.google.com200 on a Google host✓
Street address in the structured data
a service-area business legitimately omits this, so an absence is reported and not scored against you
declareddeclared for a premises business✓
Geo coordinates
coordinates are how a machine ties the entity to a place without parsing an address string
declaredlatitude and longitude✓
Opening hours
“are they open now” is one of the most common questions an assistant is asked about a local business
declaredopeningHours or openingHoursSpecification✓
Telephone in the structured data
the number an assistant reads out comes from here, not from the page
declareddeclared✓
Coordinates are usable
two decimal places is about a kilometre — fine for a city, useless for a storefront someone is being driven to
1 pair(s), full precisionin range, not 0,0, at least 3 decimal places✓
Each location has its own coordinates
one centroid copied onto every branch tells an assistant they are all the same place
—no two locations share a pointn/a
Coordinates agree with the declared country
negating the longitude puts it inside US - a missing minus sign
outside US — longitude signthe point falls inside the country the address names — reported, not scoredn/a
Coordinates match the declared address
the address and the point are two independent claims about one place, and nothing else on the web compares them
—within 500mn/a
Location entities declared
more location entities than real listings is the single most common way a local entity graph goes wrong
11 per real premises — not scoredn/a
City on the page vs declared address
the place a page markets and the place it declares are the same, so an engine has nothing to reconcile
both say Denverreported, not scoredn/a
Business name in structured data
two spellings of a legal name are two entities to a retrieval system, and it cannot tell which one you are
The Denver Postexactly one spelling✓
Schema and page agree on the phone number
with only one source there is nothing to compare, which is not the same as agreement
only structured data declares oneboth places state the same numbern/a
Phone numbers a machine can read
an assistant dictating a number has to pick one; a second number is usually an old one that still rings somewhere
+13039541010exactly one number✓
Postal address in structured data
two addresses split the entity across two places
5990 Washington St., Denver, CO, 80216one address, or none for a service-area business✓
Profiles the site claims
sameAs is an identity claim: every link says this business IS the thing at that URL
3the profiles you actually own✓
Off-site listings found by your phone number
searched by the one key a service-area business cannot hide; Yelp, Angi, Thumbtack, Houzz, Facebook refuse an outside client and are recorded as our limit, not as an absence
4 on YellowPages, DexKnows, BBBevery one you own, declaredn/a
Found off-site, not declared in sameAs
a listing you never linked is corroborated without you: whatever it prints becomes your name and category in the answer
4 — which ones, with a licence0✗
Listings printing a different business name
two names on one phone number are two entities to a retrieval system
00✓
Data Axle carries this phone number
Data Axle (formerly Infogroup / infoUSA) is one of two US business-data suppliers Google’s own provider list names (Acxiom is the other); a business it does not carry is missing from one named feed. Reported, not scored.
refused (403) — our limit, not an absencelistedn/a
Directory listings declared
reported, not scored — how many citations a business needs is a marketing judgement, not a measurement
0the ones that matter for your traden/a

Checked because this site declares a local business entity (LocalBusiness). This measures the site side of Google Business Profile alignment — whether the business points at its own profile and carries the fields a profile is matched on. It does not read the profile itself: that needs an API key, and inventing facts about a listing we cannot see would be worse than reporting nothing. Profile link found: https://maps.app.goo.gl/aK261HiEe923wgBV7.

Every profile below is one the site itself declares in sameAs. Nothing here was discovered by guessing at directories — undeclared listings need an index this scan does not have.

ProfileAnsweredYour phoneYour name
Facebook
https://www.facebook.com/denverpost
not checked——
X
https://x.com/denverpost
not checked——
LinkedIn
https://www.linkedin.com/company/denver-post-media/
not checked——

Directory coverage

Ten sources an answer engine is likely to reach for when asked about a local business, checked against what this site declares. Declared means the site names the profile in its own structured data — the only thing this scan can verify. A listing that exists but is not declared will read as missing here, and that is itself worth fixing: an engine reading your site has no way to find it either.

Yelp
the single most-cited local source in AI answers after Google itself
not declared
BBB
trust signal, and one of the few directories with a verification process an engine can lean on
not declared
Facebook
carries hours and phone, and is read by several engines as a primary source
declared
Nextdoor
hyperlocal, and disproportionately cited for home services
not declared
Thumbtack
category-specific lead surface for trades
not declared
Angi
category-specific, still heavily indexed
not declared
Bing Places
feeds Copilot and, historically, several ChatGPT retrievals
not declared
Apple Maps
the default map on every iPhone, and invisible to most SEO tooling
not declared
Trustpilot
review corpus that answer engines quote directly
not declared
Yellow Pages
low value alone, but a cheap consistency anchor
not declared

9 of 10 are not declared. Each one is a place a retrieval system could have found a second, independent statement of your name, address and phone — and the agreement between those statements is what makes any of them trustworthy.

Machine layer over time. Our own history of a site starts the first time we scanned it; the Internet Archive holds what came before. Nothing here is scored — a site’s past is not a defect, and the Archive’s coverage is uneven, so a missing snapshot says nothing about the site.

Archived robots.txt versions40 distinct, 2002-09-23 → 2004-04-12
Earliest copy named an AI crawlerno — the site names 33 today, so that policy was written after 2002-09-23
Earliest copy, first line<!-- libOpenCDA reports: OID = 36~11~ --> <html> <head> <title>The Denver Post Online - De
Archived llms.txtnone archived
Sitemap coverage — 100/100
Section score 100 / 100
MeasureCurrentlyOptimal
A sitemap was readable
a sitemap is the only place a crawler learns about pages nothing links to
790 URLs declaredat least one sitemap resolves and parses✓
Sitemaps named in robots.txt resolve
a robots.txt pointing at a dead sitemap sends every crawler to a 404
1 of 1every named sitemap returns 200✓
Homepage links are declared in the sitemap
a page missing from the sitemap is still findable by following links; a page missing from both is findable by nothing
not measured — only 20 of 3000 child sitemaps were read0 undeclaredn/a
Sampled declared URLs resolve
a sitemap that lists dead URLs spends a crawler's budget on nothing
0 of 6 broken0 broken✓
Sampled declared URLs are final
declaring the pre-redirect URL makes every crawl pay an extra hop
0 of 6 redirect0 redirects✓
URLs declared in the sitemap790
20 of 3000 child sitemaps read — the count is a floor, not a total
Internal links on the homepage193
Linked but not declarednot measured
only 20 of 3000 child sitemaps were read, so a page missing from what we read may be declared in a part we did not read
Declared URLs sampled6 checked · 0 did not resolve · 0 redirected

This compares the sitemap against the links on the homepage only, so it finds pages the sitemap omits — it cannot prove a declared page is unreachable, because that needs a full crawl. The reverse gap (declared but linked from nowhere) is real and is not measured here. A page missing from the sitemap is still findable by following links; a page missing from both is findable by nothing.

Naming and case consistency — 80/100
Section score 80 / 100
MeasureCurrentlyOptimal
Hostname case
hostnames are case-insensitive but a mixed-case host splits logs and analytics
lowercasealways lowercase✓
Internal link paths
paths ARE case-sensitive on most origins: /About and /about are two URLs, two cache keys and two index entries
all 297 lowercase98–100% lowercase✓
Trailing slash
a page linked as /x and /x/ is two URLs to a crawler; both get fetched and both compete for the same canonical
1 target linked both with and without a slash (/contact-us)one form per URL, and the other form redirects to it✗
Host form in absolute links
links to both www and the bare domain send a crawler to two copies of the site whatever the canonical says
consistentone hostname in every self-link✓
Query strings in internal links
a query string makes a distinct URL; navigation that carries one produces near-duplicate pages
noneunder 2% of internal links (and fewer than 3)✓
Image signals — 80/100
Section score 80 / 100
MeasureCurrentlyOptimal
Images on the page
counted from <img> tags in the delivered HTML
41 (3 SVG)n/a
Missing alt text
alt is the only text an engine gets from a picture; empty alt is correct for decoration, missing alt is a hole
0 missing, 1 empty (0% missing)0 missing; empty only on decorative images✓
Width and height declared
without both, the browser cannot reserve space and the page shifts as pictures load (the CLS score)
0 of 41all✗
Lazy-loaded
lazy-loading the hero image delays the largest paint; lazy-loading the rest speeds everything else
0 of 41below the fold only✓
Modern formats (WebP/AVIF)
smaller files, same picture
0 of 41mostn/a
Hero image prioritised
the one image the browser should fetch first
4 fetchpriority=high, 4 preloadedone, when the largest paint is an imagen/a
Entity image in structured data
what an engine shows next to the name; a URL in image or logo, ideally an ImageObject with dimensions
yesan image or logo on the organisation node✓
og:image
the picture a link preview uses when the page is shared
yespresent✓

What an engine can learn from the pictures. Five rows score: missing alt, declared dimensions, nothing lazy-loaded above the fold, an entity image in JSON-LD and an og:image. Empty alt, file format and hero priority are reported and never scored — empty alt is correct on a decorative image, and format is a preference, not something an engine fails to read.

What this scan could not see 23 of 24 scored sections measured · 13 of 14 identities answered · 0 stage failures · 1 vantage · 4 not observed
SubjectStateAttemptedReason
nl_entitiesskippednothe deep entity read runs for domains under continuous record; this scan used the keyless resolver
vantageunmeasurednono vantage provider is configured and no extension reading exists for this site in the last 14 days, so this was not measured
identity:BytespiderrefusedyesHTTP 403
section:commerceunmeasuredunknownno row of this section was produced for this page; it did not enter the grade

A non-observation is recorded, never scored: an unmeasured row is excluded from its section, not counted as a fail. Counts only; there is no composite confidence number.

Where each fact on this page comes from 16 facts · 21 statements · no disagreements
FactValue statedStated in
nameThe Denver PostJSON-LD LocalBusiness
meta og:site_name
agrees
telephone+13039541010JSON-LD LocalBusinessone source
emailnewsroom@denverpost.comJSON-LD LocalBusinessone source
address5990 Washington St., Denver, CO, 80216JSON-LD LocalBusinessone source
hoursMo 8AM-4PM; Tu 8AM-4PM; We 8AM-4PM; Th 8AM-4PM; Fr 8AM-4PMJSON-LD LocalBusinessone source
geo39.7392,104.9903JSON-LD LocalBusinessone source
urlhttps://www.denverpost.com/JSON-LD NewsMediaOrganization
JSON-LD LocalBusiness
agrees
logohttps://i0.wp.com/www.denverpost.com/wp-content/uploads/2020/11/denverpost.jpg?w=1200&crop=00px100630px&ssl=1JSON-LD LocalBusinessone source
foundingDate1892JSON-LD LocalBusinessone source
areaServed1 area: Denver Metropolitan Area and the state of ColoradoJSON-LD LocalBusinessone source
sameAs3 profilesJSON-LD LocalBusinessone source
titleThe Denver Post – Colorado breaking news, sports, business, weather, entertainment.<title>agrees
The Denver Postmeta og:title
descriptionColorado breaking news, sports, business, weather, entertainment.meta og:descriptionone source
canonicalhttps://www.denverpost.commeta og:urlone source
languageen-US<html lang>agrees
en_USmeta og:locale
imagehttps://www.denverpost.com/wp-content/uploads/2020/11/denverpost.jpgmeta og:image
meta twitter:image
agrees

Read from the page this scan received: its schema nodes, meta tags, link elements and page links, each value kept with the element that stated it. A fact stated in several places should say the same thing in each. Recorded, not scored.

If you would rather not do this yourself

Fix everything in this report
The 8 findings above, including 5 scored serious, corrected on your site and re-measured afterwards. You get a change receipt: the before and after, each sealed and independently verifiable on your own machine, so the fix is evidenced rather than asserted.
  • The machine layer, authored for the business. An entity map, an agents.md and an llms.txt written from the site’s own published facts rather than templated — the entity map being what lets a retrieval system tell this business apart from one with a similar name. Sold separately until now.
  • Thirty days of monitoring, included. The site is re-scanned twice a day and any crawler-visible change is posted to your webhook, so a regression is found by us rather than by you.
  • A second sealed receipt at day thirty. The first proves the fix landed; the second proves it held. Both verify on your own machine against a published root, with no need to trust us.
  • Done is defined by re-measurement, not by assertion. A finding that is still there on the re-scan is not finished work.
Get this fixed — $749
one local-business site · what that covers

The fix includes the machine layer: an entity map, an agents.md and an llms.txt authored for this business rather than templated — 6 of the 7 agent surfaces checked above are absent here. A bigger site, or several domains at once, is quoted against the report first: email hello@crawlcheck.io with this address.

Measured, not in the grade (7)

These are measured on every scan and shown in full. Each says why it carries no weight yet: most are waiting on enough sites to calibrate against, and some depend on an integration a site may not have, which must never decide a grade.

Agent surfaces (emerging, not scored) — 10 of 10 rows read

Not in the grade: these are published standards at very different stages of adoption, and several do not apply to every business.

Rows with a reading (not in the grade)10 / 10
MeasureCurrentlyOptimal
agents.md
The one surface here that pays off for an ordinary business today. Almost every agents.md on the web is a platform default its owner never wrote
absentAn authored file at /agents.md: what the business does, the area it serves, what an agent must collect before acting, and what it must not promisen/a
SKILL.md
The agent-skills convention. A 200 that carries HTML is a soft 404 and counts as absent - a naive check would call it present
absent (404)A markdown file at /SKILL.md with YAML frontmatter naming the skill, then what an agent can do here and hown/a
Markdown for agents
Cuts what an agent must parse. A toggle on some edges rather than a build
absent (text/html; charset=utf-8)Accept: text/markdown returns a markdown body; HTML stays the default for browsersn/a
Link response headers
Useful where a relation is real. Inventing one to score a point is the cargo-culting these lists encourage
present: <https://www.denverpost.com/wp-json/>; rel="https://api.w.org/", <https://wp.me/7yQli>; rel=shortlinkRegistered relations only - canonical, alternate, describedby. Never an invented rel to satisfy a checkern/a
Agent skills index
For sites that expose actions an agent can perform. A brochure site has no skills to declare
absent/.well-known/agent-skills/index.json listing each skill with a name, description, url and sha256n/a
API catalog
NOT APPLICABLE: no public API found. A catalog of nothing is not an improvement
absentNothing to publish - no API was discoverable on this siten/a
MCP server card
Applies once you run an MCP server. Publishing a card without one advertises an endpoint that does not answer
absent/.well-known/mcp/server-card.json with serverInfo, transport endpoint and capabilitiesn/a
Agent payment protocols
NOT APPLICABLE: no Product, Offer or checkout found. A scanner that docks a service business for this is measuring the wrong site
not applicableNothing to publish - no commerce signals on this siten/a
OAuth / protected-resource metadata
NOT APPLICABLE: nothing to authenticate against
not applicableNothing to publish - no protected API on this siten/a
Web Bot Auth
Identifies you as a caller, not as a destination. Irrelevant to a site that only receives traffic
informationalA JWKS at /.well-known/http-message-signatures-directory - only if THIS site sends signed agent requests to othersn/a

None of these affect the grade. They are published standards at very different stages of adoption, and several do not apply to every business — where that is true this table says so and why. A checklist that counts them all is measuring the wrong site.

Content entities (measured, not yet scored) — 4 of 4 rows read

Not in the grade: the reading is sound but the scoring rule is not settled, and a rule that may move should not move a grade.

Rows with a reading (not in the grade)4 / 4
MeasureCurrentlyOptimal
Entities read from the content
declared entities come from the page's own JSON-LD; the rest are a HEURISTIC read of the prose and are labelled as such
24the things this page is aboutn/a
Resolve to a public entity record
resolution means Wikidata holds an item for this name. It is not a ranking signal and does not mean the page ranks for it
19 of 24a named thing the public record knowsn/a
Sense is the primary one, not the page's
a bare name resolves to whatever the public record treats as its main sense — "Aurora" reads as the light display, not the Colorado city. The description is printed beside every resolution so you can see which sense you got. Resolving in CONTEXT needs a contextual model and is not measurable from a keyless lookup
not disambiguatedresolved in contextn/a
Resolved and absent from the entity graph
a subject the prose is about, that the public record recognises, and that this site's own structured data never names — the writer's gap, not the machine's
180n/a

Entities are resolved against Wikidata — free, keyless, and it answers the question a writer actually has: does this thing exist as a public entity record, or is it just a phrase. Two evidence classes, kept apart: declared entities are read from the page's own JSON-LD; detected entities are a heuristic read of the prose and will contain some noise. Resolution is not a ranking signal.

Written about, recognised by the public record, absent from this site's structured data:

  • Colorado — Q1261 · state of the United States of America · 17 mentions
  • Opinion — Q3962655 · judgement, viewpoint, or statement that is not conclusive · 7 mentions
  • Deion Sanders — Q954184 · American football and baseball player and football coach · 4 mentions
  • Aurora — Q40609 · natural light display that occurs in the sky, primarily at high latitudes (near the Arctic and Antarctic on Earth) or even on other planets · 3 mentions
  • Eastern Plains — Q5148801 · region of the U.S. state of Colorado east of the Rocky Mountains · 3 mentions
  • Hill Project — Q140360689 · video game · 3 mentions
  • Red Rocks — Q2182648 · concert venue near Morrison, Colorado, United States of America · 3 mentions
  • Amendment — Q1269627 · legal act proposed in a bill or motion, that adds, changes, substitutes, omits or abolishes clauses to one or several other acts which previously adopted or projected · 2 mentions
  • California — Q99 · state of the United States of America · 2 mentions
  • Coors Field — Q1129916 · baseball stadium in Denver, Colorado, USA; home venue of the Colorado Rockies · 2 mentions
  • Denver Broncos — Q223507 · National Football League franchise in Denver, Colorado · 2 mentions
  • Denver Pride — Q7242731 · an annual Gay pride event held each June in Denver · 2 mentions

6 more not listed.

Nothing in this section affects the grade — the corpus measures the distribution first.

Service area, as declared (measured, not scored) — 4 of 6 rows read

Not in the grade: a declared service area can be right or wrong only against facts this scan cannot see.

Rows with a reading (not in the grade)4 / 6
MeasureCurrentlyOptimal
Service areas declared in schema
the list an answer engine reads when asked "do they cover X"; a business that serves it but never declares it is invisible for that question
1every place the business actually serves, as City or Place nodesn/a
Declared radius
the radius is the one claim that bounds all the others
no GeoCirclea GeoCircle around the real centren/a
Areas that could be placed on a map
a name that no gazetteer resolves is a name an engine cannot place either
0 of 1every declared name resolves to a placen/a
Areas with a page of their own
declaring an area and having a page for it are different claims; the page is what gets cited
0 of 1a location page per area that mattersn/a
Declared areas outside the declared radius
the graph contradicts itself: the circle says no, the list says yes, and a reader has to guess which one is true
nonezeron/a
Farthest declared area
the distance the business is committing to on the record
—inside the radiusn/a
declared centre

declared, has its own page · declared in schema only · outside the declared radius · ring = GeoCircle

Not placed (no gazetteer match near the centre): Denver Metropolitan Area and the state of Colorado

Not scored. Positions are OpenStreetMap place centroids resolved once per name; a page counts when a sitemap URL contains the place name. What this shows is what the site declares, laid out so the contradictions are visible; whether the business really works there is not on the wire.

The whole site, not just the front door (measured, not scored) — 21 of 21 rows read

Not in the grade: the grade stays anchored on the homepage so it is comparable across sites and days.

Rows with a reading (not in the grade)21 / 21
MeasureCurrentlyOptimal
Pages declared, pages read
what the sitemap promises versus what this pass could fetch; the cap is stated, never hidden
774 declared, 60 read (first 60)every declared page readn/a
Pages not answering 200
a declared page that redirects, errors or is gone is a promise the sitemap breaks on every crawl
0 of 600n/a
Thin pages (under 150 words)
a page with a title and no body is indexed as nothing; location pages are the usual offenders
0 of 600n/a
Near-duplicate pages
Every audit tool accuses sites of duplicate content; almost none measure it. This compares the pages against each other after removing the text they all share, and reports a band rather than a percentage because a percentage would imply a precision the method does not have.
0 near-identical, 65 substantially overlapping, of 1770 pairs0 near-identical pairsn/a
Template text removed before comparing
Shared navigation and footers appear on every page and would put a floor under every comparison. Passages carried by more than 90% of the crawled pages are dropped first, so what is left is each page’s own copy.
303 of 38719 distinct passages—n/a
Pages with no JSON-LD
the homepage graph does not travel; a service page without its own node is anonymous to a resolver
0 of 600 on pages that describe a service or placen/a
Canonical points elsewhere
a page telling crawlers to index a different page is either a deliberate merge or a copied template
11 of 600 unintendedn/a
Orphans (declared, linked from nowhere read)
a page the sitemap declares but no page links to is reachable only by the sitemap, and weighted accordingly
0 of 600n/a
Undeclared pages (linked, not in the sitemap)
linked pages the sitemap forgot; pagination is normal, a service page is not
1000 that mattern/a
noindex on declared pages
declaring a page and then telling crawlers not to index it is two files disagreeing
0 of 600n/a
Duplicate titles
two pages with one title compete with each other for the same query
00n/a
Lane: article (59 pages)
measured against the rows that apply to article pages only: answers 200; canonical points at itself; not noindex; exactly one h1; 150+ words of its own; an Article node; a publication date. A lane is read from what each page declares (its JSON-LD types, its path, a tappable phone number), so a city page and a docs page are no longer failed with the same ruler.
11 of 59 missing canonical points at itself0 missingn/a
Lane: home (1 page)
measured against the rows that apply to home pages only: answers 200; canonical points at itself; not noindex; exactly one h1. A lane is read from what each page declares (its JSON-LD types, its path, a tappable phone number), so a city page and a docs page are no longer failed with the same ruler.
1 of 1 missing exactly one h10 missingn/a
Facts that disagree: phone number
a resolver that reads two values for the same fact on the same site has to choose or hedge; the pages carrying each value are on the record. Example: https://www.denverpost.com/ versus https://www.denverpost.com/
“+13039541010” on 1 page · “+13039541133” on 1 pageone valuen/a
Facts that disagree: email address
a resolver that reads two values for the same fact on the same site has to choose or hedge; the pages carrying each value are on the record. Example: https://www.denverpost.com/ versus https://www.denverpost.com/
“newsroom@denverpost.com” on 1 page · “cmoser@denverpostmedia.com” on 1 pageone valuen/a
Template families
pages grouped by the assets they ship (stylesheets, scripts, inline blocks of 2 KB or more), not by their words. A defect in a family is one fix, however many pages carry it.
6 families across the pages read, 2 one-offsinformationaln/a
Family 1 (27 pages: 27 article)
pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/2026/09/24/daily-horoscope-for-sept-24-2026/
48 KB inline CSS and 35 KB inline script on a typical page, 5.5% text — 3 inline blocks byte-identical on every page (63 KB each, 1709 KB across the family)a block every page carries belongs in one cached filen/a
Family 2 (15 pages: 15 article)
pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/obituaries/phillip-thomas-longo-littleton-co/
1.4 KB inline CSS and 33 KB inline script on a typical page, 3.3% text — 2 inline blocks byte-identical on every page (17 KB each, 256 KB across the family)a block every page carries belongs in one cached filen/a
Family 3 (14 pages: 14 article)
pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/obituaries/jeanette-nail-denver-co/
1.4 KB inline CSS and 32 KB inline script on a typical page, 3.2% text — 2 inline blocks byte-identical on every page (17 KB each, 239 KB across the family)a block every page carries belongs in one cached filen/a
Family 4 (2 pages: 2 article)
pages grouped by the assets they name, with page numbers normalised. A block that is byte-identical on 90% or more of the family is shipped once per page and cached never; its size times the family size is what the family pays. Example: https://www.denverpost.com/2026/09/23/best-fall-clothes-for-children/
48 KB inline CSS and 34 KB inline script on a typical page, 6.8% text — 3 inline blocks byte-identical on every page (63 KB each, 127 KB across the family)a block every page carries belongs in one cached filen/a
Median words per page
the shape of the site as delivered, not just its front door
1242 · 4.7% textinformationaln/a

60 pages read in 11 s, 2026-09-24 14:10 UTC. Read once a day after a scan somebody asked for; the cron sweeps never pay for it.

Which pages: the lists open with Watch — free for 30 days from /pricing. The counts above are the full measurement; the list is the part that saves the walk through the sitemap.

Not scored: the grade stays anchored on the homepage so it is comparable across sites and days. These counts are the same walk the depth tier runs at full length (300 pages, with the link graph) via /api/depth.

Freshness, as the site declares it (measured, not scored) — 6 of 7 rows read

Not in the grade: every date here is one the site published, so this reports what it claims rather than what is true. Calibrated 2026-09-16 on 114 records: the only row most sites carry (a sitemap that declares lastmod, 68%) is not a defect when absent, and the rows that would be defects (every lastmod stamped on one day, 8%; the sitemap and dateModified disagreeing by over a month, 2%) can be measured on too few sites to carry weight.

Rows with a reading (not in the grade)6 / 7
MeasureCurrentlyOptimal
Sitemap URLs with a change date
lastmod is the only change signal an engine can read without an account. A sitemap without dates tells a crawler nothing about what to re-fetch first
790 of 790every URL, with its real daten/a
Newest sitemap date
what the site says was touched most recently; the gap to today is a claim about activity, not a defect on its own
2026-09-24 (today)within the last month for a site that publishesn/a
Spread of sitemap dates
a spread of dates is what a real edit history looks like
26 distinct days across 790 dated URLs (2026-08-26 to 2026-09-24)many days - real edit historyn/a
Homepage Last-Modified header
the HTTP-level change date; a cache in front usually sets it, so it describes the copy served, not the edit
2026-09-24 (today)sent, and truen/a
Homepage dateModified in JSON-LD
the date the structured data asserts; an engine can quote it beside a fact, which is why a stale one is worse than none
nonepresent, and moved when the page changedn/a
Dated elements on the page
a <time datetime> element is a date a machine can read without guessing at the prose
7 (newest 2026-09-24, today)every date a person can see also marked upn/a
robots.txt history in the public archive
how long the site has been keeping a machine-readable front door; time depth is the one claim a new site cannot manufacture
first seen 2002-09-23, 40 versionsa long recordn/a

Every date in this section is one the site published about itself. The section reports the claims and where they disagree; it does not decide whether a page is out of date, because only the business knows that. None of this moves the grade.

Local intents, the Maps vocabulary (measured, not scored) — 2 of 3 rows read

Not in the grade: the intent vocabulary is recovered names with no weights, and a claimed intent without a page is a content decision, not a defect; reported until the corpus says how it distributes.

Rows with a reading (not in the grade)2 / 3
MeasureCurrentlyOptimal
Local intents the site's own declarations name
a local query is resolved into one of 446 intent types before any place is retrieved; the site's business types, Service nodes, title and headings are what name them here
1 (post offices)every intent you serve, and none you do notn/a
Claimed intents with a page linked from the homepage
an intent named only in a paragraph competes with every page on the web that names it; a page whose address carries the intent is what a query lands on
0 of 1 - no page for: post officesall of themn/a
Intents searchers used to find the Business Profile
no Business Profile search terms on record for this domain
not measuredread from the profile's search termsn/a
Conversion path on the page (measured, not scored) — 5 of 6 rows read

Not in the grade: whether a visitor can act is a business judgement, and the grade is for what a machine receives.

Rows with a reading (not in the grade)5 / 6
MeasureCurrentlyOptimal
Phone link on the page
on a phone, a tel: link is the shortest path from reading to calling; a printed number is not tappable
noneat least one tap-to-call linkn/a
Phone link before the fold
the visitor who arrived from an answer engine has already decided; the number should be where they land
noin the first quarter of the pagen/a
Quote or contact form
a form is the only path that works when the office is closed
1 found, one before the foldone form, reachable without scrollingn/a
Calls to action
counted from link and button text: quote, estimate, call, book, schedule, contact
2 (first one before the fold)one clear ask, earlyn/a
Contact or quote page linked
the page a visitor goes to when the hero did not convert
yeslinked from every pagen/a
Sticky call button (mobile)
a fixed call bar keeps the number reachable at every scroll position on a phone
not detectedpresent on service sitesn/a

The SXO layer: whether a visitor who arrived ready to act can act. Counted from the delivered HTML, so a form inside a third-party iframe shows as a form only if the iframe names one. None of this moves the grade.

Common Crawl visibility — measured, not scored

CCBot is blocked by robots.txt, so future crawls will skip this site. Checked against September 2026 Index, from a stored result.

0 of 6 pages checked were captured.

PageIn the datasetCapturedStatus
/not found——
/sitemap.xml?yyyy=2026&amp;mm=09&amp;dd=24not found——
/sitemap.xml?yyyy=2026&amp;mm=09&amp;dd=23not found——
/sitemap.xml?yyyy=2026&amp;mm=09&amp;dd=22not found——
/sitemap.xml?yyyy=2026&amp;mm=09&amp;dd=21not found——
/sitemap.xml?yyyy=2026&amp;mm=09&amp;dd=20not found——

A capture means the page is in an open dataset. It does not establish that any model trained on it, and nothing here claims otherwise.

Deep entity read — measured, not scored

the deep entity read runs for domains under continuous record; this scan used the keyless resolver

Identity fingerprint — measured, not scored

The identity this site publishes is unchanged since 2026-09-24. Same name, same phone, same address, same declared profiles.

Fingerprint 8153bbf8989417ea971c5250c7185c50… · recorded 2026-09-24

Schema type newsmediaorganization. We hold no population figure for this type, so no rarity is claimed.

A fingerprint of the identity this site publishes, recomputed on every scan. A change here is a change the site made, not a judgement about it. Anyone can recompute this from the published page: the canonical form and the hash are both public at /api/entity-drift/latest.

How this compares — 60 wordpress/publisher sites, of 866 measured
MeasureThis siteCorpus medianRank
AI visibility score (the grade shown)827862nd percentile
Section average (weighted)808230th percentile
Content ratio (% visible text)4.4452nd percentile
Stable @id coverage (%)33.381.810th percentile
Render-blocking resources246.512th percentile
Delivered page size (KB)291176.58th percentile

Percentiles come from sites this scanner has measured itself, not from a published study — so they describe this corpus, not the web. The corpus is not a random sample and is weighted toward sites that were submitted or seeded, which is why the rank is shown next to the raw number rather than instead of it. A metric is left blank rather than ranked when fewer than 8 peers carry it. This rank is against the wordpress/publisher cohort (60 sites), not the whole corpus of 866, because a WordPress local-service site and a static personal site are not the same measurement problem.

Reach, read, quote — the three stages

An engine can reach and read this site. What is left is giving it something specific to say.

StageQuestionScore
Reach
40% of the score
Can a named answer-engine crawler get your pages at all?
At least one answer engine was refused, challenged or served less than a browser.
81
Read
30% of the score
Once it has the bytes, can it find the words?
The content is buried in markup, or the files that guide a crawler are missing.
79
Quote
30% of the score
Is there a specific fact it can state and attribute?
The facts an assistant is asked for are declared and consistent.
85

Which engines, by name

From the probes in this scan — the same fetch each engine’s crawler makes, from our address.

Served the pageOAI-SearchBot · Claude-SearchBot · PerplexityBot · Googlebot · Bingbot · Applebot
Disallowed in robots.txtYahoo Slurp · Yandex (umbrella token) · ChatGPT search index · ChatGPT live fetch · Claude search index · Claude live fetch · Perplexity index · Perplexity live fetch
a policy refusal, which is a decision rather than a defect — it belongs here so it is visible, not because it is wrong

Fix the 24 items this report already lists and the same arithmetic reads 100. That is addition on the table below, not a forecast — every point has a named component behind it:

  • Answer engines allowed — now 8 of 17, worth 1.8 points
  • Search indexes allowed — now 16 of 24, worth 1.8 points
  • Cost of a miss — now 3 missing file(s) return >20KB of HTML, worth 1.3 points
  • entitymap.json — now absent (404), worth 1.3 points

The stage to fix first is read, because the three run in order: an engine cannot read what it was refused, and it cannot quote what it could not read.

What this number covers. Everything above was measured on pages this scan fetched from your own domain. An answer engine also reads sources you do not control — directory listings, review corpora, and your Google Business Profile — and none of those are in this score. The profiles this site declares are listed below but were not fetched on this scan, so nothing here reflects what those listings actually say. A site can score well here and still be described wrongly by an engine reading a listing that disagrees with it.

What this number is not. It does not measure whether ChatGPT, Perplexity or Google’s AI answers actually cite you. No server can measure that, and a score that implied otherwise would be invented. This measures whether an engine can — reach, read, quote — which is the part you control and the part that has to be true first.

Since the last scan

Nothing. Every finding on this scan was present on the previous one at the same severity, and nothing that was open has cleared. Compared against 13 recorded scans of this domain.

Change since last scan — 11 distinct states across 13 scans

Comparing against the last reading that differed, taken 2026-09-24T14:10:19.087Z.

MeasureThenNowChange
Overall score—80/100not comparable
Content ratio4.4%4.4%no change
Stable @id coverage33.3%33.3%no change
Render-blocking resources2424no change
Page size292 KB291 KB-1 KB ✓
Findings88no change

History keeps the last 24 changes on this domain — repeated identical readings collapse into one entry that records when the state began and how many scans confirmed it, so re-running a scan can never push real history out of the window. A finding that comes and goes is the pattern worth acting on: a permanent failure gets noticed, an intermittent one only shows up if something is looking at the right moment.

Since the first scan

Measured on 7 separate days between 2026-08-20 and 2026-09-24. 1 defect has stopped being measurable and 8 are still open.

What this does and does not say. Each row below is two dates: the day a defect was first measured here, and the first day it could no longer be measured. That is evidence the defect is gone. It is not evidence that we caused it to go, and nothing on this page claims otherwise — a site can be fixed for reasons that have nothing to do with a report. What is verifiable is the pair of measurements, and either date can be checked against the sealed record without asking us.

Headline measurements, first reading against latest

MeasureAt first scanNow
GradeFCworse
Overall score7382better
Section average (weighted)7780better
Visible text4.6%4.4%worse
Render-blocking resources2424no change
Page weight279 KB291 KBworse
Open findings48worse

Defects that stopped being measurable

FindingFirst measuredLast measuredDays
The sitemap returns HTML, not XML
SITEMAP_IS_HTML
2026-08-202026-09-1021

Median time from first measurement to no longer measurable: 21 days.

Still open

FindingFirst measuredDays open
Two different entities share the same identifying data
ENTITY_COLLISION
2026-08-2035
Almost nothing a machine receives from this page is readable text
PAGE_IS_MOSTLY_CODE
2026-08-2035
robots.txt rules do not apply to named agents
ROBOTS_RULES_SHADOWED
2026-09-1014
RENDER_RESOURCE_BLOCKED
RENDER_RESOURCE_BLOCKED
2026-09-1014
Published coordinates are not in the country the address names
GEO_COUNTRY_MISMATCH
2026-09-1014
No llms.txt
NO_LLMS_TXT
2026-08-2035
NAP_UNDECLARED_LISTING
NAP_UNDECLARED_LISTING
2026-09-1014
The machine files cost more than the page
MACHINE_CHAIN_HEAVY
2026-09-168

This section is reported and never scored: a site that has improved a lot and a site that was always fine should get the same grade for the same measurements today.

Traffic, enquiries and sales

Not measured here, and not measurable here. Everything else in this report was produced by fetching your pages from outside. Clicks, enquiries and orders happen where no scanner can see them — in your Search Console, your inbox and your checkout. Connect one of those and this section fills in against the dates above. Until then it stays empty rather than showing an estimate, because an invented traffic figure on a report that measures honesty would be the whole product undone.

Cite this report

This report regenerates when it is opened, so both records carry the scan timestamp as the version and the access date as the retrieval. A citation that implied fixed content would be wrong.

BibTeX download

@techreport{crawlcheck_denverpost_com_iahhc6uf5p,
  author      = {{CrawlCheck}},
  title       = {Machine-layer measurement of {denverpost.com}},
  institution = {CrawlCheck},
  type        = {Scan report},
  number      = {iahhc6uf5p},
  year        = {2026},
  url         = {https://crawlcheck.io/r/iahhc6uf5p},
  urldate     = {2026-09-28},
  note        = {Scanned 2026-09-24; report regenerates on access, figures are as of the access date}
}

RIS download

TY  - RPRT
TI  - Machine-layer measurement of denverpost.com
AU  - CrawlCheck
PB  - CrawlCheck
PY  - 2026
DA  - 2026/09/24
M1  - iahhc6uf5p
UR  - https://crawlcheck.io/r/iahhc6uf5p
Y2  - 2026/09/28
N1  - Report regenerates on access; figures are as of the access date
ER  - 

Type is RPRT — a dated report about one named site, not the aggregate dataset at /data. Note that /r/ is disallowed in our robots.txt, so this markup serves readers and reference managers, not search engines.

13

passing

7

needs work

3

failing

7

not measurable

Sections, not individual checks. Not measurable is counted separately and never as a pass — a section we could not read is not a section that passed.

Show this on your site

The badge renders the number as measured, whatever it is — it updates when this report is re-run. It is keyed on this report’s id, so only someone holding this link can produce it.

AI visibility 82 out of 100

Paste this where you want it:

<a href="https://crawlcheck.io/r/iahhc6uf5p?via=badge"><img src="https://crawlcheck.io/badge/iahhc6uf5p.svg" alt="AI visibility 82/100, measured by CrawlCheck" width="228" height="64"></a>

No landing has arrived from this badge yet. The link carries via=badge, so any that do are counted here.

Reading the transcript

CrawlCheck does not claim to predict rankings. It measures the inputs a crawler and a retrieval system must successfully receive before ranking, indexing or quoting is possible at all.

1

Observed

This scanner received this exact status, header, body and rendered-text result. Everything in the sections below is at this level unless it says otherwise — it is a fact about a real response, with the bytes attached.

2

Protocol consequence

What that response necessarily means for a compliant crawler. A disallow rule, a non-text robots.txt or a JavaScript dependency can prevent or impair access by construction — not because we ranked anything, but because the protocol says so.

3

Search implication

Possible or unknown, and labelled as such. No finding here asserts a ranking effect, a citation formula, or a weighting. If a claim cannot be reduced to level 1 or 2, it is not made.

An important limit, stated plainly: the 15 identities above are sent from this scanner’s own address wearing each crawler’s user-agent string. That is an edge-policy simulation — it shows what your edge does with a request that declares itself as that crawler. It is not a request from the real crawler, and it cannot be. A WAF may treat a verified Googlebot differently from anything merely claiming the name, which is exactly what Google recommends it do. Where a result says a crawler was refused, read it as a client declaring that name was refused — to prove what the real crawler received, check your own server logs against the operator’s published address ranges, which the free log verifier does line by line.

Transcript of the machine-file fetches

PathStatusContent typeBytesEdge
/robots.txt 200 text/plain 4,425 BYPASS ok
/sitemap.xml 200 application/xml 884,135 DYNAMIC ok
/sitemap_index.xml 404 text/html 88,178 DYNAMIC not present
/wp-sitemap.xml 200 text/html 142,160 DYNAMIC not present
/llms.txt 404 text/html 88,151 DYNAMIC not present
Reading the scores
The two scores are not a grade
Crawl clarity is whether a machine can fetch, read and tell your pages apart. User clarity is whether a reader can tell what this is and who it belongs to. They fail independently, so read them separately rather than averaging them.
A dash means not measured, not zero
Components shown as — and excluded from the score. If a scan was rate-limited or the pages render with JavaScript, that measurement is withheld. A false clean bill of health is more dangerous than a false defect.
Unconfirmed findings are separated on purpose
Anything that could describe our access rather than your site — a refusal, a timeout — is listed apart and never counted toward the grade.
Content ratio is bytes, not tokens
The share of decompressed bytes that are visible text - the body as a parser reads it, not the smaller gzipped transfer. The deep audit measures tokens per KB with a real tokenizer; the two are named differently because they are different measurements.
What to do with it
Prefer fixes that are one setting over ones that need a sprint. Inline CSS is usually the largest cheap win, because it cannot be cached between pages — a crawler pays for it again on every URL.
What this scan does NOT measure

This is a machine-layer probe of one URL and the files that describe the whole site. Everything below is outside that scope by design, and no score on this page is adjusted for it.

Not a site-wide crawler
There is no full URL graph here, no broken-link inventory across every page, and no duplicate-content comparison. Sitemap coverage is checked against the homepage’s links only. For an exhaustive crawl of thousands of URLs, use a dedicated crawler — Screaming Frog, Sitebulb, or the site audit in Ahrefs or Semrush — and use this to check the machine layer they do not read.
Not a log analysis platform
Crawler behaviour is sampled in one tightly controlled interval, not tracked over weeks. Crawl budget, bot paths, and which pages a crawler actually spends its time on can only come from your own server logs. The free log verifier on this site reads those line by line against the operators’ published address ranges.
Not a replacement for Search Console
Index coverage, canonical selection, manual actions and Google-specific diagnostics come from Search Console and nowhere else. A page can be perfectly readable to every crawler measured here and still not be indexed, for reasons only Google can report.
Extractability, and arrivals - but not citation
This measures whether the bytes an engine receives can be fetched, parsed and attributed to you. Where the first-party beacon is installed it also counts arrivals: humans who clicked through from an AI assistant, recorded from your own site rather than guessed. What it still does not do is ask ChatGPT, Claude, Perplexity or Google whether they mention you when nobody clicks. A citation that produces no click is invisible here, and measuring that needs prompt monitoring, which is a different instrument.
Our vantage point is not the crawler’s
Every request here is honest and unauthenticated: it is sent from a datacentre address under our own name and never impersonates a verified crawler’s IP. A refusal therefore describes what a client declaring that name receives — which a WAF may treat differently from the real crawler, exactly as Google recommends it should.
Not a WCAG audit
The accessibility section reads whether the structure a screen reader depends on exists in the delivered document: a language, a title, names on images, inputs, buttons and frames, a main landmark, a way past the navigation, unique ids. Contrast, focus order, keyboard traps and motion need a rendered page and a person, and are not measured.
One rendering, no JavaScript execution
What is measured is the document the server returns. Content injected by client-side JavaScript is not executed here — which is deliberate, because most crawlers measured on this page do not execute it either, but it means a JavaScript-rendered site will read thinner here than it looks in a browser.
What we could not measure — 2

Every row in this report is either a measurement of your site or nothing at all. These lookups did not return a measurement, so they were left unscored rather than counted against you. They are listed because a silent absence is indistinguishable from a clean result, and that is the difference this tool exists to keep.

This covers the third-party and enrichment lookups only — CrUX, the Google profile link, the Internet Archive, the geocoder, and the two well-known files. Your own server’s refusals are never listed here: a 403 to GPTBot is the finding, not a gap in our instrument, and it is reported as a finding above.

LookupOutcomeWhose limitationWhy
Off-site presence (directories searched by phone number)reachedyoursYelp, Angi, Thumbtack, Houzz, Facebook refuse datacentre clients and are our limit, never an absence. Google Business Profile and review counts are still not read from here
Response compressionunreachedoursour fetch decompresses the body and drops the content-encoding header before we can read it — check it in a browser network panel or with curl --compressed -I

Every figure here was measured with no cookies and no session. That is deliberate: it is what an answer engine gets. If you open your own site while logged into its admin, you are looking at a different document — caching plugins, page builders and membership tools routinely skip their output filters for signed-in users, and admin bars and edit links are added on top.

To see what this report describes, open your site in a private window. If the two disagree, the private window is the one a crawler sees.

Crawl clarity — can a machine fetch, read and tell pages apart 42 / 100
componentstatus scoremeasured
Machine filesscored 60 robots.txt, llms.txt, sitemap and entitymap as scored above
Payload — content ratioscored 26 4.4% of decompressed bytes are visible text (byte-level; the deep audit scores tokens per KB separately)
Render pathscored 0 24 render-blocking resources
Document structurescored 80 0 heading skips, 2 h1
User clarity — can a reader tell what this is and who it belongs to 57 / 100
componentstatus scoremeasured
Entity identityscored 33 33.3% of 3 entities carry a stable @id
Declared standingscored 40 no credential edges declared
Location coherencescored 100 1/1 locations have a street address
Identity claimsscored 0 1 of 1 map links point at places, not the business
Page labellingscored 100 157-char description, canonical present

0 components not measured and excluded from the score — an unmeasured component is never counted as zero.

Score composition — The letter was set by the worst finding, not by the average — that is why a high score can carry a low grade.
SectionScoreWeight
Machine layer60/1001.5
Server and hosting80/1000.6
Schema and entity graph60/1001.2
RDF and linked-data surface100/1000.8
Authorship and trust signals (E-E-A-T proxies)57/1001
Accessibility structure (from the delivered HTML)88/1000.8
Entity corroboration100/1000.6
Quotable content67/1000.8
Crawl waste100/1000.6
Page subject, the Webref question100/1000.6
AI and search agents33/1001.2
Speed a crawler sees, and field data where it exists67/1001
Payload and render path20/1001
Machine-file trust chain100/1001.5
Claim consistency100/1001.3
Agent instruction surface100/1001
What each agent receives88/1001.5
Mobile surface100/1000.8
Local presence, citations and NAP92/1001.3
Sitemap coverage100/1001.2
Internal linking100/1001
Naming and case consistency80/1000.4
Image signals80/1000.6

Not scored, and therefore excluded rather than counted as zero: Agent surfaces (emerging, not scored), Content entities (measured, not yet scored), Service area, as declared (measured, not scored), The whole site, not just the front door (measured, not scored), Freshness, as the site declares it (measured, not scored), Local intents, the Maps vocabulary (measured, not scored), Conversion path on the page (measured, not scored).

Reach, then read, then quote

Reach, then read, then quote — each stage gates the next

REACH81adequateREAD79adequateQUOTE85strong

Watch this site

These break silently and come back on their own. We re-check this site twice a day and record every change against a fingerprint of the last result, so nothing is missed between visits.

Changes are recorded from the moment you sign up, and you get one email when a result moves — nothing otherwise. Every alert carries a one-click stop link. A trial or Watch licence lets you read the record in full.

Change alerts only. Unsubscribe in one click.

Scan another domain

Export this result: JSON · CSV Same stored record as this page, so an export can never disagree with the report.