Directory · SEO suite
Semrush — what a crawler receives from semrush.com
Answer engines can reach and read this site. It needs more it can quote.
AI visibility 78/100, measured 2026-09-03 00:17 UTC from outside its network, as fifteen named crawler identities and one unnamed client. An engine can reach and read this site. What is left is giving it something specific to say.
Held at C A medium-severity finding sets the letter.
3
findings
✓
llms.txt on its host
✗
entity map on its host
6
scans on record since 2026-08-14
The three stages, in the order an engine hits a site
Each stage gates the next. An engine that cannot reach a page never reads it; one that cannot read it never quotes it. That is why reach is a ceiling and not just a weight.
Stage 1 · 40% of the score
Reach
Can a named answer-engine crawler get your pages at all?
Every named crawler was served the same page a browser gets.
Stage 2 · 30% of the score
Read
Once it has the bytes, can it find the words?
The content is buried in markup, or the files that guide a crawler are missing.
Stage 3 · 30% of the score
Quote
Is there a specific fact it can state and attribute?
An assistant would have to paraphrase your page instead of quoting a fact.
The stage to fix first is read, because the three run in order.
What to fix, in order
Findings first: each one holds the grade down until it is gone (the top one sets the letter). Then headroom: each one lifts the score by the points shown. Open a card for what we saw, why it matters and how to fix it.
Files, generated from this scan
Each file says at its top exactly what was changed or filled and what was left as a TODO. Publish, then re-scan: done is measured.
1robots.txt rules do not apply to named agents10 named User-agent groups (10 agents) declare no Disallow rules while the `*` group carries 55. A crawler obeys only its own most-specific group and ignores `*Medium
What we saw
10 named User-agent groups (10 agents) declare no Disallow rules while the `*` group carries 55. A crawler obeys only its own most-specific group and ignores `*` entirely, so those rules do not apply to any of the named agents. Worst group: `008` bypasses /blog/*?img=*, /archive/graphs.php, /forrester-marketing-silos/*, /partner/, /blog/search/*, /analytics/seomagic/, /my_reports/reports/*, /archive/, /admin/, /users/, /clients/, /custom_report/, /limit_hits/, /payment/, /no_cookies/, /double_activation/, /projects/?*, /?*/, /activation_successful*, /api.html?*. This includes a SEARCH crawler (bingbot) — the bypassed paths are being crawled and can enter the index. /robots.txt
Why it matters
A crawler reads only the User-agent group that matches it best and ignores every other group, including `*`. Naming an agent and giving it only `Allow: /` therefore deletes all of your `*` Disallow rules for that agent. The file still parses and the agent still reaches your homepage, so this is invisible on inspection — but the paths you meant to keep out of search are open, and search engines will crawl and may index them.
How to fix it
Every User-agent group that names a crawler must repeat the Disallow rules you want applied. A named group with no rules means 'allow everything' for that crawler.
ROBOTS_RULES_SHADOWED · full transcript
2Almost nothing a machine receives from this page is readable text4.3% of 197,576 decompressed bytes is visible textMedium
What we saw
4.3% of 197,576 decompressed bytes is visible text
Why it matters
A crawler pays for every byte it fetches and can only quote the text. Under 5% readable content means an answer engine downloads the whole page and comes away with almost nothing it can use — the difference between being quotable and being skipped. The content-ratio row above passes at 20%, so without this a page at 2% and a page at 19% were scored identically.
How to fix it
Move inline CSS and JavaScript into external files, drop unused page-builder styling, and make sure the words a visitor reads are in the HTML itself.
PAGE_IS_MOSTLY_CODE · full transcript
3The sitemap does not say when anything changed23 child sitemaps listed, none carrying a lastmod date.Low
What we saw
23 child sitemaps listed, none carrying a lastmod date. /sitemap.xml
Why it matters
The sitemap lists URLs but carries no lastmod dates, so it says what exists and not what changed. A crawler with a budget uses lastmod to decide what to re-fetch; without it, the choice is made for you — usually by re-crawling little and late. This is the change signal every crawler already understands, and unlike a push protocol it needs no key, no account and no participation from the search engine.
How to fix it
Add <lastmod> dates to sitemap entries so a crawler can prioritise pages that changed.
SITEMAP_NO_LASTMOD · full transcript
Headroom +19 points if every row passes
Nothing here is broken. Each row is a measure that currently fails its optimal range, and what fixing it is worth to the score.
Watch this domain
Twice-daily scans, the date each finding first appeared, and a change receipt when something moves. Free for 30 days, no card.
Start watching semrush.comGet this fixed for you
Every finding above fixed on your site, then re-measured - done is measured, not asserted. One-time, from $750.
See the fix serviceThe chain a crawler follows
The chain breaks at entity graph. Everything after that point is only reachable by a crawler guessing the conventional path.
A machine file that does not parse is worth less than one that is absent, because it looks present. Rows behind this.
Where the bytes go
Of the 197,576 decompressed bytes the homepage delivers, 4.3% is text a reader or a model can actually use. Most of what a crawler downloads here is not words.
- Readable text 4.3% · 8,534 B
- Markup & attributes 84.8% · 167,587 B
- Structured data (JSON-LD) 0.4% · 798 B
- Inline JavaScript 9% · 17,759 B
- Inline SVG 1.1% · 2,173 B
- HTML comments 0.4% · 725 B
What changed
- 2026-09-01score moved 79 → 81score-up
- 2026-09-01a new finding appeared: the sitemap does not say when anything changedappeared
- 2026-08-25a named-agent group shadowed the wildcard rulesappeared
- 2026-08-18resolved: robots.txt blocks the entire siteresolved
6 scans on record, first 2026-08-14. Score movements are shown only between readings taken under the same score version; finding codes are comparable across all of them.
Is this your site?
This listing reads the public record. Claim it and Watch keeps that record on your terms: every list this page holds back, finding age, and twice-daily re-measurement. Free for 30 days, no card.
Claim the record for semrush.comDispute or re-scan
If a reading here is wrong, it is our instrument that is wrong, and we want to know. Email hello@crawlcheck.io and the site is re-measured; the result is published as measured. Nothing about a listing, a payment or a request changes a number.
Category “SEO suite” is our label for what the product is primarily sold as; it is not scored. Back to the directory.