# CrawlCheck > CrawlCheck is a free AI visibility audit. It fetches your website the way GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot and ten other crawler identities do, checks what each one was actually served, reads the machine files that guide them (robots.txt, llms.txt, sitemaps, the entity graph), scores 22 sections organised as three stages (reach, read, quote) plus foundations, and returns an A–F grade with a ranked fix list and the evidence behind every finding. The scan is free for any domain with no account; paid plans keep a dated record per domain. It never impersonates a crawler to your logs, and a robots.txt line opts a site out of the public dataset. Pages ## Pages - https://crawlcheck.io/llms-txt-audit — llms.txt audit: does the file resolve, parse and match the pages the site declares - https://crawlcheck.io/robots-txt-for-ai-crawlers — robots.txt resolved per AI crawler, compared with what the edge actually served each one - https://crawlcheck.io/entity-map-validator — entitymap.json parsed and checked against the page JSON-LD and every sameAs target - https://crawlcheck.io/aeo-audit — AEO audit: reach, read and quote, graded A to F with a fix list - https://crawlcheck.io/globe — where scans come from: the cities that ran a scan in the last seven days, on a globe; no domain is named - https://crawlcheck.io/registry — the registry: every measured domain filed on the shelves it earned — llms.txt, agents.md, media kit, entity graph, AI access policy, agent-callable - https://crawlcheck.io/press — press kit: the facts, live numbers with dates, the vocabulary, the marks and colours, for writing about CrawlCheck - https://crawlcheck.io/app — the CrawlCheck mobile app, in build: the site's record on the phone, receipts, and a field kit of phone-only instruments; follows /tokens.css and /app.json live - https://crawlcheck.io/forms — a quote form in one script tag that records first- and last-touch source and files each lead beside the site record - https://crawlcheck.io/demo — A live report for a real site, regenerated on every visit - https://crawlcheck.io/ — Free AI visibility audit: scan any domain and see what each AI crawler was actually served, graded A-F with a fix list - https://crawlcheck.io/ai-crawlers — Every named AI crawler we check for, what each is for, and how to verify it - https://crawlcheck.io/tools — Free checks: can a machine read your files, and is a cache serving something stale - https://crawlcheck.io/extension — The browser extension: what a crawler is served from your own address and browser - https://crawlcheck.io/extension/privacy — What the browser extension does with your data — it collects nothing - https://crawlcheck.io/author — Who writes the findings posts here, and every post published. - https://crawlcheck.io/docs/api — The four supported endpoints, and why the rest are not documented. - https://crawlcheck.io/glossary — Definitions for every machine-layer term this scanner measures. - https://crawlcheck.io/guides — Every guide published here, and what each one measures. - https://crawlcheck.io/data — Aggregate findings across every site measured here - https://crawlcheck.io/blog — Measured case studies from real sites - https://crawlcheck.io/pricing — What the free scan and free video check include, and what the paid tiers add: named lists, the record, whole YouTube channels - https://crawlcheck.io/video — YouTube video check: chapters, captions, description, language and topics against what a search result can anchor to - https://crawlcheck.io/video/channel — YouTube channel check: every upload read and tallied against the video anchor model - https://crawlcheck.io/video/site — Does your site carry your videos: embeds, VideoObject schema, video spread across your own pages - https://crawlcheck.io/directory — SEO and AI-visibility tools measured on their own sites — reach, read, quote, re-measured daily; dated change ledger; CSV and JSON at /directory.csv and /directory.json - https://crawlcheck.io/proof — How every report is fingerprinted and anchored to Bitcoin - https://crawlcheck.io/policy — What this scanner requests, what it stores, and how to opt out - https://crawlcheck.io/directory/ahrefs — Ahrefs measured on its own site: what a crawler receives from ahrefs.com - https://crawlcheck.io/directory/semrush — Semrush measured on its own site: what a crawler receives from semrush.com - https://crawlcheck.io/directory/moz — Moz measured on its own site: what a crawler receives from moz.com - https://crawlcheck.io/directory/seranking — SE Ranking measured on its own site: what a crawler receives from seranking.com - https://crawlcheck.io/directory/similarweb — Similarweb measured on its own site: what a crawler receives from similarweb.com - https://crawlcheck.io/directory/screaming-frog — Screaming Frog measured on its own site: what a crawler receives from screamingfrog.co.uk - https://crawlcheck.io/directory/sitebulb — Sitebulb measured on its own site: what a crawler receives from sitebulb.com - https://crawlcheck.io/directory/lumar — Lumar measured on its own site: what a crawler receives from lumar.io - https://crawlcheck.io/directory/botify — Botify measured on its own site: what a crawler receives from botify.com - https://crawlcheck.io/directory/oncrawl — Oncrawl measured on its own site: what a crawler receives from oncrawl.com - https://crawlcheck.io/directory/seobility — Seobility measured on its own site: what a crawler receives from seobility.net - https://crawlcheck.io/directory/siteliner — Siteliner measured on its own site: what a crawler receives from siteliner.com - https://crawlcheck.io/directory/conductor — Conductor measured on its own site: what a crawler receives from conductor.com - https://crawlcheck.io/directory/brightedge — BrightEdge measured on its own site: what a crawler receives from brightedge.com - https://crawlcheck.io/directory/otterly — Otterly measured on its own site: what a crawler receives from otterly.ai - https://crawlcheck.io/directory/peec — Peec AI measured on its own site: what a crawler receives from peec.ai - https://crawlcheck.io/directory/profound — Profound measured on its own site: what a crawler receives from tryprofound.com - https://crawlcheck.io/directory/scrunch — Scrunch measured on its own site: what a crawler receives from scrunch.ai - https://crawlcheck.io/directory/athena — Athena measured on its own site: what a crawler receives from athenahq.ai - https://crawlcheck.io/directory/rankscale — Rankscale measured on its own site: what a crawler receives from rankscale.ai - https://crawlcheck.io/directory/goodie — Goodie measured on its own site: what a crawler receives from higoodie.com - https://crawlcheck.io/directory/writesonic — Writesonic measured on its own site: what a crawler receives from writesonic.com - https://crawlcheck.io/directory/surfer — Surfer measured on its own site: what a crawler receives from surferseo.com - https://crawlcheck.io/directory/clearscope — Clearscope measured on its own site: what a crawler receives from clearscope.io - https://crawlcheck.io/directory/frase — Frase measured on its own site: what a crawler receives from frase.io - https://crawlcheck.io/registry/llms — Registry shelf: llms.txt — The file resolves and its links point at pages that exist on the same host. - https://crawlcheck.io/registry/agents — Registry shelf: agents.md — An /agents.md is served and is not an HTML 404 dressed as one. - https://crawlcheck.io/registry/mediakit — Registry shelf: Media kit — A /.well-known/media-kit.json is served — machine-readable press facts, assets and usage terms. - https://crawlcheck.io/registry/entity — Registry shelf: Entity graph — At least one sameAs was fetched and found pointing back. Reciprocity tested, not declared. - https://crawlcheck.io/registry/aipolicy — Registry shelf: AI access policy — robots.txt names AI crawlers by name rather than leaving them to the wildcard. - https://crawlcheck.io/registry/agentapi — Registry shelf: Agent-callable — An API catalog, MCP server card or agent-skills index is served. - https://crawlcheck.io/blog/the-newsroom-that-published-coordinates-in-china — A LocalBusiness node on denverpost.com declares latitude 39.7392 and longitude 104.9903 - Denver's own centroid with the minus sign missing, which lands 10,668 km away in northern China. The address in the same node is correct. Our coordinate check passed it, and we explain why, what a range test can and cannot catch, and how to check your own. - https://crawlcheck.io/blog/our-scanner-counted-placeholders-not-photographs — The homepage serves 13 webp photographs. The scanner recorded 14 data URIs and zero modern-format images, because a lazyload placeholder is a perfectly good string and the fallback that would have read the real URL never ran. What the fix changed, why the controls stayed still, why no grade moved, and the check you can run on your own markup. - https://crawlcheck.io/blog/our-proofs-said-pending-for-26-days — A reader who ran the check our proof page invites got “pending” on all 26 sealed days, including days whose blocks had confirmed weeks earlier. What the two stages of a timestamp are, why only the second one proves anything, the byte-for-byte check against the reference client, and the daily pass that stops it drifting back. - https://crawlcheck.io/blog/the-viewport-was-there-all-along — Minified HTML with unquoted attributes read as having no viewport, no canonical and no meta description. All three were present. What broke, why the obvious fix was also wrong, and the invariant now watching for it. - https://crawlcheck.io/blog/our-canonical-was-right-and-the-site-was-still-served-twice — A correct canonical tag did not stop www and the apex both answering 200 with the same pages, so a crawler fetched the whole site twice. What the audit showed, why it survives so long, and the check now added to the scanner. - https://crawlcheck.io/guides/is-cloudflare-blocking-ai-crawlers-on-your-site — Cloudflare's AI crawler defaults change on 15 September 2026, including a rule that blocks multi-purpose crawlers such as Googlebot and Applebot when Training crawlers are blocked. How to check what your zone is actually doing, and the failure modes that look like a block but are not. - https://crawlcheck.io/guides/how-to-check-if-ai-crawlers-can-read-your-site — A step-by-step check of whether GPTBot, ClaudeBot and PerplexityBot can actually read your site: content type, cache state, crawler-versus-browser diff, a nonexistent-path control, and text ratio. With the failure rates read live from the public dataset. - https://crawlcheck.io/guides/ai-visibility-tools-what-each-one-measures — A comparison of AI visibility and GEO tools by what they measure rather than what they cost: prompt monitors, SEO-suite modules, free graders, machine-layer scanners and log verification. Includes what happened when 33 of these vendors were scanned with a supply-side scanner. - https://crawlcheck.io/guides/how-to-rank-in-chatgpt — What ranking in ChatGPT actually means, why there is no position to hold, and the four verifiable conditions — reachable, readable, quotable, resolvable — that decide whether an answer engine can use your site. With failure rates read live from the public dataset. - https://crawlcheck.io/guides/do-ai-crawlers-respect-robots-txt — Whether GPTBot, ClaudeBot and PerplexityBot obey robots.txt, why obedience is the wrong thing to worry about, and the policy conflict that fires when robots.txt invites a crawler the edge then refuses. - https://crawlcheck.io/guides/does-llms-txt-actually-work — Whether publishing an llms.txt has any effect: adoption figures, what the file can and cannot do, why the vendors selling the advice mostly have not taken it, and the log test that answers the question for your site specifically. - https://crawlcheck.io/guides/do-ai-crawlers-render-javascript — Whether answer-engine crawlers execute JavaScript, why the safe assumption is that they do not, and a one-command test for whether your content survives in the raw HTML response. - https://crawlcheck.io/guides/ai-crawler-list-what-each-one-does — A reference list of AI crawlers including GPTBot, ClaudeBot, PerplexityBot, Applebot and Google-Extended: what each fetches for, whether the operator publishes verifiable IP ranges, and measured forgery rates per identity. - https://crawlcheck.io/guides/geo-aeo-llm-seo-what-the-terms-mean — Definitions of generative engine optimization, answer engine optimization and LLM SEO, how they differ from each other and from SEO, and which activities under those labels produce verifiable measurements. - https://crawlcheck.io/guides/seo-audit-tools-what-each-one-actually-fetches — A comparison of SEO audit tool categories by what each one requests: desktop crawlers, platform site audits, Lighthouse and PageSpeed, and machine-layer scanners. Includes a measured 225% variance between two runs of the same page. - https://crawlcheck.io/guides/how-to-write-an-llms-txt-and-verify-it — What goes in an llms.txt, how to structure it, whether generators are worth using, and the checks that confirm the file a crawler receives is the file you published. - https://crawlcheck.io/guides/free-ai-visibility-checkers-what-they-measure — What free AI visibility checkers and AEO graders actually test, why high scores are becoming meaningless, and the checks a free tool has to run to tell you anything useful. - https://crawlcheck.io/guides/what-is-entity-seo — What entity SEO means in practice, the structured facts that make a business resolvable, the collision and identity errors that no validator reports, and how to check your own. - https://crawlcheck.io/guides/how-to-measure-ai-visibility — A method for measuring AI visibility: verifying that crawlers can fetch and parse your site, verifying which crawlers actually arrive using published IP ranges, and recording answer-engine presence in a way that can be checked. - https://crawlcheck.io/guides/how-to-block-ai-crawlers — How to block AI crawlers at both the robots.txt and network layers, why the three crawler categories should be decided separately, and what each block actually costs in citation eligibility. - https://crawlcheck.io/blog/ninety-three-percent-of-our-arrivals-sent-no-referrer — We measured arrivals on five sites we operate with our own first-party beacon. 92.9% arrived with no referrer. That number decides what any AI-attribution dashboard can honestly say. - https://crawlcheck.io/blog/the-web-server-that-gained-and-lost-43-million-sites — One survey ranks OpenResty the fourth most popular web server on earth. Another leaves it off the board entirely. Segmented by traffic, 99.81% of the detections sit outside the top million sites. - https://crawlcheck.io/blog/we-clicked-a-real-google-result-to-test-the-goto-redirect — A measured check on the google.com/goto story: what a real Denver Chrome session actually received, why the recommended detection method cannot work, and the one signal that can. - https://crawlcheck.io/blog/one-lighthouse-run-cannot-support-a-finding-225-percent-variance — Two runs, same unchanged URL, minutes apart: TBT 2,643 ms then 812 ms, FCP 2,865 ms then 4,849 ms. What that means for anyone acting on a single PageSpeed score. - https://crawlcheck.io/blog/our-audit-labelled-one-page-and-measured-another — A single trailing slash meant every page-level audit returned homepage data under the correct page's name. How we found it, and why the label agreeing with itself is not verification. - https://crawlcheck.io/blog/ai-text-watermarks-only-work-if-the-provider-opted-in — What SynthID-Text actually does, what the research says about removing it, and why 'Google will detect your AI content' does not follow. - https://crawlcheck.io/blog/a-write-that-returned-success-and-corrupted-every-newline — A backup write reported success and silently converted every line ending. The archive ran, the hashes did not match, nothing flagged it. Verify a snapshot by reading it back. - https://crawlcheck.io/blog/a-783-word-parking-page-defeats-every-thin-content-check — A domain parking page with 783 words passed every thin-content check we ran. Word count is not content. What separates a page that says something from one that fills space. - https://crawlcheck.io/blog/our-analytics-was-adding-650ms-to-every-request — Our own analytics call added 650 milliseconds to every request, and a code comment kept it there for weeks. Per-stage timing found it in one scan. Measure the stage, then blame. - https://crawlcheck.io/blog/a-cache-rule-does-not-evict-what-is-already-cached — Cache rules are not retroactive. The object was stored before the rule existed, so the rule would have taken effect in 2027. - https://crawlcheck.io/blog/a-pipeline-threw-away-our-schema-and-got-us-right-anyway — Two LLM pipelines discarded every JSON-LD node we publish and still described the business correctly. What they read instead, and where entity identity really lives on a page. - https://crawlcheck.io/blog/ai-visibility-scores-measure-an-api-not-chatgpt — Three substitutions sit between an API answer and a person's screen: the surface, the sample and the moment. A number that cannot be wrong is not a measurement. - https://crawlcheck.io/blog/we-almost-published-that-our-host-was-blocking-claude — Hundreds of ClaudeBot 403s in the logs, and we nearly published that our host blocked Anthropic. Grouped by IP, most came from one address wearing seven crawler names. - https://crawlcheck.io/blog/we-scanned-the-tools-that-audit-your-site — We ran 33 AI-visibility and SEO audit tools through our own scanner. Eleven publish an llms.txt, none publish an entity map, and several block the crawlers they grade you on. - https://crawlcheck.io/blog/the-site-that-scored-15-because-it-refused-us — A site scored 15 out of 100 because its firewall refused the scanner, not because the site was bad. A refusal must be reported as a refusal, never scored as a defect. - https://crawlcheck.io/blog/a-header-we-built-to-be-ignored — The design rule behind our self-identification header, and why letting it change a verdict would have handed every forger a bypass. - https://crawlcheck.io/blog/the-grade-that-could-not-explain-itself — An AI read one of our reports and confidently explained a grade cap we had not applied. The report had left the reason out. Now the data carries the explanation, not the reader. - https://crawlcheck.io/blog/we-counted-our-own-users-as-forgers — The tool that measures forged crawler identities was counting its own users' probes as forgeries. Here is how we found it, what the number was, and what it is now. - https://crawlcheck.io/blog/the-crawlers-nobody-can-check — Some crawler operators publish the IP ranges their bots use; some publish nothing. 251 of 950 named-crawler requests could be neither confirmed nor accused. Unverifiable is not forgery. - https://crawlcheck.io/blog/the-ai-crawler-that-wanted-your-env-file — Requests wearing AI-crawler names were probing for credential files, not pages. Telling a real crawler from a scanner using the operator's IP ranges, and why a 403 here is correct. - https://crawlcheck.io/blog/how-much-ai-crawler-traffic-is-forged — Every request claiming to be a known AI crawler, checked against the operator's published IP ranges. A third were not who they said. The per-crawler forged rate, on live traffic. - https://crawlcheck.io/blog/your-site-was-clean-the-day-you-audited-it — Most measured sites carry a crawler-visible defect today; nearly all carried one at some point on record. A one-time audit sees a day. The record sees the other days, and the live figures are on the page. - https://crawlcheck.io/blog/what-a-score-is-measuring-when-it-scores-every-site-the-same — A service business loses points for no checkout; a one-page site for no pagination. What a score measures when it scores every site the same, and the rule a fair one follows. - https://crawlcheck.io/blog/the-meta-tag-a-million-sites-have-and-the-top-ten-thousand-do-not — A tag for a browser that no longer exists is spreading across the web while the largest sites strip it out. The gap between those two lines is the whole story. - https://crawlcheck.io/blog/the-instructions-on-your-domain-you-did-not-write — Two unrelated storefronts served agents.md files 79% identical raw and 100% identical with brand names stripped. A file instructing AI agents on your behalf, written by a template. - https://crawlcheck.io/blog/eight-crawlers-eight-companies-one-decision — Eight named AI crawlers from eight companies, grouped by purpose: training, answering, search. Blocking one is a policy decision, and the robots.txt line differs for each. - https://crawlcheck.io/blog/the-two-facts-almost-nobody-publishes — Across every site scanned here, only a minority publish an llms.txt and far fewer an entity map. The measured adoption rates against BuiltWith's web-wide counts, updated live. - https://crawlcheck.io/blog/what-a-crawler-actually-receives — We measured the bytes an answer engine gets from real local business sites. On most, under 10% is text; the rest is markup and scripts the crawler pays for on every page. - https://crawlcheck.io/blog/the-purge-returned-200-and-evicted-nothing — Four successful-looking cache purges cleared zero objects. Here is the two-request test that catches it, and the script that runs it. - https://crawlcheck.io/blog/a-file-that-answered-200-and-could-not-be-read — A robots.txt that answers HTTP 200 with a bot-challenge page is read as allow-everything by most crawlers. Why every validator missed it, and the two-request test that catches it. - https://crawlcheck.io/blog/which-ai-crawlers-actually-visit-a-small-business-site — Not a probe: real, unasked-for visits to sites we operate, each checked against the operator's published IP ranges. Which AI crawlers turned up, how often, and which were forged. - https://crawlcheck.io/blog/cache-your-misses-a-lookup-that-only-caches-on-success — A lookup that only caches successes re-runs the expensive call on every miss, forever. Measured on a site with no field data: the same failing request on every scan. Cache the miss. - https://crawlcheck.io/blog/two-nodes-with-the-same-id-are-one-node-parsed-twice — Two schema nodes sharing an @id are one record read twice, not two businesses sharing a phone. How our scanner published that false accusation, and the rule that stops it recurring. - https://crawlcheck.io/blog/a-competitor-graded-us-100-percent-and-still-sold-us-the-fix — A competitor's scanner gave crawlcheck.io a perfect score, then recommended a paid fix. What a grade means when the seller also sells the cure, and how to read one. - https://crawlcheck.io/blog/one-character-of-undeliverable-email — Every published contact email on a live business site bounced because of one wrong character in the domain. How long it went unnoticed, and the check that catches it in a second. - https://crawlcheck.io/blog/an-llms-txt-that-answered-200-with-a-login-wall — A large community site answered /llms.txt with HTTP 200 and 8,401 bytes of HTML. Our host-file layer read it as absent; our checks layer read it as present. The self-audit caught the disagreement, and the fix changes what a 200 is allowed to mean. - https://crawlcheck.io/blog/the-scanner-that-cannot-fetch-its-own-host — A site's founder Person node listed crawlcheck.io/author as a sameAs profile. The page is live from any client on earth, and our corroboration check reported it dead, because a Cloudflare Worker cannot make a subrequest to its own route. Unverifiable from this vantage is not dead. - https://crawlcheck.io/blog/nan-inside-a-url-slug-a-placeholder-check-that-accused-a-newspaper — A self-audit invariant hunts stored records for placeholders that have leaked into output: NaN, undefined, [object Object]. It fired on a national newspaper's record. The three NaNs were the letters n-a-N inside a live-blog URL slug. A substring is not a token. - https://crawlcheck.io/blog/e-e-a-t-is-not-on-the-wire-what-a-scanner-can-measure-instead — We added fourteen rows that measure the machine-readable statements a site makes about who is behind it: a named person, what qualifies them, whether anyone corroborates them, dated and authored articles, an organisation that names its founder. None of them is expertise. All of them are checkable. - https://crawlcheck.io/blog/accessibility-structure-nine-things-a-screen-reader-needs-before-contrast-matters — We added an accessibility section that reads only what the server delivers: a language, a title, names on images, inputs, buttons and frames, one main landmark, a way past the navigation, unique ids. Two of the sites we operate had no main landmark at all. Two false positives were caught before it shipped. - https://crawlcheck.io/blog/we-read-every-scored-row-against-seven-sites-here-is-what-we-unscored — We pulled the pass/fail for every scored row on seven real reports into one matrix and read every row that looked wrong. Seven rows accused sites of things that are not defects; four were counting a defect twice or measuring cosmetics. What changed, what stayed, and how a score version is bumped honestly. - https://crawlcheck.io/guides/how-to-read-a-crawlcheck-report — Every CrawlCheck report opens on an Overview and keeps five more views behind tabs: the fix list, every section scored, what changed since the last scan, the evidence transcript, and how to read the scoring. This guide walks through each, explains how the grade is derived from the three stages, and shows where to dispute a finding on its bytes. - https://crawlcheck.io/guides/how-to-check-if-a-gptbot-request-in-your-logs-is-real — Requests wearing AI crawler names are routinely forged; on one site every GPTBot 403 in a week came from a single address wearing seven crawler identities. This guide shows how to verify crawler claims against the IP ranges each operator publishes, using the free verifier or the API, and why the answer has three outcomes rather than two. - https://crawlcheck.io/blog/state-of-ai-visibility-september-2026 — Seven findings from the CrawlCheck dataset as of 2 September 2026, every one sourced to the open data page or a published corpus threshold: which defects are most common, how often they cap a grade, why AI opt-outs move in lockstep, how few homepages are written to be quoted, and how much crawler traffic is not who it claims to be. - https://crawlcheck.io/blog/a-tree-service-lead-arrived-with-utm-source-chatgpt-com — On 2 September 2026 a quote form on treeservicedenverllc.com recorded a submission whose landing URL carried utm_source=chatgpt.com. This post explains where that tag comes from, why it is the only attribution that survives the trip, what the site's record looked like in the week before, and why a single lead is recorded as one data point rather than a claim. - https://crawlcheck.io/guides/youtube-seo-case-study-81-videos-rewritten-measured-before-and-after — A worked example of the YouTube check on a real channel: what the tally found, what was changed on 1 September and 5 September 2026, the windows the result will be read across, and the measurement design that will be applied on 3 October. It names what the case can and cannot show. - https://crawlcheck.io/blog/your-schema-lists-service-areas-your-sitemap-has-pages-for-a-fraction-of-them — Local service sites declare where they work in schema (areaServed) and in their pages. This scanner reads both and reports the gap. On four sites we operate: 23 declared / 21 with a page, then 19 / 7, 19 / 1 and 20 / 1. Why the over-promising direction is a real defect a schema validator will not catch, and the two ways the check itself was tightened to measure it honestly. ## Machine files - [entitymap.json](https://crawlcheck.io/entitymap.json): the entity graph for this site - [agents.md](https://crawlcheck.io/agents.md): instructions for agents reading this site ## What it measures - Whether the machine files resolve, and whether a 200 actually contains the file - What each named crawler receives compared with a browser, from one address in one second - How many of the decompressed bytes are visible text - Whether declared entities, coordinates and profile links resolve to the thing they claim