CrawlCheck
Free AI visibility audit · ChatGPT, Claude, Perplexity, Google

See what AI crawlers actually receive from your site — graded A–F, free.

Can ChatGPT, Claude, Perplexity and Google reach you, read you and safely quote you — and what to fix first.

CrawlCheck is a free AI visibility audit: it fetches your site the way AI crawlers do, grades what they get A–F and counts every defect by severity. We implement the fix list for $749.

No signup. Nothing installed. Nothing changed on your site.

The main product

Reading the list is free. Having it fixed is $749.

One site. Every finding implemented and re-measured, with a sealed before and after. Thirty days of monitoring included.

  • Reach — can AI crawlers get your pages?
  • Read — do they get real words and files that parse?
  • Quote — is there a clear fact they can repeat about you?

Paying never changes a score.

4,462

sites measured

15

crawler identities checked

20,055

real crawler visits logged

0

numbers a payment can change

Built and run by one person in Denver, on the client sites it measures — the demo report is one of ours.

Pricing

Scan free. Pay only to keep watching.

Monthly or yearly, in US dollars. Every button with a price on it is a live Stripe checkout — Agency is scoped with a person first — and paying never changes a score.

Free · any site

$0

  • Unlimited scans, no account, no card
  • A–F grade, the score, and every defect counted by severity
  • Checked as GPTBot, ClaudeBot, Googlebot and more
Scan a site free

Fix list implemented · one site

$749one-time

  • We implement the fix list from your report
  • WordPress: one plugin puts the machine-readable fixes live the same day
  • Machine files, schema and NAP, and the findings holding the grade down
  • Re-measured afterwards — a sealed before and after, verifiable without trusting us
  • Thirty days of monitoring included
  • Starts within 2 business days
Get this fixed

Bigger estates are quoted against the report — what this covers

Watch · 3 sites

$29/month

  • Every page behind each finding, named
  • Re-checked twice a day, with a dated history
  • A change receipt between any two dates — what moved, proved, not asserted
  • A lead form whose source is measured, not guessed
  • 1 YouTube channel
Start Watch

or $290 a year — 2 months free

Agency · 40 sites

$249/month

  • Scan as much as you like — the licensed lane has no daily cap
  • The API on your key, and the CI gate
  • Client reports with your name on them
  • Identity Guard: which crawlers really reached the server
  • 40 domains, everything Watch names on all of them
Start Agency

or $2,490 a year — 2 months free

How paying works

  1. 1

    Pay on Stripe

    Card, Apple Pay or Link, on Stripe’s secure checkout.

  2. 2

    Get your key

    It is shown on the next page. Save it: the key is your account, there is no password.

  3. 3

    Add your sites

    Paste the key at /manage and list your domains.

  4. 4

    Cancel any time

    Email hello@crawlcheck.io from the address you paid with.

Not ready to pay? Try Watch free for 30 days on one site, no card. More than 40 sites? Agency, quoted. Want it fixed for you? The fix list implemented, $749. Compare every feature →

What is measured

Three stages, in the order an engine hits them.

A refusal at the first makes the other two unreachable, however good they are — which is why the report scores them separately, in order, and never averages them into one number that hides which stage failed.

01 · Reach15 identities1 address<1s
66

Can an AI crawler reach you

66 on the live report above · weak
What it checks

15 client identities are sent to your homepage from one address inside one second, and every reply is kept — status, bytes, content type, cache header. Same second, same address, so a difference between them is your edge deciding, not the internet being busy.

  • GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot, Applebot, CCBot, Bytespider, Amazonbot, Meta-ExternalAgent, plus the OpenAI and Anthropic retrieval agents
  • A plain browser and a mobile browser as controls, so “everyone gets this” and “only crawlers get this” are distinguishable
  • A randomised path that cannot exist, so a wall that answers everything with 200 is told apart from a site
  • Which answer engines were served, refused at the edge, or disallowed by policy — by name
02 · Read5 files2 hosts92 agents
70

Can it read your machine files and your words

70 on the live report above · adequate
What it checks

5 machine files are fetched on both hosts — apex and www — because a redirect on one and a policy on the other is a split nobody notices. Every named agent in your robots.txt is resolved into what it may actually do.

  • robots.txt, llms.txt, sitemap.xml, sitemap_index.xml, entitymap.json — status, content type and body, not just a status code
  • 114 named agents parsed from robots.txt, each mapped to index, train or live-fetch, with shadowed rules called out
  • How much of the delivered page is readable text versus markup, scripts and inline styles; render-blocking resources; the sitemap against the pages the homepage links to
03 · QuoteNAPcitations fetchedclaims
87

Can it quote a fact about you

87 on the live report above · strong
What it checks

An engine that reaches and reads you still needs a fact it is willing to state. This stage is about whether your identity is unambiguous enough to be repeated by something that will be blamed if it is wrong.

  • Name, phone and address cross-checked between your schema and your visible page — agreeing with yourself is the whole test
  • Every directory and profile your site declares is fetched, not counted: does it load, and does it still carry your number
  • Claims a machine will repeat — years in business, founding date, service area — checked against what the page says, plus authorship and dated-article signals an engine uses for trust
Beyond the three stages, and the live report’s crawler chain and page weight

04 · Beyond the three stages

open datasetssaliencedated identity

Measured, not scored

Three rows that answer questions the grade deliberately does not. None of them moves your score — a site is not worse because an open dataset has not visited it, or because a paid API was not called on a free scan. They are here because they are true and nobody else reports them.

  • Common Crawl visibility — whether your pages are in the open dataset most models were trained from, checked per URL against the current crawl release. Three outcomes, never two: captured, absent, or could not be checked.
  • Deep entity read — which entity a language model thinks the page is about, and whether Google already holds a Knowledge Graph record for it.
  • Identity fingerprint — a SHA-256 of the name, phone, address, schema type and declared profiles your site publishes, recomputed every scan and dated, so a change becomes a moment on a record and the report names which field moved.

The chain a crawler follows

The chain breaks at llms.txt. Everything after that point is only reachable by a crawler guessing the conventional path.

robots.txtEvery crawler reads this first
sitemapNamed in robots.txt and resolves
llms.txtPoints onward, links stay on this host
entity graphParses as JSON and points home

A machine file that does not parse is worth less than one that is absent, because it looks present. Rows behind this.

Where the bytes go

Of the 298,738 decompressed bytes the homepage delivers, 4.3% is text a reader or a model can actually use. Most of what a crawler downloads here is not words.

  • Readable text 4.3% · 12,984 B
  • Markup & attributes 65.2% · 194,810 B
  • Structured data (JSON-LD) 0.7% · 2,238 B
  • Inline CSS 16% · 47,674 B
  • Inline JavaScript 12.8% · 38,096 B
  • Inline SVG 0.1% · 372 B
  • HTML comments 0.9% · 2,564 B

The individual blocks that weigh the most, so the fix is a search rather than a hunt:

  • Inline CSS · 46,114 B (15.4%) — unnamedopens @import url(https://fonts.googleapis.com/css2?family=Inter:ital,opsz,w
  • Inline JavaScript · 14,462 B (4.8%) — #digisubs-social-login-js-afteropens (function () { var logPrefix = '[digisubs-social-login]'; var readyTim
  • Inline JavaScript · 3,089 B (1%) — WordPress emoji scriptopens /*! This file is auto-generated */ var e="script#wp-emoji-settings",t=
Before you scan

Questions people ask, and what it will not do.

QWhat is AI visibility?
AI visibility is whether an answer engine — ChatGPT, Claude, Perplexity, Google’s AI answers — can reach your pages, read the words on them, and quote a specific fact about you with attribution. CrawlCheck measures those three stages in that order and reports an AI visibility score out of 100. It does not measure whether an engine actually cites you; no server can, and a score that implied otherwise would be invented.
QWhich AI crawlers does CrawlCheck test?
Every scan fetches the homepage as GPTBot, ClaudeBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot, Amazonbot, Bytespider, Meta-ExternalAgent and CCBot, plus a plain browser, a mobile browser and an unnamed client as controls. It also resolves what robots.txt says to 114 named agents across answer, search, training, regional, social and SEO groups.
QDoes it check llms.txt, robots.txt and structured data?
Yes. It reads robots.txt, llms.txt, sitemap.xml, sitemap_index.xml and entitymap.json on both the apex and www hosts, by status, content type and body — a robots.txt that returns HTTP 200 with an HTML challenge page is reported as unreadable, not present. Structured data is scored as an entity graph: stable @ids, connected nodes, credentials, locations and whether the name, phone and address agree with the visible page.
QIs CrawlCheck free?
The scan and the full report are free for any domain, with no account. Paid plans (from $29 a month for three domains) keep a dated record, measure twice a day, send change receipts and add branded reports, a defect ledger and an API lane. Nothing a licence unlocks changes a measurement, a finding or a grade.
QDoes this change anything on my site?
No. Every request is a read — the same GET a crawler makes. Nothing is submitted, no form is filled, no page is written to, and nothing is installed.
QWill my domain show up anywhere public?
No. Aggregate findings are published on the dataset page and no scanned domain is ever named there. One line in robots.txt keeps you out of the dataset entirely.
QDoes it work for local businesses and service companies?
Yes, and it was built on them. The quote stage checks street address, coordinates, hours and phone, cross-checks name, phone and city between the structured data and the visible page, and fetches every directory and profile the site declares to see whether each one still resolves and still carries the business’s number. A service business is never docked for having no checkout: a section that does not apply is excluded, not scored zero.
QHow is the grade calculated?
Each of the 24 scored sections is rated 0–100 as the share of its measures inside the optimal range. The three stages average their sections; reach is 40% of the score and also a ceiling, because nothing downstream of a refusal reaches an engine. The letter is set by that score and then capped by the worst finding: a medium finding holds a site at C, high at D, critical at F. A section whose inputs could not be measured is left out rather than counted as zero.
QIs this just an SEO audit with new words on it?
No. An SEO audit reads your page as a browser. This sends the requests an answer engine sends — as GPTBot, ClaudeBot, PerplexityBot and the rest — and reports what your edge gave each one, next to what a browser got in the same second. The findings it exists for are the ones a browser can never see: a challenge page served with a 200, a crawler handed fewer words than a visitor, a robots.txt that names a sitemap which does not resolve.
QIs this GEO, AEO or LLM SEO?
Those names all describe the same goal: being reachable, readable and quotable by answer engines. CrawlCheck is the measurement layer for that work — it tells you which stage fails and what to change — and it publishes how every number is derived. What it will not do is promise a citation, because no measurement can.
QCan I use it from code or from my own browser?
Yes. POST /api/scan takes a domain and returns the same record the report is rendered from; four endpoints are documented at /docs/api and the OpenAPI file at /openapi.json. The browser extension measures what a crawler is served from your own address and after your JavaScript runs — the half a server cannot see. A free log verifier checks whether a request claiming to be GPTBot or ClaudeBot came from that operator’s published IP ranges.
QDoes CrawlCheck check YouTube?
Yes. Paste a video URL at /video and it is read on the metadata a search result can anchor to: authored chapters, uploaded versus auto-generated captions, description depth, language agreement, topic placement. Watch adds the whole channel, tallied, and whether your own site carries your videos with VideoObject schema.
QCan I pay to improve my score?
No, and that is deliberate. Nothing a licence unlocks changes a measurement, a finding or a grade. The moment a payment can move a result, every result here is worth less — including the ones you would be paying for.
QIs there a mobile app?
One is in build for iPhone and Android from a design frozen on 5 September 2026. It reads the same record the site keeps, follows the site's own tokens and section list live, and adds a field kit of phone-only instruments. It is not on the app stores yet; a plan today is not waiting on it, because the record keeps measuring twice a day either way.
In detail

The full method, one click away.

An AI crawler audit, a machine-file check and a local-business claim check, in one report.
What this is

An AI crawler audit, a machine-file check and a local-business claim check, in one report.

Answer engines do not read your site the way a browser does. They send named crawlers, read a handful of machine files first, and quote only what they can attribute. This measures each of those from outside your network and shows you the bytes.

Reach

15 crawler identities fetch your homepage from one address in one second: GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot, Applebot and more, with a browser and a mobile browser as controls.

Read

5 machine files on 2 hosts — robots.txt, llms.txt, sitemap.xml, sitemap_index.xml, entitymap.json — by status, content type and body, and 114 named agents resolved into what each may actually do.

Quote

Structured data as an entity graph, name-phone-address agreement, authorship signals, and every directory listing the site declares fetched and checked for the business’s number.

Output

An A–F grade, an AI visibility score out of 100, 24 scored sections, every defect counted by severity, and a URL that regenerates so it stays true.

One scan, end to end
Your siteHomepage + 5 machine fileson the apex and the www host
Fetched as15 crawler identitiesGPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot… in one second from one address
Scored asReach → Read → Quote24 scored sections, each row against its optimal range
You getA–F grade + defect countfree; the findings and their evidence are what the fix service and a licence buy
Every section a report scores, in the order an engine hits a site.
Every check, by name

Every section a report scores, in the order an engine hits a site.

This is the full list of what one free scan measures. Each name below is a section on every report, with its rows, its optimal range and the evidence bytes one click down.

Reach Can a named answer-engine crawler get your pages at all?

  • What each crawler getsFifteen client identities fetched the page; this is whether they all got the same thing
  • Crawler permissionsWhat robots.txt actually tells each named AI and search crawler it may do
  • Machine-file chainrobots.txt names the sitemap, llms.txt points onward, the entity graph points home: the chain a crawler follows
  • SpeedHow fast the page answered our fetch, plus real-visitor field data where Google has it

Read Once it has the bytes, can it find the words?

  • Page weight & textHow many of the delivered bytes are readable words, and how much is markup, scripts and styling
  • Machine filesrobots.txt, llms.txt, sitemap and the entity graph: present, readable and served as the right type
  • SitemapA sitemap exists, resolves, and declares the pages the homepage links to
  • Crawl wasteFetches that produce no new page: parameter and case variants, http links, redirecting or gone declared URLs, canonical-elsewhere and noindex pages. Scored from version 23

Quote Is there a specific fact it can state and attribute?

  • Structured dataThe JSON-LD entity graph: stable @ids, connected nodes, credentials and locations declared
  • Local presence & NAPStreet address, coordinates, hours, phone, and the directory listings the site itself declares
  • Claim consistencyName, phone, city and years-in-business match between the structured data and the visible page
  • Authors & trust signalsNamed people, dated articles, an organisation with contact details: the proxies an engine uses for trust
  • Linked dataWhether the structured data parses as proper RDF and nothing dangles
  • Quotable contentWhether the sentences an answer engine would lift are on the page in a form it can take: a defining lead, short factual paragraphs, answered questions, few first-person openers
  • Entity corroborationWhether the identity links the site declares point back at it, and whether a Wikidata record names this domain
  • Page subjectWhich business entity the page is about and how every other business entity on it is related: chain, subsidiary, department, or undeclared. Scored from version 23

Foundations Scored, outside the three stages

  • Server & hostingWhat the edge and origin told us: cache state, compression, the software behind the page
  • Accessibility structureLanguage, alt text, labelled controls and one main landmark, read from the delivered HTML
  • Instructions for agentsText aimed at AI agents rather than people: present, visible, and not hidden
  • MobileViewport, zoom, icons and manifest: whether a phone gets a page it can use
  • Image signalsAlt text, declared dimensions, lazy-loading, modern formats, hero priority and the entity image in JSON-LD.
  • Commerce surfaceProduct, offer and checkout signals an agent could act on
  • Internal linksAnchors carry words, targets resolve, and the link graph is not mostly repeats
  • URL namingLowercase hosts and paths, no case-only duplicates, machine files spelt correctly

Reported, not scored Measured for information; they never move a grade

  • Agent surfacesEmerging standards (agents.md, API catalog, MCP card). Measured for information only
  • Content entitiesNamed things in the page text that a knowledge graph recognises. Measured, not scored
  • Verified crawler trafficWhat verified AI and search crawlers actually fetched from this zone in the last seven days, counted by Cloudflare at the edge. Outcome data, not scored
  • Crawlers on your serverWhich AI and search crawlers actually reached the origin, from the site’s own access log posted under the licence, each address verified against the operator’s published range. The outcome, not scored
  • Service area mapEvery place the schema declares, placed on one map around the declared centre and radius: which have a page, which fall outside the circle. Measured, not scored
  • Whole siteUp to 60 declared pages, read once a day after a scan: not-200s, thin and mostly-code pages, missing JSON-LD, canonical elsewhere, noindex, missing or duplicate h1, orphans, undeclared pages, duplicate titles, and the median words, text share and response time across them. Every count is free; Watch names which pages. Measured, not scored
  • FreshnessEvery date the site publishes about itself: sitemap lastmod, Last-Modified, dateModified, dated elements, and where they disagree. Measured, not scored
  • Local intentsWhich of the 446 recovered Maps intent types the site names about itself, and which of those have a page the homepage links to. Measured, not scored
  • Conversion pathPhone link, form and call-to-action on the page, and whether they sit before the fold. Measured, not scored
  • EngagementFirst-party views, active time, scroll depth and taps from the site’s own beacon. Measured, not scored
  • Entity map parityThe JSON entity map and the HTML version agree with each other. Measured, not scored
The same sections, scored on the live demo report — denverpost.com, grade C · 74/100

Grey tiles were measured but could not be scored, and count for nothing rather than zero. Click any tile to open that section on the live report.

A row that could not be measured says so and scores nothing — it is never counted as a zero. That is why a report on a service business does not lose points for having no checkout, and why the grade you see is the grade the evidence supports. Definitions for every term are in the glossary.

Four steps, about a minute, nothing installed.
How it works

Four steps, about a minute, nothing installed.

The scanner identifies itself honestly on every request and never impersonates a crawler to your logs; it sends the identities to your site, not to your analytics.

  1. 1

    Enter a domain

    Type any domain into the box above. No account, no card, nothing installed on your site.

  2. 2

    About 30 read-only requests are sent

    The scanner fetches the homepage as 15 client identities from one address inside one second, then reads robots.txt, llms.txt, both sitemaps and the entity map on the apex and www hosts, plus a randomised path that cannot exist as a control.

  3. 3

    Read the report

    An A–F grade, an AI visibility score out of 100, the three stages scored in order, every section scored, and how many defects you have with their severity. The fix list — what was seen, why it matters, how to fix it and what it is worth — is what the fix service and a licence buy. The URL regenerates when opened, so it stays true.

  4. 4

    Fix, re-scan, keep a record

    Change what the server sends and scan again: done is measured, not asserted. Paid plans keep a dated record per domain, measure twice a day and send a receipt when something moves.

None of these show up in a validator, a rank tracker or an uptime check.
Measured on real sites

None of these show up in a validator, a rank tracker or an uptime check.

Every site in these examples looked fine to the people who built it. No scanned domain is named here or anywhere public.

What the report reads past the first page.
Beyond the homepage

What the report reads past the first page.

Whole site, daily

Up to 60 declared pages read once a day: not-200s, thin pages, missing JSON-LD, orphans. The homepage is where it starts, not where it stops.

Verified crawler traffic

What verified AI and search crawlers fetched from your zone in the last seven days, counted at the edge and on your own server, each address checked against the operator’s published range.

Service area map

Every place your schema declares, placed around the declared centre, and whether a real page exists for it. Three of four sites we operate over-promised badly.

YouTube, checked properly

A video or a whole channel checked on the metadata a search result can anchor to — chapters, captions, description, topics — and whether your site carries your videos at all. A real channel is the case study.

A form that knows the source

One script tag that records first touch and last touch — ChatGPT, Google, a campaign tag, direct — and files each lead beside the record.

The app, in build

The record in your hand: your grade on the home screen, receipts as they happen, and a field kit of instruments only a phone can run. It follows this site’s tokens and sections live.

Anyone whose customers now ask an assistant before they search.
Who it is for

Anyone whose customers now ask an assistant before they search.

Local service businesses

Tree care, fencing, contractors, clinics, firms: the quote stage checks the address, phone, hours and directory listings an assistant repeats when someone asks who to call.

Agencies and consultants

40 client sites under one record, branded reports with the internal notes off, a defect ledger with first-seen and close dates, and receipts a client can verify without trusting you.

SaaS and publishers

Whether the docs, the pricing page and the articles reach an engine as words rather than as a JavaScript shell, and whether the entity graph gives it a fact to cite.

Developers and platform teams

A JSON record per scan, 4 documented endpoints, an OpenAPI 3.1 file, webhook signing, and evidence bytes on every finding so a fix can be argued on the bytes.

One scanner, reachable from wherever you work.
Ways to use it

One scanner, reachable from wherever you work.

Free scanAny domain, full report, no account. The form at the top of this page./ →APIPOST /api/scan returns the record the report renders from. Four documented endpoints and an OpenAPI file; a keyed lane for paid plans./docs/api →Browser extensionMeasures what a crawler is served from your own address and after your JavaScript runs. It sends nothing back./extension →Free toolsPaste access-log lines and every crawler claim is checked against the operator’s published IP ranges: verified, spoofed, or unverifiable./tools →Live badgeAn SVG that shows the current score and links to the report. It moves when the measurement moves./proof →Directory and datasetSEO and AI-visibility vendors measured on their own sites daily, with a change ledger; aggregate findings across every site measured here, CC BY./directory →AI crawler referenceEvery named agent this scanner checks, what it is for, and how to verify it from its published IP ranges./ai-crawlers →PressFacts, live numbers with dates, the product's vocabulary, and the wordmark and mark as vector files, for writing about CrawlCheck/press →AppThe CrawlCheck mobile app, in build for iPhone and Android: the site's dated record on the phone, receipts as they happen, and a field kit of phone-only instruments. Follows /tokens.css and /app.json live/app →YouTubeA video checked on the metadata a search result can anchor to: chapters, captions, description, topics. Whole channels and the site tie-in with Watch/video →MonitoringThe lists the free report holds back — which pages, which listings, who links — and the dated record per domain, with change receipts and a webhook when something moves/pricing →
Free scan, a report you can send, then a record nobody can back-fill.
What you walk away with

Free scan, a report you can send, then a record nobody can back-fill.

  1. 1

    Scan, free

    One minute, nothing installed, no address asked. 15 client identities fetch your homepage, 5 machine files are read on both hosts, and 114 named agents are resolved from your robots.txt.

    Scan a domain
  2. 2

    The report

    A grade, a score, which of the three stages fails, and how many defects you have with their severity — free, on a URL that regenerates so it stays true.

    See a live one
  3. 3

    Watch, free for 30 days

    Every list the free report holds back — which pages, which listings, who links — plus the date each finding first appeared. Measured twice a day. No card.

    Start the trial
  4. 4

    The record

    Change receipts a third party can verify, hashes anchored to Bitcoin daily, a badge that moves when the score moves. From $29 a month for 3 domains; the days you are not measuring are the days no tool can recover.

    The four tiers
Four things this is not, so you can pick the right tool.
What it is not

Four things this is not, so you can pick the right tool.

CrawlCheck measures what the server sends to a named crawler and whether an answer engine could state a fact about you from it. It does not replace an SEO audit, a rank tracker, or your own logs.

A rank trackerIt does not report where you rank in Google or how often an assistant cites you. No server can measure the second, and the first is a different product.
A JavaScript rendererIt measures what the server sends. What your page becomes after scripts run is the browser extension’s job, from your own machine.
A schema validatorValidators check syntax. This scores whether the structured data forms a usable entity graph and agrees with the visible page.
An uptime monitorIt watches meaning, not availability: a page that answers 200 with the wrong bytes is a finding here and a pass everywhere else.

Start with the free scan.

If the machine layer is already clean, you will know in a minute and pay nothing. No login, nothing changed on your site, and no domain is ever named publicly.

CrawlCheck AI visibility mark for crawlcheck.io, live score This site wears its own badge. The number is our scan of ourselves and moves when it moves.

Scan a domain — free