CrawlCheck
Free AI visibility audit · ChatGPT, Claude, Perplexity, Google

See what AI crawlers actually receive from your site — graded A–F, with a fix list.

Can ChatGPT, Claude, Perplexity and Google reach you, read you and safely quote you — and what to fix first.

CrawlCheck is a free AI visibility audit. It fetches your website the way GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot and ten other crawler identities do, checks what each one was actually served, reads the machine files that guide them (robots.txt, llms.txt, sitemaps, the entity graph), scores 22 sections organised as three stages (reach, read, quote) plus foundations, and returns an A–F grade with a ranked fix list and the evidence behind every finding.

No signup. Nothing installed. Nothing changed on your site. We only read pages, the way a crawler does — about 30 requests in total.

See a real report first — a live Denver contractor, scanned in full →

Free forever for the scan. The lists behind the counts, and the record, from $29/mocheckout in one click →

  • Reach — can GPTBot, ClaudeBot, PerplexityBot and Googlebot get your pages at all, or does the edge turn them away while browsers get 200.
  • Read — do they get real words, machine files that parse, and a sitemap that resolves — not a challenge page, a shell, or a stale cache.
  • Quote — is there a specific, consistent fact (name, phone, address, hours, a claim) an engine can repeat with attribution.

You cannot pay to improve a score. A licence adds the lists behind the counts and the record over time — never a better grade.

Where scans come from →

Last scan finished 34 minutes ago · 75 domains measured today

2,193

sites measured

15

crawler identities checked

17,617

real crawler visits logged

0

numbers a payment can change

Built and run by one person in Denver, on the client sites it measures — the demo report is one of ours. The scan is free, always. Keep a record of 3 sites for $29/mo →

Pricing · checkout in one click

The scan is free. Keeping a record starts at $29 a month.

Pick the number of domains you want kept under a dated record and start today; every button below is a live checkout. Annual pays for ten months. Nothing a licence unlocks changes a measurement. Full comparison and one-time work →

Watch · 3 domains · 1 channel

$29/month

Best for a business owner or a solo consultant with a few sites: the lists behind every count, and the record.

The free scan counts; Watch names — which pages are thin, orphaned or missing schema, which directories list you, who links to you — keeps the record that says when each defect began, and checks one YouTube channel whole: every upload against the video anchor fields, and whether your site carries your videos.

Start watch

or pay for the year, pay for ten months

Everything included →

Studio · 10 domains · 3 channels

$89/month

Best for a boutique running client sites: everything Watch shows, in a report with your name on it, plus what your server actually saw.

Everything Watch shows, on ten domains, in a client-ready report — plus Identity Guard, and the YouTube transcript layer through the browser extension: caption kind, what is said versus written, storyboard frames.

Start studio

or pay for the year, pay for ten months

Everything included →

Agency · 40 domains · 10 channels

$249/month

Best for an agency answering for a portfolio: the API lane, the portfolio view, regression alerts across it.

Forty domains, the portfolio view, regression alerts across all of them, the licensed API lane, and ten YouTube channels with a dated change record per video.

Start agency

or pay for the year, pay for ten months

Everything included →

Network · 150+ domains

From $690/month

150 domains and up, on a cadence you set, with cohort reporting across your own estate and bulk receipts. Your numbers are never published, named or pooled.

Quoted per estate, not priced

How a quote works →

Not ready to pay? Try Watch on one domain for 30 days, no card. The record it builds is real and stays readable after it ends.

Nothing a licence unlocks changes a measurement. The moment a payment can move a result, every result here is worth less — including the ones you would be paying for.

What this is

An AI crawler audit, a machine-file check and a local-business claim check, in one report.

Answer engines do not read your site the way a browser does. They send named crawlers, read a handful of machine files first, and quote only what they can attribute. This measures each of those from outside your network and shows you the bytes.

Reach

15 crawler identities fetch your homepage from one address in one second: GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot, Applebot and more, with a browser and a mobile browser as controls.

Read

5 machine files on 2 hosts — robots.txt, llms.txt, sitemap.xml, sitemap_index.xml, entitymap.json — by status, content type and body, and 92 named agents resolved into what each may actually do.

Quote

Structured data as an entity graph, name-phone-address agreement, authorship signals, and every directory listing the site declares fetched and checked for the business’s number.

Output

An A–F grade, an AI visibility score out of 100, 22 scored sections, a fix list with how-to-fix text, evidence bytes on every finding, JSON and CSV exports, and a URL that regenerates so it stays true.

One scan, end to end
Your siteHomepage + 5 machine fileson the apex and the www host
Fetched as15 crawler identitiesGPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot… in one second from one address
Scored asReach → Read → Quote22 scored sections, each row against its optimal range
You getA–F grade + fix listevidence bytes on every finding, a URL that regenerates
What is measured

Three stages, in the order an engine hits them.

A refusal at the first makes the other two unreachable, however good they are — which is why the report scores them separately, in order, and never averages them into one number that hides which stage failed.

01 · Reach15 identities1 address<1s

Can an AI crawler reach you

15 client identities are sent to your homepage from one address inside one second, and every reply is kept — status, bytes, content type, cache header. Same second, same address, so a difference between them is your edge deciding, not the internet being busy.

  • GPTBot, ClaudeBot, PerplexityBot, Googlebot, Bingbot, Applebot, CCBot, Bytespider, Amazonbot, Meta-ExternalAgent, plus the OpenAI and Anthropic retrieval agents
  • A plain browser and a mobile browser as controls, so “everyone gets this” and “only crawlers get this” are distinguishable
  • A randomised path that cannot exist, so a wall that answers everything with 200 is told apart from a site
  • Which answer engines were served, refused at the edge, or disallowed by policy — by name
02 · Read5 files2 hosts92 agents

Can it read your machine files and your words

5 machine files are fetched on both hosts — apex and www — because a redirect on one and a policy on the other is a split nobody notices. Every named agent in your robots.txt is resolved into what it may actually do.

  • robots.txt, llms.txt, sitemap.xml, sitemap_index.xml, entitymap.json — status, content type and body, not just a status code
  • 92 named agents across 23 groups parsed from robots.txt, each mapped to index, train or live-fetch, with shadowed rules called out
  • How much of the delivered page is readable text versus markup, scripts and inline styles; render-blocking resources; the sitemap against the pages the homepage links to
03 · QuoteNAPcitations fetchedclaims

Can it quote a fact about you

An engine that reaches and reads you still needs a fact it is willing to state. This stage is about whether your identity is unambiguous enough to be repeated by something that will be blamed if it is wrong.

  • Name, phone and address cross-checked between your schema and your visible page — agreeing with yourself is the whole test
  • Every directory and profile your site declares is fetched, not counted: does it load, and does it still carry your number
  • Claims a machine will repeat — years in business, founding date, service area — checked against what the page says, plus authorship and dated-article signals an engine uses for trust

04 · Beyond the three stages

open datasetssaliencedated identity

Measured, not scored

Three rows that answer questions the grade deliberately does not. None of them moves your score — a site is not worse because an open dataset has not visited it, or because a paid API was not called on a free scan. They are here because they are true and nobody else reports them.

  • Common Crawl visibility — whether your pages are in the open dataset most models were trained from, checked per URL against the current crawl release. Three outcomes, never two: captured, absent, or could not be checked.
  • Deep entity read — which entity a language model thinks the page is about, and whether Google already holds a Knowledge Graph record for it.
  • Identity fingerprint — a SHA-256 of the name, phone, address, schema type and declared profiles your site publishes, recomputed every scan and dated, so a change becomes a moment on a record and the report names which field moved.
Every check, by name

Every section a report scores, in the order an engine hits a site.

This is the full list of what one free scan measures. Each name below is a section on every report, with its rows, its optimal range and the evidence bytes one click down.

Reach Can a named answer-engine crawler get your pages at all?

  • What each crawler getsFifteen client identities fetched the page; this is whether they all got the same thing
  • Crawler permissionsWhat robots.txt actually tells each named AI and search crawler it may do
  • Machine-file chainrobots.txt names the sitemap, llms.txt points onward, the entity graph points home: the chain a crawler follows
  • SpeedHow fast the page answered our fetch, plus real-visitor field data where Google has it

Read Once it has the bytes, can it find the words?

  • Page weight & textHow many of the delivered bytes are readable words, and how much is markup, scripts and styling
  • Machine filesrobots.txt, llms.txt, sitemap and the entity graph: present, readable and served as the right type
  • SitemapA sitemap exists, resolves, and declares the pages the homepage links to

Quote Is there a specific fact it can state and attribute?

  • Structured dataThe JSON-LD entity graph: stable @ids, connected nodes, credentials and locations declared
  • Local presence & NAPStreet address, coordinates, hours, phone, and the directory listings the site itself declares
  • Claim consistencyName, phone, city and years-in-business match between the structured data and the visible page
  • Authors & trust signalsNamed people, dated articles, an organisation with contact details: the proxies an engine uses for trust
  • Linked dataWhether the structured data parses as proper RDF and nothing dangles
  • Quotable contentWhether the sentences an answer engine would lift are on the page in a form it can take: a defining lead, short factual paragraphs, answered questions, few first-person openers
  • Entity corroborationWhether the identity links the site declares point back at it, and whether a Wikidata record names this domain

Foundations Scored, outside the three stages

  • Server & hostingWhat the edge and origin told us: cache state, compression, the software behind the page
  • Accessibility structureLanguage, alt text, labelled controls and one main landmark, read from the delivered HTML
  • Instructions for agentsText aimed at AI agents rather than people: present, visible, and not hidden
  • MobileViewport, zoom, icons and manifest: whether a phone gets a page it can use
  • Image signalsAlt text, declared dimensions, lazy-loading, modern formats, hero priority and the entity image in JSON-LD.
  • Commerce surfaceProduct, offer and checkout signals an agent could act on
  • Internal linksAnchors carry words, targets resolve, and the link graph is not mostly repeats
  • URL namingLowercase hosts and paths, no case-only duplicates, machine files spelt correctly

Reported, not scored Measured for information; they never move a grade

  • Agent surfacesEmerging standards (agents.md, API catalog, MCP card). Measured for information only
  • Content entitiesNamed things in the page text that a knowledge graph recognises. Measured, not scored
  • Verified crawler trafficWhat verified AI and search crawlers actually fetched from this zone in the last seven days, counted by Cloudflare at the edge. Outcome data, not scored
  • Crawlers on your serverWhich AI and search crawlers actually reached the origin, from the site’s own access log posted under the licence, each address verified against the operator’s published range. The outcome, not scored
  • Service area mapEvery place the schema declares, placed on one map around the declared centre and radius: which have a page, which fall outside the circle. Measured, not scored
  • Whole siteUp to 60 declared pages, read once a day after a scan: not-200s, thin and mostly-code pages, missing JSON-LD, canonical elsewhere, noindex, missing or duplicate h1, orphans, undeclared pages, duplicate titles, and the median words, text share and response time across them. Every count is free; Watch names which pages. Measured, not scored
  • FreshnessEvery date the site publishes about itself: sitemap lastmod, Last-Modified, dateModified, dated elements, and where they disagree. Measured, not scored
  • Crawl wasteFetches that produce no new page: parameter and case variants, http links, redirecting or gone declared URLs, canonical-elsewhere and noindex pages. Measured, not scored
  • Conversion pathPhone link, form and call-to-action on the page, and whether they sit before the fold. Measured, not scored
  • EngagementFirst-party views, active time, scroll depth and taps from the site’s own beacon. Measured, not scored
  • Entity map parityThe JSON entity map and the HTML version agree with each other. Measured, not scored

A row that could not be measured says so and scores nothing — it is never counted as a zero. That is why a report on a service business does not lose points for having no checkout, and why the grade you see is the grade the evidence supports. Definitions for every term are in the glossary.

How it works

Four steps, about a minute, nothing installed.

The scanner identifies itself honestly on every request and never impersonates a crawler to your logs; it sends the identities to your site, not to your analytics.

  1. 1

    Enter a domain

    Type any domain into the box above. No account, no card, nothing installed on your site.

  2. 2

    About 30 read-only requests are sent

    The scanner fetches the homepage as 15 client identities from one address inside one second, then reads robots.txt, llms.txt, both sitemaps and the entity map on the apex and www hosts, plus a randomised path that cannot exist as a control.

  3. 3

    Read the report

    An A–F grade, an AI visibility score out of 100, the three stages scored in order, every section scored, and a fix list that says what was seen, why it matters, how to fix it and what it is worth. The URL regenerates when opened, so it stays true.

  4. 4

    Fix, re-scan, keep a record

    Change what the server sends and scan again: done is measured, not asserted. Paid plans keep a dated record per domain, measure twice a day and send a receipt when something moves.

Measured on real sites

None of these show up in a validator, a rank tracker or an uptime check.

Every site in these examples looked fine to the people who built it. No scanned domain is named here or anywhere public.

Beyond the homepage

What the report reads past the first page.

Whole site, daily

Up to 60 declared pages read once a day: not-200s, thin pages, missing JSON-LD, orphans. The homepage is where it starts, not where it stops.

Verified crawler traffic

What verified AI and search crawlers fetched from your zone in the last seven days, counted at the edge and on your own server, each address checked against the operator’s published range.

Service area map

Every place your schema declares, placed around the declared centre, and whether a real page exists for it. Three of four sites we operate over-promised badly.

YouTube, checked properly

A video or a whole channel checked on the metadata a search result can anchor to — chapters, captions, description, topics — and whether your site carries your videos at all. A real channel is the case study.

A form that knows the source

One script tag that records first touch and last touch — ChatGPT, Google, a campaign tag, direct — and files each lead beside the record.

The app, in build

The record in your hand: your grade on the home screen, receipts as they happen, and a field kit of instruments only a phone can run. It follows this site’s tokens and sections live.

Who it is for

Anyone whose customers now ask an assistant before they search.

Local service businesses

Tree care, fencing, contractors, clinics, firms: the quote stage checks the address, phone, hours and directory listings an assistant repeats when someone asks who to call.

Agencies and consultants

40 client sites under one record, branded reports with the internal notes off, a defect ledger with first-seen and close dates, and receipts a client can verify without trusting you.

SaaS and publishers

Whether the docs, the pricing page and the articles reach an engine as words rather than as a JavaScript shell, and whether the entity graph gives it a fact to cite.

Developers and platform teams

A JSON record per scan, 4 documented endpoints, an OpenAPI 3.1 file, webhook signing, and evidence bytes on every finding so a fix can be argued on the bytes.

Ways to use it

One scanner, reachable from wherever you work.

Free scanAny domain, full report, no account. The form at the top of this page./ →APIPOST /api/scan returns the record the report renders from. Four documented endpoints and an OpenAPI file; a keyed lane for paid plans./docs/api →Browser extensionMeasures what a crawler is served from your own address and after your JavaScript runs. It sends nothing back./extension →Free toolsPaste access-log lines and every crawler claim is checked against the operator’s published IP ranges: verified, spoofed, or unverifiable./tools →Live badgeAn SVG that shows the current score and links to the report. It moves when the measurement moves./proof →Directory and datasetSEO and AI-visibility vendors measured on their own sites daily, with a change ledger; aggregate findings across every site measured here, CC BY./directory →AI crawler referenceEvery named agent this scanner checks, what it is for, and how to verify it from its published IP ranges./ai-crawlers →PressFacts, live numbers with dates, the product's vocabulary, and the wordmark and mark as vector files, for writing about CrawlCheck/press →AppThe CrawlCheck mobile app, in build for iPhone and Android: the site's dated record on the phone, receipts as they happen, and a field kit of phone-only instruments. Follows /tokens.css and /app.json live/app →YouTubeA video checked on the metadata a search result can anchor to: chapters, captions, description, topics. Whole channels and the site tie-in with Watch/video →MonitoringThe lists the free report holds back — which pages, which listings, who links — and the dated record per domain, with change receipts and a webhook when something moves/pricing →
What you walk away with

Free scan, a report you can send, then a record nobody can back-fill.

  1. 1

    Scan, free

    One minute, nothing installed, no address asked. 15 client identities, 5 machine files on both hosts, 92 named agents.

    Scan a domain
  2. 2

    The report

    A grade, a ranked fix list with the points attached, the evidence bytes behind each finding, and a URL that regenerates so it stays true.

    See a live one
  3. 3

    Watch, free for 30 days

    Every list the free report holds back — which pages, which listings, who links — plus the date each finding first appeared. Measured twice a day. No card.

    Start the trial
  4. 4

    The record

    Change receipts a third party can verify, hashes anchored to Bitcoin daily, a badge that moves when the score moves. From $29 a month for 3 domains; the days you are not measuring are the days no tool can recover.

    The four tiers
What it is not

Four things this is not, so you can pick the right tool.

CrawlCheck measures what the server sends to a named crawler and whether an answer engine could state a fact about you from it. It does not replace an SEO audit, a rank tracker, or your own logs.

A rank trackerIt does not report where you rank in Google or how often an assistant cites you. No server can measure the second, and the first is a different product.
A JavaScript rendererIt measures what the server sends. What your page becomes after scripts run is the browser extension’s job, from your own machine.
A schema validatorValidators check syntax. This scores whether the structured data forms a usable entity graph and agrees with the visible page.
An uptime monitorIt watches meaning, not availability: a page that answers 200 with the wrong bytes is a finding here and a pass everywhere else.
Before you scan

Questions people ask, and what it will not do.

QWhat is AI visibility?
AI visibility is whether an answer engine — ChatGPT, Claude, Perplexity, Google’s AI answers — can reach your pages, read the words on them, and quote a specific fact about you with attribution. CrawlCheck measures those three stages in that order and reports an AI visibility score out of 100. It does not measure whether an engine actually cites you; no server can, and a score that implied otherwise would be invented.
QWhich AI crawlers does CrawlCheck test?
Every scan fetches the homepage as GPTBot, ClaudeBot, OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot, Applebot, Amazonbot, Bytespider, Meta-ExternalAgent and CCBot, plus a plain browser, a mobile browser and an unnamed client as controls. It also resolves what robots.txt says to 92 named agents across answer, search, training, regional, social and SEO groups.
QDoes it check llms.txt, robots.txt and structured data?
Yes. It reads robots.txt, llms.txt, sitemap.xml, sitemap_index.xml and entitymap.json on both the apex and www hosts, by status, content type and body — a robots.txt that returns HTTP 200 with an HTML challenge page is reported as unreadable, not present. Structured data is scored as an entity graph: stable @ids, connected nodes, credentials, locations and whether the name, phone and address agree with the visible page.
QIs CrawlCheck free?
The scan and the full report are free for any domain, with no account. Paid plans (from $29 a month for three domains) keep a dated record, measure twice a day, send change receipts and add branded reports, a defect ledger and an API lane. Nothing a licence unlocks changes a measurement, a finding or a grade.
QDoes this change anything on my site?
No. Every request is a read — the same GET a crawler makes. Nothing is submitted, no form is filled, no page is written to, and nothing is installed.
QWill my domain show up anywhere public?
No. Aggregate findings are published on the dataset page and no scanned domain is ever named there. One line in robots.txt keeps you out of the dataset entirely.
QDoes it work for local businesses and service companies?
Yes, and it was built on them. The quote stage checks street address, coordinates, hours and phone, cross-checks name, phone and city between the structured data and the visible page, and fetches every directory and profile the site declares to see whether each one still resolves and still carries the business’s number. A service business is never docked for having no checkout: a section that does not apply is excluded, not scored zero.
QHow is the grade calculated?
Each of the 22 scored sections is rated 0–100 as the share of its measures inside the optimal range. The three stages average their sections; reach is 40% of the score and also a ceiling, because nothing downstream of a refusal reaches an engine. The letter is set by that score and then capped by the worst finding: a medium finding holds a site at C, high at D, critical at F. A section whose inputs could not be measured is left out rather than counted as zero.
QIs this just an SEO audit with new words on it?
No. An SEO audit reads your page as a browser. This sends the requests an answer engine sends — as GPTBot, ClaudeBot, PerplexityBot and the rest — and reports what your edge gave each one, next to what a browser got in the same second. The findings it exists for are the ones a browser can never see: a challenge page served with a 200, a crawler handed fewer words than a visitor, a robots.txt that names a sitemap which does not resolve.
QIs this GEO, AEO or LLM SEO?
Those names all describe the same goal: being reachable, readable and quotable by answer engines. CrawlCheck is the measurement layer for that work — it tells you which stage fails and what to change — and it publishes how every number is derived. What it will not do is promise a citation, because no measurement can.
QCan I use it from code or from my own browser?
Yes. POST /api/scan takes a domain and returns the same record the report is rendered from; four endpoints are documented at /docs/api and the OpenAPI file at /openapi.json. The browser extension measures what a crawler is served from your own address and after your JavaScript runs — the half a server cannot see. A free log verifier checks whether a request claiming to be GPTBot or ClaudeBot came from that operator’s published IP ranges.
QDoes CrawlCheck check YouTube?
Yes. Paste a video URL at /video and it is read on the metadata a search result can anchor to: authored chapters, uploaded versus auto-generated captions, description depth, language agreement, topic placement. Watch adds the whole channel, tallied, and whether your own site carries your videos with VideoObject schema.
QCan I pay to improve my score?
No, and that is deliberate. Nothing a licence unlocks changes a measurement, a finding or a grade. The moment a payment can move a result, every result here is worth less — including the ones you would be paying for.
QIs there a mobile app?
One is in build for iPhone and Android from a design frozen on 5 September 2026. It reads the same record the site keeps, follows the site's own tokens and section list live, and adds a field kit of phone-only instruments. It is not on the app stores yet; a plan today is not waiting on it, because the record keeps measuring twice a day either way.

Start with the free scan.

If the machine layer is already clean, you will know in a minute and pay nothing. No login, nothing changed on your site, and no domain is ever named publicly.

Scan a domain — free