CrawlCheck

Guides · 2026-10-03 · By · 0 views

PageSpeed Insights and Lighthouse vs an AI crawler audit

What Google's PageSpeed Insights and Lighthouse measure, what a perfect score misses about AI crawlers, and where CrawlCheck falls short on speed.

PageSpeed Insights and Lighthouse measure performance, accessibility, best practices and basic SEO for a person in a browser, with Core Web Vitals field data from real Chrome users. They do not check whether AI crawlers such as GPTBot or ClaudeBot are allowed, served the same content, or able to read it without JavaScript. A site can score 100 and still refuse every AI crawler; a free CrawlCheck scan measures that.

Share of all scans carrying each finding named abovePAGE_IS_MOSTLY_CODE29.2%MACHINE_CHAIN_HEAVY17.8%ANSWER_ENGINE_REFUSED5.9%CONTENT_NEEDS_JAVASCRIPT1.1%CHALLENGE_SERVED_2000.6%Share of all scans carrying eachfinding named abovePAGE_IS_MOSTLY_CODE29.2%MACHINE_CHAIN_HEAVY17.8%ANSWER_ENGINE_REFUSED5.9%CONTENT_NEEDS_JAVASCRIPT1.1%CHALLENGE_SERVED_2000.6%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

PageSpeed Insights and Lighthouse are the default site check for most teams, and a green score is easy to read as “the site is fine”. They measure how fast and well-built a page is for a person in a browser. They do not measure whether AI crawlers are allowed to fetch the page, whether they are served the same content, or whether that content is quotable. A site can score 100 and be refused by every AI crawler.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

What each one measures #

PageSpeed Insights / Lighthouse (Google)CrawlCheck
Visitor it simulatesA browser (lab) and real Chrome users (field data)A browser plus AI and search crawler identities, including GPTBot, ClaudeBot, PerplexityBot and Googlebot
CategoriesPerformance, Accessibility, Best Practices, SEOReach (allowed and served), Read (parseable text), Quote (entity, machine files)
Core Web VitalsYes: LCP, INP, CLS with field data from the Chrome UX ReportNo field Core Web Vitals
robots.txtThe SEO audit checks noindex meta tags and headers that block all crawlers, not robots.txtResolved per crawler for 114 user-agents and compared with what each was served
AI crawlersNot measuredMeasured per identity
PriceFreeFree scans

Sources: About PageSpeed Insights and Lighthouse’s crawlability audit.

What a perfect Lighthouse score misses #

ProblemVisible in Lighthouse?Share of scanned sites (live)
An answer engine’s crawler refused while a browser was servedNo5.9%
Almost nothing a machine receives is readable textNo29.2%
Machine files that cost more to fetch than the pageNo17.8%
A challenge page served with status 200No0.6%
Content that only exists after JavaScript runsPartly: Lighthouse renders it, so the page looks complete1.1%

The last row is the subtle one. Lighthouse runs the page in Chrome, so a page whose text only appears after JavaScript scores normally. Most AI crawlers do not run JavaScript, and receive the empty shell. See do AI crawlers render JavaScript.

One Lighthouse run is not a measurement #

Lighthouse lab scores move between runs on the same page. We measured a 225 percent spread on repeated runs, written up in one Lighthouse run cannot support a finding. Use the field data in PageSpeed Insights where your page has enough traffic, and a median of several lab runs where it does not.

Lab data and field data are different measurements #

PageSpeed Insights shows two kinds of number on one screen, and they answer different questions. The lab score is one Lighthouse run on a simulated mid-range phone over a throttled connection. The field data, when it is shown, is the 75th percentile of real Chrome users who visited the page over the previous 28 days. A page can pass in the field and score poorly in the lab, or the reverse. Google assesses Core Web Vitals on the field data.

MetricGoodPoorWhat it measures
Largest Contentful Paint (LCP)2.5 s or lessover 4 sWhen the main content appears
Interaction to Next Paint (INP)200 ms or lessover 500 msHow quickly the page responds to input
Cumulative Layout Shift (CLS)0.1 or lessover 0.25How much the layout jumps while loading

Thresholds are from About PageSpeed Insights. None of the three applies to a crawler: a crawler does not paint, scroll or click, and it does not count toward the Chrome field data.

Check what an AI crawler receives in two minutes #

You can see the gap yourself before running any tool. Fetch the page as a browser and as GPTBot, and compare the status and size:

curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
  -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/128 Safari/537.36" https://example.com/

curl -s -o /dev/null -w "%{http_code} %{size_download}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot" https://example.com/

A 403 or 429 for the second line, or a response a fraction of the size of the first, means the crawler is treated differently. Then turn JavaScript off in your browser and reload the page: whatever disappears is what most AI crawlers never see. Two caveats keep this honest. A user-agent string is not the real crawler, so a firewall that checks the crawler’s published IP ranges may treat your test and the real GPTBot differently. And one fetch is one moment; challenges and rate limits come and go.

When the two tools disagree #

Lighthouse saysCrawlCheck saysWhat it usually means
Fast, 90+Crawler refusedA firewall or bot rule treats crawlers differently from browsers; speed is irrelevant until access is fixed
Fast, 90+Page is mostly codeThe text is built by JavaScript; server-render the main content
SlowCleanA people problem, not a machine one: fix images, scripts and server time for visitors
SlowCrawler served a challengeOften the same cause: a heavy security layer in front of the site

Where CrawlCheck is weaker #

CrawlCheck does not report field Core Web Vitals, accessibility audits at Lighthouse’s depth, or a performance waterfall. For speed and user experience, PageSpeed Insights is the reference and is free. Use it for people and CrawlCheck for machines; neither replaces the other. More: what each SEO audit tool actually fetches.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

Does PageSpeed Insights check AI crawlers?

No. PageSpeed Insights runs Lighthouse in a browser and shows Chrome user field data. It measures performance, accessibility, best practices and basic SEO for people, not whether GPTBot, ClaudeBot or PerplexityBot are allowed or served the page.

Does Lighthouse check robots.txt?

Lighthouse's crawlability audit checks for noindex meta tags and X-Robots-Tag headers that block all crawlers. It does not resolve robots.txt for individual crawlers such as GPTBot.

Can a site score 100 in Lighthouse and be invisible to ChatGPT?

Yes. A firewall can refuse AI crawlers while serving browsers, or content can exist only after JavaScript runs, which Lighthouse executes and most AI crawlers do not. Neither shows in the score.

What are the Core Web Vitals thresholds?

Per PageSpeed Insights: LCP is good at 2.5 seconds or less, INP at 200 milliseconds or less, and CLS at 0.1 or less, assessed at the 75th percentile of real users.

Should I use PageSpeed Insights or CrawlCheck?

Both, for different visitors. PageSpeed Insights measures speed and experience for people; CrawlCheck measures whether AI crawlers can reach, read and quote the site. Both are free.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works