CrawlCheck

Findings · 2026-10-04 · By · 0 views

402 Payment Required: what top sites send AI crawlers

Three of the top 1,000 sites answer AI crawlers with HTTP 402. None quotes a price: one is a refusal, one a toll, one contradicts its own robots.txt.

On 4 October 2026, three operators in the Tranco top 1,000 answered AI crawlers with HTTP 402 Payment Required: x.com, forbes.com and independent.co.uk. None quoted a price. On x.com the 402 is a refusal under another code; on Forbes it is a toll through a separate paid-access host; on The Independent it contradicts robots.txt, which allows the crawlers it charges. To a crawler, an unpriced 402 works like a block.

HTTP 402 Payment Required sat unused in the status-code list for decades. It is now how some large sites answer AI crawlers. We read the top 1,000 sites of the Tranco list on 4 October 2026, each homepage once as a browser and once as each of four AI crawlers: GPTBot, ClaudeBot, OAI-SearchBot and PerplexityBot. Three operators answered at least one of those crawlers with 402. None of the three quoted a price.

Disclosure: CrawlCheck publishes this and runs the census it comes from. Every reply described here was read by our census and then fetched again by hand on 4 October 2026. A crawler name in a request is a claim, not an identity: these are the answers a server gave to that claim from our address, which may differ from what the real crawler receives from its own published addresses.

What the census found #

Of the 991 domains read, 524 served a homepage to a plain browser request. Most of the rest are infrastructure domains with no website at all, such as CDN, DNS and update hosts; others refused the read or timed out. Among the 524, four domains answered an AI crawler with 402. Two of them are one service: twitter.com redirects to x.com.

SiteRankAnswered 402robots.txt for those crawlersWhat the 402 contained
x.com (and twitter.com)48 (15)GPTBot, OAI-SearchBot, PerplexityBot. ClaudeBot got 403disallows all four55 bytes of JSON: “Please contact the site owner for access.”
forbes.com222GPTBot, ClaudeBot, OAI-SearchBot, after a redirect to tollbit.forbes.comdisallows GPTBot and ClaudeBot, allows OAI-SearchBotJSON: no access “without a valid TollBit Token”, with a link to tollbit.dev
independent.co.uk727ClaudeBot, PerplexityBotallows all fourAn HTML page titled “Payment Required”

The x.com and Independent answers were the same in the two earlier daily runs. Forbes did not answer the census at all in those runs, so its pattern is one day old in our data, and we confirmed it by hand.

Three different things behind one status code #

The same three digits mean three different things here. On x.com, 402 is a refusal under another name: robots.txt already disallows the crawler, there is no price and no way to pay, and the message says to contact the site owner. A crawler can do nothing with it except stop.

On forbes.com, 402 is a toll. The crawler is redirected to a separate host run for a paid-access service, and that host refuses it without a token from the service. There is a path to payment, but it runs through an account with that service, not through anything in the reply itself.

On independent.co.uk, 402 contradicts robots.txt. The file invites ClaudeBot and PerplexityBot, and the server asks them for payment. GPTBot and OAI-SearchBot, also invited, got the full homepage. A site owner reading only robots.txt would believe all four crawlers can read the site.

What a priced 402 is supposed to look like #

There is a published design for 402 as a real price. Cloudflare's pay per crawl, in closed beta, answers an AI crawler with 402 and a crawler-price header. The crawler retries with crawler-exact-price or crawler-max-price, signs the request with Web Bot Auth so the payment can be tied to a verified operator, and on success gets 200 with a crawler-charged header naming what was billed. A failed attempt comes back as 402 with a crawler-error header.

None of the four replies in the top 1,000 carried a crawler-price header. So on 4 October 2026, every 402 an AI crawler got from a top-1,000 homepage was either a refusal or a toll run through another service. None was a price quoted in the reply.

Why this matters if you run a site #

Crawlers do not treat 402 as an invitation. To a search or answer crawler, a 402 on your homepage is a page it could not read, like a 403. If your goal is to be cited, a 402 does the same job as a block, whatever the intent behind it.

If a 402 comes from a setting rather than a decision, it is easy to miss. robots.txt says yes, a browser sees the page, and only the crawler sees the bill. The Independent's split, two crawlers served and two charged, is the kind of result that usually comes from a rule set per crawler at the edge or in a paid-access integration, not from anything in robots.txt.

If you mean to charge, the reply should say so in a way a crawler can act on: a price, and a way to pay that a crawler operator can use. A bare 402 with a contact address tells an automated client nothing it can do.

How to check your own site #

Request your homepage with a crawler's user-agent and look at the status line, any redirect and the headers:

curl -sI -L -A "Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)" https://www.example.com/

Repeat with each crawler you care about and once with a normal browser user-agent. Then compare:

What you seeWhat it means
200 for the browser and for the crawlerThe crawler is served. robots.txt still decides whether a polite crawler reads on.
402 with a crawler-price headerA priced offer the crawler can accept or decline.
402 with no price, or a redirect to another host that answers 402A refusal, or a toll run by a third-party service. Find which setting or integration produces it.
402 or 403 for a crawler your robots.txt allowsThe two disagree. Decide which one states your intent and change the other.

One caution. Some servers check that a request claiming to be a crawler comes from that crawler's published addresses, and refuse it from anywhere else. If your server refuses a request claiming to be Googlebot as well, your test is measuring that address check, not your crawler policy. Test from a normal connection, and read your access logs for what the real crawler got.

A CrawlCheck scan makes this comparison for every crawler at once, and the guides on whether Cloudflare is blocking AI crawlers and how to block AI crawlers on purpose cover the settings that usually sit behind a mismatch.

What this does not show #

This is one homepage per site, read from our address, with four crawler names. It does not show what the real crawlers receive from their own addresses, what any of these sites charge or to whom, or whether any crawler has paid. It shows only what the status line and the reply said when we asked. The census figures for the whole top 1,000 are on the data page.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What does HTTP 402 Payment Required mean for a crawler?

It means the server wants payment before it serves the page. A crawler that cannot pay treats it like any other refusal: the page is not read, so it cannot be indexed or cited.

Do sites actually send 402 to AI crawlers?

Yes. On 4 October 2026, three operators in the Tranco top 1,000 answered at least one AI crawler with 402: x.com, forbes.com and independent.co.uk. None of their replies quoted a price.

What does a priced 402 look like?

Cloudflare's pay per crawl, in closed beta, answers with 402 and a crawler-price header. The crawler retries with crawler-exact-price or crawler-max-price, signed with Web Bot Auth, and gets 200 with a crawler-charged header when payment succeeds.

Can robots.txt allow a crawler that the server charges?

Yes, and it happened in the top 1,000: The Independent's robots.txt allows ClaudeBot and PerplexityBot, and its server answers both with 402. robots.txt and the server are separate settings and can disagree.

How do I check what my site sends an AI crawler?

Request your homepage with curl and the crawler's user-agent, and compare the status, redirect and headers with a browser request. If a request claiming to be Googlebot is refused too, you are testing an address check, not your crawler policy.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All findings · The dataset · How the dataset works