Guides · 2026-10-02 · By VSNARY | Emmanuel Orta · 0 views
Web Bot Auth: how signed crawler requests are verified
IP lists stop working when crawlers share cloud addresses. Web Bot Auth has the crawler sign each request with a key it publishes, using HTTP Message Signatures. The three headers, the key directory, how a site verifies, and what a signature does not prove.
Web Bot Auth is an IETF draft for authenticating automated clients with HTTP Message Signatures (RFC 9421). The crawler signs each request with an Ed25519 private key and sends three headers: Signature, Signature-Input (with created, expires, keyid and tag="web-bot-auth") and Signature-Agent, which points to the operator's key directory at /.well-known/http-message-signatures-directory, served as application/http-message-signatures-directory+json. A site fetches the directory, finds the key by its JWK thumbprint, checks the signature and the time window, and then knows the request came from whoever holds that key.
The two classic ways to verify a crawler, reverse DNS and published IP ranges, both assume the operator owns its addresses. That assumption is failing. More automated traffic now comes from shared cloud infrastructure, browser-based agents and third-party fetch services, where the same IP serves many tenants. Web Bot Auth replaces the network-level check with a cryptographic one: the crawler signs each request, and the site checks the signature against a key the operator publishes. This guide covers how it works, how to verify a request, and the limits of what a valid signature tells you.
The building blocks #
Web Bot Auth is defined in IETF Internet-Drafts and builds on RFC 9421, HTTP Message Signatures, which specifies how to sign selected parts of an HTTP message. The operator generates an Ed25519 key pair, keeps the private key, and publishes the public key in a JSON Web Key Set at /.well-known/http-message-signatures-directory on its own domain, served with the media type application/http-message-signatures-directory+json. Each key is identified by its JWK thumbprint, a hash of the key's canonical form.
The three request headers #
| Header | Carries | Example content |
|---|---|---|
Signature-Agent | where the key directory lives | "https://crawler.example" |
Signature-Input | which components were signed and the parameters | ("@authority" "signature-agent");created=…;expires=…;keyid="…";alg="ed25519";tag="web-bot-auth" |
Signature | the signature bytes, base64 | sig1=:…: |
The parameters do the security work. created and expires bound the window in which the signature is valid, so a captured request cannot be replayed indefinitely. keyid names the key by thumbprint. tag set to web-bot-auth marks the purpose, so a signature made for something else cannot be reused here. A nonce can be added for replay detection within the window.
How a site verifies a request #
Parse Signature-Input and confirm the tag, algorithm and time window. Read Signature-Agent, fetch that origin's directory over HTTPS, and cache it. Find the key whose thumbprint matches keyid. Rebuild the signature base from the listed components exactly as RFC 9421 specifies, and verify the Ed25519 signature against it. If every step passes, the request was signed by whoever holds that private key within the stated window. Cloudflare runs this check for bots registered with it and treats a verified signature as a verified bot.
What a valid signature does and does not prove #
A valid signature proves possession of a key published at a particular origin. It does not prove the operator is who it says it is, what it will do with the content, or that it follows robots.txt. Those remain matters of reputation and policy. The directory domain is what you are trusting, so a signature from a domain you have never heard of verifies cleanly and tells you very little. Treat it like a TLS certificate: it authenticates an identity; it does not vouch for behaviour.
What we sign, and what we chose not to #
CrawlCheck publishes its own directory at /.well-known/http-message-signatures-directory, and requests the scanner makes under its own name are signed. The delivery comparison is deliberately not signed: it fetches pages under other crawlers’ user-agent strings to find out whether a site serves those crawlers something different. Signing those probes would let a site recognise them as ours and answer them as it answers us, which would defeat the measurement. That choice is disclosed on our crawler policy page.
Checking a signed request yourself #
Our free tools include a Web Bot Auth verifier: give it an agent origin and it fetches and checks the directory, or post a captured request and it verifies the signature, the window, the tag and the key. It flags reuse of a nonce it has seen before. For the network-level checks that still cover most crawlers today, see how to verify Googlebot, Bingbot and AI crawlers.
Key rotation and caching #
A directory can list several keys at once, which is how rotation works without breaking verification. The operator publishes the new key alongside the old one, starts signing with the new key, waits until verifiers' caches of the directory have expired, and then withdraws the old key. A verifier that caches the directory for too long will reject requests signed with a key published after its last fetch, so cache for minutes or hours, not days, and refetch once when a keyid is not found before rejecting.
Watch for one CDN behaviour in particular. A zone setting that overrides browser cache lifetimes can rewrite a directory's short max-age into a much longer one on cached responses, which defeats the point of a short-lived signed directory response. Serve the directory with explicit cache headers and check what clients actually receive, not what the origin sends.
Who signs today #
Adoption is early. Signing is concentrated among operators that run agents and fetchers on shared infrastructure, where IP verification is weakest, and among bots registering with CDNs that verify signatures. The large search crawlers continue to rely on DNS and published ranges. If you operate a crawler of your own, signing costs little: one key pair, one static directory, and a few lines of code around each request.
Should a site require signatures? #
Not yet. Most crawlers that matter for search and AI answers do not sign requests, so requiring a signature would block them. The practical stance today is to verify signatures where present, treat a verified signature as a stronger identity signal than an IP match, and keep DNS and IP verification for everything else. The drafts are still changing, so pin your implementation to a draft version and expect to update it.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What is Web Bot Auth?
An IETF draft that lets automated clients sign HTTP requests with HTTP Message Signatures (RFC 9421), so a site can verify which operator sent a request without relying on IP addresses.
Which headers does a Web Bot Auth request carry?
Signature, Signature-Input and Signature-Agent. Signature-Input lists the signed components and the created, expires, keyid, alg and tag parameters; Signature-Agent points to the key directory.
Where are the public keys published?
At /.well-known/http-message-signatures-directory on the operator's domain, as a JSON Web Key Set served with the media type application/http-message-signatures-directory+json.
What key algorithm does it use?
Ed25519 is the algorithm in common use; Cloudflare's verification supports Ed25519 keys identified by their JWK thumbprint.
Does a valid signature mean the crawler is trustworthy?
No. It proves the request was signed with a key published at a given origin. Whether that operator behaves well is a separate question.
Why does CrawlCheck leave some requests unsigned?
The delivery comparison fetches pages under other crawlers' user-agent strings to see if sites treat them differently. Signing them would let sites recognise the probes and defeat the measurement.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.