# Visitor Resolve

A user agent is a claim anyone can type. Visitor Resolve checks the claim: send the address and user agent of a request
and CrawlCheck answers whether the agent it names really sent it.

    GET https://crawlcheck.io/api/v1/visitor?ip=<address>&ua=<user agent>
    POST https://crawlcheck.io/api/v1/visitor  {"visitors": [{"ip": "...", "ua": "..."}, ...]}   (1 to 100)

Second proof (POST): add the request's Web Bot Auth headers and CrawlCheck also checks the signature against the key
directory the Signature-Agent names (up to 10 signatures per call):

    {"ip": "...", "ua": "...", "method": "GET", "path": "/page",
     "headers": {"host": "example.com", "signature-agent": ""https://agent.example"",
                 "signature-input": "sig1=(...);created=...;keyid="...";tag="web-bot-auth"", "signature": "sig1=:...:"}}

Every answer is signed (kind crawlcheck-visitor): save it and verify it later with
https://crawlcheck.io/crawlcheck-verify.mjs.

Report arrivals (Cloudflare Workers only). After your Worker has a verdict for an agent request, send one
fire-and-forget subrequest and the hit counts toward your site in the Agent Traffic Index
(https://crawlcheck.io/api/v1/traffic-index) and the network figures on /data. Your site is proven by the CF-Worker
header Cloudflare adds to the subrequest; no key, nothing to configure, no site is ever named in the index:

    ctx.waitUntil(fetch("https://crawlcheck.io/api/v1/visitor/arrival?ua=" + encodeURIComponent(ua)
      + "&path=" + encodeURIComponent(url.pathname) + "&verdict=" + verdict));

Verdicts:

- verified: the address is inside the ranges the operator publishes for that agent
- spoofed: the operator publishes ranges for that agent and the address is outside all of them
- unverifiable: the operator publishes no range file for that agent, or it could not be read
- not_a_known_agent: the user agent names none of the agents CrawlCheck tracks

Range files are read from the operators themselves (OpenAI, Anthropic, Perplexity, Google, Microsoft, Apple) and refreshed daily.

## At the edge

Cloudflare Worker example: check only requests that claim to be an agent, remember each answer for an hour, and mark
spoofed ones (block them instead if you prefer).

    const AGENT = /GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User|Googlebot|Google-Extended|Bingbot|Applebot/i;
    export default { async fetch(req, env, ctx) {
      const ua = req.headers.get("user-agent") || "", ip = req.headers.get("cf-connecting-ip") || "";
      if (!AGENT.test(ua)) return fetch(req);
      const key = new Request("https://visitor.cache/" + encodeURIComponent(ip + "|" + ua));
      let hit = await caches.default.match(key);
      if (!hit) {
        const r = await fetch("https://crawlcheck.io/api/v1/visitor?ip=" + encodeURIComponent(ip) + "&ua=" + encodeURIComponent(ua));
        hit = new Response(await r.text(), { headers: { "cache-control": "max-age=3600" } });
        ctx.waitUntil(caches.default.put(key, hit.clone()));
      }
      const v = await hit.json().catch(() => ({}));
      const res = await fetch(req), out = new Response(res.body, res);
      out.headers.set("x-agent-verdict", v.verdict || "unknown");
      return out;
    } };

## Privacy

CrawlCheck keeps no address, user agent or site name from these checks; only daily counts by verdict and claimed agent:
https://crawlcheck.io/api/v1/visitor/stats
