Findings · 2026-10-03 · By VSNARY | Emmanuel Orta · 0 views
Release notes 3 Oct 2026: Resolve, census, rulebook
What changed on crawlcheck.io between 29 September and 3 October 2026: a signed answer for AI agents, the first agentic-web census, a public rulebook, receipts, disputes and observers, and where each change can be checked.
Between 29 September and 3 October 2026, AI agents can now read one signed answer about a domain before acting (GET /api/v1/resolve, specified at /spec/resolve), with a change feed, outcome reports and a domain lifecycle record. The first agentic-web census measured 991 of the Tranco top 1,000 twice. Every rule is published at /rules, every fix ends in a signed receipt, disputes are resolved in public from more than one network, and a typed SDK, conformance fixtures and standards exports went live.
Second note in the Releases series, same rule as the first: every line names something that can be checked on the live site, with the path to check it. Nothing here is planned. This is what changed between the 29 September note and this one.
For agents: one signed answer before acting #
The largest change is a new reader. GET /api/v1/resolve?domain= (also /v1/resolve/{domain}, and POST for 1 to 100 names at once) returns one document an AI agent can read before it fetches, quotes or acts on a domain: the crawl policy as the site declares it, per crawler; what each identity was actually served; the machine files by status and hash; the capabilities the site declares agents may use, by safety class; who the site says it is and whether outside sources confirm it; open findings; freshness, observers, confidence and contradictions; and the report and manifest the answer was built from. There is deliberately no overall score. An agent decides on the section its action depends on.
Every answer is signed with a key published at /keys, and the same offline verifier that checks evidence bundles checks it with no call to this site. Each answer is also included in that day's sealed record; /api/seal/lookup says whether a given answer is in it. An unknown domain answers not_measured and is queued for measurement; a domain in the census frame answers from its census reading; an opted-out domain answers opted_out and nothing else.
Around it: GET /api/v1/changes?since= is a change feed, one event per section that moved when an answer was rebuilt, so a tool re-resolves only the domains that changed. POST /api/v1/outcome lets an agent report what happened when it acted (status, blocked, challenged, what the answer had said); outcomes are counted and published per domain and never used to change an answer. GET /api/v1/domain returns the domain's public registry record (RDAP: registrar, created, expires, nameservers) and its timeline. The whole protocol is an open, versioned specification at /spec/resolve, v1.1.0 as of today, with a JSON Schema and an MCP tool, resolve_domain, that returns the same document.
The census: 991 of the top 1,000, read twice #
The first agentic-web census is published on the dataset page. The frame is one Tranco top-1,000 list frozen on 1 October so every run measures the same sites: 991 members after nine excluded hosts, counted and never named. Each site is fetched as a browser and as the main AI crawlers, its robots.txt is read for 12 AI agents per RFC 9309, and its machine files and sitemap are read. Run 1 finished on 2 October, run 2 at 6:48 AM Mountain Time today, and every figure now carries its change between the two.
Correction, 3 October: this paragraph first said that the 644 homepages that did not answer with a 2xx "wall off datacenter traffic". The records say otherwise. 287 of them timed out, and the timeouts cluster by batch, 26 of 30 sites in one minute and none of 30 in another: that was a defect in our own census probe, not the sites. It was fixed the same morning, and run 3 is the first measured with the fix. About 244 are infrastructure names with no website (CDN, DNS and API hosts answering 530, 526, 404 and similar). About 100 refused: 403, 429, 401 or 402, or a 400. So 347 of 991 is a floor set partly by our own defect, and every delivery figure still reads refused from this vantage, never blocked. Of the 273 reachable sites that serve a robots.txt, 25.3% close the homepage to at least one of the 12 AI agents; GPTBot is the most-named at 17.6%, CCBot 19.8%, Bytespider 18.7%, ClaudeBot and Google-Extended 14.7% each. Eleven robots files carry a Content-Signal line. 23.1% of reachable sites serve an llms.txt; 2.3% serve an agents.md; 3.2% declare any agent capability at all (an MCP card, an API catalog, an agent-skills index or an A2A card). 16.7% refuse at least one AI identity where the browser got a page. Between the two runs the figures moved by tenths of a point: 4 sites became reachable and 7 stopped.
A second frame starts today: the top 10,000 of the same Tranco list, measured weekly, with its own runs, records and figures (/api/census?frame=top10k). One site's record in either frame is /api/census/records?run=N&domain=; the resolve answer now links there; it previously could point at the wrong record.
Every rule, every receipt, every dispute, in public #
/rules is the rulebook: every finding the scanner can raise, with what it means, how serious it is, how to fix it and its share of all scans. /api/rules serves the same table; a code that is not in it cannot appear in a report.
A fix now ends in a signed receipt. /api/receipt returns the before and after reports, the declared fix, the verdict and the files that changed, signed and verifiable offline. /api/outcomes lines a receipt up against the owner's connected data (Search Console, Business Profile, leads) and says plainly that a coincidence in time is not proof. /workflow shows the three most common findings fixed end to end on real sites, every step a link to its record, and names the one step those three chains cannot yet show because the fixes predate the evidence archive.
/disputes is the correction ledger. Anyone can dispute a finding; a site owner's dispute is accepted at once and anyone else's after a one-line proof at /.well-known/crawlcheck-dispute.txt. The finding is re-checked against the original evidence and a fresh audit, independent observers re-read the URL, and the verdict (upheld, corrected, fixed since, vantage-dependent, inconclusive) is signed and verifiable offline. The first resolution on the ledger is against one of our own sites and was upheld.
Behind these: /docs/lineage-coverage shows how much of the evidence path every finding family carries, and /trust shows the latest check that stored reports can be rebuilt from the archive alone.
Second and third networks #
The open observer protocol lets anyone run an independent observer: enrol a key, pull tasks, submit signed readings, and appear in /.well-known/crawlcheck-observers.json. Four observers are listed today, on separate networks, one of them the browser extension. A finding that two networks reproduce reads cross-vantage confirmed; one they disagree on reads observer disagreement and is shown as such, never merged. A reference observer and eight fixtures are published with the spec.
For developers #
The OpenAPI document now has 93 paths, every operation with a response schema and a worked example. @crawlcheck/sdk 1.0.0 is typed from it and bundles the offline verifier; it installs from this site. /conformance/ publishes 19 fixtures built only from this site's own records with the expected result of every check, so a third-party verifier can prove it agrees with ours. /api/v1/machine-record exports the same record as OTLP/JSON spans, OpenLineage run events and PROV-O JSON-LD, each validated with the standard's own tooling. /docs/api/compatibility tags every one of the 265 API fields stable (180-day notice), provisional (30-day) or internal, and a retired field announces itself with Deprecation and Sunset headers.
Two guards went in that readers will not see. Privacy floors: no median, percentile or count is shown for a cohort too small to stay anonymous, so a cohort of one cannot leak its peer. And a site owner can now export everything held about a domain as signed parts, or delete it, from /manage/data; the deletion itself is a signed record.
Security audit, round two #
An outside audit of the running service produced eight items, all closed, covering how API keys are accepted, request size limits, CSV export safety, input sanitising, report identifiers and what an evidence bundle may contain. A deploy no longer serves pages cached by the previous version.
Paid: two lines any key can carry #
Evidence ($49 a month or $490 a year, included in Agency) is the proof as one file: a signed export of every measurement in a period with its evidence bundle, every receipt, the certificate days and the seals, verified offline as a whole. Intelligence (the same price, included in Growth) is the corpus benchmark for any domain, the portfolio side by side and bulk verification. Prices are frozen until 18 October. Nothing that was free moved behind either line; every surface they gate is new.
The site #
Seventeen pages were rebuilt on the homepage's colour system, each with its own layout and a live figure drawn only from measured data. Findings and guides read as one library, most-read first, with a featured case. The glossary has 668 terms in 19 areas, each term on its own page. The navigation lost half its items; everything removed is still linked from the page it belongs to. A new mark and favicon shipped. Scoring did not change. Two rules were revised and say so on every card they touch: robots.txt matching now follows RFC 9309 wildcards and anchors, and a stale-cache finding is now confirmed before it is reported.
What changed, where to check it #
| Change | Check it at |
|---|---|
| Resolve protocol | /api/v1/resolve?domain= · /spec/resolve |
| Change feed, outcomes, lifecycle | /api/v1/changes · /api/v1/outcome · /api/v1/domain |
| Signing keys | /keys · /.well-known/http-message-signatures-directory |
| Census, two runs | /data#census · /api/census |
| Rulebook, 75 rules | /rules · /api/rules |
| Signed receipts and outcomes | /api/receipt · /api/outcomes · /workflow |
| Dispute ledger | /disputes |
| Lineage coverage, reconstruction, traces | /docs/lineage-coverage · /trust |
| Observer protocol, four observers | /docs/observer-protocol · /.well-known/crawlcheck-observers.json |
| Capability registry | /registry/capabilities |
| SDK, conformance, exports, compatibility | /sdk · /conformance/ · /api/v1/machine-record?format= · /docs/api/compatibility |
| Export or delete your data | /manage/data |
| Evidence and Intelligence | /pricing |
Live, as you read this: the corpus now holds 189,169 domains across 3,137 scans. The figures in this piece were measured on the date above; this line is not.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What does GET /api/v1/resolve return?
One signed document about a domain: crawl policy as declared per crawler, what each crawler identity was actually served, machine files by status and hash, declared agent capabilities by safety class, entity confirmation, open findings, freshness, observers, confidence, contradictions, and the evidence the answer was built from. There is no overall score.
How does an agent verify a resolve answer without calling crawlcheck.io?
The answer is signed with a key published at /keys and in the HTTP message signatures directory; the offline verifier checks the signature with no network call.
What is the agentic-web census?
A fixed frame of 991 sites from one Tranco top-1,000 list, each fetched as a browser and as the main AI crawlers, with robots.txt read for 12 AI agents and the machine files checked. It runs daily; every figure is published with its denominator and its change since the first run. No site is named.
Why did so few census sites answer in the first two runs?
In runs 1 and 2, 347 of 991 homepages answered with a page. Of the rest, 287 were timeouts caused by a defect in our own probe, fixed on 3 October; about 244 are infrastructure names with no website; about 100 refused. Delivery is reported as refused from this vantage, never as blocked.
Who can dispute a finding?
Anyone. A site owner's dispute is accepted at once; anyone else proves control with one line in a well-known file. The finding is re-checked against the original evidence, a fresh audit is taken and independent observers re-read the URL; the verdict is signed and can be verified offline.
Did the scoring change in this release?
No. Two rules were revised, robots.txt matching and the stale-cache check, and a finding that changed because a rule changed is marked rule-changed rather than counted as a change on the site.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.