Findings · 2026-09-29 · By VSNARY | Emmanuel Orta · 0 views
Release notes 29 Sep 2026: finding states, sitemap, trust
What changed on crawlcheck.io on 29 September 2026, from the scanner to the legal pages, and where each change can be checked.
On 29 September 2026 CrawlCheck shipped 24 builds. Every finding now carries a stable fingerprint (code plus path) and a lifecycle state, and the public API returns both. The sitemap became an index of five files. Every finding card shows the rule that fired and how to verify the fix. Reports in German, French, Spanish and Portuguese have no English left in their tooltips. A trust center, terms of service and a refund policy went up, and paying customers manage billing from their own dashboard.
This is the first release note in the Releases series. The rule for the series: every line names something that can be checked on the live site, with the path to check it. Nothing here is planned; everything here is deployed. Twenty-four builds went out on 29 September 2026, from the first at 1:22 PM to the last at 5:30 PM Mountain Time, each one a hashed patch applied to the running Worker only when the live code matched the hash it was built against.
Every finding has a state and a fingerprint #
A finding is identified by its code and the path it was found on. That pair is now a fingerprint that stays stable from scan to scan, and the scanner keeps a small history per domain keyed on it. Each scan assigns every finding one of seven states: new, persisted, worsened, improved, regressed, resolved, or rule-changed. The last one matters more than it looks: when a scoring rule changes and a finding stops being reported, the record says the rule changed. It does not say the site was fixed.
Before this build a severity that fell was recorded silently. Now it reads improved, and a finding that was fixed and came back reads regressed rather than new. Both were tested across four consecutive scans of one site before shipping.
The public API carries this too. PublicReportV1 findings now include fingerprint, state, first_seen, scans_seen and recurrences. They are additions, so nothing built against v1 breaks; the schema at /api/v1/schema lists them.
Every finding card shows its rule and how to verify the fix #
Each finding in a report now ends with two lines. The first names the rule: the finding code, the scoring version it was evaluated under, the fingerprint and the state above. The second, Verify the fix, says what a re-scan should show once the finding is addressed. The point is that a reader can tell a finding that disappeared because the site changed from one that disappeared because our rules did, without asking us.
The sitemap is an index of five #
/sitemap.xml used to be one file listing every URL on the site. It now points at five: pages, glossary, directory, registry and certificates. The reason is our own finding. A site whose machine files weigh more than its homepage gets MACHINE_CHAIN_HEAVY, and after the sitemap passed 1,700 URLs, crawlcheck.io earned it. The split brought the chain a crawler reads before the page from above the page's weight to below it, and the finding cleared on the next self-scan.
The certificates file is new in kind: it lists a certificate page for every host whose most recent certificate reading passed within the validity window, including the portfolio sites the scanner is built on. Those pages were indexable before and listed nowhere.
Translated reports with no English left in them #
Reports in German, French, Spanish and Portuguese have been live since 26 September. What they still carried in English was the text inside the SVG figures: the tooltips on the score spark, the section bars and the manifest line, because those strings are built from numbers at render time and no dictionary entry could key them by shape. A shape-aware pass now translates them, along with the new rule and verify lines. The measured gap is zero strings on all four languages.
The blog has four kinds #
Sixty-five posts were filed into four kinds, each with its own filter: research (measured across the corpus or the telemetry), scanner corrections (where our own scan was wrong, and the fix), case studies (one site, one defect, end to end) and releases, which this post starts. The kind is also written into each post's articleSection, so a reader of the structured data sees the same filing a reader of the page does.
Trust center, terms, refunds #
/trust is built from the running configuration rather than written about it: the retention table reads the real TTLs (reports 90 days, leads 180, the comment queue 30, Search Console and Business Profile data 400, sessions 12 hours), the security-header list is the one the Worker sends, and the mail section states what the domain actually publishes for SPF, DKIM, DMARC and MTA-STS, including that MTA-STS is still in testing mode. The daily OpenTimestamps seal is described there too.
/terms and /refunds went up the same afternoon. The refund window is seven days, the governing law is Colorado, and both pages link from the footer of every page in every language. Paying customers on a monthly or annual plan now see Manage billing on their dashboard, which opens Stripe's customer portal: update a card, download an invoice, or cancel at the end of the period without emailing anyone.
Weight and links #
The homepage went from 61.7 KB to 38.7 KB of HTML by moving every inline style block into the one fingerprinted stylesheet; the same pass then caught fourteen more inline blocks across the language switcher, the pricing grid, ten page figures and the forms page. Only the globe keeps its own styles, because it is its own document. The links section of the scanner stopped reading anchors out of script, style, noscript and template content, stopped counting anchors marked rel="nofollow", and no longer counts login, cart, checkout, account and admin paths against a site. Each of those was a false finding on our own report first.
What changed, where to check it #
| Change | Check it at |
|---|---|
| Finding state and fingerprint | any report card; /api/v1/schema |
| Rule and Verify lines | any finding card, in five languages |
| Sitemap index of five | /sitemap.xml |
| Certificate sitemap | /sitemap-certificates.xml |
| Translated SVG tooltips | any /de/r/, /fr/r/, /es/r/, /pt/r/ report |
| Blog kinds | /blog?kind=releases and the other three |
| Trust center, terms, refunds | /trust, /terms, /refunds |
| Manage billing | /manage, paid plans |
| Homepage weight, inline styles | view source on /; one <style> only on /globe |
How these builds ship #
Since this morning, changes to the Worker go through a reviewed patch file: a list of exact old-text, new-text pairs, the hash of the code it was built against and the hash it must produce. The deploy job refuses to run unless the live code matches the first hash, refuses to upload unless the result matches the second, and re-reads the live code afterwards to confirm. Every one of the 24 patches is committed beside the mirror of the running code, so the diff for any build in this note is a file, not a memory. The scanner's own report of crawlcheck.io at the end of the day: grade A, 99, zero findings, the only rows below 100 being field speed data that varies by visitor.
Live, as you read this: the corpus now holds 7,808 domains across 2,468 scans. The figures in this piece were measured on the date above; this line is not.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What is a finding fingerprint in CrawlCheck?
The finding code joined to the path it was found on. It stays the same across scans of the same domain, so the scanner can say whether a finding is new, persisted, improved, worsened, regressed, resolved or removed by a rule change.
Does a resolved finding mean the site was fixed?
Resolved means the finding was present on the previous scan and absent on this one, on a scan that could see it. A finding that disappeared because the scoring rules changed is filed as rule-changed, never as resolved.
Will the new API fields break existing integrations?
No. Within v1, fields are added and never renamed or removed. The five new finding fields are additions; a client that ignores unknown properties sees no change.
Why is the sitemap split into five files?
The single file had passed 1,700 URLs and weighed more than the homepage, which is a finding the scanner reports on other sites. Splitting it put the crawler's machine-file chain back under the page's own weight.
What does the trust center show?
Retention periods read from the live configuration, the security headers the site sends, the domain's mail authentication records, the outside services used, and how the daily OpenTimestamps seal works.
How long is the refund window?
Seven days from purchase, under Colorado law, as stated on the refund policy page linked from every page footer.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.