CrawlCheck

Guides · 2026-10-03 · By · 0 views

Rich Results Test vs Schema Markup Validator vs CrawlCheck

Three free structured-data tools, three questions: Google rich-result eligibility, schema.org validity, and whether the markup describes one coherent business machines can resolve.

Google's Rich Results Test checks which Google rich results your structured data can produce. The Schema Markup Validator checks that any schema.org markup is valid. CrawlCheck checks whether the live markup describes one coherent business: one entity per @id, name, address and phone matching the page, coordinates, and sameAs profiles that link back. All three are free; use the first two for syntax, CrawlCheck for entity coherence.

Share of all scans carrying each finding named aboveENTITY_NO_COORDINATES5.7%NAP_UNDECLARED_LISTING4.9%ENTITY_COLLISION3.5%NAP_NAME_DRIFT0.7%Share of all scans carrying eachfinding named aboveENTITY_NO_COORDINATES5.7%NAP_UNDECLARED_LISTING4.9%ENTITY_COLLISION3.5%NAP_NAME_DRIFT0.7%
Read live from the same counters the dataset page uses, at the moment this page was served. Bars are scaled to the largest value shown, not to 100%.

Three free tools look at a page’s structured data, and each answers a different question. Google’s Rich Results Test asks whether Google can generate rich results from it. The Schema Markup Validator asks whether it is valid schema.org. CrawlCheck asks whether it describes one coherent business that machines, including AI systems, can resolve: one entity per @id, consistent name, address and phone, coordinates, and outside profiles that link back.

Disclosure: CrawlCheck publishes this comparison and is one of the products in it. Every figure about another company comes from that company’s own pricing or documentation page, linked where it appears, read on 3 October 2026. Where CrawlCheck does less than a competitor, the tables say so.

What each validator checks #

Google Rich Results TestSchema Markup Validator (validator.schema.org)CrawlCheck
Question it answersWhich Google rich results this markup can produceIs this valid schema.org markupDoes this markup describe one resolvable entity
ScopeOnly types Google supports for rich resultsAll schema.org types, “without Google feature specific warnings”LocalBusiness/Organization identity, @id graph, NAP, geo, sameAs reciprocity
Cross-checks with the page and the webNoNoYes: compares schema NAP with what the page shows, tests sameAs links for a link back
Finds duplicate or colliding entitiesNoNoYes: two unlinked organisation nodes (the same @id parsed twice is correctly treated as one entity)
PriceFreeFreeFree scans

Sources: Google Search Central on structured data testing tools, validator.schema.org.

Valid markup that still confuses machines #

All of the following pass both Google tools, because each node is individually valid. Shares are across every CrawlCheck scan, read live:

ProblemShare of scanned sites
Two organisation nodes that do not link to each other (an entity collision)3.5%
A business entity with no coordinates5.7%
Directory listings the site does not declare4.9%
Business name drifting between schema and page0.7%

The checks are explained in JSON-LD @id and IRI errors and the NAP consistency audit. One trap we fell into ourselves: two records with the same @id are one record parsed twice, not two businesses, written up in two nodes with the same @id.

Which to use #

A worked example: valid markup that describes two businesses #

The most common failure we see is not a syntax error. It is two separate nodes that each describe the business, added by two different plugins, with nothing tying them together. Here is a shortened version of the pattern:

{"@context": "https://schema.org", "@type": "LocalBusiness",
 "name": "Example Fence Co", "telephone": "+1-303-555-0100"}

{"@context": "https://schema.org", "@type": "Organization",
 "name": "Example Fence Company LLC", "url": "https://example.com/"}

Both blocks are valid schema.org, so the Schema Markup Validator reports no errors. The Rich Results Test judges each block against the rich-result features it supports and has no reason to object either. A machine reading the page, though, now holds two entities with different names, one with a phone and no URL, the other with a URL and no phone, and no statement that they are the same thing. It has to guess.

The fix is to give both nodes the same @id (for example https://example.com/#business) so that they merge into one record, or to delete the weaker one and move its properties into the other. Two nodes that share an @id are one entity stated twice, which is valid and correct. CrawlCheck reports the unlinked version as an entity collision and the shared-@id version as a single entity.

A four-step check, in order #

  1. Rich Results Test on the live URL. Confirms the markup Google features parses and shows which rich results it can produce.
  2. Schema Markup Validator for everything else. Types Google does not feature (a Service, a Person, a DefinedTerm) are still read by other systems; this catches errors in them.
  3. Compare the served HTML with the rendered page. If a script or tag manager adds the JSON-LD after the page loads, a tool that renders the page will see it and a crawler that does not run JavaScript will not. View the page source (not the inspector) and search for application/ld+json.
  4. A CrawlCheck scan for coherence. One entity per business, the same name, address and phone as the visible page, coordinates on the business node, and sameAs profiles that link back to the site.

sameAs is where a page lists the other places the same business appears: a Google Business Profile, a Facebook page, a Wikidata item, a directory listing. Anyone can list any URL there, so a list on its own proves nothing. A profile that also links back to the site is a two-way statement, and it is the version machines can trust. CrawlCheck opens each declared profile and records whether it points back; neither Google tool does, because it is outside the question each answers.

Where CrawlCheck is weaker #

CrawlCheck is not a general schema validator. It does not test rich-result eligibility, it does not validate every property of every schema.org type, and it does not let you paste a code snippet to test before publishing; it reads what is live. For syntax and eligibility, Google’s two tools are the reference. Background on why entity coherence matters: what is entity SEO.

Every figure above came out of this scanner.

Point it at your own domain and see the same measurements, free.

Scan a domain — free

The main product

Found this on your own site? We fix it for $749.

Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.

Questions this post answers

What is the difference between the Rich Results Test and the Schema Markup Validator?

Google's Rich Results Test shows which Google rich results your structured data can generate and checks only the types Google supports for them. The Schema Markup Validator at validator.schema.org validates all schema.org markup without Google feature-specific warnings.

Can valid schema still be wrong?

Yes. Markup can be valid node by node and still describe two unlinked organisations, omit coordinates, or state a name and phone that differ from the page. Neither Google tool checks for that; CrawlCheck does.

Does CrawlCheck validate schema.org markup?

Not as a general validator. CrawlCheck checks whether live structured data describes one coherent, resolvable business: one entity per @id, consistent name, address and phone, coordinates, and sameAs profiles that link back. Use Google's tools for syntax and rich-result eligibility.

Which schema tool should I use for AI search?

Run the Rich Results Test and Schema Markup Validator for syntax first, then a CrawlCheck scan to confirm the entity is coherent and matches what the page shows. AI systems resolve businesses from consistent identity signals, not from rich-result eligibility.

Are these schema tools free?

Yes. Google's Rich Results Test, the Schema Markup Validator and CrawlCheck scans are all free.

Related findings

How anything measured in this article was measured15client identitiesone second, one address5machine filesapex and www114named agentsresolved from robots.txt24sections scoredreach, read, quoteHow anything measured here was measured15 client identities5 machine files114 named agents24 sections scoredone second, one addressapex and wwwresolved from robots.txtreach, read, quote
No account, nothing installed, and the same sequence on every domain — which is what makes one scan comparable to another. Run it on your own site.

Comments

Comments are read before they appear. Nothing is published automatically, and no account is needed.

Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.

All guides · The dataset · How the dataset works