Guides · 2026-09-30 · By VSNARY | Emmanuel Orta · 0 views
JSON-LD @id and IRI errors: what breaks and how to fix it
Six identifier defects in JSON-LD that no vocabulary validator reports: blank nodes, path @ids that 404, Unicode normalisation, raw IDN hosts, encoding drift and rival roots.
A JSON-LD @id is the IRI that identifies a node in the RDF graph a consumer builds; two nodes with the same @id are one node. CrawlCheck reads every @id, url and sameAs on the homepage and reports six defects vocabulary validators miss: blank-node identifiers (_:), path-form @ids that can 404, values not in Unicode NFC, raw internationalised hostnames, percent-encoding drift on one path, and rival roots (one business under two @ids). Fix: one fragment-form @id per entity on the canonical host, NFC-normalised, referenced rather than restated everywhere.
Every JSON-LD node on a page can carry an @id, and most pages that carry one got it from a plugin default rather than a decision. That matters because @id is not a label. It is an IRI, an internationalised URI, and the RDF graph a consumer builds from the page treats two nodes with the same @id as one node and two nodes with different @ids as two things. Every mistake in this guide comes from that rule. The scanner reads the @id, url and sameAs values of every JSON-LD block on the homepage and reports six specific defects, none of which any validator that checks schema.org vocabulary will mention, because the vocabulary is fine; it is the identifiers that are wrong.
What an @id is for #
A consumer that reads JSON-LD, whether a search engine, an answer engine or a knowledge-graph pipeline, converts it to triples: subject, predicate, object. The subject of every triple is the node's @id. A node without one gets a blank node, an identifier that exists only inside that one document and can never be referred to from anywhere else. A node with @id https://example.com/#organization can be referenced from every other page on the site, from an EntityMap, from a Wikidata entry and from another site's sameAs, and all of those references resolve to the same subject. That is the entire value of the field: it lets the world agree on which thing you are.
Six things the scanner checks #
Blank-node identifiers. An @id beginning _: is a blank node written out by hand. It is legal JSON-LD and it is useless as an identity, because _:b0 in your document and _:b0 in anyone else's are unrelated. The report lists the blank identifiers it found. The fix is to give the node a real IRI on your own domain.
Fragment versus path. The scanner counts top-level nodes whose @id is on the site's own host and splits them into fragment form, https://example.com/#org, and path form, https://example.com/org. Both are valid IRIs. The difference is what happens when someone fetches them. A path IRI is a URL that should return something, and if /org returns a 404, the identifier points at nothing. A fragment IRI resolves to the page that carries the node, which always exists. Fragment form is the safer default for entities that do not have a page of their own.
Unicode normalisation. An IRI containing accented characters can be written in composed form (one code point per accented letter) or decomposed form (base letter plus combining mark). They render identically and compare unequal. The scanner tests every identifier against NFC and counts the ones that are not normalised, with examples. The fix is to normalise to NFC before writing the value, which every language runtime can do in one call.
Raw internationalised hostnames. A hostname with non-ASCII characters in an @id or url is counted separately. Consumers differ on whether they convert it to punycode before comparing, so the same business can appear under two hosts. Write the punycode form, xn--, in identifiers.
Percent-encoding drift. The same path written with two different encodings, /caf%C3%A9 and /café, decodes to one path and compares as two identifiers. The scanner decodes every path, groups the raw spellings, and reports any path with more than one spelling on the same page, showing the pair.
Rival roots. Two top-level organisation-like nodes with the same normalised name but different @ids are the same business declared as two entities. This is the one that does the most damage, because a consumer has no way to merge them, and the corroboration that one node earns from a directory or a review site does not accrue to the other. The report names the business and lists the identifiers. The opposite case, one node parsed twice under the same @id, is not a defect and is covered in a separate post.
| Defect | Looks like | Consequence | Fix |
|---|---|---|---|
| blank node | "@id": "_:b0" | cannot be referenced from anywhere | own-domain IRI |
| path @id that 404s | https://x.com/org | identifier points at nothing | fragment form or a real page |
| not NFC | é as e + combining acute | same string compares unequal | normalise to NFC |
| raw IDN host | https://müller.de/#org | two hosts to some consumers | punycode host |
| encoding drift | /caf%C3%A9 and /café | one path, two identifiers | one spelling site-wide |
| rival roots | two Organization nodes, one name, two @ids | corroboration split across two entities | one @id, referenced everywhere |
The sameAs side #
The same reading covers sameAs values, with one extra check for Wikidata. A Wikidata link can be written as https://www.wikidata.org/wiki/Q42, the human page, or http://www.wikidata.org/entity/Q42, the entity IRI. Both are commonly written; only the entity form is the identifier of the thing in Wikidata's own graph. The scanner counts each form separately, so a site can see which one it is using. The broader question of whether a sameAs target links back is measured elsewhere in the report, under entity corroboration, and explained in the entity SEO guide.
How to fix it in one pass #
Choose one @id per entity, in fragment form on your canonical host: https://example.com/#organization for the business, https://example.com/#website for the site, https://example.com/page/#webpage for pages. Write it identically on every page, NFC-normalised, punycode host, one encoding. Reference it, rather than restating the node, wherever another node needs to point at the business: a publisher field that contains {"@id": "https://example.com/#organization"} and nothing else is a reference, and the graph merges it with the full node. The free generator at /api/fix/jsonld (?domain= your host) reads a page's existing JSON-LD and emits a corrected block with these identifiers applied, leaving the vocabulary untouched. What it cannot decide for you is which of two rival roots is the real one; that is a business fact, and it is the one you fix by hand.
Checking a page by hand #
View the page source, not the rendered DOM, and copy every <script type="application/ld+json"> block into a JSON parser. List every @id with its host, its form (fragment or path) and its @type. Any identifier beginning _: goes on the fix list at once. For each path-form identifier on your own host, fetch it and note the status; a 404 is a dangling identifier. Sort the remaining list by normalised name and look for the same name under two identifiers, which is the rival-roots case. Then run one more pass over the raw bytes for % sequences and for any non-ASCII character in a hostname. The whole pass takes a few minutes on a typical page and finds the defects the vocabulary validators skip, because those tools check that telephone is a string and never ask what @id points at.
What this is not #
None of this is a schema.org vocabulary check. A page can pass every vocabulary validator and still have two rival roots and a blank node, and a page can have a warning about a recommended property and be perfectly identified. The scanner reports the identifier findings as measured rows without moving the grade, because the rule needs corpus calibration before it can score; the share of scans carrying each row is on the dataset page. The finding that does score, ENTITY_COLLISION, is the case where two business nodes share a phone number, an @id or a coordinate pair, and it is described in the report guide.
Every figure above came out of this scanner.
Point it at your own domain and see the same measurements, free.
The main product
Found this on your own site? We fix it for $749.
Scan free to see where you stand. The fix is one site, every finding implemented and re-measured, with a sealed before and after.
Questions this post answers
What is an @id in JSON-LD?
The IRI that identifies the node as a subject in the RDF graph a consumer builds from the page. Two nodes with the same @id are one node; two nodes with different @ids are two things, whatever their names say.
Is a blank node identifier like _:b0 an error?
It is valid JSON-LD but useless as an identity: it exists only inside that document and cannot be referenced from any other page, EntityMap or sameAs. Replace it with an IRI on your own domain.
Should @id use a fragment or a path?
Fragment form, such as https://example.com/#organization, resolves to the page carrying the node and always exists. A path form is a URL that should return something, and one that returns 404 points at nothing.
Why does Unicode normalisation matter in an identifier?
Composed and decomposed forms of an accented character render identically and compare as different strings, so the same identifier written both ways is two identifiers to a consumer. Normalise to NFC before writing the value.
What are rival roots?
Two top-level organisation-like nodes on one page with the same name and different @ids: the same business declared as two entities. Corroboration earned by one does not accrue to the other, and a consumer cannot merge them.
Which Wikidata URL form belongs in sameAs?
The entity form, http://www.wikidata.org/entity/Q…, is the identifier of the thing in Wikidata's graph. The /wiki/ form is the human page. The scanner counts both so a site can see which it uses.
Related findings
Comments
Comments are read before they appear. Nothing is published automatically, and no account is needed.
Writing about this? Facts, live figures and marks — every number on that page is dated and traceable to a scan.