CrawlCheck

Glossary · area 4 of 19

Entities and schema

Declaring who you are in a form a resolver can use without guessing.

39 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.

39terms in this area
5with a live finding rate

structured data

Machine-readable statements about a page, usually JSON-LD, describing what the page is and what it is about. Validity is cheap to check; agreement with the visible page is the part that is usually wrong.

JSON-LD

The JSON serialisation of linked data, embedded in a script tag. It is the format search and answer engines read most reliably, which also means an error in it is an error stated confidently to machines.

stable @id

A durable identifier that lets separate statements be understood as being about the same thing. Two records sharing an @id are the same record twice; without one, every page's claims start again from nothing.

entity collision

Two distinct entities in a site's structured data sharing an identifier, or one entity split across identifiers. Either way a machine reading the graph cannot tell how many things are being described.

Measured: ENTITY_COLLISION 3.8%

reciprocated sameAs

A sameAs link that the destination profile returns. Counting sameAs links measures how many claims a site makes about its own identity; fetching each destination and looking for the domain measures how many of those claims anyone else confirms.

sameAs

A schema property linking an entity to its profiles elsewhere. Its value is entirely in whether the target links back; without that it is a claim about yourself.

knowledge graph

A store of entities and relationships an engine keeps independently of any page. Publishing correct markup makes a claim available for reconciliation; it does not put anything into a graph by itself.

entity resolution

Deciding whether two records describe the same thing. Everything a site can do to help is about giving one identity one stable identifier and never contradicting itself.

entity disambiguation

Distinguishing your entity from others sharing a name. A shared name with no distinguishing identifier is the ordinary case for local businesses, not an edge case.

Wikidata

An open structured database used as an authority link target. A sameAs pointing at Wikidata is only corroboration if the Wikidata item points back; an unreciprocated link is a claim, not a confirmation.

NAP

Name, address and phone number. Its value is consistency across every place the business appears, so the failure mode worth measuring is disagreement between sources rather than absence from any one.

Measured: NAP_NAME_DRIFT 0.5% · NAP_UNDECLARED_LISTING 5.4%

local citation

A directory or profile listing a business. Note the collision: in the answer-engine sense a citation is a source credited in an answer. This glossary keeps them separate because conflating them makes both numbers meaningless.

Organization schema

Markup declaring the publishing entity. The single most reusable node on a site, because everything else can reference it by identifier rather than repeating the same facts.

LocalBusiness

Markup for a business with a service area or premises. Coordinates and a street address that disagree are a defect a resolver cannot arbitrate.

FAQPage

Schema marking question-and-answer content. The answers must match what a visitor sees; declaring answers not present on the page is the ordinary way this markup becomes a liability.

structured data validation

Checking markup parses and matches the visible page. Valid markup that contradicts the visible text is a worse position than no markup, because it makes a confident wrong claim available for reuse.

@context

The JSON-LD declaration binding your keys to an external vocabulary. Without a valid context, name and description are just strings — the linked-data meaning is carried by the context, not by the key.

agentic knowledge graph

A knowledge graph that agents read from and write to, rather than one a person curates and queries. Two senses are in circulation and they are not the same claim: the practical one, where a graph stores entities and relations that a retrieval agent traverses instead of, or alongside, embedding search; and a speculative one, where the graph is generated by a model as it reasons and passed between agents as a reasoning artifact. The first is a database design and is in production; the second is a research and blog-post idea with no standard behind it. Neither is something a website publishes. Your site can contribute nodes to somebody else graph by declaring stable identifiers and reciprocated sameAs edges, but you do not control what the graph makes of them.

microformats (h-card)

Contact and identity markup carried in HTML class attributes rather than in JSON-LD: a name, organisation, email, URL and locality that any parser can read from the page itself. Measured and shown in the schema section, and compared against the NAP where both exist. Not scored: adoption outside a few communities is a rounding error, and absence is not a defect.

RDF

The Resource Description Framework: facts expressed as subject, predicate, object statements (triples) with resolvable identifiers. JSON-LD is one serialisation of it, so every schema.org block on a page is already a small RDF graph, and it can be checked as one - whether its identifiers resolve, whether references point at nodes that exist, whether it parses at all.

triple

One statement in an RDF graph: this subject has this property with this value. A JSON-LD node with a name, a URL and a type is three triples. Counting them is the honest size of what a linked-data reader receives; a page can carry a large script block and declare almost nothing.

dangling @id

A reference to an @id that no node on the page declares. The reference is syntactically valid, so validators pass it, and it is a statement about nothing: the resolver drops the edge and the entity loses the relation it was meant to carry.

RDFa and Microdata

Two ways of embedding structured data in HTML attributes rather than a script block. Read by some parsers, ignored by others, and a second place for the facts to disagree with the JSON-LD. Counted alongside JSON-LD and not scored there, because one serialisation kept accurate is worth more than three that drift. When a page declares no JSON-LD node at all, its schema.org Microdata and RDFa are read as the page's graph and every entity check runs on them (from score version 20).

founder edge

The founder or employee property on an Organization node, pointing at a Person by @id. It is the statement that connects the business graph to the people graph. The target must be declared on the same page or the edge dangles; a reference is not a declaration.

foundingDate

The date property on an Organization node. It is the one claim behind every years-of-experience counter that a reader can check against a registry or a domain record, so it should be a date the operator can defend: incorporation, or the day the site first went live.

knowsAbout

A property on a Person or Organization listing the subjects the entity claims expertise in, ideally as DefinedTerm nodes with a sameAs to a public identifier. Limited to what the site's own content evidences, it is one of the few machine-readable expertise statements; padded, it is a list of keywords nobody will believe.

unreciprocated profile

A sameAs target that is reachable and does not link back. It says nothing about whether the profile is real; some platforms never link a website out. It is shown, not failed. Only a dead target, one that does not answer from a client that can read it, fails a corroboration row.

entity graph

The web of nodes and typed edges a resolver builds from your declarations plus everyone else's: your Organization node, the Person it names as founder, the Wikidata item that names you back. It is distinct from the entity map, which is the file pair a site publishes, and from a knowledge graph, which is the store an engine keeps. A site contributes nodes and edges to a graph; it does not own the graph, and a claim nobody else corroborates is an edge with one end.

DefinedTerm

The schema.org type for a term that has its own identity, used inside knowsAbout so that a claimed subject carries a sameAs to a public identifier instead of sitting there as a bare string. It gives a resolver something to match against Wikidata or a vocabulary. It is a claim about a subject, not about the organisation: a DefinedTerm pointing at a real Wikidata item does not corroborate the organisation's own identity, and a scanner that counts it as an unreciprocated sameAs on the business is reading the wrong node.

Person schema

The node that Organization, LocalBusiness, founder edge and author references all point at, and the authorship half of an entity graph. One Person, one stable @id declared on every page that references it, a name, a jobTitle, a url, and a sameAs to at least one profile that links back. It can support a claim that a named human stands behind the site. It cannot support expertise: a Person node with a credentials list is a declaration, and the scanner's E-E-A-T proxy rows read only whether it is declared and corroborated, not whether it is true.

named entity recognition

How a model extracts people, organisations, places and products from unmarked prose when no structured data exists. It is why a page with no schema is still partly understood, and why structured data is an advantage rather than a requirement. What it cannot give is a stable identity: a name found in text is a string, matched by guess to whichever entity seems likeliest, and two organisations with similar names are confused exactly where a stable @id would have separated them.

NAP consistency

Agreement of a business's name, address and phone number across its own site and every directory that lists it. The phone number does most of the work, because it is written the same way everywhere and rarely shared; a listing that prints a different name on the same number is a second entity to a retrieval system.

Measured: NAP_NAME_DRIFT 0.5% · NAP_UNDECLARED_LISTING 5.4%

name drift

A directory listing reached by a business's phone number that prints a name different from the one the site declares. The older name usually carries more corroboration, so an engine resolving the entity may prefer it. It is found by searching directories by number, not by name, which would only find namesakes.

Measured: NAP_NAME_DRIFT 0.5%

geocoding

Converting a text address into a latitude and longitude. A business record that gives an address and no coordinates forces every resolver to geocode, which is where confusion between nearby or similarly named businesses enters. Declaring the point removes the guess, and a declared point can still be wrong by a sign.

Measured: ENTITY_NO_COORDINATES 4.1% · GEO_COUNTRY_MISMATCH 0.1%

rival roots

Two top-level organisation-like nodes on one page with the same name and different @ids: the same business declared as two entities. A consumer cannot merge them, and corroboration earned by one does not accrue to the other. The usual cause is a theme emitting its own business node beside a plugin's.

speakable

A schema.org property naming which parts of a page are suitable to be read aloud or quoted as an answer, by CSS selector or XPath. It points at existing text; it does not make the text quotable. A page that marks a whole article speakable has told an assistant nothing it did not know.

FAQPage schema

Structured data listing question-and-answer pairs that appear on the page. Google restricts rich results for it to government and health sites since 2023, but the pairs remain readable to answer engines, which is the audience that still uses them. Answers must match the visible text; a pair that exists only in the markup is a claim without a page.

OpeningHoursSpecification

The schema.org shape for declaring business hours: day of week, opens and closes, optionally a validity window. It is one of the least-adopted properties on local business pages and one of the most often read by an assistant answering a when question; hours that disagree with a directory listing are resolved by the directory, not the site.

service area

The places a business serves rather than the place it sits, declared in schema as areaServed and in a Google Business Profile as a list of localities. A site whose schema lists forty areas and whose sitemap has pages for four has declared reach it cannot show; the declared list is a claim, the pages are the evidence.

← Machine files  ·  Indexing and discovery →

All 668 terms across 19 areas.