index coverage
How many of a site pages an engine has actually indexed. Distinct from sitemap coverage, which is only what the site declares. The two are compared, never substituted.
Glossary · area 5 of 19
Whether a page can be found and kept, separately from whether it can be fetched.
31 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
How many of a site pages an engine has actually indexed. Distinct from sitemap coverage, which is only what the site declares. The two are compared, never substituted.
A declaration of which URL is the preferred version. A canonical pointing away from a page is a request not to index that page, and a sitemap declaring a URL whose canonical points elsewhere is a contradiction the site is publishing about itself.
Measured: HOST_DUPLICATE_200 3.3% · DOMAIN_IS_ALIAS 0.9%
An instruction to keep a page out of search results, sent as a header or a meta tag. A crawler obeys either, so setting it in one place and not the other still removes the page.
An HTTP header carrying indexing directives. Invisible in the page source, which is why a page can look perfectly indexable in a browser while being excluded.
Measured: INSTRUCTION_CONFLICT 0.3%
A page no internal link points to. Findable by a crawler only if it is declared in a sitemap, and findable by nobody if it is in neither.
Experience, expertise, authoritativeness, trustworthiness. A framework from Google rater guidelines, not a measurable score. Nothing on a page can be tested to return an E-E-A-T value.
The machine-readable statements a site makes that a reader could use as evidence of experience, expertise, authoritativeness and trust: a named person and what qualifies them, corroborating profiles, typed authors and dates on articles, who runs the organisation, how to reach it, and whether the about, contact and policy pages are linked. Each is measurable; none of them IS the quality it stands in for.
The idea that covering a subject thoroughly improves standing across it. Widely held and hard to test on one site, because coverage and quality usually change together.
A page that answers HTTP 200 while telling a human the content is missing. The status line and the content disagree, and the crawler believes the status line until something else corrects it.
Whether the site's own links and sitemap entries agree on how a URL is written: lowercase paths, one trailing-slash form, one host form, no navigation through query strings. To a crawler /About and /about are two URLs, and a site that links to both has published every page twice. Scored from the hrefs and sitemap entries the site itself serves, never from the paths the scanner chose to request.
What can be read about accessibility from the document a server returns, without rendering it: a lang attribute, a title, alt on images, names on form controls and buttons, titles on frames, a main landmark, a skip link or landmarks, unique ids. A screen reader depends on all of these before contrast or focus order ever matters. Measured here; a WCAG audit is not.
An HTML region a screen-reader user can jump to directly: main, nav, header, footer, aside, or an element with the matching role. One main landmark is the target for skipping the navigation; two is as unusable as none.
The text assistive technology announces for a control: visible text, a label, aria-label or aria-labelledby, an image's alt, a frame's title. A button with none is announced as "button" and nothing else; an input with none as "edit text". A placeholder is not a name.
A link, usually first in the document, that jumps past the navigation to the main content. Without it, or without landmarks that serve the same purpose, every page starts by tabbing through the whole menu.
The element, main or role=main, that marks the primary content of a page so a screen reader can jump to it. Exactly one per page. A theme that prints it on its own templates often prints nothing on a page builder's, so the pages that matter are the ones to check.
The text alternative on an image. Present and empty marks the image as decorative and is correct; present and descriptive is read aloud; absent entirely makes a screen reader announce the file name. The defect is absence, and a check that fails on empty alt reports every decorative icon as a fault.
What assistive technology announces for an input: a label element pointing at its id, a wrapping label, aria-label, aria-labelledby, or a title. A placeholder is not a name; it disappears on the first keystroke and is not announced as the field's purpose.
Two elements on one page with the same id attribute. Labels, aria-labelledby, aria-describedby and skip links all resolve by id and land on the first match, which is usually the wrong element. Invalid HTML that validators tolerate and screen readers do not.
Measured: HOST_DUPLICATE_200 3.3%
A push protocol in which the site tells participating search engines that a URL changed, instead of waiting to be recrawled: a key file on the host, then one request per changed URL. It shortens the delay for engines that participate, chiefly Bing, Yandex, Naver and Seznam. Google does not participate, and the AI answer engines largely do not, so a submission proves the site announced a change; it does not put the page in front of any answer engine.
Filter combinations, size, colour, price, sort order, that each generate a URL, so that a catalogue of a hundred products can expose millions of addresses. It is the ordinary-case cause of a crawler trap, and the usual reason crawl budget is spent on pages that all show the same items. The fix is a rule about which combinations are canonical, which are noindexed, and which are blocked; case consistency in the parameter names is the related failure, since ?Color= and ?color= are two more URLs.
A page whose visible text is too short, too generic or too duplicated to answer the query it targets. The word count alone does not define it: a 783-word parking page is thin and a 200-word answer to a precise question is not. Engines judge it against the intent, and so should any check that claims to measure it.
The same or substantially the same text reachable at more than one URL, on one site or across sites. Engines pick one copy to index and the choice is theirs unless a canonical, a redirect or a noindex makes it. Most audit tools accuse sites of it by heuristics; measuring it means comparing text, which almost none of them do.
Two pages whose core text overlaps substantially once shared navigation, footer and boilerplate are removed. Measured by comparing sets of overlapping word sequences, and reported in bands rather than percentages, because the estimate quantises and a precise-looking number would overstate what the method knows.
The text that appears on most pages of a site: navigation, footer, cookie notice, sidebar. It puts a floor under every duplicate-content comparison, so it must be removed before pages are compared, and it inflates a page's word count without adding anything an engine would quote.
A technique for estimating how similar two sets are from small fixed-size sketches of each, so that page overlap can be compared without storing the pages. With 64 hashes the estimate is within a few percent of the exact value, which is enough to place a pair in a band and not enough to publish as a number.
A run of consecutive words, typically eight, taken as one unit when comparing texts. Two pages' shingle sets overlap in proportion to how much text they share in the same order; single-word overlap would count every page in the same language as similar.
A link to a fragment, such as /pricing#fix, whose target id does not exist on the page. The page loads and the browser stays at the top, so nothing reports an error; a link check that follows the URL and ignores the fragment passes it. It is found by reading the target page's ids.
Measured: BROKEN_INTERNAL_LINK 2%
Text that was never meant to ship: lorem ipsum, 'Test content', an empty template section, a string like NaN or undefined leaked from code. A check for it must match whole tokens, because a substring match accuses any URL that happens to contain the letters.
Measured: PLACEHOLDER_CONTENT 0.2%
A transparent pixel, a data URI or a tiny SVG served where a photograph will be inserted later by script, usually by a lazy loader. To a reader of the delivered bytes, which is what a crawler is, the page has that many images and none of them are photographs. Counting placeholders as images is a measurement error.
The text alternative for an image, read by screen readers and by any client that does not render images. Its absence is an accessibility failure; its presence is not evidence of a photograph, because a placeholder can carry alt text too. Decorative images should carry an empty alt rather than none.
Metadata embedded in a photograph by the camera or editor: capture time, device, sometimes coordinates. Its presence is weak evidence that an image is an original photograph rather than stock; its absence proves nothing, because most pipelines strip it. Coordinates in it can leak a location the page did not mean to publish.