Rule reference
The catalog of check rules referenced by AIEO findings — categories, check types, scope and limitations, with links to the guides that teach each fix.
Findings in AIEO reports and demos reference stable rule IDs. This page is the definition of every rule: what problem it describes, how it is checked, where it applies, and which guide teaches the fix.
Note
This is a catalog, not a detection engine. It defines vocabulary and meaning. The three check types have different trust levels — keep them apart when reading any report.
Check types
- Rule-based check — machine-verifiable against the recorded evidence. Two systems following the same rule reach the same verdict.
- Human judgment — the evidence is shown, but the conclusion depends on context only a person can weigh.
- Sampled observation — a measurement over a stated sample, valid only with its sample size, conditions, and failure handling.
Crawling & access
CRAWL-ROBOTS-001
robots.txt does not block key resourcesRule-based checkrobots.txt disallows paths that contain primary content, or the site unknowingly blocks AI crawler user agents. Blocking is a decision — the problem is doing it without knowing.
- Scope
- Parses robots.txt directives against a list of key paths and known crawler user agents. It cannot detect server-level blocks (WAF, rate limits).
- Basis
- The Robots Exclusion Protocol is a documented standard; major AI providers publish their crawler user agents and how to control them.
CRAWL-RENDER-001
Primary content present in server HTMLRule-based checkThe main content is only built by client-side JavaScript, so a plain HTTP fetch (no browser) returns an empty shell. Some crawlers render pages, others do not — content that is not in the HTML is at risk.
- Scope
- Compares text volume and key phrases between the raw HTML response and the rendered page. It flags risk; it cannot guarantee how any specific crawler processes your site.
- Basis
- Crawler documentation across search engines describes rendering pipelines and their limits; client-only content is a well-known failure mode.
Metadata & indexing
SEO-TITLE-001
Unique, descriptive title tagRule-based checkThe page is missing a title element, repeats another page’s title, or has a title that does not describe the page content. Search engines and AI systems use titles as a primary label for the page.
- Scope
- Applies to every indexable HTML page. Does not evaluate whether the wording is persuasive — only presence, uniqueness and rough relevance.
- Basis
- Title elements are a long-standing, documented indexing signal for search engines and are commonly quoted by AI answers when naming a source.
SEO-METADESC-001
Meta description present and meaningfulRule-based checkThe page has no meta description, or it is a template string that does not summarize the page. A missing description leaves the summary shown in search results to chance.
- Scope
- Applies to indexable pages. Presence and non-generic content are checked; actual snippet rendering in search engines is not guaranteed and is out of scope.
- Basis
- Meta descriptions are a documented way to influence result snippets; they are frequently reused verbatim or in part by answer engines.
SEO-CANONICAL-001
Consistent self-referencing canonical URLRule-based checkThe page declares a canonical URL that conflicts with itself (e.g. different protocol/host), or different variants of the same page declare different canonicals. This splits signals and confuses crawlers about which URL to keep.
- Scope
- Applies to pages that should be indexed under one URL. Parameters, pagination and cross-domain syndication need explicit decisions per case.
- Basis
- rel=canonical is a documented link element supported by all major search engines.
Page structure
SEO-HEADING-001
Single h1 with a logical heading orderRule-based checkThe page has zero or multiple h1 elements, or heading levels jump around (h2 → h4). Structure signals help both readability and machine parsing of sections.
- Scope
- Mechanically checkable: counts and level ordering. Whether heading texts are meaningful requires human review.
- Basis
- Heading structure is documented in HTML semantics and is used by parsers to build document outlines.
Content quality
CONTENT-EVIDENCE-001
Claims are backed by evidenceHuman judgmentThe page makes claims (numbers, comparisons, "best", "proven") without sources, data or examples. Answers built from such pages inherit weak grounding, and readers have no reason to trust them.
- Scope
- Requires human judgment: what counts as sufficient evidence depends on the claim. The check can only surface claims that appear unsupported.
- Basis
- Answer engines weight corroborated content when selecting sources to cite; supported claims are also more likely to be quoted accurately.
CONTENT-STRUCTURE-001
Self-contained sections that answer real questionsHuman judgmentSections depend on context scattered across the page, or headings do not correspond to questions people actually ask. Answers and snippets are built from fragments — fragments that stand alone get quoted.
- Scope
- Human judgment over heading wording and section completeness. No automated verdict is possible; a check can only organize the review.
- Basis
- Answer systems extract passage-level content; writing for standalone passages is a documented content practice.
CONTENT-FRESHNESS-001
Visible, truthful update datesRule-based checkThe page shows no date, or a date that changes on every deploy without content changes. Fake freshness erodes trust with both readers and systems that compare versions.
- Scope
- Detectable: presence of a date and whether it matches the actual content revision history.
- Basis
- Visible dates are a common trust signal; search documentation warns against misleading date manipulation.
AI answer observations
OBS-MENTION-001
Brand mention rate in sampled AI answersSampled observationAcross a fixed question set, answers do not mention the brand, or mention competitors instead. This is an observation about samples — it has a denominator and conditions, not a verdict.
- Scope
- Always reported with: question set, sample size, valid response count, language/region, assistant identity and date. A low rate in one sample does not prove a brand cannot be seen by AI.
- Basis
- Mention counts over repeated samples are the only honest way to talk about "AI visibility"; single anecdotes are not measurement.
OBS-ACCURACY-001
Accuracy of AI descriptions of the brandSampled observationAnswers that mention the brand get facts wrong: wrong positioning, outdated pricing, confused product names. Inaccuracy usually traces back to stale or contradictory public sources.
- Scope
- Human review of sampled answers against current official sources. Reports which fact was wrong and where the correct value is published.
- Basis
- Answer systems synthesize from public sources; the actionable fix is on the source pages, not on the assistant.
OBS-CITATION-001
Citation of owned sources in AI answersSampled observationWhen the brand is mentioned, the answer cites third parties only — or nothing at all. Citations are the measurable link between answers and your site.
- Scope
- Counted over the same samples as mentions. Whether a specific assistant cites links at all is a platform behavior, not something a site can control.
- Basis
- Citation presence varies by assistant and answer type; treating it as a site-side failure would misread the measurement.
Related guides
- How to complete your first site checkA repeatable, evidence-first procedure for checking the basics of any website — crawlability, server HTML, titles, canonicals, headings and dates. No special tools required.
- AIEO, GEO and SEO — how they relateWhat SEO, GEO and AIEO actually mean, what changes with AI answers, what does not, and what honest optimization can and cannot promise.
- Understanding AI crawlers and access controlWho fetches your pages — search bots, AI crawlers, browsing assistants — how robots.txt and server-level controls apply, and how to make access decisions deliberately.
Last updated on
How to complete your first site check
A repeatable, evidence-first procedure for checking the basics of any website — crawlability, server HTML, titles, canonicals, headings and dates. No special tools required.
Understanding AI crawlers and access control
Who fetches your pages — search bots, AI crawlers, browsing assistants — how robots.txt and server-level controls apply, and how to make access decisions deliberately.