How to complete your first site check
A repeatable, evidence-first procedure for checking the basics of any website — crawlability, server HTML, titles, canonicals, headings and dates. No special tools required.
This guide walks through a first pass over a website you control. It takes 30–60 minutes for a small site, requires only a browser and a terminal (or "view source"), and produces something more valuable than a score: a written list of findings with evidence.
The method works the same whether the site has five pages or five thousand — for larger sites, sample representative page types (home, article, product, landing) instead of checking everything.
Before you start: how to record findings
For each check below, write down three things:
- What you observed — the actual value or snippet, not a judgment. (
<title>= "Home — ACME", not "title is bad".) - The URL and date you observed it on.
- A status, chosen from: open, needs review, or not checked. "Not checked" is a legitimate result — never quietly record it as "passing".
This record is what turns checking into verification later: after you change something, you compare against the recorded evidence, not against a memory.
Step 1 — Can crawlers reach your pages?
Fetch https://your-site.com/robots.txt. Read it actively, looking for two things:
Disallow:rules that cover paths with real content (for example a broadDisallow: /left from staging).- User-agent specific blocks, especially ones naming AI crawlers you did not consciously decide about. See AI crawlers and access control for who these are.
If the file is missing or empty, that means "no restrictions" — acceptable, as long as it is a decision and not an accident.
What you want to end up with is a short robots.txt where every Disallow line maps to a deliberate exclusion (admin paths, search results, API endpoints) and no rule shadows a content directory.
About this evidence
Record the exact robots.txt lines you are relying on, with the date. Robots files change; your record is what makes later comparisons meaningful.
Step 2 — Is your content in the server HTML?
Open a page and view its source (Ctrl+U / right-click → "View Page Source" — not the DOM inspector). Search the raw HTML for a phrase that only exists in your main content.
- Found in source — crawlers that fetch without rendering see your content. Good baseline.
- Only in the rendered DOM — your content depends on client-side JavaScript. Some crawlers render, some do not; content that only exists after JS is at risk across the board.
The robust fix is server-side rendering (or at least server-injected key content), not "optimizing for specific bots".
Limitation
This manual test samples one page. Client-side-only content often hides in template variations — check one page per template type, not just the homepage.
Step 3 — Titles and meta descriptions
For each sampled page, extract the <title> and <meta name="description"> from the source.
Check:
- Unique — no two important pages share a title.
- Descriptive — the title states what the page is, not just the brand name.
- Present and meaningful description — not a template string repeated site-wide.
These remain the primary labels that search results and many answer systems use when naming your page.
Step 4 — Canonical URLs
Look for <link rel="canonical"> in the source.
- It should point to the URL you want this page indexed as — usually itself.
- Watch for canonicals pointing to a different protocol (
http://), a different host (www.vs apex), or all pages pointing at the homepage — the last one effectively tells systems your other pages do not exist as separate documents.
If you serve the same content under parameterized URLs, canonicals are how you nominate the clean version; make sure the nomination is consistent from every variant.
Step 5 — Heading structure
From the source (or the rendered outline), extract the headings in order.
- Exactly one
h1that states the page's subject. - Heading levels that do not skip (an
h2followed by anh4breaks the outline). - Headings whose text reads like the questions or topics people actually have — this is what passage-level extraction keys on.
Whether heading wording is good is a human judgment; the structure check above is mechanical.
Step 6 — Dates and freshness signals
Find where the page shows a date, if anywhere.
- A visible "last updated" date that matches real content changes builds trust with readers and systems that compare versions.
- A date that refreshes on every deploy — with no content change — is worse than no date.
If your build system stamps dates automatically, make sure the stamp comes from the content's version history (git), not from the build time.
After the check: prioritize and act
Rank your findings:
- Access problems first (robots, server HTML) — they invalidate everything downstream.
- Metadata and structure — cheap, mechanical fixes.
- Content evidence — slower work, human-judgment territory.
Then do the same pass on the sample report in the interactive demo — it shows how these findings look when a product organizes them: evidence, explanation, recommendation, and the scope of each claim.
Note
One check is a snapshot. The loop this site teaches — check, find, improve, verify — needs a second observation after your changes. Keep the first record; you will need it.
Rules referenced in this guide
Related guides
- AIEO, GEO and SEO — how they relateWhat SEO, GEO and AIEO actually mean, what changes with AI answers, what does not, and what honest optimization can and cannot promise.
- Understanding AI crawlers and access controlWho fetches your pages — search bots, AI crawlers, browsing assistants — how robots.txt and server-level controls apply, and how to make access decisions deliberately.
- Rule referenceThe catalog of check rules referenced by AIEO findings — categories, check types, scope and limitations, with links to the guides that teach each fix.
Last updated on
AIEO, GEO and SEO — how they relate
What SEO, GEO and AIEO actually mean, what changes with AI answers, what does not, and what honest optimization can and cannot promise.
Rule reference
The catalog of check rules referenced by AIEO findings — categories, check types, scope and limitations, with links to the guides that teach each fix.
AIEO, GEO and SEO — how they relate
What SEO, GEO and AIEO actually mean, what changes with AI answers, what does not, and what honest optimization can and cannot promise.
Understanding AI crawlers and access control
Who fetches your pages — search bots, AI crawlers, browsing assistants — how robots.txt and server-level controls apply, and how to make access decisions deliberately.