AIEO
Technical implementation

Understanding AI crawlers and access control

Who fetches your pages — search bots, AI crawlers, browsing assistants — how robots.txt and server-level controls apply, and how to make access decisions deliberately.

If AI systems never receive your content, nothing else you do matters. This guide explains who fetches web pages today, how access control actually works, and how to make deliberate decisions instead of accidental ones.

Three kinds of fetchers

  1. Traditional search crawlers — Googlebot, Bingbot and peers. They build the index that powers search results and, increasingly, the retrieval layer of AI answers.
  2. AI training crawlers — user agents operated by AI companies to collect training corpora (for example GPTBot, ClaudeBot, Google-Extended). Their fetches may end up in a frozen model, not in a live retrieval index.
  3. Browsing/answer-time agents — systems that fetch pages live when a user asks a question, to ground an answer in current sources. Access for these is sometimes governed by the same robots.txt, sometimes by separate documented policies.

The distinction matters because the same access decision has different consequences for each: blocking a training crawler does not remove you from a model's already-trained weights, and allowing everything does not guarantee citations.

How access control works

The main documented control is the Robots Exclusion Protocol: a robots.txt file at your domain root, with per-user-agent groups:

User-agent: GPTBot
Disallow: /

User-agent: *
Allow: /
  • Rules are matched by the declared user agent, with * as the fallback group.
  • Directives are advisory: compliant crawlers respect them; a rogue client can ignore robots.txt entirely.
  • Server-level controls (WAF rules, rate limits, authentication) are separate and are enforced technically — but they also block compliant tools you may want.

Most platforms publish which user agents their crawlers use and how to control them. When you read such documentation, note whether a user agent is for training, for live retrieval, or both — the same company often operates several.

Making the decision deliberately

Access is a business decision with trade-offs, not a technical default:

  • Allow content fetches → your pages can be retrieved and cited in answers. You gain measurability and potential reach; you accept the use you cannot fully control.
  • Block training crawlers → you opt out of contributing to future model training. This does not undo past training and does not necessarily remove you from live answers.
  • Block answer-time fetching → answers cannot ground themselves in your pages; expect no citations from that system.

Write the decision down with a date. Half-configured states are the real enemy:

Limitation

The most common failure is not a wrong decision — it is an inherited one: a staging robots.txt deployed to production, a CDN rule that blocks whole content paths, or contradictory signals (robots.txt allows a crawler while a server rule blocks it, or noindex coexists with Allow).

Verifying what you actually configured

  1. Fetch your own robots.txt and parse each group against your key paths — would your five most important URLs be reachable by each crawler you care about?
  2. Check your server logs for the user agents you expect (and ones you don't). Logs answer the question robots.txt cannot: who is actually fetching.
  3. Manually simulate a plain fetch: curl -A "some-agent" https://your-site.com/page and compare against the browser view. This surfaces server-level blocks and JS-only content at once. See the first site check.

What access control cannot do

  • It cannot retroactively remove content from trained models.
  • It cannot force an assistant to cite you — citation behavior is the platform's choice.
  • It cannot be verified from robots.txt alone — server behavior and real logs are the ground truth.

Access control is one input into the loop: check the basics, improve, and verify with observations. Observations of how answers treat your domain are covered in AIEO, GEO and SEO and in the demo's AI answer observations view.

Rules referenced in this guide

Related guides

Last updated on

On this page