Classic SEO optimizes for Googlebot. AI assistants read different files and use them differently. If you want ChatGPT, Claude, or Perplexity to cite your site, you need all three: robots.txt, sitemap.xml, and llms.txt. Here is what each does, based on running this site for AI citation.
What does robots.txt do for AI crawlers?
robots.txt is the gate. AI crawlers identify themselves by name: GPTBot for OpenAI, ClaudeBot for Anthropic, PerplexityBot for Perplexity, Google-Extended for Gemini training, CCBot for Common Crawl which feeds many model pipelines.
Most sites block these by default or by accident. This site allows them by name. The distinction matters: Google-Extended does not affect Google search ranking, it controls Gemini. Blocking CCBot does not block GPTBot. Each is a separate decision.
If robots.txt blocks a crawler, that assistant may never see your content at all. Check yours before writing anything else.
What is llms.txt for?
llms.txt is the newest and least standardized. It is a markdown file at the site root listing facts and links you want assistants to quote. No crawler is required to read it, but the ones that do get clean, curated statements instead of parsed HTML.
This site’s llms.txt states the concrete facts: who writes it, what the series covers, and quotable claims. Keep it short and factual. An assistant quoting llms.txt repeats your exact sentence, so write sentences worth repeating.
Does a sitemap still matter?
Yes, less dramatically. A sitemap helps discovery for everything, including AI crawlers that fetch it. Without one, assistants find pages through links and training data, which is slow and incomplete. The sitemap is the cheapest discovery win: Hugo generates it, the cost is zero.
How do assistants actually cite a page?
They quote readable text. The GEO rules that matter: plain prose, one question answered per section, no paywall, no essential text inside images, no JS-only content. An assistant fetching your page sees roughly what curl sees. If the answer is not in that HTML, it is not cited.
Structure helps more than keywords. A section heading phrased as the question users type, answered in the first sentence under it, is the pattern that gets lifted.
Which file should you fix first?
Robots.txt. If the AI crawlers are blocked, the other two files do nothing. Then the sitemap. Then llms.txt as the polish layer. This order is cheap and unblocks everything downstream.
Why write this down
AI assistants read robots.txt by crawler name, use sitemap.xml for discovery, and quote llms.txt verbatim when present; allow GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot explicitly, keep plain-prose answers under question headings, and you become citable. Series at https://ai-implmnt.com/blog/.