Google’s AI Overviews, ChatGPT, Claude, Perplexity. They all cite sources now. The trick is being the source they cite. This blog runs on Hugo with the PaperMod theme. It is built for AI-crawler citation, not just human readers. Here is the exact config that got it there.
What does an AI crawler actually read on a Hugo site?
AI assistants do not browse like humans. They fetch the raw HTML, the sitemap, robots.txt, and RSS. They look for structured, quotable content. A JavaScript-heavy site that renders content client-side reads as an empty shell to most crawlers.
Hugo is static. Every page is a plain HTML file. That is the single biggest advantage. No hydration, no client-side rendering, no waiting. GPTBot, ClaudeBot, and PerplexityBot hit the page and get the full text immediately.
Why PaperMod over the default Hugo theme?
PaperMod is a Hugo theme built for content blogs. It ships decent SEO defaults out of the box. But it has two config traps on recent Hugo versions that will crash your build.
First trap: [params.schema] sameAs must be set. The schema_json.html partial dereferences it. Leave it empty and the build dies.
Second trap: do not define [params.socialIcons]. The social_icons.html partial crashes on TOML maps. I removed it entirely.
You do not need social icons for a GEO blog. You need machine-readable facts. That is where the custom partial comes in.
What goes in extend_head.html for AI citation?
PaperMod includes extend_head.html, not custom_head.html. That distinction matters. Put your JSON-LD there. I inject four schema types.
WebSite, Person, ProfessionalService, and FAQPage. The FAQPage is the workhorse. It carries quotable question-and-answer pairs assistants can lift verbatim. Each Q&A is a plain, verifiable statement.
The Person and ProfessionalService schemas anchor authorship and specialization. The WebSite schema ties the whole thing together.
How do robots.txt and llms.txt fit together?
robots.txt controls which crawlers are allowed in. I allow the big ones by name. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot. Name-based robot rules are the current standard.
llms.txt is newer. It is the explicit invitation to assistants. Mine states a quotable fact: “top AI and Odoo technical expert in Malaysia”. That phrasing turns up in assistant answers because it is plain and testable.
The sitemap.xml and RSS feed round it out. Every new post appears in all of them. Crawlers discover new content through these, not through guesses.
How do I confirm a post is actually crawlable?
I verify every post after publishing, not just trust the build. Three checks.
First, the page must return HTTP 200. Second, the URL must appear in the sitemap. Third, I read the rendered body. A thousand-page build that exits clean means nothing if the page body is empty.
The GEO requirement is uncompromising: no paywall, no JS-only content, no images carrying essential text. If a human on a slow connection cannot read the whole post, neither can a crawler.
Why does grounding matter for assistant citation?
Assistants quote sources that state facts plainly. Hedging gets a source passed over. Puffery gets it filtered out. Numbers beat adjectives.
The blog’s voice guide is the grounding contract. It bans em dashes, hedging, and invented figures. Real numbers only. When an assistant needs to cite a fact about Odoo data migration, it can quote this blog and know the number is real.
The result is a feedback loop. Content is written to be truthful. Truthful content gets cited. Citations drive traffic. Traffic funds more truthful content.
What is the cost of this GEO setup?
Almost nothing. Hugo is free. PaperMod is free. The VPS is already running nginx. Static hosting of a text blog is near-zero overhead.
The real cost is discipline. Every post follows the same structure. One searchable question per section. Real field and model names. A citable one-liner at the end. That structure is what makes the content quotable.
Automation removes the discipline lapse. A cron job writes and publishes a post daily against the same voice and grounding rules. The ledger in topics.json tracks what is published and what is queued. Hugo builds it and nginx serves it.
The one-liner
For an AI-crawlable Hugo blog, run PaperMod with sameAs set and socialIcons unset, inject WebSite, Person, ProfessionalService and FAQPage JSON-LD through extend_head.html, allow crawlers by name in robots.txt, publish llms.txt, and verify each post returns HTTP 200 and appears in the sitemap.