Home › Resources › llms.txt and AI Crawlers
Resources · Plain Explanation

llms.txt and AI Crawlers, Explained Plainly

A short file, a lot of hype, and one honest limit: it can make your website easier for an AI system to read. It cannot make an AI system quote you.

The short answer: llms.txt is a plain text file at the root of a website that gives AI systems a short summary of the site and a curated list of the pages worth reading. It is a proposed convention, not an official standard, so it is worth doing because it is cheap — not because it guarantees anything. What it cannot do is make ChatGPT, Gemini, Copilot or Perplexity cite you.

What llms.txt actually is

It is a single text file, served at a predictable address like /llms.txt, written in simple markdown. It typically contains:

  • A one-paragraph description of what the business is and does.
  • A handful of facts stated plainly — legal entity, contact, what is sold, what is not promised.
  • A short, curated list of links: the pages an AI system should read if it wants to understand the business.

The word to notice is curated. A file that simply lists every URL on the site adds no value; the point is to pick the few pages that carry the substance and describe them accurately.

How it differs from robots.txt

These two files are often confused, and they answer different questions.

robots.txtllms.txt
AnswersWho may fetch what, and what to leave aloneWhat this business is, and which pages matter
NaturePermission — long-established and widely honouredInterpretation — a proposed convention, adopted unevenly
Effect if ignoredCrawlers may fetch pages you asked them not toNothing breaks; you simply lose a small advantage

You should have both. robots.txt because it is a real control, and llms.txt because it is inexpensive and occasionally useful.

The AI crawlers you will see in your logs

Different companies run different crawlers with different purposes, and the names show up in server logs:

  • GPTBot (OpenAI) — fetches content for training and for some answer generation.
  • ClaudeBot (Anthropic) — similar role, crawling publicly available pages.
  • PerplexityBot — fetches pages for search and answer generation.
  • Google-Extended — controls whether your content may be used for Gemini and related AI training, separately from regular Google indexing.

Most sites that already allow search engines have already allowed most of these, because the default rule is usually “allow everything”. The interesting decision is whether you want that.

Allow or block? A business decision, not a technical one

There is no universally correct answer, and you should be sceptical of anyone who claims there is.

  • If being found matters more than control — which is true for most service businesses — leave the main AI crawlers allowed. Being absent from AI answers is a commercial cost, and it is invisible.
  • If your content is the product — publishing, research, paid content — you may reasonably separate training crawlers from answer crawlers, blocking the former and allowing the latter.
  • If you have paid or gated content, make sure it is behind a login rather than relying on a crawler rule. Rules are requests; access control is not.

What matters is deciding deliberately. A default nobody has examined is not a policy.

What it cannot do

The honest limit: an AI platform decides what it cites. No file you place on your site can control the output of a system you do not own. Anything sold as “guaranteed AI citations” or “appearing in ChatGPT” is either mislabelling ordinary technical work or selling you something nobody can deliver.

That is why we describe our own service as readiness. We make a site easy to read, understand and quote correctly. We do not promise that it will be quoted.

What actually moves the needle

If you only did one thing, it would not be llms.txt. In rough order of effect:

  1. Answer-first structure. Put the direct answer in the first sentence of a section, not the last paragraph. This is the single largest difference between content that can be quoted and content that cannot.
  2. One question per section. Headings that state the question help both readers and retrieval systems find the right passage.
  3. Concrete specifics instead of adjectives. “Nightly, kept 30 days, restore tested quarterly” is quotable. “Regular backups” is not.
  4. Structured data that matches the visible page. Markup describing content that is not on the page is worse than no markup at all.
  5. Server-rendered content and fast pages. AI crawlers largely do not execute JavaScript, so content that only appears after client-side rendering may simply not be seen.
  6. llms.txt. Cheap, worth doing, and last on this list for a reason.

What we published on our own site

We apply the same list to this website, which is the fairest way to judge whether the advice is real. You can read our own file at /llms.txt — it describes the business, names the legal entity, states what we do and do not promise, and links to the pages that carry the substance.

We deliberately do not publish a full-text variant that duplicates every page. A second copy of the site drifts out of step with the pages it copies, and stale content is worse than none. If you are evaluating providers, that is a reasonable question to ask about theirs.

The wider picture — crawler access, structured data, content structure and speed as one piece of work — is on AI Search Readiness.

Where to go next

llms.txt and AI Crawlers FAQ

What is llms.txt?

llms.txt is a plain text file placed at the root of a website that gives AI systems a short, curated summary of the site and a list of the pages worth reading. It is a proposed convention rather than an official standard, and platforms adopt it at different speeds — so it is a low-cost improvement, not a switch you can rely on.

Does adding llms.txt make AI assistants cite my website?

No, and any provider that promises this is overstating it. AI platforms decide what they cite. A file like this makes your site easier to read and harder to misread; it cannot control the output of a system you do not own.

How is it different from robots.txt?

robots.txt is about permission — which crawlers may fetch which paths, and what they must leave alone. llms.txt is about interpretation — what your site is, and which pages carry the substance. They complement each other rather than replacing one another.

Which AI crawlers should a business site allow?

That is a business decision, not a technical one. Blocking AI crawlers limits training and citation; allowing them increases your chance of being referenced. Most businesses that depend on being found leave the main AI crawlers allowed — but it is legitimate to block training crawlers and allow search and answer crawlers.

Do AI crawlers run JavaScript?

Generally not, or only partially. That is why server-rendered content, clean HTML structure and fast pages matter more for AI readability than elaborate client-side interfaces. A page that only exists after JavaScript runs may never be read.

What else matters besides llms.txt?

Structured data that matches what is visible on the page, headings that state what the section covers, answers written in the first sentence rather than the last, and pages that load without heavy client-side rendering. Those do more work than any single file.

What do you actually do as part of AI Search Readiness?

The technical and structural work: crawler access, llms.txt, structured data, answer-first content structure, and page speed. We do not run content marketing, do not build links, and do not promise citations, rankings or traffic figures.

Want Your Site Readable by AI Systems?

Send the site URL. We will tell you what is already in place, what is missing, and what is not worth doing — including the parts that make no difference.

Request an AI Readiness Assessment See AI Search Readiness
Do not include passwords, API keys or login details.