What is llms.txt, and should my site have one?

What is llms.txt, and should my site have one?

What is llms.txt, and should my site have one?

THE SHORT ANSWER

An llms.txt file is a plain-text file at the root of your domain that tells AI systems what your site is, what it offers and which pages matter most. It is not an official standard and no engine is obliged to read it, but it costs an hour to write, removes ambiguity about what your business does, and is trivially easy to keep current.

What goes in it

A short statement of what the business does and for whom. The commercial model in plain language, including pricing if you publish it. A list of the pages that carry your substantive answers, with one line describing each. Contact and company details. Nothing more.

Write it as prose and links, not as marketing copy. The audience is a system trying to establish facts, not a visitor being persuaded.

What it is not

It is not a ranking mechanism, and it does not make an engine cite you. It is closer to a README for your domain: it reduces the chance of being described incorrectly, and it gives a retriever a shortcut to your best pages.

It also does not replace a sitemap. Keep both; they answer different questions.

Whether it is worth the hour

For a business whose category is easy to misdescribe, yes. If an assistant might confuse an exclusive lead supplier with a marketplace, or a local installer with a national franchise, one paragraph of unambiguous self-description is cheap insurance.

For a site with three pages and no proprietary information, the marginal value is small.

What a working example looks like

The version on flockleads.com opens with a single sentence stating that Flock Leads sells exclusive home-improvement leads to installers in twelve European markets, that every lead goes to one buyer only, and that packages run from ten to ninety leads a month. Then it lists the answer pages by theme: cost per lead by trade and country, how exclusivity is enforced, what qualification means in practice, and how lead-to-customer ratios differ by trade.

Each line is one URL plus one clause describing what question that page settles. A retriever reading it can go straight to the page that answers the prompt it was given, instead of guessing from a navigation menu built for human eyes.

How to keep it from going stale

The failure mode is not a missing file, it is a file that contradicts the site. If your pricing page says one thing and llms.txt says another, you have manufactured exactly the ambiguity the file was supposed to remove.

Two habits prevent that. Reference pricing by band rather than by exact figure where the figure moves, and put the file on the same review cycle as the pricing page itself, so the two can never drift more than one release apart. If you generate pages programmatically, generate the file from the same source of truth and the problem disappears entirely.

How to check whether anything reads it

Server logs are the only honest answer. Filter requests for the exact path and group by user agent: GPTBot, PerplexityBot, ClaudeBot, Google-Extended and Bytespider are the agents worth watching. If you see no hits after a month, the file is doing nothing for retrieval, though it still costs you nothing to keep.

The second test is behavioural rather than technical. Ask four assistants to describe your business in one sentence, before and a month after publishing. If the description sharpens, something in your structured self-description landed, whether it was this file, your schema markup or your answer pages.

Related answers

Frequently asked questions

Is llms.txt an official standard?

No. It is a community convention. No engine is required to read it, and some do not.

How long should the file be?

Short enough to read in a minute. A few hundred words of description plus a list of your substantive pages. Length is not a quality signal here.

Should I list every page on the site?

No. List the pages that settle a question. Cart, login and legal pages add noise and make the useful lines harder to find.

Can I block AI crawlers and still publish llms.txt?

You can, and it is contradictory. If robots.txt disallows the agents, the file will not be fetched. Decide which side of that line you want to be on first.

Where does the file go?

At the root of the domain, served as plain text, in the same place a robots.txt file would live.

Does llms.txt improve my rankings?

No. It clarifies what your site is and points to your key pages. Ranking and citation are decided elsewhere.

Do I still need a sitemap?

Yes. A sitemap enumerates URLs for crawlers; llms.txt describes the business and highlights the pages that matter. They do not substitute for each other.

NEED A CLEARER PLAN?

Let’s turn your next move into momentum.

Talk to us →