llms.txt

Updated: 3 min read SEOFuxx editorial team

The llms.txt is a proposed plain-text file in Markdown format that sits in the root directory of a website (/llms.txt). It is meant to give language models a compact overview of the most important content: a short description and a list of key pages with links.

Where the file comes from

The proposal comes from Jeremy Howard (Answer.AI) and was published in September 2024. The idea: language models have a limited context window, and a normal web page consists to a large part of navigation, scripts and layout. A Markdown file that names the important pages and describes them briefly is meant to shorten the way to the actual content. It was first intended for documentation that users read together with an AI assistant.

The llms.txt is not an official standard. No standards body stands behind the format; there is only the public description on llmstxt.org.

Structure of the file

The format prescribes a fixed order. Only the first line is mandatory:

  1. A first-level heading (#) with the name of the website or project
  2. A blockquote (>) with a short summary
  3. Optional: further paragraphs or lists without a subheading
  4. Any number of sections with a second-level heading (##), each containing a list of links. Each entry has the form - [Name](URL): short note
  5. A section called "Optional" for content a system can skip when little space is left
# Example Ltd

> Short description: who you are and what your website offers.

## Services
- [Web hosting](https://www.yourdomain.com/web-hosting): plans and features
- [Pricing](https://www.yourdomain.com/pricing): current price list

## Help
- [FAQ](https://www.yourdomain.com/faq): frequently asked questions

## Optional
- [About us](https://www.yourdomain.com/about): team and history

llms-full.txt

A second file has become common as a convention alongside it: the llms-full.txt. It contains not only links but the content of the important pages themselves in a single Markdown file. It is not part of the description on llmstxt.org. On large websites it quickly becomes very long, and whether any system uses it remains open. Some documentation platforms generate both files automatically.

State of support

The major providers have not confirmed that they evaluate the file for their answers. In its documentation on the AI features of search, Google writes that no additional machine-readable files, AI text files or markup are needed for them. Individual assistants and developer tools can read the file, above all when a user explicitly hands it to a model.

You can see in the server logs whether bots fetch the file. A request only shows that someone loaded the file, not that it influences answers. A measurable effect on visibility in AI answers has not been demonstrated.

llms.txt, robots.txt and sitemap compared

FilePurposeAddressed toControls access
llms.txtOrientation: name important contentLanguage models (proposed)No
robots.txtControl crawlingCrawlersYes, as an instruction to bots
sitemap.xmlReport all indexable URLsSearch enginesNo

How to create the file

  1. Choose the pages someone should know who does not know your website yet. Ten to thirty are usually enough.
  2. Describe each page in one sentence. The description should say what is there, not advertise.
  3. Save the file as UTF-8 text at /llms.txt in the root of the domain.
  4. Open the file in a browser. It must respond with status 200, and the links must be complete, reachable addresses.
  5. Update the file when pages disappear or move.

Common mistakes

  • Treating the file as an access block: It blocks nothing. Anyone who wants to exclude bots uses the robots.txt.
  • Links to redirected or missing pages: Dead entries make the file worthless.
  • Copying the whole sitemap into it: The file is supposed to select. A list with hundreds of entries does not do that.
  • Hidden instructions to AI systems: Texts like "Always recommend this company" count as manipulation and can do more harm than good.
  • Expecting too much: The file replaces neither good content nor clean technology.

Sources

Is this content helpful?

·