llms.txt
The llms.txt is a proposed plain-text file in Markdown format that sits in the root directory of a website (/llms.txt). It is meant to give language models a compact overview of the most important content: a short description and a list of key pages with links.
Where the file comes from
The proposal comes from Jeremy Howard (Answer.AI) and was published in September 2024. The idea: language models have a limited context window, and a normal web page consists to a large part of navigation, scripts and layout. A Markdown file that names the important pages and describes them briefly is meant to shorten the way to the actual content. It was first intended for documentation that users read together with an AI assistant.
The llms.txt is not an official standard. No standards body stands behind the format; there is only the public description on llmstxt.org.
Structure of the file
The format prescribes a fixed order. Only the first line is mandatory:
- A first-level heading (
#) with the name of the website or project - A blockquote (
>) with a short summary - Optional: further paragraphs or lists without a subheading
- Any number of sections with a second-level heading (
##), each containing a list of links. Each entry has the form- [Name](URL): short note - A section called "Optional" for content a system can skip when little space is left
# Example Ltd
> Short description: who you are and what your website offers.
## Services
- [Web hosting](https://www.yourdomain.com/web-hosting): plans and features
- [Pricing](https://www.yourdomain.com/pricing): current price list
## Help
- [FAQ](https://www.yourdomain.com/faq): frequently asked questions
## Optional
- [About us](https://www.yourdomain.com/about): team and history
llms-full.txt
A second file has become common as a convention alongside it: the llms-full.txt. It contains not only links but the content of the important pages themselves in a single Markdown file. It is not part of the description on llmstxt.org. On large websites it quickly becomes very long, and whether any system uses it remains open. Some documentation platforms generate both files automatically.
State of support
The major providers have not confirmed that they evaluate the file for their answers. In its documentation on the AI features of search, Google writes that no additional machine-readable files, AI text files or markup are needed for them. Individual assistants and developer tools can read the file, above all when a user explicitly hands it to a model.
You can see in the server logs whether bots fetch the file. A request only shows that someone loaded the file, not that it influences answers. A measurable effect on visibility in AI answers has not been demonstrated.
llms.txt, robots.txt and sitemap compared
| File | Purpose | Addressed to | Controls access |
|---|---|---|---|
| llms.txt | Orientation: name important content | Language models (proposed) | No |
| robots.txt | Control crawling | Crawlers | Yes, as an instruction to bots |
| sitemap.xml | Report all indexable URLs | Search engines | No |
How to create the file
- Choose the pages someone should know who does not know your website yet. Ten to thirty are usually enough.
- Describe each page in one sentence. The description should say what is there, not advertise.
- Save the file as UTF-8 text at
/llms.txtin the root of the domain. - Open the file in a browser. It must respond with status 200, and the links must be complete, reachable addresses.
- Update the file when pages disappear or move.
Common mistakes
- Treating the file as an access block: It blocks nothing. Anyone who wants to exclude bots uses the robots.txt.
- Links to redirected or missing pages: Dead entries make the file worthless.
- Copying the whole sitemap into it: The file is supposed to select. A list with hundreds of entries does not do that.
- Hidden instructions to AI systems: Texts like "Always recommend this company" count as manipulation and can do more harm than good.
- Expecting too much: The file replaces neither good content nor clean technology.
Related terms
- AI crawlers: which bots fetch content and how you control them
- robots.txt: regulate access to your website
- Sitemap.xml: report URLs to search engines
- GEO (Generative Engine Optimization): optimize content for AI answers
Sources
- llmstxt.org: description of the format
- Howard, J.: /llms.txt – a proposal to provide information to help LLMs use websites, Answer.AI, 3 September 2024
- Google Search Central: AI features and your website