llms.txt: What It Is, What It Does, and Whether You Need One
What llms.txt is, where it came from, and how it differs from robots.txt and sitemap.xml.
The honest state of adoption: who has said they use it, who has said they don't, and why that matters for your priorities.
A practical template, the mistakes I see most, and where llms.txt fits in a real GEO plan.
KEY TAKEAWAYS
- check_circlellms.txt is a proposed standard: a Markdown file at your site root that gives language models a short, curated map of your most important content.
- check_circleIt is not a crawling control. It doesn't block or allow anything. robots.txt still does that job.
- check_circleAdoption by the big AI platforms is unclear and uneven. Google has publicly said it doesn't use llms.txt for Search. Treat it as low cost and uncertain benefit.
- check_circleThe file is most useful for documentation-heavy sites, developer tools and knowledge bases, where a clean map of canonical pages genuinely helps.
- check_circleWriting one takes about an hour. Keep it short, curated and accurate. A dump of every URL defeats the point.
- check_circlellms.txt won't fix weak content or missing third-party mentions. It's a tidy extra, not a GEO strategy.
INSIDE THIS GUIDE
8 chapters. Jump to any of them.
CHAPTER 01
What llms.txt Actually Is
llms.txt is a plain text file, written in Markdown, that lives at the root of your domain, like example.com/llms.txt. Its purpose is to give large language models a concise, human-curated overview of your site: who you are, and links to the pages that matter most, each with a short description.
The idea was proposed publicly in 2024 by Jeremy Howard of Answer.AI. The reasoning is sensible: web pages are full of navigation, scripts and layout noise, and context windows are limited. A clean summary file could help a model find the right content faster.
The basic structure
- An H1 with the site or project name.
- A short blockquote summary of what the site is.
- Optional paragraphs with key context.
- H2 sections with lists of links, each link followed by a one-line description.
- An optional section for secondary resources that can be skipped when context is tight.
A curated map, not a crawl list
llms.txt is closer to a table of contents written by an editor than to a sitemap generated by a plugin. Its value comes from what you choose to include and how clearly you describe it.
CHAPTER 02
llms.txt vs robots.txt vs sitemap.xml
People confuse these constantly, so let me separate them cleanly.
- robots.txt tells crawlers what they may or may not fetch. It's an access instruction. If you want to allow or block AI crawlers, this is the file that matters, along with any platform-specific controls.
- sitemap.xml lists URLs you want search engines to discover, usually generated automatically and often containing every indexable page.
- llms.txt offers a curated, readable summary for language models. It doesn't grant or deny access, and it isn't meant to list everything.
warningWATCH OUT
Don't put crawl rules in llms.txt and expect them to be followed. If you want to control AI crawler access, use robots.txt directives for the specific user agents, and check each platform's documentation.
lightbulbPRO TIP
Your crawl and index foundation still decides whether content can be found at all. The technical SEO play covers that layer.
CHAPTER 03
Who Actually Reads llms.txt
This is the chapter most llms.txt guides skip, because the honest answer is less exciting than the hype.
Many developer-focused companies publish llms.txt files, and some AI coding tools and agents can use them when pointed at documentation. That's real, practical usage.
The big AI search and assistant platforms are a different story. Adoption has been unclear and uneven, and Google has publicly said it doesn't use llms.txt for Search, with one of its search advocates comparing it to the old keywords meta tag. Platforms can change their approach, but you shouldn't plan around a benefit nobody has confirmed.
If a tactic takes an hour and might help, do it. If someone sells it to you as the key to AI visibility, walk away.Shmul
Low cost, uncertain benefit
That's the correct mental model. llms.txt is cheap to create and harmless when accurate. It is not a proven ranking or citation lever, so it shouldn't take time from work that is.
CHAPTER 04
Who Should Bother, and Who Can Skip It
Worth doing
- Developer tools and APIs. Documentation sites benefit most, because AI coding assistants and agents are often pointed at docs.
- Large knowledge bases. When you have hundreds of help articles, a curated map of the canonical ones reduces confusion.
- Sites with complex product lines. A clear summary of what each product is can reduce mix-ups.
- Anyone who already has solid fundamentals and wants to cover small bases.
Fine to skip for now
- Small local businesses with a handful of pages.
- Sites that still have crawlability, indexing or content quality problems. Fix those first.
- Teams with no capacity to keep the file accurate as the site changes.
lightbulbPRO TIP
If you run a documentation site, consider also offering clean Markdown versions of key pages. That's often more useful to agents than the index file alone.
CHAPTER 05
How to Write a Good llms.txt in an Hour
- 1Write the one-paragraph summary. Who you are, what you offer, who it's for. Plain language, no slogans.
- 2Pick your canonical pages. Core product or service pages, the most important guides, pricing, about, and key documentation. Usually 15 to 60 links, not 600.
- 3Group them. Use H2 sections like Products, Guides, Documentation, Company.
- 4Describe each link in one line. Say what the page answers, not marketing copy.
- 5Add a skippable section for secondary resources.
- 6Publish at the root as
/llms.txt, served as plain text. - 7Put a review date in your calendar. An outdated map is worse than none.
Example
# Acme Analytics
> Acme Analytics is web analytics software for small ecommerce stores, focused on privacy-friendly tracking and simple revenue reports.
## Product
- [Features](https://example.com/features/): What Acme tracks and the reports it includes.
- [Pricing](https://example.com/pricing/): Plans, limits and what each plan includes.
## Guides
- [Set up tracking](https://example.com/docs/setup/): Install the tracking script on common store platforms.
warningWATCH OUT
Keep descriptions factual and current. If your llms.txt says a feature exists that you removed last year, you've created a new source of wrong answers.
CHAPTER 06
The Mistakes I See Most
- Dumping the whole sitemap. Thousands of links make the file meaningless. Curate.
- Marketing language. "Industry-leading, best-in-class platform" tells a model nothing. Describe what the page actually contains.
- Linking to redirected or broken URLs. Point to final canonical URLs.
- Setting and forgetting. Products change. The file should too.
- Treating it as access control. It isn't. Use robots.txt for that.
- Expecting it to fix visibility. If you aren't cited today, the reason is almost never a missing text file.
Accuracy beats completeness
A short, accurate llms.txt is better than a long, stale one. When in doubt, cut.
CHAPTER 07
Controlling AI Crawler Access, Properly
Since llms.txt doesn't control access, it's worth being clear about what does. The practical control for crawling is still robots.txt, together with any platform-specific settings the AI companies document.
- Named user agents. AI companies publish the user agent names their crawlers use, and some separate crawlers for training from crawlers used to fetch pages for live answers.
- Robots.txt rules per agent. You can allow or disallow specific agents, rather than all bots at once.
- Google-specific controls. Google documents a separate token for controlling use of content in some of its AI products, distinct from Googlebot crawling for Search.
- Server-level blocking. Some sites block by IP or at the firewall, which is more forceful but needs maintenance.
Decide what you actually want
Blocking all AI crawlers may protect content from training, but it can also reduce your chances of being cited in AI answers that fetch live pages. Allowing everything maximizes visibility but gives up control. Many businesses allow retrieval for answers and make a separate decision about training.
warningWATCH OUT
Check each platform's current documentation before writing rules. User agent names and policies change, and an outdated rule can quietly block the crawler you wanted to allow.
Access first, maps second
If a crawler can't reach your pages, a perfect llms.txt won't matter. Get access rules right, then worry about the nice-to-haves.
CHAPTER 08
Where llms.txt Fits in a Real GEO Plan
Here's how I'd rank the work if your goal is showing up in AI answers.
- 1Crawlable, indexable, fast pages. If systems can't access your content, nothing else matters.
- 2Clear, specific, answer-first content on the topics you want to be cited for. See getting cited in ChatGPT.
- 3Consistent third-party mentions on the sources models lean on: reviews, comparisons, publications.
- 4Entity clarity across your site and profiles. See entity SEO.
- 5Measurement of citations over time. See measuring LLM citations.
- 6Nice-to-haves, including llms.txt.
llms.txt is the bow on the present. Make sure there's a present first.Shmul
Frequently asked
What is llms.txt?expand_more
Does Google use llms.txt?expand_more
Is llms.txt the same as robots.txt?expand_more
Should my site have an llms.txt file?expand_more
How many links should llms.txt include?expand_more
Will llms.txt get me cited in ChatGPT?expand_more
Want this done for you?
I help brands win on Google and get cited in AI search. Tell me about your project.
ABOUT THE AUTHOR

Shmulik Dorinbaum (Shmul)
SEO and GEO consultant with 20 years measuring search. He has trained more than 1,200 marketers and advised brands including Duty Free Israel, Isrotel and Wix. Shmul writes the Playbook to help teams win on Google and get cited in AI search.