· 7 min read
Is llms.txt worth it in 2026? What Google, OpenAI, Anthropic and Perplexity document
What the llms.txt proposal is, what each company’s own documentation says about it, what Ahrefs’ crawl of 137,210 domains found, and what to do instead.
llms.txt is a proposed convention: a Markdown file at /llms.txt that summarises a website and links to clean, LLM-friendly versions of its pages, so an AI agent can find what it needs without parsing HTML. In 2026 it is worth publishing for the readers that actually use it, mostly coding agents reading software documentation, and not as a way into AI search: Google says its search ignores the file, the crawler documentation of OpenAI, Anthropic and Perplexity doesn't mention it, and Ahrefs found that 97% of the files published across 137,210 domains received no requests at all in May 2026.
I'm a senior software engineer in Mumbai, and technical SEO for search and AI visibility is part of my work, so this file comes up in it. This site publishes one. Below is what each company actually documents, what the best public data shows, and where I'd spend the time instead. Every claim about a company's behaviour links to that company's own page, read on 11 October 2026; these pages change, so check the dates.
What is llms.txt, exactly?
Jeremy Howard proposed it on 3 September 2024, and the proposal, now in a second version dated 10 August 2026, sets out the format. The file lives at the root of a site or at any path within it, covering the pages under that path. It contains, in order:
- an H1 with the name of the site or project, the only required part;
- a blockquote with a short summary;
- optional paragraphs or lists with more detail;
- optional sections under H2 headings, each a list of links, with an "Optional" section for links an agent can skip.
The proposal also suggests publishing a Markdown version of each useful page at the same URL with .md added, and, in version 2, pointing to both with standard link relations (rel="alternate" with type="text/markdown", and rel="describedby" for the llms.txt file). Its author writes that he expected the file to be used for inference rather than training, and that it is used most heavily for software documentation, where coding agents follow it to API references and tutorials.
What do the companies actually document?
| Company | What its documentation says | Source |
|---|---|---|
| Google Search | You don't need "machine readable files, AI text files, markup, or Markdown" to appear in Search, "as Google Search itself doesn't use them"; creating an llms.txt "will neither harm nor help" your visibility | Optimizing for generative AI features (updated 10 July 2026) |
| Google Chrome (Lighthouse) | An agentic-browsing audit checks the file: a server error when fetching it is flagged, a missing file (404) is marked "Not Applicable", as providing it "is optional at the moment" | Lighthouse: llms.txt (updated 5 May 2026) |
| OpenAI | Describes OAI-SearchBot (ChatGPT search), GPTBot (training) and ChatGPT-User (user-initiated visits), each managed through robots.txt; sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers". Nothing on llms.txt | OpenAI crawlers |
| Anthropic | Describes ClaudeBot (training), Claude-User (fetches for a user's question) and Claude-SearchBot (search quality), each respecting robots.txt. Nothing on llms.txt | Anthropic's crawlers |
| Perplexity | Describes PerplexityBot (search results) and Perplexity-User (fetches on request, which "generally ignores robots.txt rules"). Nothing on llms.txt | Perplexity crawlers |
There is one twist worth knowing. All four companies publish an llms.txt for their own developer documentation; I fetched developers.openai.com/llms.txt, platform.claude.com/llms.txt, docs.perplexity.ai/llms.txt and ai.google.dev/gemini-api/docs/llms.txt while writing this, and each answered with a file. That fits the proposal's own account of where it is useful: developers and their coding agents reading documentation. It says nothing about the companies' crawlers reading yours.
What does the traffic data show?
The largest public study I know of is Ahrefs' analysis of 137,210 domains, published on 15 June 2026 and based on the server logs and traffic of sites using Ahrefs Web Analytics. Its main findings:
- 28% of the domains publish an llms.txt; Ahrefs notes its customers skew technical, so treat that as an upper bound.
- Of the roughly 38,000 valid files, 97% received no requests in May 2026.
- In the 3% that were fetched, 96% of requests came from bots, and 19.5% from named AI tools. GPTBot fetched the most, and Claude Code, Anthropic's coding agent, came second, ahead of every AI search and assistant bot.
- OAI-SearchBot, PerplexityBot and Claude's search crawler together made only a couple of hundred fetches across thousands of sites.
- Not one AI bot requested an llms.txt that didn't exist; the requests for missing files came almost entirely from people, presumably checking competitors.
Ahrefs is careful about the limits: "fetched" doesn't mean "read", so every figure is a ceiling on real use. It also found a research crawler identifying itself as a prompt-injection survey, which is a reminder that a file agents are designed to trust is also a target.
So is it worth it?
It depends on who you expect to read it.
| Your site | My verdict |
|---|---|
| API, library or product documentation | Yes. Coding agents are the readers the data shows; publish llms.txt and Markdown versions of the pages, generated from the docs so they never drift |
| A docs platform generates one for you | Keep it, and review what it says |
| A company, consulting or marketing site hoping for AI search visibility | It won't help in Google, and nothing documented or measured says it helps ChatGPT or Perplexity search. Fine to keep if it costs nothing; don't spend a sprint on it |
| A site with sensitive or fast-changing content | Only with care: anything instruction-shaped in it is an invitation to the wrong reader |
What should you do instead?
The things the same companies do document:
- Let the search crawlers in. Allow OAI-SearchBot, Claude-SearchBot and PerplexityBot in robots.txt, and make sure your CDN or firewall isn't blocking them. Decide the training opt-outs (GPTBot, ClaudeBot and Google's Google-Extended token) separately; that is a content-licensing choice, not a visibility one.
- Be indexable and snippet-eligible. Google's AI features documentation says there are "no additional requirements" for AI Overviews or AI Mode: a page must be indexed and eligible to appear with a snippet. Server-rendered HTML, honest canonicals and a clean sitemap do more than any extra file.
- Write pages worth quoting. Google's guide asks for "non-commodity content" and a unique point of view, such as a first-hand account. Put the answer in the first sentence under a question-shaped heading, use real tables, and cite primary sources.
- Keep facts consistent. One name, one role and one description of what you do, the same on your site, your profiles and your structured data.
- If you publish llms.txt, treat it like code. Ahrefs' own advice: keep it in version control, restrict who can edit it, alert on changes, keep it to plain links and descriptions with nothing instruction-shaped, link only to resources you control, and review anything a platform generates for you.
What do I do on this site?
This site publishes /llms.txt and /llms-full.txt, both generated at build time from the same content as the pages, so they can't drift from what the site says, and its robots.txt allows OpenAI's, Anthropic's and Perplexity's crawlers by name, among others. That is the whole investment: a build step I wrote once. The time goes into the pages, the structured data and the sources, which are what the companies' own documentation points to.
How I do this
Technical SEO and AI visibility is the offer behind this: I audit how search engines and AI crawlers see your site (robots rules, rendering, canonicals, structured data, what answer engines can quote), fix it in your code and, last, add llms.txt if your readers are agents. Pages that load fast help every reader, so it often runs alongside a performance audit. The engineering glossary defines the terms used here.
Sources
- The /llms.txt file proposal (llmstxt.org)
- Google Search Central: optimizing for generative AI features
- Google Search Central: AI features and your website
- Chrome for Developers: the Lighthouse llms.txt audit
- OpenAI: overview of OpenAI crawlers
- Anthropic: how Anthropic crawls the web and how site owners can block it
- Perplexity: Perplexity crawlers
- Ahrefs: llms.txt study of 137K sites (June 2026)
Related service
- Service: Technical SEO and AI search visibility · Performance and reliability
Sites that search engines and AI assistants can crawl, understand and cite, with the checks built into every deploy.
- Service: Performance and Core Web Vitals audit · Performance and reliability
Find out in five working days exactly why your application is slow, and what to fix first.
Related case study
- Case study: TopGear India CMS · CMS migration and modernization
- page load time, before → after
- 1–3 min → 3 s
- Case study: InfluencerX · Real-time voting platform
- users over the campaign
- 1.6M+