Tellr · AI Search

llms.txt Explained: Does Your Site Need One?

llms.txt helps AI tools find your best pages—but most sites don’t need one. Learn when it matters, what it does, and how to implement it safely.

By Tellr Editorial TeamPublished 5 October 2026

Most sites only need an llms.txt file if AI assistants or coding agents regularly answer questions about their product, documentation or policies. The file is plain markdown at the root of a website (yoursite.com/llms.txt), and it gives large language models a short, curated map of your most important pages. If your buyers use ChatGPT, Perplexity, Claude or an IDE assistant to learn how your product works, the file is a cheap, low-risk way to give those systems clean context. If you run a five-page brochure site, it will change very little.

The proposal has spread through documentation teams and now appears on standards and reference sites too. As of October 2026, no major AI search engine has publicly confirmed that it uses llms.txt to rank or cite pages. The sections below cover what the file does and does not do, which site types benefit, and how to build, test and maintain one without creating new risks.

Key takeaways

  • llms.txt is a curated markdown index that helps language models find the right pages on a site, and it is not an access-control or ranking file.
  • Documentation sites, developer platforms and SaaS companies with complex products get the most value from llms.txt, while small brochure sites get very little.
  • A good llms.txt file links to clean markdown versions of pages, uses specific descriptions and stays short enough for a model to read in one pass.
  • Listing a URL in llms.txt makes it more discoverable, so gated, internal or staging content must never appear in the file.
  • llms.txt supports AI visibility work but does not replace the content that AI engines actually choose to cite.

What llms.txt is and why HTML is a poor format for LLMs

HTML pages built for browsers bury their content in markup that a model has to strip out. llms.txt is a proposed convention that gets around this by giving AI systems a short, human-written summary of a site and links to its most useful content.

When an agent fetches a typical page, most of what it receives is noise:

  • Navigation menus, footers, cookie banners and repeated calls to action.
  • JavaScript-rendered content that a simple fetcher may never see.
  • Markup and inline scripts that consume tokens in a limited context window.
  • No signal about which of, for example, 4,000 URLs is the canonical answer to a question.

A model working within a fixed context budget does better with a short index and clean text. The file tells the model what the site is, which pages matter and where to find text versions of them. It complements how models already choose sources, which we cover in how large language models decide what to cite.

How llms.txt compares with existing standards

File or standardMain audienceWhat it doesWhat it does not do
llms.txtLLMs and AI agents at inference timeCurated summary and prioritised links, usually to markdownBlock crawlers, grant permissions or guarantee citations
robots.txtCrawlers, including AI crawlers such as GPTBotAllows or disallows crawling by user agent and pathExplain content or rank pages by importance
sitemap.xmlSearch engine crawlersLists every indexable URL with optional last-modified datesSummarise pages or separate key content from long-tail pages
Structured data (schema.org)Search engines and parsersLabels entities on a page: product, FAQ, organisation, reviewProvide a site-level reading order
Markdown page versions (.md)LLMs and agentsClean text copy of a single page, often at the same URL plus .mdTell the model which pages to read first

A sitemap is exhaustive; llms.txt is selective. A sitemap with 20,000 URLs tells a model nothing about which ten pages explain your pricing model.

Does your site need llms.txt? A decision framework by site type

Your site needs llms.txt if people regularly ask AI tools questions your own pages answer best, and your site has enough depth that a model could pick the wrong page.

You probably need llms.txt if… you publish product or API documentation, developers use coding assistants with your SDK, your support content is large, or prospects ask AI tools to compare you with competitors.

You probably don't need it if… your site has fewer than about 20 pages, has no documentation, and AI assistants rarely come up in how customers research you.

Site typeVerdictWhat to include
Documentation and developer platformsStrong yesQuickstarts, API reference, SDK guides, versioned docs; consider llms-full.txt
B2B SaaS marketing sitesYesProduct overview, pricing model, integrations, security and trust pages, comparisons
Ecommerce support centres and marketplacesYes, for policy contentShipping, returns, warranty, seller rules; not every product URL
PublishersSelectiveEditorial standards, topic hubs, licensing contact; weigh against your content licensing strategy
Universities and schoolsUsefulAdmissions, tuition and aid, programme catalogue, academic calendar
Legal and healthcareUseful, with carePractice areas, service lines, locations, reviewed patient or client information
Agencies, portfolios and personal brandsOptionalAbout, services, case studies, contact
Small brochure sitesLow valueA five-line file is fine, but expect little effect

For example, a useful file for an ecommerce support centre can be this small:

# Northwind Outdoor Help Centre
> Policies and support for orders shipped from Northwind Outdoor (US and Canada).

## Policies, [Returns](https://help.northwind.com/returns.md): 30-day returns, exclusions, refund timing, [Shipping](https://help.northwind.com/shipping.md): carriers, delivery windows, international rules, [Warranty](https://help.northwind.com/warranty.md): coverage by product category and how to claim

What to expect

  • Agents that read the file get cleaner retrieval context.
  • Coding assistants pointed at your docs answer more accurately.
  • Support workflows give more consistent answers.

What not to expect

  • A ranking boost in search results.
  • Guaranteed citations in Google AI Overviews.
  • Any change in crawler behaviour.

Showing up in ChatGPT answers still depends on the factors in our guide to ChatGPT SEO and visibility.

What goes in an llms.txt file

The proposal sets out a fixed order for the file:

  1. An H1 with the project or company name (the only required element).
  2. A blockquote with a short summary of the facts a model needs to understand everything else.
  3. Optional paragraphs or lists with extra context, such as terminology or caveats.
  4. H2 sections, each containing a markdown list of links in the form [name](url): description.
  5. An optional section titled "Optional" for secondary links a model can skip when context is tight.

Because the structure is fixed, both people and models can read the file, and classical tools such as parsers and regex can process it reliably.

# Acme Cloud Security
> Acme is a cloud security platform for AWS, Azure and Google Cloud workloads. This file lists canonical sources for product facts, documentation and security posture.

Product names: "Acme Posture" (CSPM) and "Acme Runtime" (workload protection). Prices on this site are list prices.

## Product, [Platform overview](https://acme.com/platform.md): Capabilities, supported clouds, deployment model, [Acme vs alternatives](https://acme.com/compare.md): Feature comparison with other CNAPP vendors

## Docs, [Quickstart](https://docs.acme.com/quickstart.md): Connect an AWS account with a read-only IAM role, [API reference](https://docs.acme.com/api.md): REST endpoints, authentication, rate limits

## Trust, [Security and compliance](https://acme.com/trust.md): Certifications, data residency, subprocessors

## Optional, [Changelog](https://docs.acme.com/changelog.md): Release notes by month

What each part of the example does:

  • The blockquote states category, scope and purpose in two sentences, so a model can answer "what is Acme?" from the summary alone.
  • The context line defines product names that models often confuse.
  • Every link points to a .md version instead of the full HTML page.
  • Descriptions say what each page answers, beyond its title.
  • The changelog sits under "Optional" because it is long and rarely needed.

Markdown versions of linked pages

The proposal recommends serving each page's markdown at the same URL as the original, either with .md appended (page.html.md) or with the extension replaced (page.md). URLs without a file name append index.html.md or index.md. You do not need markdown mirrors of every page. Create them first for the pages listed in the file. For large documentation sets, a companion llms-full.txt with the full text of key docs is a common pattern.

Common llms.txt mistakes

  • Listing too many links. A 600-link file is a sitemap in markdown and defeats the purpose.
  • Vague labels such as "Resources" or "Learn more" with no description.
  • Linking to script-heavy HTML pages when markdown versions exist.
  • Duplicate sections listing the same URLs under different headings.
  • Letting the file grow too large. As a working guideline, keep the index well under about 10,000 tokens and move full content into llms-full.txt.

How to implement llms.txt on WordPress, Webflow, static sites and headless stacks

On any stack, the job is to serve a UTF-8 text file at /llms.txt and, ideally, publish markdown versions of the pages it links to. How you do that depends on the platform:

  1. WordPress: upload llms.txt to the web root through SFTP or your host's file manager, or use a plugin that serves custom root files. Generate markdown versions with a build script or export, and confirm caching plugins do not rewrite the file.
  2. Webflow: check whether your plan and hosting settings allow root file uploads. If not, serve /llms.txt through a reverse proxy or edge function, such as a Cloudflare Worker, that returns the file for that single path.
  3. Static sites (Next.js, Astro, Hugo, Eleventy): place the file in the public or static folder, or better, generate it at build time from your content collection so new pages appear automatically.
  4. Headless CMS (Contentful, Sanity, Strapi): add an "include in llms.txt" boolean and a short description field to your content model, then have the front-end build render the file and .md routes from those fields.
  5. Docs platforms: check your platform's settings for llms.txt generation before writing one by hand. If you host docs on a subdomain, place a separate file at docs.example.com/llms.txt.
  6. Custom stacks: add a route that returns text/plain or text/markdown with charset UTF-8, and a route that serves a .md rendition of any page at its URL plus .md.

Help agents discover the files

Standard link relations tell clients where the files live. rel="alternate" type="text/markdown" points to the markdown version of a page, and rel="describedby" points to the llms.txt file that covers it. Add them as HTML <link> elements, or as an HTTP response header such as Link: <https://example.com/llms.txt>; rel="describedby". The header form also works for non-HTML resources, including the markdown files themselves, and you can set it in web server or CDN configuration without editing any pages.

Root versus subpath: an llms.txt file describes every page under its path, so the root file covers the whole site and /docs/llms.txt covers everything in /docs/. Multiple files work well for distinct properties, such as a marketing site, a docs subdomain and a developer portal, as long as the root file links to the others.

How to test, maintain and govern the file

A working llms.txt loads correctly, links only to pages that return clean text, and lets a model answer your core questions accurately. Reference sites are adopting the file too. MITRE's D3FEND change log records the addition of llms.txt files alongside other site updates, which suggests the convention is maturing beyond developer tools.

Validation checklist

  • Run curl -I https://example.com/llms.txt and confirm a 200 status, a text content type and no redirect chain.
  • Confirm the file is not blocked by robots.txt rules, a login wall or a bot challenge from your WAF or CDN.
  • Run a link checker (for example lychee as a CI step) over every URL and fail the build on 404s.
  • Open each .md link and confirm it contains body text without navigation, scripts or cookie text.
  • Paste the file into ChatGPT or Claude, ask ten real buyer or support questions, and check each answer cites the right page.

Maintenance

  • Assign one owner, usually SEO for marketing sites and docs engineering for documentation.
  • Regenerate the file on every deploy where possible, and review the hand-written summary at least quarterly.
  • Update it whenever you rename products, restructure URLs or retire pages.
  • For versioned docs, point to the current version and list older versions under "Optional".

Security and governance

llms.txt advertises URLs publicly and blocks nothing. Never list staging environments, internal wikis, customer-specific pages, gated content URLs or anything with tokens in the query string. To keep content away from AI crawlers, use robots.txt rules for specific user agents, noindex where appropriate, and authentication for anything private.

The file also carries no licensing weight. Media coverage summarised in Wikipedia's Signpost reported arguments that training LLMs on copyrighted content is largely covered by fair use, so legal and content teams should set their AI access policy separately from llms.txt.

Where llms.txt fits in an enterprise AI visibility program

llms.txt is a technical hygiene step. Enterprise AI visibility depends far more on the comparison pages, reviews, answer-shaped articles and Reddit threads that Google, ChatGPT and Perplexity actually cite. Tellr runs that work as one governed program, with approvals and an audit trail, following the approach in our enterprise guide to generative engine optimization. It is a managed program built for marketing teams spending $10k+ a month at companies worth $500M+ or with 200+ employees, so smaller teams will usually get better value from self-serve tools. A senior team works on Tellr's own platform across four products:

  • Reddit: subreddit mapping, a daily thread radar and guideline-checked replies behind an approval gate.
  • Content: comparison pages, reviews and answer-shaped articles built to be quoted by AI engines, published to your CMS.
  • Answer visibility: weekly tracking of who Google and its AI Overviews cite for your category's queries.
  • Paid media: category ad intelligence and ready-to-run creative.

The bottom line on llms.txt

Documentation sites, developer platforms and complex SaaS products should publish an llms.txt file, while small brochure sites can safely skip it. Build it as a short, curated index, link to clean markdown, announce it with link relations, test it with real prompts and keep private URLs out. Then put most of your effort where AI answers are decided: the pages, reviews and threads that models choose to quote. Treated that way, llms.txt is a small, sensible addition to a serious AI visibility program. It cannot stand in for one.

FAQ

What is llms.txt and what does it actually do?

llms.txt is a plain markdown file placed at the root of a site that gives large language models a short, curated map of the most important pages. It helps models find cleaner context, but it does not block crawlers, grant permissions, improve rankings or guarantee citations.

Does every website need an llms.txt file?

No. llms.txt is most useful for documentation sites, developer platforms, support-heavy sites and complex SaaS products where AI tools may otherwise choose the wrong page. Small brochure sites with few pages and no documentation usually get little value from it.

What should you include in an llms.txt file?

A good llms.txt file includes an H1 with the site or product name, a short blockquote summary, optional context notes, and H2 sections with markdown links and specific descriptions. It should stay short, selective and focused on canonical pages that answer important questions.

Can llms.txt create security or privacy risks?

Yes. Because llms.txt is public, any URL listed in it becomes more discoverable. You should never include staging sites, internal pages, customer-specific content, gated URLs or links with tokens in query strings.