August 19, 2026

What Is llms.txt, Should You Use It, and Do Google and OpenAI Actually Read It?

Adding llms.txt has become a checklist item almost everyone recommends, but do Google or OpenAI actually read it? A clear look at what the file does, Google's on the record confirmation that it plays no role in ranking for AI Overviews, and the Ahrefs data showing most published files go completely unread.

Somewhere in the last year and a half, adding an llms.txt file went from a niche suggestion buried in developer forums to a line nearly every AI search checklist insists on. The pitch sounds almost too easy: drop one small markdown file at your site's root, and AI systems will suddenly understand your business better. Once you actually look at what Google and OpenAI have said on the record, and what independent server log data shows about who requests the file today, that pitch holds up considerably less well than the checklists suggest.

This piece works through llms.txt without leaning on either the hype or the dismissal. It covers what the file is and how it works, whether Google and OpenAI actually use it in practice, Google's own confirmation that the file plays no role in ranking for AI Overviews, the specific evidence, Ahrefs' large scale study chief among it, showing that the overwhelming majority of published files sit completely unread, and what still makes building one worthwhile for the narrower set of sites it genuinely helps.

What Is llms.txt and How Does It Work? Should You Use It?

Start with the plain definition, since half the confusion around this file comes from skipping it. llms.txt is a short markdown document that a site publishes at its own root folder, addressable the same way robots.txt already is, meant to give an AI system a fast, human curated orientation to what the site covers and which pages are worth its time, rather than forcing it to wade through a full HTML page of menus and scripts to find that out on its own.

Credit for the idea goes to Jeremy Howard, whose work through Answer.AI and fast.ai led him to publish the proposal on September 3, 2024, hosting the specification at llmstxt.org. Some corners of the industry now call it the llms.txt standard, though that word is doing more work than it should, since no official body governs the format the way one does for HTTP or HTML, it remains a community maintained convention. The problem Howard was solving was concrete: a model only has so much context to spend, and burning a chunk of that budget parsing navigation and cookie banners just to find the two paragraphs that actually matter is wasteful in a way a short, deliberately written index is not.

Structurally, the file is simple by design. It opens with a single H1 naming the site, followed by a short blockquote summary, then a series of H2 sections grouping markdown links by topic, each with a one line description of what that page covers. A full breakdown of how the format is structured and how to build one well goes deeper into the editorial choices that separate a genuinely useful file from a hastily generated one.

Should you use it? That question is really two separate questions wearing one trench coat: does it cost anything meaningful to publish, and does anything meaningful read it once published. The first answer is close to no, a well curated file takes an afternoon and a few kilobytes of disk space. The second answer is where this article actually lives, and it is considerably less flattering than most llms.txt SEO advice lets on. The rest of this piece works through what Google and OpenAI have actually said and done, what the crawler logs actually show, and what, if anything, still makes it worth the afternoon.

Do Google and OpenAI Actually Use llms.txt Files?

The assumption behind most llms.txt advice is that the major AI platforms are quietly reading these files somewhere in the pipeline, even if nobody has confirmed exactly how. That assumption does not hold up equally well across every platform, and the two names attached to the largest share of AI search traffic today, Google and OpenAI, deserve separate answers rather than one shared shrug.

Start with OpenAI llms.txt questions, since this side of the story gets discussed far less than Google's despite ChatGPT being the platform most site owners actually care about. OpenAI has not published documentation stating that ChatGPT, GPTBot, or OAI-SearchBot parse llms.txt or treat it as a preferred input during crawling, indexing, or live browsing. Server level testing of what GPTBot and its sibling crawlers actually request shows them fetching standard HTML the way search crawlers always have, not sending the markdown preferring headers that would suggest the file is being sought out. Site owners occasionally spot GPTBot or a Microsoft associated bot hitting the llms.txt path in raw server logs, but a single fetch is not the same as a documented, repeatable behavior built into how ChatGPT answers a question. Fetching a file and using what it says are two different claims, and only the first one has any real evidence behind it for OpenAI's crawlers today.

Anthropic and Perplexity sit in a similar spot to OpenAI rather than a clearly different one. Some independent testing has observed occasional, inconsistent access to llms.txt files by tools associated with these platforms, and all three companies publish their own llms.txt files for their developer documentation, which understandably gets read as a soft endorsement. Publishing one for your own docs is not the same claim as confirming your assistant reads everyone else's. As of today, no major AI provider, OpenAI, Anthropic, or Perplexity included, has published anything stating its consumer facing assistant reads llms.txt at the moment it answers a question. Google is the exception worth walking through in full, mainly because the company has actually gone on the record about it directly.

Google Has Confirmed llms.txt Is Not a Ranking Signal for AI Overviews

Where OpenAI has simply stayed quiet, the paper trail on Google llms.txt questions is unusually direct, and walking through it in order matters, because the story has a real wrinkle near the end.

Start with the plainest data point: at a Search Central Live session in July 2025, Google's Gary Illyes told attendees, without much hedging, that the company has no intention of building llms.txt support into Search. John Mueller picked up the same thread separately and drew a pointed comparison, likening the file to the long retired meta keywords tag, a field where a site got to describe itself and search engines eventually learned to ignore, since self reporting like that gets gamed the moment it starts to matter. His characterization went further still, describing llms.txt as something closer to a workaround a handful of coding tools lean on to save a few tokens than a genuine search signal.

Then, in December 2025, an llms.txt file briefly turned up on one of Google's own developer documentation properties. For a few hours, part of the SEO community read it as a quiet reversal. It was gone by the end of the day. Mueller explained what had actually happened: an internal system used to publish that particular property had generated the file automatically as a default behavior, not because anyone on the search team had reconsidered. Google's own guidance for showing up in generative AI features went live on May 15, 2026 and has been revised since, most recently on July 10, 2026, which settles the question about as explicitly as a company ever does in writing. Inside a section aimed squarely at busting common myths, it draws the line plainly: neither Google Search nor Google AI Overviews and AI Mode need llms.txt to show your pages, since Google Search itself does not use the file. A later revision softened the delivery without changing the substance, adding that maintaining llms.txt for other platforms that do rely on it causes no harm either way, Google Search just does not read it.

Here is the wrinkle. Around the same time Google Search was closing that door, Google's own Chrome team was building an audit that checks whether a site has an llms.txt file at all. The Agentic Browsing category added to Lighthouse in May 2026, covered in full in our breakdown of that audit , includes a check for exactly this file, treating its presence as one marker of machine readiness for browsing agents. That is not really a contradiction so much as two teams solving two different problems. Search ranking is one question, and Google has answered it clearly: llms.txt does not move it. Whether a browsing agent can orient itself on a site is a separate question, and Chrome's tooling treats that as worth checking, even while Search treats it as beside the point for who shows up in a generated answer. Confusing the two is exactly how a Lighthouse checklist item ends up mistaken for a ranking factor.

Recommended Tool

See how AI actually understands your website.

Don't guess whether ChatGPT, Claude, Gemini, or Perplexity can access your content. Analyze your site in seconds.

Check AI Visibility Try Prompt Finder

No signup required • Instant results

Ahrefs' 97% Reality Check: Why Most llms.txt Files Are Currently Unread Decoration

Google's public statements explain what the company says it does. Server log data explains what is actually happening, and on that front the clearest picture so far comes from Ahrefs.

Published in June 2026, Ahrefs' analysis covered 137,210 domains with measurable traffic in May 2026. About 28 percent of them had published a valid llms.txt file, a figure worth treating as an upper bound rather than a web wide average, since Ahrefs' own customer base skews more technical and AI aware than the average site owner. Of the roughly 38,000 domains with a valid file, only about 1,100 received any request for it at all during the entire month. That means 97 percent sat completely untouched: no bots, no humans, nothing.

The remaining 3 percent does not rescue the story much either. Of the requests that did land, 96 percent came from bots rather than people, and named AI tools accounted for only around 19.5 percent of that already small bot share. The rest broke down roughly like this:

SEO audit tools checking whether a file exists, the single largest category at about 21 percent

Unidentified or unclassified bots, around 14 percent

General web crawlers such as Googlebot, around 13 percent

Technology profiling tools like BuiltWith, around 11 percent

Even within the AI bot share, GPTBot and Claude-Code topped the list of named fetchers, both of them tools tied to training and coding workflows rather than the live retrieval bots that actually generate a citation inside a chat answer. Ahrefs also found that AI bots essentially never probed for the file on domains where it did not exist, meaning crawlers are not quietly checking in the background, they are simply not looking. Even Chrome's own Lighthouse audit, mentioned above, accounted for roughly one in a thousand of the total fetches Ahrefs recorded.

This is not an isolated finding. SE Ranking's separate analysis of close to 300,000 domains found llms.txt adoption sitting at about 10 percent, spread fairly evenly across site sizes, with the largest, most authoritative domains actually slightly less likely to have one than mid sized sites. Running a statistical model and a machine learning classifier against citation frequency, SE Ranking found no meaningful correlation between having the file and getting cited, and removing it from the model as a variable improved the model's accuracy rather than hurting it, a fairly clean sign that the file was adding noise rather than signal. A separate review by ALLMO.ai of roughly 94,600 cited URLs pulled from real AI answers found the file present in essentially none of them. One independent operator who ran a 90 day test logged just 84 llms.txt requests out of more than 62,000 total AI bot visits, a shade over one tenth of one percent.

Put together, the honest read is that llms.txt today is mostly being checked by the industry studying itself, SEO tools, GEO tools, curious researchers, rather than by the retrieval systems that actually decide who gets cited in a generated answer. That does not make the file worthless. It makes the marketing around it considerably ahead of the evidence.

Creating llms.txt Anyway? Here's How to Do It Properly

None of this means skip it automatically. It means going in with the right expectations: low cost, unproven upside for AI search citations specifically, but a genuinely useful orientation layer for the narrower crowd that does treat it as useful today: coding assistants working inside a live repository, IDE plugins, and anything wired into the Model Context Protocol that needs to get its bearings on a codebase or a set of docs fast.

If that narrower audience matters to your site, particularly if you run developer documentation, an API, or a product with a technical buyer, a properly built file is worth the afternoon. Following llms.txt best practices does not require much more than five honest steps:

Shortlist 20 to 50 pages that genuinely define the site, not everything you have ever published.

Group them into 3 to 6 clear sections organized the way an outside visitor would look for them.

Write an honest, specific one or two sentence summary as the opening blockquote, since some tools lift this line directly.

Give each link a real, specific description rather than a generic one, and order sections by actual priority.

Save it as plain text named exactly llms.txt at the site's root, confirm it resolves without a login wall or redirect, and revisit it whenever the site changes meaningfully.

That covers the mechanics, and the full walkthrough linked near the top of this article goes deeper into the editorial judgment behind a genuinely useful file, including a complete example. Whatever time gets spent on llms.txt, spend more on the fundamentals the evidence above actually points toward. Clean heading hierarchy, genuine FAQ schema, and content structured the way AI systems actually parse it does more to earn a citation than a well curated llms.txt file ever will on its own, precisely because those elements sit inside the pages a crawler is already fetching by default, rather than a separate file most crawlers currently skip. Whether any of it, llms.txt included, is actually moving the needle is a measurement question rather than a matter of intuition. Tracking citation rate, mention rate, and share of voice across ChatGPT, Gemini, and Perplexity over time is what actually shows whether a change like this, or any structural change to a site, moved anything at all.

AI Visibility Live preview

Try it with your own website

Check whether AI engines can reach, read and recommend your pages.

Frequently Asked Questions

Does OpenAI's ChatGPT read llms.txt files?

Not based on anything OpenAI has put on the record. No documentation from the company states that ChatGPT, GPTBot, or its search indexing crawler treat llms.txt as a file worth prioritizing, and independently observed crawler behavior backs that up: these bots keep asking for ordinary HTML rather than requesting the lighter markdown format the file is written in. A stray server log entry showing one visit every so often does not add up to a built in habit, and that distinction is really the whole answer here.

Is it worth creating an llms.txt file in 2026?

If the goal is purely more AI search citations, the honest answer leans no, since every independent dataset covered above landed on the same flat result: sites with the file are not getting cited any more often than sites without one. The calculation changes for sites with real documentation, a public API, or a technical buyer base, since coding assistants and Model Context Protocol tools are the audience actually reading the file today. Given how little it costs to build one properly, most teams are better off treating it as low priority housekeeping rather than agonizing over the decision either way.

Why does Google's Chrome Lighthouse tool check for llms.txt if Search ignores it?

Because it is answering a narrower, different question than Search does. The Agentic Browsing check is not asking whether a page deserves to rank, it is asking whether an autonomous agent can find its footing on a site before it starts clicking around, and it treats the file's presence as one small piece of evidence toward that. A box ticked inside an experimental Lighthouse category is not the same thing as a ranking factor, even when the same company built both.

Does having an llms.txt file hurt my SEO or AI visibility?

No published evidence suggests it. The file is best understood as neutral rather than risky: it does not appear to improve citation rates, but nothing in the available research or Google's own statements suggests it penalizes a site either. The one practical caution worth keeping in mind is that a poorly maintained file, one pointing at retired pages or describing a product line that no longer exists, can quietly mislead whatever tool does bother to read it, which is more of a housekeeping risk than an SEO one.

What should I focus on instead of llms.txt to improve AI visibility?

Structure and citability inside the pages AI systems are already crawling by default. That means a clean heading hierarchy, genuine FAQ and Article schema, direct answers stated early rather than buried three paragraphs in, and content specific enough to be worth quoting. Those elements sit inside the HTML every crawler already requests, unlike a root file the data above shows most of them never bother to fetch, which is why the research consistently points to fundamentals over llms.txt as the better use of time.

Table of Contents

What Is llms.txt and How Does It Work? Should You Use It?

Do Google and OpenAI Actually Use llms.txt Files?

Google Has Confirmed llms.txt Is Not a Ranking Signal for AI Overviews

Ahrefs' 97% Reality Check: Why Most llms.txt Files Are Currently Unread Decoration

Creating llms.txt Anyway? Here's How to Do It Properly

Frequently Asked Questions

Summarize this article

Summarize with ChatGPT Summarize with Claude Summarize with Perplexity

About the Author

S

Satvik Mishra

Co Founder of Verseodin

Satvik Mishra is the Co Founder of Verseodin, an AI visibility platform that tracks brand citations across ChatGPT, Gemini, Claude, and Perplexity. He writes about generative engine optimization strategy and what actually works for brands trying to earn visibility in AI powered search.

Newsletter

Stay Updated

Get the latest AI Visibility insights, GEO research, product updates, and SEO strategies delivered straight to your inbox.

Share this article

Enjoyed this article?

Continue improving your AI visibility.

Prompt Finder

Discover what users are asking AI about your industry and uncover prompt opportunities.

Open Prompt Finder

AI Visibility Report

Analyze whether AI search engines can discover and cite your website.

Run Free Report

Related Articles

AI SEO 11 min read

What Signals Influence Brand Visibility in AI Search? How AI Systems Decide Which Brands to Surface

The five signals AI systems use to understand, validate, and retrieve brands, and how trust, authority, and data consistency decide which brands get surfaced in AI answers.

Read article

AI SEO 11 min read

AEO Tools for Tracking Product Mentions in ChatGPT: Measuring Visibility Across Your Product Portfolio

A brand can be mentioned constantly in ChatGPT while its actual products stay invisible. Here is how AEO tools track individual product mentions in ChatGPT, separate them from parent brand visibility, and measure coverage across an entire product portfolio, from mapping aliases to setting up scheduled tracking.

Read article

AI SEO 14 min read

Answer Engine Optimization Website Structure: H1, BLUF, and FAQ Strategy for 2026

H1 hierarchy, BLUF formatting, and FAQ strategy are the three structural levers deciding whether AI engines can extract and cite your content in 2026, plus a simple test for whether a section is actually citation ready.

Read article

Want more AI SEO insights?

Monthly research on AI search, GEO and citation trends. No noise.

[ partner program ]

Become a Verseodin Partner

Join our partner network, whether you're an agency, consultant, or reseller. Leave your email and we'll reach out with details.

Revenue share

Earn a cut of every referral you bring on, renewing month over month.

Co-marketing

Joint webinars, case studies, and content with the Verseodin team.

Priority support

Direct line to our team plus early access to new features.

What Is llms.txt, Should You Use It, and Do Google and OpenAI Actually Read It? | VerseOdin