PHII Labs
2026-12-21AI Search6 min read

llms.txt: what it does, doesn't do, and who needs it

llms.txt explained plainly: what the file is, which engines actually read it, what to put in it, and why it is a cheap signal rather than a ranking lever

Sergei Suvorin · Co-founder, PHII Labs

A markdown index file being read by an AI crawler

llms.txt is a plain-text markdown file at your domain root that hands an AI crawler a curated index of who you are, the facts that define you, and the pages worth reading. No major engine requires it and Google explicitly does not use it. On this site it took under an hour to write and lives at /llms.txt as documentation hygiene, not a ranking lever

Which engines actually read llms.txt?

None of the engines that answer questions at scale consume it today. Google states that AI Overviews and AI Mode need nothing beyond normal indexing, that "there are no additional technical requirements," and that no special file or markup is required (Google Search Central: AI Overviews and AI Mode). OpenAI, Perplexity and Anthropic publish retrieval materials that describe crawling, chunking and passage extraction, and none of them lists an llms.txt step. What actually happens is that a few tooling and research projects fetch the file where it exists, and some companies inside AI-adjacent products read it opportunistically.

So the honest position is: the file is a convenience for whoever chooses to read it, and nobody is required to. It does no harm, and it cannot move a page into an answer on its own. The retrieval pipeline that decides citations is the prose on your pages, not the index file at the root

What is the proposal behind llms.txt?

The idea came from Jeremy Howard, who drafted a spec in mid-2024 as a suggested standard (llmstxt.org). The pitch is to give models something like robots.txt but for meaning rather than access: instead of a set of blocking rules, a short markdown file that lists the organization's name, a one-line description, canonical facts, and links to the pages that matter. The format is deliberately simple so a model can parse it without a heavy pipeline. Our own file at /llms.txt follows that shape: a header with who we are, a canonical-facts block with the legal entity name, licence authority and address, then grouped links to services, projects and top articles, and a line about AI-crawler policy.

The gap between the pitch and the practice is real. The spec was written and promoted, several sites publish the file, and the ecosystem that consumes it is thin. Jamie Zawinski's warning from another context applies here fairly directly: the file solves a problem most people do not actually have, and it creates one (the appearance of a lever) that the evidence does not support

What should you put in an llms.txt file?

Keep it short and factual, and let it mirror the entity consistency your structured data already carries (structured data for AI search). This is the annotated file we serve:

# PHII Labs                     <- brand name, exactly as used everywhere

> PHII Labs is an AI automation agency serving businesses in Dubai
> and the wider UAE. It builds AI agents, internal systems, AI products,
> document automation and AI search visibility. 20+ systems shipped.
> Legal entity: PHII LABS (FZC), SRTIP Free Zone, Sharjah.

## Canonical facts
- Brand name: PHII Labs
- Legal name: PHII LABS (FZC)      <- matches schema.org Organization node
- Licence authority: SRTIP Free Zone
- Website: https://phiilabs.com

## Pages
- [Home](https://phiilabs.com/): overview, cases, process, audit
- [Services](https://phiilabs.com/services): four service areas
- [Projects](https://phiilabs.com/projects): shipped systems
- [Blog](https://phiilabs.com/blog): articles by topic

## AI crawler policy
robots.txt is permissive: GPTBot, ClaudeBot, PerplexityBot, Google-Extended allowed

Three lines matter and the rest is decoration. The brand name in the first line should match your footer, your Organization schema, and your social profiles, because a model that sees one name everywhere resolves you to one entity. The canonical facts should be the same legal name, address and licence that your structured data lists. The page links should point at pages that are actually good at being cited, not your whole sitemap

A line about crawler policy is optional and worth including only because it documents a decision. If you allow AI crawlers, say so. If you block them, that is a deliberate opt-out from the channel, and an llms.txt that claims to court models while robots.txt refuses them is contradictory

Does llms.txt help or hurt SEO?

It does neither. Crawlers fetch, parse and index your HTML and your structured data; a text file at the root is not part of that pipeline. The caution some agencies sell, that an unmaintained llms.txt "confuses" engines, has no mechanism behind it. Search engines treat it as an ordinary file they may ignore. Google's own guidance on AI features confirms no special file is needed, which is the stronger statement: the absence has no cost, and the file's presence has no ranking effect to reverse

The one place it could matter is brand hygiene rather than search. A model that happens to read the file gets a clean, unambiguously yours statement of identity, which reduces the chance of four name variants being mistaken for four companies. For a UAE business that trades under a marketing name, holds a RERA-registered legal name, and appears under different spellings on Property Finder and Bayut, that consistency is where the value sits: the same reasoning behind entity-first structured data, expressed in plain markdown

Where should you draw the line on believing the hype?

We put llms.txt in the cheap, unproven column, the same place we put no-code experiment claims and guaranteed-citation sales pitches. We publish one because it cost under an hour and a crawler that does read it benefits, and we report honestly that we have not measured a citation effect. Anyone who sells llms.txt as a "GEO lever" with a price tag is selling a file and a story, not a result

Note

Treat llms.txt as documentation, not optimization. What makes content citable is independent of this file: self-contained passages, sourced numbers, named quotes and one consistently spelled brand. The pattern is the same on every engine

In the citation logs we run across ChatGPT, Perplexity, Gemini and AI Overviews, we have never once seen an llms.txt entry cited as a source. The passages that get quoted come from prose and tables on real pages. We ship the file anyway because it is cheap and it forces the entity facts to stay consistent

Sergei Suvorin · Co-founder, PHII Labs

The productive time is the same either way: write passages an engine can quote, keep one brand name everywhere, publish tables nobody else has, and measure with a frozen prompt set three times per engine per run. The measurement method, including our per-citation log format, is in how to get cited by ChatGPT and Perplexity, and the layer-by-layer split between search, answer engines and generative citations sits in SEO vs AEO vs GEO. Entity consistency, the only real effect llms.txt touches, is the subject of the structured data guide

We have seen the entity-consistency payoff fail in a concrete way on a real site. On the Trusted Real Estate build, the same business was listed under three names across its portal profiles and footer, and organic visibility only started moving once every surface agreed on one canonical name. That fix was prose and schema, not an index file. Retrieval is stage one of every AI answer, and pages Google already trusts are the pages AI engines find first, which is why classic SEO is not obsolete

How to keep an llms.txt from going stale

A stale index is worse than none only in the sense that it wastes a reader's time, so the maintenance bar is low but real. Review the file on the same schedule you review your robots.txt and your Organization schema. When you launch a service page, add it to the Pages section. When the legal entity, licence authority or contact address changes, update the Canonical facts block the same day you update the footer. If you rename the brand, change every surface at once and list the old name nowhere new. Read the file top to bottom once a quarter and delete anything a model would be worse for reading. That is the whole routine; it takes minutes, and it keeps the file consistent with the structured data every engine already sees

Should a UAE business bother at all?

For a UAE service business the file is optional and the question is what else it would displace. A Dubai agency or brokerage that trades under one marketing name and a separate RERA-registered legal name gets more consistency value from fixing its portal profiles, footer and schema than from writing an index file, because the engines that decide citations read those surfaces. llms.txt is a byproduct of that consistency, not a replacement for it. The one case where it earns its place is a brand actively courting AI visibility: a plain file that states the canonical name, legal entity and key pages in one place gives any crawler or tool that does fetch it a clean starting point. Publish it, keep it true, and measure nothing from it directly

The takeaway

Publish llms.txt because it is a cheap, honest index of who you are, and stop there. It is not a ranking factor, no major engine requires it, and Google's guidance confirms nothing beyond normal indexing is needed. The work that moves citations is the passages themselves, the numbers behind them and the one brand name you keep everywhere. If your brand sits at four different spellings across the web today, an llms.txt is the least useful of the three fixes you could make to that gap

Book the free audit and we will show you where your brand stands across ChatGPT, Perplexity, Gemini and AI Overviews, with your prompt set, the logs and the first five changes worth making

FAQ

What is llms.txt?

A markdown file at your domain root listing canonical facts about your organization and links to your most important pages — a curated index for language models that fetch it

Do Google or ChatGPT actually read llms.txt?

No major engine requires or officially consumes it today. Some tools fetch it opportunistically. Publish it as cheap documentation hygiene, not as an optimization lever

What should go into llms.txt?

Your canonical company name and description, key facts, and links to services, top articles and contact — the same entity consistency your structured data carries

Can llms.txt hurt SEO?

No — it is a plain text file unrelated to crawling or indexing directives. It neither blocks nor promotes anything

Sources

Want systems like this?

We build and ship AI systems for real operations