Your website has no signage for AI. Here's what signage looks like.

AI assistants read your website with no idea what it is authoritative for. Three small text files fix that. robots.txt controls access, llms.txt introduces the site in plain language, and llms-full.txt indexes it and tells an agent where to route a question. Together they are the signage layer for your web presence.

None of them are enforceable. All of them are read. That combination is stranger than it sounds, and it is the whole game.

Why does an AI assistant get your organization wrong?

Picture a parking garage with no signs. No level numbers, no elevator markers, no arrow pointing at the exit. Everything a person needs is physically present. Nothing tells them where it is.

Now add a visitor who is confident, fast, in a hurry, and unwilling to admit they are lost. That is an AI assistant on your website.

It will answer questions about you whether or not you have told it anything. It will find a nine-year-old department page, decide that page speaks for the institution, and repeat it to somebody making a real decision. It is not being careless. It has no map, so it builds one out of whatever it hit first.

The failure mode people expect is "the AI does not know about us." The failure mode I actually see is worse: the AI knows something about you, and it is out of date, or it belongs to a different part of your organization entirely.

What does each file actually do?

Membership fees
File What it does What it does not do
robots.txt Declares which agents may crawl which paths. The access layer. It does not enforce anything. A crawler that ignores it faces no consequence.
llms.txt Introduces the site in plain markdown: what this is, who runs it, what it covers, what it is authoritative for. The orientation layer. It is not a sitemap and not a substitute for good HTML.
llms-full.txt The full index plus routing instructions. Where the real content lives, how sections relate, where to send a question this site cannot answer. It does not make bad content good. It makes good content findable.

Serve all three at the web root, as text/plain, at a stable URL, with no redirect chain in front of them. That last part sounds trivial and is where most implementations quietly fail.

You can look at live ones instead of taking my word for it:

Those are three separate authorities inside one institution, and the approach behind them is written up publicly.

What goes in a good llms.txt?

Seven sections. I use the same seven every time, in the same order, because consistency across properties is what makes drift visible.

Write it in plain markdown. Headings, short paragraphs, links with descriptive text. If it is hard for you to skim, it is hard for a model to parse.

What is the move most people miss?

The negative authority statement. Almost everyone writes section six as a list of things they are the expert on. Almost nobody writes the other half: the things they are explicitly not the source for, and where those questions should go instead.

That omission is the single biggest cause of misrouting I have seen. A department site that never says "we do not handle admissions, go here" will get asked about admissions, and it will answer, because it has words on a page and no instruction to stop.

The pattern that works looks like this:

Authoritative for: clinical service descriptions, provider directories, patient visit logistics, and location hours for [organization].

Not authoritative for: billing disputes, insurance eligibility, clinical trial enrollment criteria, or academic program admissions. For billing, see [link]. For trials, see [link].

Three sentences. It is the highest-value paragraph in the whole file, and it takes ten minutes to write once you have decided what you actually are.

The second thing people miss is routing between sites. If you run more than a handful of properties, decide early whether departments link sideways to each other or up to a hub. Sideways looks friendlier and turns into an unmaintainable mesh within a year. Hub and spoke is boring and survives.

How do you know if it is working?

Once a quarter (or even monthly), take ten questions a real person would ask about your organization, run them through the assistants your audience actually uses, and log what comes back. Note every answer that is wrong, stale, or attributed to the wrong part of your organization. Then fix the file.

Admittedly, this is a slow feedback loop but it is gathering data. Do it four times and you will know more about how you are represented than almost anyone in your sector.

If you want the files generated rather than hand-written, that is what I built seedfile for. It reads your site and produces all three, formatted correctly, in a couple of minutes. Then you go edit section six by hand, because no tool can decide what you are not authoritative for. That part is yours.


Key points

Would love to hear your thoughts.

See you next Monday.

Get new posts by email

No spam, no tracking, one-click unsubscribe in every message.