How I wrote llms.txt for 200+ sites without losing my mind
Once you are governing AI signage across hundreds of sites, it stops being a writing problem and becomes a federation problem. You need one specification, one delivery mechanism, and one rule for how sites link to each other. Here is the architecture I used at an academic medical center with more than 200 public web properties, including the parts that broke.
Why is 200 sites different from one site?
Writing a good llms.txt for a single site takes an afternoon. You know what the site is, you know what it is authoritative for, you write six links and a scope statement and you are done.
Now do it 200 times, for sites you do not own, run by people who do not report to you, on a schedule nobody controls. That is the actual job, and once you have more than a handful of properties, three things go wrong at once. Sites drift away from the format. They claim authority over things they do not own. And each one links to whatever its owner happened to think of that week, which produces a web of cross-references that sends AI agents in circles.
The architecture where I work looks like this. Several CMS Multisite installs. One of them uses a subdomain model, where each subsite is its own hostname. The pillar Multisites use a subdirectory model, so medicine.uams.edu, cancer.uams.edu, and uamshealth.com each carry dozens of department subdirectories underneath them.
In a subdirectory model, each subdirectory is an independent authority. A department page is a separate publisher that happens to share a hostname with the pillar site, and if you treat it as a chapter of that site, your scope statements will be wrong everywhere.
Why do static files in the web root lose?
Because they win, and that is the problem. The first instinct is to drop a robots.txt and an llms.txt in the web root and move on. On a Multisite, that file is now shared by every subsite on that install, which means every subsite is describing itself using another subsite's identity.
Worse, if you later build a smart per-site delivery mechanism, that stray static file will beat it. Nginx serves the file off disk before your rewrite rules ever fire. Your plugin is running, your routes are registered, and none of it matters, because the static file answered first.
We serve the triad through virtual routes instead. A network-activated plugin, "Per-Site robots.txt, llms.txt, llms-full.txt," registers the three paths and serves content stored per subsite in the CMS options store. Nothing lives on disk, so there is nothing to forget to delete when a site moves.
Two details took longer to get right than the plugin itself:
- The response has to send
Content-Type: text/plain. Serve it as HTML and some agents ignore it entirely. - The canonical redirect will try to add a trailing slash to
/llms.txt. It has to be suppressed on those three routes, or you are serving a 301 to a URL that does not exist.
What has to be in every file, every time?
One specification, enforced identically everywhere, because uniformity makes drift visible. If every file has the same seven sections, a missing section is a bug you can see from a list.
| Section | What it does | The failure mode it prevents |
|---|---|---|
| Single H1 display name | Names the publisher, exactly once | Agents guessing the entity name from the URL |
| Blockquote summary | One paragraph on what this site is | Summaries assembled from whatever page ranked |
| Metadata (Version, Last Updated, Status) | Makes staleness legible | Agents citing a file nobody has touched in two years |
| Targeted Content (3 to 6 links) | The pages you want cited | A crawler picking your six least useful pages |
| Related Resources | Where to go for adjacent needs | Dead ends |
| Authority and Scope | Positive and negative statements | Misrouting, which is the expensive one |
| Full Index link | Points to llms-full.txt | Agents assuming the short file is the whole map |
Three to six targeted links, not thirty. The short file is the lobby directory; llms-full.txt is the building.
The section that does the real work is Authority and Scope, and only if you write both halves. Everyone writes "we are authoritative for X." Almost nobody writes "we are not authoritative for Y, go here instead." The negative statement is what stops a research center from being cited as the place to book an appointment.
llms-full.txt is where the volume goes. For uamshealth.com we went from a hand-curated index of 15 URLs to a structured directory of 1,441 entries. Nobody typed that. We decided the structure first and generated the entries from the CMS, and if you are hand-typing your full index, you have already lost.
You can look at all of this. The llms.txt for uamshealth.com, the llms.txt for cancer.uams.edu, and the llms.txt for communications.uams.edu are public, and the AI wayfinding governance documentation lives on the communications site.
What about robots.txt, the file everyone already has?
It is the third leg and the one people treat as solved because it has existed since 1994.
In a federation it is not solved. A per-site robots.txt is where you decide, site by site, which AI user agents get in. That decision is not uniform, and it should not be. A public patient-information site and an internal-facing administrative site have different answers, and a single network-wide file cannot express both.
When a new agent shows up, and one shows up every few months, you need to add it in one place and have it propagate to every property. If your answer to "add this user agent everywhere" involves opening files by hand, you do not have a policy. You have two hundred opinions. Propagation is the whole argument for dynamic delivery.
How should sites link to each other?
This is the rule I would put on a slide if I only got one.
The pillar root is the hub. Departments link up, never sideways.
Every department triad declares exactly three kinds of links:
- An up-link to its pillar root. Always.
- A clinical destination link to the patient-facing site, if the department has patient-facing operations.
- One or two topic-specific cross-pillar links, where a genuine subject relationship exists across pillars.
And one prohibition: peer-department cross-links inside the same pillar are forbidden. If cardiology and radiology need to be connected, that connection lives at the pillar root, not in either department's file.
I did not start here. I started by letting departments link to whoever they collaborated with, which is generous and human and produced a graph with no center. An agent entering at any department could wander laterally for six hops without ever meeting the authority that could answer the question.
Hub and spoke is worse for expressing how the organization feels and much better for how it routes. Reciprocity belongs at the root. Departments do not owe each other links.
What actually broke?
In rough order of how much time each one cost me.
- A leftover static
robots.txtin a web root, beating the plugin on one install for about a week and a half. Everything looked fine in the admin. The internet saw something else. - Trailing-slash redirects on
/llms.txt, turning a working file into a 301 chain. - Content-Type served as HTML, which some agents will parse and some will skip.
- Scope statements copied between departments, so three units all claimed authority over the same program. Copy and paste is how a federation rots.
- Version and Last Updated fields going stale. Nothing technical broke there; nobody owned the review date, so the file became a lie.
What would I do differently on day one?
Write the specification before writing a single file. I wrote about five of them first and then discovered the pattern, which meant rewriting five files and, worse, defending inconsistencies I had personally created.
Decide the cross-link rule before anyone asks for an exception. Once one department has a sideways link, every department has a reason.
Build the delivery mechanism before the content. I had good content sitting in the wrong place for longer than I want to admit.
And put a review date in the file itself, where anyone can see it. If the date is not visible, the review does not happen.
In short
- At scale, AI signage is a federation problem: one specification, one delivery mechanism, one cross-link rule.
- Serve the triad dynamically per site. A static file in the web root will beat your rewrite rules and you will not notice.
- Enforce the same seven sections everywhere so that drift becomes visible.
- Use hub-and-spoke cross-linking. Departments link up to their pillar root and never sideways to peers.
- Write both halves of the authority statement. The negative half is what prevents misrouting.
If you are doing this for one site or ten, Seedfile will generate the triad and keep the structure consistent so you are editing content instead of arguing about format. If you are doing it for two hundred, start with the specification and steal mine. I have it written up as a three-page checklist with the seven required sections and the cross-link rule.
I manage web for an academic medical center. Everything I write about that work comes from what is already public. Opinions are mine.