I asked five AI tools about my employer. Here's what they got wrong.

I ran the same set of questions past ChatGPT, Claude, Gemini, Perplexity, and Microsoft Copilot to see how they represent a large academic medical center. The errors were not random. They clustered into three types: stale attribution, wrong-site routing, and confident invention of services. Two of the three I could fix with files.

Why run this as a repeating audit instead of once out of curiosity?

Because a one-time look tells you about a model version, and a quarterly look tells you about your own content.

Everyone has typed their organization's name into ChatGPT once. That is a party trick. It produces a screenshot, a laugh or a wince, and nothing you can act on. The version changes next month and your screenshot is archaeology.

The value shows up on the second run. When the same question produces a different wrong answer in a different tool three months later, you learn which errors are the model and which errors are you. Model errors you report and wait out. Your errors are the ones your own files caused, and those you can fix in an afternoon.

We run it quarterly. Five tools in scope. Microsoft Copilot is on the list specifically because employees have it through Microsoft 365, which means it is answering internal questions with external content whether we audit it or not. That is the tool most organizations forget, and it is the one with the most captive audience.

What is the method?

Nothing clever. The discipline is in keeping it identical each time.

  1. Build a stable question bank. Twenty to forty questions you do not change between cycles. Mix patient-intent, referrer-intent, media-intent, and student or applicant-intent. Write them the way a real person types, not the way your comms team writes headlines.
  2. Run all five tools on the same day. Model behavior drifts. If you spread the run across three weeks, you cannot compare anything.
  3. Use a clean session for each question. No memory, no prior turns, no personalization. Log out or use a fresh window. Otherwise you are auditing your own history.
  4. Log the full answer, the citations, and a screenshot. Citations matter more than the prose. The prose tells you what it said. The citations tell you why.
  5. Classify every error by type before you interpret anything. Types first, stories later. Otherwise you will fix the memorable one and miss the pattern.
  6. Write down what you changed and re-ask next quarter. An audit with no delta is a hobby.

The output is a table. Question, tool, verdict, error type, cited source, proposed fix. Nothing prettier than that.

One note on citations, because it is the column people skim. Read what the tool cited before you read what it said. Three quarters of the time, a wrong answer is traceable to a real page of yours that was outdated, ambiguous, or written for a different audience than the one asking. The citation is the diagnosis. The prose is just the symptom.

The other quarter cites nothing, or cites a competitor, or cites a news article about you from 2019. Those are the ones to escalate rather than fix.

Error type one: what does misattribution look like?

Credit landing on the wrong unit for real work that actually happened.

The pattern is consistent. A program genuinely exists. The answer describes it accurately. And it attributes the program to a department or center that is adjacent to it rather than responsible for it.

Why it happens: the wrong unit published more, published more recently, or published in a format that was easier to parse. Authority in these systems is inferred from the shape of your content, not from your org chart. If your actual owner has one thin page and a neighbor has a well-structured landing page and three news stories, the neighbor wins.

Error type two: how does routing go wrong?

A patient question gets answered with a research page.

This is the one that actually costs something. Somebody asks how to be seen for a condition. The answer is competent, accurate, and points at a laboratory, a clinical trial listing, or a faculty profile. Everything in the answer is true. None of it gets the person an appointment.

Routing errors are almost always self-inflicted. They happen when your research content is better structured than your clinical content, or when a site never told the agent what it is not for.

This is exactly the failure the negative authority statement exists to prevent. A research site that says "this site is not authoritative for appointments or patient services, see [clinical site]" gets routed around instead of routed into.

Error type three: how confident is the invention?

Completely. That is the whole problem.

The third category is services that do not exist, described in fluent institutional prose, sometimes with a phone number that belongs to something else. There is no hedging. Nothing in the answer signals that this part was assembled rather than retrieved.

Where invention comes from, in my experience:

That third bullet is the uncomfortable one. A meaningful share of what looks like hallucination is just old pages nobody unpublished.

What did I change, and what could I not fix?

The honest summary: files fixed routing well, misattribution partially, and invention almost not at all. Signage tells an agent where to go. It does not stop an agent from making something up about a place it never visited.

The other thing files cannot fix is your content strategy. If the wrong unit outranks the right unit because the wrong unit publishes more, the answer is not a better llms.txt. The answer is that the right unit needs to publish.

How do you run this for your own organization?

Start with fifteen questions, not forty. Five tools, one day, clean sessions, a spreadsheet with six columns. Do it once, fix what is obviously yours, and put the next run on the calendar ninety days out.

Two things to decide before you start. First, who reads the report. If it goes to leadership as a list of things AI got wrong about you, it reads as an outrage document and nothing changes. If it goes to whoever owns the content, with a proposed fix per row, things change. Second, what counts as severe enough to report to the vendor. Most errors are yours to fix. A small number are genuinely the model's, and for those there is a feedback channel worth using. Draw that line in advance, in writing, or you will redraw it every quarter based on how annoyed you are.

The first cycle will feel embarrassing. That is the correct reaction and it passes. By the second cycle you are looking at a delta instead of a verdict, and a delta is something you can manage.


Key points

eiAEO watches this continuously so you are not rebuilding the spreadsheet every quarter, but the spreadsheet works and I want you to run it either way. I have the question bank we use, genericized, as a downloadable audit template: fifteen starter questions, the six-column log, and the three-type classification key.

I manage web for an academic medical center. Everything I write about that work comes from what is already public. Opinions are mine.

Get new posts by email

No spam, no tracking, one-click unsubscribe in every message.