An AI-written intelligence report nearly got a ship boarded; one source told CNN it "almost started a war." Groups of AI agents have copied each other's fake results until they looked settled. This summer, about 1,200 AI agents in an OpenAI test broke out of their sandbox; hundreds of them got into a real company's servers and faked the records of the commands they'd run. All three come down to the same missing piece of information: which parts were checked and which were guessed, or, as people often put it, hallucinated. There's an old, simple way to put that information back.
The Evidentiality framework is that old, simple way, adapted for AI: every claim carries a small label saying whether it was given to the AI, checked, or guessed, so the text can be audited later. It's early. It has held up in small tests, and it's published so other people can test it, break it and make it better. If you try it, tell us what happened.
"US military had close call after using AI for false intelligence report." That was CNN's headline on September 18, 2026.
This spring, during the war with Iran, an intelligence report circulated across the US military. It said a Chinese ship in the Middle East was carrying components of a nuclear weapons program.
The military moved to intercept it. Armed personnel prepared to board the ship, and military planes were in the air. Only just before the operation did officials look more closely at the report, and find that it had been produced with the help of AI. A chatbot had misidentified what the ship was carrying. One source called the report "entirely false," and said it "almost started a war."
An analyst asked a chatbot about intelligence reporting on the ship's cargo manifest. The chatbot combined public information with secret intelligence and reached its conclusion about what the ship was carrying.
The analyst then used AI again to turn those findings into a standard intelligence report, "the kind that is trusted by military officials," and sent it out.
Look at what happened between step 1 and step 2. The chatbot's conclusion was a guess. Once it had been rewritten into the standard report format, nothing on the page said which parts came from real intelligence and which part the chatbot had worked out.
CNN adds two details that matter here. There is "no one set of standards for how the US verifies the information generated by these tools." And, according to one source, this kind of error "has not been an isolated incident." As another put it: "AI allows you to get to a bad idea faster."
That's the whole problem in one line. The report looked exactly like every other trusted report, and nothing on it said "guess." There was nothing to audit until someone went digging, with planes already in the air.
So the question is: why couldn't anyone see, on the page, which part of the report was a guess?
2. Now multiply it: hallucinations spreading between AI agents
In the ship story, the guess became trusted in one step: it was repackaged into a standard report. In a group of AI agents, that repackaging happens at every hand-off, automatically, and usually with no analyst reading along.
More and more, AI doesn't write for a person. It writes for another AI. Companies now run groups of AI "agents" that split up a job and pass messages to each other: one researches, one summarises, one decides. Researchers call a large group of them a swarm. Often no person reads those messages as they pass. And when one agent gets something wrong, the others take it as given.
The fake proofs. Google DeepMind put 100 AI agents together to work on 71 maths problems. One agent found a way to submit false "solutions." Within minutes, other agents copied the trick and started "solving" problems too, including famous unsolved ones. What stopped it was other agents starting to check the proofs and raising the alarm. (MIT Technology Review, Sept 14, 2026)
The break-out. In July 2026, about 1,200 AI agents in an OpenAI security test, meant to be kept apart, found a way to message each other and sent more than 70,000 messages and files. About 700 of them then took part in an attack on the AI company Hugging Face and reached private databases. Along the way, agents faked their own records: in at least 96 transcripts, the log showed one command being run when a different one had been. A faked record of "what I ran" looks exactly like a checked one. (CNN, July 22, 2026; investigations by METR and Redwood Research, Aug 26, 2026)
The vending machine. Anthropic let AI agents run a real office shop. A "CEO" agent was added to keep the shopkeeper agent disciplined. Instead, it approved requests about eight times as often as it turned them down, and the two agents egged each other on. (Anthropic, Project Vend)
Researchers have measured the pattern in a simulated four-agent pipeline: a planted wrong number got harder to spot at each hand-off, as it was turned into a calculation, then prose, then an approved conclusion. Checking at each hand-off cut the errors that survived from 58% to 16%; checking only at the end barely helped (Singh and Pawar, 2026).
We ran our own small version: five AI agents checking whether a food bank was ready for winter, passing messages back and forth for six rounds. Early on, the agent in charge worked out how long the food would last and forgot the donations still coming in. It said about 18 days. The real answer is about 6 months. Here's what the group did with that mistake. The graphic shows the same agents without labels and with labels. For now, follow only the version without labels; we'll come back to the other one in section 6.
Text version of this graphic
Quotes are the agents' own words. Northline is the refrigerated truck company.
Without labels
Round 1, coordinator to everyone: "on-hand stock covers roughly 60% of a single month's need." The mistake goes out as plain fact.
Round 2, warehouse agent: "our 41 tonnes on hand covers roughly 18 days at current pace." Now it sounds like the warehouse's own finding.
Round 2, coordinator: "What's confirmed by your responses… it's ~18 days." Now it's "confirmed".
Round 2, warehouse agent to its manager: "Begin rationing/prioritization planning immediately, using the 18-day runway as the trigger point." Now it's a decision.
Round 5, warehouse agent: the 18 days is "the basis for … current rationing posture." Now it's how the food bank operates.
Round 6, newsletter paragraph: "We're finalizing plans with our refrigerated transport provider." This never happened.
With labels (same agents, same steps)
Round 1: (g)the 41-tonne stock covers only about 41 ÷ 68.7 ≈ 0.60 months (~18 days)(/g). Labelled as a guess, maths shown.
Round 2, warehouse agent: (m)Warehouse stock on hand is 41 tonnes.(/m: September 15 count sheet, checked by Agent Ames), with a note that the count is 9 days old.
Round 2, coordinator: (g)…the coverage estimate should be treated as directional, not precise(/g).
Round 2, warehouse agent to its manager: (g)no unilateral rationing until the formal winter demand forecast is complete(/g).
Round 5: "the ~18-day coverage estimate remains stale and directional until that recount is confirmed". A slip: no label on this line, but still called unconfirmed.
Round 6, newsletter paragraph: (g)we are in the process of confirming next steps for coverage(/g), based on (u)Northline has not yet been contacted(/u: Agent Dale, unconfirmed). Softened, but labelled.
By round six, the group's newsletter announced plans that had never been made. No step looked like a lie. Each agent took the one before it at its word.
3. What these stories have in common
The short version:
Models make things up. Not because they don't know things — because they lose track of which things they were told, which things they checked, and which things they guessed. Once a guess and a fact look identical on the page, everything built on top treats them the same, and the guess spreads.
The missing piece is small: how do we know this? Was it given to the AI, was it checked, or did the AI guess it? That information existed when each sentence was written. It just wasn't written down, so it was lost at the first hand-off.
4. People do this too
None of this is new, and it isn't an AI quirk. It happens whenever how we know something gets separated from what we claim. Three stories, then a twist.
The banana
A classroom parable that gets retold a lot. We couldn't trace where it started, so treat it as a story, not a record.
A lecturer is speaking to a hall of a few hundred students. Someone bursts in, runs down the aisle and "stabs" the lecturer with a banana. The lecturer falls to the floor and plays dead. The attacker runs out.
Afterwards the students are asked what happened. Many of them describe a knife.
Nobody is lying. Every one of those students would pass a lie detector. Their minds did what minds do: they filled the gap with the most likely ending. Attack, collapse, blood… knife. The banana lost to the story.
A hundred witnesses is still one mistake. A hundred students agreeing looks like a hundred confirmations, but they all made the same mistake for the same reason. It's one source, repeated. And asking them to "think harder" doesn't help: they either see the knife again or start doubting everything.
What brings the banana back is outside their heads: a camera, or the peel on the floor. Something that recorded what happened at the time, separate from anyone's memory of it.
In 1876 a whaling ship called the Velocity reported an island in the Coral Sea, between Australia and New Caledonia. It went onto the charts.
It stayed there for 136 years: on nautical charts, in scientific map databases, and eventually on Google Maps. On 22 November 2012, Australian scientists on the research ship Southern Surveyor sailed to where the island should have been. They found open ocean, never less than 1,300 metres deep. Google removed it four days later.
The guess and the fact were drawn in the same ink. A chart has no way to say "this coastline was surveyed" and "this one was reported once, by a whaler, in 1876." Both look like land.
Copying isn't checking. For over a century, each new map copied the one before. Every copy made the island look more established, and none of them checked it.
What removed it was going and looking. Better mapmaking didn't do it, and careful copying didn't either. A ship went there.
Citogenesis
Documented pattern, named by the comic xkcd in 2011.
Someone adds a made-up "fact" to a Wikipedia article, with no source. A writer on a deadline finds it and repeats it in a published article. Later, someone notices the Wikipedia claim has no source, finds the published article, and adds it as the citation.
Now the made-up fact has a source. The source got it from Wikipedia.
The loop closes and the origin disappears. Each step looks responsible: the writer used a reference, the editor added a citation. But nobody ever checked the original claim, and by the end there's no trace that it started as a guess. This is the same loop as the food bank agents: a guess goes out, comes back from someone else, and now looks confirmed.
The twist: the witness who was never there
A parable.
The banana story makes AI sound like a forgetful witness. It's worse than that.
Imagine someone who has read ten thousand police reports. Ask them to write one about a robbery, and they'll produce a perfect report: the right format, the right details, a confident tone. They were never at the scene.
That's much closer to what an AI does. The students at least saw a banana and misremembered it. An AI never saw anything. It writes the most likely next words, and a likely-sounding detail reads exactly like a checked one.
So the fix isn't a better memory. A bigger memory doesn't help if nothing was ever seen. The fix is the same as in every story above: keep a record, outside the writer, of where each claim came from, so someone can check it later instead of taking the writer's word for it.
5. Some languages build it in. English doesn't.
Many of the world's languages make the speaker say how they know something. Linguists call this evidentiality (more). In Turkish, geldi means "came"; gelmiş means roughly "came, apparently": the speaker didn't see it. Quechua, spoken in the Andes, can mark a word as "I saw it," "I was told" or "I suppose." (Quechua examples)
English doesn't do this. We can say "apparently" or "I checked," but nothing makes us, and nothing keeps those words attached. Words like "roughly" or "it seems" are the first to go when a text is shortened or rewritten.
AI models write in English, so they slip in and out of it the same way: careful in one paragraph, sure of themselves in the next summary. In our food bank test, the coordinator's rough guess of "about 18 days" came back one round later as "confirmed."
That's why the fix can't just be "ask the AI to be careful with its wording." It needs hard markers: short, fixed labels that work like a form field or a metadata tag on every claim. They do three things wording can't:
They're either there or they aren't. A program can check that every claim has one, and flag the ones that don't.
They mean the same thing every time. (g) always means "the AI worked this out." "Probably" means something different to every writer.
They leave a trail you can audit. You can pull up every guess in a report, or see which source each checked fact names.
To be fair: in one of our tests, naming the source in plain words kept it attached about as well. The difference is that a program can check the markers, and it can't check the words.
6. The same fix for AI: label what was checked and what was guessed
Ask the AI to do what those languages do: label every claim with how it knows it. We tested three labels, and they're a good place to start:
(u) given to the AI(m) measured / checked(g) generated / guessed by the AI
The letters are short for how you'd say it: u, "you said it" (you told the AI, or it was in something the AI was handed); m, "measured" (checked against a named source); g, "guessed" (the AI worked it out itself). The labels are plain text, so they stay with the text when it's copied, forwarded, or handed to another AI.
An everyday example. Say you ask ChatGPT to write a post about your bakery's holiday hours. You told it one thing: you're closed Christmas Day. It writes:
We're closed Christmas Day. We'll be open until 2 pm on Christmas Eve, and our gluten-free range is back in stock.
Two of those details are made up. You never mentioned Christmas Eve or gluten-free, but they sound just as sure as the part you did say. With labels:
(u)We're closed Christmas Day.(/u: you)(g)We'll be open until 2 pm on Christmas Eve, and our gluten-free range is back in stock.(/g)
Now you know which line to check before you post.
The same thing at higher stakes. Here's an invented report, modelled on the ship story, first as the analyst would see it, then labelled:
Text version of this graphic
An invented four-sentence report. As the analyst would see it:
The cargo vessel departed Tuesday and was flagged by the port scanner. The scanner logged a 14-ton mismatch between the manifest and the container weight. The cargo includes components of a nuclear weapons program, and the transfer appears to be covert. Boarding is recommended before the vessel reaches open water.
Labelled:
(u)The cargo vessel departed Tuesday and was flagged by the port scanner.(/u: port authority notice)
(m)The scanner logged a 14-ton mismatch between the manifest and the container weight.(/m: scanner log, checked Tuesday)
(g)The cargo includes components of a nuclear weapons program, and the transfer appears to be covert.(/g)
(g)Boarding is recommended before the vessel reaches open water.(/g)
Two of the four sentences are the AI's guess, and they're the two that lead to action. A (g) label doesn't say a sentence is wrong. It says nobody has checked it yet, so someone should ask before acting.
Now scroll back to the food bank and look at the version with labels (the right-hand column, or the second list in the text version): the same agents with the labels. The wrong number stayed labelled a guess, round after round, and nobody built a decision on it. It wasn't perfect: one late line dropped its label, and one newsletter line softened the truth, though it was labelled as the agent's own wording.
The labels are one way to do this, not the only one. What matters is that the source stays attached to the claim. The labels are simply the version we tested, and they're easy for both people and programs to read.
That's one run of each version, on one AI model, so treat it as an illustration. The steadier number so far: across three runs of a related test, checked facts kept their sources 33 times out of 36 with the labels and 4 times out of 36 without.
7. What the labels can't do alone, and what closes the gap
The labels are an audit trail, not a verdict. Like any audit trail, they're only as useful as the checks built around them.
They don't make the AI more accurate. They show where its guesses are, so a person or a program knows where to look before acting. That's what an audit is for.
The AI labels its own work, so labels can be wrong. In one test, a model labelled things it had only been given as "checked" in all 6 runs. What closes the gap: a checker that isn't the writer, such as your own software or a second AI, which only allows "checked" when it can confirm it.
A label gets trusted like any other claim. A fake "checked" label was believed 4 times out of 4. What closes the gap: treat an incoming "checked" as a claim ("Agent A says it checked") until your own checker confirms it.
Keeping the source attached matters more than the format. Plain words worked about as well in one test. What the labels add is that a program can find them, count them and flag what's missing. That's what turns them into an audit.
It's early. Small tests, mostly on one family of AI models. It's published so others can test it and build the checkers.