Evidentiality framework for AI
Download .md

Evidence and limits · early findings

Does labelling AI claims stop hallucinations spreading? Early results

Who this page is for: anyone deciding how much weight to give this. Every number here is from a small test; treat it as a lead, not a rate.

This page labels its own claims.

(u) given to the AI(m) measured / checked(g) generated / guessed by the AI

We label our results (m) because we checked them against our own test logs. To you, they're our report, which makes them (u) until you check. The logs for the six-round test are published; others are available on request.

Held up

Didn't hold up, or not shown yet

Fair objections

"The AI grades its own homework."

True. The labels are self-applied, and we haven't measured how often they're right. (g)The next step we’d test is a checker that isn’t the writer: the application, or a second model, applies "checked" only to what it can verify, and never upgrades a label it receives.(/g)

"It labels guesses; it doesn't reduce them."

True, and that's the aim. The labels give the reader more information, not a verdict.

"Is it the labels, or just the careful instructions?"

We don't know yet. The instructions include rules like "a guess stays a guess" as well as the labels. Separating the two is the test we most want run.

Limits of these tests

Related research