Contribute
Take it further
This is a working idea with early evidence, published so other people can test it, break it and build on it.
Report a result
Open an issue in the GitHub repository with the AI model and version, the instructions version, which test you ran, how many runs per version, what you counted, and the outputs if you can share them. Negative results are as useful as positive ones.
Open problems
- How often are the labels right? Nobody has measured it.
- A checker that isn't the writer. Can an application or second model apply or verify "checked"?
- Labels or instructions? Run the test with the instructions but without the labels.
- Other models. Almost all our runs used one model family.
- Rates, not examples. Enough runs of the five-agent test to report how often drift happens.
- Fixes to the instructions. The known issues on How it works.
Reuse
Text: CC BY 4.0. Scripts: MIT. Adapt the notation, rename the labels, build tools on it. Please credit the source and share what you learn.