The benchmark, in full and with its method

This is the whole benchmark, not a highlight reel: the reproducible method, the current numbers, and the one figure we retracted. A marketing number you cannot reproduce is debt, not an asset, so the method is written down and a third party can run it.

The current run (commit 957372c, N=5 per cell, four arms) measures the cost of orienting once in a repo. Orienting with the atlas took 62% fewer tokens and 58% fewer turns than plain exploration, at full adoption. The open risks (queued tasks, alerts, repeat regressions) scored 0.6 to 0.8 with the atlas and 0.0 in git, even disciplined git, because they are not in git.

It also publishes what did not work. The atlas_symbol_lookup trigger raised adoption but made its adopters spend more than the arm without the atlas, so it was retired. And on 2026-07-19 a 52x claim went out in a paid campaign and was retracted the same day: it divided an agent's consumed tokens by the size of a response, two different units. That section stays, because verifiable honesty is the point.