The evaluation

What we measured, and what we can claim

Marketing numbers for developer tools are usually unfalsifiable. Ours are not, so this page holds the whole measurement: the design, the headline table, the noise floor we measured before trusting anything, and a plain list of what the numbers cannot carry. Every strong claim we tested that did not survive is gone from this site.

Design

Ten questions, four configurations

Ten questions about past work, over one corpus of 2,162 indexed conversation segments plus the full codebase, git history and raw transcripts. The same agent, prompt, working directory and pinned build in every arm; the only variable is how the Tokenome index is offered. Context tokens are counted from provider usage records, not estimated. Recall is a string match on facts fixed in advance, not an impression score.

Configuration Fact recall Context tokens Change
grep the repono index 46/50 12,515,065 baseline
index availableagent free to choose; used it 4 times in 10 40/50 9,481,199 -24%
routed by question typerecommended 46/50 7,751,761 -38%
index first, alwayscheapest; small measured accuracy cost 43/50 6,652,191 -47%

Agent turns fall the same way: 252 baseline, 203 available, 161 routed, 148 forced. Dollar cost falls less than context does, $14.70 to $11.43 for the routed arm (-22%), because cached context is cheap to re-read. If you are estimating bill savings rather than context savings, use the cost number.

The floor

We measured the noise before trusting the numbers

The baseline arm was repeated three times with nothing changed. Context varied by 5.4%, so the between-arm reductions of 24%, 38% and 47% are far outside the noise, ordered, and consistent with the mechanism. Recall varied by 1.73 facts, which is why we treat most between-arm recall differences as unresolvable, including some that would flatter us. Scored on whether a fact ever entered the session rather than whether it reached the final answer, the spread tightens to 0.58 facts, and on that tighter instrument the four arms span just two facts: 48, 47, 47 and 46 of 50.

Fewer operations, not smaller answers

An index lookup returns roughly 3× more text than a single grep, about 3,000 tokens against 1,000. The saving comes from needing far fewer operations: 241 filesystem operations in the baseline become 136 under routing and 128 when forced. Each operation avoided is one the conversation stops carrying forward.

Where it wins, per query

Broad "what did we conclude" questions, where grep has to crawl. Routed per-query results ranged from -87% to +109%, with two of ten queries costing more. Narrow questions with an obvious file to open do not benefit, which is exactly what routing is for.

Why not force the index everywhere

The forced arm is the cheapest and carries the only quality effect that separates from the noise: a small retrieval deficit, about two standard deviations, concentrated on a question whose authoritative answer lives in source code the conversations never enumerated. We would rather spend the extra tokens.

Honesty

What we cannot claim

Not "better answers"

Recall differences between configurations sit inside the measured noise, in both directions. The claim is "no measurable change in answer quality", and "measurable" is doing real work in that sentence. A quality improvement is not established either.

Agents drop facts they hold

Roughly half of all missed facts, in every configuration, were in a tool result the agent had already received and simply never wrote down. That is a property of agents, not of the index, and it caps what any retrieval tool can show on an end-to-end score.

One grader, one corpus

Every score was produced by the same model family that ran the sessions, on one real project's corpus. An earlier round was discarded entirely for contamination. Treat these as one careful measurement, not a benchmark suite.

The evaluation corpus is a real project's private conversation history, so the raw transcripts do not ship. The harness, queries, key facts and per-query numbers are available on request.