Design
Ten questions, four configurations
Ten questions about past work, over one corpus of 2,162 indexed conversation segments
plus the full codebase, git history and raw transcripts. The same agent, prompt, working
directory and pinned build in every arm; the only variable is how the Tokenome index is
offered. Context tokens are counted from provider usage records, not estimated. Recall
is a string match on facts fixed in advance, not an impression score.
Configuration
Fact recall
Context tokens
Change
grep the repono index
46/50
12,515,065
baseline
index availableagent free to choose; used it 4 times in 10
40/50
9,481,199
-24%
routed by question typerecommended
46/50
7,751,761
-38%
index first, alwayscheapest; small measured accuracy cost
43/50
6,652,191
-47%
Agent turns fall the same way: 252 baseline, 203 available, 161 routed, 148 forced.
Dollar cost falls less than context does, $14.70 to $11.43 for the routed arm (-22%),
because cached context is cheap to re-read. If you are estimating bill savings rather
than context savings, use the cost number.