The evidence
A 38% cut in context, no measurable change in answer quality
And because your AI can retrieve what it already learned instead of re-deriving it,
it uses 38% less context in our tests. Ten questions about past work, asked four
ways by the same agent over the same corpus; forty controlled sessions, scored on
recall of facts fixed in advance.
Configuration
Fact recall
Context tokens
Change
grep the repono index
46/50
12,515,065
baseline
index availableagent free to choose; used it 4 times in 10
40/50
9,481,199
-24%
routed by question typerecommended
46/50 within noise of baseline
7,751,761
-38%
index first, alwayscheapest; small measured accuracy cost
43/50
6,652,191
-47%
The saving is fewer operations, not smaller answers: a lookup returns about three times
more text than a grep, but the routed configuration needed 136 operations where grep
needed 241, and every operation avoided is one the conversation stops carrying forward.
Routing is the default we recommend: forcing the index everywhere saves more, but in our
testing it also missed material that reading the source would have caught, and we would
rather spend the extra tokens. Left free to choose, the agent reached for the index in
only four of ten sessions, which is why the routing rule ships with the product.
The quality differences between configurations sit inside the measurement noise;
the token savings do not.
Read the full evaluation, including what we cannot claim.