The backstory

Your code says what changed. Tokenome says why.

The reasoning behind your code is trapped in chat history: siloed by tool, disconnected from the artifact, invisible to your team and your agents. Tokenome links every line to the conversation that produced it, on your machine, so your team and your agents stop paying to re-derive what they already learned.

The problem

The reasoning behind your code is trapped in chat history

It all still exists, somewhere. It is just not where the work happens: not linked to the code, not visible to the team, not available to an agent. Three moments where that turns directly into hours.

01

You solve it twice

A tradeoff settled with an AI in June gets re-litigated in September because nobody can find the thread. You pay for the same thinking twice: once in engineer hours, and again in tokens when the AI re-derives an answer it already gave you.

02

The "why" is lost

AI now writes much of the code, and git blame shows a commit message, never the conversation where the decision was made. New hires spend ramp time reconstructing context that was written down at the time and then thrown away.

03

Your agents start from zero

An agent re-derives, at full price, a conclusion another session already reached, because nothing connects it to what it already learned. And "why is this written this way?" still goes to Slack and waits for whoever remembers.

Signature feature

git blame for the "why"

Point at any line of code. Tokenome blames it to learn when it was last edited, then surfaces the conversations from exactly that change. Permanent, searchable, and yours. Try the interactive example on the home page.

  • Answer "why is this function written this way" in one click
  • Read the conversation behind a change under review
  • Recover the tradeoffs your commit message never kept

How it works

Capture to recall, on your machine

Five stages, all local. Claude Code, Claude Desktop, ChatGPT and Gemini today, with a pluggable adapter for anything else. Only new or changed turns are re-embedded, so a live session costs about one turn of work per poll.

  1. CaptureAuto-discover and poll your AI tools
  2. SegmentSplit threads into question and answer documents
  3. EmbedOn-device ONNX vectors, no GPU, no cloud
  4. IndexTypesense: BM25 fused with vector similarity
  5. ServeCLI, web UI, and MCP for AI agents
Tokenome's architecture in one drawing. A boundary marked Your machine encloses the AI coding tools, capture and segmentation, on-device embedding, the local Typesense index, code provenance, and every way you and your agent search it. One gate crosses the boundary, marked opted in and secrets redacted, and it leads to a team server you run.
Everything inside the line runs on your own machine and is free. The gate is the only way out, and only a project you opt in goes through it.
The same diagram in words

Your AI coding tools, Claude Code, Claude Desktop, ChatGPT and Gemini, write their transcripts to disk as you work. Tokenome reads them where they land.

On the same machine it splits the threads into turns, computes the embeddings on your own hardware, and indexes everything locally in Typesense. Code provenance links a line of code back to the conversation that produced it.

You reach the index through the app, the command line and a local web UI, and an agent reaches it over MCP. All of that is free, and none of it leaves the machine.

One gate crosses the boundary. When you opt a project in, its segments are redacted on the laptop and then mirrored to a team server your organization runs, where a teammate's search finds them with your name on them. Tokenome does not host that server.

For developers

Command line

Fast search from your terminal. Filter by platform, model, project or date. JSON output for scripting.

tokenome search "type hints" --since 30d

For everyone

Web UI

Search, journal and FAQ, the code-context finder, and a dashboard that shows the tokens the index saved your agents. A local dark-mode app.

tokenome app → localhost:8741

For AI agents

MCP server

Your AI queries its own past. Claude Desktop and Claude Code call ask_faq, why_was and search_memory directly.

tokenome mcp → stdio
Hybrid searchTool-action captureDaily journal + running FAQ Incremental ingestPython + uvTypesense fastembed / ONNXMCP500+ tests

The agent's view

And the agent can ask too

Through the MCP server, Claude queries the same memory while it works. Here it is talking a developer out of reintroducing a bug that was settled three months earlier.

A real claude -p session against the Tokenome MCP server. Eight turns, 27.5 seconds, no network calls out of the machine.

Provenance for decisions

A day of decisions, with receipts

The same discipline as the code link, applied to decisions: once a day, Tokenome reads the conversations you had and writes down what you decided and why. Every reason quotes the sentence it came from and names the turn. If a claim cannot be quoted, it is not written down. Ephemeral AI interactions become durable organizational memory, automatically.

  • Rationale you never wrote down, kept without writing it down
  • Options you rejected, recorded as rejected rather than forgotten
  • A running FAQ where a reversal stays visible instead of overwriting the answer
  • Plain markdown on your disk. Grep it, commit it, read it anywhere

In the free tier. On a team, each machine writes its own; folding them into one team FAQ is coming after the beta.

A Tokenome journal entry for 14 June 2026, showing three decisions each with the
                verbatim conversation quote and citation behind it.
Sample data. The renderer, the citations and the format are the real ones.

The evidence

A 38% cut in context, no measurable change in answer quality

And because your AI can retrieve what it already learned instead of re-deriving it, it uses 38% less context in our tests. Ten questions about past work, asked four ways by the same agent over the same corpus; forty controlled sessions, scored on recall of facts fixed in advance.

Configuration Fact recall Context tokens Change
grep the repono index 46/50 12,515,065 baseline
index availableagent free to choose; used it 4 times in 10 40/50 9,481,199 -24%
routed by question typerecommended 46/50 within noise of baseline 7,751,761 -38%
index first, alwayscheapest; small measured accuracy cost 43/50 6,652,191 -47%

The saving is fewer operations, not smaller answers: a lookup returns about three times more text than a grep, but the routed configuration needed 136 operations where grep needed 241, and every operation avoided is one the conversation stops carrying forward. Routing is the default we recommend: forcing the index everywhere saves more, but in our testing it also missed material that reading the source would have caught, and we would rather spend the extra tokens. Left free to choose, the agent reached for the index in only four of ten sessions, which is why the routing rule ships with the product.

The quality differences between configurations sit inside the measurement noise; the token savings do not. Read the full evaluation, including what we cannot claim.