I pulled my own coding-agent usage ledger. I expected output tokens to dominate. They don’t, and it isn’t close.
Across 3,823 requests to one model, 835,335,405 of 879,234,185 tokens were cache_read. That’s 95.0%. Call it 218,500 replayed tokens per request against roughly 11,500 new ones. Every heavy row in that table sits between 95% and 99%.
So the cost of a session tracks the length of the context. And the context gets resent every turn. What I let into it is the expensive decision.
Why this surprised me
An agent turn feels simple. Ask a question, get an answer. I was pricing the answer.
Here’s what actually happens. The conversation so far gets sent again. It becomes the prefix of the next request. Every file I read. Every command output. Every tool result.
I opened a 400-line file once to check one function. That isn’t a one-time cost. I pay for it on every turn after that. Twenty turns later I’ve paid twenty times. Nineteen of those I didn’t quote a word of it.
That changes how I use tools. Grep for a symbol and read four lines? Cheap forever. Read three files to find out that none of them mattered? Expensive forever, and the finding was “no.”
Two changes
First one is mechanical. My compaction threshold was effectively unset. On a million-token model that means it almost never fired. I set thresholdTokens to 160000. That models out at roughly 5× less cache and input per turn on long sessions. I’ve since taken it to 120000. Compaction costs me the summarised detail. It also removes that detail from every later turn, which is the point.
Second change is behavioral. It’s worth more. Investigation moves out of the main conversation.
Searching unfamiliar code, locating callsites, answering “where” or “how many”. The output is one sentence. The byproduct is a pile of file contents. I run that in a sub-session now. Its context dies when it returns. What comes back is the sentence.
Here’s the rule I settled on:
- Delegate anything where I’ll read a lot and quote a little.
- Keep anything that needs the conversation’s own history. Same for design calls, and for the edits themselves.
- Ask for the answer, not the transcript. A subagent that returns “auth is in
src/auth/jwt.ts:40-90, three callers” costs me twenty tokens. One that returns the files costs me the files, on every turn, forever.
What it doesn’t fix
Delegation moves tokens. It doesn’t delete them. The sub-session still pays to read what it reads.
A subagent pays that cost once. The parent would have paid it on every turn after. That only helps when the summary is much smaller than the material. If it isn’t, I shouldn’t have been reading it in the main loop either.
Two things to know if you go measure this yourself.
My ledger refreshes on demand, not continuously. Query it right after a session and the numbers are stale. Force a refresh first.
And the timestamps are epoch milliseconds. WHERE timestamp/1000 > strftime('%s','now','-7 days') returns nothing in SQLite. The right side is text. The left side is an integer. The comparison never matches. Cast it. An empty result set looks like “no usage”. That’s a bad place to start a decision.