seatledger guide · checked

Why did my Claude Code token use jump?

Five things move Claude Code's token use, and the transcripts on your machine record each of them for every request: how much context a request carries, whether the prompt cache was read or written and at which lifetime, how many requests one prompt turns into, which model answered, and which client version sent it. Find the day the rate stepped and what moved with it, and you can tell a change in your own work from one that arrived with a client update. npx seatledger rates does that from the transcripts already on your machine.

What each cause looks like in the transcripts

CauseWhat movesNote
Bigger requests: a larger repository, more tool, skill or MCP definitions, a longer sessiontokens per requestEvery request carries more context.
Cache lifetime: writes at 5 minutes instead of 1 hour, or idle past the lifetimecache-read share, cache writes per read, and the 1-hour share of cache writesEach response records its cache writes split by lifetime, so this one is observed, not inferred.
More requests per prompt: subagents, agent teams, scheduled tasks, goal check-instokens per user turnEach prompt turns into more requests.
A different modelthe model columnseatledger keeps each model's series apart, so a switch shows as a new series, not a step.
A client updatethe client version each request was sent byA finding says whether its first day was the first on a new version.

The method

For each client and model, seatledger takes every complete day with at least 20 requests and computes those metrics. A day is flagged when the median of it and up to six active days after it differs from the median of up to seven active days before it by at least 30% (10 points for shares), at least three quarters of the days sit on their own side, and the difference is more than three times the day-to-day spread. It reports the date, the client version and whether that day was the first on it, the values before and after, and what the step is consistent with. It sees requests on your machine, not Anthropic's servers, so it never names a cause; when no version change coincides, it says a change in your own work is as likely.

$ npx seatledger rates
  • From 30 Sep (Claude Code 2.1.230, the first day on it), cache reads per request fell 37%, cache writes per request rose 6.8x, the share of cache writes at the 1-hour lifetime went from 100% to 0% and the cache-read share of prompt tokens went from 93% to 56% on claude-opus-5-5 — the transcript itself shows writes moving from the 1-hour to the 5-minute cache lifetime, so a pause of more than five minutes now pays for a new cache write. It starts with Claude Code 2.1.230 (previously 2.1.229).
Synthetic history: the real output of npx seatledger demo on generated transcripts with a step built in, not anyone's usage (as of 2026-10-07).

The cache lifetime, as documented on 7 October 2026

Anthropic's cost docs give the lifetime as one hour on a subscription, dropping to five minutes once you are drawing on usage credits, and five minutes by default on an API key or a cloud provider. A first message after a break longer than the lifetime misses the cache and processes the whole context again. The changelog shows how much of this has moved: ENABLE_PROMPT_CACHING_1H and FORCE_PROMPT_CACHING_5M (2.1.108), a fix for "1-hour prompt cache TTL being silently downgraded to 5 minutes" (2.1.129), the promptCacheTtl and subagentPromptCacheTtl settings (2.1.243), and a likely cause for cache misses in /cost and the status line (2.1.260). A step in the 1-hour share of cache writes that starts on the first day of a new version is worth reading that version's changelog for.

What people reported

In #46829 (12 April 2026) a user showed, from three months of transcripts on two machines, cache writes moving from the 1-hour to the 5-minute lifetime in early March 2026. It was closed as not planned; a collaborator replied the next day that the 1-hour cache is used in some places and the 5-minute one in others, such as subagents. In #46917 a user reported about 20K more cache-creation tokens per request from v2.1.100 than v2.1.98 for the same payload. Both were found by hand, from the same transcript fields seatledger reads.

What Claude Code itself shows

On a Pro, Max, Team or Enterprise plan, /usage flags behaviours such as long context or cache misses when one accounts for 10% or more of recent usage, over the last 24 hours or 7 days on this machine, and its Session block has a prompt-cache line for the current conversation. Use them for today. seatledger is for the date a step started, across 90 days and more, including history Claude Code has already cleaned up.

Run it

npx seatledger rates for the step changes; npx seatledger for today and the week. With a team ledger (npx seatledger team create), a step on anyone's machine is emailed to the lead on Team and Business.

$ npx seatledger

Free and MIT. It reads the transcripts on this machine, makes no network request and sends nothing unless you create or join a team ledger.

Sources

  • Claude Code docs: Manage costs effectively — read 7 October 2026: why usage climbs, the cache lifetime per plan, /usage's breakdown and prompt-cache line
  • Claude Code CHANGELOG — read 7 October 2026 at 2.1.293: the cache entries quoted on this page, by version
  • anthropics/claude-code#46829 — cache writes moving from the 1-hour to the 5-minute lifetime, opened and closed (not planned) 12 April 2026; collaborator reply 13 April 2026; 340 reactions on 7 October 2026
  • anthropics/claude-code#46917 — cache creation about 20K tokens higher from v2.1.100 than v2.1.98 for the same payload, opened 12 April 2026, open on 7 October 2026
  • anthropics/claude-code#38335 — Max plan session limits exhausted abnormally fast since 23 March 2026, open on 7 October 2026