which-model · Claude Code and Codex

Which model actually answered your coding agent, from the transcripts on your own machine.

$ npx --allow-git=root github:agentwares/which-model

It reads the transcripts Claude Code and Codex already keep and prints which models answered by day and session, the switches nothing you did explains, the reasoning tokens per model with any spike at a fixed value, and step changes in context ceiling or tokens per task with the client version where each starts. MIT, no account, no telemetry; it makes no network request. Source

What it prints

$ npx --allow-git=root github:agentwares/which-model
which-model — which model actually answered, from this machine's transcripts
All history on this machine

Models that answered, last 7 active days (Codex: the model it requested)
  Wed 7 Oct   Claude Code  claude-fable-5-1 29 (100%)
  Tue 6 Oct   Claude Code  claude-fable-5-1 32 (84%) · claude-haiku-4-5-20251001 6 (16%)
              Codex        gpt-5.6-sol 28 (100%)
  Mon 5 Oct   Claude Code  claude-fable-5-1 32 (84%) · claude-haiku-4-5-20251001 6 (16%)
              Codex        gpt-5.6-sol 28 (100%)
  Sun 4 Oct   Claude Code  claude-fable-5-1 32 (100%)
              Codex        gpt-5.6-sol 28 (100%)
  Sat 3 Oct   Claude Code  claude-fable-5-1 32 (84%) · claude-haiku-4-5-20251001 6 (16%)
              Codex        gpt-5.6-sol 28 (100%)
  Fri 2 Oct   Claude Code  claude-fable-5-1 32 (100%)
              Codex        gpt-5.6-sol 28 (100%)
  Thu 1 Oct   Claude Code  claude-fable-5-1 32 (100%)
              Codex        gpt-5.6-sol 28 (100%)
  --json lists every day and every session.

Served-model switches nothing you did explains (Claude Code, main thread): 2 in 1 session
  16 Jul 02:42  session 63c29f12  Claude Code 2.1.197
    claude-fable-5 → claude-opus-4-8 after a fallback notice the client recorded (claude-fable-5 →
    claude-opus-4-8, trigger: refusal); served 308 responses over 21.5 h, until the next switch.
  17 Jul 00:15  session 63c29f12  Claude Code 2.1.197
    claude-opus-4-8 → claude-fable-5 with nothing recorded before it (no /model, no fallback notice,
    same client version); served 12 responses over 0.8 h, to the end of the session.
  Not listed: 1 after your own /model.
  Answered by a different model than the one last recorded as requested: 308 responses in 1 session.

Reasoning tokens per response, by model
  model                                    responses  zero  median    p90    p99    max  spikes
  Codex · gpt-5.6-sol                          2,744    0%     260    853  2,182  5,830  none
  Claude Code · claude-fable-5                 1,958    0%     461  1,211  2,550  5,385  none
  Claude Code · claude-fable-5-1               1,181    0%     415  1,203  2,746  5,203  none
  Codex · gpt-5.5                                648    0%     516  1,034  1,994  3,929  516 (49%), 1,034 (7%)
  Claude Code · claude-opus-4-8                  308    0%     328    848  1,537  2,001  none
  Claude Code · claude-haiku-4-5-20251001        264    0%     123    347    732    912  none
  • Codex · gpt-5.5: 317 responses (49%) used exactly 516 reasoning tokens, over 100x as many as the
    counts within 5% either side, 1 Jul–31 Aug (0.144.3) — consistent with a cap or budget at that
    value; it cannot show whether the reasoning was cut short.
  • Codex · gpt-5.5: 46 responses (7%) used exactly 1,034 reasoning tokens, over 100x as many as the
    counts within 5% either side, 1 Jul–28 Aug (0.144.3) — consistent with a cap or budget at that
    value; it cannot show whether the reasoning was cut short.

Step changes (seatledger's rolling baseline)
  • From 13 Jul (Codex 0.144.3, the first day on it), the largest prompt of the day on gpt-5.6-sol
    fell 27% (343K → 251K) — consistent with a lower context ceiling, earlier compaction, or shorter
    sessions in your own work. It starts with Codex 0.144.3 (previously 0.144.2).

Context window, as Codex reported it
  • From 13 Jul (Codex 0.144.3, the first day on it), Codex reported a context window of 258,400
    tokens for gpt-5.6-sol, down from 353,400.
  • Codex reported a context window of 258,400 tokens for gpt-5.5 from 1 Jul (Codex 0.144.3), the
    first reading in this history.
  • Codex reported a context window of 353,400 tokens for gpt-5.6-sol from 1 Jul (Codex 0.144.2),
    the first reading in this history.

Codex records the model it requested for each turn, not the model that answered, so a Codex reroute
cannot be seen from its rollouts.
A switch with nothing recorded before it is not proof of a reroute: a model picked in a way the
transcript does not record (a picker, or a relaunch with --model on the same version) looks the
same.
Read 143 Claude Code transcripts and 125 Codex rollouts on this machine. Nothing was sent anywhere.
Synthetic history: the real output of npx --allow-git=root github:agentwares/which-model demo, which writes generated transcripts in the real Claude Code and Codex formats and reads them. None of it is anyone's usage (as of 2026-10-07).

What it answers

  • Did my session switch models? Claude Code writes the model that answered on every response. which-model finds each switch in a session's main thread and says what the transcript records just before it: your /model, a fallback notice the client wrote, a client update, or nothing; and how many responses and hours the other model then served.
  • Is reasoning stuck at a fixed number? The reasoning tokens of every response, per model, with any exact count that holds far more responses than the values around it, such as 516.
  • Did the context ceiling or the cost of a task change? Step changes in reasoning per response, tokens per task, the largest prompt of the day and the auto-compaction point, by the same rolling baseline as seatledger, and every change in the context window Codex reports.

Each finding says what moved, when, on which version, and what it is consistent with. It never says why: it sees your machine, not the vendor's servers.

What it reads, and what it never sends

ClientWhereWhat it takes
Claude Code~/.claude/projectseach response's model (the one that answered), client version, time, session id and token counts, including thinking tokens; whether it came from a subagent; your /model choices; the fallback notices the client writes; auto-compaction points
Codex~/.codex/sessionsthe model Codex requested for each turn (Codex does not record the one that answered), client version, time, thread id, token counts including reasoning tokens, and the context window Codex reported
  • Never sent: anything. which-model makes no network request and has no telemetry.
  • Never kept or printed: prompts, responses, tool output, file paths or project names. The output holds model ids, client versions, dates, counts and session ids, so you can paste it into an issue as it is.
  • Kept on your machine: counts and model ids per transcript in ~/.which-model, so a second run is fast and history the clients delete is not lost. Delete the directory to forget it.

Guides

Is Claude Code switching my model? — It can, and your transcripts record it: each response carries the model that answered. How to find every switch, the notice before it, and how long it lasted.

Why are Codex reasoning tokens stuck at 516? — Codex writes each response's reasoning tokens to disk. See whether yours cluster at 516 or 1,034, on which model and since when, and what that cannot show.

For a team

which-model is one machine. seatledger's team ledger shows the whole team's model mix today, by model and by day, free for up to 3 developers: seatledger.

A team alert when the served model changes on anyone's machine does not exist yet. If you would use it, one click tells us; nothing about you is recorded beyond a page view.

Tell me when there is a team alert for served-model changes