which-model guide · checked
Why are Codex reasoning tokens stuck at 516?
Whether yours are is something you can check on your own machine; why is not. Codex writes the reasoning tokens of every response into its rollouts under ~/.codex/sessions, and the 516 clusters were found by people adding those numbers up. In openai/codex#30364 (27 June 2026) a user aggregated them and found gpt-5.5 responses piling up at exactly 516, 1,034 and 1,552 tokens; the issue was closed as completed on 23 August 2026, and later comments ask why. npx --allow-git=root github:agentwares/which-model shows your own distribution per model, any exact count that holds far more responses than the values around it, and the dates and Codex versions it spans. It reads locally and sends nothing.
What Codex writes
| Field | What it tells you |
|---|---|
| token_count events: last_token_usage.reasoning_output_tokens | the reasoning tokens of the response just finished |
| token_usage_record: usage.reasoning_output_tokens (newer versions) | the same, one record per response |
| turn_context.model | the model Codex requested for the turn (not the one that answered) |
| token_count: model_context_window | the context window Codex reported for that model |
| session_meta.cli_version | the Codex version that wrote the rollout |
The method
For each model, which-model collects the reasoning tokens of every response the rollouts record and reports the share at zero, the median, the 90th and 99th percentiles and the maximum. A spike is an exact count of at least 64 tokens that holds at least 1% of that model's responses (and at least 10) and at least eight times the average count of the values within 5% either side; neighbouring spikes collapse to the tallest. A spike is consistent with a cap or budget at that value. It cannot show that any reasoning was cut short, and a lower effort setting moves these numbers too.
Separately, it looks for step changes in the median reasoning tokens per response, day by day, with the rolling baseline seatledger uses for token rates, and says which Codex version each starts on. Every change in the context window Codex reports is listed exactly, because that is a value Codex wrote, not an estimate.
$ npx --allow-git=root github:agentwares/which-model
Reasoning tokens per response, by model
model responses zero median p90 p99 max spikes
Codex · gpt-5.6-sol 2,744 0% 260 853 2,182 5,830 none
Claude Code · claude-fable-5 1,958 0% 461 1,211 2,550 5,385 none
Claude Code · claude-fable-5-1 1,181 0% 415 1,203 2,746 5,203 none
Codex · gpt-5.5 648 0% 516 1,034 1,994 3,929 516 (49%), 1,034 (7%)
Claude Code · claude-opus-4-8 308 0% 328 848 1,537 2,001 none
Claude Code · claude-haiku-4-5-20251001 264 0% 123 347 732 912 none
• Codex · gpt-5.5: 317 responses (49%) used exactly 516 reasoning tokens, over 100x as many as the
counts within 5% either side, 1 Jul–31 Aug (0.144.3) — consistent with a cap or budget at that
value; it cannot show whether the reasoning was cut short.
• Codex · gpt-5.5: 46 responses (7%) used exactly 1,034 reasoning tokens, over 100x as many as the
counts within 5% either side, 1 Jul–28 Aug (0.144.3) — consistent with a cap or budget at that
value; it cannot show whether the reasoning was cut short.What people reported
In #30364 the reporter took Codex's own token metadata from February to June 2026 and showed reasoning output stacking at fixed values; other users reported the same counts from their own rollouts in the thread. In #32806 (13 July 2026) the context window Codex reported for gpt-5.6-sol fell from 353,400 to 258,400 tokens. Both are visible in the rollouts on your own machine, if they reach back that far.
What it cannot tell you
It sees your machine, not OpenAI's servers, so it never says why a count clusters or a window changed. Codex records the model it requested for each turn, not the one that answered, so a reroute to another model is not visible in Codex rollouts at all.
Run it
npx --allow-git=root github:agentwares/which-model --client codex for Codex alone; --since 30d to narrow it; --json for every count.
$ npx --allow-git=root github:agentwares/which-modelFree and MIT. It reads the transcripts on this machine, makes no network request and sends nothing. Its output holds model ids, versions, dates, counts and session ids, never a prompt, a path or a project name, so you can paste it into an issue.
For a team
seatledger's team ledger shows the whole team's model mix today, by model and by day: seatledger. A team alert when the served model or its reasoning changes does not exist yet; one click tells us you would use it.
Tell me when there is a team alert for served-model changesSources
- openai/codex#30364 — gpt-5.5 reasoning tokens clustered at 516, 1,034 and 1,552, opened 27 June 2026, closed as completed 23 August 2026; 289 👍 and 188 comments on 8 October 2026
- openai/codex#32806 — the context window Codex reported for gpt-5.6-sol fell from 353,400 to 258,400 tokens, opened and closed 13 July 2026; read 8 October 2026
- Codex source: codex-rs/protocol/src/protocol.rs — read 8 October 2026 (latest release rust-v0.161.0): TokenUsage.reasoning_output_tokens and TokenUsageInfo.model_context_window