Model retirements · Vertex AI (Gemini Enterprise Agent Platform) · checked
gemini-2.5-pro (Vertex AI) shutdown on 20 October 2026
Google Cloud shuts down gemini-2.5-pro on Vertex AI on 20 October 2026 and recommends gemini-3.8-flash or gemini-3.5-flash as the replacement. At list prices gemini-3.8-flash costs 0.6× gemini-2.5-pro's input price and 0.38× its output price per token until 31 December 2026. Google Cloud documents three changes that can break code: temperature, top_p and top_k are deprecated and should be removed; thinking_budget gives way to thinking_level; a function response that does not match its call returns an empty answer, not an error. Record your answers before 20 October 2026: after that nobody can ask gemini-2.5-pro again.
Find every place your code and config name it, with this page's facts beside each:
$ npx --allow-git=root github:agentwares/model-migrate scan
npx --allow-git=root github:agentwares/model-migrate scan lists every model ID in your code and config that has a shutdown date, with its replacement, price ratio and this page's link. It needs no account and no API key, reads your files on your machine and sends nothing; its one request fetches this data. Free and MIT licensed.
The facts
- Platform
- Vertex AI (Gemini Enterprise Agent Platform)
- Model ID
gemini-2.5-pro- Announced
- no date on the page
- Shutdown
- 20 October 2026
- Recommended replacement
gemini-3.8-flashorgemini-3.5-flash- List price per million tokens, input / output
gemini-3.8-flash: $0.75 / $3.75 against $1.25 / $10, so 0.6× input, 0.38× output (until 31 Dec 2026)gemini-3.8-flash: $1.50 / $7.50 against $1.25 / $10, so 1.2× input, 0.75× output (from 1 Jan 2027)gemini-3.5-flash: $1.50 / $9 against $1.25 / $10, so 1.2× input, 0.9× output
Prices are the provider's list prices for the standard tier, text input and the shortest context tier, read on 8 October 2026. A per-token ratio leaves out how many tokens each model spends on the same prompt; only your own prompts show that.
What changes when you move
- For Gemini 3.x models, temperature, top_p and top_k are deprecated: Google says to remove them from every request and to use a system instruction for determinism. source
- thinking_budget is deprecated on Gemini 3.x in favour of thinking_level. source
- Every FunctionResponse must carry the id and name of its FunctionCall, exactly one per call. Google says the API does not yet return an error for a mismatch; the model returns an empty response with finish_reason STOP instead. source
- On the Gemini API (ai.google.dev) the same model ID had no shutdown date announced when checked on 8 October 2026; Google limits 2.5 access there to projects that used it before. This date is for Vertex AI, now Gemini Enterprise Agent Platform.
- Google says retirement dates on this page may be extended but will not be moved earlier.
How to test the replacement on your own prompts
- Find every call that names
gemini-2.5-pro:npx --allow-git=root github:agentwares/model-migrate scanprints each file and line. - Before 20 October 2026, record what gemini-2.5-pro answers on a sample of your real prompts, with the token counts Google Cloud reports, on your own key:
npx --allow-git=root github:agentwares/model-migrate capture prompts.jsonl --model gemini-2.5-pro. After the shutdown that record cannot be made. It calls the same model ID through the Gemini API with a Gemini API key; Vertex AI credentials are not supported. - Run the same prompts on
gemini-3.8-flash, a cheaper sibling and a model from another provider, and compare with checks that need no judge: exact or normalised match, JSON validity and schema, tool-call names and arguments, refusals and length.npx --allow-git=root github:agentwares/model-migrate comparedoes this on your keys. - Price the move with the token counts each provider reports for those prompts, not with the list price alone;
comparereports the cost per 1,000 calls. - Make the changes listed above before you switch.
Keep testing until the date
capture and compare run on your keys and keep everything on your machine. To re-run the same prompts against your own endpoint whenever a provider ships a model, with the history kept, npx --allow-git=root github:agentwares/model-migrate export --agentcheck writes them as traces that agentcheck imports onto a target you monitor there. On its Pro plan they become checks that re-run when a provider ships a model, with 90 days of history.
Migration watch is not built. It would re-run your captured prompts on each replacement and each new model until 20 October 2026, with no endpoint needed, and email you what changed. If you would pay for that, say so with one click. The click is counted; nothing else is sent or stored.
Counted. Thank you; nothing else was sent.Also shutting down on 20 October 2026
- gemini-2.5-flash (Vertex AI (Gemini Enterprise Agent Platform))
- gemini-2.5-flash-lite (Vertex AI (Gemini Enterprise Agent Platform))
Other Vertex AI (Gemini Enterprise Agent Platform) retirements
- gemini-3.6-flash — 19 Nov 2026
- gemini-3.7-flash — 28 Jan 2027
- gemini-2.0-flash — 1 Jun 2026
- gemini-2.0-flash-lite — 1 Jun 2026
- gemini-1.5-flash-002 — 24 Sep 2025
- gemini-1.5-pro-002 — 24 Sep 2025
Every announced shutdown, by date
Sources
- Vertex AI (Gemini Enterprise Agent Platform) deprecations — shutdown date and replacement, read 8 Oct 2026
- Vertex AI (Gemini Enterprise Agent Platform) pricing — prices, read 8 Oct 2026
- docs.cloud.google.com/gemini-enterprise-agent-platform/models/migrate — the changes above
Built only from the providers' own pages. No vendor intent is implied: a date is what the page says on the day it was read. Data: /models/retirements.json.
model-migrate is a free tool from agentcheck, which re-runs an agent's checks when a provider ships a model and replays recorded traces nightly, so a model change that breaks it arrives as one diff by email.