Model retirements · Vertex AI (Gemini Enterprise Agent Platform) · checked

gemini-2.5-flash (Vertex AI) shutdown on 20 October 2026

Google Cloud shuts down gemini-2.5-flash on Vertex AI on 20 October 2026 and recommends gemini-3.8-flash, gemini-3.5-flash-lite, or gemini-3.1-flash-lite as the replacement. At list prices gemini-3.8-flash costs 2.5× gemini-2.5-flash's input price and 1.5× its output price per token until 31 December 2026. Google Cloud documents three changes that can break code: temperature, top_p and top_k are deprecated and should be removed; thinking_budget gives way to thinking_level; a function response that does not match its call returns an empty answer, not an error. Record your answers before 20 October 2026: after that nobody can ask gemini-2.5-flash again.

Find every place your code and config name it, with this page's facts beside each:

$ npx --allow-git=root github:agentwares/model-migrate scan

npx --allow-git=root github:agentwares/model-migrate scan lists every model ID in your code and config that has a shutdown date, with its replacement, price ratio and this page's link. It needs no account and no API key, reads your files on your machine and sends nothing; its one request fetches this data. Free and MIT licensed.

The facts

Platform
Vertex AI (Gemini Enterprise Agent Platform)
Model ID
gemini-2.5-flash
Announced
no date on the page
Shutdown
20 October 2026
Recommended replacement
gemini-3.8-flash or gemini-3.5-flash-lite or gemini-3.1-flash-lite
List price per million tokens, input / output

gemini-3.8-flash: $0.75 / $3.75 against $0.3 / $2.50, so 2.5× input, 1.5× output (until 31 Dec 2026)

gemini-3.8-flash: $1.50 / $7.50 against $0.3 / $2.50, so 5× input, 3× output (from 1 Jan 2027)

gemini-3.5-flash-lite: $0.3 / $2.50 against $0.3 / $2.50, so 1× input, 1× output

gemini-3.1-flash-lite: $0.25 / $1.50 against $0.3 / $2.50, so 0.83× input, 0.6× output

Prices are the provider's list prices for the standard tier, text input and the shortest context tier, read on 8 October 2026. A per-token ratio leaves out how many tokens each model spends on the same prompt; only your own prompts show that.

What changes when you move

How to test the replacement on your own prompts

  1. Find every call that names gemini-2.5-flash: npx --allow-git=root github:agentwares/model-migrate scan prints each file and line.
  2. Before 20 October 2026, record what gemini-2.5-flash answers on a sample of your real prompts, with the token counts Google Cloud reports, on your own key: npx --allow-git=root github:agentwares/model-migrate capture prompts.jsonl --model gemini-2.5-flash. After the shutdown that record cannot be made. It calls the same model ID through the Gemini API with a Gemini API key; Vertex AI credentials are not supported.
  3. Run the same prompts on gemini-3.8-flash, a cheaper sibling and a model from another provider, and compare with checks that need no judge: exact or normalised match, JSON validity and schema, tool-call names and arguments, refusals and length. npx --allow-git=root github:agentwares/model-migrate compare does this on your keys.
  4. Price the move with the token counts each provider reports for those prompts, not with the list price alone; compare reports the cost per 1,000 calls.
  5. Make the changes listed above before you switch.

Keep testing until the date

capture and compare run on your keys and keep everything on your machine. To re-run the same prompts against your own endpoint whenever a provider ships a model, with the history kept, npx --allow-git=root github:agentwares/model-migrate export --agentcheck writes them as traces that agentcheck imports onto a target you monitor there. On its Pro plan they become checks that re-run when a provider ships a model, with 90 days of history.

Migration watch is not built. It would re-run your captured prompts on each replacement and each new model until 20 October 2026, with no endpoint needed, and email you what changed. If you would pay for that, say so with one click. The click is counted; nothing else is sent or stored.

Also shutting down on 20 October 2026

Other Vertex AI (Gemini Enterprise Agent Platform) retirements

Every announced shutdown, by date

Sources

Built only from the providers' own pages. No vendor intent is implied: a date is what the page says on the day it was read. Data: /models/retirements.json.

model-migrate is a free tool from agentcheck, which re-runs an agent's checks when a provider ships a model and replays recorded traces nightly, so a model change that breaks it arrives as one diff by email.