Model retirements · Vertex AI (Gemini Enterprise Agent Platform) · checked
gemini-2.0-flash (Vertex AI) was shut down on 1 June 2026
Google Cloud shut down gemini-2.0-flash on Vertex AI on 1 June 2026 and recommends gemini-3.1-flash-lite as the replacement. No price ratio: the retiring model is not listed on the Agent Platform pricing page when checked on 8 October 2026, so no ratio is given. Google Cloud documents three changes that can break code: temperature, top_p and top_k are deprecated and should be removed; thinking_budget gives way to thinking_level; a function response that does not match its call returns an empty answer, not an error. It can no longer be asked, so its answers cannot be recorded now; compare the replacement against the outputs you still have.
Find every place your code and config name it, with this page's facts beside each:
$ npx --allow-git=root github:agentwares/model-migrate scan
npx --allow-git=root github:agentwares/model-migrate scan lists every model ID in your code and config that has a shutdown date, with its replacement, price ratio and this page's link. It needs no account and no API key, reads your files on your machine and sends nothing; its one request fetches this data. Free and MIT licensed.
The facts
- Platform
- Vertex AI (Gemini Enterprise Agent Platform)
- Model ID
gemini-2.0-flash- Announced
- no date on the page
- Shutdown
- 1 June 2026
- Recommended replacement
gemini-3.1-flash-lite- List price per million tokens, input / output
- The retiring model is not listed on the Agent Platform pricing page when checked on 8 October 2026, so no ratio is given.
Prices are the provider's list prices for the standard tier, text input and the shortest context tier. A per-token ratio leaves out how many tokens each model spends on the same prompt; only your own prompts show that.
What changes when you move
- For Gemini 3.x models, temperature, top_p and top_k are deprecated: Google says to remove them from every request and to use a system instruction for determinism. source
- thinking_budget is deprecated on Gemini 3.x in favour of thinking_level. source
- Every FunctionResponse must carry the id and name of its FunctionCall, exactly one per call. Google says the API does not yet return an error for a mismatch; the model returns an empty response with finish_reason STOP instead. source
How to test the replacement on your own prompts
- Find every call that names
gemini-2.0-flash:npx --allow-git=root github:agentwares/model-migrate scanprints each file and line. - gemini-2.0-flash can no longer answer, so use the outputs you already logged, or your current outputs, as the baseline.
- Run the same prompts on
gemini-3.1-flash-lite, a cheaper sibling and a model from another provider, and compare with checks that need no judge: exact or normalised match, JSON validity and schema, tool-call names and arguments, refusals and length.npx --allow-git=root github:agentwares/model-migrate comparedoes this on your keys. - Price the move with the token counts each provider reports for those prompts, not with the list price alone;
comparereports the cost per 1,000 calls. - Make the changes listed above before you switch.
Keep testing until the date
capture and compare run on your keys and keep everything on your machine. To re-run the same prompts against your own endpoint whenever a provider ships a model, with the history kept, npx --allow-git=root github:agentwares/model-migrate export --agentcheck writes them as traces that agentcheck imports onto a target you monitor there. On its Pro plan they become checks that re-run when a provider ships a model, with 90 days of history.
Migration watch is not built. It would re-run your captured prompts on each replacement and each new model until 1 June 2026, with no endpoint needed, and email you what changed. If you would pay for that, say so with one click. The click is counted; nothing else is sent or stored.
Counted. Thank you; nothing else was sent.Also shutting down on 1 June 2026
- gemini-2.0-flash (Gemini API)
- gemini-2.0-flash-001 (Gemini API)
- gemini-2.0-flash-lite (Gemini API)
- gemini-2.0-flash-lite-001 (Gemini API)
- gemini-2.0-flash-lite (Vertex AI (Gemini Enterprise Agent Platform))
Other Vertex AI (Gemini Enterprise Agent Platform) retirements
- gemini-2.5-flash — 20 Oct 2026
- gemini-2.5-flash-lite — 20 Oct 2026
- gemini-2.5-pro — 20 Oct 2026
- gemini-3.6-flash — 19 Nov 2026
- gemini-3.7-flash — 28 Jan 2027
- gemini-1.5-flash-002 — 24 Sep 2025
Every announced shutdown, by date
Sources
- Vertex AI (Gemini Enterprise Agent Platform) deprecations — shutdown date and replacement, read 8 Oct 2026
- docs.cloud.google.com/gemini-enterprise-agent-platform/models/migrate — the changes above
Built only from the providers' own pages. No vendor intent is implied: a date is what the page says on the day it was read. Data: /models/retirements.json.
model-migrate is a free tool from agentcheck, which re-runs an agent's checks when a provider ships a model and replays recorded traces nightly, so a model change that breaks it arrives as one diff by email.