Model retirements · Vertex AI (Gemini Enterprise Agent Platform) · checked

gemini-3.7-flash (Vertex AI) shutdown on 28 January 2027

Google Cloud shuts down gemini-3.7-flash on Vertex AI on 28 January 2027 and recommends gemini-3.8-flash as the replacement. At list prices gemini-3.8-flash costs 1× gemini-3.7-flash's input price and 1× its output price per token. Google Cloud's deprecation page documents no API changes for this move. Record your answers before 28 January 2027: after that nobody can ask gemini-3.7-flash again.

Find every place your code and config name it, with this page's facts beside each:

$ npx --allow-git=root github:agentwares/model-migrate scan

npx --allow-git=root github:agentwares/model-migrate scan lists every model ID in your code and config that has a shutdown date, with its replacement, price ratio and this page's link. It needs no account and no API key, reads your files on your machine and sends nothing; its one request fetches this data. Free and MIT licensed.

The facts

Platform
Vertex AI (Gemini Enterprise Agent Platform)
Model ID
gemini-3.7-flash
Announced
no date on the page
Shutdown
28 January 2027
Recommended replacement
gemini-3.8-flash
List price per million tokens, input / output

gemini-3.8-flash: $1.50 / $7.50 against $1.50 / $7.50, so 1× input, 1× outputAgent Platform lists both models at $0.75 / $3.75 through 31 December 2026 and $1.50 / $7.50 from 1 January 2027.

Prices are the provider's list prices for the standard tier, text input and the shortest context tier, read on 8 October 2026. A per-token ratio leaves out how many tokens each model spends on the same prompt; only your own prompts show that.

What changes when you move

Google Cloud's deprecation page lists no parameter or API changes for this move. The model's answers still change: a replacement is a different model.

How to test the replacement on your own prompts

  1. Find every call that names gemini-3.7-flash: npx --allow-git=root github:agentwares/model-migrate scan prints each file and line.
  2. Before 28 January 2027, record what gemini-3.7-flash answers on a sample of your real prompts, with the token counts Google Cloud reports, on your own key: npx --allow-git=root github:agentwares/model-migrate capture prompts.jsonl --model gemini-3.7-flash. After the shutdown that record cannot be made. It calls the same model ID through the Gemini API with a Gemini API key; Vertex AI credentials are not supported.
  3. Run the same prompts on gemini-3.8-flash, a cheaper sibling and a model from another provider, and compare with checks that need no judge: exact or normalised match, JSON validity and schema, tool-call names and arguments, refusals and length. npx --allow-git=root github:agentwares/model-migrate compare does this on your keys.
  4. Price the move with the token counts each provider reports for those prompts, not with the list price alone; compare reports the cost per 1,000 calls.

Keep testing until the date

capture and compare run on your keys and keep everything on your machine. To re-run the same prompts against your own endpoint whenever a provider ships a model, with the history kept, npx --allow-git=root github:agentwares/model-migrate export --agentcheck writes them as traces that agentcheck imports onto a target you monitor there. On its Pro plan they become checks that re-run when a provider ships a model, with 90 days of history.

Migration watch is not built. It would re-run your captured prompts on each replacement and each new model until 28 January 2027, with no endpoint needed, and email you what changed. If you would pay for that, say so with one click. The click is counted; nothing else is sent or stored.

Other Vertex AI (Gemini Enterprise Agent Platform) retirements

Every announced shutdown, by date

Sources

Built only from the providers' own pages. No vendor intent is implied: a date is what the page says on the day it was read. Data: /models/retirements.json.

model-migrate is a free tool from agentcheck, which re-runs an agent's checks when a provider ships a model and replays recorded traces nightly, so a model change that breaks it arrives as one diff by email.