Skip to main content

October 3, 2026

Thinking budget on Google and Vertex AI routes

A Gemini thinking cap under providerOptions.google is ignored when the gateway answers as Vertex. Mirror the same thinkingConfig on both provider keys.

Sander Korf3 min read
llmtypescriptai-sdk

A guided movement-discovery web session runs with an on-screen AI facilitator: voice plus scripted surfaces, not therapy. Replies go through Vercel AI Gateway, a router that picks which provider serves a named model. The same Flash-class Gemini model can come back from Google AI Studio or Vertex AI, so a provider outage does not kill the turn.

When a tester waits for the guide to speak, latency is the product. Hidden reasoning tokens burn before the first word lands. I wanted a hard ceiling on that thinking phase for the live Flash-class model.

A Vercel AI SDK call carries a providerOptions bag keyed by provider id (google, vertex, gateway). Only the provider that actually answers reads its slice. A thinking budget is the cap on those hidden reasoning tokens before the first spoken token (thinkingConfig.thinkingBudget on Gemini).

The trap in review

Code review looked fine. thinkingBudget: 256 lived under providerOptions.google. Diff was green.

Production still felt sluggish on 1 October 2026. The gateway had routed the same model id through Vertex AI while failover or routing shifted. Vertex never saw the google key. The budget was ignored and turns still thought for seconds before speech. I have not re-checked chatTurns.reasoningTokens after the dual-key deploy to prove the cap landed in metrics. Treat that as a post-deploy check, not a claim in this post.

Earlier, until 28 September, the setting was thinkingLevel: 'low' because Gemini 3.8 Flash refuses minimal. Live turns that saved a memory still spent up to about 1,479 reasoning tokens, roughly six to sixteen seconds. You cannot set a numeric budget and a level together. I switched to the numeric budget and then hit the provider-key mismatch.

Mirror thinking on both keys

Reuse one object for both provider entries so review cannot drift:

const thinking = { thinkingConfig: { thinkingBudget: 256 } }
 
export const chatModelSettings = {
  maxOutputTokens: 2000,
  providerOptions: {
    gateway: { models: fallbackModelIds },
    google: thinking,
    vertex: thinking,
  },
}

A unit test should assert providerOptions.vertex deep-equals providerOptions.google for the thinking slice. That assertion is what keeps failover routing from silently dropping the cap again.

Use the demo below. Toggle Route Google vs Route Vertex, and Google only vs Both keys. Watch whether the 256 budget applies or the turn thinks uncapped.

providerOptions vs routed provider

Same Flash-class model id through an AI gateway. Only the provider that actually answers reads its keyed bag.
ignored this turnproviderOptions.google

thinkingBudget: 256

routed provider reads thisproviderOptions.vertex

(no thinkingConfig on this key)

Gateway routed to
vertex
256 budget applies?
No
Uncapped thinkingvertex deep-equals google: false

Toy latency to first spoken word

Routed provider is vertex, but providerOptions.vertex has no thinkingBudget. The turn still thinks uncapped (~9.2s here).

Routed provider
Option bag wiring

What I watch next

Any gateway that can answer as more than one provider id needs the same option shape on every id that might win the route. gateway.models fallbacks are not a substitute for per-provider thinking config.

After deploy I still owe myself a pass over stored turn metrics to see whether reasoning token counts dropped. Until that query runs, the fix is structural: same thinking object on google and vertex, plus a test that keeps them aligned.

If you cap Gemini thinking for voice UX, grep your option bags for a single-provider key and route logs for which provider actually served the turn.


Happy coding! Sander