AI Gateway: model selection as configuration, not code


The storefront chat has a Free / Paid toggle. Free gets a small fast model; paid gets a bigger brain. Here’s the thing: the application doesn’t know which model answers. It sends tier: free or tier: paid as metadata and names a route — the gateway’s dynamic route picks the actual model. Model policy is configuration, editable in the dashboard without a redeploy.

The story

The request path is one line:

const res = await env.AI.run(
  `dynamic/${env.DYNAMIC_ROUTE_NAME}`,
  { messages: routed },
  { gateway: { id: env.AI_GATEWAY_ID, metadata: { tier: plan } } }
);

The app targets the route by name and attaches metadata. In the AI Gateway dashboard, the dynamic route ai_gateway_routes holds a conditional like “if metadata.tier == paid then use the big model, else the small one”. The response comes back OpenAI-shaped and carries the model that actually answered — which is what the chat’s meta line displays (Lumina AI • paid • @cf/meta/...), so the routing is visible, not magic.

The war story. This demo originally used AI Gateway’s custom routes — an older mechanism that swapped models transparently based on metadata. In August 2026 an AI Gateway unification retired it: dynamic routes must now be named in the request, and the model selection lives in the gateway config. One afternoon of wiring later, the design ended up better: the app is honest about calling a route, and the gateway owns the policy. The lesson is not “migrate” — it’s that routing policy in config beats branching in code whenever the policy is business-level (tiers, models, fallbacks) rather than application logic.

One sharp edge learned the hard way: the compatibility endpoint leaked the model key into the Workers AI input, which strict model schemas reject with a 400. The binding transport (env.AI.run with the route name) passes a clean {messages} body and works. If a dynamic route 400s inexplicably, check for stray keys in the request body.

What the demo shows

Chat → toggle Free/Paid → the meta line shows a different model per tier, chosen by the gateway. Then, in AI Gateway → Dynamic Routes, flip the model behind a tier → Save + Deploy → re-send — the storefront changes models with no redeploy of the Worker, because the policy changed in the gateway, not the app.

The agent path (next post) bypasses the dynamic route and names the 70B model explicitly with tier: "agent" metadata — so all three tiers are visible in the same gateway logs, but only the two simple tiers are gateway-routed. That’s deliberate: the agent’s model choice is application logic, the tiers are business policy.

Evidence: what to capture

  • Storefront chat with Free selected → meta line Lumina AI • free • <model A> → 05-meta-free.png
  • Same with Paid → meta line shows <model B> → 05-meta-paid.png
  • AI Gateway dashboard → Dynamic Routes → the ai_gateway_routes conditional config → 05-dynamic-route-config.png
  • Flip the model behind a tier, Save + Deploy, re-send the same prompt, meta line changes → 05-model-flip.png
  • AI Gateway → Logs → both requests listed with tier metadata visible → 05-gateway-logs.png

Key takeaways

  • Send tier/context as gateway metadata and name a dynamic route: model policy becomes dashboard config.
  • The response carries the actual model — surface it in the UI so routing is auditable at a glance.
  • Gateway logs give per-request observability (models, tokens, blocks) without any APM wiring.