Model Routing & Fallback
One endpoint, several models
Applications that call provider SDKs directly often duplicate the same concerns: credential selection, retry budgets, timeouts, spend attribution, and handling overload responses such as 429, 503, or provider-specific codes. An AI gateway can centralise those concerns.
Concretely it can hold provider credentials, apply per-tenant rate limits and spend caps, and choose which model receives the request. Prompt or completion logging requires deliberate redaction, access control, retention, and consent; centralising it without those controls centralises sensitive data too.
The gateway also gives you a place to stand when a provider changes. Model names get deprecated, prices change, a new tier appears that is better on your workload. If every service names a model directly, that is a fleet-wide migration; if the gateway names it, it is a config change with a canary.
┌── auth, keys, quota, spend caps
│
client ──▶ AI gateway ──┬──▶ fast tier (most requests)
│ │
│ └──▶ deep tier (the ones that need it)
│
└── logging, evals, cost attributionThe gateway is on the request path of every model call you make, which makes it both the best place to put policy and a single point of failure worth designing carefully.
4 components3 connections0:00
Recording…