Skip to content
Learn/

Model Routing & Fallback

1 / 7

One endpoint, several models

Applications that call provider SDKs directly often duplicate the same concerns: credential selection, retry budgets, timeouts, spend attribution, and handling overload responses such as 429, 503, or provider-specific codes. An AI gateway can centralise those concerns.

Concretely it can hold provider credentials, apply per-tenant rate limits and spend caps, and choose which model receives the request. Prompt or completion logging requires deliberate redaction, access control, retention, and consent; centralising it without those controls centralises sensitive data too.

The gateway also gives you a place to stand when a provider changes. Model names get deprecated, prices change, a new tier appears that is better on your workload. If every service names a model directly, that is a fleet-wide migration; if the gateway names it, it is a config change with a canary.

                 ┌── auth, keys, quota, spend caps
                 │
client ──▶ AI gateway ──┬──▶ fast tier    (most requests)
                 │      │
                 │      └──▶ deep tier    (the ones that need it)
                 │
                 └── logging, evals, cost attribution

The gateway is on the request path of every model call you make, which makes it both the best place to put policy and a single point of failure worth designing carefully.

4 components3 connections0:00

Traffic
100req/s
p50
3.46s
p99
7.46s
Errors
0.06%
Dropped
0.1req/s
Cost
$6.55M/mo