From a client in the same metro
Free capacity right now
Reservable on the Dedicated plan
Use this in routing rules
Peering is direct to the two largest local transit providers, so a request from a client in the metro reaches the gateway in about 12 ms before any inference happens at all. Shared capacity only. If you need a replica reserved for one workload, the nearest region offering it is a routing rule away.
Every model in the catalogue that fits on a single card is warm here. The largest mixture-of-experts models are not — those live in six regions and requests are routed to the nearest of them, which is stated plainly in the response headers rather than hidden.
regions: ["dfw1"] residency: "strict" fallback: false