Ashburn

Ashburn, US

11ms

From a client in the same metro

62%

Free capacity right now

Dedicated

Reservable on the Dedicated plan

iad1

Use this in routing rules

Peering is direct to the two largest local transit providers, so a request from a client in the metro reaches the gateway in about 11 ms before any inference happens at all. Dedicated capacity is reservable here: a replica held for your workload alone, billed monthly rather than per token.

Every model in the catalogue that fits on a single card is warm here. The largest mixture-of-experts models are not — those live in six regions and requests are routed to the nearest of them, which is stated plainly in the response headers rather than hidden.

Pin a workload here
regions: ["iad1"]
residency: "strict"
fallback: false
strict means fail rather than route elsewhere

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.