Portland

Portland, US

12ms

From a client in the same metro

77%

Free capacity right now

Shared

Reservable on the Dedicated plan

pdx1

Use this in routing rules

Peering is direct to the two largest local transit providers, so a request from a client in the metro reaches the gateway in about 12 ms before any inference happens at all. Shared capacity only. If you need a replica reserved for one workload, the nearest region offering it is a routing rule away.

Every model in the catalogue that fits on a single card is warm here. The largest mixture-of-experts models are not — those live in six regions and requests are routed to the nearest of them, which is stated plainly in the response headers rather than hidden.

Pin a workload here
regions: ["pdx1"]
residency: "strict"
fallback: false
strict means fail rather than route elsewhere

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.