Pricing

Per-token, dedicated, or committed. Egress, failed requests and the first million tokens each month are not metered.

Three plans, and the only one most teams ever need is the first. There is no seat cost, no minimum commitment and no enterprise-only feature gate on anything that affects reliability — SSO, audit logs, region pinning and private networking are on every plan, because charging for those is charging for a security posture.

Per-token

You are billed for the tokens you actually served, per model, priced at the rate on the model’s page. Nothing is billed for a model sitting idle. A model that has not been called in an hour still answers in about 40 ms, so keeping something warm is not a thing you have to pay for.

Model class Input / M Output / M Cold start
Small — up to 15B $0.06 $0.09 18 ms
Mid — 15B to 40B $0.14 $0.20 27 ms
Large — 40B and up $0.34 $0.48 38 ms
Mixture of experts $0.22 $0.58 58 ms
Embeddings $0.01 6 ms

Dedicated

A replica held for one workload, in a region you choose, billed monthly at $2,400 per replica rather than per token. It is never shared and never cold. Worth it somewhere north of roughly 400 million tokens a month on a single model — below that, per-token is cheaper and we will say so if you ask.

Committed

An annual commitment against per-token usage, discounted on a published schedule rather than a negotiated one. Twelve per cent at $50k, eighteen at $150k, twenty-four at $500k. Unused commitment rolls forward one quarter.


What is not metered

  • Egress. Responses leaving our network are not billed. Neither are the images you send to a vision model.
  • Failed requests. A 5xx from us is not billed. Neither is a request that we routed out of its pinned region — if we broke the residency guarantee, you do not pay for the call.
  • The first million tokens each month. Enough to build something and decide.

How the invoice reads

One line per model per region, with the token counts that produced it, and a downloadable CSV of every request that contributed. If our number and your number disagree, the CSV is how you find out which of us is wrong.

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.