live All models chatcodelong-contextreasoning

DeepSeek V3

The largest thing we serve, in six regions rather than thirty-one.

$argosy up --model deepseek-v3
58ms

p50 from the nearest warm region

132tok/s

Single stream, no batching

128K

Full window, no sliding truncation

$0.58

Per million tokens served

The largest thing we serve. It needs eight cards per replica, so it lives in six regions rather than thirty-one, and requests from elsewhere are routed to the nearest of those.

Deploymentdep_c20ad4
Parameters671B MoE
Quantisationfp8
LicenceDeepSeek
Home regionsin1
Extra replicas2
p50 latency62ms
p99 latency148ms
Requests / s640
Figures are illustrative — this is a demonstration site.
Last 30 minutessin1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.