DeepSeek V3
The largest thing we serve, in six regions rather than thirty-one.
58ms
p50 from the nearest warm region
132tok/s
Single stream, no batching
128K
Full window, no sliding truncation
$0.58
Per million tokens served
The largest thing we serve. It needs eight cards per replica, so it lives in six regions rather than thirty-one, and requests from elsewhere are routed to the nearest of those.
Deploymentdep_c20ad4
Parameters671B MoE
Quantisationfp8
LicenceDeepSeek
Home regionsin1
Extra replicas2
p50 latency62ms
p99 latency148ms
Requests / s640
Figures are illustrative — this is a demonstration site.
Last 30 minutessin1
cache hitqueue 0ms
errors 0