DeepSeek R1 Distill 32B
Thinks out loud before it answers, which changes what its first token means.
34ms
p50 from the nearest warm region
244tok/s
Single stream, no batching
128K
Full window, no sliding truncation
$0.20
Per million tokens served
Emits a reasoning trace before the answer. Time to first token is honest but misleading here — the first token is the opening of the trace, not the opening of the reply.
Deploymentdep_c51ce4
Parameters32B
Quantisationfp8
LicenceMIT
Home regioniad1
Extra replicas4
p50 latency36ms
p99 latency88ms
Requests / s1580
Figures are illustrative — this is a demonstration site.
Last 30 minutesiad1
cache hitqueue 0ms
errors 0