live All models codelong-contextreasoning

DeepSeek R1 Distill 32B

Thinks out loud before it answers, which changes what its first token means.

$argosy up --model deepseek-r1-distill-32b
34ms

p50 from the nearest warm region

244tok/s

Single stream, no batching

128K

Full window, no sliding truncation

$0.20

Per million tokens served

Emits a reasoning trace before the answer. Time to first token is honest but misleading here — the first token is the opening of the trace, not the opening of the reply.

Deploymentdep_c51ce4
Parameters32B
Quantisationfp8
LicenceMIT
Home regioniad1
Extra replicas4
p50 latency36ms
p99 latency88ms
Requests / s1580
Figures are illustrative — this is a demonstration site.
Last 30 minutesiad1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.