live All models chatcodelong-context

Qwen2.5 72B Instruct

The multilingual choice, and clearly the best of these at Chinese.

$argosy up --model qwen2-5-72b-instruct
41ms

p50 from the nearest warm region

198tok/s

Single stream, no batching

128K

Full window, no sliding truncation

$0.36

Per million tokens served

Stronger than the Llama at multilingual work and noticeably better at Chinese. Our Frankfurt and Singapore regions carry most of its traffic.

Deploymentdep_c9f0f8
Parameters72B
Quantisationfp8
LicenceQwen
Home regionfra1
Extra replicas4
p50 latency44ms
p99 latency104ms
Requests / s1210
Figures are illustrative — this is a demonstration site.
Last 30 minutesfra1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.