Qwen2.5 72B Instruct
The multilingual choice, and clearly the best of these at Chinese.
41ms
p50 from the nearest warm region
198tok/s
Single stream, no batching
128K
Full window, no sliding truncation
$0.36
Per million tokens served
Stronger than the Llama at multilingual work and noticeably better at Chinese. Our Frankfurt and Singapore regions carry most of its traffic.
Deploymentdep_c9f0f8
Parameters72B
Quantisationfp8
LicenceQwen
Home regionfra1
Extra replicas4
p50 latency44ms
p99 latency104ms
Requests / s1210
Figures are illustrative — this is a demonstration site.
Last 30 minutesfra1
cache hitqueue 0ms
errors 0