Mixtral 8x7B Instruct
Eight experts, two of them awake per token. Cheaper to serve than the parameter count implies.
31ms
p50 from the nearest warm region
276tok/s
Single stream, no batching
32K
Full window, no sliding truncation
$0.22
Per million tokens served
A mixture of experts, so the memory footprint is large but only two experts run per token. Cheaper to serve than its parameter count suggests, which is why the price looks wrong at first glance.
Deploymentdep_6512bd
Parameters46.7B MoE
Quantisationfp8
LicenceApache 2.0
Home regionams1
Extra replicas3
p50 latency33ms
p99 latency82ms
Requests / s1760
Figures are illustrative — this is a demonstration site.
Last 30 minutesams1
cache hitqueue 0ms
errors 0