live All models chat

Mixtral 8x7B Instruct

Eight experts, two of them awake per token. Cheaper to serve than the parameter count implies.

$argosy up --model mixtral-8x7b-instruct
31ms

p50 from the nearest warm region

276tok/s

Single stream, no batching

32K

Full window, no sliding truncation

$0.22

Per million tokens served

A mixture of experts, so the memory footprint is large but only two experts run per token. Cheaper to serve than its parameter count suggests, which is why the price looks wrong at first glance.

Deploymentdep_6512bd
Parameters46.7B MoE
Quantisationfp8
LicenceApache 2.0
Home regionams1
Extra replicas3
p50 latency33ms
p99 latency82ms
Requests / s1760
Figures are illustrative — this is a demonstration site.
Last 30 minutesams1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.