live All models chatlong-context

Llama 3.1 8B Instruct

Cheap enough to put inside a loop. Where classification and routing work usually ends up.

$argosy up --model llama-3-1-8b-instruct
19ms

p50 from the nearest warm region

486tok/s

Single stream, no batching

128K

Full window, no sliding truncation

$0.06

Per million tokens served

Small, fast, and cheap enough to put in a loop. Most people who start on the 70B and then measure end up moving classification and routing work down to this one.

Deploymentdep_8f14e4
Parameters8B
Quantisationfp8
LicenceLlama 3.1 Community
Home regioniad1
Extra replicas11
p50 latency21ms
p99 latency52ms
Requests / s4100
Figures are illustrative — this is a demonstration site.
Last 30 minutesiad1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.