live All models chatlong-contextreasoning

Llama 3.3 70B Instruct

The default for anything that has to sound like a person and still be checkable.

$argosy up --model llama-3-3-70b-instruct
38ms

p50 from the nearest warm region

212tok/s

Single stream, no batching

128K

Full window, no sliding truncation

$0.34

Per million tokens served

The default choice for anything that needs to sound like a person and still be checkable. Serves out of every region; we keep six warm replicas on the US east and west coasts because that is where the traffic is.

Deploymentdep_167909
Parameters70B
Quantisationfp8
LicenceLlama 3.3 Community
Home regioniad1
Extra replicas6
p50 latency41ms
p99 latency96ms
Requests / s1420
Figures are illustrative — this is a demonstration site.
Last 30 minutesiad1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.