live All models chat

Gemma 2 27B

An 8K window that will not be extended, and a low price that reflects it.

$argosy up --model gemma-2-27b
24ms

p50 from the nearest warm region

332tok/s

Single stream, no batching

8K

Full window, no sliding truncation

$0.16

Per million tokens served

An 8K window, which is short by current standards, and it will not be extended. Good at short-form generation and cheap to keep warm.

Deploymentdep_aab323
Parameters27B
Quantisationfp8
LicenceGemma
Home regionlhr1
Extra replicas3
p50 latency26ms
p99 latency64ms
Requests / s2410
Figures are illustrative — this is a demonstration site.
Last 30 minuteslhr1
cache hitqueue 0ms errors 0

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.