Company

Forty-one people who buy GPUs, put them in thirty-one metros, and publish the numbers that come out.

Argosy exists because the gap between “this model works on my laptop” and “this model answers ten thousand people in Jakarta in under a second” turned out to be almost entirely operations work, and almost none of it interesting.

What we actually do

We buy GPUs, put them in thirty-one metros, and keep weights memory-mapped on local NVMe so that a model nobody has called in an hour still answers as fast as one called a second ago. Everything else on this site is a consequence of that one decision.

What we do not do

We do not train models. We do not fine-tune them for you. We do not resell somebody else’s hosted API behind our own logo — if a model is in the catalogue, the weights are on our hardware, and you can pin which hardware.

How we publish numbers

Latency figures on this site are p50 measured at the gateway from a client in the same metro, not from our own datacentre to itself. When a number gets worse we change the number. There is no marketing pass over the status page.

If our figure and your figure disagree, ours is the one that should move.

The company

Forty-one people, of whom twenty-six are engineers, across nine time zones. We have never had a sales team and do not intend to build one; the pricing page is the pricing. Investors are a seed round and a Series A, neither of which gets to see your usage data.

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.