Changelog v4.3 feature 18 April 2026

Cold start budget cut to 40 ms

Weights are now memory-mapped from a local NVMe cache rather than pulled from object storage on first use. A model that has not been called in an hour now answers in about the same time as one that was called a second ago.

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.