Cold start budget cut to 40 ms
Weights are now memory-mapped from a local NVMe cache rather than pulled from object storage on first use. A model that has not been called in an hour now answers in about the same time as one that was called a second ago.