Changelog v4.7 fix 28 June 2026

Streaming responses no longer buffer the first chunk

A proxy in front of three regions was buffering until 4 KB had accumulated, which added between 60 and 300 ms to the first visible token depending on how chatty the model was. It only affected sin1, hkg1 and bom1, and only for streamed requests.

If your time-to-first-token graphs have a step down in late June, this is why.

Put the model next to the user.

One command, thirty-one regions, and an invoice that matches what you actually served.