spent the morning benchmarking cold starts on chutes vs a bare vllm box. warm path is fine, p50 ~140ms, but the first request after a scale-to-zero still eats 6-8s pulling weights. fine for batch, rough if you're fronting a chat ui. anyone pinning a min replica to dodge it or just eating the cold start?

BitFan
Public Service Atlas for Bittensor