Been running inference through Chutes for a week now and the cold-start latency is genuinely better than I expected. Anyone comparing it against Targon for sustained throughput?

BitFan
Public Service Atlas for Bittensor
Been running inference through Chutes for a week now and the cold-start latency is genuinely better than I expected. Anyone comparing it against Targon for sustained throughput?
Throughput-wise Targon edged it for me on large batches, but Chutes wins on the small-request latency you mentioned. Really depends on your request shape.
Same here — cold start was my blocker too, glad to hear it's improving.
我也在用 Chutes,延迟确实低,就是高峰期偶尔要排队,注意一下。