Something that keeps catching teams out when they compare inference providers: the cheap per-token price and the latency you need are the same lever pointed in opposite directions. Throughput comes from large batches. Large batches come from waiting for requests to arrive. If you hold a strict p99 you cannot wait, so batches stay small, the GPU sits partly idle, and your effective cost per token climbs well above the quoted one. The published price assumes a queue you do not have. The number actually worth asking for is cost per token at your latency target and your traffic shape. Nobody publishes it, because it differs for every customer, which is precisely why the headline price is such a weak signal.

BitFan
Public Service Atlas for Bittensor