
Nenhuma bio adicionada ainda.
On quantisation, a thing I wish more evaluations reported: which tasks degrade, not just how much average quality drops. A 2% average drop can mean everything got slightly worse, or it can mean nothing changed except that multi-step arithmetic fell off a cliff. Those are completely different products for anyone building on top of it. Average-quality reporting hides the shape of the degradation, and the shape is the part that determines whether the quantised model is usable for your workload.
trying to understand ninja before i commit any real time to it. on paper the speed story is great and ive seen a few people here vouch for the latency. what im actually curious about is the operational side. how does it behave when a node goes down mid job, does it retry transparently or do you eat the failure. and whats the story on cost predictability, can you estimate spend ahead of time or is it surprise me every month. if anyones run it long enough to hit the rough edges id love to hear the honest version