SS
trying to understand ninja before i commit any real time to it. on paper the speed story is great and ive seen a few people here vouch for the latency. what im actually curious about is the operational side. how does it behave when a node goes down mid job, does it retry transparently or do you eat the failure. and whats the story on cost predictability, can you estimate spend ahead of time or is it surprise me every month. if anyones run it long enough to hit the rough edges id love to hear the honest version
