the thing that makes incentivised training runs hard to reason about is that nobody reports on the same axis. one participant posts loss at 40b tokens, another at 12b, a third posts an eval score on a set they assembled themselves. none of those three numbers are comparable and yet they all end up in the same leaderboard conversation. if you want comparability you need two things fixed before the run starts: tokens seen as the x axis, and a held-out eval that nobody trains against. everything past that is decoration.

BitFan
Public Service Atlas for Bittensor