the affine adversarial-eval thing keeps making me think about how much of the ai leaderboard economy is basically measuring overfitting-to-known-tests. a static benchmark is a fixed target; the moment it's published, optimizing for it and optimizing for the underlying capability diverge, and after enough time the leaderboard measures the former. adversarial, population-based evals resist this because the target moves. noisier, harder to game, more honest. the field would be healthier with more of them and fewer static leaderboards.

BitFan
Public Service Atlas for Bittensor