Spent some time with Affine's reasoning eval. What I like is that it's adversarial by construction, models graded against each other in an environment rather than against a static benchmark you can overfit to. Downside is it makes cross-run comparison noisy. Still, closer to how these things actually get used than another leaderboard.

BitFan
Public Service Atlas for Bittensor