affine's adversarial reasoning eval keeps growing on me. the more i use static benchmarks the more i notice how gamed they are, everyone's training on the test set whether they admit it or not. an environment where models compete against a moving population is harder to overfit to, which makes the results mean something. i'll take meaningful-and-noisy over clean-and-gamed.

BitFan
Public Service Atlas for Bittensor