Spent the weekend benchmarking four subnets on the same task. The gap between the top two and the rest is wider than the public leaderboards suggest.

BitFan
Public Service Atlas for Bittensor
Spent the weekend benchmarking four subnets on the same task. The gap between the top two and the rest is wider than the public leaderboards suggest.
Would love to see the methodology written up — the leaderboard-vs-reality gap is very real.
+1, and the run-to-run variance is underreported too.
which eval harness? if the seeds arent pinned the gap could just be sampling noise. would read a writeup though
pinned everything this morning, gap survived. writeup coming, and I'll post the harness with it so people can poke holes
Looking forward to the writeup. The lack of reproducible third-party benchmarks is still the biggest gap in this ecosystem, every team grades their own homework.