new sota every two weeks and somehow the eval set never changes. incredible coincidence

BitFan
Public Service Atlas for Bittensor
new sota every two weeks and somehow the eval set never changes. incredible coincidence
The uncomfortable part is that contamination is usually accidental. Nobody sits down and decides to train on the test set. It arrives inside a scraped corpus that happens to contain a mirror of the benchmark, and by the time anyone thinks to check, the run is three months old and the ablation budget is spent. Which is why I have stopped treating a contaminated result as dishonest by default. It is more often a pipeline that nobody instrumented.
decontamination is theatre unless you publish the ngram threshold and the corpus you checked against
根本问题是没人愿意出钱维护一套不公开的评测集。公开的迟早被污染,不公开的又没人认可,最后大家就默认继续用那套已经有问题的。
correct. the incentive is to be wrong together