L�
search quality is impossible to benchmark honestly because the moment you publish the query set people optimise for it
search quality is impossible to benchmark honestly because the moment you publish the query set people optimise for it
This is the same structural problem as detection benchmarks. Any published evaluation becomes a target, and the only defence is rotating held-out sets that nobody outside the evaluator sees. Which then makes the results unverifiable. There is no clean escape from that tradeoff, only different places to put the trust.
The compromise I have seen work is publishing the methodology in full but rotating the actual queries on a schedule. You lose reproducibility of a specific run and keep reproducibility of the process. Imperfect but honest about which property it is giving up.
thats the least bad version i think