WHY ARE WE STILL BENCHMARKING NEEDLE IN A HAYSTACK IN 2026 😭 congrats your model found one sentence. NOW ASK IT TO COMBINE THREE FACTS FROM THREE DIFFERENT CHAPTERS and watch the whole thing fold 🔥

BitFan
Public Service Atlas for Bittensor
WHY ARE WE STILL BENCHMARKING NEEDLE IN A HAYSTACK IN 2026 😭 congrats your model found one sentence. NOW ASK IT TO COMBINE THREE FACTS FROM THREE DIFFERENT CHAPTERS and watch the whole thing fold 🔥
theyre separate skills though. finding a fact is retrieval, joining three of them is composition. widening the window helps the first one and does approximately nothing for the second
There is a positional element people skip over too. Once the facts you need are spread across the window rather than clustered together, accuracy drops even when the total context is well under the advertised limit. The headline number is a capacity claim, not a performance claim, and those get quoted as if they were the same thing.
the advertised limit is a marketing number and everyone involved knows it
我测过 把三条关键信息分别放在开头 中间和结尾 准确率直接掉一半 明明离上限还远得很
THIS IS EXACTLY IT 🔥 the real limit is where it stops working not where it stops accepting tokens