ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
… Inspired by Allen-Zhu and Li [20], this paper proposes LongBioBench, employing fictional biographies as a playground to examine existing long-context language models. In our …
Long-context language models have become a major focus in NLP, with claims of handling hundreds of thousands of tokens. However, rigorous evaluation of their actual capabilities remains scarce. This paper addresses that gap by introducing LongBioBench, a synthetic benchmark that allows precise control over context length and the position of relevant information. The key insight is that many models that claim long-context proficiency fail dramatically when tested systematically.
The use of fictional biographies is clever: it avoids data contamination issues and enables unlimited scaling of context length. This makes the benchmark both fair and extensible. The paper's findings are sobering—even state-of-the-art models show significant degradation beyond 4K tokens, suggesting that current architectures have fundamental limitations in utilizing very long contexts.
This paper provides a much-needed reality check for the long-context claims made by many model providers. By offering a simple, reproducible benchmark, it enables the community to track progress in long-context understanding. The findings highlight that simply extending the context window is insufficient—models need better attention mechanisms or memory architectures to truly leverage long contexts. This work will likely influence both future model design and evaluation standards in the field.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba