ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… This rapid growth in context window capacity presents new challenges for evaluating these long-context language models (LCLMs). A majority of existing benchmarks focus on tasks …
Long-context language models (LCLMs) have seen rapid growth in context window capacity, but existing benchmarks often fail to test the true long-range capabilities required for real-world tasks. Many benchmarks focus on short-range understanding or retrieval, missing the need for coherent generation over extended sequences. Longproc addresses this gap by introducing a benchmark centered on procedural generation, which inherently requires maintaining long-range dependencies, following multi-step instructions, and producing consistent outputs over thousands of tokens.
This is significant because procedural generation tasks—such as writing complex code, generating detailed plans, or creating structured narratives—are representative of practical applications where LCLMs are expected to excel. By focusing on these tasks, Longproc provides a more realistic assessment of LCLM capabilities, pushing the field toward evaluating models on tasks that truly stress long-context reasoning and generation.
The abstract does not include specific numerical results, but the benchmark is designed to reveal performance trends as context length increases. It is expected that current LCLMs will show degradation in performance on longer tasks, particularly in maintaining consistency and following complex instructions. The benchmark likely provides a more challenging test than existing ones, potentially showing that models with large context windows still struggle with long-range procedural generation.
Longproc has the potential to influence the development of LCLMs by highlighting the need for improved long-range coherence and instruction following. It provides a benchmark that is closer to real-world applications, encouraging researchers to focus on architectural and training improvements that enhance long-context generation. This could lead to more capable models for tasks like automated code generation, long-form content creation, and complex planning, ultimately expanding the practical utility of LCLMs.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba