ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at https://earth-insights.github.io/OVEarth-bench.
Open-vocabulary Earth observation (EO) is a rapidly advancing field where models must localize geospatial concepts from natural language descriptions, moving beyond fixed label sets. However, existing benchmarks often suffer from narrow category vocabularies and limited query forms, which hampers reliable evaluation and progress. OVEarth-Bench addresses this critical gap by introducing a benchmark that significantly expands both category breadth and query diversity, offering a more realistic and challenging testbed for current and future methods.
The paper's findings are particularly significant: they reveal that current methods, including both general and EO-specific ones, still have substantial room for improvement. The observation that MLLM-based methods lead the field, while EO-specific methods often lag behind general models, challenges the assumption that domain-specific training is always beneficial. This insight is crucial for researchers and practitioners, as it suggests that leveraging large multimodal models may be a more promising direction for open-vocabulary EO tasks.
The evaluation reveals that current methods achieve limited performance on OVEarth-Bench, indicating that the benchmark is challenging and that existing approaches are not yet robust enough for diverse open-vocabulary EO tasks. Notably, broader category coverage leads to more stable model rankings, suggesting that narrow benchmarks may give misleading results. MLLM-based methods achieve the strongest overall performance, demonstrating the power of large-scale pretraining and multimodal understanding. In contrast, EO-specific methods generally underperform general models and rarely match the strongest methods, highlighting a potential gap in domain-specific model design.
OVEarth-Bench sets a new standard for evaluating open-vocabulary EO systems, pushing the field toward more realistic and diverse benchmarks. Its findings provide actionable guidance for future method development, particularly the importance of leveraging MLLMs and the need for more robust domain-specific approaches. By releasing the data and evaluation package, the authors enable the community to build upon this work, fostering further innovation. This benchmark is likely to become a key reference point for researchers aiming to advance open-vocabulary understanding in Earth observation and related fields.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba