Preprint
Machine Learning

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Kaiyu Li, Zepeng Xin, Zixuan Jiang, Jing Fu, Lanxuan Xue, Lingyu Zhang, Xiangyong Cao
July 29, 2026

0

Citations

0

Influential Citations

Venue

2026

Year

Abstract

Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at https://earth-insights.github.io/OVEarth-bench.

Analysis

Why This Paper Matters

Open-vocabulary Earth observation (EO) is a rapidly advancing field where models must localize geospatial concepts from natural language descriptions, moving beyond fixed label sets. However, existing benchmarks often suffer from narrow category vocabularies and limited query forms, which hampers reliable evaluation and progress. OVEarth-Bench addresses this critical gap by introducing a benchmark that significantly expands both category breadth and query diversity, offering a more realistic and challenging testbed for current and future methods.

The paper's findings are particularly significant: they reveal that current methods, including both general and EO-specific ones, still have substantial room for improvement. The observation that MLLM-based methods lead the field, while EO-specific methods often lag behind general models, challenges the assumption that domain-specific training is always beneficial. This insight is crucial for researchers and practitioners, as it suggests that leveraging large multimodal models may be a more promising direction for open-vocabulary EO tasks.

Technical Contributions

  • Category Breadth: OVEarth-Bench introduces broad hierarchical category coverage, including positive and negative expressions, which tests a model's ability to distinguish between similar concepts and understand negation.
  • Query Diversity: The benchmark includes three query types—vocabulary, referring, and reasoning—each requiring different levels of understanding, from simple label matching to complex spatial and logical reasoning.
  • Unified Zero-Shot Protocol: The benchmark supports both mask and box localization under a consistent zero-shot evaluation framework, enabling fair comparisons across diverse methods.
  • Comprehensive Evaluation: The authors evaluate a wide range of general and EO-specific methods, providing a detailed analysis of their strengths and weaknesses.

Results

The evaluation reveals that current methods achieve limited performance on OVEarth-Bench, indicating that the benchmark is challenging and that existing approaches are not yet robust enough for diverse open-vocabulary EO tasks. Notably, broader category coverage leads to more stable model rankings, suggesting that narrow benchmarks may give misleading results. MLLM-based methods achieve the strongest overall performance, demonstrating the power of large-scale pretraining and multimodal understanding. In contrast, EO-specific methods generally underperform general models and rarely match the strongest methods, highlighting a potential gap in domain-specific model design.

Significance

OVEarth-Bench sets a new standard for evaluating open-vocabulary EO systems, pushing the field toward more realistic and diverse benchmarks. Its findings provide actionable guidance for future method development, particularly the importance of leveraging MLLMs and the need for more robust domain-specific approaches. By releasing the data and evaluation package, the authors enable the community to build upon this work, fostering further innovation. This benchmark is likely to become a key reference point for researchers aiming to advance open-vocabulary understanding in Earth observation and related fields.