Preprint2026
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
Xu Wang, Kaixiang Yao, Miao Pan, et al.
Proposes ProVisE, a benchmark-agnostic framework to evaluate spatial cognition in image-generation models by parsing pixel-space outputs into structured predictions, and introduces SpatialGen-Bench for unified comparison with text-output VLMs.
0Jul 23, 2026ReasoningBenchmarks
arXiv