SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
Sirun Li, Minghao Liu, Ling Dai, et al.
Introduces SIGNPOST-Bench, a counterfactual benchmark with 25,555 image variants to evaluate how multimodal LLMs resolve text-vision conflicts, showing adversarial text can increase geolocation error 4.8-fold.