Graph-of-Mark: Spatial Reasoning via Visual Prompting (2026) logo

Graph-of-Mark: Spatial Reasoning via Visual Prompting (2026)

Free

Overlays scene graphs onto input images at the pixel level to model object relationships — up to +11 percentage points on VQA and localization across 4 datasets, zero-shot

FreeFree tier
Inputs: image
Type
Open Source

About Graph-of-Mark: Spatial Reasoning via Visual Prompting (2026)

Graph-of-Mark (GoM) is the first pixel-level visual prompting technique for multimodal language models (MLMs) that overlays scene graphs onto input images to model object relationships for spatial reasoning tasks. Proposed as a training-free enhancement over methods like Set-of-Mark, GoM encodes relational information directly into the image by adding graph-based visual marks, improving the model's ability to interpret object positions and relative directions. Evaluated across three open-source MLMs and four datasets, GoM achieves consistent zero-shot improvements, boosting accuracy in visual question answering and localization by up to 11 percentage points. The method also supports auxiliary graph descriptions in the text prompt. The paper is presented at AAAI 2026.

Key Features

First pixel-level visual prompting technique that overlays scene graphs onto input images
Training-free, zero-shot enhancement for multimodal language models
Improves interpretation of object positions and relative directions
Supports auxiliary graph descriptions in the text prompt for improved grounding
Evaluated across 3 open-source MLMs and 4 different datasets
Extensive ablations on drawn components and impact of graph descriptions

Pros & Cons

Pros
  • Consistently improves zero-shot performance across multiple MLMs and datasets
  • Up to 11 percentage points improvement on VQA and localization
  • Training-free, easy to integrate with existing pre-trained MLMs
  • First method to overlay scene graphs at pixel level for spatial reasoning
  • Extensive ablations and analysis provided in the paper
Cons
  • Requires pre-extracted scene graphs, which adds a preprocessing step
  • Pixel-level overlays may introduce computational overhead compared to simpler marking techniques
  • Evaluation limited to 4 datasets and 3 MLMs

Best For

Visual question answering requiring spatial understandingObject localization and grounding in complex scenesSpatial reasoning tasks in multimodal AI systemsZero-shot image understanding with relational scene graphs

FAQ

What is Graph-of-Mark?
Graph-of-Mark (GoM) is a training-free visual prompting technique that overlays scene graphs onto input images at the pixel level to help multimodal language models better understand object relationships and spatial reasoning.
Does GoM require fine-tuning or training?
No, GoM is a training-free method. It works by augmenting the input image with scene graph overlays before feeding it to a pre-trained multimodal language model.
How much improvement does GoM provide?
GoM improves base accuracy in visual question answering and localization by up to 11 percentage points across four datasets, evaluated zero-shot.
What experiments were conducted?
GoM was evaluated on 3 open-source multimodal language models and 4 different datasets, with ablations on drawn components and the impact of auxiliary graph descriptions in the text prompt.
Where was Graph-of-Mark published?
The paper was presented at AAAI 2026 and appears in the conference proceedings.