Preprint
Large Language Models

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
July 22, 2026

0

Citations

0

Influential Citations

Venue

2026

Year

Abstract

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.

Analysis

Why This Paper Matters

Contextual entrainment—the tendency of a model to let auxiliary context in its input pull its output regardless of relevance or truth—has been identified and mechanistically explained in unimodal language models. However, its manifestation in vision-language models (VLMs) remains largely unexamined. This paper fills that gap by arguing that the move to VLMs is substantive, not incremental: entrainment becomes a dual phenomenon drivable independently by textual and visual context, and it introduces a veracity distinction (context false of the depicted scene yet possible in the world) that has no counterpart in unimodal settings. The authors provide a purpose-built instrument, ENTRAP-VL, to enable rigorous investigation.

Technical Contributions

  • Taxonomy-driven dataset: ENTRAP-VL contains 1,500 manually curated items across eight categories, organized by a taxonomy spanning two axes: association of context with the item and its relationship to truth.
  • Dual-modality streams: The dataset is split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions), allowing independent probing of each modality's influence.
  • Veracity distinction: The taxonomy introduces a novel dimension—context that is false of the depicted scene yet possible in the world—which is absent in unimodal formulations.
  • Evaluation protocols: The paper provides protocols for using the probe, enabling standardized comparisons across VLMs.

Results

The paper does not present empirical results from model evaluations. Instead, it focuses on the design and release of the dataset and its documentation. The authors explicitly state they do not claim to measure entrainment in any particular model; they provide the instrument for the community to use.

Significance

ENTRAP-VL addresses a critical gap in multimodal AI evaluation by providing a standardized, taxonomically grounded probe for contextual entrainment. This enables researchers to systematically study how VLMs are influenced by irrelevant or misleading context, which has implications for robustness, fairness, and reliability. The dual-modality design and veracity distinction open new avenues for understanding model behavior beyond what unimodal benchmarks can capture. The public release of the dataset will facilitate broader community engagement and reproducibility.