ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
Contextual entrainment—the tendency of a model to let auxiliary context in its input pull its output regardless of relevance or truth—has been identified and mechanistically explained in unimodal language models. However, its manifestation in vision-language models (VLMs) remains largely unexamined. This paper fills that gap by arguing that the move to VLMs is substantive, not incremental: entrainment becomes a dual phenomenon drivable independently by textual and visual context, and it introduces a veracity distinction (context false of the depicted scene yet possible in the world) that has no counterpart in unimodal settings. The authors provide a purpose-built instrument, ENTRAP-VL, to enable rigorous investigation.
The paper does not present empirical results from model evaluations. Instead, it focuses on the design and release of the dataset and its documentation. The authors explicitly state they do not claim to measure entrainment in any particular model; they provide the instrument for the community to use.
ENTRAP-VL addresses a critical gap in multimodal AI evaluation by providing a standardized, taxonomically grounded probe for contextual entrainment. This enables researchers to systematically study how VLMs are influenced by irrelevant or misleading context, which has implications for robustness, fairness, and reliability. The dual-modality design and veracity distinction open new avenues for understanding model behavior beyond what unimodal benchmarks can capture. The public release of the dataset will facilitate broader community engagement and reproducibility.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba