Preprint
Large Language Models

Disentangling the Factors of Convergence between Brains and Computer Vision Models

Jos'ephine Raugel, Marc Szafraniec, Huy V. Vo, Camille Couprie, Patrick Labatut, Piotr Bojanowski, Valentin Wyart, Jean-R'emi King
August 25, 2025arXiv.org16 citations

16

Citations

0

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors that drive this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we trained a family of self-supervised vision transformers (DINOv3) that systematically varied these different factors. We compare their representations of images to those of the human brain recorded with both fMRI and MEG, providing high resolution in spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on overall representational similarity, topographical organization, and temporal dynamics. We show that all three factors - model size, training amount, and image type - independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by both structural and functional properties of the human cortex: the representations that are acquired last by the models specifically align with the cortical areas with the largest developmental expansion, thickness, least myelination, and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world.

Analysis

Why This Paper Matters

This paper addresses a fundamental question in AI and neuroscience: why do some neural networks develop representations that resemble the human brain? While previous work has shown that models trained on natural images can exhibit brain-like representations, the specific factors driving this similarity have remained unclear. By systematically varying model size, training data amount, and image type in a controlled family of DINOv3 vision transformers, the authors provide a causal, disentangled analysis of these factors. This is crucial for understanding the conditions under which AI models become 'brain-like' and for informing the design of more human-aligned models.

The study also introduces a developmental perspective, showing that brain-like representations emerge in a specific chronological order during training. This parallels findings in human brain development, where sensory cortices mature earlier than prefrontal regions. The link between training dynamics and cortical structural properties (expansion, thickness, myelination, timescales) is novel and suggests that AI training can serve as a model for human neurodevelopment. This interdisciplinary approach bridges AI and neuroscience, offering a promising framework for future research.

Technical Contributions

  • Systematic variation of factors: The authors trained a family of DINOv3 self-supervised vision transformers with controlled variations in model size (e.g., small to large), training amount (e.g., number of epochs or data size), and image type (natural images vs. human-centric images). This allows for factorial analysis of each factor's contribution.
  • Multi-modal brain comparison: They used both fMRI (spatial resolution) and MEG (temporal resolution) to compare model representations to human brain activity, enabling high-resolution spatial and temporal analyses.
  • Three complementary metrics: They employed representational similarity analysis (RSA) for overall similarity, topographical organization metrics (e.g., alignment of spatial maps), and temporal dynamics metrics (e.g., correlation of time courses). This multi-metric approach captures different aspects of brain-model alignment.
  • Developmental trajectory analysis: They tracked how brain similarity evolves during training, identifying a specific chronology: early alignment with sensory cortices, later alignment with prefrontal regions. They also correlated this trajectory with cortical structural and functional properties.

Results

The paper reports that all three factors—model size, training amount, and image type—independently and interactively affect brain similarity metrics. Specifically, the largest DINOv3 models trained with the most human-centric images achieve the highest brain similarity across all metrics. The developmental analysis reveals that models first align with early representations of sensory cortices, and only with considerably more training do they align with late and prefrontal representations. This trajectory is indexed by cortical properties: the representations acquired last by the models align with cortical areas that have the largest developmental expansion, greater thickness, least myelination, and slowest timescales. These findings are supported by quantitative comparisons using the three metrics, though specific numerical values are not provided in the abstract.

Significance

This research provides a clear framework for understanding how architecture and experience shape brain-like representations in AI. By identifying the independent and interactive roles of model size, training data, and image type, it offers practical guidance for building vision models that align with human perception. The developmental trajectory finding is particularly impactful, as it suggests that AI training can be used as a model to study human cortical development, potentially leading to insights into neurodevelopmental disorders. Moreover, the methodology of combining fMRI and MEG with multiple similarity metrics sets a new standard for evaluating brain-model alignment. This work will likely influence both AI researchers aiming to create more human-like models and neuroscientists seeking to understand the principles of brain organization.