ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
16
Citations
0
Influential Citations
arXiv.org
Venue
2025
Year
Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors that drive this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we trained a family of self-supervised vision transformers (DINOv3) that systematically varied these different factors. We compare their representations of images to those of the human brain recorded with both fMRI and MEG, providing high resolution in spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on overall representational similarity, topographical organization, and temporal dynamics. We show that all three factors - model size, training amount, and image type - independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by both structural and functional properties of the human cortex: the representations that are acquired last by the models specifically align with the cortical areas with the largest developmental expansion, thickness, least myelination, and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world.
This paper addresses a fundamental question in AI and neuroscience: why do some neural networks develop representations that resemble the human brain? While previous work has shown that models trained on natural images can exhibit brain-like representations, the specific factors driving this similarity have remained unclear. By systematically varying model size, training data amount, and image type in a controlled family of DINOv3 vision transformers, the authors provide a causal, disentangled analysis of these factors. This is crucial for understanding the conditions under which AI models become 'brain-like' and for informing the design of more human-aligned models.
The study also introduces a developmental perspective, showing that brain-like representations emerge in a specific chronological order during training. This parallels findings in human brain development, where sensory cortices mature earlier than prefrontal regions. The link between training dynamics and cortical structural properties (expansion, thickness, myelination, timescales) is novel and suggests that AI training can serve as a model for human neurodevelopment. This interdisciplinary approach bridges AI and neuroscience, offering a promising framework for future research.
The paper reports that all three factors—model size, training amount, and image type—independently and interactively affect brain similarity metrics. Specifically, the largest DINOv3 models trained with the most human-centric images achieve the highest brain similarity across all metrics. The developmental analysis reveals that models first align with early representations of sensory cortices, and only with considerably more training do they align with late and prefrontal representations. This trajectory is indexed by cortical properties: the representations acquired last by the models align with cortical areas that have the largest developmental expansion, greater thickness, least myelination, and slowest timescales. These findings are supported by quantitative comparisons using the three metrics, though specific numerical values are not provided in the abstract.
This research provides a clear framework for understanding how architecture and experience shape brain-like representations in AI. By identifying the independent and interactive roles of model size, training data, and image type, it offers practical guidance for building vision models that align with human perception. The developmental trajectory finding is particularly impactful, as it suggests that AI training can be used as a model to study human cortical development, potentially leading to insights into neurodevelopmental disorders. Moreover, the methodology of combining fMRI and MEG with multiple similarity metrics sets a new standard for evaluating brain-model alignment. This work will likely influence both AI researchers aiming to create more human-like models and neuroscientists seeking to understand the principles of brain organization.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba