ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… In this paper, we present the first comprehensive survey of Vision–Language–Action models for Autonomous Driving (VLA4AD), bridging the gap between classic AV reviews and the …
This paper addresses a critical gap in the autonomous driving literature. While there have been numerous surveys on classic perception, prediction, and planning modules, the recent surge of Vision-Language-Action (VLA) models—which unify perception, reasoning, and control into a single framework—has not been comprehensively reviewed. The authors position VLA4AD as a new paradigm that leverages large-scale pretrained vision-language models to enhance driving capabilities, enabling more interpretable and interactive behavior.
The significance lies in the timing: VLA models are rapidly gaining traction in both academia and industry, with applications ranging from end-to-end driving to explainable decision-making. By providing the first comprehensive survey, this paper offers a structured overview that helps researchers navigate the fragmented landscape of VLA4AD, understand common design choices, and identify open challenges. It bridges the gap between traditional AV reviews and the emerging VLA paradigm, making it a valuable resource for both newcomers and experts.
The survey's key technical contributions include:
As a survey, the paper does not introduce new experimental results. Instead, it synthesizes findings from the literature, summarizing the performance of various VLA models on tasks like driving command prediction, trajectory planning, and visual question answering. The survey highlights that VLA models have shown promising results in simulation and closed-loop settings, but their real-world deployment remains limited due to safety and reliability concerns. The authors note that no single model has yet achieved full autonomy, and there is a trade-off between interpretability and performance.
The broader impact of this survey is substantial. It establishes VLA4AD as a distinct research area, encouraging cross-disciplinary collaboration between computer vision, natural language processing, and robotics. By systematically organizing the field, it lowers the barrier to entry for new researchers and helps practitioners identify best practices. The survey also underscores the potential of VLA models to make autonomous driving more transparent and interactive, which is crucial for public trust and regulatory acceptance. As the field evolves, this survey will likely serve as a foundational reference, guiding future innovations and fostering the development of safer, more capable autonomous vehicles.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba