Preprint
Computer Vision

A survey on vision-language-action models for autonomous driving

January 1, 2025

0

Citations

0

Influential Citations

Venue

2025

Year

Abstract

… In this paper, we present the first comprehensive survey of Vision–Language–Action models for Autonomous Driving (VLA4AD), bridging the gap between classic AV reviews and the …

Analysis

Why This Paper Matters

This paper addresses a critical gap in the autonomous driving literature. While there have been numerous surveys on classic perception, prediction, and planning modules, the recent surge of Vision-Language-Action (VLA) models—which unify perception, reasoning, and control into a single framework—has not been comprehensively reviewed. The authors position VLA4AD as a new paradigm that leverages large-scale pretrained vision-language models to enhance driving capabilities, enabling more interpretable and interactive behavior.

The significance lies in the timing: VLA models are rapidly gaining traction in both academia and industry, with applications ranging from end-to-end driving to explainable decision-making. By providing the first comprehensive survey, this paper offers a structured overview that helps researchers navigate the fragmented landscape of VLA4AD, understand common design choices, and identify open challenges. It bridges the gap between traditional AV reviews and the emerging VLA paradigm, making it a valuable resource for both newcomers and experts.

Technical Contributions

The survey's key technical contributions include:

  • Taxonomy of VLA models: The authors categorize VLA4AD models based on their architecture (e.g., modular vs. end-to-end), input modalities (vision, language, action), and output representations (trajectory, control signals).
  • Analysis of training strategies: They discuss various training paradigms, including pretraining on large-scale multimodal data, fine-tuning on driving datasets, and reinforcement learning for policy optimization.
  • Review of datasets and benchmarks: The paper compiles existing datasets and evaluation metrics used for VLA4AD, highlighting their strengths and limitations.
  • Identification of key challenges: It outlines open problems such as data scarcity, sim-to-real transfer, interpretability, and safety guarantees.
  • Future directions: The authors propose potential research avenues, including leveraging foundation models, improving generalization, and developing standardized benchmarks.

Results

As a survey, the paper does not introduce new experimental results. Instead, it synthesizes findings from the literature, summarizing the performance of various VLA models on tasks like driving command prediction, trajectory planning, and visual question answering. The survey highlights that VLA models have shown promising results in simulation and closed-loop settings, but their real-world deployment remains limited due to safety and reliability concerns. The authors note that no single model has yet achieved full autonomy, and there is a trade-off between interpretability and performance.

Significance

The broader impact of this survey is substantial. It establishes VLA4AD as a distinct research area, encouraging cross-disciplinary collaboration between computer vision, natural language processing, and robotics. By systematically organizing the field, it lowers the barrier to entry for new researchers and helps practitioners identify best practices. The survey also underscores the potential of VLA models to make autonomous driving more transparent and interactive, which is crucial for public trust and regulatory acceptance. As the field evolves, this survey will likely serve as a foundational reference, guiding future innovations and fostering the development of safer, more capable autonomous vehicles.