Preprint
Reinforcement Learning

Real-world robot applications of foundation models: A review

Kento Kawaharazuka, T. Matsushima, Andrew Gambardella, Jiaxian Guo, Chris Paxton, Andy Zeng
January 1, 2024127 citations

127

Citations

3

Influential Citations

Venue

2024

Year

Abstract

… In Section 2, we overview the characteristics of foundation models and introduce … of foundation models in robotics. In Section 4, we introduce prior work on creating foundation models …

Analysis

Why This Paper Matters

This paper addresses a critical juncture in robotics: the integration of large-scale foundation models—originally developed for language and vision—into physical robotic systems. As the field moves from narrow, task-specific controllers toward generalist agents, understanding how these models can be leveraged is essential. The review is timely because it synthesizes a rapidly growing body of work, offering clarity on what has been tried, what works, and what remains unsolved.

The significance lies in its focus on real-world deployment, not just simulation. Many foundation model approaches succeed in controlled environments but fail on physical robots due to issues like latency, safety, and domain shift. By cataloging real-world applications, the paper highlights practical constraints that are often overlooked in academic research, making it valuable for both academic researchers and industry practitioners.

Technical Contributions

The paper's main contribution is a structured taxonomy of foundation model usage in robotics. Key categories include:

  • Perception and State Estimation: Using pre-trained vision-language models (VLMs) for object detection, scene understanding, and affordance prediction.
  • Policy Learning and Control: Leveraging foundation models as priors for reinforcement learning or as zero-shot planners that generate action sequences.
  • Reward and Task Specification: Using language models to define reward functions or task goals, enabling more flexible and intuitive robot programming.
  • Data Generation and Augmentation: Employing generative models to create synthetic training data or to augment real-world datasets.

The paper also discusses the characteristics of foundation models—such as scale, multimodality, and emergent capabilities—and how these properties can be exploited or become obstacles in robotics. It emphasizes the need for robot-specific foundation models that incorporate physical understanding and embodiment.

Results

As a review, the paper does not present new experimental metrics. Instead, it synthesizes findings from 127 cited works, identifying trends such as the dominance of perception-focused applications and the relative scarcity of end-to-end foundation model policies. The authors note that while some works achieve impressive zero-shot generalization in manipulation and navigation, most still require fine-tuning or adaptation for reliable real-world performance. They also highlight that no single foundation model currently serves as a complete robot brain, and that hybrid approaches combining classical control with foundation models are common.

Significance

The broader impact of this review is its potential to align research efforts toward the most promising directions. By clearly mapping the landscape, it helps avoid redundant work and encourages the development of foundation models that are inherently robot-aware—incorporating action spaces, physical constraints, and safety. For the AI field, it underscores the importance of embodiment and real-world validation, pushing foundation model research beyond text and images into interactive, physical domains. This could accelerate the deployment of general-purpose robots in homes, factories, and other unstructured environments, ultimately making AI more useful in the physical world.