ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2026
Year
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.
This survey addresses a critical juncture in robot learning, where two competing paradigms are emerging: one that bakes competence into frozen weights via vision-language-action (VLA) models, and another that enables robots to write and refine their own executable skills as code. By organizing the field around this weights-versus-skills axis, the paper provides a much-needed conceptual clarity that is often missing in the rapidly growing literature. The central analytical contribution—a deep-dive into code-as-policy methods arranged by degree of self-improvement—highlights a clear progression from zero-shot program synthesis to closed-loop self-repair and persistent skill memory, culminating in a sparsely populated cell where execution feedback, skill memory, and evolutionary search combine into an open-ended loop. This taxonomy not only helps researchers position their work but also reveals a significant gap: only a few very recent systems (ASPIRE, ENPIRE, and RoboClaw) occupy that advanced cell, indicating a promising frontier for future research.
The paper also connects the taxonomy to the emerging skill economy, where commercial robot-skill marketplaces are beginning to distribute one-tap skills across robots. However, these marketplaces currently ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardization. This is a timely observation, as the commercialization of robot skills is still in its infancy, and the survey provides a structured way to think about these challenges. By examining 77 representative systems across six technique families, the survey offers a focused yet comprehensive overview that is both accessible and insightful.
The survey's key technical contributions include:
The survey does not present new experimental results but rather synthesizes existing work. Its main findings include the identification of a sparsely populated cell in the self-improvement spectrum, occupied by only a few recent systems (ASPIRE, ENPIRE, and RoboClaw). This suggests that open-ended self-improvement is still a nascent area. Additionally, the survey reveals that commercial skill marketplaces currently offer only static playback, lacking the dynamic adaptation and verification needed for robust real-world deployment. The analysis of 77 systems provides a comprehensive overview, but no quantitative metrics are reported, as this is a survey paper.
This survey has significant implications for the AI and robotics community. By providing a clear taxonomy and contrast tables, it helps researchers navigate the fragmented landscape of robot learning and identify underexplored areas, particularly the open-ended self-improvement cell. The distinction between weights and skills, and the clarification of the term 'skill', could lead to more precise communication and collaboration across subfields. Furthermore, the connection to the skill economy highlights practical challenges that must be addressed for commercial viability, potentially guiding industry efforts. Overall, this survey serves as a valuable resource for both academic researchers and practitioners, offering a structured framework that could accelerate progress toward more autonomous and adaptable robots.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba