AI Models

Xiaomi Robotics Model Shows Data Beats Model Size

Xiaomi released Xiaomi-Robotics-1, an AI model for robots that improves with more training data. To build a large dataset, the team used handheld grippers instead of physical robots, collecting over 100,000 hours of motion recordings. Tests showed that adding data boosted performance far more than increasing model size, with success rates in unfamiliar environments rising from about 25 percent to 75 percent.

Neura News

Neura News

Neura Market Editorial

July 21, 20265 min read
Xiaomi Robotics Model Shows Data Beats Model Size

Xiaomi has released Xiaomi-Robotics-1, a robot AI model that improves with more data, after collecting over 100,000 hours of motion recordings using portable handheld grippers instead of physical robots. The Chinese electronics company says the model posted the best results to date across standard robot AI benchmarks, following the same scaling pattern as large language models where performance improves with more training data. The model was trained on a dataset that includes more than 1,700 different environments, from kitchens and offices to factory floors and outdoor spaces, all captured without a single physical robot present during the initial recording phase.

The Data Problem in Robot AI

Robot AI faces a fundamental data problem. Large language models train on the vast public internet, but robot movement data is scarce. The standard method has humans remotely guide a physical robot through every movement, a process that is slow, expensive, and produces repetitive data. Xiaomi mostly ditched real robots during data collection, using portable handheld grippers with attached cameras instead. These grippers are operated by hand and record manipulation tasks in kitchens, offices, stores, factory floors, and outdoor spaces without a robot present. The result was over 100,000 hours of motion recordings across more than 1,700 different environments.

Labeling all that data by hand was impractical, so Xiaomi used another AI model to describe each motion segment in text. The team says it labeled the full dataset in about two weeks. Training was then transferred to physical robots, including wheeled models and dual-arm systems. The model had to account for differences between the handheld gripper and a robot arm. Post-training combines Xiaomi's own recordings from real apartments with open-source robot datasets and annotated UMI data. The approach is designed to overcome the limits of teleoperation, which a May 2026 survey of World Action Models found to be precise but costly and limited to few settings. That survey reviewed about 100 studies and concluded that teleoperation data is rarely collected in diverse environments, a gap Xiaomi's handheld grippers were built to fill.

More Data Beats More Compute

Tests showed that a larger model improved performance, but more training data produced much bigger gains than more compute. Researchers say progress in robot AI will depend mainly on collecting larger and more varied datasets. This finding matches what researchers showed for visual data in early March 2026. For language models, a long-standing rule is that model size and data volume should grow at roughly the same rate to make best use of a fixed compute budget. That balance changes when a model processes images: adding data helps far more than increasing model size. The same pattern appeared in tests with physical robots.

As Xiaomi increased training data, the model's success rate in unfamiliar environments rose from about 25% to 75%. Researchers say they haven't reached the point where more data stops improving performance. On the RoboCasa365 leaderboard, Xiaomi-Robotics-1 leads by a wide margin, especially on unseen composite tasks. On the RoboDojo benchmark, Xiaomi-Robotics-1 scores about 58% higher than the runner-up, though absolute success rates remain low. The model also showed strong results on the LIBERO-90 benchmark, where it outperformed previous models on tasks requiring long-horizon planning and precise manipulation.

Real-World Performance and Adaptation

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

In a demonstration, a robot packed a suitcase without human help. The task took more than ten minutes and required the robot to move across an entire room. The model adapts to new tasks with little training data. Tested on four tasks—packaging a phone, loading laundry into a washing machine, feeding paper into a printer, and packing items into a box—the model reached an average success rate of 75% with less than ten hours of training data per task. A competing model from Physical Intelligence managed only 40% in the same test. Xiaomi says its model performed especially well with soft, deformable materials such as paper and tasks requiring moving around a room. The robot's ability to handle deformable objects is notable because such materials are notoriously difficult for robot AI systems to grasp and manipulate without tearing or dropping them.

The model also demonstrated adaptability in dynamic environments. In one test, the robot was asked to retrieve a specific item from a cluttered drawer it had never seen before. It succeeded on the first attempt by recognizing the object's shape and texture from its training data, even though the exact drawer arrangement was novel. Researchers attribute this to the diversity of the 1,700 environments used during data collection, which exposed the model to a wide range of clutter patterns and lighting conditions.

Alternative Approaches in the Field

Other researchers are taking different routes. Nvidia, Carnegie Mellon University, and UC Berkeley used the ENPIRE project, where AI coding agents autonomously taught a fleet of eight robots to grasp objects. Nvidia also wants to turn robotics' data problem into a compute problem by generating synthetic training data. The Beijing Academy of Artificial Intelligence (BAAI) developed the Orca world model, which learns from 125,000 hours of video without any action labels. Orca matches the specialized pi0.5 model on five manipulation tasks. Xiaomi automates labeling and applies it at scale, while Orca tries to eliminate labels entirely. Both approaches aim to solve the same bottleneck: the scarcity of labeled robot training data.

Physical Intelligence drew criticism in April 2026 for its pi0.7 model after claiming the model could generalize when it may have been recalling similar training data. Xiaomi's finding that more data matters more than more compute makes that question harder to resolve. If performance gains come primarily from larger and more diverse datasets, then distinguishing true generalization from memorization becomes more difficult. Researchers caution that the field needs better evaluation protocols to separate these two phenomena.

Release and Future Plans

Xiaomi introduced MiMo-V2 models in March 2026 and said those agents could eventually control robots. The same team recently released Xiaomi-Robotics-0 as open source, focusing on fast, real-time operation. Xiaomi plans to release Xiaomi-Robotics-1 and its code on GitHub and Hugging Face. The company says it will also release the full dataset of over 100,000 hours of motion recordings, along with the labeling pipeline, to encourage further research. The open-source release is expected to include pre-trained weights, training scripts, and evaluation benchmarks so that other labs can reproduce the results and build on the work.

Related on Neura Market

More from Neura News

Product Launch

Acer Unveils Veriton RI110 Mini Workstation for Local Agentic AI

Acer unveiled the Veriton RI110 AI Mini Workstation on September 2, 2026, in Berlin. This compact desktop, featuring an Intel Core Ultra X7 processor and Intel Arc B390 graphics, supports local inference of AI models up to 120 billion parameters. It is designed for hybrid agentic AI workloads, combining local processing with cloud resources, and includes the Qubi Claw software suite for secure, autonomous AI tasks. The system offers up to 96 GB of LPDDR5X memory, 4 TB of SSD storage, and extensive connectivity options including OCuLink, Wi-Fi 7, and dual LAN ports. Availability begins in North America in Q4 2026 and EMEA in Q1 2027.

Sep 2·4 min read