AI Models

Xiaomi Robotics Model Shows Data Beats Model Size

Xiaomi released Xiaomi-Robotics-1, an AI model for robots that improves with more training data. To build a large dataset, the team used handheld grippers instead of physical robots, collecting over 100,000 hours of motion recordings. Tests showed that adding data boosted performance far more than increasing model size, with success rates in unfamiliar environments rising from about 25 percent to 75 percent.

Neura News

Neura News

Neura Market Editorial

July 21, 20265 min read

Originally reported by the-decoder.com

Xiaomi Robotics Model Shows Data Beats Model Size

Xiaomi has introduced an AI model for robots that follows the same scaling pattern seen in large language models. The model's performance improves as it trains on more data. To build that dataset, Xiaomi largely avoided using physical robots.

The Data Problem in Robot AI

Robot AI faces a data problem that language models do not. LLMs can train on huge parts of the public internet, while useful data on robot movement is scarce. Robots usually have to learn from scratch how to grip, lift, and put away objects. The standard method has people remotely guide a physical robot through every movement. That process is slow and expensive, and it often produces repetitive data from the same tasks in the same settings.

Xiaomi-Robotics-1 aims to close that gap. The model is designed to follow spoken or written commands in unfamiliar environments without prior exposure and adapt to new tasks with little extra training.

Handheld Grippers Replace Expensive Robots

To get around the data bottleneck, Xiaomi mostly ditched real robots during data collection. Instead, the team used portable handheld grippers with attached cameras that a person simply picks up and operates by hand. This setup lets you record manipulation tasks in kitchens, offices, stores, factory floors, and outdoor spaces without a robot even being present. The result was over 100,000 hours of motion recordings.

A dataset that large creates another problem because each recording needs a description the model can learn from. Labeling it all by hand was not practical, so Xiaomi used another AI model to describe each motion segment in text. The team says it labeled the full dataset in about two weeks.

Xiaomi then transferred that training to physical robots, including wheeled models and dual-arm systems. The model still had to account for the differences between a handheld gripper and a robot arm.

More Training Data Matters More Than Model Size

Tests showed that a larger model improved performance, but more training data produced much bigger gains than more compute. The researchers say progress in robot AI will depend mainly on collecting larger and more varied datasets.

The finding matches what researchers showed for visual data in early March. For language models, the long-standing rule is that model size and data volume should grow at roughly the same rate to make the best use of a fixed compute budget. That balance changes when a model must process images rather than just text. In that case, adding data helps far more than increasing model size.

The same pattern appeared in tests with physical robots. As Xiaomi increased the amount of training data, the model's success rate in unfamiliar environments rose from about 25 percent to 75 percent. The researchers say they have not reached a point where more data stops improving performance.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Xiaomi says Xiaomi-Robotics-1 posted the best results to date across standard robot AI benchmarks. In one demo, a robot reportedly packed a suitcase without human help. The task took more than ten minutes and required the robot to move across an entire room.

Adapting to New Tasks With Little Data

A foundation model is most useful when it can adapt to new tasks quickly. Xiaomi tested Xiaomi-Robotics-1 on four tasks, including packaging a phone, loading laundry into a washing machine, feeding paper into a printer, and packing items into a box.

With less than ten hours of training data for each task, the model reached an average success rate of 75 percent. A competing model from Physical Intelligence managed only 40 percent in the same test. Xiaomi says its model performed especially well with soft, deformable materials such as paper and on tasks that required moving around a room.

Different Approaches to Robot Data

Robot data remains a major focus of robotics research, and teams are pursuing very different solutions. A May survey of World Action Models reviewed about 100 studies and found that teleoperation data is precise but costly and limited to a few settings. Xiaomi's handheld grippers are designed to address those limits.

Nvidia, Carnegie Mellon University, and UC Berkeley used the ENPIRE project to have AI coding agents autonomously teach a fleet of eight robots how to grasp objects. Nvidia also wants to turn robotics' data problem into a compute problem by generating synthetic training data.

China's BAAI research institute is taking the opposite route. Its Orca world model learns from 125,000 hours of video without any action labels and still matches the specialized pi0.5 model on five manipulation tasks. Xiaomi automates labeling and applies it at scale, while Orca tries to eliminate labels entirely.

Physical Intelligence serves as the reference point in both cases. The company drew criticism for pi0.7 in April after claiming the model could generalize when it may have been recalling similar training data. Xiaomi's finding that more data matters more than more compute makes that question harder to resolve.

The work also fits Xiaomi's broader plans. When the company introduced its MiMo-V2 models in March, it said those agents could eventually control robots.

Xiaomi plans to release the model and code on GitHub and Hugging Face. The same team recently released Xiaomi-Robotics-0 as open source, with a focus on fast, real-time operation.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read